How to Burn Subtitles Into a Video Permanently (Free, No Upload)

Burning subtitles into a video means rendering the caption text into the picture frames themselves, so the words travel with the video and appear in every player, editor and upload. Aihangsoft does this for free in the browser: drop in your clip and an SRT, VTT, ASS or SSA file, style the text, and export an MP4.

The difference that matters is hard versus soft subtitles. Soft subtitles are a separate track you attach to a video file; the player has to switch them on. Hard subtitles, also called burned in or permanent subtitles, are drawn onto the pixels, so they cannot be turned off and cannot be ignored. Burning is the right choice for social uploads, projectors, kiosks and any older player that refuses to read a subtitle track. The trade is that the text is fixed once exported: to change wording, timing or colour you burn the clip again. Export always re-encodes the video.

Subtitle files are fragile in a way that surprises people. You download an SRT, drop it next to the video, and it plays perfectly on your machine. Then you send the pair to a colleague, upload the video to a social network, or plug the file into a projector, and the words are not there. The video is fine; the caption track was ignored. Burning the subtitles into the picture removes that dependency.

Add Subtitles

Burn an SRT, VTT, ASS or SSA file into any clip, with control over size, colour, outline and position. The render happens in the browser, so the video is never uploaded.

Open Add Subtitles

Hard subtitles and soft subtitles, side by side

The word subtitle covers two different mechanisms. A soft subtitle is a separate stream, inside the container or in a file beside it, that the player overlays on demand. A hard subtitle is drawn into the picture when the video is encoded, so there is nothing left for a player to switch on or ignore.

The table below is the short version, and the choice is hard to reverse once a video is out in the world.

Hard subtitles (burned in)Soft subtitles (subtitle track)
Where the text livesDrawn into the picture pixelsA separate track in the file or a sidecar file
Can viewers turn them offNo, they are part of the imageYes, from the player menu
Do they show on social platformsYes, alwaysOften stripped or ignored
Can you edit them laterOnly by burning the clip againYes, edit the file and re-attach it
Cost of adding themA full re-encode of the videoA fast remux, or no change at all
Best forDistribution, uploads, kiosks, projectionArchiving, editing, multi-language masters

In practice most projects want both: a soft version as an editable master, and a hard version for anything that leaves your machine.

When you have to burn subtitles into the video

Soft subtitles fail in a handful of predictable situations. If any of the following describes your destination, burning is the only way the text survives.

Social platforms and messaging

Most social video, from short-form feeds to in-app players, plays the picture and the sound and nothing else. Subtitle tracks that were fine in a desktop player are dropped on upload, or hidden behind an option viewers never open. Burned in captions are just pixels, so they appear for everyone, on every device. For a feed where most people watch with the sound off, that difference is the whole point of adding captions.

Projectors, kiosks and offline playback

A file played from a USB stick in a projector, a digital sign or a museum kiosk is at the mercy of whatever software the venue installed. Some read subtitle files, some only read tracks embedded a certain way, and some read nothing at all. Burning removes the question entirely, because the text is in the picture.

Old players and embedded devices

Set-top boxes, older TVs, car displays and embedded players often have no subtitle pipeline worth trusting. They may handle a plain SRT in one encoding and choke on the same file with different line endings, so hard subtitles sidestep those variables at the price of a re-encode.

How to burn subtitles into a video

The whole job runs in the browser tab, with no account and no upload.

  1. Add your video and your subtitle file. Drop in a clip in MP4, MOV, WebM or MKV, then add an SRT, VTT, ASS or SSA file. The cues are read on your device and the tool reports how many lines it found.
  2. Style the captions. Set the font size, the text colour, the outline colour and weight, then choose bottom or top and fine tune the vertical margin. The preview shows the clip itself; the captions are drawn in during the export, not in that player.
  3. Burn in and download. Press Burn in subtitles and the export writes the text into every frame. When it finishes, the result appears with a download button, and the file is saved as an MP4 with subtitled added to its name.

The settings that decide whether captions are readable

Legibility is mostly about contrast and clearance, not about picking a pretty font. Captions sit over moving pictures, so the background behind the words is never guaranteed to be dark or light.

ControlRangeDefaultWhen to change it
Font size16 to 64 pixels28 pxRaise it for phone viewing, lower it for a dense lower third
Text colourAny colourWhitePick a colour that stays distinct from the footage behind it
Outline colourAny colourBlackPair it with the text colour so the edge separates words from picture
Outline weight0 (none) to 4 (thick)2Increase it over bright or busy footage, set it to 0 for a flat look
PositionBottom or topBottomMove to the top when the lower frame carries a logo or interface
Vertical margin5 to 100 pixels30 pxPush the text clear of a channel bug or a burned in logo

The default of white text with a black outline survives the widest range of footage, which is why broadcast captions have used it for decades. The vertical margin is the setting people forget: a caption 30 pixels from the bottom edge collides with a watermark or a lower third in the same band, and there is no fixing it after the burn.

Why burning subtitles means re-encoding

The tool draws captions with the subtitles filter from FFmpeg, backed by libass, the same library a desktop player uses for ASS files. That library needs a font file, and the browser sandbox has no system fonts, so a bundled Latin typeface is written into the engine's virtual filesystem before the first frame renders. The consequence: styling is limited to that one font, which is why decorative ASS layouts do not survive.

Colour is handled in the ASS convention rather than in CSS. A value such as #ffffff becomes &HAABBGGRR, where the channels run blue, green, red, the reverse of CSS order, and the leading pair is transparency, where 00 means fully opaque. The visible effect is simple: the colour picked in the panel is the colour that comes out.

Because the picture changes on every frame, the video stream cannot be copied. It is decoded, the caption is composited, and the result is encoded again with H.264 at CRF 20 and a veryfast preset, with audio passed through as AAC at 128 kbps. That is why the export takes real time.

What this tool does not do

These are the limits, stated up front.

  • Subtitles cannot be changed after the burn. Once exported, the text is in the pixels: no off switch, no edit and no restyle. To change a word or the timing, return to the original clip and subtitle file and burn again, so keep both.
  • Every export re-encodes. Burning touches the picture, so there is no fast stream copy path, and the render runs for a good fraction of the clip length.
  • The first run downloads the engine. The FFmpeg engine is about 30 MB of code, fetched on the first export. That pause is normal on a slow connection, and the engine is cached afterwards so later jobs start at once.
  • Large files are memory bound. The pipeline runs inside the browser, so a few hundred megabytes is comfortable on a desktop while a phone or a low memory laptop can run out sooner; split the clip first if a big file fails.
  • No karaoke, no animation, no multiple tracks. The tool draws one plain text track with a single style, so karaoke timing, complex motion, several languages at once and complex ASS positioning tags are out of scope.
  • One bundled Latin font. English and accented Latin letters are covered; Chinese, Arabic, Hebrew, Cyrillic and Greek are not, because the sandbox has no system fonts. There is no workaround inside this tool.
  • The output is always MP4. The input container can be anything common, but the file you get back is H.264 in MP4. Use the Video Converter if a different container is required.
  • It adds subtitles, it does not remove them. Reading hard subtitles back off a finished video is a different job that this tool does not attempt.

If timing rather than format is the problem, fix the file first with the Subtitle Converter, and the guide to converting SRT to VTT covers the format differences and how to shift a whole timeline. If you are also marking the same clips, the guide on adding a watermark to a video for free covers text and logo overlays.


Frequently asked questions

Hard subtitles are drawn into the video frames, so they are part of the picture and always visible. Soft subtitles are stored as a separate track inside the container or as a sidecar file next to the video, and the player decides whether to show them. Hard subtitles cannot be switched off or restyled after export, but they survive every platform and every player, including social apps that strip subtitle tracks. Soft subtitles stay editable and can be turned on and off, but they disappear the moment a service ignores the track. As a rule, use soft subtitles for archiving and editing, and hard subtitles for distribution, where the words must appear for everyone without any action. The term caption usually means the same thing as subtitle, though captions sometimes include sound descriptions such as music or a door slamming, while subtitles cover dialogue. This tool treats the two the same: whatever text is in the file is burned into the frame.
Four formats are accepted: SRT, VTT, ASS and SSA. Whatever you upload is parsed on your device and rewritten as clean SRT before it is handed to the caption renderer, because SRT is the most predictable format for the bundled libass engine that draws the text. Timecodes are normalised along the way, so a file that uses dots instead of commas, or hour, minute, second and centisecond fields as ASS does, still lands on the right frame. If the parser finds no cues, the tool reports that and stops rather than exporting a video with no captions at all. Decorative ASS override tags, in curly braces, are stripped rather than reproduced, and markup such as VTT bold or italic tags is removed, so what you burn in is the plain text under the styling you choose in the panel.
Expect it to take a noticeable amount of time, because subtitles are drawn onto every frame and the clip is fully re-encoded rather than copied. As a rule of thumb, the export runs for between half of the video length and the full length, and it depends on the resolution, the frame rate and the machine doing the work. A short phone clip often finishes in under a minute, while a long 4K recording can take several minutes. The first run also downloads the FFmpeg engine, roughly 30 MB of code, which is cached afterwards so later jobs start straight away. Trim the clip first if you want a faster result, since every second you remove is a second the encoder does not have to draw captions onto and encode again. Keeping the tab in front while it runs is a good idea, because the render happens in the page and a sleeping tab slows down.
No. Burning writes the words into the pixels, so once the file is exported there is no separate caption layer to edit. Changing a typo, retiming a line or switching from bottom to top means going back to the source clip and the subtitle file and burning again. That is why you should keep both the original video and the caption file until the export looks right, and why it is worth checking a short test section before running a long render. The styling you can choose before the burn covers font size from 16 to 64 pixels, any text and outline colour, an outline weight from none to thick, a bottom or top position, and a vertical margin from 5 to 100 pixels. After the burn, none of those settings can be revisited.
Three limits are worth knowing. First, everything runs in browser memory, so very large files can hit WebAssembly limits; a few hundred megabytes is comfortable on a desktop, while a phone or a low memory laptop may struggle sooner. Second, the renderer ships one bundled Latin font, so English and accented Latin letters work but Chinese, Arabic, Hebrew, Cyrillic and Greek text is not covered. Third, karaoke effects, complex animations and multiple subtitle languages burned at the same time are out of scope: the tool draws one plain text track with a single style. If a large file fails, export a lower resolution copy first and try again, or split the clip into parts with the Video Trimmer before burning. Burning is also one way traffic: the tool adds captions, it does not read them back off a video that already has hard subtitles. If the words are already in the picture, there is nothing here that can pull them out.
Never. The subtitle parser and the FFmpeg engine both run locally in the browser through WebAssembly, so your video and your caption file are read from disk, processed in memory and written back as a download. There is no upload step, no server, no queue and no account, and the exported file carries no watermark or branding of ours. That matters when the footage is an internal screen recording, a customer call, a paid course video or anything under an embargo, because those are exactly the files you would rather not hand to a random converter site. You can confirm it for yourself: open the browser developer tools, switch to the network panel, and watch while a clip is burned in. No request will carry your video. The engine download on the first run is our code, not your media, and it is cached for later jobs.

Put your captions in the picture

Free, private and unlimited. No account needed.

Open Add Subtitles