Video to Subtitles

Turn the speech in a video or audio file into a timed subtitle file. Drop in a clip, run the recogniser, then download the transcript as SRT, VTT or plain text, with the timecodes already in place.

Add your video or audio

MP4 · MOV · MKV · MP3 · WAV · M4A

Drop a file here or click to browse

One file at a time, up to 30 MB. For a long recording, pull the audio out with Video to MP3 first, it uploads far faster and covers more minutes.

MP4MOVMKVMP3WAVM4A

Subtitles

Add a file to start

Your transcript will appear here

How to generate subtitles from a video

Add your file

Drop a video or an audio file. The limit is 30 MB, so on a long recording it is worth extracting the audio track first: the same minute of speech is a fraction of the size.

Run the recogniser

Pick the spoken language if you know it, and turn on speaker labels if more than one person is talking. Then start the run and leave the tab open, the page shows how long you have waited.

Check, then download

Read the transcript and correct anything the recogniser misheard. Then save it as SRT for a player or an editor, VTT for the web, or TXT when you only want the words.

Why use Aihangsoft Video to Subtitles

Timed, not just transcribed

Every line comes back with a start and end timecode, so the result drops straight into a subtitle track. There is no manual lining up of text against the audio.

Three formats, one run

The same transcript is exported as SRT, VTT or plain text, so you do not have to run it again to get a different format. Pick whichever the platform asks for.

Nothing to install

No app, no plugin and no account needed to try it. Runs on the phone in your pocket as well as on a desktop, which matters because that is where the recording usually is.

What people use it for

Captions for social video

Most short video is watched with the sound off, and the platforms expect a caption file rather than text typed by hand. Generate the SRT here and upload it alongside the clip.

Use case: Reels, Shorts, TikTok, YouTube

Transcripts of meetings and interviews

Turn a recorded call, a lecture or an interview into searchable text instead of listening back through the whole thing. Switch on speaker labels and you can see who said what.

Use case: meetings, lectures, journalism

Text from a recording you own

The TXT export strips the timecodes and leaves plain prose, which is what you want for show notes, a blog draft, a set of study notes or anything you plan to search later.

Use case: notes, show notes, drafts

Speaker labels for a conversation

On a two person interview the recogniser can tag each line with the speaker, so the transcript reads as a conversation rather than one long block of text.

Use case: interviews, panels, podcasts

Video to subtitles FAQ

Yes, and there is no account needed to try it. Every connection gets one free run a day of up to 2 minutes of media. A free account gets 5 credits a day, which covers about one minute of speech, and Pro subscribers get 1500 credits a month with a limit of 30 minutes per run. A run costs 5 credits per minute of media, rounded up, so a 2 minute clip costs 10 credits and a 10 minute clip costs 50. Signing in raises both the daily allowance and the length limit. Credits are charged once the file has actually been processed, which means a clip that turns out to have no recognisable speech still costs the credits for its length, and so does one that contains only music. A run that fails before processing, or that the service rejects, costs nothing, and a repeated poll of a finished job is never charged twice.
Accuracy depends on the recording rather than on the tool. Clear speech recorded close to a microphone, with one person speaking at a time, transcribes very well, including names and technical terms. Accuracy drops when several people talk over each other, when loud music or traffic sits behind the voice, when the microphone is far away, and on heavy regional accents. The recogniser currently supports simplified Chinese and English only, so speech in other languages is not usable. Numbers, brand names and proper nouns are the most common mistakes, which is why the result is editable before you download it. Treat the output as a fast first draft that saves you most of the typing, not as a file you publish without reading.
You can upload audio as MP3, WAV, M4A, AAC, FLAC or OGG, or video as MP4, MOV, MKV, WebM, M4V or MPEG. The upload limit is 30 MB. The length limit depends on your plan: guests can transcribe up to 2 minutes in one run, free accounts up to 10 minutes, and Pro up to 30 minutes. Because video is much heavier than audio at the same length, a 30 MB video may be shorter than the length limit allows. For long recordings, pull the audio out first with our Video to MP3 tool, which is also browser based, then upload the much smaller audio file. If a file is too long in one piece, split it and run each part separately.
Yes, this tool works differently from our browser based tools. Your file is uploaded to our server and forwarded to Volcengine AI MediaKit, where the speech recognition runs, and the recognised text is sent back to your browser. The uploaded file is not stored as a media file on our server and no download link is generated for it. This means the tool cannot work offline and it cannot work without a connection. If your recording contains anything confidential, remember that it does leave your device. For work that must never be uploaded, use our browser based tools, which run entirely in your own browser, for example the Subtitle Converter, or a local transcription program.
Yes. When the run finishes you get a preview of the whole transcript with its timecodes, and you can fix names, numbers and wording directly in that box before saving. The download buttons always export whatever is currently in the box, so the SRT, VTT and TXT files all match your edits. If you only need the text, download the TXT version or copy it straight out of the preview. If you need the subtitles burned into the picture so they are always visible, download the SRT here and then use our Add Subtitles tool, which runs in your browser and can also restyle the font, colour and position before exporting the video.
SRT is the most widely accepted subtitle format. Nearly every video player, editor and social platform reads it, which is why it is the default for uploading captions. VTT is the format used by web video and by the HTML video element, and it supports a few extra styling features, so it is the right pick for a player on a website or for a platform that asks for a WebVTT file. TXT is just the spoken words in reading order with the timecodes stripped out, which is what you want when the text is going into a document, a blog post, a set of notes or a search index rather than back into a video. All three are generated from the same transcript, so the wording is identical and only the packaging changes.

Get subtitles from your video

Timed text you can download as SRT, VTT or TXT. No account needed to try it.

Upload a file