“Converting” MP4 to SRT is really transcription
Worth being precise about, because it explains why the process takes minutes rather than seconds and why the output needs a read-through. An MP4 is a container holding compressed video and audio streams. An SRT is a text file. There is no encoded form of the dialogue sitting inside the MP4 waiting to be unpacked — the words exist only as sound. So a converter has to run automatic speech recognition over the audio and produce the text and the timings from scratch.
The one exception: some MP4s, typically ones exported from a professional editor or ripped from a disc, carry a
real subtitle track alongside the video. Those can be demuxed, and the text comes out exactly as the
author wrote it. If you are not sure which case you are in, ffmpeg -i yourfile.mp4 lists every
stream; a line reading Subtitle: mov_text means there is text to extract.
Containers, codecs and muxing covers this properly.
What you get: the anatomy of an SRT file
Each cue is four lines, and the fourth is blank:
1
00:00:01,480 --> 00:00:04,120
Right, so the first thing we need to do
2
00:00:04,120 --> 00:00:07,900
is work out where the audio actually starts.
The sequence number, the timing line with a --> arrow, one or more lines of text, then an empty
line to close the cue. Two details trip people up: SRT uses a comma before the milliseconds
(00:00:01,480), not a period — that is WebVTT’s convention — and the blank line
between cues is structural, not cosmetic. Delete it and most parsers will merge or drop cues.
This matters when you edit the output. The subtitle box in the tool is a plain text area holding exactly the SRT above, so you can fix a misheard name in place and watch the video preview update. Just leave the numbers, the timing lines and the blank lines where they are.
Step by step
- Choose a model. Nothing downloads until you click. Base (~145 MB) is the right default. Small (~480 MB) is noticeably better on accented or non-English speech and worth the wait if the audio is difficult. Tiny (~75 MB) is for clean English where speed matters more than the last few percent of accuracy. The model comparison has the detail.
- Set the spoken language if you know it. Auto-detect works, but it decides from the opening audio, so a video that starts with music or silence can be misidentified. Naming the language removes that failure mode — and it is required if you want to translate afterwards.
- Drag the MP4 in. The page reads it with the File API. There is no upload request, because there is no server to receive one.
- Transcribe. The audio is decoded and resampled to 16 kHz mono — the only input Whisper accepts — then processed in 30-second windows inside a Web Worker, so the tab stays responsive. Speed depends on model size and whether your browser has WebGPU.
- Read it through, then download. Click Download .srt. Proper nouns, acronyms and technical jargon are where speech recognition reliably fails, so they are what to scan for.
Fixing the three things that usually go wrong
Repeated or invented lines over silence
If a stretch with no speech comes back as a looping phrase, you have hit Whisper’s best-known failure mode. The decoder is trained to produce text for every window, so given near-silence it emits whatever is most probable — often a caption-style stock phrase, or a repeat of the previous line. Delete those cues; they carry no information. A larger model hallucinates less, but none are immune. Why this happens.
Cue boundaries in the wrong place
Whisper segments on its own prediction of where a sentence ends, which is not necessarily where a reader wants a line break. If a cue runs long, split it: duplicate the cue, renumber, and divide the time range between the two halves. Broadcast guidelines put the comfortable ceiling at roughly 17 characters per second and no more than two lines on screen, which is a useful target even for informal video.
The whole track is offset
A constant offset across the entire file is almost never a transcription error. It is usually a frame-rate mismatch introduced later — the classic case being a 23.976 fps master treated as 24 fps, which drifts by about 3.6 seconds per hour. Any subtitle editor can apply a global shift or a frame-rate rescale.
SRT or VTT?
| SRT (SubRip) | VTT (WebVTT) | |
|---|---|---|
| Decimal separator | Comma — 00:00:01,480 | Period — 00:00:01.480 |
| Header | None | WEBVTT on the first line, required |
| Styling | None (some players honour basic HTML tags) | Cue settings plus CSS ::cue |
| Use it for | Editors, VLC, YouTube, social uploads | HTML5 <track> on a web page |
Both are offered. If you are unsure, take the SRT — it is accepted almost everywhere, and converting SRT to VTT later is a trivial text transformation. The full comparison, including ASS and the broadcast formats, is in the guide.
What this does not do
Stated plainly, so nothing is a surprise:
- It needs a modern browser with WebAssembly and Web Workers. Opening the HTML as a
file://URL will not work. - The model download is 75–480 MB the first time you pick a size. Your browser caches it afterwards.
- Everything runs on your own CPU or GPU, so an old laptop will be slow on a long video. That is the trade for not uploading anything.
- Translation of the finished subtitles currently targets English and Swedish only.
- Speaker labels (diarisation) are not produced — you get the words, not who said them.
Related
- Add subtitles to an MP4
Keeping the SRT as a separate file versus burning it into the picture, and when each is the right call.
- Which Whisper model to use
Tiny, base and small compared on speed, size and the kind of audio each copes with.
- Subtitle formats and video basics
SRT, VTT and ASS in detail, plus containers, codecs and frame-rate sync drift.