MP4 to SRT, free and without uploading the video

An .srt file is plain text: numbered cues, a start and end time, and the words spoken in between. Getting one out of an MP4 means transcribing the speech, and the tool on this site does that with OpenAI’s Whisper model running inside your browser tab. The video stays on your disk, and the finished subtitles are editable before you download them.

Open the converter

Free, no account, no watermark, no file leaves your machine.

Convert an MP4 to SRT →

“Converting” MP4 to SRT is really transcription

Worth being precise about, because it explains why the process takes minutes rather than seconds and why the output needs a read-through. An MP4 is a container holding compressed video and audio streams. An SRT is a text file. There is no encoded form of the dialogue sitting inside the MP4 waiting to be unpacked — the words exist only as sound. So a converter has to run automatic speech recognition over the audio and produce the text and the timings from scratch.

The one exception: some MP4s, typically ones exported from a professional editor or ripped from a disc, carry a real subtitle track alongside the video. Those can be demuxed, and the text comes out exactly as the author wrote it. If you are not sure which case you are in, ffmpeg -i yourfile.mp4 lists every stream; a line reading Subtitle: mov_text means there is text to extract. Containers, codecs and muxing covers this properly.

What you get: the anatomy of an SRT file

Each cue is four lines, and the fourth is blank:

1
00:00:01,480 --> 00:00:04,120
Right, so the first thing we need to do

2
00:00:04,120 --> 00:00:07,900
is work out where the audio actually starts.

The sequence number, the timing line with a --> arrow, one or more lines of text, then an empty line to close the cue. Two details trip people up: SRT uses a comma before the milliseconds (00:00:01,480), not a period — that is WebVTT’s convention — and the blank line between cues is structural, not cosmetic. Delete it and most parsers will merge or drop cues.

This matters when you edit the output. The subtitle box in the tool is a plain text area holding exactly the SRT above, so you can fix a misheard name in place and watch the video preview update. Just leave the numbers, the timing lines and the blank lines where they are.

Step by step

  1. Choose a model. Nothing downloads until you click. Base (~145 MB) is the right default. Small (~480 MB) is noticeably better on accented or non-English speech and worth the wait if the audio is difficult. Tiny (~75 MB) is for clean English where speed matters more than the last few percent of accuracy. The model comparison has the detail.
  2. Set the spoken language if you know it. Auto-detect works, but it decides from the opening audio, so a video that starts with music or silence can be misidentified. Naming the language removes that failure mode — and it is required if you want to translate afterwards.
  3. Drag the MP4 in. The page reads it with the File API. There is no upload request, because there is no server to receive one.
  4. Transcribe. The audio is decoded and resampled to 16 kHz mono — the only input Whisper accepts — then processed in 30-second windows inside a Web Worker, so the tab stays responsive. Speed depends on model size and whether your browser has WebGPU.
  5. Read it through, then download. Click Download .srt. Proper nouns, acronyms and technical jargon are where speech recognition reliably fails, so they are what to scan for.

Fixing the three things that usually go wrong

Repeated or invented lines over silence

If a stretch with no speech comes back as a looping phrase, you have hit Whisper’s best-known failure mode. The decoder is trained to produce text for every window, so given near-silence it emits whatever is most probable — often a caption-style stock phrase, or a repeat of the previous line. Delete those cues; they carry no information. A larger model hallucinates less, but none are immune. Why this happens.

Cue boundaries in the wrong place

Whisper segments on its own prediction of where a sentence ends, which is not necessarily where a reader wants a line break. If a cue runs long, split it: duplicate the cue, renumber, and divide the time range between the two halves. Broadcast guidelines put the comfortable ceiling at roughly 17 characters per second and no more than two lines on screen, which is a useful target even for informal video.

The whole track is offset

A constant offset across the entire file is almost never a transcription error. It is usually a frame-rate mismatch introduced later — the classic case being a 23.976 fps master treated as 24 fps, which drifts by about 3.6 seconds per hour. Any subtitle editor can apply a global shift or a frame-rate rescale.

SRT or VTT?

 SRT (SubRip)VTT (WebVTT)
Decimal separatorComma — 00:00:01,480Period — 00:00:01.480
HeaderNoneWEBVTT on the first line, required
StylingNone (some players honour basic HTML tags)Cue settings plus CSS ::cue
Use it forEditors, VLC, YouTube, social uploadsHTML5 <track> on a web page

Both are offered. If you are unsure, take the SRT — it is accepted almost everywhere, and converting SRT to VTT later is a trivial text transformation. The full comparison, including ASS and the broadcast formats, is in the guide.

What this does not do

Stated plainly, so nothing is a surprise:

Related