The three models, compared
All three are the same architecture trained on the same data; they differ in how many parameters they have, which is to say how much nuance they can hold. Nothing is downloaded until you pick one, and your browser caches it afterwards, so choosing again later is instant.
| Tiny | Base | Small | |
|---|---|---|---|
| Download | ~75 MB | ~145 MB | ~480 MB |
| Relative speed | Fastest | Roughly half tiny | Slowest by a wide margin |
| Clear English | Usable | Good | Marginally better |
| Accents, noise | Struggles | Mixed | Clearly best |
| Non-English | Weak | Workable | Clearly best |
| Hallucination on silence | Most prone | Less | Least, but not immune |
How to choose, in one line each
- Clean English, one speaker, a good microphone — tiny or base. The accuracy difference on easy audio is small, and you will not wait long.
- Anything you intend to publish — base at minimum. You will be editing the output either way, and base leaves less to edit.
- Accented speech, background noise, several speakers, or any non-English language — small. This is where model size actually earns its download; the gap is much wider here than on easy audio.
- A long recording you want to leave running — small, and accept the wall-clock time. Re-transcribing an hour of audio because tiny mangled it costs more than doing it properly once.
One thing worth knowing: model size is not a cure for bad audio. A clip-on microphone improves transcription accuracy more than stepping up two model sizes does. If you have any control over the recording, spend the effort there.
Two settings that matter more than people expect
Name the spoken language
Whisper can detect the language itself, and it does so by looking at the opening stretch of audio. If your video begins with music, room noise or silence, that detection can land on the wrong language and the entire transcript comes out as nonsense — or worse, as a plausible-looking translation. Selecting the language explicitly removes the guess. It is also required if you want to translate the subtitles afterwards.
WebGPU, if your browser has it
Whisper inference is mostly large matrix multiplications, which a GPU does far better than a CPU. Where the browser exposes WebGPU, those operations are dispatched to the graphics hardware, and transcription is substantially faster. This happens automatically — there is no setting. Current Chrome and Edge have it on desktop; Firefox and Safari support is still arriving. Without it, everything still works on the CPU, just slower.
Browser Whisper versus the hosted API
Being straight about the trade-off, because it is a real one.
| In your browser | Hosted Whisper API | |
|---|---|---|
| Model size | tiny / base / small | large-v2 or large-v3 |
| Precision | 8-bit quantised | Full precision |
| Accuracy on hard audio | Lower | Higher |
| Your video | Never leaves the device | Uploaded to a third party |
| Cost | None | Per minute of audio |
| Setup | Open a tab | Account, API key, billing |
| Speed on a long file | Bounded by your hardware | Bounded by their hardware, usually faster |
The honest summary: a hosted large model will transcribe difficult audio better than a quantised small model in a browser tab. If you are captioning broadcast material in a language small handles poorly, that difference matters. For the far more common case — a talk, an interview, a tutorial, a meeting recording, reasonably clear speech — base or small plus five minutes of your own editing gets you to the same place, for nothing, without the file leaving your machine. Why that last part is not just a privacy nicety.
What quantisation actually costs
The models here are stored with 8-bit integer weights rather than 32-bit floats. That is a four-fold reduction in download size and memory footprint, which is the only reason a 480 MB model is feasible in a browser at all. The cost is a small loss of numerical precision, which shows up as slightly more errors on marginal audio — exactly the audio that was already hard. On clear speech it is close to imperceptible. How Whisper works goes into the architecture and the quantisation in depth.
Where it will let you down
- Silence. The decoder is trained to emit text for every 30-second window, so over near-silence it produces its best guess — typically a repeated phrase or a caption-style stock line. Delete those cues.
- Proper nouns and jargon. Names, product names, acronyms and domain terms are the reliable failure cases. Always scan for them.
- Overlapping speakers. Whisper transcribes words, not speakers. Crosstalk comes out garbled and there are no speaker labels.
- Timestamp precision. Cue boundaries are predicted, not measured, and round to roughly 20 ms. Fine for viewing, not frame-accurate.
Related
- How Whisper works
Log-Mel spectrograms, the encoder-decoder transformer, the multitask token format, and why it hallucinates.
- MP4 to SRT
Getting a subtitle file out, what it contains, and how to correct the output.
- What “free subtitle generator” usually means
Minute caps, watermarks and account walls — and why none apply when the model runs on your machine.