A Whisper subtitle generator that runs in your browser

Whisper is OpenAI’s open-source speech recognition model, released under the MIT licence. Normally you get at it by installing Python and a few gigabytes of dependencies, or by paying an API per minute of audio. Neither is necessary: the smaller Whisper models are compact enough to download into a browser tab and run there, on your own CPU or GPU. That is what this tool does, and this page is about choosing the right model for your audio.

Run Whisper now

No install, no Python, no API key, no per-minute cost.

Open the Whisper subtitle generator →

The three models, compared

All three are the same architecture trained on the same data; they differ in how many parameters they have, which is to say how much nuance they can hold. Nothing is downloaded until you pick one, and your browser caches it afterwards, so choosing again later is instant.

 TinyBaseSmall
Download~75 MB~145 MB~480 MB
Relative speedFastestRoughly half tinySlowest by a wide margin
Clear EnglishUsableGoodMarginally better
Accents, noiseStrugglesMixedClearly best
Non-EnglishWeakWorkableClearly best
Hallucination on silenceMost proneLessLeast, but not immune

How to choose, in one line each

One thing worth knowing: model size is not a cure for bad audio. A clip-on microphone improves transcription accuracy more than stepping up two model sizes does. If you have any control over the recording, spend the effort there.

Two settings that matter more than people expect

Name the spoken language

Whisper can detect the language itself, and it does so by looking at the opening stretch of audio. If your video begins with music, room noise or silence, that detection can land on the wrong language and the entire transcript comes out as nonsense — or worse, as a plausible-looking translation. Selecting the language explicitly removes the guess. It is also required if you want to translate the subtitles afterwards.

WebGPU, if your browser has it

Whisper inference is mostly large matrix multiplications, which a GPU does far better than a CPU. Where the browser exposes WebGPU, those operations are dispatched to the graphics hardware, and transcription is substantially faster. This happens automatically — there is no setting. Current Chrome and Edge have it on desktop; Firefox and Safari support is still arriving. Without it, everything still works on the CPU, just slower.

Browser Whisper versus the hosted API

Being straight about the trade-off, because it is a real one.

 In your browserHosted Whisper API
Model sizetiny / base / smalllarge-v2 or large-v3
Precision8-bit quantisedFull precision
Accuracy on hard audioLowerHigher
Your videoNever leaves the deviceUploaded to a third party
CostNonePer minute of audio
SetupOpen a tabAccount, API key, billing
Speed on a long fileBounded by your hardwareBounded by their hardware, usually faster

The honest summary: a hosted large model will transcribe difficult audio better than a quantised small model in a browser tab. If you are captioning broadcast material in a language small handles poorly, that difference matters. For the far more common case — a talk, an interview, a tutorial, a meeting recording, reasonably clear speech — base or small plus five minutes of your own editing gets you to the same place, for nothing, without the file leaving your machine. Why that last part is not just a privacy nicety.

What quantisation actually costs

The models here are stored with 8-bit integer weights rather than 32-bit floats. That is a four-fold reduction in download size and memory footprint, which is the only reason a 480 MB model is feasible in a browser at all. The cost is a small loss of numerical precision, which shows up as slightly more errors on marginal audio — exactly the audio that was already hard. On clear speech it is close to imperceptible. How Whisper works goes into the architecture and the quantisation in depth.

Where it will let you down

Related