Model internals

How Whisper works: inside the speech model that writes your subtitles

Updated 5 October 2026 · ~12 min read

Whisper is not an exotic piece of engineering. It is the same encoder–decoder transformer that powers machine translation, pointed at a picture of sound, and trained on an enormous pile of imperfectly-labelled audio from the web. Almost everything surprising about its behaviour — including the strange habit of writing subtitles over total silence — falls out of those three facts.

Step 1: sound becomes a picture

Audio arrives as a long list of amplitude samples — 16,000 numbers per second of mono audio, in Whisper’s case.[1] Raw samples are a terrible input for a recognition model: two recordings of the same word, shifted by a few milliseconds, look numerically unrelated.

So the audio is converted into a log-Mel spectrogram, which is built in four moves:

  1. Window it. Chop the signal into short overlapping frames (Whisper uses a 25 ms window stepped every 10 ms), because speech is only approximately stationary over tens of milliseconds.
  2. Transform each frame. A short-time Fourier transform converts each frame from amplitude over time into energy per frequency.[2]
  3. Warp the frequency axis. Group those frequencies into Mel bands — 80 of them for Whisper. The Mel scale is approximately linear below 1 kHz and logarithmic above it, modelling the fact that human hearing discriminates low frequencies far more finely than high ones.[3] A 100 Hz difference is obvious at 200 Hz and inaudible at 8 kHz.
  4. Take the logarithm. Perceived loudness is roughly logarithmic in intensity, so a log compresses the enormous dynamic range into something a neural network can train on stably.

The result is a 2-D array: 80 frequency bands × 3,000 time steps for a 30-second window. Visually it is a heat map where formants show up as horizontal bands and consonant bursts as vertical streaks. This representation, and the pipeline that produces it, long predates deep learning — Mel-frequency features were the backbone of speech recognition for decades.[3] Whisper kept the front end and replaced everything after it.

Step 2: the encoder reads the picture

Two strided 1-D convolutions first compress the 3,000 time steps to 1,500 and project each into a vector. A fixed sinusoidal positional encoding is added so the model knows frame order, and the sequence goes through a stack of transformer encoder blocks.[1]

Each block does two things. Self-attention lets every time step look at every other time step and decide which ones matter: that is how the model uses the end of a word to disambiguate its beginning, or a speaker’s pitch range several seconds earlier to interpret a vowel now.[4] A position-wise feed-forward network then transforms each step independently. Residual connections and layer normalisation around both keep gradients healthy through dozens of layers.[5]

The crucial property is that attention is global and unmasked in the encoder. The model sees the whole 30-second window at once, both directions in time. This is why Whisper handles disfluency and accent better than streaming systems that must commit to a word before hearing what follows.

The encoder’s output is 1,500 vectors describing what is happening acoustically at each moment. No text yet.

Step 3: the decoder writes text

The decoder is an autoregressive language model. It predicts one token at a time, where a token is a subword fragment from a byte-pair-encoded vocabulary — common words are single tokens, rare ones split into pieces, so any string in any language is representable without an unbounded vocabulary.[6]

Each decoder block contains three sublayers:

At each step the model produces a probability distribution over the whole vocabulary. Picking the single most likely token every time (greedy decoding) is fast but gets trapped by locally-attractive choices; keeping the k best partial sequences and extending them in parallel — beam search — finds higher-probability complete sequences at proportionally more compute.[7] Whisper’s reference implementation also uses a temperature fallback: if the output looks degenerate (compression ratio too high, average log-probability too low) it retries that window with increasing randomness.[8]

The multitask trick: tasks as tokens

This is Whisper’s defining design choice. Rather than train one model for transcription, one for language identification, one for translation and a separate aligner for timestamps, all four are folded into the decoder’s output sequence as special tokens.[1] A sequence looks roughly like:

<|startoftranscript|> <|de|> <|transcribe|> <|0.00|>
   Guten Abend, meine Damen und Herren.
<|3.44|> ... <|endoftext|>

The implications of that one format are large:

The underlying idea — recast every task as text-to-text prediction so one model and one objective cover them all — is the same one behind text models like T5,[9] and it is why the interface to Whisper is so uniform.

30-second windows and long audio

Whisper is architecturally fixed to 30 seconds of audio. Shorter clips are zero-padded to 30; longer ones are not simply fed in.[1] Your 40-minute recording is processed as a sequence of windows, and the obvious naive approach — cut every 30 seconds exactly — slices words in half at every boundary.

The reference implementation instead advances the window using the model’s own last reliable timestamp, so each window starts where the previous one genuinely finished mid-silence rather than mid-syllable. It also optionally feeds the previous window’s text back in as a prompt, giving the decoder textual context across the boundary: useful for consistent spelling of names, but a liability if the previous window went wrong, because the error gets conditioned on and can propagate for minutes.

An alternative is chunked batch inference: overlapping windows transcribed independently and stitched by matching the overlap. That loses cross-window context but is embarrassingly parallel and cannot propagate an error forward — which is why it is the common choice in browser and server implementations where throughput matters.

Where the timestamps come from — and why they are approximate

Because timestamps are predicted tokens rather than measured alignments, they are a model opinion about when speech happened. They are usually good to a few hundred milliseconds at segment level and distinctly unreliable at word level.

For subtitles this matters at the edges: a cue that starts 300 ms early is unobjectionable, but one that drifts a second late reads as broken. Systems needing tighter sync run forced alignment afterwards — taking the transcript as known and finding the most probable time alignment against the audio with a separate phoneme-level model. If you only need readable subtitles, Whisper’s native timestamps plus manual nudging is enough; if you need frame-accurate karaoke, you need an aligner.

Why weak supervision made it robust

The architecture is ordinary. The data is what made Whisper notable: 680,000 hours of audio paired with transcripts harvested from the internet, 117,000 hours of it non-English, filtered only heuristically — machine-generated transcripts were detected and removed, to avoid training a model to imitate another ASR system’s mistakes.[1]

Academic ASR had long been trained on clean, carefully-annotated corpora, and models trained that way score superbly on their own test set and degrade badly on real-world audio. Whisper’s training set is instead full of the mess that real audio contains: room reverb, telephone bandwidth, music beds, overlapping speakers, every accent. The paper’s central result is about zero-shot generalisation — competitive accuracy on benchmarks it was never fine-tuned on, approaching human robustness across varied conditions.[1] Scale and diversity of weakly-labelled data bought generalisation that clean data at smaller scale did not.

The same data explains the uneven language coverage: performance per language tracks that language’s share of the 680,000 hours. English dominates; a language contributing a few hundred hours is served far worse. There is no architectural fix for that, only more data.

Hallucination, and why it happens

Whisper will sometimes produce fluent text for audio that contains no speech — a plausible sentence over silence, a stretch of music transcribed as dialogue, or a phrase like a subtitle-site credit line that was frequent in its training data. Users report it as a bug. It is a direct consequence of the design.

The decoder is a language model that must emit tokens for every window. Over clear speech, cross-attention constrains it tightly. Over silence or noise there is no acoustic evidence to constrain anything, so the language-model prior takes over and generates what is linguistically likely, which over silence means whatever boilerplate was common in the training transcripts. The model has no mechanism for “I am not confident”; its output distribution is always normalised to sum to one.

Looping is the same phenomenon: having emitted a phrase, a repetition of it becomes the highest-probability continuation, and the model falls into a cycle. The standard mitigations all amount to detecting the condition from outside the model:

This is the strongest argument for keeping subtitles editable rather than treating ASR output as finished: the failures are not random noise, they are confident prose, and a human spots them in seconds.

tiny vs. base vs. small: what you actually trade

Whisper ships in a family of sizes sharing one architecture, differing in layer count, width and attention heads.[1] Three fit comfortably in a browser:

ModelDownloadBest forWatch out for
tiny~75 MBFast drafts of clean English speech; low-powered devicesWeak on accents, noise, proper nouns; leans English
base~145 MBThe sensible default; handles non-English audio competentlyStill struggles with jargon and heavy overlap
small~480 MBAccents, background noise, technical vocabulary, non-English audioSeveral times slower; a real download

Two rules of thumb hold up well in practice. First, the accuracy gain from going up a size is much larger for non-English audio than for clean English — if you are transcribing English podcast audio, base is often within a point or two of small; if you are transcribing Hindi or Turkish, the step up is substantial. Second, if auto-detect picks the wrong language, selecting the language explicitly helps more than a bigger model does, and costs nothing.

How a transformer runs in a browser tab

Four pieces of infrastructure make this practical without a server:

Once fetched, the weights sit in the browser’s cache, so the second transcription starts immediately. The tool on this site deliberately downloads nothing until you pick a model.

What Whisper cannot do

References

  1. Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C. & Sutskever, I. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356. The primary source for everything structural here: the 680k-hour dataset, the multitask token format, the 30-second window, the model-size table and the per-language results. (PDF)
  2. Wikipedia. Short-time Fourier transform and Spectrogram.
  3. Wikipedia. Mel scale and Mel-frequency cepstrum.
  4. Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762 — scaled dot-product attention, multi-head attention, the encoder–decoder stack and sinusoidal positional encoding, all of which Whisper uses essentially unmodified. See also Wikipedia: Transformer.
  5. Goodfellow, I., Bengio, Y. & Courville, A. Deep Learning (MIT Press, 2016), free online — background on residual connections, normalisation and sequence models.
  6. Sennrich, R., Haddow, B. & Birch, A. (2015). Neural Machine Translation of Rare Words with Subword Units. arXiv:1508.07909 — byte-pair encoding. See also Wikipedia: Byte pair encoding.
  7. Wikipedia. Beam search.
  8. OpenAI. openai/whisper — the reference implementation; transcribe.py contains the sliding-window logic, temperature fallback, and the compression-ratio and log-probability thresholds used to detect degenerate output.
  9. Raffel, C. et al. (2019). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 — the “every task is text-to-text” framing Whisper applies to speech.
  10. W3C. WebAssembly — overview and specifications.
  11. Hugging Face. Transformers.js documentation and ONNX Runtime Web.
  12. W3C. WebGPU specification; see also MDN: WebGPU API.
  13. MDN. Using Web Workers.
  14. Wikipedia. Quantization; for the neural-network case see Jacob, B. et al. (2017), Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, arXiv:1712.05877.
  15. Also useful: Jurafsky, D. & Martin, J. H., Speech and Language Processing, 3rd ed. draft — the ASR and feature-extraction chapters cover the spectrogram front end and evaluation in textbook depth.

Related: how AI translates video puts this model in the context of a full subtitling pipeline, and subtitle formats and video basics covers what happens to its output afterwards. Or try all three model sizes on your own video — nothing is uploaded.