Why hosted tools have to put limits somewhere
This is not a complaint about other services — it is arithmetic. Running speech recognition on a server costs the operator real money: GPU time, bandwidth to receive your video, storage while it is processed. Those costs scale directly with minutes of audio transcribed. A service that gave away unlimited transcription would be paying, per user, without limit.
So the cost gets recovered, and there are only so many ways to do it:
- A minute cap. Thirty minutes a month free, then a subscription.
- A watermark. The export is free but carries branding, which is advertising you are paying for in pixels.
- An export paywall. Transcribe for free, pay to download the SRT. The sting is finding out after you have done the work.
- An account wall. Free, but your email address becomes the product.
- Resolution or length limits on the exported video.
What changes when the model runs on your machine
The entire cost structure above disappears, because none of it is being paid. There is no server receiving your video — the page is static HTML, and the Whisper model is downloaded into your browser and executed there, on your own CPU or GPU. Transcribing ten hours of audio costs the site exactly as much as transcribing ten seconds: nothing.
So the honest claim list is short and checkable:
| Here | |
|---|---|
| Minute or length cap | None. The limit is your hardware and your patience. |
| Watermark on export | None, on the video or in the subtitle file. |
| Account / email | Not required. There is nothing to sign up to. |
| Payment or card details | Never requested. There is no paid tier. |
| Export paywall | None. SRT, VTT and burned-in MP4 are all free. |
| Your video | Read from disk by the page. Not uploaded, because there is no server to upload to. |
| Tracking / analytics | None on the page. |
That last row is worth a note: the page does fetch two things from third-party CDNs — the Whisper model weights, and the WebAssembly builds of the inference runtime and ffmpeg. Those are downloads to you. Nothing about your video, your audio or your subtitles is sent anywhere. How to verify that yourself in about thirty seconds.
What you give up
Every tool has a catch, and pretending otherwise would be the sort of claim this page is complaining about. The catch here is that you supply the compute, and that has four concrete consequences.
- A model download, once. 75 MB, 145 MB or 480 MB depending on which size you pick. Your browser caches it, so it is a one-time cost per model — but it is a real wait the first time, and a real amount of data on a metered connection.
- Speed is your hardware’s speed. A recent laptop with WebGPU is quick. An old machine running on the CPU is not. A hosted service with a rack of GPUs will beat your laptop on wall-clock time, and there is no way around that.
- Smaller models than a paid API runs. Browser-practical Whisper tops out at small; a paid API runs large. On clear speech the gap is modest. On accented, noisy or overlapping speech it is noticeable. The comparison, with numbers.
- You will edit the output. True of every automatic subtitle tool at any price, but worth setting the expectation: proper nouns, acronyms and jargon need a pass, and stretches of silence sometimes produce invented lines that need deleting.
There is also a hard requirement rather than a trade-off: it needs a reasonably current browser with WebAssembly
and Web Workers, and it will not run from a file:// URL.
When a paid service is the right answer
Worth saying plainly. If you are captioning a large volume of difficult audio to a professional standard, on a deadline, a paid service running a large model with human review is the correct tool and the money is well spent. This is not that. This is the right tool when you have a video, you want decent subtitles, you do not want to upload the file, and you do not want to pay or sign up for either.
Related
- MP4 to SRT
The conversion itself, what an SRT contains, and how to fix the output.
- Add subtitles to an MP4
Sidecar file or burned in — which route each platform needs.
- Which Whisper model to use
Tiny, base and small compared, and browser Whisper versus the hosted API.