When uploading is genuinely not an option
This is not an abstract privacy argument. There are specific, ordinary situations where handing a video to a third-party service is either prohibited or plainly unwise:
- Material under NDA. A client’s unreleased product demo, a pitch recording, footage from a confidential project. Most NDAs do not carve out an exception for uploading the material to a captioning vendor.
- Recorded meetings and internal calls. These contain other people’s voices, names, salaries, customer details and strategy. Your organisation may have a policy about where that can go; if it does, a free web tool is almost certainly not on the approved list.
- Anything involving personal data. Under the GDPR, a recording of identifiable people is personal data, and sending it to a processor is a processing activity that needs a lawful basis and usually a data-processing agreement. Clicking “upload” on a consumer site does not create one.
- Medical, legal and HR recordings. Patient consultations, client interviews, investigation footage. The sensitivity is obvious and the rules are strict.
- Unreleased creative work. A film cut, a course you are about to sell, a music video under embargo. A leak here is commercially expensive and there is no undo.
- Personal video. Family footage, home recordings. No rule applies; you simply may not want it on someone else’s disk.
The common thread is that an upload is irreversible in a way that matters. Once the file is on a server you do not control, you are relying on that operator’s retention policy, their security, their subprocessors and their future change of terms — and “we delete after 24 hours” is a promise, not a mechanism.
Read the terms, specifically the training clause
The clause worth finding in any free captioning service is the one about how your content may be used. Free tiers quite often reserve the right to use submitted content to improve the provider’s models. That is a defensible business model — you are paying with data instead of money — but it is categorically incompatible with confidential footage, and it is usually several screens into a document nobody reads.
Three questions answer most of it:
- Is submitted content used for model training or product improvement, and can that be turned off?
- How long is the file retained after processing, and is deletion automatic or on request?
- Which subprocessors and which jurisdictions does the file pass through?
A service that answers all three clearly is being straight with you. The point of local processing is that none of the three questions arises.
The four ways to caption video locally
1. In the browser, with WebAssembly
Modern browsers can execute compiled code at close to native speed through WebAssembly, and can reach the GPU through WebGPU. That is enough to run the smaller Whisper speech-recognition models directly in a tab. The page downloads the model weights to you, reads your video off disk with the File API, and does the inference locally. Nothing is sent back, because there is nothing to send it to.
This is the lowest-friction option by a wide margin: no install, no command line, works on a locked-down work laptop where you cannot install software. The trade is that browser-practical models are smaller than what a server would run, and speed depends on your hardware. It is how the tool on this site works, and the model comparison sets out what the smaller models can and cannot do.
2. Whisper on your own machine, from the command line
If you are comfortable with a terminal, running Whisper natively gets you the full-size models at full precision,
which is the accuracy ceiling. whisper.cpp is a compact C++ implementation that runs well on CPU and
on Apple Silicon; the reference Python implementation and faster-whisper are the other common
routes. Best accuracy, most setup, and you need somewhere to put several gigabytes of model.
3. A desktop application
Several subtitle editors now bundle local speech recognition. Subtitle Edit on Windows is the long-standing free one and can drive Whisper locally. The advantage over the command line is a proper cue editor for the part you will actually spend time on — fixing timings and line breaks.
4. Your video editor
DaVinci Resolve and Premiere Pro both have automatic transcription. Check which one you have, though: Resolve does it on-device, while Premiere’s speech-to-text has historically run in the cloud. For confidential material that distinction is the whole question, and it is worth confirming against current documentation for your version rather than assuming.
How to verify a browser tool really is local
You do not have to take anyone’s word for it, including this site’s. The browser will tell you.
- Open the page, then open developer tools (
F12, orCmd+Option+Ion a Mac) and go to the Network tab. - Load your video and run the transcription.
- Sort the request list by size, and look at the method column. Model weights and WebAssembly runtimes arrive as large
GETrequests — those are downloads to you and are expected. What you are checking for is aPOSTorPUTrequest with a payload the size of your video. If your file is 200 MB and nothing left the browser remotely near that size, it was not uploaded.
The stricter version of the test, which is conclusive:
- Load the page and pick a model, so the weights are fetched and cached.
- Disconnect from the network entirely, or tick Offline in the Network tab.
- Transcribe.
If it completes with no connection, the work is unambiguously happening on your machine. A tool that silently uploaded would simply fail. This works on the tool here once a model is cached, and it is the test worth applying to any product making a local-processing claim.
What “nothing is uploaded” should and should not mean
Precision matters here, because the phrase is sometimes used loosely. A local browser tool still makes network requests — it has to fetch the page, and it fetches model weights and WebAssembly binaries, often from a CDN. Those are downloads to you. The claim worth making, and the one you can verify, is narrower and more useful:
- Your video file is never transmitted.
- Your audio is never transmitted — including the extracted 16 kHz mono track the model actually consumes, which is the sneaky case, since it is far smaller than the video and easier to miss in a network log.
- Your transcript is never transmitted.
Treat any tool that blurs the distinction between “we download a model to you” and “we do not receive your file” with suspicion. They are different claims and only one of them protects you.
The practical compromise
If you do end up needing a hosted service for its accuracy on difficult audio, you can still narrow the exposure: extract and send only the audio rather than the video, since the picture is usually the more sensitive part; strip identifying metadata; and use a paid tier, which more often comes with contractual retention and no-training terms than a free one does.
But for the common case — a talk, an interview, a tutorial, a meeting recording with reasonably clear speech — a local model plus ten minutes of your own editing lands in the same place as a hosted service, and the file never leaves your desk. Getting an SRT out of an MP4 is the practical walkthrough.
Related reading
- How Whisper works
What the model does to your audio, why it runs in 30-second windows, and what quantisation costs.
- How AI translates video
The six stages between a video file and a translated subtitle track, and where each one fails.
- What “free subtitle generator” usually means
Minute caps, watermarks and account walls, and why local processing has different economics.