How to generate subtitles without uploading your video

Almost every automatic captioning service works the same way: you upload the file, a server transcribes it, you download the subtitles. For a lot of video that is perfectly fine. For some of it, it is the one thing you cannot do — and the alternatives are better understood than most people realise.

When uploading is genuinely not an option

This is not an abstract privacy argument. There are specific, ordinary situations where handing a video to a third-party service is either prohibited or plainly unwise:

The common thread is that an upload is irreversible in a way that matters. Once the file is on a server you do not control, you are relying on that operator’s retention policy, their security, their subprocessors and their future change of terms — and “we delete after 24 hours” is a promise, not a mechanism.

Read the terms, specifically the training clause

The clause worth finding in any free captioning service is the one about how your content may be used. Free tiers quite often reserve the right to use submitted content to improve the provider’s models. That is a defensible business model — you are paying with data instead of money — but it is categorically incompatible with confidential footage, and it is usually several screens into a document nobody reads.

Three questions answer most of it:

  1. Is submitted content used for model training or product improvement, and can that be turned off?
  2. How long is the file retained after processing, and is deletion automatic or on request?
  3. Which subprocessors and which jurisdictions does the file pass through?

A service that answers all three clearly is being straight with you. The point of local processing is that none of the three questions arises.

The four ways to caption video locally

1. In the browser, with WebAssembly

Modern browsers can execute compiled code at close to native speed through WebAssembly, and can reach the GPU through WebGPU. That is enough to run the smaller Whisper speech-recognition models directly in a tab. The page downloads the model weights to you, reads your video off disk with the File API, and does the inference locally. Nothing is sent back, because there is nothing to send it to.

This is the lowest-friction option by a wide margin: no install, no command line, works on a locked-down work laptop where you cannot install software. The trade is that browser-practical models are smaller than what a server would run, and speed depends on your hardware. It is how the tool on this site works, and the model comparison sets out what the smaller models can and cannot do.

2. Whisper on your own machine, from the command line

If you are comfortable with a terminal, running Whisper natively gets you the full-size models at full precision, which is the accuracy ceiling. whisper.cpp is a compact C++ implementation that runs well on CPU and on Apple Silicon; the reference Python implementation and faster-whisper are the other common routes. Best accuracy, most setup, and you need somewhere to put several gigabytes of model.

3. A desktop application

Several subtitle editors now bundle local speech recognition. Subtitle Edit on Windows is the long-standing free one and can drive Whisper locally. The advantage over the command line is a proper cue editor for the part you will actually spend time on — fixing timings and line breaks.

4. Your video editor

DaVinci Resolve and Premiere Pro both have automatic transcription. Check which one you have, though: Resolve does it on-device, while Premiere’s speech-to-text has historically run in the cloud. For confidential material that distinction is the whole question, and it is worth confirming against current documentation for your version rather than assuming.

How to verify a browser tool really is local

You do not have to take anyone’s word for it, including this site’s. The browser will tell you.

  1. Open the page, then open developer tools (F12, or Cmd+Option+I on a Mac) and go to the Network tab.
  2. Load your video and run the transcription.
  3. Sort the request list by size, and look at the method column. Model weights and WebAssembly runtimes arrive as large GET requests — those are downloads to you and are expected. What you are checking for is a POST or PUT request with a payload the size of your video. If your file is 200 MB and nothing left the browser remotely near that size, it was not uploaded.

The stricter version of the test, which is conclusive:

  1. Load the page and pick a model, so the weights are fetched and cached.
  2. Disconnect from the network entirely, or tick Offline in the Network tab.
  3. Transcribe.

If it completes with no connection, the work is unambiguously happening on your machine. A tool that silently uploaded would simply fail. This works on the tool here once a model is cached, and it is the test worth applying to any product making a local-processing claim.

What “nothing is uploaded” should and should not mean

Precision matters here, because the phrase is sometimes used loosely. A local browser tool still makes network requests — it has to fetch the page, and it fetches model weights and WebAssembly binaries, often from a CDN. Those are downloads to you. The claim worth making, and the one you can verify, is narrower and more useful:

Treat any tool that blurs the distinction between “we download a model to you” and “we do not receive your file” with suspicion. They are different claims and only one of them protects you.

The practical compromise

If you do end up needing a hosted service for its accuracy on difficult audio, you can still narrow the exposure: extract and send only the audio rather than the video, since the picture is usually the more sensitive part; strip identifying metadata; and use a paid tier, which more often comes with contractual retention and no-training terms than a free one does.

But for the common case — a talk, an interview, a tutorial, a meeting recording with reasonably clear speech — a local model plus ten minutes of your own editing lands in the same place as a hosted service, and the file never leaves your desk. Getting an SRT out of an MP4 is the practical walkthrough.

Related reading