Add Captions to Video
Drop a video, choose a language and quality (or import an SRT/VTT), fix any mistakes, pick a style, then download burned-in captions or a subtitle file.
Drop a video here
or click to browse · paste from clipboard
Accepts .MP4, .MOV, .WEBM, .AVI, .MKV · One video
How it works
Drop your video
MP4, MOV, WebM, AVI, or MKV. The file stays on your device — nothing is uploaded.
Transcribe or import
Pick a language and quality, then Whisper runs locally — or import an SRT/VTT and skip transcription. Phones default to Fast (~40 MB). Desktop Balanced is ~74 MB, cached after the first run.
Edit, style, and preview
Fix any transcript mistakes, pick a style (One Word, Outline, Bar, Classic, Follow), choose a font, set colors and position. The video previews your changes live.
Burn and download
Burn captions into the picture, or mux a soft subtitle track without re-encoding. Download the finished MP4. Files never leave your browser.
Frequently asked questions
No. Every step runs entirely in your browser. The Whisper transcription model downloads once to your device and runs locally. The ffmpeg caption burn happens in-browser via WebAssembly. Your video never leaves your computer.
Set the spoken language before transcribing — that is the biggest accuracy win. Fast uses Whisper Tiny (~40 MB), Balanced uses Base (~74 MB), and Accurate uses Small (~244 MB). All three still stumble on names, jargon, accents, and noise. Review the transcript, or import an SRT/VTT you already trust.
Yes. After you drop a video you can import an existing SRT or VTT and skip Whisper entirely. After editing you can download SRT or VTT without burning captions into the file — useful for YouTube and players that read sidecar subtitles.
One Word shows a single bold word at a time, centered on screen. Each word appears when it is spoken and disappears when the next word starts. Outline is the same timing with all-caps and a heavier stroke. Bar uses a dark background behind full lines. Follow keeps the line on screen and changes the color of the spoken word.
Yes. After you have a transcript, use “Soft captions” to mux an SRT track into the MP4. The video and audio are stream-copied — usually a few seconds — and players like YouTube, VLC, and QuickTime can turn the track on or off. Burn-in is still required for most social apps.
Yes — two ways. You can upload a .ttf or .otf font file, or on Chrome and Edge you can click "Your computer" to browse your installed fonts. The selected font is used for both the preview and the burned-in output.
Input: MP4, MOV, WebM, AVI, MKV. Output is always MP4. The burned-in captions are permanently part of the video — no separate subtitle file needed.
Two steps: transcription (typically 2–5× real-time, so a 2-minute video takes ~1 minute) and burning (1–4× real-time depending on your device). Longer videos and slower devices take more time. Keep the tab open and in the foreground while processing — iPhone Safari will reload the page if the tab is backgrounded or runs out of memory. On phones, Fast is the default for that reason.
Yes. Click a word to jump the preview. Double-click to edit text. Select a word to split, merge, insert, or nudge its start and end times. The preview updates as you edit.
Classic shows 1–2 lines of white text with a black shadow, like traditional subtitles. Follow keeps the current line on screen and changes the color of the word being spoken.