Transcribe Audio & Video
Powered by OpenAI Whisper, running entirely in your browser. 50+ languages.
Drop files here
or click to browse · paste from clipboard
Accepts .MP3, .MP4, .WAV, .OGG, .FLAC, .M4A, .WEBM, .MOV · Up to 1,000 files
Audio never touches a server. Verify in DevTools — all processing runs locally via WebAssembly.
Frequently asked questions
No. Whisper runs locally in your browser via transformers.js — your audio files are never sent to a server. The model downloads once on first use and is cached in your browser for all future sessions.
Balanced mode (Whisper-base) achieves ~90%+ word accuracy on clear, accented-neutral English speech. Accurate mode (Whisper-small) is noticeably better for accented speech and technical vocabulary. Both models struggle with heavy background noise, multiple overlapping speakers (crosstalk), and domain-specific jargon.
Whisper supports 50+ languages including English, Spanish, French, German, Japanese, Portuguese, Arabic, and Hindi. Select a language manually or leave it on Auto-detect for automatic language identification.
Yes. Upload MP4, WebM, or MOV files — the tool extracts the audio track in-browser before transcribing. The video itself is never sent anywhere.
TXT is plain text — the full transcript as one continuous block. SRT is a subtitle format with timestamps for each segment, suitable for adding captions to video in any editor (DaVinci Resolve, Premiere, CapCut, etc.).
Files up to 30 minutes are supported. For longer recordings, split them into segments first using an audio or video editor, then transcribe each segment.