Turn audio into text, privately
Whisper runs inside your browser, so the recording never leaves your device. Get plain text back, plus subtitles timed to the speech.
- Runs on your device
- 99 languages
- Free no sign-up
1Choose a model
The first run downloads the model that is the size shown above, and it only happens once. Your browser caches it, and everything after that starts immediately. The audio itself is never uploaded; it is transcribed on this device, which is why there is no length limit.
Drop an audio file here, or choose one
MP3, WAV, M4A, WebM, FLAC nothing is uploaded
Transcribe audio without uploading it
Drop in a recording and get the text back, along with subtitles timed to the speech. Whisper runs inside your browser, so interviews, voice notes and anything else you would rather not hand to a server stay on your own machine free, and without an account.
How it works
- 1
Choose a model
Fast for clear speech, Accurate for difficult audio. It downloads once from Hugging Face and your browser keeps it, so every transcription after the first starts immediately.
- 2
Drop in the audio
MP3, WAV, M4A, WebM or FLAC. The file is decoded and transcribed on this device it is never uploaded, which is also why there is no length limit.
- 3
Edit and export
Fix anything Whisper misheard, then take the plain text, or SRT and WebVTT subtitles timed to the speech.
Questions
- Is my audio uploaded?
- No. Transcription runs inside your browser using OpenAI's Whisper model. The only thing downloaded is the model itself; the recording never leaves your device.
- Why is the first run slow?
- The model has to be fetched once between 40 MB and 250 MB depending on which you pick. Your browser caches it, so later transcriptions skip that step entirely.
- Which languages does it handle?
- Whisper covers around 99 languages and detects the spoken one automatically. Accuracy varies by language and by how clear the recording is.
- How long can the recording be?
- There is no limit imposed here, because nothing is being sent to a server. Long recordings simply take longer, and a larger model takes longer still.
- Can I get subtitles from it?
- Yes. Whisper reports when each phrase was spoken, so the transcript exports as SRT or WebVTT with the timings already in place.
- Can I turn the transcript back into speech?
- Yes Open in the studio carries the text across, so you can edit it and generate it in any of the available voices.