Audio to Text Converter (Free, No Upload)
Turn meeting, lecture, interview or voice-memo recordings into editable text. Whisper AI runs on your device, so your file never leaves it.
- Only upload files you have the right to use. Copying or sharing other people's music, videos or recordings without permission can violate copyright law. FreeSoundTools does not store your files on a server.
- Don't upload secret recordings of conversations you were not part of. Recording other people without consent can be illegal.
Loading the tool...
Typing up a recording by hand takes hours. This converter runs OpenAI's Whisper speech recognition model inside your browser, so you get a first draft in minutes without uploading anything.
When to use it
- Turn a meeting recording into a draft for minutes
- Transcribe lectures, interviews or podcasts
- Convert voice memos into searchable notes
- Get English subtitles or a translated transcript of a recording
How to use
- Pick a model size and the spoken language (English is selected by default).
- Drop an audio or video file. The model downloads once and is then kept in your browser.
- Wait for recognition. You'll see progress, elapsed time and an estimate of the time left.
- Fix mistakes in the editor, use Find & Replace, or save frequent fixes to "My word list".
- Export as TXT, SRT, VTT, Word (DOCX) or JSON, or copy everything.
Supported formats and limits
- Input: mp3 · wav · m4a · ogg · mp4 · webm · mov · flac and more
- Models: tiny (about 41MB), base (about 77MB), small (about 250MB) and, on WebGPU PCs, large-v3 turbo (about 560MB)
- Languages: English, Korean, Japanese, Chinese or auto detect; optional translation to English
- Output: TXT, TXT with timestamps, SRT, VTT, two-line SRT (original + English), Word (DOCX), JSON
- Speaker labels are not added here; use AI Meeting Notes for that
Tips
- Clear recordings with little background noise give the best results. Run Noise Remover first if needed.
- Bigger models are more accurate but slower. On phones, pick base and keep files under about 10 minutes.
- "Skip quiet parts" speeds up recordings with long pauses.
FAQ
Is my audio uploaded anywhere?
No. Recognition runs in your browser. Only the AI model files are downloaded, from our own model storage (models.dagotools.com), with Hugging Face as a fallback.
How accurate is it?
Accuracy depends on the recording and the model. Treat the result as a draft and fix it in the editor. The large-v3 turbo option on WebGPU PCs is the most accurate.
Can it translate to English?
Yes. Choose "Translate to English" or "Original + English (two lines)" for bilingual subtitles. Translation takes longer.
Can I export subtitles?
Yes. Save SRT or VTT for video players and YouTube, or a two-line SRT with the original and English.
Is there a limit on length or number of files?
There is no usage limit. Very long files can run out of memory on phones, so split them if processing stops.