Most online transcription services require an account, charge per minute, or upload your files to their servers. FreeSoundTools' Speech-to-Text tool is different: it runs an open-source Whisper model right inside your browser. Your audio never leaves your device, there is no sign-up, and there is no per-minute limit.
How to Transcribe
- Open the Speech-to-Text tool.
- Upload your audio file. Supported formats include MP3, WAV, M4A, OGG, WebM, and more.
- Choose a model size:
- Tiny/Base: Fastest, lower accuracy. Good for quick drafts.
- Small: Balanced speed and accuracy for most recordings.
- Medium/Large: Highest accuracy, slower. Best for important documents or noisy audio.
- Click Transcribe. The model downloads once (cached for future use) and then processes the audio.
- When finished, the text appears with timestamps. Copy it, or download as plain text or SRT subtitle file.
Getting Better Accuracy
- Clean the audio first: Use Noise Removal to reduce background noise before transcribing.
- Use a larger model: The medium and large models handle accents, jargon, and fast speech much better.
- Record clearly: Speak close to the microphone, avoid overlapping conversations, and minimize background noise.
- Specify the language: If auto-detection misidentifies the language, manually select it for better results.
Common Uses
- Meeting notes: Record a meeting, transcribe it, and turn it into action items. For speaker-labeled notes, try the Meeting Notes tool.
- Lecture summaries: Transcribe lectures and highlight key points for study notes.
- Interview transcripts: Convert long interviews into searchable text.
- Subtitle creation: Transcribe and export as SRT for use in video editors. See also the Subtitle Generator.
- Voice memos: Turn quick voice recordings into readable text.
Privacy & Limits
- All processing happens in-browser. No server upload, no data collection.
- Long files (over 30 minutes) may take a while depending on your device's processing power.
- The model is open-source (OpenAI Whisper). Results are not 100% perfect — always proofread important transcripts.
FAQ
How accurate is the transcription?
Accuracy depends on audio quality and the model size. Clear single-speaker recordings in English can reach 95%+ accuracy with the large model. Noisy or multi-speaker recordings will be less accurate.
What languages are supported?
The underlying Whisper model supports 100+ languages including English, Korean, Japanese, Chinese, Spanish, French, German, and many more.
Is my audio uploaded to a server?
No. The speech recognition model runs entirely in your browser. Your audio stays on your device and is never sent to any server.