FreeSoundTools runs OpenAI's Whisper speech recognition model inside your browser. Even with the same Whisper, results change a lot depending on the model size you pick.
What we measured
We tested 4 public read-aloud Korean samples (about 9 to 12 seconds per sentence). Character error rate is counted on characters only, without spaces and punctuation.
| Model | Average character error rate | Processing time | Notes |
|---|---|---|---|
| tiny | Not usable | About 0.3 to 1.4× the audio length | Repeated phrases or produced unrelated text |
| base | 12 to 67% depending on the sentence | About 0.6 to 1.3× | One sentence stopped halfway |
| small | About 4% | About 1.4 to 4× (first run includes the model download) | The most stable |
With only 4 clearly read sentences, real meetings or phone calls will likely have more errors. These numbers come from a single PC.
Our test used Korean. Whisper was trained on far more English audio than any other language, so English transcripts are usually more accurate than the numbers above, but the same pattern holds: bigger models make fewer mistakes.
Why transcription gets words wrong
- Smaller models are weaker, especially in languages with less training data.
- Background noise, people talking over each other, jargon and foreign names all cause errors.
- Numbers and units are written inconsistently (for example "15 m" vs "15 meters").
How to get better results
- On a PC, use the larger model (small).
- If there is a lot of noise, run the Voice Noise Remover first.
- Listen back in the editor and fix the text yourself.
- Do not use automatic transcripts as-is for important documents (contracts, legal or medical records).
Sources and date checked
Checked: October 3, 2026
- OpenAI Whisper — https://github.com/openai/whisper
- FreeSoundTools test log (4 Google FLEURS read-aloud Korean samples, 2026-10-03)
This article is general information. Devices and services can change over time.
Frequently asked questions
Which Whisper model is most accurate?
Of the three we offer, small was by far the most accurate in our tests (about 4% character error rate), while tiny was not usable.
Why is the first run slow?
The first run downloads the model to your browser. After that it is cached, so later runs start faster.
Is English more accurate than other languages?
Usually yes, because Whisper was trained on much more English audio. Noise, overlapping speech and jargon still cause mistakes.
Can I trust the transcript for legal or medical use?
No. Always check important documents against the original recording.