Text to Speech (TTS, Free)
Paste your text, pick a natural AI voice and save it as WAV or MP3 — or use your device voices instantly.
- The copyright in a text belongs to its author. If you plan to publish someone else's writing, check that you have permission.
- The "High-quality AI voice" runs Supertone's Supertonic 3 model directly in this browser. Your text and audio stay on this device, but if you publish the result, disclose that it is AI-generated.
- Don't use the AI voice to impersonate someone, create deepfakes, do anything illegal, harass anyone, or spread disinformation. See the OpenRAIL-M license for details.
- The AI voice downloads about 400MB of model files the first time you use it (Wi-Fi recommended). On the "Device voices" tab, voices marked "online" may send your text to the browser or OS provider's servers.
Loading the tool...
Hearing text read aloud helps with proofreading, presentation practice and accessibility. This tool offers two modes. The "High-quality AI voice" runs Supertone's open-source Supertonic 3 model directly in this browser, producing natural, human-like narration that you can save as a WAV or MP3 file. The "Device voices" tab uses the speech synthesis built into your browser, so you can use it right away without installing anything or downloading a model.
What is the High-quality AI voice (Supertonic 3)?
Supertonic 3 is an on-device text-to-speech model that supports 31 languages, including Korean, English and Japanese. It runs inference with ONNX Runtime Web entirely inside this browser, so your text never leaves your device. Choose from 5 male and 5 female voices, a speed from 0.8x to 1.5x, and a quality setting of 4, 8 or 16 denoising steps — higher steps sound more natural but take a little longer. You can paste up to 5,000 characters at once; longer text is split into paragraph-sized chunks automatically, generated in order and stitched into a single audio file. The model files (about 400MB) download only the first time; after that they're cached in this browser. If your device supports WebGPU, it's used automatically for much faster generation, with an automatic fallback to WebAssembly (WASM) otherwise.
When to use it
- Proofread a draft or practice a talk by listening to it
- Listen to a document when reading on screen is hard
- Get a rough idea of how text in another language sounds
- Quickly make a temporary voice-over for a video
- Try different speeds and pitches for a different feel
How to use (High-quality AI voice)
- Open the "High-quality AI voice" tab and pick a language (Korean, English, Japanese or others) and a voice (Male 1-5, Female 1-5).
- Choose a speed (0.8 to 1.5x) and quality (4, 8 or 16 steps), then paste your text (up to 5,000 characters).
- Press "Pre-download" to fetch the model ahead of time, or just press "Generate AI voice" to start right away. The model (about 400MB) downloads only once.
- The result plays automatically; save it as a WAV or MP3 file.
How to use (Device voices)
- On the "Device voices" tab, paste the text into the box.
- Choose a voice (from the voices installed on your device), a speed (0.5 to 2.0x) and a pitch (0.5 to 2.0x).
- Press the play button to listen.
- To save the sound as a file, press "Record tab audio (experimental)" and turn on "Share tab audio" in the sharing window, or the sound won't be recorded.
Supported formats and limits
- High-quality AI voice: 31 languages, 5 male and 5 female voices, speed 0.8x-1.5x, quality 4/8/16 steps, up to 5,000 characters (auto-chunked), saves as WAV or MP3, 44.1kHz sample rate
- Device voices: uses the browser's built-in speech synthesis (Web Speech API). Only the voices installed on your device and browser appear, so the list differs from device to device. Speed 0.5x-2.0x, pitch 0.5x-2.0x
- Device voices don't provide a built-in way to save the speech as a file. The experimental "Record tab audio" feature uses screen sharing (with tab audio) to record the speech as a webm file
- The AI voice's model files come from Hugging Face and need about 400MB on first use; after that they're cached in your browser
Open source used
- Supertonic 3 (model weights, OpenRAIL-M license) — Supertone Inc.
- Supertonic inference sample code (MIT license) — Supertone
- ONNX Runtime Web (MIT license) — Microsoft
Tips
- Use "Pre-download" on Wi-Fi the first time so the AI voice is ready to generate audio instantly later.
- Pick 4 steps when you just need a quick result, or 16 steps when naturalness matters most. The default, 8 steps, is a good balance of speed and quality.
- For presentation practice, set the AI voice speed a little slower (0.9-1.0x) than normal speech so you can hear each word clearly.
- If your device has no voice for the language you need, install a voice pack in your operating system settings, or just use the AI voice tab, which needs no installation.
- When you use "Record tab audio", always turn on "Share tab audio" in the sharing window. Otherwise the file may be silent.
Troubleshooting
- If the AI model download is slow or fails, check your connection (or switch to Wi-Fi) and press "Pre-download" again.
- If AI voice generation feels slow, try 4 steps, or your device may not support WebGPU (it then falls back automatically to WASM, which is slower).
- If you see a message that speech synthesis isn't supported, open the page in the latest Chrome, Edge or Safari.
- On the Device voices tab, if it says no English voice was found, your device has no English voice installed. Add one in your operating system settings or try another browser.
- If the recorded file is silent, "Share tab audio" wasn't turned on. Turn it on when you try again.
FAQ
Can I save the speech as an audio file?
Yes. Audio made in the "High-quality AI voice" tab can be saved directly as a WAV or MP3 file. The "Device voices" tab has no built-in save feature, so use the screen-recording method instead.
Which model powers the AI voice?
It runs Supertone's open-source Supertonic 3 model directly in this browser (ONNX Runtime Web). It supports 31 languages and lets you pick from 5 male and 5 female voices.
Does the AI voice need an internet connection?
Only the first time, to download the model files (about 400MB). After that they're cached in this browser, so you can use it right away next time.
Is my text sent to a server?
No. Your text and the generated audio are processed entirely on this device. Only the model files come from Hugging Face; the text itself is never sent anywhere.
Why do device voices look different for different people?
The device voice list comes from the speech engines installed on each device and browser, so it differs from device to device. The AI voice always offers the same 10 voices for everyone.
Can I publish audio made with the AI voice?
Yes, but you must disclose that it's AI-generated, and you can't use it to impersonate someone, create deepfakes, do anything illegal, harass anyone, or spread disinformation. See the OpenRAIL-M license for details.
Can it read several paragraphs at once?
Yes. The AI voice accepts up to 5,000 characters; longer text is split into paragraph-sized chunks automatically, read in order and stitched back together.
Learn more
For more detail, see the guide: How to Use Text to Speech and Save It as a File.