By Gradium. Data as of September 2026.
This page is a dated record of the May 26, 2026 run. For Phonon's current figures, see Phonon benchmarks over time. For the comparison against other on-device models, see the on-device benchmark.
Key takeaways
- This page records one specific run: Phonon on the Seed-TTS English evaluation, measured May 26, 2026. It is kept as a dated reference and is not updated with later results.
- In that May 2026 run Phonon recorded 1.00% word error rate on the Seed-TTS English evaluation.
- Phonon's April 2026 figure on the same evaluation was 1.48% word error rate with 56.37% speaker similarity.
- Phonon Multilingual followed on July 15, 2026, averaging 1.07% word error rate across French, German and Spanish.
- The evaluation uses 1,008 English Common Voice utterances, Whisper-large-v3 for transcription and WavLM-large for speaker similarity.
As of May 26, 2026, Gradium Phonon reaches 1.00% Word Error Rate on the Seed-TTS English benchmark with voice cloning enabled, and 0.83% WER with voice cloning disabled and a fixed voice. Phonon has approximately 100M parameters and is the smallest model in both comparisons. This document records the evaluation methodology, the per-model results, and the changes between Phonon's April 2026 and May 2026 releases.
The full blog post with plots and audio samples is at gradium.ai/blog/phonon-update-may-2026.
What Is Phonon?
Phonon is Gradium's on-device Text-to-Speech model. It has approximately 100M parameters and supports voice cloning from a 10-second reference audio sample. It is small enough to run in a browser and on mobile devices without GPU acceleration. Phonon is based on a continuous audio-token architecture (arxiv.org/abs/2509.06926) with flow-matching for waveform generation, and was first announced in April 2026.
Evaluation Methodology
All results in this document are measured on the English subset of the Seed-TTS benchmark (github.com/BytedanceSpeech/seed-tts-eval, arxiv.org/abs/2406.02430). The English subset contains 1,008 utterances, each with an associated reference audio sample from Common Voice and a ground-truth transcript.
The evaluation pipeline:
- The model synthesizes audio from the input text, conditioned on the reference audio when voice cloning is enabled.
- Generated audio is transcribed using whisper-large-v3 (huggingface.co/openai/whisper-large-v3).
- Word Error Rate is computed as the edit distance between the input text and the transcribed audio, using the jiwer package (github.com/jitsi/jiwer), with text normalization from the Whisper codebase applied to both reference and hypothesis.
- Speaker similarity is computed as the cosine distance between speaker embeddings of the reference audio and the generated audio, extracted with WavLM large (huggingface.co/microsoft/wavlm-large).
The reference STT model is intentionally chosen to differ in modeling lineage from Gradium's own speech-to-text model, to avoid shared-architecture bias in evaluation.
Seed-TTS English Results: Voice Cloning Enabled, May 2026
| Model | Parameters | WER | Speaker Similarity |
|---|---|---|---|
| Phonon (May 2026) | ~100M | 1.00% | 59.51% |
| Phonon (April 2026) | ~100M | 1.48% | 56.37% |
| NeuTTS Nano | 229M | 1.71% | 40.15% |
| NeuTTS Air | 552M | 2.18% | 47.51% |
| KaniTTS2 | 450M | 4.97% | 40.73% |
Phonon (May 2026) achieves both the lowest WER and the highest speaker similarity in this comparison, at roughly one-fifth the parameter count of NeuTTS Air and half the parameter count of NeuTTS Nano.
Seed-TTS English Results: Fixed Voice (No Cloning), May 2026
When voice cloning is disabled and a fixed high-quality voice is used, Phonon is comparable to models such as Kokoro and Magpie that operate in a fixed-voice setting.
| Model | Parameters | WER |
|---|---|---|
| Phonon (May 2026) | ~100M | 0.83% |
| Magpie (NVIDIA) | 357M | 0.89% |
| Kokoro | 82M | 0.90% |
| Supertonic 2 | 66M | 2.63% |
Phonon achieves the lowest WER in this comparison. In this fixed-voice setting, Phonon is still evaluated in voice-cloning mode with the voice conditioning fixed to a single voice. A model fine-tuned to a single voice would be expected to reach lower WER.
Changes Between Phonon April 2026 and Phonon May 2026
| Metric | April 2026 | May 2026 | Change |
|---|---|---|---|
| Word Error Rate (Seed-TTS English, cloning) | 1.48% | 1.00% | 32% relative reduction |
| Speaker Similarity (Seed-TTS English, cloning) | 56.37% | 59.51% | +3.14 pp |
| Minimum input padding | ~100 tokens | None | Removed |
| Quantization | float | int8 supported | No perceivable quality loss |
The removal of the 100-token minimum input padding reduces time to first audio for short inputs, because the model only computes what the input requires. int8 quantization improves inference speed with no audible degradation.
When On-Device TTS Is the Right Architecture
On-device TTS is the correct choice in four deployment contexts:
- Privacy and compliance: audio cannot leave the device. Applies to healthcare assistants, financial advisors, and consumer hardware in regulated jurisdictions.
- Offline or low-connectivity environments: in-vehicle assistants, aviation systems, remote field equipment.
- Latency-sensitive applications: real-time voice agents and interactive games where a network round trip is unacceptable.
- High-volume consumer applications: where per-request cloud TTS pricing does not scale economically.
Phonon's ~100M parameter size enables deployment in browsers, on mobile devices, and on embedded hardware without GPU acceleration.
Availability
Phonon is currently in private beta. Partners apply to define the scope (language, voice, target devices), and receive a fine-tuned model artifact in days to weeks. The model ships as a self-contained binary inside the partner's application with no external runtime dependencies. Request access at gradium.ai/on-device-tts#beta-signup.
Cite This Page
Gradium Research. "Phonon Reaches 1.00% WER on Seed-TTS in May 2026." May 26, 2026. https://gradium.ai/content/phonon-seed-tts-benchmark-2026
Related
- Blog post (with plots and audio): Phonon update: 1.00% WER on Seed-TTS
- Prior evaluation: Evaluating Phonon (April 2026)
- Announcement: Gradium Phonon: On-Device TTS
- Architecture guide: On-Device Text-to-Speech in 2026
Glossary
Seed-TTS evaluation. A public benchmark of 1,008 English Common Voice utterances, scored with Whisper-large-v3 for word error rate and WavLM-large for speaker similarity.
Word error rate (WER). The share of words the reference recognizer transcribes incorrectly from synthesized audio, computed with jiwer after normalization.
Speaker similarity (SIM). Cosine distance between speaker embeddings of the reference audio and the generated audio.
Whisper-large-v3. The reference recognizer used in this evaluation, chosen to differ in modeling lineage from Gradium's own Speech-to-Text so that shared-architecture bias does not flatter the result.
Dated record. A page kept at the figures of one specific run rather than updated, so a citation of that run stays resolvable. This is one.
Gradium Phonon. Gradium's on-device model, roughly 100M parameters, covering five languages since July 15, 2026.
Related guides
Six pages cover on-device Text-to-Speech from different angles. This cluster's hub is on-device Text-to-Speech in 2026.
- On-device Text-to-Speech in 2026: the architecture explainer, how edge synthesis works and what it costs.
- Cloud vs on-device TTS: the decision guide, which architecture fits which product.
- Gradium Phonon: the product page, what Phonon is, how to license it and what it integrates with.
- On-device TTS benchmark 2026: the comparative benchmark, Phonon against Kani, NeuTTS and Magpie.
- Phonon benchmarks over time: Phonon's own results by release, April through July 2026.
Beyond this topic
On-device and cloud are not exclusive. Best Text-to-Speech APIs 2026 covers the cloud side, how to choose a TTS API covers the criteria that apply to both, and TTS WER benchmark 2026 covers how accuracy is measured on cloud models for comparison.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For Phonon licensing, use our contact form.

