Voice Design. Live today.

Benchmark

TTS WER Benchmark 2026: Pronunciation Accuracy

Gradium9 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • On the Coval TTS leaderboard, Gradium's word error rate sat at about 5% through the week of September 6 to 10, 2026, 5.29% on the September 10 read of 27 models, where the lowest reading was Soniox tts-rt-v2 at about 4.4%.
  • On the MiniMax Multilingual TTS Test Set, evaluated in April 2026 across English, French, Spanish, Portuguese and German, Gradium recorded a 1.11% average word error rate, the lowest of the six systems in that test.
  • On a 500-sentence hard-case set covering the same five languages, rated by independent native speakers in August 2026, Gradium TTS passed 81.0% of sentences against Cartesia Sonic 3.6 at 75.1% and ElevenLabs Eleven v3 Conversational at 65.4%.
  • Those three numbers measure different things. Clean-text English word error rate is close to saturated in 2026; the differences that matter now show up on numbers, acronyms, emails and code-switching.
  • Word error rate depends on the reference Speech-to-Text model and the text normalizer as much as on the Text-to-Speech model, so figures from two benchmarks are never one series.

How is Text-to-Speech accuracy measured?

Word error rate is the standard proxy. A Text-to-Speech model synthesizes audio from a reference text, a reference Speech-to-Text model transcribes that audio back to text, both sides are normalized, and the edit distance between them is divided by the number of reference words. Insertions, deletions and substitutions all count. Lower is better.

Normalization is doing more work in that pipeline than it looks. It lowercases, strips punctuation and converts numbers and abbreviations to a common form, so that "3" against "three" or "Dr." against "doctor" is not scored as an error when the pronunciation was correct. It is also language-specific and imperfect. Two benchmarks that use different normalizers will report different word error rates from identical audio, which is why a figure without its methodology attached is not comparable to anything.

What word error rate does not capture: naturalness, prosody, expressiveness, or whether a voice is pleasant to listen to. A model can score well and still sound robotic. For that dimension, see the preference-based Elo boards discussed at the end of this page.

What did the Coval board show for word error rate in September 2026?

Coval runs a continuous, independent test against production endpoints, transcribing with a reference Speech-to-Text model and normalizing deterministically with jiwer across currency, ordinals, dates and times. The prompt set holds 30 texts, 10 sampled per run, 480 samples to a 1-day window. The board is English only.

Source: benchmarks.coval.ai/tts, 1-day window, read September 8, 2026 at about 14:00 Paris time, 480 samples per model, 26 models on the board. Average word error rate, lower is better. Median perceived time to first audio shown for context.

Model Avg WER P50 TTFA
Soniox TTS Rt v2 4.0% 255 ms
ElevenLabs Eleven v3 Conversational 4.3% 320 ms
Inworld TTS 2 4.5% 170 ms
Fish Audio S2.1 Pro 4.7% 293 ms
Gradium TTS 4.9% 214 ms
OpenAI GPT-4o mini TTS 4.9% n/a
Deepgram Aura-2 5.0% 290 ms
Rime Mist v3 5.0% 256 ms
Fluxions vui 5.3% 51 ms
Inworld TTS Flash 2 5.3% 75 ms
Cartesia Sonic 3.6 5.3% 440 ms
Palabra TTS v1 5.7% 103 ms
Cartesia Sonic 3.5 5.8% 269 ms
ElevenLabs Flash v2.5 6.5% 185 ms

Gradium reads 4.9% there, within one point of the lowest figure on the board.

Why this page says "about 5%" and not a decimal

Because the daily figure moves more than the gaps between models. On the internal tracker fed by Coval's public series API, Gradium's Coval word error rate read 5.38% on September 6, 6.18% on September 7, 4.81% on September 8, 5.25% on September 9 and 5.29% on September 10, 2026. That is a 1.37-point swing inside five days, on a board where fifth place and fourteenth place are less than a point apart.

Quoting the September 8 reading of 4.81% as a standing result would be a selection, not a measurement. The honest forms are a range with its week, or the 7-day window figure. On the September 10, 2026 read, with 27 models on the board, Gradium sat at 5.29%; the lowest figure was Soniox tts-rt-v2 at 4.37%.

What does a controlled multilingual test show?

Gradium published a per-language word error rate benchmark on April 29, 2026 in Word Error Rate Evaluations, run on the public MiniMax Multilingual TTS Test Set (MiniMaxAI/TTS-Multilingual-Test-Set, 24 languages, of which five were evaluated). Reference Speech-to-Text was Qwen3-ASR. Normalizers were the Whisper English normalizer for English, kyutai/tts_longeval for French and the Whisper basic normalizer for Spanish, Portuguese and German.

Source: Gradium, April 29, 2026, self-reported on a public test set. Word error rate in percent, lower is better. Bold marks the best figure in each column. Cartesia's entry is Sonic-3, the model current at the time of the test.

Model Avg EN FR ES PT DE
Gradium 1.11 0.41 2.16 0.40 2.02 0.54
ElevenLabs Flash v2.5 1.52 0.36 2.45 0.99 3.18 0.61
Cartesia Sonic-3 1.56 0.83 2.66 1.19 2.74 0.37
Mistral Voxtral 1.59 0.88 2.48 1.01 2.87 0.69
ElevenLabs Multilingual v2 1.68 0.37 2.06 1.93 3.34 0.72
Qwen3 TTS 1.98 0.82 2.18 2.61 3.96 0.35

Two readings. First, English is saturated: the top three systems sit inside 0.05 points of each other, so an English-only word error rate no longer separates them. Second, the spread widens sharply on Spanish and Portuguese, where Gradium recorded 0.40% and 2.02% in that April 2026 test against 0.99% and 3.18% for ElevenLabs Flash v2.5. If your agent serves those markets, the per-language column is the one to read, not the average.

This test is dated April 2026 and describes the models as they stood then. Cartesia has since shipped Sonic 3.6, and ElevenLabs has since removed Turbo v2.5 from its models page.

Which models get the hard cases right?

Clean prose is not where production Text-to-Speech fails. Order references, phone numbers, email addresses, ticket IDs and currency amounts are.

Gradium built a 500-sentence set for exactly those, 100 per language across English, French, Spanish, Portuguese and German, and had independent native-speaker raters score the output. Audio was loudness-normalized and randomized, and each rater was capped at 40 comparisons across two sessions. The pass condition is strict: a sentence passes only if an independent native-speaker rater hears every element pronounced correctly and completely.

The ten criteria are seven atomic ones over 350 rows (Spelling, Acronyms, Alphanumerical Tokens, Date, Regular Numbers, Floating and Large Numbers, Email) and three composite ones over 150 rows (Orders, IT Ticket, Claims).

Source: Gradium, "Gradium TTS: latency and accuracy", August 31, 2026, data collected August 28, 2026. 500 sentences, five languages, independent native-speaker raters. Percentage of sentences where every element was heard correctly and completely.

Model Hard-case pass rate
Gradium TTS 81.0%
Cartesia Sonic 3.6 75.1%
ElevenLabs Eleven v3 Conversational 65.4%
Inworld Realtime TTS-2 61.5%
Fish Audio S2.1 Pro 49.5%

Bar chart of hard-case pronunciation pass rate across five languages, from Gradium's August 2026 test of 500 sentences rated by independent native speakers. Gradium TTS leads at 81.0%, shown in blue, followed by Cartesia Sonic 3.6 at 75.1%, ElevenLabs Eleven v3 Conversational at 65.4%, Inworld Realtime TTS-2 at 61.5% and Fish Audio S2.1 Pro at 49.5%.

The four comparison models are the ones Gradium selected, not a full board. The text set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608, so the test can be rerun against any provider.

Note what the spread says: nearly one sentence in five still fails for the best model in that August 2026 test. Hard-case accuracy is not solved, and it is where the remaining headroom is.

Why do these three numbers disagree?

They measure different things, and each is valid inside its own definition.

  • Corpus. Coval uses 30 English prompts. The MiniMax set is multilingual research text. The hard-case set is deliberately adversarial customer-support content.
  • Reference model. Coval uses its own reference Speech-to-Text. The April 2026 test used Qwen3-ASR. The hard-case test used human raters, no Speech-to-Text at all.
  • Normalizer. Coval normalizes with jiwer over currency, ordinals, dates and times. The April test used three different normalizers across five languages. Human raters normalize nothing.
  • Window. Coval is a rolling live measurement that moves daily. The other two are fixed dated runs.

The practical rule: quote one number, name its source, its date and its window, and never average across benchmarks. For the latency side of the same boards, see the TTS latency benchmark.

Which accuracy number should drive your decision?

  • Building an English voice agent on clean scripted text. Clean-text word error rate will not separate the top systems. Test on your own script instead.
  • Reading back order numbers, amounts, codes or addresses. Use a hard-case test. The 500-sentence set is public; run it against your shortlist with your own raters.
  • Serving Spanish, Portuguese, French or German. Read the per-language column of a multilingual test, not the average.
  • Needing accuracy on names and technical terms specifically. No benchmark substitutes for a pronunciation dictionary. See fixing mispronounced names and acronyms and pronunciation dictionaries.
  • Optimizing for how a voice sounds rather than what it says. Word error rate is the wrong metric. Artificial Analysis runs blind pairwise Elo boards; on September 8, 2026 Gradium TTS held 1,149 Elo on the 92-model provider-voice board, and on the controlled-voice per-language boards it held rank 6 of 23 in French, rank 7 of 23 in Portuguese and rank 8 of 24 in German on September 10, 2026. Elo values drift a few points a day.

Glossary

Word error rate (WER). Insertions plus deletions plus substitutions, divided by the number of words in the reference text, after normalization. The standard proxy for Text-to-Speech intelligibility.

Normalization. The deterministic rewriting applied to both the reference text and the transcript before scoring, so that different spellings of the same spoken form do not count as errors. Language-specific, and a major source of disagreement between benchmarks.

Reference Speech-to-Text. The recognizer used to transcribe synthesized audio back to text. Its own errors are counted against the Text-to-Speech model, so two benchmarks with different recognizers are not comparable.

Hard-case pass rate. The share of sentences in an adversarial test set where a native-speaker rater hears every element pronounced correctly and completely. Unlike word error rate it is all-or-nothing per sentence, which is closer to how a caller experiences a wrong digit.

Elo (speech arenas). A preference score derived from blind pairwise votes on which of two samples sounds better. It measures preference, not accuracy, and the two can move in opposite directions.

Code-switching. Changing language inside a single utterance, for example a French sentence containing an English product name. A common source of mispronunciation, covered in stopping accent switches mid-sentence.

References

  1. Coval TTS leaderboard, live board: benchmarks.coval.ai/tts
  2. Coval benchmark harness and methodology, Apache-2.0: github.com/coval-ai/benchmarks
  3. Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
  4. Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
  5. Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608
  6. MiniMax Multilingual TTS Test Set: huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set
  7. Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
  8. Artificial Analysis Speech Arena: artificialanalysis.ai

Part of the Gradium benchmark cluster, hub at TTS latency benchmark 2026. Its siblings:

Beyond this topic

Accuracy is settled at generation time, but most of the fixes are upstream of the model. Text normalization edge cases covers the input side, pronunciation dictionaries the override side, and json_config the parameters that control both.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions