By Gradium. Data as of September 2026.
Key takeaways
- On the Coval TTS leaderboard read on September 8, 2026 (1-day window, 480 samples per model, 26 models), Gradium TTS records a median perceived time to first audio of 214 ms with a 31 ms P25 to P75 spread, on a board whose medians ran from 51 ms to 440 ms.
- Gradium TTS measures 0 ms leading silence on the Coval leaderboard (September 10, 2026, 27 models), the only model on the board at zero.
- Its P25 to P75 spread is 31 ms on the same September 8, 2026 read, among the tightest on the board; P95 is 346 ms.
- Coval redefined time to first audio on June 3, 2026 to include leading silence. Every Coval figure published before that date is round-trip only and is not comparable with the numbers on this page.
- Coval word error rate for Gradium moves day to day and sat at about 5% through the week of September 6 to 10, 2026; the lowest reading on that board was Soniox tts-rt-v2 at about 4.4%.
What did the Coval TTS leaderboard show on September 8, 2026?
Coval is an independent voice AI evaluation platform that continuously tests production Text-to-Speech endpoints. It is not affiliated with any provider, its harness is open source under Apache-2.0 at github.com/coval-ai/benchmarks, and the board refreshes roughly every 30 minutes across 1-day, 7-day and 30-day windows.
The table below is a single read of the live dashboard. It is a snapshot, not a standing result: check the board itself for the current numbers.
Source: benchmarks.coval.ai/tts, 1-day window, read September 8, 2026 at about 14:00 Paris time, 480 samples per model, 26 models on the board. All latency values are perceived time to first audio in milliseconds. Lower is better on every column.
| Model | P25 | P50 | P75 | P95 | Avg WER |
|---|---|---|---|---|---|
| Fluxions vui | 27 | 51 | 66 | 96 | 5.3% |
| Inworld TTS Flash 2 | 70 | 75 | 87 | 111 | 5.3% |
| Palabra TTS v1 | 98 | 103 | 115 | 144 | 5.7% |
| Inworld TTS 2 | 152 | 170 | 190 | 230 | 4.5% |
| ElevenLabs Flash v2.5 | 174 | 185 | 199 | 225 | 6.5% |
| Gradium TTS | 213 | 214 | 244 | 346 | 4.9% |
| Soniox TTS Rt v2 | 214 | 255 | 286 | 311 | 4.0% |
| Rime Mist v3 | 246 | 256 | 267 | 295 | 5.0% |
| Cartesia Sonic 3.5 | 245 | 269 | 291 | 346 | 5.8% |
| Deepgram Aura-2 | 186 | 290 | 417 | 522 | 5.0% |
| Fish Audio S2.1 Pro | 268 | 293 | 335 | 586 | 4.7% |
| ElevenLabs Eleven v3 Conversational | 294 | 320 | 343 | 393 | 4.3% |
| Cartesia Sonic 3.6 | 380 | 440 | 551 | 696 | 5.3% |
| OpenAI GPT-4o mini TTS | n/a | n/a | n/a | 3,927 | 4.9% |
Reading it: in this September 8, 2026 read Gradium TTS records 214 ms median perceived time to first audio on a board whose medians ran from 51 ms to 440 ms, and its 31 ms P25 to P75 spread is among the tightest on the board. Its 4.9% word error rate sits within one point of the lowest figures in the same read: Soniox TTS Rt v2 at 4.0%, ElevenLabs Eleven v3 Conversational at 4.3%, Inworld TTS 2 at 4.5% and Fish Audio S2.1 Pro at 4.7%.
A separate read on September 10, 2026 at 09:10 UTC, with 27 models on the board, put Gradium TTS at 0 ms leading silence, rank 1 of 27 and the only model at zero, with 214 ms perceived time to first audio, 214 ms round-trip time to first audio and 5.29% average word error rate. Fluxions vui led perceived time to first audio at 53 ms and round trip at 25 ms; Soniox tts-rt-v2 led word error rate at 4.37%.
Those two rows are the whole story of this page: Gradium's perceived and round-trip figures are identical because it ships no leading silence at all, so every millisecond of its 214 ms is audio the listener hears.
For the latency figures behind this section, and how Coval has measured perceived time to first audio since June 3, 2026, see How a benchmark change produced a faster TTS model.
What is perceived time to first audio?
Perceived time to first audio is the delay a listener actually experiences before hearing speech. Coval defines it, verbatim, as "TTFA = (first audio chunk arrival - synthesis start) + leading silence inside the stream before the first audible sample".
The second term is the part most benchmarks miss. A model can return its first audio chunk quickly and then spend 200 ms of that stream on silence before the first audible sample. The stopwatch stops early; the listener still waits.
Coval finds the onset with a fixed threshold rather than a voice activity model. Verbatim: "The perceived onset of speech is then the first 10 ms window where RMS exceeds 0.01 on a normalized signal, everything before is leading silence."
Across the whole board, leading silence accounts for 28% of time to first audio. That is the measurement gap between what vendors publish and what callers hear.
Why did Gradium's Coval figure move from 171.9 ms to 429.6 ms on June 3, 2026?
Because the definition changed, not the model. Coval shipped perceived time to first audio to the leaderboard on June 3, 2026. On that date Gradium's time to first audio on the Coval TTS leaderboard moved from 171.9 ms to 429.6 ms with no code or infrastructure change on Gradium's side. The previous model carried a 225 ms median leading silence that the old metric never counted.
The model Gradium released on August 31, 2026 removes it. Its leading silence measures 0 ms, and its median perceived time to first audio is 214 ms, roughly 170 ms faster than the model it replaced. The full account, including Coval's own framing of the change, is in the joint Coval and Gradium post How a benchmark change produced a faster TTS model, published September 9, 2026.
The practical consequence for anyone reading older comparisons: every Coval time to first audio figure published before June 3, 2026 is round-trip only. That includes the 155 ms and 158 ms medians, the 2 ms interquartile range, the 171.9 ms figure and the OpenAI TTS-1-HD 2,295 ms figure that circulated widely in the spring of 2026. None of them are comparable with the numbers in the table above. Coval also stopped tracking TTS-1-HD; OpenAI now appears on the board as GPT-4o mini TTS.
What was the last comparable snapshot before this one?
Gradium published a five-model comparison drawn from Coval's 1-day window on August 28, 2026, alongside its hard-case accuracy results. On that date Gradium TTS measured 216 ms median perceived time to first audio with a 30 ms P25 to P75 spread over 480 runs. Inworld Realtime TTS-2 measured 166 ms, Fish Audio S2.1 Pro 291 ms, ElevenLabs Eleven v3 Conversational 329 ms, and Cartesia Sonic 3.6 454 ms with a 165 ms spread.
Those five are the models Gradium chose to compare against. They are not the full Coval board, which tracked 26 to 30 models through that period.
What does Coval measure, and what does it leave out?
Coval reports time to first audio at P50, P75, P95 and P99 plus a P25 to P75 latency range, with a dashboard slider exposing P25 through P100. Word error rate comes from a reference Speech-to-Text pass after deterministic normalization with jiwer, covering currency, ordinals, dates and times. The text set holds 30 prompts, 10 sampled per run; 480 samples make up one 1-day window. The board is English only.
Two things it deliberately does not measure: voice cloning, and the pre-connection handshake. TLS plus WebSocket upgrade costs roughly 50 to 200 ms and is excluded uniformly across every model, so the board compares synthesis and streaming rather than connection setup. If your agent opens a fresh connection per turn, add that cost back. If it holds a connection open, see WebSocket multiplexing on Gradium TTS.
Coval now exposes three latency series per model: perceived time to first audio at P50, which is the headline, round-trip time to first audio at P50, and leading silence at P50. Comparing a vendor's published number against the board requires knowing which of the three it corresponds to.
Coval was founded in 2024 by Brooke Hopkins, previously tech lead for testing and evaluation on self-driving at Waymo. It went through Y Combinator in summer 2024 and has raised a $3.3M seed from MaC VC and a $28M Series A led by Norwest in 2026.
Why quote Coval word error rate as "about 5%"?
Because a single-day decimal is noise. On the internal tracker fed by Coval's public series API, Gradium's Coval word error rate read 5.38% on September 6, 6.18% on September 7, 4.81% on September 8, 5.25% on September 9 and 5.29% on September 10, 2026. Quoting any one of those as a stable figure would misrepresent the board.
The honest form is a range with a date, or the 7-day figure. Through the week of September 6 to 10, 2026 Gradium's Coval word error rate sat at about 5% on a 27-model board.
Coval word error rate is also not the same measurement as a controlled multilingual test. On the MiniMax Multilingual TTS Test Set, evaluated with Qwen3-ASR across English, French, Spanish, Portuguese and German, Gradium recorded a 1.11% average word error rate in April 2026, against ElevenLabs Flash v2.5 at 1.52% and Cartesia Sonic-3 at 1.56%. The methodology and per-language numbers are in Word Error Rate Evaluations and broken down further on the TTS WER benchmark page. Different corpus, different reference Speech-to-Text, different normalizer, so the two figures should never be presented as one series.
What does Gradium's own latency benchmark measure?
Gradium published a controlled time to first audio benchmark on March 24, 2026, in Time to First Audio. Measured from Paris, 100 queries per model, first five discarded for a warm state, 15 to 25 word inputs, the same output format and sample rate for every provider.
Source: Gradium, March 24, 2026, self-reported. Round-trip time to first audio in milliseconds, measured from Paris on a multiplexed connection. The ElevenLabs model named here is the one tested in March 2026; ElevenLabs has since removed Turbo v2.5 from its models page.
| Model | P25 | P50 | P95 |
|---|---|---|---|
| Gradium (excluding connection setup) | 212 | 214 | 228 |
| Gradium (including connection setup) | 255 | 258 | 274 |
| ElevenLabs Turbo v2.5 (excluding connection setup) | n/a | 257 | n/a |
| ElevenLabs Turbo v2.5 (including connection setup) | n/a | 304 | n/a |
| ElevenLabs Flash v2.5 (excluding connection setup) | n/a | 277 | n/a |
| ElevenLabs Flash v2.5 (including connection setup) | n/a | 324 | n/a |
This is a round-trip measurement taken before the June 3, 2026 methodology change, so it is not comparable with the perceived figures at the top of this page. It is useful for one thing: the gap between the two Gradium rows is the cost of opening a connection per turn, roughly 44 ms at the median.
The same post documents how to measure the metric correctly. For WAV, discard the 44-byte header. For Ogg/Opus, skip the identification and comment header pages. For MP3, skip ID3 tags and detect the first valid MPEG frame. A benchmark that stops its clock on the first byte is timing metadata delivery, not speech.
What latency should a voice agent budget for?
Human conversation is the reference. Across ten languages, the gap between one speaker finishing and the next beginning averages 208 ms (Stivers et al., PNAS, 2009). Gradium's own guidance puts the Text-to-Speech share of the budget at 200 to 300 ms.
Distribution shape matters as much as the median. As the May 2026 Coval post puts it, "a model with 200 ms P50 and 600 ms P75 will feel worse than one with 250 ms P50 and 300 ms P75." A tight spread means every turn feels the same; a wide one means a visible fraction of turns feel broken.
Read the September 8, 2026 table with that in mind:
- Lowest median on the board. Fluxions vui at 51 ms and Inworld TTS Flash 2 at 75 ms are the two fastest medians in that read.
- Tightest spread. Inworld TTS Flash 2 and Palabra TTS v1 at 17 ms P25 to P75, Rime Mist v3 at 21 ms, ElevenLabs Flash v2.5 at 25 ms and Gradium TTS at 31 ms are the tightest in that read; Deepgram Aura-2 at 231 ms and Cartesia Sonic 3.6 at 171 ms are the widest.
- Lowest word error rate. Soniox TTS Rt v2 at 4.0% led that read, with Eleven v3 Conversational at 4.3% next.
- Not viable for real-time. OpenAI GPT-4o mini TTS at a P95 of 3,927 ms belongs in batch generation.
No model won every column, which is the point of reading a live board rather than a vendor page. For a framework that turns these columns into a decision, see how to choose a TTS API; for the head-to-head between the named providers, see ElevenLabs vs Cartesia vs OpenAI vs Gradium.
Glossary
Time to first audio (TTFA). The elapsed time between sending a synthesis request and the first playable audio sample arriving. Distinct from time to first byte, which can be satisfied by a container header carrying no audio.
Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. This is what a listener experiences and what the Coval leaderboard has reported since June 3, 2026.
Leading silence. The stretch of inaudible samples at the start of a synthesized stream, measured by Coval as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal. It accounts for 28% of time to first audio across the Coval board.
P25 to P75 spread. The gap between the 25th and 75th percentile of a latency distribution, also called the interquartile range. A direct measure of how consistent latency is; the median alone cannot show it.
Word error rate (WER). The share of words a reference Speech-to-Text system transcribes incorrectly from synthesized audio, after deterministic text normalization. A proxy for intelligibility, sensitive to the reference model and the normalizer used.
Delayed Streams Modeling (DSM). The streaming sequence-to-sequence architecture Gradium's models are built on, described in "Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling" (arXiv:2509.08753, Kyutai).
References
- Coval TTS leaderboard, live board: benchmarks.coval.ai/tts
- Coval benchmark harness and methodology, Apache-2.0: github.com/coval-ai/benchmarks
- Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
- Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
- Gradium, "Time to First Audio", March 24, 2026: gradium.ai/blog/time-to-first-audio
- Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
- Gradium, "Coval TTS benchmarks", May 13, 2026: gradium.ai/blog/coval-tts-benchmarks-may-2026
- Stivers et al., "Universals and cultural variation in turn-taking in conversation", PNAS 106(26):10587-92, 2009: doi.org/10.1073/pnas.0903616106
- Kyutai, "Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling", arXiv:2509.08753: arxiv.org/abs/2509.08753
Related guides
This page is the hub of the Gradium benchmark cluster. Its siblings:
- TTS WER benchmark 2026: accuracy rather than latency, with the multilingual and hard-case results in full.
- Gradium vs ElevenLabs for voice agents: the two-provider head-to-head on latency, accuracy and pricing.
- STT API benchmark 2026: the same treatment for Speech-to-Text.
- On-device TTS benchmark 2026: Phonon against Kani, NeuTTS and Magpie, where no network is involved.
- Best low-latency TTS APIs 2026: the latency cluster hub, covering the full pipeline rather than the Text-to-Speech stage alone.
Beyond this topic
Latency is one stage of a voice agent. The turn cannot start until the agent knows the caller has stopped speaking, which is the subject of turn-taking and VAD in voice agents. Once the text arrives, holding one connection open across turns is the cheapest latency win available, covered in WebSocket multiplexing. And the numbers only matter if the words come out right, which is where pronouncing numbers, dates and currencies picks up.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

