By Gradium. Data as of September 2026.
Key takeaways
- This page scores three providers against seven criteria and totals the scores under two weightings. For the four-way comparison including OpenAI, see ElevenLabs vs Cartesia vs OpenAI vs Gradium.
- Under a live-agent weighting Gradium and ElevenLabs tie at the top, for opposite reasons. Under a content-production weighting of the same September 2026 data ElevenLabs scores highest. Same data, different weights.
- ElevenLabs Flash v2.5 had the lowest median perceived time to first audio of the three on the Coval board read September 8, 2026, at 185 ms against Gradium's 214 ms and Cartesia Sonic 3.6's 440 ms.
- Cartesia Sonic 3.6 led the Artificial Analysis provider-voice board on September 8, 2026 at 1,282 Elo, rank 1 of 92.
- Gradium TTS led the August 2026 hard-case pronunciation test at 81.0%, against 75.1% for Cartesia Sonic 3.6 and 65.4% for ElevenLabs Eleven v3 Conversational.
Why these three, and how they are scored
The three are Gradium, ElevenLabs and Cartesia: the providers that appear on both the Coval latency board and the Artificial Analysis preference board, and that were all three in Gradium's August 2026 hard-case test. That overlap is what makes a like-for-like score possible. Providers that appear on only one board are covered in best Text-to-Speech APIs 2026.
Each criterion is scored 3, 2 or 1 by rank among these three on a named, dated measurement. No criterion is scored on impression. Where a provider has several models, the best-placed one on that criterion is used and named.
The scoring table
Scores are ranks among these three providers only. Sources: Coval leaderboard, 1-day window, read September 8, 2026 (latency, spread); Gradium hard-case evaluation, August 2026, 500 sentences, native-speaker raters (accuracy); Artificial Analysis provider-voice board, September 8, 2026 (preference); vendor pricing pages, September 8, 2026 (cost); vendor documentation (deployment, languages, cloning).
| Criterion | Gradium | ElevenLabs | Cartesia | Measured value behind the score |
|---|---|---|---|---|
| Median perceived time to first audio | 2 | 3 | 1 | 214 ms / 185 ms (Flash v2.5) / 440 ms (Sonic 3.6) |
| P25 to P75 latency spread | 2 | 3 | 1 | 31 ms / 25 ms (Flash v2.5) / 171 ms (Sonic 3.6) |
| Hard-case pronunciation | 3 | 1 | 2 | 81.0% / 65.4% (Eleven v3 Conversational) / 75.1% (Sonic 3.6) |
| Blind-preference Elo | 1 | 2 | 3 | 1,149 / 1,210 (Eleven v3 Conversational) / 1,282 (Sonic 3.6) |
| Cost per 1M characters | 3 | 2 | 1 | $35.90 to $57.80 by plan / $50 to $100 by model / credit-priced, no clean conversion |
| Deployment range | 3 | 1 | 1 | Cloud, dedicated, self-hosted, on-premise, on-device / cloud / cloud |
| Language breadth | 1 | 3 | 3 | Five languages / substantially longer list / substantially longer list |
Two totals, because the weighting is the decision. Under the live-agent weighting, median time to first audio, latency spread and hard-case pronunciation count double and the other four criteria count once. Under the content-production weighting, preference Elo and language breadth count double and the other five count once.
| Weighting | Gradium | ElevenLabs | Cartesia |
|---|---|---|---|
| Live agent | 22 | 22 | 16 |
| Content production | 17 | 20 | 18 |
Gradium and ElevenLabs tie on the live-agent weighting, and they get there from opposite directions: ElevenLabs on median latency and spread, Gradium on hard-case accuracy, cost and deployment range. That tie is the honest result, and it means the choice between those two comes down to which failure mode your product cannot afford, not to a total. Cartesia trails on that weighting and closes most of the gap on the content one, where its preference Elo counts double.
The scores are a device for making the tradeoff visible, not a verdict. Read the measured-value column, not the digits.
The measurements behind the scores
Latency
Source: benchmarks.coval.ai/tts, 1-day window, read September 8, 2026 at about 14:00 Paris time, 480 samples per model. Perceived time to first audio in milliseconds, which since June 3, 2026 includes leading silence inside the stream. Lower is better.
| Model | P25 | P50 | P75 | P95 | Avg WER |
|---|---|---|---|---|---|
| ElevenLabs Flash v2.5 | 174 | 185 | 199 | 225 | 6.5% |
| Gradium TTS | 213 | 214 | 244 | 346 | 4.9% |
| Cartesia Sonic 3.5 | 245 | 269 | 291 | 346 | 5.8% |
| ElevenLabs Eleven v3 Conversational | 294 | 320 | 343 | 393 | 4.3% |
| Cartesia Sonic 3.6 | 380 | 440 | 551 | 696 | 5.3% |
Cartesia's newer model measures slower than its older one on this board. Its own published figure for Sonic 3.6 is "sub-90 ms", which is model inference time excluding network, not the end-to-end measurement Coval reports.
Gradium's 214 ms carries 0 ms leading silence, rank 1 of 27 on the September 10, 2026 board and the only model at zero, which is why its perceived and round-trip figures are identical at 214 ms: none of the stream is spent on silence. Board-wide, leading silence accounts for 28% of time to first audio. Any figure dated before June 3, 2026 measures round trip only; see How a benchmark change produced a faster TTS model.
Accuracy
Source: Gradium, "Gradium TTS: latency and accuracy", August 31, 2026, data collected August 28, 2026. 500 sentences, five languages, ten criteria, independent native-speaker raters. A sentence passes only if a rater hears every element pronounced correctly and completely.
| Model | Hard-case pass rate |
|---|---|
| Gradium TTS | 81.0% |
| Cartesia Sonic 3.6 | 75.1% |
| ElevenLabs Eleven v3 Conversational | 65.4% |
On clean multilingual text the gap is much smaller. In the April 2026 MiniMax Multilingual TTS Test Set run, Gradium averaged 1.11% word error rate, ElevenLabs Flash v2.5 1.52% and Cartesia Sonic-3 1.56%, with the English column a tie inside 0.05 points.
The hard-case set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608.
Voice preference
Source: Artificial Analysis Speech Arena, provider-voice board, fetched September 8, 2026. Elo from blind pairwise votes on English audio. Values drift a few points per day.
| Model | Elo | Rank of 92 |
|---|---|---|
| Cartesia Sonic 3.6 | 1,282 | 1 |
| ElevenLabs Eleven v3 Conversational | 1,210 | 8 |
| Gradium TTS | 1,149 | 18 |
Preference and correctness point in opposite directions across these three, which is the single most useful thing this page has to say. A product whose failure mode is "sounds slightly flat" should weight this table. A product whose failure mode is "read the confirmation code wrong" should weight the previous one.
Cost
Source: vendor pricing pages, read September 8, 2026. Text-to-Speech only.
| Provider | Rate per 1M characters |
|---|---|
| Gradium L plan | $35.90 effective |
| Gradium M plan | $37.80 effective |
| Gradium S plan | $47.80 effective |
| ElevenLabs Flash v2.5, Eleven v3 Conversational | $50 |
| Gradium XS plan | $57.80 effective |
| ElevenLabs Eleven v3, Multilingual v2 | $100 |
Cartesia prices in credits ($5 for 100k on Pro, $49 for 1.25M on Startup, $299 for 8M on Scale), which does not convert cleanly to this column. ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026.
Deployment, languages and cloning
Gradium offers cloud API, dedicated instances, self-hosted, on-premise and on-device, with EU and US regional endpoints and Zero Data Retention self-serve on all paid plans since September 2, 2026. ElevenLabs and Cartesia are cloud APIs.
Gradium covers English, French, Spanish, Portuguese and German with mid-sentence code-switching. Cartesia and ElevenLabs both cover substantially longer lists; check their models pages for current counts.
All three offer instant and professional cloning tiers. Gradium's Instant Voice Clone takes 10 seconds of audio, with 5 clones on the free tier for non-commercial use and 1,000 per month on paid plans; its Pro Voice Clone needs 30 minutes of clean audio minimum from the M plan up.
Which of the three, in one line each
Gradium if the agent reads structured content aloud, if on-premise or on-device is on the requirements list, or if per-character cost at volume matters.
ElevenLabs if median latency is the binding constraint, if the catalogue of ready-made voices is part of the product, or if you need a language outside Gradium's five.
Cartesia if the voice itself is the product and preference Elo is the metric you are optimizing.
Glossary
Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.
Leading silence. Inaudible samples at the start of a synthesized stream, measured as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal.
P25 to P75 spread. The gap between the 25th and 75th percentile of latency. A direct measure of consistency.
Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely.
Elo (speech arenas). A preference score from blind pairwise votes. Preference, not accuracy.
Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys.
References
- Coval TTS leaderboard: benchmarks.coval.ai/tts
- Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
- Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
- Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
- Artificial Analysis Speech Arena: artificialanalysis.ai
- Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608
Related guides
This is the scored three-way head-to-head in the Text-to-Speech selection cluster:
- Best Text-to-Speech APIs 2026: the broad ranking across ten-plus providers, the cluster hub.
- ElevenLabs vs Cartesia vs OpenAI vs Gradium: the same exercise with OpenAI added.
- How to choose a TTS API: the criteria framework behind the scoring.
- Best TTS API 2026: the short verdict page, one pick per use case.
- Best Text-to-Speech API for voice agents: the voice-agent hub.
- Best speech APIs 2026: Text-to-Speech and Speech-to-Text from one vendor.
Beyond this topic
Once you have picked one of the three, best API to build an AI voice agent covers the orchestration around it, TTS latency benchmark 2026 covers the latency measurement in full, and TTS WER benchmark 2026 covers the accuracy side.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

