By Gradium. Data as of September 2026.
Key takeaways
- This is the four-way head-to-head. For the criteria framework behind it, see how to choose a TTS API; for the full field of ten-plus providers, see best Text-to-Speech APIs 2026.
- On the Coval leaderboard read September 8, 2026 (1-day window), median perceived time to first audio was 185 ms for ElevenLabs Flash v2.5, 214 ms for Gradium TTS, 269 ms for Cartesia Sonic 3.5, 320 ms for ElevenLabs Eleven v3 Conversational and 440 ms for Cartesia Sonic 3.6.
- On the Artificial Analysis provider-voice Elo board read September 8, 2026, Cartesia Sonic 3.6 led at 1,282 of 92 models, with Eleven v3 Conversational eighth at 1,210 and Gradium TTS at 1,149; Gradium held top-8 controlled-voice Elo in French, Portuguese and German on September 10, 2026.
- On a 500-sentence hard-case set rated by native speakers in August 2026, Gradium TTS passed 81.0%, Cartesia Sonic 3.6 75.1% and ElevenLabs Eleven v3 Conversational 65.4%.
- OpenAI's Text-to-Speech is the cheapest of the four at about $15 per 1M characters for tts-1, and the slowest: Coval measured GPT-4o mini TTS at a P95 of 3,927 ms on September 8, 2026.
Which of the four should you pick?
Each of the four wins a different column, and none wins all of them.
Cartesia takes voice preference. On the Artificial Analysis provider-voice board read September 8, 2026, Sonic 3.6 sat first of 92 at 1,282 Elo. It also has the broadest language list of the four.
ElevenLabs takes median latency and voice catalogue. Flash v2.5 was the fastest of the four on the Coval board read September 8, 2026, at 185 ms, and its library of ready-made voices is the largest here.
Gradium takes hard-case pronunciation and deployment range. It passed 81.0% of a 500-sentence adversarial set in August 2026, 15.6 points ahead of Eleven v3 Conversational, and it is the only one of the four offering on-premise and on-device alongside a cloud API.
OpenAI takes price and convenience. About $15 per 1M characters for tts-1, one API key alongside the rest of an OpenAI stack, and no real-time option worth planning around.
The rest of this page is the evidence, column by column.
How do the four compare on latency?
Source: benchmarks.coval.ai/tts, 1-day window, read September 8, 2026 at about 14:00 Paris time, 480 samples per model, 26 models on the board. Perceived time to first audio in milliseconds, which since June 3, 2026 includes leading silence inside the stream. Lower is better.
| Model | P25 | P50 | P75 | P95 | Avg WER |
|---|---|---|---|---|---|
| ElevenLabs Flash v2.5 | 174 | 185 | 199 | 225 | 6.5% |
| Gradium TTS | 213 | 214 | 244 | 346 | 4.9% |
| Cartesia Sonic 3.5 | 245 | 269 | 291 | 346 | 5.8% |
| ElevenLabs Eleven v3 Conversational | 294 | 320 | 343 | 393 | 4.3% |
| Cartesia Sonic 3.6 | 380 | 440 | 551 | 696 | 5.3% |
| OpenAI GPT-4o mini TTS | n/a | n/a | n/a | 3,927 | 4.9% |
Three things worth reading out of that table.
Cartesia's newer model is slower than its older one on this board. Sonic 3.6 shipped August 27, 2026 and measures 440 ms median against Sonic 3.5's 269 ms. Cartesia's own vendor figure for Sonic 3.6 is "sub-90 ms", which is model inference latency excluding network, a different measurement from what Coval reports.
ElevenLabs publishes the same kind of caveat. Its latency documentation notes that its inference-time figure excludes network, and that real time to first audio is always higher once network (20 to 200 ms) and player buffer are counted.
OpenAI is not in the real-time conversation. Its standard endpoint does not stream over WebSocket, and a P95 near four seconds puts it in batch territory.
One methodology note that governs every number above: Coval redefined time to first audio on June 3, 2026 to include leading silence inside the stream. Any comparison of these four dated before then measures round trip only. Gradium's own board figure moved from 171.9 ms to 429.6 ms on that date with no change on its side, and the model it shipped on August 31, 2026 measures 0 ms leading silence, rank 1 of 27 on the September 10, 2026 board. The full account is in How a benchmark change produced a faster TTS model.
How do the four compare on accuracy?
Two tests, because clean-text word error rate no longer separates the leaders.
Source: Gradium, "Word Error Rate Evaluations", April 29, 2026, on the MiniMax Multilingual TTS Test Set with Qwen3-ASR. Word error rate in percent, lower is better. Cartesia's entry is Sonic-3, the model current at the time of the test.
| Model | Avg | EN | FR | ES | PT | DE |
|---|---|---|---|---|---|---|
| Gradium | 1.11 | 0.41 | 2.16 | 0.40 | 2.02 | 0.54 |
| ElevenLabs Flash v2.5 | 1.52 | 0.36 | 2.45 | 0.99 | 3.18 | 0.61 |
| Cartesia Sonic-3 | 1.56 | 0.83 | 2.66 | 1.19 | 2.74 | 0.37 |
| ElevenLabs Multilingual v2 | 1.68 | 0.37 | 2.06 | 1.93 | 3.34 | 0.72 |
The English column is a tie inside 0.05 points. Spanish and Portuguese are where the four separate.
Source: Gradium, "Gradium TTS: latency and accuracy", August 31, 2026, data collected August 28, 2026. 500 sentences, five languages, ten criteria, independent native-speaker raters. A sentence passes only if a rater hears every element pronounced correctly and completely.
| Model | Hard-case pass rate |
|---|---|
| Gradium TTS | 81.0% |
| Cartesia Sonic 3.6 | 75.1% |
| ElevenLabs Eleven v3 Conversational | 65.4% |
OpenAI was not in that test. The text set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608, so any of the four can be scored against it.
This is the column that decides support and telephony products. An agent that reads back an order reference or a confirmation code fails visibly when it gets one character wrong, and clean-prose word error rate does not predict that.
How do the four compare on how the voice sounds?
Source: Artificial Analysis Speech Arena, provider-voice board, fetched September 8, 2026. Elo from blind pairwise votes on English audio. Higher is better. Elo values drift a few points per day and ranks move by one or two.
| Model | Elo | Rank of 92 |
|---|---|---|
| Cartesia Sonic 3.6 | 1,282 | 1 |
| ElevenLabs Eleven v3 Conversational | 1,210 | 8 |
| ElevenLabs Eleven v3 | 1,175 | 14 |
| Gradium TTS | 1,149 | 18 |
Elo measures preference, not correctness, and the two dimensions genuinely disagree here: Cartesia ranked first on the preference board on September 8, 2026, while Gradium ranked first on the hard-case test in August 2026. Which one matters is a product question, not a benchmark question. A narration product should weight Elo; a support agent reading account numbers should weight the pass rate.
On cloned voices the picture differs again. In a blinded A/B study over 3,220 voice pairs across English, French, Spanish and German, published January 22, 2026, Gradium held the highest live Elo in every language from a 10-second source sample. That test compared against ElevenLabs Flash only.
How do the four compare on price?
Source: vendor pricing pages, read September 8, 2026. Text-to-Speech only; Speech-to-Text is billed separately by all four.
| Provider and model | Listed rate |
|---|---|
| OpenAI tts-1 | about $15 per 1M characters |
| OpenAI tts-1-hd | about $30 per 1M characters |
| Gradium L plan ($1,615/mo) | $35.90 per 1M characters effective |
| Gradium M plan ($340/mo) | $37.80 per 1M characters effective |
| Gradium S plan ($43/mo) | $47.80 per 1M characters effective |
| ElevenLabs Flash v2.5 and Eleven v3 Conversational | $50.00 per 1M characters |
| Gradium XS plan ($13/mo) | $57.80 per 1M characters effective |
| ElevenLabs Eleven v3 and Multilingual v2 | $100.00 per 1M characters |
Cartesia prices in credits rather than characters, so it does not convert to this column directly: Free at $0 for 20k credits, Pro at $5 for 100k, Startup at $49 for 1.25M and Scale at $299 for 8M, with Text-to-Speech concurrency of 2, 3, 5 and 15 by tier.
Two adjustments to make before comparing: ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026, and OpenAI's GPT-4o mini TTS is token-priced (about $0.015 per minute) rather than per character. Gradium's rates are the effective per-character cost of each plan's credit allocation at 1 credit per character; annual billing gives twelve months for eleven. The full method is in how to compare TTS pricing across providers.
What else separates them?
Languages. Gradium covers English, French, Spanish, Portuguese and German with mid-sentence code-switching. Cartesia and ElevenLabs both cover substantially longer lists; check their models pages for current counts. If your market is outside Gradium's five, that decides it before any benchmark does.
Speech-to-Text in the same platform. Gradium's Speech-to-Text draws on the same credit pool as Text-to-Speech at 3 credits per second, with native semantic voice activity detection. ElevenLabs bills Scribe v2 separately at $0.22 per hour, or $0.39 per hour for Scribe v2 Realtime. Cartesia shipped Ink-2 on July 9, 2026. OpenAI offers transcription separately.
Voice cloning. Gradium's Instant Voice Clone takes 10 seconds of audio, with 5 clones on the free tier for non-commercial use and 1,000 per month on paid plans; its Pro Voice Clone needs 30 minutes of clean audio minimum, 2 hours recommended, from the M plan up. ElevenLabs and Cartesia both publish instant and professional tiers.
Deployment. Gradium is the only one of the four offering cloud API, dedicated instances, self-hosted, on-premise and on-device, with EU and US regional endpoints and Zero Data Retention self-serve on all paid plans since September 2, 2026.
Concurrency. Gradium allows 2, 5, 5, 10 and 15 concurrent Text-to-Speech streams from Free through L, custom on Enterprise. ElevenLabs allows 4, 6, 10, 20, 30 and 30 on Flash by tier. Cartesia allows 2, 3, 5 and 15.
Session length. Gradium allows 3,000 seconds per session on both Text-to-Speech and Speech-to-Text.
Which constraint points where?
| If your hard constraint is | Start with |
|---|---|
| Lowest median latency among these four (September 2026) | ElevenLabs Flash v2.5 |
| Reading back numbers, codes, emails without errors | Gradium |
| Highest blind-preference Elo | Cartesia Sonic 3.6 |
| Lowest cost per character | OpenAI tts-1 |
| A language outside English, French, Spanish, Portuguese and German | Cartesia or ElevenLabs |
| On-premise or on-device deployment | Gradium |
| Speech-to-Text and Text-to-Speech on one credit pool | Gradium |
| Staying inside an existing OpenAI stack | OpenAI |
Glossary
Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.
Leading silence. Inaudible samples at the start of a synthesized stream. Coval measures it as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal, and it accounts for 28% of time to first audio across the board.
Elo (speech arenas). A preference score derived from blind pairwise votes on which of two samples sounds better. Preference, not accuracy; the two can and do move in opposite directions.
Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.
Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys. The only way to compare a credit-based plan against a per-character list price.
State Space Model (SSM). The architecture family behind Cartesia's Sonic line. Cartesia claims lower time to first audio, lower real-time factor and higher throughput from it.
References
- Coval TTS leaderboard: benchmarks.coval.ai/tts
- Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
- Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
- Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
- Artificial Analysis Speech Arena: artificialanalysis.ai
- ElevenLabs, "Understanding latency": elevenlabs.io/docs/eleven-api/concepts/latency
- Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608
Related guides
This page is the four-way head-to-head in the Text-to-Speech selection cluster. Its siblings each answer a different shape of the same question:
- Best Text-to-Speech APIs 2026: the broad ranking, ten-plus providers, the cluster hub.
- How to choose a TTS API: the criteria framework, before you have a shortlist.
- Best TTS API 2026: the short verdict page, one pick per use case.
- Top 3 Text-to-Speech solutions 2026: the three-way head-to-head with a scoring table.
- Best Text-to-Speech API for voice agents: the voice-agent hub.
- Best AI voice generators 2026: consumer and creator tools rather than APIs.
- Best speech APIs 2026: Text-to-Speech and Speech-to-Text from one vendor.
Beyond this topic
Once the provider is chosen, the work moves to the stack around it: best API to build an AI voice agent covers orchestration, TTS latency benchmark 2026 covers the measurement in depth, and best low-latency TTS APIs covers the pipeline budget rather than the synthesis stage alone.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

