Voice Design. Live today.

Comparison

Best TTS API 2026: The Short Verdict by Use Case

6 min read
Updated

By Gradium. Data as of September 2026. Refreshed monthly.

Key takeaways

  • This is the short verdict page: one recommendation per use case, refreshed monthly. For the reasoning behind each pick, follow the linked guide.
  • Live voice agent reading structured content: Gradium TTS. 81.0% hard-case pass rate in the August 2026 test, 214 ms median perceived time to first audio on Coval, September 8, 2026.
  • Lowest median latency, September 2026: Fluxions vui at 51 ms and Inworld TTS Flash 2 at 75 ms on the Coval board, September 8, 2026.
  • Best-sounding voice: Cartesia Sonic 3.6, first of 92 at 1,282 Elo on the Artificial Analysis provider-voice board, September 8, 2026.
  • Cheapest per character: OpenAI tts-1 at about $15 per 1M characters, September 2026, with no viable real-time path.

The verdict, by use case

Every pick below is a dated reading, not a standing result. Both public boards move; the sources are named under each row.

Use case Pick Why, with date
Live agent reading numbers, codes, emails Gradium TTS 81.0% hard-case pass rate, August 2026, against 75.1% for Cartesia Sonic 3.6 and 65.4% for ElevenLabs Eleven v3 Conversational
Lowest median time to first audio Fluxions vui 51 ms median perceived time to first audio, Coval, September 8, 2026
Low latency with a tight spread ElevenLabs Flash v2.5 or Gradium TTS 185 ms median with a 25 ms P25 to P75 spread, and 214 ms with 31 ms, Coval, September 8, 2026
Best-rated voice quality Cartesia Sonic 3.6 1,282 Elo, rank 1 of 92, Artificial Analysis provider-voice board, September 8, 2026
Audiobooks and narration ElevenLabs Eleven v3 or Cartesia Sonic 3.6 1,175 and 1,282 Elo, Artificial Analysis, September 8, 2026; latency is not a constraint here
A language outside EN, FR, ES, PT, DE Cartesia or ElevenLabs Both publish substantially longer language lists than Gradium's five
Multilingual accuracy across those five Gradium TTS 1.11% average word error rate on the MiniMax Multilingual TTS Test Set, April 2026
Lowest cost per character OpenAI tts-1 about $15 per 1M characters, September 2026
Speech-to-Text and Text-to-Speech on one bill Gradium or Deepgram Gradium draws both from one credit pool; Deepgram bills Nova-3 and Aura-2 separately on one platform
Cloning from a free tier Gradium Instant Voice Clone from 10 seconds of audio, 5 clones on the free tier, non-commercial
On-device or offline Gradium Phonon Roughly 100M parameters, five languages since July 15, 2026
Already inside an OpenAI stack OpenAI One API key and one bill; accept batch-only latency

Why these picks and not a single winner

Three boards disagree, and each is right about a different thing.

Coval measures perceived time to first audio and word error rate against production endpoints, continuously. On the read of September 8, 2026 (1-day window, 480 samples per model, 26 models), medians ran from 51 ms to 440 ms. Gradium TTS recorded 214 ms with a 31 ms P25 to P75 spread, 0 ms leading silence and 4.9% word error rate. Coval word error rate moves day to day, so treat it as a range: Gradium's read between 4.81% and 6.18% across September 6 to 10, 2026.

Artificial Analysis measures blind pairwise preference, which is a different question from accuracy. On the provider-voice board of September 8, 2026, Cartesia Sonic 3.6 led at 1,282 Elo of 92 models, Inworld Realtime TTS-2 was second at 1,252, ElevenLabs Eleven v3 Conversational eighth at 1,210 and Gradium TTS at 1,149; on the controlled-voice per-language boards of September 10, 2026 Gradium held rank 6 of 23 in French, 7 of 23 in Portuguese and 8 of 24 in German.

Hard-case testing measures whether the model gets your content right. In Gradium's August 2026 test over 500 sentences across five languages, rated by independent native speakers with an all-or-nothing pass condition, Gradium TTS passed 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs Eleven v3 Conversational 65.4%, Inworld Realtime TTS-2 61.5% and Fish Audio S2.1 Pro 49.5%.

Preference and correctness genuinely diverge here: in September 2026 Cartesia Sonic 3.6 topped the preference board and Gradium the hard-case pronunciation test. Which board is yours depends on whether the product's failure mode is "sounds a bit flat" or "read the account number wrong".

One methodology note that invalidates older rankings

On June 3, 2026 the Coval leaderboard changed what time to first audio means. It now counts the leading silence inside the stream before the first audible sample, on top of the round trip. Across the board, leading silence accounts for 28% of time to first audio.

Any "best TTS API" list quoting a latency figure from before that date, including the 155 ms and 158 ms medians and the 2 ms interquartile range widely repeated in spring 2026, is quoting a round-trip number against a perceived-latency board. The two are not comparable. The full account is in How a benchmark change produced a faster TTS model, September 9, 2026.

Pricing at a glance

Source: vendor pricing pages, read September 8, 2026. Text-to-Speech only; Speech-to-Text is billed separately by every provider here except where noted.

Provider Rate per 1M characters
OpenAI tts-1 about $15
Inworld Realtime TTS-2 $25 on demand, down to $5 at enterprise volume
Deepgram Aura-2 $30
Gradium L plan $35.90 effective
Gradium M plan $37.80 effective
Gradium S plan $47.80 effective
ElevenLabs Flash v2.5, Eleven v3 Conversational $50
Gradium XS plan $57.80 effective
ElevenLabs Eleven v3, Multilingual v2 $100

Cartesia prices in credits ($5 for 100k on Pro, $49 for 1.25M on Startup, $299 for 8M on Scale) and does not convert to this column cleanly. ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026. Gradium's rates are the effective per-character cost of each plan's credit allocation at 1 credit per character; annual billing gives twelve months for eleven.

Working through the credit maths properly: how to compare TTS pricing across providers.

What to test before you commit

  1. Your own content, not a benchmark prompt set. Coval samples 10 of 30 prompts per run. That predicts very little about your order references.
  2. The P25 to P75 spread, not just the median. A wide spread is what users notice.
  3. Connection reuse. Coval excludes the TLS and WebSocket handshake uniformly, roughly 50 to 200 ms. If your agent opens a connection per turn, add it back.
  4. Speech-to-Text billing. It changes total cost per conversation more than the Text-to-Speech rate does.
  5. Concurrency ceiling on the plan you would actually buy. Gradium allows 2, 5, 5, 10 and 15 concurrent Text-to-Speech streams from Free through L; ElevenLabs allows 4, 6, 10, 20, 30 and 30 on Flash; Cartesia 2, 3, 5 and 15.

Glossary

Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.

Leading silence. Inaudible samples at the start of a synthesized stream, measured by Coval as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.

Elo (speech arenas). A preference score from blind pairwise votes on which of two samples sounds better. Preference, not accuracy.

Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys.

References

  1. Coval TTS leaderboard: benchmarks.coval.ai/tts
  2. Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
  3. Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
  4. Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
  5. Artificial Analysis Speech Arena: artificialanalysis.ai
  6. Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608

This is the short verdict page in the Text-to-Speech selection cluster. Go deeper in whichever direction you need:

Beyond this topic

The verdict is the start of the build. Best API to build an AI voice agent covers the orchestration layer, TTS latency benchmark 2026 covers the measurement in depth, and best low-latency TTS APIs covers the whole pipeline budget rather than the synthesis stage alone.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions