NewOur fastest Text-to-Speech model yet ・ Sub-50ms latency

Comparison

How to Choose a TTS API: The Five Criteria

8 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • This page is the criteria framework, for use before you have a shortlist. For the four-way head-to-head, see ElevenLabs vs Cartesia vs OpenAI vs Gradium; for a one-line verdict per use case, see best TTS API 2026.
  • Five criteria decide it: perceived time to first audio, pronunciation accuracy on your content, streaming transport, voice cloning, and cost at your projected volume. Rank them before you look at a single benchmark.
  • Latency benchmarks changed definition on June 3, 2026. Coval now counts leading silence inside the stream, so any figure published before that date measures round trip only and cannot be compared with a current one.
  • Clean-text English word error rate no longer separates the leaders. In the April 2026 MiniMax test the top three sat within 0.05 points; the August 2026 hard-case test spread the same field from 49.5% to 81.0%.
  • Test on your own content. A 30-prompt public benchmark predicts almost nothing about how a model reads your order references.

What are the five criteria?

No provider wins all five. The work is deciding which one your product fails on first, then choosing for that.

1. Perceived time to first audio

Time to first audio is the delay between sending text and the first audible sample. For a live agent it is the criterion; for a narration pipeline it is close to irrelevant.

The reference point is human conversation. Across ten languages, the gap between one speaker finishing and the next beginning averages 208 ms (Stivers et al., PNAS, 2009). Gradium's published guidance puts the Text-to-Speech share of the turn budget at 200 to 300 ms.

Two things to get right before quoting any number.

First, the definition. Since June 3, 2026 the Coval leaderboard reports perceived time to first audio: round trip plus the leading silence inside the stream before the first audible sample. A model can return a chunk quickly and still make the listener wait. Across the board, leading silence accounts for 28% of time to first audio. Every figure published before that date measures round trip only. For how these Coval latency figures are measured since June 3, 2026, see How a benchmark change produced a faster TTS model.

Second, the distribution. A median tells you how a good turn feels; the P25 to P75 spread tells you how many turns feel wrong. On the Coval board read September 8, 2026 (1-day window, 26 models), medians ran from 51 ms for Fluxions vui to 440 ms for Cartesia Sonic 3.6, with ElevenLabs Flash v2.5 at 185 ms and Gradium TTS at 214 ms. Gradium's 31 ms spread was among the tightest; Deepgram Aura-2's was 231 ms on the same read.

Weight this criterion heavily if you are building a live agent, an IVR or a phone bot. Weight it lightly if you are producing files.

Detail: TTS latency benchmark 2026.

2. Pronunciation accuracy on your content

Word error rate measures how often a reference Speech-to-Text model hears something different from what you sent. It is a proxy, and which proxy you use changes the answer.

On clean English prose the leaders have converged. In the April 2026 MiniMax Multilingual TTS Test Set run, the top three systems sat inside 0.05 points on English, with Gradium at 1.11% average across five languages against ElevenLabs Flash v2.5 at 1.52% and Cartesia Sonic-3 at 1.56%.

On adversarial content they have not. In the August 2026 hard-case test over 500 sentences rated by independent native speakers, pass rates ran from 49.5% for Fish Audio S2.1 Pro to 81.0% for Gradium TTS, with Cartesia Sonic 3.6 at 75.1% and ElevenLabs Eleven v3 Conversational at 65.4%. A sentence passed only if a rater heard every element pronounced correctly and completely.

If your agent reads order references, amounts, dates, email addresses or codes, the second test is the one that predicts your failure rate. The text set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608, so you can score your shortlist on it directly.

One caution on live boards: Coval word error rate moves day to day. Gradium's read between 4.81% and 6.18% across September 6 to 10, 2026. Quote it as a range with its week, never as a fixed decimal.

Detail: TTS WER benchmark 2026.

3. Streaming transport

Genuine streaming returns audio before synthesis of the full input is complete. An API that builds a complete file first cannot reach a conversational budget regardless of model quality.

WebSocket is the transport that matters for multi-turn agents, because a persistent connection removes the per-turn handshake. That handshake is not free and it is not in the benchmarks: Coval excludes the TLS and WebSocket upgrade uniformly, roughly 50 to 200 ms. Gradium's own March 2026 measurement put the difference between opening a connection per turn and reusing one at about 44 ms at the median, 258 ms against 214 ms round trip.

Ask three questions of any candidate: does it stream over WebSocket, can one connection carry several sessions, and what is the maximum session length. Gradium allows 3,000 seconds per session on both Text-to-Speech and Speech-to-Text.

Detail: streaming TTS audio in real time and WebSocket multiplexing.

4. Voice cloning

If the product needs a branded voice, a preserved user voice or per-character voices, cloning is not optional and the tiers differ sharply.

The dimensions that matter are minimum sample length, whether the clone streams at the same latency as a stock voice, and which plan unlocks it. Gradium's Instant Voice Clone takes 10 seconds of audio, with 5 clones on the free tier for non-commercial use and 1,000 per month on paid plans; its Pro Voice Clone needs 30 minutes of clean audio minimum, 2 hours recommended, and is available from the M plan up at a one-time training fee of 1M credits plus 1.2 credits per character. ElevenLabs and Cartesia both publish instant and professional tiers on paid plans.

On quality, the comparison Gradium published on January 22, 2026 ran 3,220 blinded voice pairs across English, French, Spanish and German, 890 sentences and 20 voices per language, from a 10-second source sample. Gradium held the highest live Elo in every language in that test, against ElevenLabs Flash.

Detail: best voice cloning APIs 2026 and instant vs pro cloning.

5. Cost at your projected volume

Headline price per million characters is the wrong unit if the vendors bill differently. Convert everything to cost per production turn, or per output hour, and check whether Speech-to-Text is inside or outside the same bill.

Source: vendor pricing pages, read September 8, 2026. Text-to-Speech only.

Provider Text-to-Speech rate
OpenAI tts-1 about $15 per 1M characters
Deepgram Aura-2 $30 per 1M characters
Gradium L plan $35.90 per 1M characters effective
Gradium M plan $37.80 per 1M characters effective
Gradium S plan $47.80 per 1M characters effective
ElevenLabs Flash v2.5, Eleven v3 Conversational $50 per 1M characters
Gradium XS plan $57.80 per 1M characters effective
ElevenLabs Eleven v3, Multilingual v2 $100 per 1M characters

Cartesia prices in credits rather than characters ($5 for 100k on Pro, $299 for 8M on Scale), and Inworld prices Realtime TTS-2 from $25 per 1M on demand down to $5 at enterprise volume, so neither converts to this column cleanly.

Three costs that hide outside the table: Speech-to-Text billing (Gradium draws it from the same credit pool at 3 credits per second; ElevenLabs bills Scribe v2 separately at $0.22 per hour), concurrency ceilings by plan, and whether unused credits roll over (Gradium plan credits do not; pay-as-you-go credits do).

Detail: how to compare TTS pricing across providers.

How do the criteria map to use cases?

Real-time voice agents. Rank perceived time to first audio and hard-case accuracy first, cost third. Check the P25 to P75 spread, not just the median. Start at best Text-to-Speech API for voice agents.

Phone and IVR. Same as above plus output sample rate: telephony codecs expect 8 or 16 kHz, and a provider that only emits one high rate forces a resampling step into the path. See voice AI APIs for phone-based agents.

Audiobooks and long-form narration. Time to first audio drops out entirely. Rank preference Elo, prosody over long passages and voice consistency across sessions. On the Artificial Analysis provider-voice board read September 8, 2026, Cartesia Sonic 3.6 led at 1,282 Elo of 92 models, with Inworld Realtime TTS-2 second at 1,252 and ElevenLabs Eleven v3 Conversational eighth at 1,210. See building an audiobook agent and keeping a voice consistent across sessions.

Multilingual products. Read the per-language column, not the average, and check code-switching behaviour explicitly. Gradium covers English, French, Spanish, Portuguese and German with mid-sentence code-switching; Cartesia and ElevenLabs cover longer lists. See best multilingual TTS APIs 2026.

Mobile and offline. A cloud API is the wrong shape when the network cannot be assumed or the text cannot leave the device. Gradium Phonon is an on-device model of roughly 100M parameters covering English, French, German, Spanish and Portuguese since July 15, 2026. See cloud vs on-device and gradium.ai/on-device-tts.

Which criterion points where?

Rankings below are dated readings, not standing results. Both boards move; check them before deciding.

If your binding constraint is Start with Source and date
Lowest median perceived time to first audio Fluxions vui, then Inworld TTS Flash 2 Coval, September 8, 2026
A tight latency spread at a usable median Gradium TTS, ElevenLabs Flash v2.5 Coval, September 8, 2026
Numbers, codes and emails read correctly Gradium TTS Hard-case test, August 2026
Highest blind-preference Elo Cartesia Sonic 3.6 Artificial Analysis, September 8, 2026
A language outside EN, FR, ES, PT, DE Cartesia or ElevenLabs Vendor model pages
Lowest cost per character OpenAI tts-1 Vendor pricing, September 8, 2026
Cloning from a free tier Gradium Vendor pricing, September 8, 2026
On-device or offline Gradium Phonon Product page
Speech-to-Text and Text-to-Speech on one bill Gradium, Deepgram Vendor pricing, September 8, 2026

Glossary

Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.

Leading silence. Inaudible samples at the start of a synthesized stream, measured by Coval as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal.

P25 to P75 spread. The gap between the 25th and 75th percentile of latency, also called the interquartile range. A direct measure of consistency that the median cannot show.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.

WebSocket multiplexing. Carrying several synthesis sessions over one persistent connection, so no turn pays the handshake cost.

Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys. The unit that makes credit plans and per-character list prices comparable.

This page is the criteria framework in the Text-to-Speech selection cluster. Its siblings answer different shapes of the same question:

Beyond this topic

Choosing the API is the first decision, not the last. Best API to build an AI voice agent covers the orchestration layer around it, turn-taking and VAD covers the stage that decides when the model is even called, and adding Text-to-Speech to an app covers the integration itself.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions