By Gradium. Data as of September 2026.
Key takeaways
- This is the voice-agent hub of the Text-to-Speech selection cluster. It weights the turn budget and pronunciation of structured content above voice preference. For the broad ranking across all use cases, see best Text-to-Speech APIs 2026.
- A turn's latency is the sum of Speech-to-Text, LLM and Text-to-Speech, so the synthesis stage is a budget line. Gradium's guidance puts it at 200 to 300 ms; human turns change hands in about 208 ms on average.
- On the Coval leaderboard read September 8, 2026, Gradium TTS recorded 214 ms median perceived time to first audio with a 31 ms P25 to P75 spread and 0 ms leading silence, rank 1 of 27 on that last metric.
- For an agent reading order numbers and codes, hard-case accuracy predicts failures better than clean-text word error rate: 81.0% for Gradium TTS in August 2026 against 65.4% for ElevenLabs Eleven v3 Conversational.
- The spread matters as much as the median. A wide P25 to P75 range makes a visible fraction of turns feel broken even when the median looks fine.
TL;DR: No provider led every criterion in September 2026. On the Coval leaderboard read September 8, 2026, Gradium TTS recorded 214 ms median perceived time to first audio with a 31 ms P25 to P75 spread and 0 ms leading silence, and ElevenLabs Flash v2.5 recorded 185 ms. On the Artificial Analysis provider-voice board the same day, Cartesia Sonic 3.6 held 1,282 Elo and Gradium 1,149. On the August 2026 hard-case pronunciation test the order changed again: Gradium 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs Eleven v3 Conversational 65.4%. Gradium is the pick when the agent reads structured content aloud and when on-premise or on-device deployment is a requirement. The five criteria that decide it for a voice agent are latency, voice quality, pronunciation robustness, stability under load, and deployment flexibility. How Coval measures perceived time to first audio, and the June 3, 2026 change that introduced the metric, is explained in How a benchmark change produced a faster TTS model.
How to choose a TTS API for voice agents
On the Coval leaderboard read September 8, 2026 (1-day window, 26 models), Gradium TTS recorded 214 ms median perceived time to first audio with a 31 ms P25 to P75 spread; in the blinded cloning study of January 2026 it held the highest Elo across four languages against ElevenLabs Flash. But latency and quality alone do not determine whether a Text-to-Speech API works for production voice agents. A TTS API that performs well in isolated tests can fail under real workloads. For voice agents handling real conversations at scale, five criteria separate a viable engine from a liability:
Latency. Time To First Audio (TTFA) determines whether a voice agent feels conversational or sluggish. In natural dialogue, the gap between turns averages 208 ms across ten languages (Stivers et al., PNAS, 2009). TTS is the last stage in a cascaded STT → LLM → TTS pipeline, so every millisecond it adds is a millisecond the user waits. The target is sub-300ms TTFA.
Voice quality. Users judge a voice agent within the first few seconds. The TTS needs to produce expressive, natural-sounding speech with accurate prosody, emotion, and pacing across different accents, speaking styles, and languages. A catalog of diverse preset voices and high-fidelity voice cloning from short audio samples are both important: preset voices cover common use cases quickly, while cloning enables branded or personalized agents.
Pronunciation robustness. Production voice agents encounter dates, email addresses, phone numbers, currency amounts, alphanumeric codes, and domain-specific terminology. A TTS that stumbles on a confirmation number or mispronounces a street name erodes trust. Reliable pronunciation across edge cases, without requiring manual preprocessing, is essential.
Stability under load. A single-request benchmark is not a production benchmark. Voice agents run hundreds or thousands of concurrent calls. The TTS must maintain consistent TTFA and audio quality at scale, with tight P95 latency and no degradation as concurrency increases.
Flexible deployment. Regulatory requirements, data sovereignty, and infrastructure constraints vary by customer. The TTS provider needs to offer cloud, private cloud, and on-premises options so the same engine works for a startup prototype and a regulated enterprise deployment.
This post evaluates Gradium against ElevenLabs, OpenAI, and Cartesia across all five criteria, using published benchmarks, blind listening tests, and production deployment data.
How to measure TTS latency correctly
Most TTS latency benchmarks report Time To First Byte (TTFB). This is misleading. The first bytes a streaming TTS API returns are container metadata (WAV headers, Ogg identification pages, MP3 ID3 tags), not playable audio. A server might respond with headers in 50ms while the first actual audio samples arrive 200ms later. A naive benchmark reports 50ms. The user experiences 250ms.
Time To First Audio (TTFA) measures the delay from request submission to the first chunk of playable audio reaching the client. This is the metric that correlates with perceived responsiveness.
TTS API performance comparison
The table below shows the March 2026 round-trip measurements across Gradium, ElevenLabs and OpenAI models. For the current independent figures, see TTS latency benchmark 2026. All measurements were taken from Paris using WebSocket APIs where available (POST for OpenAI, which does not offer WebSocket on its TTS API). Input: standardized 15-25 word sentence, same output format and sample rate across providers. 100 queries per model, first 5 discarded (warm state). Network ping: ~5ms to Gradium and ElevenLabs endpoints, ~3ms to OpenAI. Full methodology in Time to First Audio: Measuring and reducing TTS latency in voice agents.
Source: Gradium, "Time to First Audio", March 24, 2026, self-reported, measured from Paris. These are round-trip figures taken before the Coval methodology change of June 3, 2026, so they are not comparable with perceived time to first audio. The model names are the ones tested in March 2026; ElevenLabs has since removed Turbo v2.5 from its models page. For current independent figures see TTS latency benchmark 2026.
| Model | P25 | P50 | P75 | P95 | Instant voice cloning | Languages |
|---|---|---|---|---|---|---|
| Gradium | 255ms | 258ms | 263ms | 274ms | Yes (10s sample) | EN, FR, DE, ES, PT |
| ElevenLabs Turbo v2.5 | 294ms | 304ms | 311ms | 324ms | Yes | Broad |
| ElevenLabs Flash v2.5 | 317ms | 324ms | 333ms | 351ms | Yes | Broad |
| OpenAI GPT-4o Mini TTS | 400ms | 420ms | 439ms | 483ms | Limited (Custom Voices, restricted access) | Multiple |
| ElevenLabs Multilingual v2 | 690ms | 706ms | 720ms | 742ms | Yes | 29 languages |
| OpenAI TTS-1 | 722ms | 969ms | 1,232ms | 1,807ms | Limited (Custom Voices, restricted access) | Multiple |
With WebSocket multiplexing (persistent connection, no per-turn connection overhead), latency drops further.
Source: Gradium, "Time to First Audio", March 24, 2026, self-reported, measured from Paris. Round-trip figures from before the Coval methodology change of June 3, 2026. Turbo v2.5 is the ElevenLabs model tested at that date and is no longer on the ElevenLabs models page.
| Model | P25 | P50 | P75 | P95 |
|---|---|---|---|---|
| Gradium | 212ms | 214ms | 219ms | 228ms |
| ElevenLabs Turbo v2.5 | 248ms | 257ms | 263ms | 278ms |
| ElevenLabs Flash v2.5 | 271ms | 277ms | 284ms | 302ms |
| ElevenLabs Multilingual v2 | 643ms | 657ms | 672ms | 688ms |
Two things stand out. First, Gradium's P50 at 258ms sits under the 300ms conversational threshold, and its P95 at 274ms stays there too, meaning latency is typically predictable under load. With multiplexing, P50 drops to 214ms. Second, Gradium is the only provider in this group that combines sub-300ms TTFA with the highest measured voice cloning fidelity across four languages, word-level streaming, robust pronunciation, and native multilingual support without a latency penalty.
Streaming architecture built for voice agent pipelines
Gradium's TTS generates audio incrementally as text tokens arrive. In a typical voice agent pipeline, the LLM streams tokens to the TTS, and the TTS begins generating audio from the first tokens without waiting for the full sentence.
This matters in practice because:
- Audio playback starts before the LLM has finished generating its response, reducing perceived end-to-end latency.
- Word-level streaming allows precise synchronization between text and audio, which is necessary for features like real-time captions, lip sync, and barge-in detection.
- The tight spread between P25 (255ms) and P95 (274ms) in the benchmarks above indicates consistent generation times per request. For concurrent workloads, Gradium's architecture uses batched generation with CUDA graph optimization, which is designed to maintain this profile as concurrency scales.
For developers integrating Gradium into agent frameworks, the API supports WebSocket streaming and is compatible with LiveKit, Pipecat, and other major orchestration frameworks out of the box.
How Gradium compares on streaming: Gradium, ElevenLabs, and Cartesia all provide WebSocket streaming on their TTS APIs. OpenAI's TTS API uses POST-based streaming only, with no WebSocket option. Cartesia streams via WebSocket with the lowest raw latency in this group (Sonic 3 at ~90ms TTFA). Where Gradium differentiates is the combination of word-level timestamps and connection multiplexing (reusing a single WebSocket across conversation turns, saving ~50ms per turn). Word-level sync enables real-time captions, lip sync, and precise barge-in detection. Multiplexing matters in multi-turn voice agents where connection overhead adds up across dozens of exchanges.
Voice cloning: highest speaker similarity from 10 seconds of audio
Voice agents need two things from their TTS: a voice that sounds natural, and a voice that sounds like the right person. The first is about the model's ability to produce expressive, human-like speech across a wide range of accents, speaking styles, emotional registers, and pacing patterns. The second is about identity: can the system reproduce a specific speaker's characteristics from a short audio sample, and maintain that identity consistently across sessions, languages, and content types?
These two capabilities are closely related. A model that cannot handle diverse accents and speaking styles in its base generation will also produce flat, generic voice clones. The quality of cloned voices reflects the model's underlying ability to represent the full range of human speech variation, not just its ability to copy a waveform.
Gradium evaluated this through a live ELO ranking system: 3,220 blind A/B listening tests across four languages (EN, FR, DE, ES), where human evaluators judged overall voice quality, naturalness, and speaker similarity. The results:
| Language | Gradium ELO | ElevenLabs Flash 2.5 ELO | Gap |
|---|---|---|---|
| English | ~1950 | ~1880 | +70 |
| French | ~2040 | ~1780 | +260 |
| German | ~2170 | ~1790 | +380 |
| Spanish | ~2030 | ~1880 | +150 |
Gradium held the highest Elo in all four languages tested in that January 2026 study, against ElevenLabs Flash, with the gap widening outside English. Note the comparison set: this was a two-provider test on cloned voices, not a board. For blind preference on stock voices see the Artificial Analysis figures above, where Gradium TTS ranked 18th of 92 at 1,149 Elo on September 8, 2026. The January 2026 result reflects a perceptual quality advantage in non-English voice cloning. Gradium's cloning preserves micro-traits that other systems tend to smooth out: vocal fry, breathiness, pitch dynamics, accent characteristics, and speaking style. This holds across languages, meaning a voice cloned from an English sample maintains its identity when generating French or German speech.
The minimum audio requirement is 10 seconds for instant cloning. No fine-tuning step, no per-voice training job. Clone a voice via the API, and it is available for synthesis immediately across all supported languages.
For use cases that require the highest possible fidelity, Gradium offers Pro Voice Clones: a higher-tier cloning option that produces even more accurate voice reproduction, suitable for branded agents, public-facing products, or any application where the cloned voice needs to be indistinguishable from the original speaker.
Gradium also provides a catalog of pre-built voices spanning a range of ages, accents, tones, and genders, so developers can match a voice to their brand or audience without recording custom audio. For cases where the catalog does not fit, instant cloning fills the gap: supply 10 seconds of any target voice, and the system produces a clone that can be used across all languages and sessions. Between the catalog and cloning, there is no constraint on voice identity. Customer service agents, in-game NPCs, branded assistants, and accessibility tools can each use a voice that matches their context.
How Gradium compares on voice cloning: ElevenLabs, Cartesia, and Gradium all offer voice cloning from short audio samples (10-15 seconds). OpenAI has announced Custom Voices but access remains restricted to selected partners as of early 2026. The difference among available options is measurable quality. In the blind Elo study Gradium published on January 22, 2026, over 3,220 voice pairs across four languages, Gradium scored above ElevenLabs Flash by +70 in English rising to +380 in German. The gap was largest in non-English languages, where accent preservation and speaker identity are hardest. Note the comparison set: that test ran against ElevenLabs Flash only. Cartesia and OpenAI publish no independent per-language blind cloning evaluation at that scale, so no comparable figure exists for them.
Multilingual support: same latency, same quality, every language
Gradium natively supports five languages within a single model architecture: English, French, German, Spanish, and Portuguese. This is not a separate model per language. The same model handles all five while maintaining speaker identity.
Most TTS providers now support multiple languages, but language coverage and localization quality are different things. A model can generate speech in 30+ languages while still sounding flat or generic in non-English accents. For voice agents, the goal is not just generating French or German speech, but generating it in a way that sounds local: correct accent, natural prosody, and consistent speaker identity. Gradium's architecture handles all five languages at the same latency with no speed penalty for switching between them.
The ELO voice cloning benchmarks above reinforce this. Gradium's quality advantage is largest in non-English languages, which means voice agents operating across European markets get natural-sounding output in every supported language, not just English.
This is particularly relevant for:
- Voice agents serving multilingual users. A single deployment handles callers in any supported language without routing to separate TTS instances or accepting higher latency.
- Live translation workflows. Gradium has partnered with Acolad, a global language services provider, to integrate real-time voice into multilingual enterprise workflows.
- Pan-European and Latin American deployments at scale. A contact center operating across France, Germany, Spain, Portugal, and Brazil uses one API, one model, one consistent voice identity, with no per-language latency tradeoff. This simplifies infrastructure and reduces operational cost compared to maintaining separate TTS providers or model variants per region.
How Gradium compares on multilingual: ElevenLabs Flash v2.5 supports 32 languages, Cartesia covers 40+, and OpenAI handles multiple languages (Custom Voices announced but access is restricted). Raw language count is not the differentiator. What matters for localization is how well the cloned voice preserves accent, identity, and naturalness in each target language. Gradium's ELO advantage is largest in non-English languages (+260 French, +380 German, +150 Spanish), which means a voice agent localized for European markets sounds more natural and more like the original speaker. Other providers support the same languages on paper, but the perceptual quality gap widens outside English. For teams building voice agents that need to sound local in every market they serve, that gap is what determines whether users trust the voice.
Pronunciation accuracy: handling real-world voice agent inputs
Most TTS demos use clean, well-formed sentences. Production voice agents read back confirmation numbers, email addresses, dates in regional formats, currency amounts, phone numbers, URLs, and mixed alphanumeric strings. A single mispronunciation in a booking confirmation or account number erodes user trust.
Gradium handles these cases natively without requiring text preprocessing or special formatting. The API also accepts a custom pronunciation dictionary for domain-specific terms (medical terminology, brand names, product codes, internal acronyms) that a general model would not encounter in training data.
This matters because some competing providers require developers to manually format inputs for correct pronunciation (inserting spaces in phone numbers, spelling out abbreviations, reformatting dates). With Gradium, the same raw text the LLM generates can be sent directly to the TTS without an intermediate normalization step. In testing, pronunciation remains consistent across long-form generations without degradation.
How Gradium compares on pronunciation: ElevenLabs and OpenAI handle standard text well but can require preprocessing for structured data like phone numbers, alphanumeric codes, and mixed-format strings. Cartesia's pronunciation handling is less documented. Gradium handles these edge cases natively, plus offers a custom pronunciation dictionary API for domain-specific terms. For voice agents that read back confirmation numbers, addresses, and account details, this eliminates a preprocessing step and reduces the surface area for errors.
TTS API pricing: free tier to enterprise
Gradium offers tiered pricing designed to scale from prototyping through production:
| Plan | Price | TTS hours (approx.) | Concurrent requests | Voice clones |
|---|---|---|---|---|
| Free | $0/month | ~1 hour | 3 | Limited |
| XS | $13/month | ~5 hours | 5 | 1,000 |
| S | $43/month | ~20 hours | 5 | 1,000 |
| M | $340/month | ~200 hours | 10 | 1,000 |
| L | $1,615/month | ~1,000 hours | 15 | Unlimited |
| Enterprise | Custom | Custom | Unlimited | Unlimited |
Gradium prices by millions of characters processed, not by audio duration. The "TTS hours" column above is an approximation based on typical speaking rates (~150 words per minute, ~5 characters per word). Actual hours vary depending on the content: dense technical text with abbreviations and numbers yields fewer audio hours per character than conversational dialogue.
At the L tier, the effective rate is approximately $1.62/hour of synthesized speech. Pay-as-you-go overflow is available at $3.80 per 100k additional credits.
The Free tier provides enough usage for meaningful evaluation (~1 hour of TTS with 3 concurrent requests), which is more generous than most competing real-time TTS APIs at the free level.
Voice cloning is included in all paid plans. There are no separate charges for streaming, WebSocket access, or API features.
For high-volume production deployments, enterprise pricing includes custom concurrency limits, SLA guarantees, and on-premises deployment options. Contact Gradium for details.
How Gradium compares on pricing: ElevenLabs' pricing starts higher for equivalent usage tiers and charges separately for some features (voice cloning quality tiers, certain API access modes). OpenAI's TTS pricing is competitive per-character but deployment is cloud-only. Cartesia offers competitive pricing with fast latency but fewer deployment options. Gradium includes voice cloning, WebSocket streaming, and multiplexing in all paid plans with no feature gating. The Free tier (3 concurrent requests, ~1 hour of TTS) is sufficient for meaningful evaluation before committing.
Who is using Gradium in production
Voice agents and conversational AI. This is Gradium's primary design target. Production applications include customer service bots, AI receptionists, appointment booking systems, market research and survey calls, sales automation, outbound calling, and IVR systems. Wonderful, which builds AI voice agents for real-world conversational use cases, runs on Gradium's streaming TTS infrastructure. In regulated industries (healthcare, finance), on-premises deployment ensures audio data never leaves the customer's environment.
Gaming and interactive entertainment. Sub-300ms latency enables dynamic NPC dialogue that responds to player input without breaking immersion. Ten-second voice cloning means hundreds of unique character voices can be generated programmatically rather than recorded in a studio. InteractionLabs uses Gradium to bring expressive, real-time voice AI to its Ongo robot.
Live translation and interpretation. Native multilingual support with voice identity preservation across languages makes Gradium suitable for real-time speech translation and conference interpretation. Acolad, a global language services provider, has partnered with Gradium to integrate real-time voice into multilingual enterprise workflows.
Accessibility. Gradium powers Invincible Voice, an open-source assistive system helping people with ALS and speech loss communicate in real time using their own cloned voice.
Framework integrations. Gradium has close partnerships with Pipecat and LiveKit, the two most widely used open-source frameworks for building real-time voice agent pipelines. Both offer native Gradium plugins maintained in collaboration with the Gradium team. Beyond Pipecat and LiveKit, Gradium integrates with any voice agent framework that supports WebSocket or REST-based TTS. The API is framework-agnostic: if your orchestration layer can send text and receive streaming audio, Gradium works with it.
TTS deployment options: cloud to on-premises
Gradium supports five deployment models, depending on latency requirements, data constraints, and infrastructure:
- Cloud API. The fastest way to get started. Hosted API with endpoints in multiple regions. Suitable for prototyping and production workloads where data residency is not a constraint.
- Inference partner deployments. Gradium deploys its API on infrastructure partners in multiple locations worldwide, allowing colocations with existing LLM providers to minimize inter-service latency.
- Dedicated instances. Reserved compute with guaranteed capacity and deterministic latency. No shared-infrastructure variance.
- Private cloud. Self-hosted inference on your own GPU infrastructure. You manage the deployment; Gradium provides the model and support.
- On-premises. Full deployment within your environment for strict data sovereignty or regulatory constraints (healthcare, financial services). Audio data never leaves your infrastructure.
All deployment models support the same API surface, the same models, and the same voice cloning capabilities. Enterprise plans include unlimited concurrent requests, multi-region deployment, auto-scaling, 99.9% uptime SLA, and direct engineering support.
Developer tooling includes REST API, official Python and Rust SDKs, comprehensive documentation at docs.gradium.ai, and native integration with LiveKit and Pipecat.
How Gradium compares on deployment: ElevenLabs and OpenAI are cloud-only. Cartesia offers cloud plus some self-hosted options. Gradium provides five deployment models (cloud, inference partners, dedicated instances, private cloud, on-premises), all with the same API surface and model capabilities. For regulated industries (healthcare, finance, government) or teams with strict data sovereignty requirements, the range from shared cloud to full on-prem under one provider is often the deciding factor.
Overall comparison: Gradium vs the field
No single metric tells the full story. A TTS API can win on latency but lack voice cloning. It can offer broad language support but at 3x the latency. The table below summarizes how each provider performs across the five criteria that matter for production voice agents.
| Criteria | Gradium | ElevenLabs | OpenAI | Cartesia |
|---|---|---|---|---|
| Perceived TTFA (Coval, September 8, 2026) | 214ms P50, 31ms P25 to P75 spread, 0ms leading silence | 185ms P50 (Flash v2.5), 320ms (Eleven v3 Conversational) | P95 3,927ms (GPT-4o mini TTS) | 440ms P50 (Sonic 3.6), 269ms (Sonic 3.5) |
| Voice quality (cloning) | Highest Elo in all four languages tested against ElevenLabs Flash, 3,220 blind pairs, January 2026 | Instant and professional tiers, broad language list | Preset voices; no self-serve cloning listed | Instant and professional tiers, broad language list |
| Pronunciation | Native handling + custom dictionary API | Good, some preprocessing for edge cases | Standard | Good |
| Latency spread (Coval, September 8, 2026) | 31ms P25 to P75 | 25ms (Flash v2.5), 49ms (Eleven v3 Conversational) | n/a | 171ms (Sonic 3.6), 46ms (Sonic 3.5) |
| Deployment flexibility | Cloud, partners, dedicated, private cloud, on-prem | Cloud only | Cloud only | Cloud + limited self-hosted |
Note: these Gradium, ElevenLabs and OpenAI figures are from Gradium's March 2026 benchmark under identical conditions, and are round-trip measurements taken before the Coval methodology change of June 3, 2026. Cartesia's sub-90 ms figure is the vendor's own model inference time excluding network and was not tested in the same benchmark. For independent end-to-end figures across all four, see TTS latency benchmark 2026.
Every provider in this table is production-capable. The question is where each one forces a tradeoff.
Gradium vs Cartesia
Cartesia publishes a sub-90 ms figure for Sonic 3.6, which is model inference time excluding network. On the Coval end-to-end read of September 8, 2026, Sonic 3.6 recorded 440 ms median perceived time to first audio and Sonic 3.5 recorded 269 ms. The vendor figure is not comparable with the board (~90ms TTFA). For applications where latency is the only priority, Cartesia is a strong choice. Where Gradium differentiates: the highest measured voice cloning fidelity across four languages in blind evaluations, a custom pronunciation dictionary API, and five deployment models including full on-premises. Cartesia's voice cloning quality across non-English languages has not been independently benchmarked at the same scale. For a full comparison covering TTS pronunciation, semantic VAD, voice cloning, and deployment, see Cartesia alternative: why developers choose Gradium.
Gradium vs ElevenLabs
ElevenLabs is the closest competitor across all five criteria. On the Coval read of September 8, 2026, Flash v2.5 recorded 185 ms median perceived time to first audio with a 25 ms P25 to P75 spread; Gradium recorded 214 ms with a 31 ms spread and 0 ms leading silence. Its language list is substantially longer than Gradium's five. The tradeoffs run the other way on two axes. On the August 2026 hard-case pronunciation test, Eleven v3 Conversational passed 65.4% against Gradium's 81.0%. And in the January 2026 blinded cloning study, ElevenLabs Flash scored below Gradium by 70 points in English rising to 380 in German. ElevenLabs is also cloud-only, with no on-premise or on-device option. For a full comparison covering TTS, STT, voice cloning, platform neutrality, and pricing, see ElevenLabs alternative: why developers choose Gradium.
Gradium vs OpenAI
OpenAI offers the simplest integration for teams already inside its ecosystem. The tradeoffs are latency and flexibility: on the Coval read of September 8, 2026, GPT-4o mini TTS recorded a P95 of 3,927 ms and no reportable median, its standard endpoint does not stream over WebSocket, deployment is cloud only, and it publishes no cross-language voice cloning benchmark. For teams that need sub-300ms TTFA, voice cloning available today, WebSocket streaming, or deployment flexibility, Gradium fills the gaps OpenAI leaves.
Gradium's differentiator is not winning any single metric in isolation. It is the only provider where sub-300ms latency with a 19ms P25-P95 spread, the highest measured cloning fidelity across languages, robust pronunciation, and five deployment options all ship in the same API. That combination is what matters when the voice agent goes to production.
Beyond latency: related reading
This post focused on choosing a TTS API for voice agents. For deeper technical context on the topics covered here:
- Time to First Audio: Measuring and reducing TTS latency in voice agents covers the full benchmarking methodology, TTFB vs TTFA measurement, and WebSocket multiplexing optimization.
- Optimizing quality vs. latency in real-time TTS AI models explains the Delayed Streams Modeling architecture, RVQ codebook tradeoffs, and how Gradium balances audio quality with inference speed.
- Why your voice cloning sounds fake (and how to fix it) details the cross-attention cloning architecture, CFG scale tuning, and the full ELO evaluation methodology across 3,220 blind tests.
Getting started
Gradium offers a free tier for evaluation. Sign up at gradium.ai, generate an API key, and start streaming TTS in minutes. Documentation and quickstart guides are available at docs.gradium.ai.
For enterprise evaluations or technical questions, use our contact form or visit gradium.ai.
Glossary
Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.
Time to first byte (TTFB). When the first bytes arrive, which for a streaming API is usually container metadata carrying no sound. Not a useful proxy for responsiveness.
Leading silence. Inaudible samples at the start of a synthesized stream. Coval measures it as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal.
P25 to P75 spread. The gap between the 25th and 75th percentile of latency. A direct measure of consistency the median cannot show.
Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.
Turn budget. The total latency a conversational turn may consume across Speech-to-Text, the LLM and Text-to-Speech before it stops feeling conversational.
Semantic voice activity detection. End-of-turn detection from the meaning of what was said rather than silence duration.
Related guides
Eight pages cover choosing a Text-to-Speech provider, each from one angle. This cluster's hub is best Text-to-Speech APIs 2026.
- Best Text-to-Speech APIs 2026: the broad ranking across ten-plus providers, the cluster hub.
- How to choose a TTS API: the criteria framework, for use before you have a shortlist.
- ElevenLabs vs Cartesia vs OpenAI vs Gradium: the four-way head-to-head.
- Top 3 Text-to-Speech solutions 2026: the three-way comparison with a scoring table.
- Best TTS API 2026: the short verdict page, one pick per use case, refreshed monthly.
- Best speech APIs 2026: Text-to-Speech and Speech-to-Text from one vendor.
- Best AI voice generators 2026: consumer and creator tools rather than APIs.
Beyond this topic
Narrower questions have their own guides: best low-latency TTS APIs, best multilingual TTS APIs, best voice cloning APIs and how to compare TTS pricing across providers.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

