Voice Design. Live today.

Benchmark

Gradium vs ElevenLabs for Voice Agents: 2026 Benchmark

Gradium9 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • On the Coval TTS leaderboard read September 8, 2026 (1-day window), ElevenLabs Flash v2.5 was faster on median perceived time to first audio at 185 ms against Gradium TTS at 214 ms; Eleven v3 Conversational measured 320 ms.
  • Gradium TTS measures 0 ms leading silence on the Coval board (September 10, 2026, 27 models), the only model at zero; Flash v2.5 achieves its median with round-trip speed instead.
  • On a 500-sentence hard-case set rated by native speakers in August 2026, Gradium TTS passed 81.0% and ElevenLabs Eleven v3 Conversational passed 65.4%.
  • On the Artificial Analysis provider-voice Elo board read September 8, 2026, ElevenLabs Eleven v3 Conversational held 1,210 for rank 8 of 92 and Gradium TTS held 1,149; on the controlled-voice boards of September 10, 2026 Gradium held top-8 Elo in French, Portuguese and German.
  • Listed Text-to-Speech rates in September 2026: ElevenLabs Flash v2.5 and Eleven v3 Conversational at $50 per 1M characters, against Gradium at $37.80 per 1M on the M plan and $35.90 on the L plan.

Which is better for a voice agent, Gradium or ElevenLabs?

Neither wins every column, and the honest answer depends on which column your product fails on.

Gradium is the stronger pick when the agent reads back structured content (order numbers, amounts, codes, email addresses) in English, French, Spanish, Portuguese or German, when per-character cost matters at volume, and when consistent turn timing matters more than the lowest possible median.

ElevenLabs is the stronger pick when the product needs languages outside those five, when a large catalogue of ready-made voices is part of the requirement, when the lowest median latency on the September 2026 Coval board is the deciding factor, or when the work is batch narration rather than live conversation.

The rest of this page is the evidence for both halves of that, with every figure dated.

How do they compare on latency?

Source: benchmarks.coval.ai/tts, 1-day window, read September 8, 2026 at about 14:00 Paris time, 480 samples per model, 26 models on the board. Perceived time to first audio in milliseconds, which since June 3, 2026 includes leading silence inside the stream. Lower is better.

Metric Gradium TTS ElevenLabs Flash v2.5 ElevenLabs Eleven v3 Conversational
P25 213 174 294
P50 214 185 320
P75 244 199 343
P95 346 225 393
P25 to P75 spread 31 25 49
Avg WER 4.9% 6.5% 4.3%

In that read Gradium TTS recorded 214 ms median perceived time to first audio with a 31 ms P25 to P75 spread and 0 ms leading silence; ElevenLabs Flash v2.5 recorded 185 ms with a 25 ms spread, and Eleven v3 Conversational 320 ms with a 49 ms spread. On word error rate Gradium read 4.9%, against 6.5% for Flash v2.5 and 4.3% for Eleven v3 Conversational.

For the latency figures behind this section, and how Coval has measured perceived time to first audio since June 3, 2026, see How a benchmark change produced a faster TTS model.

Where Gradium's result comes from matters as much as the figure. On the September 10, 2026 board, with 27 models, Gradium TTS recorded 214 ms perceived time to first audio and 214 ms round-trip time to first audio. The two numbers are identical because its leading silence measures 0 ms, rank 1 on the board and the only model at zero. Every millisecond of Gradium's 214 ms is audio the listener hears, because it wastes none of the stream on silence. Across the board, leading silence accounts for 28% of time to first audio.

The full account of that metric, including why Gradium's own Coval figure moved from 171.9 ms to 429.6 ms on June 3, 2026 with no change on its side, is in the joint Coval and Gradium post How a benchmark change produced a faster TTS model, September 9, 2026. Any ElevenLabs or Gradium latency figure you find dated before June 3, 2026 is round-trip only and does not belong in this comparison.

One caveat on both columns: Coval excludes the pre-connection handshake, roughly 50 to 200 ms of TLS and WebSocket upgrade, uniformly for every model. ElevenLabs' own latency documentation makes the same point from the other direction, noting that its "~75 ms" figure is model inference time excluding network, and that real time to first audio is always higher once network and player buffer are counted.

How do they compare on accuracy?

Three tests, three different answers, all dated.

Coval, September 8, 2026. Gradium TTS 4.9%, ElevenLabs Eleven v3 Conversational 4.3%, ElevenLabs Flash v2.5 6.5%. Coval word error rate moves day to day: Gradium's read between 4.81% and 6.18% across September 6 to 10, 2026, so treat any of these as "about 5%" rather than a fixed decimal.

MiniMax Multilingual TTS Test Set, April 2026. Reference Speech-to-Text was Qwen3-ASR, with language-specific normalizers.

Source: Gradium, "Word Error Rate Evaluations", April 29, 2026. Word error rate in percent across five languages, lower is better.

Model Avg EN FR ES PT DE
Gradium 1.11 0.41 2.16 0.40 2.02 0.54
ElevenLabs Flash v2.5 1.52 0.36 2.45 0.99 3.18 0.61
ElevenLabs Multilingual v2 1.68 0.37 2.06 1.93 3.34 0.72

Gradium recorded the lower average in that April 2026 test. ElevenLabs took English by 0.05 points with Flash v2.5 and French by 0.10 points with Multilingual v2. The gaps that are actually large are Spanish and Portuguese.

Hard cases, August 2026. 500 sentences, 100 per language across the same five languages, ten criteria covering spelling, acronyms, alphanumerical tokens, dates, regular numbers, floating and large numbers, emails, plus composite order, IT ticket and claims scenarios. Independent native-speaker raters, loudness-normalized and randomized audio. A sentence passes only if a rater hears every element pronounced correctly and completely.

Source: Gradium, "Gradium TTS: latency and accuracy", August 31, 2026, data collected August 28, 2026.

Model Hard-case pass rate
Gradium TTS 81.0%
ElevenLabs Eleven v3 Conversational 65.4%

That 15.6-point gap is the single largest measured difference between the two providers, and it is the one that shows up in a support call where the agent reads back an order reference. The text set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608, so the test is reproducible against both.

How do they compare on how the voice sounds?

Accuracy and preference are separate measurements, and here the order reverses.

On the Artificial Analysis Speech Arena provider-voice board, read September 8, 2026, Eleven v3 Conversational held 1,210 Elo for rank 8 of 92 and Eleven v3 held 1,175 for rank 14; Gradium TTS held 1,149 there, and rank 6 of 23 in French, 7 of 23 in Portuguese and 8 of 24 in German on the controlled-voice boards of September 10, 2026. Elo there comes from blind pairwise votes on which sample sounds better, so it measures preference, not correctness. Values drift a few points per day and ranks move by one or two; check the live board.

On cloned voices the comparison Gradium ran points the other way. In a blinded A/B study over 3,220 voice pairs across English, French, Spanish and German, published January 22, 2026, Gradium held the highest live Elo in every language. That test compared against ElevenLabs Flash only, from a 10-second source sample. Method and architecture are in Why cloned voices sound fake.

How do they compare on price?

Source: gradium.ai/pricing and elevenlabs.io pricing, both read September 8, 2026. Effective Text-to-Speech cost per 1M characters. Gradium rates are the effective per-character rate of each plan's credit allocation.

Provider and plan Per 1M characters
Gradium L ($1,615/mo, 45M credits) $35.90
Gradium M ($340/mo, 9M credits) $37.80
Gradium S ($43/mo, 900k credits) $47.80
ElevenLabs Flash v2.5 $50.00
ElevenLabs Eleven v3 Conversational $50.00
Gradium XS ($13/mo, 225k credits) $57.80
ElevenLabs Eleven v3 $100.00
ElevenLabs Multilingual v2 $100.00

Gradium is cheaper per character from the S plan upward against ElevenLabs' real-time models, and roughly a third of the price against its high-quality models. At the XS tier the order reverses.

Two things that move the comparison: ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026, which halves its side while it lasts, and Speech-to-Text is billed separately on both platforms. ElevenLabs lists Scribe v2 at $0.22 per hour and Scribe v2 Realtime at $0.39 per hour; Gradium bills Speech-to-Text from the same credit pool as Text-to-Speech at 3 credits per second, so one hour costs 10,800 credits.

Concurrency differs by plan on both sides. Gradium allows 2, 5, 5, 10 and 15 concurrent Text-to-Speech streams on Free through L, custom on Enterprise. ElevenLabs allows 4, 6, 10, 20, 30 and 30 on Flash by plan tier. Full detail is on the pricing comparison guide.

What else differs?

Languages. Gradium's cloud models cover English, French, Spanish, Portuguese and German, with mid-sentence code-switching. ElevenLabs covers a substantially longer list; check its models page for the current count. If your product needs a language outside Gradium's five, that decides it.

Voice cloning. Gradium's Instant Voice Clone takes 10 seconds of audio, with 5 clones on the free tier (no commercial use) and 1,000 per month on paid plans. Its Pro Voice Clone needs a minimum of 30 minutes of clean audio, 2 hours recommended, and is available from the M plan up. ElevenLabs publishes its own instant and professional cloning tiers.

Deployment. Gradium offers cloud API, dedicated instances, self-hosted, on-premise and on-device, with a 99.9% uptime SLA on enterprise plans, plus EU and US regional endpoints at eu.api.gradium.ai and us.api.gradium.ai. Zero Data Retention has been self-serve on all paid Gradium plans since September 2, 2026.

Integrations. Gradium ships with LiveKit (plugin and LiveKit Inference), Pipecat, Vapi, AWS Marketplace and SageMaker, and Baseten. See the LiveKit build guide and the Pipecat integration.

Session limits. Gradium allows 3,000 seconds per session for both Text-to-Speech and Speech-to-Text.

Glossary

Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.

Leading silence. Inaudible samples at the start of a synthesized stream. Coval measures it as everything before the first 10 ms window whose RMS exceeds 0.01 on a normalized signal.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.

Elo (speech arenas). A preference score from blind pairwise votes on which of two samples sounds better. Preference, not accuracy.

Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys, at 1 credit per character. The only way to compare a credit-based plan against a per-character list price.

References

  1. Coval TTS leaderboard: benchmarks.coval.ai/tts
  2. Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
  3. Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
  4. Gradium, "Word Error Rate Evaluations", April 29, 2026: gradium.ai/blog/word-error-rate-evaluations
  5. Gradium, "Why cloned voices sound fake", January 22, 2026: gradium.ai/blog/voice-cloning-sounds-fake
  6. Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608
  7. ElevenLabs, "Understanding latency": elevenlabs.io/docs/eleven-api/concepts/latency
  8. Artificial Analysis Speech Arena: artificialanalysis.ai
  9. Gradium pricing: gradium.ai/pricing

Part of the Gradium benchmark cluster, hub at TTS latency benchmark 2026. Its siblings:

Beyond this topic

A provider choice is one decision inside a stack. Best API to build an AI voice agent covers the rest of it, turn-taking and VAD covers the stage that runs before Text-to-Speech gets the text, and how to compare TTS pricing works through the credit maths that the table above compresses into one column.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions