By Gradium. Data as of September 2026.
Key takeaways
- This is the alternatives cluster hub: the field ranked on voice quality against price. For the measurement view of one pairing, see Gradium vs ElevenLabs.
- On the Artificial Analysis provider-voice board read September 8, 2026, Cartesia Sonic 3.6 led at 1,282 Elo of 92 models, Inworld Realtime TTS-2 was second at 1,252, and ElevenLabs Eleven v3 Conversational eighth at 1,210.
- ElevenLabs lists $50 per 1M characters for Flash v2.5 and Eleven v3 Conversational, and $100 for Eleven v3 and Multilingual v2. It ran a 50%-off-for-life API promotion through September 11, 2026.
- Elo measures blind preference, not correctness. On the August 2026 hard-case pronunciation test the order changed: Gradium TTS 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs Eleven v3 Conversational 65.4%.
- The per-model sections below are a dated May 2026 snapshot, kept as a record. Check the live board before deciding.
ElevenLabs is one of the most well-known names in Text-to-Speech, but it is far from the only option worth considering in 2026. Depending on your use case, quality requirements, and budget, several alternatives now match or exceed ElevenLabs models on independent benchmarks.
This article ranks the best ElevenLabs alternatives using two data sources: the Artificial Analysis ELO Speech Arena (a continuously updated leaderboard based on human preference votes) and verified pricing data per million characters. All ELO scores and pricing figures referenced below come directly from the Artificial Analysis leaderboard as of May 2026.
Why Look for ElevenLabs Alternatives?
Source: Artificial Analysis Speech Arena, provider-voice board, fetched September 8, 2026. Elo from blind pairwise votes on English audio. Values drift a few points per day and ranks move by one or two, so this is a dated reading rather than a standing result.
| Model | Elo | Rank of 92 | Price per 1M characters |
|---|---|---|---|
| Cartesia Sonic 3.6 | 1,282 | 1 | credit-priced |
| Inworld Realtime TTS-2 | 1,252 | 2 | $25 down to $5 |
| Alibaba Qwen-Audio-3.0-TTS-Plus | 1,241 | 3 | see vendor pricing |
| ElevenLabs Eleven v3 Conversational | 1,210 | 8 | $50 |
| Google Gemini 3.1 Flash TTS | 1,208 | 9 | $36.60 |
| ElevenLabs Eleven v3 | 1,175 | 14 | $100 |
| Gradium TTS | 1,149 | 18 | $35.90 to $57.80 by plan |
ElevenLabs' strongest entry in that read, Eleven v3 Conversational, ranked eighth at 1,210 Elo, listed at $50 per million characters; Eleven v3 ranked fourteenth at 1,175 and $100 per million. That price point is justified for content creation, audiobooks, or dubbing workflows where voice quality is the primary constraint.
For teams building real-time voice agents, running high-volume workloads, or operating under budget constraints, the combination of $100/1M pricing and an architecture that was not designed for sub-200 ms streaming can be a limiting factor. Several alternatives now deliver competitive or higher ELO scores at a lower cost. For the dedicated head-to-head, see ElevenLabs Alternative: Why Developers Choose Gradium for Real-Time Voice AI.
Which Are the Best ElevenLabs Alternatives in 2026, Ranked by Voice Quality?
The per-model sections and the wide table below are a May 2026 snapshot of the Artificial Analysis Speech Arena, kept as a dated record. For the current top of the board see the table above, read September 8, 2026. Elo reflects human preference votes from blind pairwise comparisons; Artificial Analysis shows a 95% confidence interval per model. The Speech Arena evaluates English audio only, so the ELO scores below reflect opinions on English voices. Results may differ when voice cloning is used or in other languages.
Inworld Realtime TTS-2 (May 2026 entry: Realtime TTS 1.5 Max, Elo 1,208)
Inworld Realtime TTS 1.5 Max held the top position on the Artificial Analysis leaderboard in the May 2026 snapshot at 1,208 Elo. Inworld has since released Realtime TTS-2 and TTS-2 Flash, generally available September 2, 2026, and on the September 8, 2026 read Realtime TTS-2 ranked second of 92 at 1,252 Elo. TTS-2 is listed at $25 per 1M characters on demand, down to $5 at enterprise volume. Its price of $35 per million characters is significantly lower than ElevenLabs Eleven v3 ($100/1M) while scoring higher on perceived voice quality.
Inworld positions this line for real-time applications. As of September 2026, teams using ElevenLabs Eleven v3 for quality should evaluate Inworld Realtime TTS-2, which ranked second of 92 at 1,252 Elo on the September 8, 2026 read, rather than the earlier TTS 1.5 Max.
Google Gemini 3.1 Flash TTS: ELO 1,206, $36.6/1M
Google Gemini 3.1 Flash TTS ranks #2 with an ELO of 1,206. At $36.6 per million characters, it offers near-top-tier voice quality at roughly one-third of the ElevenLabs Eleven v3 price.
For teams already integrated with the Google Cloud ecosystem, this is a natural starting point for evaluation.
StepAudio 2.5 TTS: ELO 1,187
StepAudio 2.5 TTS from StepFun (a Chinese AI lab) ranks #3 with an ELO of 1,187, scoring above every ElevenLabs model on the leaderboard, including Eleven v3 (ELO 1,178). Pricing is not published per-character; see StepFun's commercial terms. Public API documentation and developer tooling are less mature than Western alternatives at this stage.
MiniMax Speech 2.8 HD: ELO 1,164, $100/1M
MiniMax Speech 2.8 HD ranks #5 with an ELO of 1,164. It matches ElevenLabs Eleven v3 in pricing at $100/1M but scores slightly lower in ELO. It is a strong option for teams looking for a direct quality-to-quality comparison in the premium tier.
Fish Audio S2 Pro: ELO 1,128, $15/1M (Open Weights)
In the May 2026 snapshot, Fish Audio S2 Pro ranked 11th with an Elo of 1,128 and a price of $15 per million characters. It is an open-weights model, which gives teams the option of self-hosting. At $15 per 1M via API, it scored above the ElevenLabs entries in that snapshot on perceived voice quality. Note that ElevenLabs has since removed Turbo v2.5 at a fraction of the cost.
Azure AI Speech HD 2.5: ELO 1,123, $22/1M
In the May 2026 snapshot, Microsoft Azure HD 2.5 ranked 12th with an Elo of 1,123. At $22/1M, it is well below ElevenLabs pricing and integrates directly with Azure infrastructure. For enterprise teams running workloads on Azure, this is a strong candidate. See the dedicated Azure TTS alternative comparison for the Gradium head-to-head.
OpenAI TTS-1: ELO 1,102, $15/1M
In the May 2026 snapshot, OpenAI tts-1 ranked 17th with an Elo of 1,102 over 7,548 evaluation samples, one of the higher sample counts on the board at that date. At $15/1M, it is among the most affordable options for teams that want a well-established provider. Voice quality sits above ElevenLabs Turbo v2.5 and Flash v2.5.
OpenAI TTS-1 HD: ELO 1,098, $30/1M
In the May 2026 snapshot, OpenAI tts-1-hd ranked 19th with an Elo of 1,098 over 3,123 evaluation samples. At $30/1M, it targets teams that want higher quality than TTS-1 but prefer to stay within the OpenAI ecosystem. Its Elo was comparable to the ElevenLabs entries in that snapshot.
Gradium TTS: ELO 1,072, from $35.9/1M
Gradium TTS ranks #24 on the Artificial Analysis leaderboard with an ELO of 1,072 and 323 evaluation samples. As a newer model on the leaderboard, its sample count is still growing and the confidence interval is wider than more established rankings.
Gradium's leaderboard position does not capture what differentiates it from other providers: it was built specifically for real-time voice agent infrastructure. On the independent Coval leaderboard, 1-day window read September 8, 2026, Gradium TTS records 214 ms median perceived time to first audio with a 31 ms P25 to P75 spread, and 0 ms leading silence, rank 1 of 27 on that metric on the September 10, 2026 read. On a 500-sentence hard-case set rated by independent native speakers in August 2026 it passed 81.0%. It is a streaming, WebSocket-first API with voice cloning from 10 seconds of audio and a Free tier available. The measurement behind these latency figures, including the June 3, 2026 switch to perceived time to first audio, is covered in How a benchmark change produced a faster TTS model.
For teams whose primary use case is conversational AI, call automation, or voice agents at scale, Gradium's latency and streaming architecture are relevant factors that do not appear in ELO rankings. See Best Text-to-Speech API for Voice Agents for the deeper voice-agent treatment.
Cartesia Sonic 3.6 (May 2026 entry: Sonic-3, Elo 1,070)
Cartesia Sonic-3 ranked 25th at 1,070 Elo in the May 2026 snapshot. Cartesia released Sonic 3.6 on August 27, 2026, and on the September 8, 2026 read it ranked first of 92 at 1,282 Elo, the top of the board. At $39/1M, it positions itself as a streaming-capable alternative to ElevenLabs. Its ELO is slightly below Gradium TTS with a higher sample count, making it a statistically more established comparison point at this tier. See the dedicated Cartesia alternative comparison.
How Do the Top ElevenLabs Alternatives Compare in One Table?
| Model | Rank (May 2026 snapshot) | Elo (May 2026) | Price (per 1M chars) | Samples |
|---|---|---|---|---|
| Inworld Realtime TTS 1.5 Max | #1 | 1,208 | $35 | 1,851 |
| Google Gemini 3.1 Flash TTS | #2 | 1,206 | $36.6 | 1,890 |
| StepAudio 2.5 TTS | #3 | 1,187 | see StepFun pricing | 1,341 |
| ElevenLabs Eleven v3 | #4 | 1,178 | $100 | 3,753 |
| MiniMax Speech 2.8 HD | #5 | 1,164 | $100 | 3,512 |
| Fish Audio S2 Pro | #11 | 1,128 | $15 | 1,115 |
| Azure AI Speech HD 2.5 | #12 | 1,123 | $22 | 1,133 |
| ElevenLabs Multilingual v2 | #15 | 1,107 | $100 | 8,371 |
| OpenAI TTS-1 | #17 | 1,102 | $15 | 7,548 |
| ElevenLabs Turbo v2.5 | #18 | 1,099 | $50 | 7,804 |
| OpenAI TTS-1 HD | #19 | 1,098 | $30 | 3,123 |
| ElevenLabs Flash v2.5 | #21 | 1,086 | $50 | 5,875 |
| Gradium TTS | #24 | 1,072 | from $35.9 (pricing) | 323 |
| Cartesia Sonic-3 | #25 | 1,070 | $39 | 2,808 |
Elo scores and pricing from the Artificial Analysis Speech Arena, snapshot May 2026, kept as a dated record. The board updates continuously and the ranks above have all moved since; the September 8, 2026 top of the board is in the table at the top of this page. Verify current rankings on artificialanalysis.ai.
How Should You Choose the Right ElevenLabs Alternative?
The right choice depends on what you are optimizing for.
If voice quality is the only constraint: on the September 8, 2026 read, Cartesia Sonic 3.6 (1,282 Elo, rank 1 of 92), Inworld Realtime TTS-2 (1,252, rank 2) and Google Gemini 3.1 Flash TTS (1,208, rank 9) all ranked above ElevenLabs Eleven v3 (1,175, rank 14), and Inworld and Google at a lower price.
If budget is the primary constraint: in the May 2026 snapshot, Fish Audio S2 Pro ($15 per 1M, 1,128 Elo) and OpenAI tts-1 ($15 per 1M, 1,102 Elo) offered competitive quality in the lowest price tier, both above the ElevenLabs entries of that date. OpenAI tts-1 remains about $15 per 1M as of September 2026.
If you are building real-time voice agents: ELO rankings measure voice preference in single-prompt comparisons, not production latency. Gradium TTS (214 ms median perceived time to first audio with a 31 ms P25 to P75 spread and 0 ms leading silence, Coval, September 8, 2026) and Cartesia are specifically architected for streaming workloads. Standard ELO scores do not capture latency, interruption handling, or WebSocket streaming behavior.
If you need open weights or self-hosting: in the May 2026 snapshot, Fish Audio S2 Pro (rank 11, open weights, $15 per 1M) had the highest Elo among open-weights models in the top 15.
If you are on Azure infrastructure: Azure AI Speech HD 2.5 integrates natively and, in the May 2026 snapshot, ranked 12th at 1,123 Elo and $22 per 1M.
About the Artificial Analysis ELO Methodology
All ELO scores referenced in this article come from the Artificial Analysis ELO Speech Arena. The leaderboard uses pairwise human preference comparisons: evaluators listen to two anonymous audio samples and select the one that sounds more natural. ELO scores update continuously as new votes are collected.
Models with fewer than 500 evaluation samples carry a wider 95% confidence interval and should be interpreted with more caution than models with 3,000 or more samples. For the most current rankings, refer directly to the Artificial Analysis leaderboard.
Glossary
Elo (speech arenas). A preference score from blind pairwise votes on which of two samples sounds better. Measures preference, not accuracy, and the two frequently disagree.
Provider-voice board. An Artificial Analysis board where each provider picks its own voice, so voice casting is part of what is being judged.
Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely.
Open weights. Model weights published for download, so the model can be self-hosted rather than only called through an API.
Effective rate per 1M characters. A plan's monthly price divided by the characters its credit allocation buys, the only way to compare credit plans against per-character list prices.
Related guides
This page is the hub of the provider alternatives cluster. Its siblings:
- ElevenLabs alternative: Gradium: the migration view for teams already on ElevenLabs.
- Cartesia alternative: Gradium: the migration view for teams already on Cartesia.
- Deepgram alternative: Gradium: the migration view for teams already on Deepgram.
- Azure TTS alternative: Gradium: the migration view for teams on Azure AI Speech.
- Gradium vs ElevenLabs for voice agents: the measurement view rather than the migration view.
- How to compare TTS pricing across providers: the credit maths behind every price in this cluster.
Beyond this topic
A migration is a measurement problem before it is an engineering one. TTS latency benchmark 2026 and TTS WER benchmark 2026 cover how the numbers here are produced, and how to choose a TTS API covers the criteria to rank before you switch.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

