By Gradium. Data as of September 2026.
Key takeaways
- This page is written for creators and individual users generating voiceovers, narration and character voices. If you are choosing an API to build software on, start at best Text-to-Speech APIs 2026 instead.
- For creator work the metric that matters is blind listener preference. On the Artificial Analysis provider-voice board read September 8, 2026, Cartesia Sonic 3.6 led at 1,282 Elo of 92 models, with Inworld Realtime TTS-2 second at 1,252 and ElevenLabs Eleven v3 Conversational eighth at 1,210.
- Latency is close to irrelevant here. A 200 ms difference in time to first audio changes nothing when you are exporting a file.
- Cloning your own voice takes 10 seconds of audio on an instant tier and 30 minutes or more on a professional tier. Check the commercial-use terms of the free tier before you publish anything.
- Free tiers as of September 2026: Gradium gives 45,000 credits a month with no credit card, roughly one hour of audio, with 5 instant clones restricted to non-commercial use.
What actually matters when you are generating voice, not building with it
The benchmarks that dominate Text-to-Speech coverage in 2026 are built for engineers shipping live agents. Time to first audio, latency spread, streaming transport: none of it changes your work if you are writing a script, generating a take and dropping the file into an edit.
Five things do.
How the voice sounds to a listener. The only measurement that captures this is blind pairwise preference, where people hear two samples and pick one without knowing the source. That is what an Elo board reports.
Whether it reads your script correctly. Names, brand terms, numbers and acronyms are where generated audio goes wrong in a way an audience notices. A retake costs you time; a missed error costs you credibility.
Whether you can use your own voice. Cloning tiers differ in sample length, output quality and, critically, in what the licence lets you publish.
What a month actually costs at your volume. Per-character pricing is opaque until you convert it: roughly 750 characters is a minute of speech, so an hour of finished audio is about 45,000 characters.
Language and accent coverage, if you publish in more than one.
Which voices do listeners prefer?
Source: Artificial Analysis Speech Arena, provider-voice board, fetched September 8, 2026. Elo derived from blind pairwise votes on English audio. Higher is better. Values drift a few points per day and ranks move by one or two, so treat this as a dated reading.
| Rank of 92 | Model | Elo |
|---|---|---|
| 1 | Cartesia Sonic 3.6 | 1,282 |
| 2 | Inworld Realtime TTS-2 | 1,252 |
| 3 | Alibaba Qwen-Audio-3.0-TTS-Plus | 1,241 |
| 8 | ElevenLabs Eleven v3 Conversational | 1,210 |
| 9 | Google Gemini 3.1 Flash TTS | 1,208 |
| 14 | ElevenLabs Eleven v3 | 1,175 |
| 18 | Gradium TTS | 1,149 |
Two cautions on reading this table. Elo is measured on English audio only, so it says nothing about how a model sounds in French or Portuguese. And it measures preference on a short sample, not consistency across a forty-minute narration, which is a different property entirely.
For per-language preference, the controlled-voice boards are the ones to check. On September 10, 2026 Gradium TTS held rank 6 of 23 in French at 1,209 Elo, rank 7 of 23 in Portuguese at 1,240 and rank 8 of 24 in German at 1,152, on boards led by Cartesia models.
Does it read your script correctly?
Preference and accuracy are separate measurements, and in 2026 they disagree.
Source: Gradium, "Gradium TTS: latency and accuracy", August 31, 2026, data collected August 28, 2026. 500 sentences across five languages, ten criteria covering spelling, acronyms, alphanumeric tokens, dates, numbers and emails, rated by independent native speakers. A sentence passes only if a rater hears every element pronounced correctly and completely.
| Model | Hard-case pass rate |
|---|---|
| Gradium TTS | 81.0% |
| Cartesia Sonic 3.6 | 75.1% |
| ElevenLabs Eleven v3 Conversational | 65.4% |
| Inworld Realtime TTS-2 | 61.5% |
| Fish Audio S2.1 Pro | 49.5% |
The model that listeners preferred most is not the model that got the most sentences right. If your script is plain prose, weight the Elo table. If it carries product names, prices, dates, model numbers or URLs, weight this one, because every failure here is a retake.
Even at the top of this table, roughly one sentence in five needed a fix in that August 2026 test. Plan on proofing the audio, whichever tool you pick.
Practical fixes that work across tools: spell tricky names phonetically in the script, break long numbers into groups, and use a pronunciation dictionary where the tool offers one. See fixing mispronounced names and acronyms and pronunciation dictionaries.
Can you use your own voice?
Two tiers exist almost everywhere, and the difference is sample length against fidelity.
Instant cloning takes seconds of audio and returns a usable voice immediately. Gradium's Instant Voice Clone needs 10 seconds. Inworld's takes 5 to 15 seconds. ElevenLabs and Cartesia both offer an instant tier on paid plans.
Professional cloning trains on much more audio and holds up better across long passages. Gradium's Pro Voice Clone needs a minimum of 30 minutes of clean audio, with 2 hours recommended, and is available from the M plan up at a one-time training fee of 1M credits plus 1.2 credits per character. Inworld's professional tier starts at 30 minutes.
On quality, the comparison Gradium published on January 22, 2026 ran 3,220 blinded voice pairs across English, French, Spanish and German, 890 sentences and 20 voices per language, all from a 10-second source sample. Gradium held the highest live Elo in every language in that test, against ElevenLabs Flash. The architecture behind it, cross-attention layers that attend directly to the speaker recording, is described in Why cloned voices sound fake, which also explains why micro-traits like vocal fry, rasp and breathiness survive the clone.
Check the licence before you publish. Gradium's free tier includes 5 instant clones restricted to non-commercial use; paid plans allow 1,000 per month. Every provider draws this line somewhere, and it is the detail most likely to catch out a creator who prototyped on a free account.
If you want an original voice rather than a copy of one, Gradium's Voice Designer generates a voice from a text description at studio.gradium.ai/voices/design. More on the tradeoff in instant vs pro voice cloning and regional accents in voice cloning.
What does a month cost?
Source: vendor pricing pages, read September 8, 2026. About 45,000 characters is one hour of finished audio at roughly 750 characters per minute.
| Provider | Per 1M characters | Roughly per hour of audio |
|---|---|---|
| OpenAI tts-1 | about $15 | about $0.68 |
| Inworld Realtime TTS-2 | $25 on demand | about $1.13 |
| Deepgram Aura-2 | $30 | about $1.35 |
| Gradium L plan | $35.90 effective | about $1.62 |
| Gradium M plan | $37.80 effective | about $1.70 |
| Gradium S plan | $47.80 effective | about $2.15 |
| ElevenLabs Flash v2.5, Eleven v3 Conversational | $50 | about $2.25 |
| Gradium XS plan | $57.80 effective | about $2.60 |
| ElevenLabs Eleven v3, Multilingual v2 | $100 | about $4.50 |
Cartesia prices in credits ($5 for 100k on Pro, $49 for 1.25M on Startup), which does not convert to this column cleanly. ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026.
For creator volumes the per-hour column is the one to read, and the differences are small in absolute terms. What matters more is the entry point: Gradium's free tier is 45,000 credits a month with no credit card, which is about one hour of audio and does not expire, and Cartesia's free tier is 20,000 credits. Annual billing on Gradium gives twelve months for eleven.
Which tool for which kind of work?
| If you are making | Start with | Why, with date |
|---|---|---|
| Short-form video and social voiceover | Cartesia Sonic 3.6 or Inworld Realtime TTS-2 | Ranks 1 and 2 on Artificial Analysis, September 8, 2026 |
| Product demos, tutorials, explainers with names and numbers | Gradium TTS | 81.0% hard-case pass rate, August 2026 |
| Long-form narration and audiobooks | ElevenLabs Eleven v3 or Cartesia Sonic 3.6 | 1,175 and 1,282 Elo, September 8, 2026; test consistency across a full chapter first |
| Content in French, Portuguese or German | Check the per-language board | Gradium held ranks 6, 7 and 8 on those controlled-voice boards, September 10, 2026 |
| A cloned version of your own voice | Gradium or ElevenLabs | Instant tier from 10 seconds; check commercial-use terms |
| An original voice from a description | Gradium Voice Designer | Generates a voice from a text prompt |
| The cheapest possible bulk generation | OpenAI tts-1 | About $0.68 per hour of audio, September 2026 |
| Generation on your own machine, offline | Gradium Phonon | On-device model, roughly 100M parameters, five languages since July 15, 2026 |
Glossary
AI voice generator. A tool that turns written text into spoken audio, either through a web interface or an API. The same models usually power both.
Elo (speech arenas). A preference score derived from blind pairwise votes on which of two samples sounds better. Measures preference, not accuracy, and the two frequently disagree.
Instant voice cloning. Creating a usable synthetic voice from a few seconds of audio, with no training step. Gradium's takes 10 seconds.
Professional voice cloning. Training a higher-fidelity voice on 30 minutes or more of clean audio. Holds up better across long passages.
Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.
Voice design. Generating an original voice from a written description rather than cloning an existing speaker.
References
- Artificial Analysis Speech Arena: artificialanalysis.ai
- Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
- Gradium, "Why cloned voices sound fake", January 22, 2026: gradium.ai/blog/voice-cloning-sounds-fake
- Hard-case evaluation set, CC BY 4.0: huggingface.co/datasets/gradium/tts-eval-customer-support-202608
- Gradium pricing: gradium.ai/pricing
- Gradium Voice Designer: studio.gradium.ai/voices/design
Related guides
This is the creator-facing page in the Text-to-Speech selection cluster. Its siblings are written for people building software:
- Best Text-to-Speech APIs 2026: the broad API ranking, the cluster hub.
- How to choose a TTS API: the criteria framework for engineering decisions.
- ElevenLabs vs Cartesia vs OpenAI vs Gradium: the four-way head-to-head.
- Best TTS API 2026: the short verdict page.
- Top 3 Text-to-Speech solutions 2026: the scored three-way comparison.
- Best voice cloning APIs 2026: the cloning cluster hub.
- Best multilingual TTS APIs 2026: publishing in more than one language.
Beyond this topic
If the work grows into a product rather than a file, the questions change: adding Text-to-Speech to an app or website covers the first integration, building an audiobook agent covers long-form pipelines, and keeping a voice consistent across sessions covers the problem every serial creator hits.
Gradium's free tier includes 45,000 credits per month with no credit card. Try voices in the browser at studio.gradium.ai. For enterprise evaluations, use our contact form.

