Launching Voice Design: prompt the voice your agent needs
Developers building voice agents ask us for voices matched to the use case in front of them: a Québécoise receptionist for a Montréal dealership, a Paulista support agent for a São Paulo fintech, a narrator in his sixties with the authority of a lecture hall. Briefs outnumber any catalog.
Voice cloning inherits several constraints: each voice costs a speaker to source, consent to obtain and a licence to sign. It is not scalable.
Voice Design starts from a prompt. Write one or two sentences and complete new voices are created in a few seconds. Keep the one you want and use it with the same Text-to-Speech endpoint as any catalog voice.
Voice Design is live today in the API and the Studio. It outperforms all the public APIs for voice design that we could test. Native speakers picked Gradium voices over ElevenLabs, Inworld, MiniMax and Fish Audio in 72.6% of blind pairwise comparisons on accent prompts, 13.6 points clear of the next system and first in every language tested.

Win rate against the field on accent prompts. Blind pairwise human ratings, 7,627 comparisons, September 2026. 50% is par.
Hear it first
Four voices, each generated from a single written description. Text is the only input.
A gruff Bristolian English male pirate voice, 45 to 60, for game and character narration: weathered low pitch, gravelly timbre with heavy vocal fry, strong projection, at a steady, unhurried pace, boisterous and commanding energy.
An Irish English male voice, 40 to 55, for customer service: calm and crisp, with mid-low pitch, steady natural pacing, medium energy and warm rounded resonance. Ideal for reassuring walkthroughs, empathic de-escalation and complex IT support.
A British female voice, 20 to 30, glossy and confident, with girly chatter, high pitch, fast pacing, high energy and bright sparkling resonance. Ideal for a friendly receptionist or assistant.
An American English male voice, 55 to 65: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.
Custom voices for voice agents: accent, age and character
Customer support callers bring their own accents, ages and registers, and agents land better when they sound like they come from the same place. Developers thus ask us for Austrian German, Swiss French, Mexican Spanish and many other accents, as well as specific age bands, genders, pitch, pace, energy, register and intended use.
Creators also ask for specific voices to match descriptions from games, audiobooks and in-product characters. Voice cloning is the usual answer to all of this, and it means sourcing, auditioning and licensing a real speaker for every single voice in the product, with a cost and a timeline that grow with each one you add.
How Voice Design works
Describe a voice in a sentence, in our Studio or via our API, get up to five candidates back in seconds, and keep the one you want as a permanent voice that runs on every Gradium endpoint. Full reference in the docs.
How Voice Design compares: 7,627 blind comparisons, judged by humans and by a model
We tested the match between the prompt and the voices. We ran that as a blind pairwise listening test across six voice-design systems available via public APIs and five languages: Gradium Voice Design, ElevenLabs eleven_ttv_v3 (once as default, once with "Studio quality." appended), Inworld Voice Design, MiniMax Voice Design and Fish Audio voice-design-1. Every system received the same English description.
Prompts. 50 accent descriptions across the five languages, each naming a region, country or diaspora. Example:
Paulista Brazilian woman in her thirties. Open vowels, retroflex R, relaxed and musical, mid pitch, unhurried. Warm, friendly customer-support agent, clear and professional. São Paulo Brazilian, not European Portuguese.
Audio. Every system spoke the same line, written for that description in the target language.
Raters. Native speakers of the language being rated, recruited through crowd platforms. Each rater saw the description in their own language, heard two clips from two systems, and picked the one that matched the description better, or a tie. System names, category and the shared line were hidden. Thirty comparisons per session, order randomised, pairs balanced.
Scoring. Win rate is (wins + ½ ties) / comparisons, so 50% is par.
Results
Gradium is first in all five languages, and the ranking is stable with strong performance across languages.

Win rate on accent prompts. Blind pairwise human ratings, 7,627 comparisons, September 2026.
Here are the strongest cases, by accent:
| Language | Accent requested | Win rate vs field |
|---|---|---|
| French | Québécois | 97% |
| Spanish | Rioplatense (Argentina) | 86% |
| German | Bavarian | 85% |
| Spanish | Colombian | 83% |
| Portuguese | African Portuguese | 83% |
What the comparisons sound like
Five comparisons from the test set, one per language. Each is one English description, one line spoken in the target language, and audio clips returned by systems in the test. Raters heard them unlabelled, with no system names.
Scottish Highland woman, soft and lilting.
“Listen close. The mist is rolling in thick over the glen tonight, and I can feel the chill in my bones. It is a fine evening for a quiet cup of tea.”
Québécois accent, woman in her fifties. Vowels broken into diphthongs, T and D affricated before front vowels, informal and unhurried. Montreal rather than rural.
“C'est correct, relaxe. On va aller magasiner au centre-ville plus tard, mais pour l'instant, on va juste finir notre café tranquillement sur la.”
Austrian German, woman in her forties. Distinctly Viennese melody, lengthened vowels, softened plosives. Must not come back as standard High German.
“Schau dir das an. Der Kaffee im Kaffeehaus ist heute wieder einmal eine Frechheit, aber der Apfelstrudel ist wenigstens noch halbwegs essbar, wenn.”
Rioplatense accent, man, sing-song.
“Che, escuchame una cosa. ¿Viste lo que pasó en el Obelisco? No te lo puedo creer, fue una locura total, me quedé mudo mirando todo lo que pasaba.”
European Portuguese from Lisbon, woman in her forties. Unstressed vowels reduced almost to nothing, SH-coloured S at word ends, tightly packed and quick.
“Não tenho tempo para isto. O autocarro passa já daqui a dois minutos e eu ainda não encontrei as chaves. Se não sair agora, vou chegar atrasada ao.”
LLM as a judge
We ran the same comparisons through a model judge to measure how far it can stand in between runs.
Method. Gemini 3.1 Pro received the voice description and a single generated clip as audio, with system names removed, and assigned a rating from 1 to 5. The same prompt set was used across systems, and the model was asked to rate each clip independently using the same scoring criteria.
Note. Unlike the previous human evaluation, which used pairwise preferences to avoid differences in how annotators interpret and use rating scales, we use a fixed 1–5 scale for the LLM judge. This allows the LLM to assign an absolute score to each clip using the same scale across all systems and runs, giving us a consistent basis for comparison.
Results. Each score is the average rating over 50 accent prompts (English, French, German, Spanish and Portuguese), one clip per prompt per system, for the four systems below. A 5 means every accent detail the prompt asked for is audibly there while a 2 means the requested accent is largely absent.

Gemini 3.1 Pro evaluation of accent following, mean score over 50 accent prompts, September 2026.
| Model | Clips rated ≥ 4/5 | Clips ≤ 2/5 |
|---|---|---|
| Gradium | 72% | 22% |
| ElevenLabs | 58% | 24% |
| Inworld | 52% | 36% |
| Fish Audio | 24% | 38% |
Agreement with human raters. The Gemini judge produced the same ranking as the human evaluation.
Try it now
Voice Design is free in the API and in Studio, so you can iterate on voice candidates until one fits.
Frequently Asked Questions
Cloning starts from a recording of a real speaker and reproduces that person. Voice Design starts from a written description and samples a new voice from text alone. Use cloning to reproduce a specific person, and Voice Design when the brief exists before the speaker does.
A designed voice is fully synthetic. No speaker was recorded, so there is no voice actor licence, no consent to renew and no royalty to track. You own the audio you generate and you can use it commercially, in production, in any market, with no per-voice licensing.
Blind pairwise ratings. 7,627 comparisons across six voice-design systems and five languages, each on a shared English description and a shared line in the target language, judged by native speakers of the rated language. Win rate is (wins + ½ ties) / comparisons, so 50% is par.
Gradium Voice Design, ElevenLabs eleven_ttv_v3 (as written, and with "Studio quality." appended per ElevenLabs' prompting guidance), Inworld Voice Design, MiniMax Voice Design and Fish Audio voice-design-1, all through their public APIs in September 2026.
English, French, German, Spanish and Portuguese, with regional accents inside each. Name the region rather than the language family: Colombian rather than Latin American, Québécois rather than French Canadian.
Yes. A kept voice_id runs on the same streaming Text-to-Speech endpoint as any catalog voice or clone, with the same latency and output formats.
New voices. The model samples rather than retrieves, so every request returns something fresh. When you find one you want, keep it, and the voice_id stays stable for the life of the voice.

