New

Voice Design. Live today.

← Back to Blog

Launching Voice Design: prompt the voice your agent needs

6 min read

Developers building voice agents ask us for voices matched to the use case in front of them: a Québécoise receptionist for a Montréal dealership, a Paulista support agent for a São Paulo fintech, a narrator in his sixties with the authority of a lecture hall. Briefs outnumber any catalog.

Voice cloning inherits several constraints: each voice costs a speaker to source, consent to obtain and a licence to sign. It is not scalable.

Voice Design starts from a prompt. Write one or two sentences and complete new voices are created in a few seconds. Keep the one you want and use it with the same Text-to-Speech endpoint as any catalog voice.

Voice Design is live today in the API and the Studio. It outperforms all the public APIs for voice design that we could test. Native speakers picked Gradium voices over ElevenLabs, Inworld, MiniMax and Fish Audio in 72.6% of blind pairwise comparisons on accent prompts, 13.6 points clear of the next system and first in every language tested.

Bar chart of win rate against the field on accent prompts, with error bars. Gradium 72.6%, ElevenLabs v3 59.0%, Inworld 44.8%, Fish Audio 36.7%, MiniMax 31.7%. The 50% par line runs across the chart.

Win rate against the field on accent prompts. Blind pairwise human ratings, 7,627 comparisons, September 2026. 50% is par.

Hear it first

Four voices, each generated from a single written description. Text is the only input.

Pirate
Prompt

A gruff Bristolian English male pirate voice, 45 to 60, for game and character narration: weathered low pitch, gravelly timbre with heavy vocal fry, strong projection, at a steady, unhurried pace, boisterous and commanding energy.

0:00 / 0:00
Fionn
Prompt

An Irish English male voice, 40 to 55, for customer service: calm and crisp, with mid-low pitch, steady natural pacing, medium energy and warm rounded resonance. Ideal for reassuring walkthroughs, empathic de-escalation and complex IT support.

0:00 / 0:00
Freya
Prompt

A British female voice, 20 to 30, glossy and confident, with girly chatter, high pitch, fast pacing, high energy and bright sparkling resonance. Ideal for a friendly receptionist or assistant.

0:00 / 0:00
Desmond
Prompt

An American English male voice, 55 to 65: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.

0:00 / 0:00

Custom voices for voice agents: accent, age and character

Customer support callers bring their own accents, ages and registers, and agents land better when they sound like they come from the same place. Developers thus ask us for Austrian German, Swiss French, Mexican Spanish and many other accents, as well as specific age bands, genders, pitch, pace, energy, register and intended use.

Creators also ask for specific voices to match descriptions from games, audiobooks and in-product characters. Voice cloning is the usual answer to all of this, and it means sourcing, auditioning and licensing a real speaker for every single voice in the product, with a cost and a timeline that grow with each one you add.

How Voice Design works

Describe a voice in a sentence, in our Studio or via our API, get up to five candidates back in seconds, and keep the one you want as a permanent voice that runs on every Gradium endpoint. Full reference in the docs.

How Voice Design compares: 7,627 blind comparisons, judged by humans and by a model

We tested the match between the prompt and the voices. We ran that as a blind pairwise listening test across six voice-design systems available via public APIs and five languages: Gradium Voice Design, ElevenLabs eleven_ttv_v3 (once as default, once with "Studio quality." appended), Inworld Voice Design, MiniMax Voice Design and Fish Audio voice-design-1. Every system received the same English description.

Prompts. 50 accent descriptions across the five languages, each naming a region, country or diaspora. Example:

Paulista Brazilian woman in her thirties. Open vowels, retroflex R, relaxed and musical, mid pitch, unhurried. Warm, friendly customer-support agent, clear and professional. São Paulo Brazilian, not European Portuguese.

Audio. Every system spoke the same line, written for that description in the target language.

Raters. Native speakers of the language being rated, recruited through crowd platforms. Each rater saw the description in their own language, heard two clips from two systems, and picked the one that matched the description better, or a tie. System names, category and the shared line were hidden. Thirty comparisons per session, order randomised, pairs balanced.

Scoring. Win rate is (wins + ½ ties) / comparisons, so 50% is par.

Results

Gradium is first in all five languages, and the ranking is stable with strong performance across languages.

Ranked dot plot of win rate on accent prompts for five voice design systems, with margins. Gradium 72.6%, plus or minus 5 points. ElevenLabs 59.0%, plus or minus 9. Inworld 44.8%, plus or minus 9. Fish Audio 36.7%, plus or minus 12. MiniMax 31.7%, plus or minus 8.

Win rate on accent prompts. Blind pairwise human ratings, 7,627 comparisons, September 2026.

Here are the strongest cases, by accent:

Language Accent requested Win rate vs field
French Québécois 97%
Spanish Rioplatense (Argentina) 86%
German Bavarian 85%
Spanish Colombian 83%
Portuguese African Portuguese 83%

What the comparisons sound like

Five comparisons from the test set, one per language. Each is one English description, one line spoken in the target language, and audio clips returned by systems in the test. Raters heard them unlabelled, with no system names.

EnglishScottish Highland
Prompt

Scottish Highland woman, soft and lilting.

Spoken

Listen close. The mist is rolling in thick over the glen tonight, and I can feel the chill in my bones. It is a fine evening for a quiet cup of tea.

Gradium
ElevenLabs v3
Fish Audio
MiniMax
FrenchQuébécois
Prompt

Québécois accent, woman in her fifties. Vowels broken into diphthongs, T and D affricated before front vowels, informal and unhurried. Montreal rather than rural.

Spoken

C'est correct, relaxe. On va aller magasiner au centre-ville plus tard, mais pour l'instant, on va juste finir notre café tranquillement sur la.

Gradium
ElevenLabs v3
Fish Audio
Inworld
GermanViennese
Prompt

Austrian German, woman in her forties. Distinctly Viennese melody, lengthened vowels, softened plosives. Must not come back as standard High German.

Spoken

Schau dir das an. Der Kaffee im Kaffeehaus ist heute wieder einmal eine Frechheit, aber der Apfelstrudel ist wenigstens noch halbwegs essbar, wenn.

Gradium
ElevenLabs v3
Fish Audio
MiniMax
SpanishRioplatense
Prompt

Rioplatense accent, man, sing-song.

Spoken

Che, escuchame una cosa. ¿Viste lo que pasó en el Obelisco? No te lo puedo creer, fue una locura total, me quedé mudo mirando todo lo que pasaba.

Gradium
ElevenLabs v3
Fish Audio
Inworld
PortugueseLisbon
Prompt

European Portuguese from Lisbon, woman in her forties. Unstressed vowels reduced almost to nothing, SH-coloured S at word ends, tightly packed and quick.

Spoken

Não tenho tempo para isto. O autocarro passa já daqui a dois minutos e eu ainda não encontrei as chaves. Se não sair agora, vou chegar atrasada ao.

Gradium
ElevenLabs v3
Inworld
MiniMax

LLM as a judge

We ran the same comparisons through a model judge to measure how far it can stand in between runs.

Method. Gemini 3.1 Pro received the voice description and a single generated clip as audio, with system names removed, and assigned a rating from 1 to 5. The same prompt set was used across systems, and the model was asked to rate each clip independently using the same scoring criteria.

Note. Unlike the previous human evaluation, which used pairwise preferences to avoid differences in how annotators interpret and use rating scales, we use a fixed 1–5 scale for the LLM judge. This allows the LLM to assign an absolute score to each clip using the same scale across all systems and runs, giving us a consistent basis for comparison.

Results. Each score is the average rating over 50 accent prompts (English, French, German, Spanish and Portuguese), one clip per prompt per system, for the four systems below. A 5 means every accent detail the prompt asked for is audibly there while a 2 means the requested accent is largely absent.

Bar chart of mean Gemini 3.1 Pro score on a 1 to 5 scale for four voice design systems. Gradium 4.06, ElevenLabs 3.86, Inworld 3.64, Fish Audio 3.51.

Gemini 3.1 Pro evaluation of accent following, mean score over 50 accent prompts, September 2026.

Model Clips rated ≥ 4/5 Clips ≤ 2/5
Gradium 72% 22%
ElevenLabs 58% 24%
Inworld 52% 36%
Fish Audio 24% 38%

Agreement with human raters. The Gemini judge produced the same ranking as the human evaluation.

Try it now

Voice Design is free in the API and in Studio, so you can iterate on voice candidates until one fits.

Frequently Asked Questions