New voices: how we pick the one that wins

Choosing a voice is a casting decision. Ask customers how they make it and it comes down to gut feeling. They play a few candidates for a room of stakeholders and someone eventually says "that's the one." Roughly one candidate in a thousand has what we call the hit factor. A big catalog doesn't help either because everyone in the room hears "good" differently, and the more candidates there are, the harder that first listen gets.
We can still measure subjective preferences. Gradium's voice pipeline generates candidates on demand against a specific use case, then surfaces the one that performs best for it.
How we decide which voices to build
Every voice starts as a written target. Demand sets which targets we pick up: deals in flight, customer feedback, gaps in locale coverage. A target category names six things:
| Axis | Example |
|---|---|
| Language and locale | en-gb-x-georgie is British English from Newcastle, not "English" |
| Gender | female |
| Age band | 30 to 45 |
| Vocal traits | mid-low pitch, warm timbre, little vocal fry |
| Use case | inbound customer support, 20 to 40 second turns |
| Delivery style | steady pacing, medium energy, no upward inflection |
Use case moves the result more than any other axis. Ask US English listeners to rate voices for customer support and more effusive, higher-pitched female voices tend to win. Ask the same listeners about long-form narration and it inverts, with lower-pitched male voices and slower, more even delivery coming out on top.
These are learned cultural associations rather than acoustic facts. Decades of IVR systems, audiobook casting and broadcast convention built them. The bias also shows up in what customers ask for, which is why we treat use case as a first-class axis.
“Before removing the block, I need to confirm a payment of $97.60 at Central Pharmacy on March 30th 2026, card number ending with 1719.”
Write the target yourself: a five-step guide
For enterprise customers we run a workshop, and the first half goes to defining the target.
1. Name the listener, not the buyer. Draw three rings on a wall: the people who hear the voice in the centre, the people who shape the experience in the middle, decision-makers on the outside. Populate the centre first and be specific. "New hire on day one of onboarding" is a target. "Employee" is not. The pitfall is designing for whoever signs off.
2. Force specificity on the persona. "35-year-old professional" gives you nothing. "38-year-old logistics manager who doesn't have time to listen twice and got burned by a confusing AI system last month" gives you pacing, energy and register. Context gives you personality: what they are doing while they listen, and what they are already annoyed by.
3. Write the "not us" list. Have the room agree on the adjectives that are wrong for the brand, and keep that list visible for the whole exercise. Rejected adjectives do more work than chosen ones, because they stop a group converging on a voice that is safe and generic.
4. Convert adjectives into acoustic terms, in the target locale. "Friendly" is an adjective, and it doesn't survive translation. Pitch, pacing, energy and resonance are instructions. Do this step with a native speaker, for the reason in the next section.
5. Write the scripts you will judge on. You will pick a voice by listening, so samples have to be real. Use language your brand actually uses, include varied vocabulary and pronunciation cases, keep it under 200 words, and write separate scripts per language rather than translating verbatim.
You leave with a written target category and a script set.
An American English male voice, 55 to 65, for research paper podcasts: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for engaging an expert audience, projecting academic authority and narration.
“Imagine a problem so stubborn that it went unsolved for decades, and then, one afternoon, a graduate student notices something nobody else had ever thought to question.”
An Irish English male voice, 40 to 55, for customer service: calm and crisp, with mid-low pitch, steady natural pacing, medium energy and warm rounded resonance. Ideal for reassuring walkthroughs, empathic de-escalation and complex IT support.
“I completely understand your concern, and I want you to know we'll take care of this properly, so just relax and I'll walk you through every step from here.”
How we select and rank flagship voices
What "flagship" means
"Flagship" is the set we recommend first, drawing on feedback from linguists and human evaluators. When a customer opens the studio and needs a voice for German customer support, the flagship voices are the ones to try before anything else, and we rank them within their category. A voice becomes flagship by beating the voices already holding that slot on a specific, measured goal.
A brief narrows the target without picking the voice. The hit factor means most candidates are fine and one is right, so a category needs hundreds of candidates, and we generate them all rather than recording them.
Generating voice at scale gives us studio-quality speech from a text prompt alone, which a recording studio can't do. One target yields hundreds of candidates, so we can select on measured listener preference rather than a producer's taste. And a generated voice doesn't belong to anyone, so there is no likeness, no release and no consent risk.
The four narrowing stages
Selection is the actual work, and it runs in four narrowing stages:
- Generate a deliberately diverse candidate pool, hundreds of voices per category, then cluster it so the survivors differ from one another rather than being near-copies.
- Shortlist about a dozen candidates by listener vote. Crowd raters hear every candidate read the same in-language scripts and keep or discard each one (the "keeper test").
- Rank head-to-head against the voices already flagship in that category, using an ELO benchmark, so candidates compete directly with the incumbents.
- Validate with native speakers, who confirm accent authenticity and flag pronunciation problems.
Every keeper test and ELO round follows the same controls: normalize for loudness, randomize the order, cap raters at 40 comparisons and 2 sessions in a row, and require a 5-minute break to limit fatigue.
The final bar is statistical. A candidate earns the slot only if it lands in the top 5 of the ELO and its lower 95% confidence bound clears the median ELO of the existing flagship voices in its category. We cut statistical ties. A voice that is merely as good as what we already recommend would only add noise to a list whose job is to be short.
A British female voice, 20 to 30, glossy and confident, with girly chatter, high pitch, fast pacing, high energy and bright sparkling resonance. Ideal for a friendly receptionist or assistant.
“Okay babes, gather round, because I have found the most gorgeous little cafe tucked down a side street in Brighton and honestly you are going to obsess over it, I promise!”
A German female voice, 20 to 30, warm and lively like a kind big sister, with mid pitch, gentle unhurried pacing, medium energy and soft rounded resonance. Ideal for greeting customers on the support line.
“Schön, dass du da bist! Komm rein, zieh die Schuhe aus und mach es dir gemütlich, ich hab dir schon einen Tee aufgesetzt.”
EN · Lovely that you're here! Come in, take off your shoes and make yourself comfortable, I've already put on some tea for you.
Why requirements don't travel between locales
A prompt is locale-specific
Early on we treated generation prompts as roughly portable across languages. Describe the voice you want, swap the language label, translate the scripts, generate. So we took a prompt that produced good English customer-support voices, warm and enthusiastic with high energy and upward inflection, and reused it for French.
French raters rejected most of what came back. One French female category took around 400 candidates to yield a usable shortlist. They are harsher across the board (a 0.59 keep-rate baseline against 0.69 in English), so the fair comparison is each descriptor against its own language. On that basis the words that saturate the English recipe, warm, enthusiastic and lively, sat above baseline in English and below it in French.
In English it climbs to 70% of winners; in French raters filter it out at the top.
The control: rewarded and rising in both languages, so it isn’t “energy” that fails.
Climbs to half of English winners; generated in French too, but French selection dropped it.
The same descriptor resolves to a different acoustic target per locale. French listeners don't want a lower-pitched voice as a rule, and they haven't rejected high energy wholesale: energetic and bright travel intact. "Warm" in a French customer-service context reads as calmer and more grounded, where the same word in US English pulls toward brightness and high energy.
Brooklyn
A warm, effusive young American voice with girly charm and contagious laughter, greeting customers and making everyone feel at home.
Noémie
A direct, expressive Parisian voice with a quick, playful pace and mid-level pitch with delicate high notes, a care specialist who stays bright and witty.
Accents live in specific places inside a sentence
An accent surfaces in particular vowels, consonant clusters, sentence positions and prosodic contours rather than spreading evenly across speech. Two things decide whether an evaluation catches it: what the voice reads, and who listens to it.
The sentences. If your test sentences don't exercise the positions where an accent lives, your evaluation can't tell an authentic accent from a shallow one. A voice passes a clip that never asked it to produce the distinguishing sound.
The listeners. Automatic labelling can't reliably separate an accent from its nearest well-covered neighbour. Accent strength is the harder part, because many samples carry only a trace of the accent and no automatic metric grades a trace well. So we gate every accent category on native-speaker validation. Bavarian ships today because native reviewers signed it off, and it holds up for mainstream customer-service delivery well before it holds up for strong rural dialect.
The sentences below land on the vowels and clusters where Irish and British English separate, with Aoife and Freya reading identical text.
“She said she can keep some cash in a big bag, and we can see if his friends can come and pick us up when we finish.”
Where the catalog stands, and what's next
Model quality is table stakes now, so the outcome turns on picking the right voice for the right locale and use case.
| Language | Voices | % Female | Flagship | Regional variants (flagship) |
|---|---|---|---|---|
| English | 140+ | 49% | 20 | US, British, Irish |
| French | 80+ | 48% | 12 | Metropolitan, Québécois |
| German | 60+ | 55% | 12 | High German, Austrian, Bavarian |
| Portuguese | 40+ | 53% | 12 | Brazilian, European |
| Spanish | 40+ | 48% | 13 | Castilian, Mexican, Colombian |
| Total | 360+ | 50% | 69 | 13 regional variants |
Two things are next. We are opening access to the voice design model itself, the same model behind every flagship voice above, as a limited beta. And we are shipping new locales and use cases on a regular cycle, with better voice discovery in the studio so the ranking reaches you as a recommendation rather than a list you have to know about.
Working at scale, or need a custom or brand voice with requirements we haven't covered here? Talk To Our Team.
Related posts
Gradium TTS, upgraded: more accurate Text-To-Speech
Gradium TTS now runs on a new model: more natural prosody and substantially more accurate pronunciation on the cases that break voice agents in production, including spelling, acronyms, emails, phone numbers, and codes. It wins head-to-head against our previous model in all five languages, and leads real-time competitors (Cartesia Sonic 3.5, Inworld TTS 1.5 Max, ElevenLabs Flash v2.5 and Multilingual v2) on the hardest pronunciation cases. Available now as the new default, with custom voices carried over.
The most accurate multilingual text-to-speech, by the numbers
How we measure WER for TTS at Gradium: text normalization, jiwer alignment, results on the MiniMax Multilingual benchmark across English, French, Spanish, Portuguese and German — and why the standard metric is starting to saturate.
Why Your Voice Cloning Sounds Fake (And How to Fix It)
Discover how Gradium's instant voice cloning achieves superior speaker similarity to ElevenLabs. Benchmark results across 4 languages with 3,220 human evaluations.

