Public Beta ・ Our New Gradium Text-to-Speech Model Is Live

← Back to Blog

New voices: how we pick the one that wins

9 min read
How Gradium chooses new voices. A row of classical marble busts floats through a field of coloured particles, each one tagged with a rating card reading warmth / naturalness or personality / emotion, as we audition and score candidate voices.

Choosing a voice is a casting decision. Ask customers how they make it and it comes down to gut feeling. They play a few candidates for a room of stakeholders and someone eventually says "that's the one." Roughly one candidate in a thousand has what we call the hit factor. A big catalog doesn't help either because everyone in the room hears "good" differently, and the more candidates there are, the harder that first listen gets.

We can still measure subjective preferences. Gradium's voice pipeline generates candidates on demand against a specific use case, then surfaces the one that performs best for it.

How we decide which voices to build

Every voice starts as a written target. Demand sets which targets we pick up: deals in flight, customer feedback, gaps in locale coverage. A target category names six things:

Axis Example
Language and locale en-gb-x-georgie is British English from Newcastle, not "English"
Gender female
Age band 30 to 45
Vocal traits mid-low pitch, warm timbre, little vocal fry
Use case inbound customer support, 20 to 40 second turns
Delivery style steady pacing, medium energy, no upward inflection

Use case moves the result more than any other axis. Ask US English listeners to rate voices for customer support and more effusive, higher-pitched female voices tend to win. Ask the same listeners about long-form narration and it inverts, with lower-pitched male voices and slower, more even delivery coming out on top.

These are learned cultural associations rather than acoustic facts. Decades of IVR systems, audiobook casting and broadcast convention built them. The bias also shows up in what customers ask for, which is why we treat use case as a first-class axis.

Inbound customer support

Before removing the block, I need to confirm a payment of $97.60 at Central Pharmacy on March 30th 2026, card number ending with 1719.

Russel_6Aslh2DxfmnRLmP
0:00 / 0:00
NolanShNU5dzYbQnT1B5z
0:00 / 0:00
Long-form narration

It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness.

Dickens, A Tale of Two Cities

Russel_6Aslh2DxfmnRLmP
0:00 / 0:00
NolanShNU5dzYbQnT1B5z
0:00 / 0:00

Write the target yourself: a five-step guide

For enterprise customers we run a workshop, and the first half goes to defining the target.

1. Name the listener, not the buyer. Draw three rings on a wall: the people who hear the voice in the centre, the people who shape the experience in the middle, decision-makers on the outside. Populate the centre first and be specific. "New hire on day one of onboarding" is a target. "Employee" is not. The pitfall is designing for whoever signs off.

2. Force specificity on the persona. "35-year-old professional" gives you nothing. "38-year-old logistics manager who doesn't have time to listen twice and got burned by a confusing AI system last month" gives you pacing, energy and register. Context gives you personality: what they are doing while they listen, and what they are already annoyed by.

3. Write the "not us" list. Have the room agree on the adjectives that are wrong for the brand, and keep that list visible for the whole exercise. Rejected adjectives do more work than chosen ones, because they stop a group converging on a voice that is safe and generic.

4. Convert adjectives into acoustic terms, in the target locale. "Friendly" is an adjective, and it doesn't survive translation. Pitch, pacing, energy and resonance are instructions. Do this step with a native speaker, for the reason in the next section.

5. Write the scripts you will judge on. You will pick a voice by listening, so samples have to be real. Use language your brand actually uses, include varied vocabulary and pronunciation cases, keep it under 200 words, and write separate scripts per language rather than translating verbatim.

You leave with a written target category and a script set.

Desmond

bwRhQrJel4IuvxLF
Prompt

An American English male voice, 55 to 65, for research paper podcasts: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for engaging an expert audience, projecting academic authority and narration.

0:00 / 0:00
Transcript

Imagine a problem so stubborn that it went unsolved for decades, and then, one afternoon, a graduate student notices something nobody else had ever thought to question.

Fionn

aokR6uVODdJ2FZSQ
Prompt

An Irish English male voice, 40 to 55, for customer service: calm and crisp, with mid-low pitch, steady natural pacing, medium energy and warm rounded resonance. Ideal for reassuring walkthroughs, empathic de-escalation and complex IT support.

0:00 / 0:00
Transcript

I completely understand your concern, and I want you to know we'll take care of this properly, so just relax and I'll walk you through every step from here.

How we select and rank flagship voices

What "flagship" means

"Flagship" is the set we recommend first, drawing on feedback from linguists and human evaluators. When a customer opens the studio and needs a voice for German customer support, the flagship voices are the ones to try before anything else, and we rank them within their category. A voice becomes flagship by beating the voices already holding that slot on a specific, measured goal.

A brief narrows the target without picking the voice. The hit factor means most candidates are fine and one is right, so a category needs hundreds of candidates, and we generate them all rather than recording them.

Generating voice at scale gives us studio-quality speech from a text prompt alone, which a recording studio can't do. One target yields hundreds of candidates, so we can select on measured listener preference rather than a producer's taste. And a generated voice doesn't belong to anyone, so there is no likeness, no release and no consent risk.

The four narrowing stages

Selection is the actual work, and it runs in four narrowing stages:

  1. Generate a deliberately diverse candidate pool, hundreds of voices per category, then cluster it so the survivors differ from one another rather than being near-copies.
  2. Shortlist about a dozen candidates by listener vote. Crowd raters hear every candidate read the same in-language scripts and keep or discard each one (the "keeper test").
  3. Rank head-to-head against the voices already flagship in that category, using an ELO benchmark, so candidates compete directly with the incumbents.
  4. Validate with native speakers, who confirm accent authenticity and flag pronunciation problems.

Every keeper test and ELO round follows the same controls: normalize for loudness, randomize the order, cap raters at 40 comparisons and 2 sessions in a row, and require a 5-minute break to limit fatigue.

The final bar is statistical. A candidate earns the slot only if it lands in the top 5 of the ELO and its lower 95% confidence bound clears the median ELO of the existing flagship voices in its category. We cut statistical ties. A voice that is merely as good as what we already recommend would only add noise to a list whose job is to be short.

Female · English GBELO ±95% CI
1#187Freya★ promoted2058 ±40
2Tilly2058 ±39
3#48Elodie-Rose★ promoted2058 ±40
4#193cut2052 ±40
5#92cut2037 ±40
6#40cut2031 ±42
7Maeve2020 ±42
8#1712011 ±41
9#382007 ±41
10Imogen2006 ±40
candidate flagship in catalog median 2015
Female · German DEELO ±95% CI
1#55Lorena★ promoted2191 ±41
2#107Jette★ promoted2127 ±46
3#29Femke★ promoted2099 ±42
4#143cut2064 ±47
5Annika2049 ±47
6Ronja2043 ±49
7Mila2037 ±48
8#76cut2012 ±49
9#162011 ±49
10#992010 ±52
candidate flagship in catalog median 2043
Example of ELO scores with candidates ranked head-to-head against catalog voices in English GB and three in German DE.

Freya#187

GgfEkEJtxZR7gnpy
Prompt

A British female voice, 20 to 30, glossy and confident, with girly chatter, high pitch, fast pacing, high energy and bright sparkling resonance. Ideal for a friendly receptionist or assistant.

0:00 / 0:00
Transcript

Okay babes, gather round, because I have found the most gorgeous little cafe tucked down a side street in Brighton and honestly you are going to obsess over it, I promise!

Lorena#55

aBNlTApBeOlVKa23
Prompt

A German female voice, 20 to 30, warm and lively like a kind big sister, with mid pitch, gentle unhurried pacing, medium energy and soft rounded resonance. Ideal for greeting customers on the support line.

0:00 / 0:00
Transcript

Schön, dass du da bist! Komm rein, zieh die Schuhe aus und mach es dir gemütlich, ich hab dir schon einen Tee aufgesetzt.

EN · Lovely that you're here! Come in, take off your shoes and make yourself comfortable, I've already put on some tea for you.

Why requirements don't travel between locales

A prompt is locale-specific

Early on we treated generation prompts as roughly portable across languages. Describe the voice you want, swap the language label, translate the scripts, generate. So we took a prompt that produced good English customer-support voices, warm and enthusiastic with high energy and upward inflection, and reused it for French.

French raters rejected most of what came back. One French female category took around 400 candidates to yield a usable shortlist. They are harsher across the board (a 0.59 keep-rate baseline against 0.69 in English), so the fair comparison is each descriptor against its own language. On that basis the words that saturate the English recipe, warm, enthusiastic and lively, sat above baseline in English and below it in French.

warmDoesn't travel
0204060Batch(200)Shortlist(20)Top-10ELO70%40%
US keep 0.76FR keep 0.53

In English it climbs to 70% of winners; in French raters filter it out at the top.

energeticTravels intact
0204060Batch(200)Shortlist(20)Top-10ELO20%10%
US keep 0.79FR keep 0.70

The control: rewarded and rising in both languages, so it isn’t “energy” that fails.

girlyUS-signature
0204060Batch(200)Shortlist(20)Top-10ELO50%10%
US keep 0.72FR keep 0.61

Climbs to half of English winners; generated in French too, but French selection dropped it.

English (US) French (FR)y = % of prompts at that stage
Share of prompts containing each descriptor at every stage of the funnel, batch → shortlist → top-10 by ELO.

The same descriptor resolves to a different acoustic target per locale. French listeners don't want a lower-pitched voice as a rule, and they haven't rejected high energy wholesale: energetic and bright travel intact. "Warm" in a French customer-service context reads as calmer and more grounded, where the same word in US English pulls toward brightness and high energy.

Female · English US

Brooklyn

Prompt

A warm, effusive young American voice with girly charm and contagious laughter, greeting customers and making everyone feel at home.

0:00 / 0:00
Female · French FR

Noémie

Prompt

A direct, expressive Parisian voice with a quick, playful pace and mid-level pitch with delicate high notes, a care specialist who stays bright and witty.

0:00 / 0:00

Accents live in specific places inside a sentence

An accent surfaces in particular vowels, consonant clusters, sentence positions and prosodic contours rather than spreading evenly across speech. Two things decide whether an evaluation catches it: what the voice reads, and who listens to it.

The sentences. If your test sentences don't exercise the positions where an accent lives, your evaluation can't tell an authentic accent from a shallow one. A voice passes a clip that never asked it to produce the distinguishing sound.

The listeners. Automatic labelling can't reliably separate an accent from its nearest well-covered neighbour. Accent strength is the harder part, because many samples carry only a trace of the accent and no automatic metric grades a trace well. So we gate every accent category on native-speaker validation. Bavarian ships today because native reviewers signed it off, and it holds up for mainstream customer-service delivery well before it holds up for strong rural dialect.

The sentences below land on the vowels and clusters where Irish and British English separate, with Aoife and Freya reading identical text.

Sentence 1 · Irish vs British on identical text

She said she can keep some cash in a big bag, and we can see if his friends can come and pick us up when we finish.

Aoife · IrishvimnD4UQG_36P43U
0:00 / 0:00
Freya · BritishGgfEkEJtxZR7gnpy
0:00 / 0:00
Sentence 2 · Irish vs British on identical text

The three brothers thought they'd take the boat out on Tuesday, but the water by the harbour wall was rough.

Aoife · IrishvimnD4UQG_36P43U
0:00 / 0:00
Freya · BritishGgfEkEJtxZR7gnpy
0:00 / 0:00

Where the catalog stands, and what's next

Model quality is table stakes now, so the outcome turns on picking the right voice for the right locale and use case.

Language Voices % Female Flagship Regional variants (flagship)
English 140+ 49% 20 US, British, Irish
French 80+ 48% 12 Metropolitan, Québécois
German 60+ 55% 12 High German, Austrian, Bavarian
Portuguese 40+ 53% 12 Brazilian, European
Spanish 40+ 48% 13 Castilian, Mexican, Colombian
Total 360+ 50% 69 13 regional variants

Two things are next. We are opening access to the voice design model itself, the same model behind every flagship voice above, as a limited beta. And we are shipping new locales and use cases on a regular cycle, with better voice discovery in the studio so the ranking reaches you as a recommendation rather than a list you have to know about.

Working at scale, or need a custom or brand voice with requirements we haven't covered here? Talk To Our Team.

Related posts

Frequently Asked Questions