NewOur fastest Text-to-Speech model yet with 1000+ voices ・ Sub-50ms latency

1,000+ new voices, our most natural set yet

6 min read

Developers now have 1,000+ new voices that are more expressive, built for more use cases, and fluent in more accents. They cover English, French, Spanish, German and Portuguese, extend beyond customer service into narration, ads, social media and characters, and speak in 29 accents, from Scottish and Indian English to Quebecois French, Rioplatense Spanish and Swiss German.

This post explains how the new voices improve in naturalness and expressivity, which new use cases and accents are now covered, and how the voices were chosen.

1,000+

New voices

Added in this release

1.4k

In the catalog

Up from ~380

29

Accents

Covered by the new voices

4

Use cases

Customer service, narration, ads, characters

More natural voices for customer service and voice agents

Customer service and voice agents were the catalog's first use case, and this release raises the bar there first. Native listeners compared the new customer service voices with the ones we already recommended, reading the same support lines, and chose which they preferred. New voices came out ahead.

Previous best voiceNew voiceMedian
Voice
Score vs median
BeforeAfter
Tilly→Elsie
Standard British · en-gb
−27 1949+118 2094
Tilly
Elsie
Declan→Eoin
Irish · en-ie
+5 1922+151 2068
Declan
Eoin
Marcos→Nuno
Castilian Spanish · es-es
−101 1945+76 2122
Marcos
Nuno
Ximena→Itzel
Mexican Spanish · es-mx
+125 2050+256 2181
Ximena
Itzel
Maude→Rosalie
Quebecois French · fr-ca
−12 1955+126 2093
Maude
Rosalie
Augustin→Armand
Parisian French · fr-fr
±0 1993+121 2114
Augustin
Armand
Annika→Nele
Standard German · de-de
+25 2031+139 2145
Annika
Nele
Ricardo→Duarte
European Portuguese · pt-pt
+73 1868+336 2131
Ricardo
Duarte

The scores are ELO ratings from those comparisons: a voice climbs when listeners pick it over another, so they rank preference and nothing else. Bars are 95% confidence intervals, and the table shows the widest gap in each accent. The previous voices stay in the catalog, so integrations that use them keep working.

How the new voices were chosen

In August we described how we pick the voice that wins: generate many candidates against a written target, then let listeners decide. The method has been upgraded since, so it scales to thousands of voices. Automated judges score every clip for audio quality, pacing, expressiveness and accent; native listeners then compare the finalists head to head against our existing voices.

  1. 16,000+voices designed with the voice design model
  2. ~3,000shortlisted by automated judges
  3. Human listeningagainst existing catalog voices
  4. 1,000+added to the catalog
From designed voice to catalog voice. Figures are rounded.

Within each language, accent, gender and use case, up to five voices are flagship: the ones we recommend you try first, decided by those same listener comparisons. Flagship is a shortcut to a good first choice. The other voices in a category passed the same review, and a different brand, script or audience often calls for one of them.

Every voice in this release was designed from a written prompt with the voice design model, the same model you can use in the studio. If the catalog does not have the voice you need, describe it and design your own.

New use cases and accents in the voice catalog

The new voices are cast for four use cases, with a clear step up in naturalness and expressivity: 257 for customer service and voice agents, 221 for narration, 228 for ads and social media, and 329 character voices. Narration and ads now have flagship voices of their own, so a content team gets the same recommended starting point a voice agent team already had.

Customer service and voice agents257
English 105 · Spanish 64 · French 34 · German 31 · Portuguese 23
Narration221
English 72 · Spanish 62 · French 30 · German 31 · Portuguese 26
Ads and social media228
English 77 · Spanish 70 · French 27 · German 30 · Portuguese 24
Characters329
English 137 · Spanish 74 · French 44 · German 40 · Portuguese 34
New voices per use case, with the split by language. 1,000+ in total.

A use case changes what a voice has to do, and listeners judge each one on different things. A support voice has to stay clear and patient across short turns on a phone line. A narrator needs steady breath, clauses grouped by sense, and sentence endings that land. An ad read needs forward energy and crisp consonants that cut through a music bed. As we wrote in August, use case moves listener preference more than any other axis, so every new voice is cast for one from the start.

Every language grows, and Spanish grows the most.

in the catalog before new voices
English
148 → 539 +391
Spanish
40 → 310 +270
French
86 → 221 +135
German
69 → 201 +132
Portuguese
44 → 151 +107
Voices in the catalog per language, before and after this release. 387 before, 1,000+ added, 1,422 now.

29 accents across five languages

Teams building voice agents for a specific market ask for the accent their callers speak, and content teams localizing a campaign ask for the same.

English10 accents
  • Standard USen-us
  • Standard Britishen-gb
  • Australianen-au
  • Canadianen-ca
  • Scottishen-gb-x-scottish
  • Irishen-ie
  • Indian Englishen-in
  • New Zealanden-nz
  • Southern Americanen-us-x-southern
  • Californianen-us-x-western
Spanish7 accents
  • Castilian Spanishes-es
  • Mexican Spanishes-mx
  • Rioplatense Spanishes-ar
  • Chilean Spanishes-cl
  • Colombian Spanishes-co
  • Dominican Spanishes-do
  • Puerto Rican Spanishes-pr
French5 accents
  • Parisian Frenchfr-fr
  • Quebecois Frenchfr-ca
  • Southern Frenchfr-fr-x-southern
  • Moroccan Frenchfr-ma
  • Senegalese Frenchfr-sn
German4 accents
  • Standard Germande-de
  • Austrian Germande-at
  • Swiss Germande-ch
  • Bavariande-de-x-bavarian
Portuguese3 accents
  • Brazilian Portuguesept-br
  • European Portuguesept-pt
  • Brazilian Nordestinopt-br-x-nordestino

Listen to the new voices by language

Pick an accent to hear two of its voices side by side: a female and a male voice, or two different use cases. Each card plays the voice's own preview line, so no two cards say the same sentence. Lines in French, Spanish, German and Portuguese carry an English gloss underneath. The ID on each card is the voice_id you pass to the API, and Try this voice opens that voice in the studio catalog.

English voices: Standard US, Southern US, British, Scottish, Irish, Indian, Australian and New Zealand

NarrationFemale

Nadia

Voice

Low and resonant with real chest weight under every line, holding even dynamics and a settled fall at each sentence end.

Transcript

“Judging by the worn thresholds, thousands of children passed through these doors.”

3kcyRn3zOtdObedMTry this voice
NarrationMale

Warren

Voice

Low, chest-deep resonance with controlled breath and sense-grouped phrasing that holds structure across long paragraphs.

Transcript

“Glaciers retreat slowly, leaving behind ridges of gravel that mark each year of loss.”

2rvLkgYrTjtRKncjTry this voice

French voices: Parisian, Quebecois and Southern French

NarrationFemale

Ombeline

Voice

Low and chest-resonant, with even dynamics and a definite fall at the close of every sentence.

Transcript

“Chaque chapitre s'ouvre sur une question simple, puis déroule patiemment ses conséquences.”

EN · Each chapter opens on a simple question, then patiently works through its consequences.

wdpWBtcPIzpNd1voTry this voice
Customer serviceMale

Baptiste

Voice

Lively and expressive, with pitch that lifts and dips across a phrase and a playful edge to every line.

Transcript

“Votre abonnement reste actif, la coupure vient d'une simple erreur de facturation.”

EN · Your subscription is still active; the cut-off came from a simple billing error.

rqqeErpASHVQXHDdTry this voice

Spanish voices: from Castilian to Rioplatense and Caribbean Spanish

NarrationFemale

Itziar

Voice

Low and resonant, with chest weight and steady breath, it groups clauses by sense and lands the end of every sentence.

Transcript

“Cuando el río bajaba de nivel, los vecinos cruzaban por las piedras sin mojarse los pies.”

EN · When the river ran low, the neighbours crossed on the stones without getting their feet wet.

egJV4Gdu316e71AbTry this voice
NarrationMale

Gonzalo

Voice

Low and chest-resonant, steady in dynamics, with clauses grouped by sense and a real fall at the end of each sentence.

Transcript

“El viajero anotó en su cuaderno el nombre de cada arroyo, cada puente y cada campanario.”

EN · The traveller wrote down the name of every stream, every bridge and every bell tower in his notebook.

fE97ZAg6XnASr1yxTry this voice

German voices: Standard, Swiss, Austrian and Bavarian

NarrationFemale

Mareike

Voice

Low and chest-resonant with even dynamics and a definite fall at the close of each sentence, authority coming from steadiness rather than volume.

Transcript

“Was geschieht eigentlich mit einer Sprache, wenn ihre letzten Sprecher verstummen?”

EN · What actually happens to a language when its last speakers fall silent?

bNIsaleqsPU3P5FBTry this voice
Customer serviceFemale

Nele

Voice

Rounded, open tone with gentle lift at the ends of phrases and an easy, flowing rhythm.

Transcript

“Möglicherweise liegt es an Ihrem Browser, versuchen Sie es bitte in einem anderen Fenster.”

EN · It may be your browser; please try again in another window.

2A2fFf8SkDEJ8nzxTry this voice

Portuguese voices: European, Brazilian and Nordestino

NarrationFemale

Leonor

Voice

Low, chesty resonance with settled breath control, real pauses at paragraph breaks and a clean fall at every full stop.

Transcript

“Neste capítulo, vamos observar como as plantas transformam a luz em alimento.”

EN · In this chapter, we'll look at how plants turn light into food.

1FQtniM72t8vgsPRTry this voice
Customer serviceMale

Duarte

Voice

Rounded and resonant, with open vowels and a gentle lift at the end of phrases.

Transcript

“Prefere o reembolso por transferência, ou desconto o valor na próxima fatura?”

EN · Would you prefer the refund by bank transfer, or shall I take the amount off your next invoice?

YHUvqFRUUFmtoTBDTry this voice

Character voices for games, animation and stories

The catalog also gets character voices in all five languages: pixies, dragons, ogres, pirates and sorcerers. Each one sounds like a human actor playing the role, with the gravel, the pauses and the timing of a performance. Our next model iteration will add fictional and cartoon-like character voices. Below, each language has two of its best-rated character voices, each a different character type.

DragonMale

OrmagarStandard Britishen-gb

Voice

Vast chest weight and a dark, gravelled low register, with slow deliberate phrasing and long controlled pauses.

Transcript

“Three kingdoms burned before your grandmother drew breath, and I remember each of them.”

WZnkNRg670mFcD18Try this voice
PixieFemale

MabScottishen-gb-x-scottish

Voice

Sparkling and airy, with a forward smiling placement and quick, bouncing melody that darts between tiny pauses.

Transcript

“Follow the glow, quick now, before the thistles close over the path for the night.”

EXZnOC5HOBMl0lLJTry this voice

How to find the right voice in the catalog

In the studio, open the voice catalog and filter by language, accent, gender, age and use case. The filters combine, so "Irish, female, customer service" is one view. To link a single voice, add its ID to the catalog URL: https://studio.gradium.ai/voices/library?voice=<voice_id>. That is what the Try this voice buttons above do.

From code, pass the same ID as voice_id in a TTS request. The docs cover the request format and streaming.

What's next: regional accents within a language

Most of the 29 accents in this release are national or broad regional standards: the English of Scotland, the French of Quebec, the Spanish of Buenos Aires. The next step goes finer, to city and regional accents inside a country, where a caller hears the difference within a few words.

That work happens in the voice design model. Each accent it learns opens a new set of voices to design, and native listeners still decide which ones ship. Accent granularity is where we are putting the effort, because a team serving one region needs a voice from that region, and a national standard does not cover it.

Further reading:

Need an accent, a language or a brand voice we do not cover yet? Tell us what you are building.

Related posts

Frequently Asked Questions