NewOur fastest Text-to-Speech model yet with 1000+ voices ・ Sub-50ms latency

Tutorial

Pronunciation Dictionaries in Gradium TTS

4 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • A pronunciation dictionary maps written text to how it should be spoken, applied at synthesis time without altering the input string.
  • It is the right fix for names, brand terms, acronyms and domain vocabulary. Rewriting the input phonetically also works and leaks into logs, transcripts and anything downstream that reuses the text.
  • Create a dictionary in Gradium Studio, then pass its pronunciation_id on the setup message. It is a top-level field, a sibling of voice_id, not inside json_config.
  • The same mechanism doubles as a content-moderation layer: map an unwanted word to an acceptable substitute and the model speaks the substitute.
  • Pronunciation is where a voice agent loses trust fastest. On the August 2026 hard-case test, spelling and acronyms were two of ten criteria, and Gradium TTS passed 81.0% of the 500-sentence set.

What is a pronunciation dictionary?

A set of rules telling the model how to handle specific pieces of text before they are spoken. Each rule maps an original string to a replacement pronunciation or word. Once created, the dictionary has a pronunciation_id that you pass with every request that should use it.

The key property: the input text is unchanged. You still send Kyutai or SQL or your customer's surname, and the substitution happens inside synthesis. That matters more than it first appears, because the text you send to a Text-to-Speech API is usually the same text you log, display as a caption, store in a transcript and feed to analytics. Phonetic respellings in that string corrupt all of them.

When do you need one?

Normalization and dictionaries solve different problems, and reaching for the wrong one wastes time.

Problem Fix
12/31/2026 read as digits rewrite_rules, see text normalization
A phone number run together rewrite_rules
A brand name stressed wrongly Pronunciation dictionary
An acronym spelled out that should be a word, or the reverse Pronunciation dictionary
A surname the model has never seen Pronunciation dictionary
A drug name, ticker or part number with a house pronunciation Pronunciation dictionary
A word that must never be spoken aloud Pronunciation dictionary, as a moderation rule

Rule of thumb: if the token has a format, normalization handles it. If it has a name, a dictionary does.

How do you create one?

In Gradium Studio:

  1. Open the Pronunciation tab.
  2. Create a dictionary and name it for its scope, for example brand-terms or clinical-vocab.
  3. Add rules. Each is an original text and a replacement pronunciation.
  4. Copy the pronunciation_id.

Write replacements the way you would sound the word out to a colleague, using hyphens to mark syllable boundaries and ordinary spelling for each piece. The model reads the replacement as text, so it works best when the replacement is unambiguous in the target language.

Original Replacement Why
xoxo ex-o-ex-o Letter sequence that would otherwise be guessed at
SQL sequel An initialism your house style speaks as a word
Kyutai kyoo-tye Proper noun with no English spelling cue
AAPL A-A-P-L Ticker that must stay letter by letter
mg milligrams Unit abbreviation, expanded rather than spelled

Keep one dictionary per domain rather than one per product. A rule that is right for a clinical script is often wrong for a marketing one, and dictionaries are selected per request, so splitting them costs nothing.

How do you use one in the API?

Pass pronunciation_id on the setup object:

import asyncio
import gradium

async def main():
    client = gradium.client.GradiumClient()

    result = await client.tts(
        setup={
            "voice_id": "<VOICE_ID>",
            "model_name": "default",
            "output_format": "wav",
            # Top level, a sibling of voice_id. Not inside json_config.
            "pronunciation_id": "<DICTIONARY_ID>",
            "json_config": {
                "rewrite_rules": "en",
            },
        },
        text="Your SQL migration ships in 250 mg increments, per Kyutai's spec.",
    )

    with open("output.wav", "wb") as f:
        f.write(result.raw_data)

asyncio.run(main())

The single most common mistake is nesting it:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "json_config": {
    "pronunciation_id": "<DICTIONARY_ID>"
  }
}

That fails silently. The session opens, synthesis succeeds, the audio comes back, and no rule is applied. If a dictionary appears to do nothing, check this first. The correct shape:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "model_name": "default",
  "output_format": "wav",
  "pronunciation_id": "<DICTIONARY_ID>",
  "json_config": {
    "rewrite_rules": "en"
  }
}

The LiveKit plugin exposes pronunciation_id as a service option, so an agent built on LiveKit sets it once at session construction rather than per request. See building a voice AI agent with Gradium and LiveKit.

Using a dictionary for content moderation

The same substitution mechanism works as an output filter. Instead of correcting a pronunciation, map an unwanted word to an acceptable one and the model speaks the substitute.

This matters on any product that reads user-generated content aloud, where the input text is not yours to control. A rule replacing a profanity with a neutral word keeps spoken output inside platform guidelines regardless of what arrives.

Two limits worth stating plainly. This is a substitution list, not a moderation system: it catches exactly the strings you enumerate, and nothing else. And it changes only the audio, so if you also display the text, filter that separately.

Testing a dictionary

  1. Generate the same script with and without pronunciation_id. If the two files are byte-identical, the field is in the wrong place.
  2. Test each term in a sentence, not alone. Prosody around a word changes how it lands, and a replacement that reads well in isolation can sound stressed wrongly in context.
  3. Test in every language you ship. A replacement spelled for English readers will be read with the wrong vowels in French or German.
  4. Test the terms next to normalization. Rules and normalizers both act before synthesis; a term that contains digits can interact with rewrite_rules.
  5. Have a native speaker listen. This is the method behind Gradium's own hard-case evaluation: a sentence passes only if a rater hears every element pronounced correctly and completely. Gradium TTS passed 81.0% of that 500-sentence set in August 2026, so even a strong model needs the check.

Dictionaries across languages

A replacement is read as text, so it is read in the language of the session. kyoo-tye produces the intended sound with English vowels and something else entirely with French ones. That has two consequences.

Keep one dictionary per language, not one global dictionary applied everywhere. The same brand name needs a differently spelled replacement per language, and a single shared rule will be right in one of them.

Watch loanwords specifically. A product name that is English inside a French sentence is exactly the case where the model is most likely to apply the wrong vowel system, and where a dictionary is most valuable. This overlaps with language detection: if whole clauses are flipping accent rather than single terms, the fix is rewrite_rules as a language lock, covered in stopping accent switches mid-sentence.

Gradium's cloud models cover English, French, Spanish, Portuguese and German, so a product shipping in all five wants five dictionaries selected per request, not one dictionary with five sets of rules in it.

Glossary

Pronunciation dictionary. A stored set of rules mapping written text to a spoken form, applied at synthesis time without altering the input string.

pronunciation_id. The setup-message field naming which dictionary to apply. Top level, a sibling of voice_id, not inside json_config.

Initialism. An acronym spoken letter by letter, as against one spoken as a word. Spelling does not distinguish the two, which is why dictionaries exist.

Text normalization. Rewriting structured tokens such as dates and numbers into a speakable form before synthesis. Handles formats; dictionaries handle names.

Grapheme-to-phoneme. Converting written characters into speech sounds. Where proper nouns and loanwords break, because the rules are language-specific and the word may not be.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely. All-or-nothing per sentence.

Part of the pronunciation and accuracy cluster, hub at making Text-to-Speech pronounce numbers, dates and currencies. Its siblings:

Beyond this topic

Names and codes are what a phone agent reads most, and a caller has no screen to catch a mistake: voice AI APIs for phone-based agents covers that case. Best Text-to-Speech API for voice agents covers how to weight pronunciation against latency when choosing a provider.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions