NewOur fastest Text-to-Speech model yet with 1000+ voices ・ Sub-50ms latency

Tutorial

Text Normalization for TTS Edge Cases

4 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • Text normalization rewrites structured tokens into a speakable form before the model sees them, so dates, numbers, emails and codes are not left to guesswork.
  • One field turns it on: rewrite_rules inside json_config. Pass a language code for the preset bundle, or a comma-separated rule list for control.
  • Rules are applied word by word and only the first matching rule applies to each word, so ordering a custom list most-specific-first is not optional.
  • Normalization handles formats. Names, brand terms and acronyms need a pronunciation dictionary instead.
  • Structured tokens are where production Text-to-Speech fails visibly. On the August 2026 hard-case test, seven of the ten criteria were exactly these token types.

What is text normalization?

Models are good at phrasing and prosody. They are weaker on tokens that are written in one form and spoken in another, because the mapping is conventional rather than phonetic. 12/31/2020 is a date in the United States and a nonsense date almost everywhere else. 2500000 is nine characters that are spoken as five words. AB12CD34 is a code that has to be chunked to be heard at all.

Normalization resolves those before synthesis. Instead of asking the model to interpret 12/31/2020, it receives something unambiguous, and spends its capacity on delivery rather than on decoding.

The token types it covers: dates, times, numbers, emails, URLs, phone numbers and alphanumeric codes.

How do you turn it on?

Add rewrite_rules to json_config on the setup message:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "model_name": "default",
  "output_format": "wav",
  "json_config": {
    "rewrite_rules": "en"
  }
}

en is a language alias: a preset bundle of the common rules for English. Accepted values are en, fr, de, es, pt and none. Use an alias when a session's input is mostly one language, which is the usual case.

For tighter control, name the rules:

{
  "json_config": {
    "rewrite_rules": "TimeEn,Date,NumberEn,EmailEn"
  }
}

Two behaviours govern how a list is evaluated:

  • Rules are applied word by word.
  • Only the first matching rule applies to each word.

So order matters, and the correct order is most specific first. A general number rule placed ahead of a currency rule will consume the amount and the currency rule will never fire.

What does each normalizer change?

Source: Gradium documentation, read September 8, 2026. Spoken forms shown as the normalized text handed to the model.

Type Input Spoken as
Date 12/31/2020 12-31 2020
Time 3:45PM! 3.45PM!
Number 2500000 2 million 500 thousand
Email foo.bar@gmail.com foo dot bar at gmail dot com
URL any URL spelled out character by character, with language-appropriate handling
Phone number any phone number grouped according to the country convention
Alphanumeric code AB12CD34 A-B 1-2 C-D 3-4

The alphanumeric row is the one that most often decides whether a support agent works. A confirmation code read as an undifferentiated run of characters is unusable on a phone call, where the listener has no screen and no scrollback.

A worked example

Take a line a support agent actually reads:

Order AB12CD34 for $1,249.50 ships 12/31/2026. Confirmation to foo.bar@gmail.com.

Without normalization the model has to infer four separate conventions in one sentence: how to chunk a code, how to read a currency amount with cents, which of two date orders applies, and how to speak an address containing a dot and an at sign.

With "rewrite_rules": "en" each of those arrives resolved. The model's remaining job is prosody: where to breathe, which clause to stress, how to close the sentence. That is what it is good at.

import asyncio
import gradium

async def main():
    client = gradium.client.GradiumClient()
    result = await client.tts(
        setup={
            "voice_id": "<VOICE_ID>",
            "model_name": "default",
            "output_format": "wav",
            "json_config": {"rewrite_rules": "en"},
        },
        text=(
            "Order AB12CD34 for $1,249.50 ships 12/31/2026. "
            "Confirmation to foo.bar@gmail.com."
        ),
    )
    with open("output.wav", "wb") as f:
        f.write(result.raw_data)

asyncio.run(main())

Alias or explicit rules?

Use a language alias when the session's input is mostly one language and you want sensible defaults. This is the right answer for most products, and it stays correct as Gradium adds rules to the bundle.

Use an explicit rule list in three situations: input mixes languages inside one session, one normalizer is doing something you do not want, or you need a specific evaluation order. The cost is that you now own the list, and a rule added to the preset bundle later will not reach you.

Use none when your pipeline already normalizes upstream. Running two normalizers over the same text is a reliable way to produce doubly-expanded output.

Ordering an explicit rule list

Because only the first matching rule fires per word, a custom list is an ordered pipeline rather than a set. Two orderings of the same rules produce different audio.

Take a line containing both a plain quantity and a currency amount. With a general number rule first, it consumes the digits and the currency rule never sees them, so the amount is spoken as a bare number and the unit is lost. Put the more specific rule first and each token is handled by the rule written for it.

The working order is specificity, descending:

  1. Rules matching a compound shape: email, URL, phone number, alphanumeric code
  2. Rules matching a formatted value: currency, date, time
  3. Rules matching a bare value: number, ordinal

Two practical consequences. Adding a rule to an existing list means deciding where it goes, not appending it. And a list is a snapshot: rules Gradium later adds to the language bundle will not reach a session that names its rules explicitly, so revisit an explicit list when you upgrade.

What normalization does not fix

Formats, yes. Names, no.

Normalization operates on token shape. It has no way to know that your company spells SQL as a word, that a surname takes stress on the second syllable, or that a ticker must stay letter by letter. Those need a pronunciation dictionary, which maps specific strings to specific spoken forms.

The two compose. A typical production setup runs rewrite_rules for formats and a dictionary for vocabulary, both on the same setup message:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "model_name": "default",
  "output_format": "wav",
  "pronunciation_id": "<DICTIONARY_ID>",
  "json_config": {
    "rewrite_rules": "en"
  }
}

Note that pronunciation_id sits at the top level and rewrite_rules inside json_config. Getting that wrong fails silently.

How much does this actually matter?

Enough to be the main axis on which production models differ. Gradium built a 500-sentence evaluation set for exactly these cases: 100 sentences per language across English, French, Spanish, Portuguese and German, scored by independent native speakers, where a sentence passes only if every element is heard correctly and completely.

Seven of its ten criteria are token types normalization touches: spelling, acronyms, alphanumerical tokens, dates, regular numbers, floating and large numbers, and emails. The remaining three are composite scenarios (orders, IT tickets, claims) built from the same material.

In the August 2026 run, Gradium TTS passed 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs Eleven v3 Conversational 65.4%, Inworld Realtime TTS-2 61.5% and Fish Audio S2.1 Pro 49.5%. The set is public under CC BY 4.0 at huggingface.co/datasets/gradium/tts-eval-customer-support-202608, so you can score your own configuration on it.

Note the ceiling: roughly one sentence in five still failed for the strongest model in that test. Normalization moves the number a long way and does not finish the job. Proof your own content.

Glossary

Text normalization. Rewriting structured tokens into a speakable form before synthesis, so the model receives unambiguous input.

rewrite_rules. The json_config field that controls normalization. Takes a language alias, a comma-separated rule list, or none.

Language alias. A preset bundle of normalizers for one language. en, fr, de, es, pt.

Normalizer. One rule handling one token type, for example Date or EmailEn. Composable into an explicit list.

Alphanumeric code. A mixed letters-and-digits token such as an order or ticket reference. Needs chunking to be heard correctly, which is what the normalizer does.

Prosody. The rhythm, pitch and stress of speech. What the model handles well once normalization has removed the guesswork.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely.

Part of the pronunciation and accuracy cluster, hub at making Text-to-Speech pronounce numbers, dates and currencies. Its siblings:

Beyond this topic

Normalization is the cheapest accuracy win in a voice pipeline, but it sits inside a bigger integration. Adding Text-to-Speech to an app or website covers the paths in, and voice AI APIs for phone-based agents covers why these tokens matter most on a call.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions