NewOur fastest Text-to-Speech model yet with 1000+ voices ・ Sub-50ms latency

Tutorial

json_config in Gradium: TTS and STT Parameters

5 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • json_config is one object on the setup message. It works for both Text-to-Speech and Speech-to-Text, and it is where per-session behaviour is tuned.
  • Four Text-to-Speech fields: rewrite_rules (normalization language), temp (0.0 to 1.4, default 0.7), cfg_coef (1.0 to 4.0, default 2.0) and padding_bonus (-4.0 to 4.0, default 0.0).
  • pronunciation_id is a common mistake: it is a top-level field on the setup message, not inside json_config.
  • Codebook depth is not a json_config parameter. It is an architecture-level tradeoff Gradium tunes when it trains a model.
  • Keep json_config identical across sessions in a long render. Changing temp or padding_bonus between segments is audible even with the same voice.

What is json_config?

Gradium's Text-to-Speech model handles most contextual terms and characters without help. json_config is for the cases where "most" is not enough: a date read the wrong way round, a phone number run together, delivery that is too fast for an audiobook or too flat for a character.

It is a single JSON object attached to the setup message that opens a session. Everything in it applies for the life of that session.

A minimal setup with no configuration at all:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "model_name": "default",
  "output_format": "wav"
}

The same setup with normalization enabled:

{
  "type": "setup",
  "voice_id": "<VOICE_ID>",
  "model_name": "default",
  "output_format": "wav",
  "json_config": {
    "rewrite_rules": "en"
  }
}

That one field is the highest-value change on this page. It is what turns 12/31/2020 and foo.bar@gmail.com into something a listener parses on the first pass.

What are the Text-to-Speech parameters?

Source: docs.gradium.ai/guides/voice-settings, read September 8, 2026.

Parameter Range Default Controls
rewrite_rules en, fr, de, es, pt, none, or a comma-separated rule list none Text normalization applied before synthesis
temp 0.0 to 1.4 0.7 Variation between generations. Lower is more repeatable
cfg_coef 1.0 to 4.0 2.0 Similarity to the target voice. 2.0 to 3.0 covers most cloning work
padding_bonus -4.0 to 4.0 0.0 Delivery pace. Negative is faster, positive is slower

rewrite_rules

rewrite_rules turns on text normalization: the rewriting that happens before the model sees your text, so structured tokens arrive in a speakable form.

Pass a language code for the preset bundle:

"json_config": {
  "rewrite_rules": "en"
}

Or name individual rules for tighter control, which is what you want when one session carries more than one language:

"json_config": {
  "rewrite_rules": "TimeEn,Date,NumberEn,EmailEn"
}

Two behaviours to plan around: rules are applied word by word, and only the first matching rule applies to each word. Order the list most specific first. Full detail in text normalization edge cases.

temp

temp is sampling temperature. At 0.3 the same input produces near-identical audio every time, which is what you want for a render pipeline that may need to regenerate one segment without it standing out:

"json_config": { "temp": 0.3 }

Higher values give more variation in delivery between generations from the same voice, which suits character work and anything where repeated phrases would otherwise sound mechanical.

cfg_coef

cfg_coef is the classifier-free guidance coefficient: how tightly the output tracks the target voice.

"json_config": { "cfg_coef": 3.0 }

Higher values increase similarity and, past a point, introduce artifacts. Lower values give the model more range at the cost of drifting from the reference. For cloned voices, 2.0 to 3.0 covers most work. Try it across several voices in Gradium Studio rather than picking a number analytically: the right value depends on the source recording.

padding_bonus

padding_bonus shifts delivery pace without resampling the audio.

"json_config": { "padding_bonus": 2.0 }

Positive slows the speaker down, which suits long-form narration where clarity beats pace. Negative tightens delivery, which suits a live agent where the turn budget is the constraint. Small steps: the range is only eight wide and 2.0 is already clearly slower.

What are the Speech-to-Text parameters?

json_config works the same way on the Speech-to-Text setup message:

{
  "type": "setup",
  "model_name": "default",
  "input_format": "pcm",
  "json_config": {
    "language": "en",
    "delay_in_frames": 10
  }
}
Parameter Default Controls
language auto-detected The spoken language. Setting it explicitly improves accuracy on short or noisy audio
delay_in_frames not set How much audio the recognizer buffers before emitting. One frame is 80 ms

language is worth setting whenever you know it. Detection is reliable on long monolingual audio and much less so on a three-word answer.

delay_in_frames is the responsiveness against accuracy dial. More frames means more context before the recognizer commits, which improves the transcript and delays it. Ten frames is 800 ms of buffer. Tune it against your own audio rather than against a default, because the right value depends on utterance length.

What is not in json_config

Two things commonly assumed to be json_config fields and are not.

pronunciation_id is a top-level field on the setup message, a sibling of voice_id, not a child of json_config. Putting it in the wrong place fails silently: the session opens, synthesis works, and your dictionary is simply not applied.

import asyncio
import gradium

async def main():
    client = gradium.client.GradiumClient()
    result = await client.tts(
        setup={
            "voice_id": "<VOICE_ID>",
            "model_name": "default",
            "output_format": "wav",
            # Top level, a sibling of voice_id, not inside json_config.
            "pronunciation_id": "<DICTIONARY_ID>",
            "json_config": {
                "rewrite_rules": "en",
                "temp": 0.5,
                "cfg_coef": 2.5,
                "padding_bonus": 1.0,
            },
        },
        text="Your order 4B-2291 ships on 12/31/2026.",
    )
    with open("output.wav", "wb") as f:
        f.write(result.raw_data)

asyncio.run(main())

Codebook depth is not a customer setting. The 8, 16, 24 and 32 codebook tradeoff is an architecture-level decision Gradium makes when training and shipping a model, measured in Optimizing quality vs latency (February 11, 2026, NVIDIA RTX 4080 Super, batch 8): 160.3 ms time to first audio at 8 codebooks rising to 228.4 ms at 32, with the real-time factor falling from 7.71x to 4.43x. Interesting context, not a dial you turn.

Which values for which use case?

Use case rewrite_rules temp cfg_coef padding_bonus
Live agent reading order numbers your language code 0.5 2.0 to 2.5 -1.0
Audiobook or long-form narration your language code 0.4 2.5 1.0 to 2.0
Character or expressive work your language code 0.9 to 1.1 2.0 0.0
Repeatable render pipeline your language code 0.3 2.5 0.0
Multilingual input in one session explicit rule list 0.7 2.0 0.0

These are starting points, not settings to ship unheard. Generate the same script at two or three values and listen.

One rule that is not a starting point: in a multi-session render, json_config must be byte-identical across every session. Gradium's session limit is 3,000 seconds, so a two-hour audiobook spans at least three sessions, and a temp that differs between them is audible at the seams. See keeping a voice consistent across sessions.

Glossary

json_config. The configuration object on a Gradium setup message. Applies for the life of one session, to Text-to-Speech or Speech-to-Text.

Setup message. The first message on a Gradium WebSocket session. Carries voice_id, model_name, output_format, pronunciation_id and json_config.

Text normalization. Rewriting structured tokens (dates, numbers, emails, URLs) into a speakable form before synthesis. Controlled by rewrite_rules.

Temperature (temp). Sampling variability. Lower values make repeated generations more alike; higher values vary delivery.

Classifier-free guidance (cfg_coef). How tightly generation tracks the target voice. Higher is more similar and, past a point, more prone to artifacts.

Frame. The unit of delay_in_frames on Speech-to-Text. One frame is 80 ms of audio.

Codebook depth. How many residual quantizer levels the model generates. An architecture-level tradeoff Gradium tunes, not a request parameter.

Part of the pronunciation and accuracy cluster, hub at making Text-to-Speech pronounce numbers, dates and currencies. Its siblings:

Beyond this topic

Configuration sits inside a larger integration. Adding Text-to-Speech to an app or website covers the four integration paths, streaming TTS audio in real time covers the transport, and TTS WER benchmark 2026 covers how the accuracy these settings protect is measured.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions