By Gradium. Data as of September 2026.
Key takeaways
json_configis one object on thesetupmessage. It works for both Text-to-Speech and Speech-to-Text, and it is where per-session behaviour is tuned.- Four Text-to-Speech fields:
rewrite_rules(normalization language),temp(0.0 to 1.4, default 0.7),cfg_coef(1.0 to 4.0, default 2.0) andpadding_bonus(-4.0 to 4.0, default 0.0). pronunciation_idis a common mistake: it is a top-level field on the setup message, not insidejson_config.- Codebook depth is not a
json_configparameter. It is an architecture-level tradeoff Gradium tunes when it trains a model. - Keep
json_configidentical across sessions in a long render. Changingtemporpadding_bonusbetween segments is audible even with the same voice.
What is json_config?
Gradium's Text-to-Speech model handles most contextual terms and characters without help. json_config is for the cases where "most" is not enough: a date read the wrong way round, a phone number run together, delivery that is too fast for an audiobook or too flat for a character.
It is a single JSON object attached to the setup message that opens a session. Everything in it applies for the life of that session.
A minimal setup with no configuration at all:
{
"type": "setup",
"voice_id": "<VOICE_ID>",
"model_name": "default",
"output_format": "wav"
}
The same setup with normalization enabled:
{
"type": "setup",
"voice_id": "<VOICE_ID>",
"model_name": "default",
"output_format": "wav",
"json_config": {
"rewrite_rules": "en"
}
}
That one field is the highest-value change on this page. It is what turns 12/31/2020 and foo.bar@gmail.com into something a listener parses on the first pass.
What are the Text-to-Speech parameters?
Source: docs.gradium.ai/guides/voice-settings, read September 8, 2026.
| Parameter | Range | Default | Controls |
|---|---|---|---|
rewrite_rules |
en, fr, de, es, pt, none, or a comma-separated rule list |
none | Text normalization applied before synthesis |
temp |
0.0 to 1.4 | 0.7 | Variation between generations. Lower is more repeatable |
cfg_coef |
1.0 to 4.0 | 2.0 | Similarity to the target voice. 2.0 to 3.0 covers most cloning work |
padding_bonus |
-4.0 to 4.0 | 0.0 | Delivery pace. Negative is faster, positive is slower |
rewrite_rules
rewrite_rules turns on text normalization: the rewriting that happens before the model sees your text, so structured tokens arrive in a speakable form.
Pass a language code for the preset bundle:
"json_config": {
"rewrite_rules": "en"
}
Or name individual rules for tighter control, which is what you want when one session carries more than one language:
"json_config": {
"rewrite_rules": "TimeEn,Date,NumberEn,EmailEn"
}
Two behaviours to plan around: rules are applied word by word, and only the first matching rule applies to each word. Order the list most specific first. Full detail in text normalization edge cases.
temp
temp is sampling temperature. At 0.3 the same input produces near-identical audio every time, which is what you want for a render pipeline that may need to regenerate one segment without it standing out:
"json_config": { "temp": 0.3 }
Higher values give more variation in delivery between generations from the same voice, which suits character work and anything where repeated phrases would otherwise sound mechanical.
cfg_coef
cfg_coef is the classifier-free guidance coefficient: how tightly the output tracks the target voice.
"json_config": { "cfg_coef": 3.0 }
Higher values increase similarity and, past a point, introduce artifacts. Lower values give the model more range at the cost of drifting from the reference. For cloned voices, 2.0 to 3.0 covers most work. Try it across several voices in Gradium Studio rather than picking a number analytically: the right value depends on the source recording.
padding_bonus
padding_bonus shifts delivery pace without resampling the audio.
"json_config": { "padding_bonus": 2.0 }
Positive slows the speaker down, which suits long-form narration where clarity beats pace. Negative tightens delivery, which suits a live agent where the turn budget is the constraint. Small steps: the range is only eight wide and 2.0 is already clearly slower.
What are the Speech-to-Text parameters?
json_config works the same way on the Speech-to-Text setup message:
{
"type": "setup",
"model_name": "default",
"input_format": "pcm",
"json_config": {
"language": "en",
"delay_in_frames": 10
}
}
| Parameter | Default | Controls |
|---|---|---|
language |
auto-detected | The spoken language. Setting it explicitly improves accuracy on short or noisy audio |
delay_in_frames |
not set | How much audio the recognizer buffers before emitting. One frame is 80 ms |
language is worth setting whenever you know it. Detection is reliable on long monolingual audio and much less so on a three-word answer.
delay_in_frames is the responsiveness against accuracy dial. More frames means more context before the recognizer commits, which improves the transcript and delays it. Ten frames is 800 ms of buffer. Tune it against your own audio rather than against a default, because the right value depends on utterance length.
What is not in json_config
Two things commonly assumed to be json_config fields and are not.
pronunciation_id is a top-level field on the setup message, a sibling of voice_id, not a child of json_config. Putting it in the wrong place fails silently: the session opens, synthesis works, and your dictionary is simply not applied.
import asyncio
import gradium
async def main():
client = gradium.client.GradiumClient()
result = await client.tts(
setup={
"voice_id": "<VOICE_ID>",
"model_name": "default",
"output_format": "wav",
# Top level, a sibling of voice_id, not inside json_config.
"pronunciation_id": "<DICTIONARY_ID>",
"json_config": {
"rewrite_rules": "en",
"temp": 0.5,
"cfg_coef": 2.5,
"padding_bonus": 1.0,
},
},
text="Your order 4B-2291 ships on 12/31/2026.",
)
with open("output.wav", "wb") as f:
f.write(result.raw_data)
asyncio.run(main())
Codebook depth is not a customer setting. The 8, 16, 24 and 32 codebook tradeoff is an architecture-level decision Gradium makes when training and shipping a model, measured in Optimizing quality vs latency (February 11, 2026, NVIDIA RTX 4080 Super, batch 8): 160.3 ms time to first audio at 8 codebooks rising to 228.4 ms at 32, with the real-time factor falling from 7.71x to 4.43x. Interesting context, not a dial you turn.
Which values for which use case?
| Use case | rewrite_rules |
temp |
cfg_coef |
padding_bonus |
|---|---|---|---|---|
| Live agent reading order numbers | your language code | 0.5 | 2.0 to 2.5 | -1.0 |
| Audiobook or long-form narration | your language code | 0.4 | 2.5 | 1.0 to 2.0 |
| Character or expressive work | your language code | 0.9 to 1.1 | 2.0 | 0.0 |
| Repeatable render pipeline | your language code | 0.3 | 2.5 | 0.0 |
| Multilingual input in one session | explicit rule list | 0.7 | 2.0 | 0.0 |
These are starting points, not settings to ship unheard. Generate the same script at two or three values and listen.
One rule that is not a starting point: in a multi-session render, json_config must be byte-identical across every session. Gradium's session limit is 3,000 seconds, so a two-hour audiobook spans at least three sessions, and a temp that differs between them is audible at the seams. See keeping a voice consistent across sessions.
Glossary
json_config. The configuration object on a Gradium setup message. Applies for the life of one session, to Text-to-Speech or Speech-to-Text.
Setup message. The first message on a Gradium WebSocket session. Carries voice_id, model_name, output_format, pronunciation_id and json_config.
Text normalization. Rewriting structured tokens (dates, numbers, emails, URLs) into a speakable form before synthesis. Controlled by rewrite_rules.
Temperature (temp). Sampling variability. Lower values make repeated generations more alike; higher values vary delivery.
Classifier-free guidance (cfg_coef). How tightly generation tracks the target voice. Higher is more similar and, past a point, more prone to artifacts.
Frame. The unit of delay_in_frames on Speech-to-Text. One frame is 80 ms of audio.
Codebook depth. How many residual quantizer levels the model generates. An architecture-level tradeoff Gradium tunes, not a request parameter.
Related guides
Part of the pronunciation and accuracy cluster, hub at making Text-to-Speech pronounce numbers, dates and currencies. Its siblings:
- Making Text-to-Speech pronounce numbers, dates and currencies: the cluster hub, what
rewrite_rulesis for. - Text normalization edge cases: every normalizer, with before and after.
- Pronunciation dictionaries in Gradium TTS: where
pronunciation_idcomes from. - Fixing mispronounced names, acronyms and technical terms: when normalization is not enough.
- Stopping accent switches mid-sentence:
rewrite_rulesas a language lock. - Keeping a voice consistent across sessions: why these values must not drift between sessions.
- WebSocket multiplexing on Gradium TTS: the other setup-message fields, for connection reuse.
Beyond this topic
Configuration sits inside a larger integration. Adding Text-to-Speech to an app or website covers the four integration paths, streaming TTS audio in real time covers the transport, and TTS WER benchmark 2026 covers how the accuracy these settings protect is measured.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

