By Gradium. Data as of September 2026.
Key takeaways
- Word error rate says nothing about naturalness. A voice can be word-perfect and still flat.
- Naturalness lives in micro-traits, vocal fry, rasp, breathiness and pitch dynamics, which latency-optimized models compress away.
- Two
json_configlevers do most of the work:cfg_coef(1.0 to 4.0, default 2.0, with 2.0 to 3.0 covering most cloning) andpadding_bonus(-4.0 to 4.0, default 0.0). - In a blinded study over 3,220 voice pairs across English, French, Spanish and German published January 22, 2026, Gradium held the highest Elo in every language against ElevenLabs Flash from a 10-second sample.
- Structured content breaks the illusion fastest. Gradium TTS passed 81.0% of a 500-sentence hard-case set in August 2026.
A TTS model can have near-perfect pronunciation accuracy and still sound robotic. Gradium's WER post states the limit directly: "WER does not capture voice naturalness, prosody, or expressiveness. A TTS system can have low WER while still sounding robotic, or high WER while sounding natural in casual speech."[1]
Naturalness in a synthesized voice is a combination of micro-level vocal traits, prosodic patterns, expressiveness appropriate to the context, and consistent identity across a session. Each can be addressed independently, and improving any one of them moves the output closer to a human voice.
This guide covers what makes voices sound robotic, which technical levers fix it, and how Gradium exposes those levers.
Why AI voices sound robotic
What WER does not measure
The most common TTS evaluation metric is word error rate: how often the model pronounces the input text correctly. WER captures pronunciation accuracy. It captures nothing about how the voice sounds.
A voice that says every word correctly with flat intonation, uniform pacing, and no variation in pitch or energy sounds robotic. The words are right. The voice is wrong. This is the most common failure mode for TTS models optimized for accuracy benchmarks without being evaluated on naturalness.
Naturalness also includes appropriateness. A voice that delivers a billing question with the same energy as breaking news is technically correct and contextually wrong. Emotion has to be present and has to fit the moment.
The micro-trait problem
Human voices carry information at multiple levels simultaneously. The words carry the semantic content. Underneath, constant micro-level variation carries identity, emotion, and naturalness: the slight roughness of a voice that has been talking for a while, the breathiness at the end of a phrase, the way pitch shifts between a question and a statement.
Most TTS systems optimize for the dominant signal and discard these micro-level features, because modeling them adds complexity. Models optimized for latency alone tend to compress voice identity cues. When these cues are lost, the voice sounds clean but generic, without the specific character of a real speaker.
Context mismatch
A voice that sounds natural in one context sounds wrong in another. Gradium's voice cloning post describes three registers: "Personalized news delivery requires a polished, radiophonic quality with professional gravitas and clear enunciation. Customer care applications need semi-professional voices that balance competence with approachability and warmth. Video games and entertainment require expressive, even cartoonish character voices with exaggerated emotional range and personality."[2]
A voice built for audiobook narration applied to a customer care interaction sounds over-produced. A voice built for quick conversational replies applied to long-form narration sounds thin. Naturalness is context-dependent, and a voice that misses the register of the interaction sounds off even when the technical quality is high.
The technical levers for more natural AI voices
Micro-trait preservation in voice cloning
The most direct path to a natural-sounding voice is starting from a real human voice and preserving it at a fine enough level of detail that the clone retains what makes that person sound like themselves.
Gradium's voice cloning uses cross-attention layers that attend to speaker recordings during training, which gives the model freedom to respect speaker identity while generating new text. This differs from the prefixing technique used by many providers, which "processes the voice sample as if it has already generated that audio, treating the new text as a continuation".[2]
The result is preservation of what Gradium calls micro-traits: "vocal fry, rasp, breathiness, and pitch dynamics, even when switching languages or scripts".[2] These are the features that make a cloned voice sound like a specific person rather than a generic voice in the same category.
Voice similarity control (classifier-free guidance)
One of the most practical tools for controlling voice naturalness is the classifier-free guidance (CFG) scale, exposed in Gradium as voice similarity: the cfg_coef parameter in json_config (range 1.0 to 4.0, default 2.0) and the Voice Similarity slider in Gradium Studio.[3]
CFG controls how strongly the model follows the voice characteristics of the reference audio during synthesis. In the notation of Gradium's cloning post, with guidance scale α:[2]
- α = 1: standard model output
- α between 0 and 1: interpolation between the baseline voice and the cloned voice characteristics
- α above 1: extrapolation that amplifies the distinctive characteristics of the speaker
Above 1, the system moves beyond standard cloning into emphasizing what makes the voice specific: its quirks, its texture, its micro-level patterns. The tradeoff is that extreme values can over-amplify characteristics like hesitations or background sounds from the reference recording. Gradium's docs recommend 2.0 to 3.0 for most voice cloning work.[3]
In production, this parameter lets developers tune the balance between voice fidelity and generation quality. A clone that sounds too generic can be pushed toward identity specificity by raising cfg_coef. A clone that over-amplifies artefacts from the reference audio can be pulled back by lowering it.
Speed and pacing control
Unnatural pacing is one of the most immediately perceptible signs of a robotic voice. Human speech is delivered at variable speed: emphasized words slow down, transitions between clauses speed up, hesitations appear at specific points.
Gradium exposes speed control through the padding_bonus parameter in json_config (range -4.0 to 4.0, default 0.0). Negative values produce faster speech, positive values produce slower speech.[3] This lets developers calibrate delivery pace to the register of the interaction. A customer service agent confirming a booking should pace differently from a voice reading a news headline.
For the full parameter reference, see json_config in Gradium.
Speaking style matching
Naturalness is a property of the match between the voice and its context, as well as of the voice itself. Gradium's voice library includes voices designed for specific registers: conversational warmth for customer care agents, broadcast authority for media production.
For applications where the voice is cloned from a real speaker, the reference audio carries both the identity and the speaking style. A clone created from a casual conversation recording carries that conversational register. A clone from a professional broadcast recording carries the broadcast register. Choosing reference audio that matches the intended deployment context determines part of whether the resulting voice sounds appropriate.
Where no suitable speaker exists, an original voice can be generated from a written description instead of cloned. Gradium's Voice Designer at studio.gradium.ai/voices/design produces a voice from a text prompt, which lets the register be specified directly (warm, brisk, authoritative) rather than inherited from whoever was recorded.
Structured content
Mispronounced numbers, codes, and email addresses break the illusion of a natural speaker faster than flat prosody does. Two controls help: the rewrite_rules parameter in json_config applies language-specific text normalization rules before synthesis, and pronunciation dictionaries (passed as pronunciation_id, a top-level setup field) fix recurring proper nouns and product names.[3] The normalizers themselves are covered in text normalization edge cases. Gradium's August 2026 hard-case evaluation, on 500 sentences across five languages built from real voice-agent failures, is the reference for how far this goes: Gradium TTS passed 81.0% of the set.[4]
Benchmark evidence: what naturalness looks like in evaluation
Gradium's voice cloning benchmark used a three-level sentence complexity structure to test naturalness across contexts:[2]
- Level 1: simple conversational questions and common interactions ("I have an opening at 3:00 PM")
- Level 2: moderate complexity with names, dates, and structured content
- Level 3: complex numbers, URLs, emails, addresses, and rare named entities ("Your onomatopoeic description of the sound is helpful, but I need the alphanumeric error code e.g., 0x80070057")
The benchmark covered English, French, Spanish, and German, with 890 sentences per language, 20 unique voices per language, and a 10-second source sample per voice. Human evaluators conducted blinded A/B listening tests, comparing anonymized clones with original recordings through a live Elo ranking. In total, 3,220 voice pairs were evaluated against ElevenLabs Flash, and Gradium achieved the highest Elo in all four languages.[2]
The three-level structure matters for naturalness specifically: a voice can sound natural on Level 1 sentences and robotic on Level 3 sentences, which is what happens when a model handles simple conversational speech well but stumbles on the structured data that real-world voice agents produce constantly.
A checklist for more natural AI voice output
| Property | What causes robotic sound | How to address it |
|---|---|---|
| Micro-traits | Model compresses voice identity cues | Use a cloning architecture that preserves vocal fry, rasp, breathiness, pitch dynamics |
| Voice similarity | Clone sounds generic rather than like the specific speaker | Raise cfg_coef (Voice Similarity) toward 2.5 to 3.0 to amplify speaker characteristics |
| Pacing | Uniform speed regardless of emphasis or clause structure | Use padding_bonus in json_config to calibrate speech rate per deployment context |
| Register match | Voice designed for one context deployed in another | Select reference audio or catalogue voices that match the intended interaction register |
| Structured content | Mispronounced numbers, codes, emails break naturalness | Use rewrite_rules and pronunciation dictionaries for structured inputs |
| Cross-language consistency | Voice sounds different across languages | Use cross-lingual cloning that retains micro-traits across language switches |
Glossary
Micro-traits: Fine-grained vocal characteristics including vocal fry, rasp, breathiness, and pitch dynamics that carry speaker identity and contribute to naturalness. Often compressed by TTS models optimized for latency. Preserved by Gradium's cross-attention cloning architecture across languages and sessions.
Classifier-free guidance (CFG) scale: A mechanism that controls how strongly a voice cloning model follows the reference speaker's characteristics during synthesis. Exposed as Voice Similarity in Gradium Studio and as cfg_coef (1.0 to 4.0, default 2.0) in the API. Higher values amplify distinctive speaker traits; lower values move toward the model's baseline generation.
Prosody: The patterns of stress, intonation, rhythm, and tempo in speech. A voice with flat prosody sounds robotic even if every word is pronounced correctly. Prosody is not captured by WER and requires separate evaluation through human preference or perceptual tests.
Prefixing: A voice cloning technique where the model processes the reference audio as a prefix and treats the new text as a continuation. Requires a clean transcript of the reference audio. Gradium uses cross-attention layers instead, which gives the model more flexibility to respect speaker identity throughout generation.
Speaking style: The register and delivery characteristics of a voice, distinct from speaker identity. Includes conversational warmth, broadcast authority, expressive character, and similar dimensions. Naturalness depends on matching the speaking style to the deployment context as well as on the technical quality of the voice.
padding_bonus: A json_config parameter in Gradium's TTS API that controls speech rate, range -4.0 to 4.0, default 0.0. Negative values produce faster speech, positive values produce slower speech. Used to calibrate delivery pace to the intended interaction register.
References
[1] Gradium, "The most accurate multilingual text-to-speech, by the numbers", April 29, 2026, gradium.ai/blog/word-error-rate-evaluations
[2] Gradium, "Why Your Voice Cloning Sounds Fake (And How to Fix It)", January 22, 2026, gradium.ai/blog/voice-cloning-sounds-fake
[3] Gradium docs, Voice settings (json_config), docs.gradium.ai/guides/voice-settings
[4] Gradium, "Gradium TTS: you no longer have to choose between latency and accuracy", August 31, 2026, gradium.ai/blog/gradium-tts-latency-and-accuracy
Related guides
Part of the voice cloning cluster, hub at best voice cloning APIs 2026. Its siblings:
- How to Clone a Voice or Create a Custom Voice with Gradium: the two cloning tiers, step by step.
- Best voice cloning APIs 2026: the cluster hub, providers compared on sample length, quality and licensing.
- json_config in Gradium:
cfg_coef,padding_bonusand the rest of the setup message. - Pronunciation dictionaries: fixing the names that break the illusion.
- Fixing mispronounced names and acronyms: diagnosing which fix a term needs.
- Voice cloning across regional accents: keeping an accent through a clone.
Beyond this topic
Naturalness starts from the voice you clone. How to Clone a Voice or Create a Custom Voice with Gradium covers the two cloning tiers and their audio requirements, TTS WER benchmark 2026 covers how accuracy is measured and what it misses, and keeping a voice consistent across sessions covers holding it steady over a long render.
Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

