Voice Design. Live today.

Blog

What can we help youfind today?

Discover our articles and videos

September 9, 20265 minEngineering

How a benchmark change produced a faster TTS model

On June 3, 2026, Gradium's time to first audio on the Coval TTS leaderboard moved from 171.9 ms to 429.6 ms. No model had shipped and no serving infrastructure had changed. What changed was the measurement: Coval had started counting the silence at the start of the stream.

Discover
September 9, 20263 minAnnouncement

Gradium now available for free on LiveKit Inference

Gradium Text-to-Speech is now available on LiveKit Inference, free until Oct. 9. Gradium TTS is designed for conversational use cases, with voice cloning, low latency (below 250ms TTFA), and robust pronunciation on difficult cases.

Discover
September 7, 20266 minProduct

Launching Voice Design: prompt the voice your agent needs

Developers building voice agents ask us for voices matched to the use case in front of them: a Québécoise receptionist for a Montréal dealership, a Paulista support agent for a São Paulo fintech, a narrator in his sixties with the authority of a lecture hall. Briefs outnumber any catalog. Voice Design starts from a prompt: write one or two sentences, get complete new voices back in seconds, keep the one you want.

Discover
September 2, 20265 minProduct

Build voice agents with EU and US residency and Zero Data Retention

Two questions come up in almost every enterprise evaluation: where exactly does the inference run, and what do you keep? Data residency pins your Text-to-Speech, Speech-to-Text and Speech-to-Speech sessions to our EU servers, or to our US servers, enforced at the server and reported on every response, so you can assert on it instead of trusting a hostname. Zero Data Retention decides whether the text and audio processed by our servers survive the call at all.

Discover
August 31, 20266 minResearch

Gradium TTS: you no longer have to choose between latency and accuracy

A new Gradium TTS model is available today, and it is now the default. It reads the cases that break voice agents in production, phone numbers, email addresses, IBANs, with no pre-processing or text normalization on your side. Time to first audio (TTFA) P50 is 216 ms on Coval, 170 ms faster than the model it replaces, with a 30 ms p75 - p25 spread, the tightest of the five models tested. Here is the hard-case evaluation set, the head-to-head samples, and how to hear the difference yourself.

Discover
August 19, 20263 minProduct

Gradium voices are now live on AudioStack

AudioStack, the agentic audio production platform for media and advertising, has added Gradium to its roster of voice providers. Our voices are live on the platform now, across multiple languages and regional accents.

Discover
August 4, 20269 minProduct

New voices: how we pick the one that wins

Picking a voice works like casting: one candidate in a thousand has a hit factor you cannot explain. Here is the methodology behind every flagship voice in the Gradium catalog, from generating hundreds of candidates per category to ranking them head-to-head against the incumbents on an ELO benchmark, why the same prompt does not travel across locales, and how to write the target yourself.

Discover
July 30, 20263 minResearch

Gradium TTS public beta: open for your hardest cases

A new Gradium TTS model is available today in public beta. It keeps improving on the current production model, handling the complex cases natively, phone numbers, email addresses, IBAN numbers, time expressions and more, with no pre-processing or text normalization on your side. Anyone can test it through the API, and everyone who sends us feedback during the beta receives 1M Gradium credits.

Discover
July 21, 20267 minEngineering

Keyword boosting: teaching Speech-To-Text your vocabulary

Speech-To-Text models rarely see your product's vocabulary, brand names, drugs, athletes, and places, so they fall back on a similar-sounding common word. Keyword boosting raises the probability of the terms you care about while decoding, in real time, with no retraining. Here is how the boost parameter behaves, what a sweep across hard clips shows, and how to add it in a few lines.

Discover
July 15, 20265 minResearch

Smaller, clearer, better: Phonon Multilingual beats Magpie and NeuTTS Nano in French, German, and Spanish

Phonon, our 100M-parameter on-device Text-To-Speech model, produces up to 3.5x fewer word errors than NVIDIA's Magpie TTS in French, German, and Spanish at 3.6x fewer parameters, and leads NeuTTS Nano on both word error rate and speaker similarity. It clones a reference voice across five languages: English, French, German, Spanish, and Portuguese.

Discover