# Gradium > Gradium develops voice AI models for natural, expressive, ultra-low latency voice interactions at scale. Gradium provides Text-to-Speech (TTS), Speech-to-Text (STT), and speech-to-speech AI models optimized for real-time voice agent applications. For the full version of this document with complete article content, see: https://gradium.ai/llm-full.txt ## Key Resources - [API Documentation](https://docs.gradium.ai): Complete API reference for Gradium's voice AI platform - [Pricing](https://gradium.ai/pricing): Pricing plans and credit allocations - [Blog](https://gradium.ai/blog): Technical articles and company updates - [Blog RSS Feed](https://gradium.ai/blog/feed.xml): Subscribe to new blog posts - [Studio](https://studio.gradium.ai/): Voice playground and management interface - [Voices](https://gradium.ai/voices): Every accent Gradium maps, the voices in the library, and the prompt that designs one - [Voice Designer](https://studio.gradium.ai/voices/design): Generate an original voice from a text description ## Technical Content - [How to Choose a TTS API in 2026: A Developer's Decision Guide](https://gradium.ai/content/how-to-choose-a-tts-api): How to choose a TTS API: the five criteria that matter in production (TTFA, WER, streaming, cloning, cost) and which providers lead on each. - [How to Pick a TTS Provider in 2026: ElevenLabs, Cartesia, OpenAI, and Gradium Compared](https://gradium.ai/content/how-to-pick-tts-provider-elevenlabs-cartesia-openai): ElevenLabs, Cartesia, OpenAI, or Gradium? How to pick a TTS provider in 2026 based on latency, WER, voice quality, cloning, and cost. Data-backed comparison. - [How to Compare TTS Pricing Across Providers in 2026](https://gradium.ai/content/how-to-compare-tts-pricing-across-providers-2026): TTS pricing in 2026: per-character, credit-based, token-based models explained. Full comparison of Gradium, ElevenLabs, Cartesia, OpenAI, and more. - [How to Add Text-to-Speech to Your App, Website, or Voice Agent](https://gradium.ai/content/how-to-add-text-to-speech-to-app-website-voice-agent): How to add TTS to your app, website, or voice agent: streaming WebSocket, LiveKit, Pipecat, and direct API integration with Gradium in under 100 lines. - [How to Connect TTS to an LLM in a Voice Pipeline (STT → LLM → TTS)](https://gradium.ai/content/how-to-connect-tts-to-llm-voice-pipeline): How to connect TTS to an LLM in a voice pipeline: cascade architecture, streaming optimization, end-of-turn detection, and total latency budget explained. - [How to Stream TTS Audio in Real Time: Bidirectional Streaming Guide](https://gradium.ai/content/how-to-stream-tts-audio-real-time-bidirectional-streaming): How to stream TTS in real time: WebSocket lifecycle, bidirectional streaming, LLM-TTS interleaving, and multiplexing explained with Gradium's API. - [How to Fix TTS Mispronouncing Names, Acronyms, and Technical Terms](https://gradium.ai/content/how-to-fix-tts-mispronouncing-names-acronyms-technical-terms): How to fix TTS mispronouncing names, acronyms, and technical terms: pronunciation dictionaries in Gradium Studio and the Python SDK, step by step. - [How to Make TTS Pronounce Numbers, Dates, and Phone Numbers Correctly](https://gradium.ai/content/how-to-make-tts-pronounce-numbers-dates-currencies-phone-numbers): How to make TTS pronounce numbers, dates, currencies, and phone numbers correctly: native normalization, rewrite_rules, and pronunciation dictionaries explained. - [How to Keep a Voice Consistent Across Long-Form or Multi-Session Audio](https://gradium.ai/content/how-to-keep-voice-consistent-long-form-multi-session-audio): How to keep a voice consistent across long-form or multi-session audio: voice cloning tiers, micro-traits, session architecture, and cross-session identity. - [How to Stop TTS from Switching Accents or Languages Mid-Sentence](https://gradium.ai/content/how-to-stop-tts-switching-accents-languages-mid-sentence): Why TTS switches accents or languages mid-sentence, and how to fix it: native code-switching, single-model architecture, and language depth explained. - [Cloud vs On-Device TTS: How to Choose the Right Architecture](https://gradium.ai/content/cloud-vs-on-device-tts-how-to-choose): Cloud vs on-device TTS: when each architecture is the right choice, what Gradium Phonon delivers on-device, and how many products use both. - [Best API to Build an AI Voice Agent in 2026](https://gradium.ai/content/best-api-build-ai-voice-agent-2026): Best API to build an AI voice agent in 2026: Gradium leads on TTFA and WER on Coval. Compare voice layer APIs and full-stack platforms in one guide. - [Why Generic TTS Fails on Regional Accents](https://gradium.ai/content/voice-cloning-regional-accents-2026): Regional accent voice AI: generic TTS flattens Bavarian, Argentinian, and Quebecois speech into one default. Gradium preserves accent from a 10-second sample. - [Gradium Phonon: On-Device TTS Benchmarks in 2026](https://gradium.ai/content/gradium-phonon-on-device-tts-benchmarks-2026): Gradium Phonon on-device TTS reaches 1.00% WER with voice cloning and 0.83% WER fixed voice on Seed-TTS, at 100M parameters. Full benchmarks and architecture. - [Semantic VAD for Voice Agents: Turn Detection 2026](https://gradium.ai/content/semantic-vad-voice-agents-turn-detection-2026): Semantic VAD explained: why voice agents interrupt users mid-sentence, how turn detection actually works, and how to configure delay_in_frames in Gradium STT. - [Gradium TTS in Pipecat: Setup and Integration Guide](https://gradium.ai/content/gradium-pipecat-native-tts-integration-voice-agents): Gradium TTS integration for Pipecat: install GradiumTTSService, configure voice and model settings, and stream low-latency speech in your voice agent. - [STT API Benchmark 2026: Latency and Accuracy for Voice Agents](https://gradium.ai/content/stt-api-benchmark-2026-latency-accuracy): Independent Coval benchmark comparing 5 STT APIs on TTFT and WER for voice agents: Gradium, Deepgram Nova-3, Nova-2, AssemblyAI Universal Streaming, and ElevenLabs Scribe v2. Latency vs accuracy tradeoffs explained. - [Turn-Taking in Voice Agents: Why Rule-Based VAD Is Broken and What Comes Next](https://gradium.ai/content/turn-taking-voice-agents-vad): Why voice activity detection rules make every cascaded voice agent feel unnatural. The turn-taking problem explained, how full duplex solves it, and where the architecture is heading in 2026. - [Cascaded Voice Agents vs Speech-to-Speech: Architecture Tradeoffs in 2026](https://gradium.ai/content/cascaded-voice-agent-vs-speech-to-speech-2026): Cascade (STT + LLM + TTS) vs speech-to-speech for voice agents: latency, modularity, paralinguistic information, and LLM flexibility compared. Which architecture fits which use case in 2026. - [Phonon Reaches 1.00% WER on Seed-TTS in May 2026: Smallest On-Device TTS Model in the Comparison](https://gradium.ai/content/phonon-seed-tts-benchmark-2026): Phonon, Gradium's 100M-parameter on-device Text-to-Speech model, reaches 1.00% WER on the Seed-TTS English benchmark in May 2026 with voice cloning, and 0.83% WER with a fixed voice. Outperforms NeuTTS Air (552M), KaniTTS2 (450M), NeuTTS Nano (229M), Kokoro (82M), Magpie (357M), and Supertonic 2 (66M). - [On-Device Text-to-Speech in 2026: When Edge TTS Is the Right Architecture](https://gradium.ai/content/on-device-text-to-speech-2026): When on-device TTS is the right architecture in 2026: offline apps, high-volume consumer products, and privacy-constrained deployments. Gradium Phonon benchmarks and use case guide. - [Gradium Phonon: On-Device TTS for Mobile Apps, NPCs, and Offline Products](https://gradium.ai/content/gradium-phonon-on-device-tts): Gradium Phonon is an on-device TTS model that runs on CPU across Android, iOS, and browser with no network dependency. 1.48% WER, 56.37% speaker similarity on Seed-TTS. Built for mobile apps, game NPCs, and offline products. - [On-Device TTS Benchmark 2026: Phonon vs Kani-TTS2 vs NeuTTS on Seed-TTS](https://gradium.ai/content/on-device-tts-benchmark-2026): Independent on-device TTS benchmark 2026: Gradium Phonon vs Kani-TTS2 vs NeuTTS Air vs NeuTTS Nano on Seed-TTS English. WER and speaker similarity results, methodology, and what they mean for edge deployment. - [Best AI Voice Generators in 2026: APIs Ranked by Voice Quality, Latency, and Price](https://gradium.ai/content/best-ai-voice-generators-2026): Compare the best AI voice generator APIs in 2026: voice quality (Artificial Analysis ELO), latency, pricing, and production benchmarks. Includes Gradium, ElevenLabs, Inworld, Google, OpenAI, and more. - [Best Speech APIs in 2026: TTS, STT Compared](https://gradium.ai/content/best-speech-apis-2026): Compare the best speech APIs in 2026 for text-to-speech and speech-to-text. Verified pricing, latency benchmarks, and production data. - [Best ElevenLabs Alternatives in 2026: Top TTS APIs Ranked by Voice Quality and Price](https://gradium.ai/content/best-elevenlabs-alternatives-2026-tts-apis-voice-quality-price): Looking for ElevenLabs alternatives in 2026? Compare the top TTS APIs ranked by independent voice quality ELO score (Artificial Analysis) and pricing per million characters, including Gradium, Inworld, Google, OpenAI, and more. - [Azure TTS Alternative: Gradium for Real-Time Voice AI](https://gradium.ai/content/azure-tts-alternative-gradium): Microsoft Azure TTS alternative for voice agents: Gradium delivers streaming TTS and STT with semantic VAD, sub-300 ms TTFA, instant voice cloning, and transparent pricing. - [Best Low-Latency TTS APIs in 2026: TTFA, P99 and Pipeline Impact](https://gradium.ai/content/best-low-latency-tts-apis-2026): Compare the best low-latency TTS APIs in 2026, benchmarked by TTFA (P50/P99), WebSocket architecture, codebook configs and full voice agent pipeline impact. - [Best Multilingual TTS APIs in 2026: Coverage, Quality, Code-Switching](https://gradium.ai/content/best-multilingual-tts-apis-2026): Compare the best multilingual TTS APIs in 2026: Gradium, ElevenLabs, Cartesia, Deepgram on language coverage, code-switching, voice cloning and latency. - [Best Text-to-Speech APIs in 2026: Developer Guide](https://gradium.ai/content/best-text-to-speech-apis-2026): Compare the best TTS APIs in 2026 by latency, voice cloning, languages, and pricing. Covers Gradium, ElevenLabs, Cartesia, Deepgram Aura-2, and OpenAI TTS. - [Best Voice Cloning APIs in 2026: Instant Cloning, Fine-Tuning, Benchmarks](https://gradium.ai/content/best-voice-cloning-apis-2026): Compare the best voice cloning APIs in 2026: Gradium, ElevenLabs, Cartesia. Instant vs. professional cloning, speaker similarity benchmarks, pricing. - [Gradium vs ElevenLabs for Voice Agents: TTFA, WER and IQR Compared (2026 Coval Data)](https://gradium.ai/content/gradium-vs-elevenlabs-voice-agents-benchmark): Gradium vs ElevenLabs for voice agents in 2026. Independent Coval benchmark data on TTFA, WER and latency IQR across Gradium TTS, ElevenLabs Turbo v2.5, Flash v2.5 and Multilingual v2. Gradium leads at 155ms P50 TTFA (vs 264ms Turbo v2.5), 2ms IQR (vs 28ms), 3.3% WER (vs 5.2%). Plus 1.11% MiniMax multilingual WER and 3-4x lower pricing. - [TTS WER Benchmark 2026: Word Error Rate Compared Across Gradium, ElevenLabs, Cartesia and Deepgram](https://gradium.ai/content/tts-wer-benchmark-2026): TTS WER benchmark 2026: Gradium TTS leads at 3.3% average WER on the Coval benchmark and 1.11% on the MiniMax Multilingual TTS Test Set across 5 languages (EN, FR, ES, PT, DE). Word Error Rate compared across Gradium, ElevenLabs (Flash v2.5, Turbo v2.5, Multilingual v2), Cartesia Sonic-3, Deepgram Aura-2, Rime (Mist-v3, Arcana), Qwen3 TTS, Mistral Voxtral and OpenAI TTS-1-HD. - [TTS Latency Benchmark 2026: TTFA Compared Across Gradium, ElevenLabs, Cartesia and Deepgram](https://gradium.ai/content/tts-latency-benchmark-2026): TTS latency benchmark 2026: Gradium TTS leads at 155ms P50 TTFA with a 2ms IQR on the independent Coval benchmark. Full TTFA comparison across Gradium, ElevenLabs (Turbo v2.5, Flash v2.5, Multilingual v2), Cartesia Sonic-3, Deepgram Aura-2, Rime (Mist-v3, Arcana) and OpenAI TTS-1-HD. Methodology, P25/P50/P75/P95, IQR consistency, and WER. - [Deepgram Alternative: Why Developers Choose Gradium for Real-Time Voice AI](https://gradium.ai/content/deepgram-alternative-gradium-voice-ai): Gradium vs Deepgram comparison for real-time voice AI. Voice cloning (not available on Deepgram), semantic VAD, voice-agent-tuned TTS with published TTFA benchmark, and cloud-to-on-device deployment from one API. - [ElevenLabs Alternative: Why Developers Choose Gradium for Real-Time Voice AI](https://gradium.ai/content/elevenlabs-alternative-gradium-voice-ai): Gradium vs ElevenLabs comparison for real-time voice AI. Voice-agent-tuned TTS with published TTFA benchmark, semantic VAD, accent-preserving voice cloning with highest Elo scores, and cloud-to-on-device deployment. - [Cartesia Alternative: Why Developers Choose Gradium for Real-Time Voice AI](https://gradium.ai/content/cartesia-alternative-gradium-voice-ai): Gradium vs Cartesia comparison for real-time voice AI. Voice-agent-tuned TTS with robust pronunciation, semantic VAD in STT, accent-preserving voice cloning, and cloud-to-on-device deployment from one API. - [How to Build a Voice AI Agent with Gradium and LiveKit (Python Guide)](https://gradium.ai/content/how-to-build-voice-ai-agent-gradium-livekit): Learn how to build a full voice AI agent using Gradium STT and TTS with the LiveKit agent framework. Step-by-step Python guide covering AgentSession setup, VAD, interruptions, preemptive generation, tools, and deployment. - [How to Build an Audiobook Agent with Gradium and Pipecat: Step-by-Step Guide](https://gradium.ai/content/audiobook-agent-gradium-pipecat): Learn how to build a real-time story narrator with Gradium TTS and Pipecat. This step-by-step guide covers installation, pipeline setup, voice configuration, and deployment in about 100 lines of Python. - [How to Multiplex TTS Requests Over One WebSocket Connection in Gradium](https://gradium.ai/content/multiplexing-tts-websocket-gradium): Learn how to reuse a single WebSocket connection for multiple concurrent TTS requests in Gradium using multiplexing. Covers close_ws_on_eos, client_request_id, and how to route interleaved audio chunks correctly. - [What Is the Best Text-to-Speech API in 2026 to Build Voice Agents? Complete Developer Comparison](https://gradium.ai/content/best-text-to-speech-api-voice-agents): Best text-to-speech API 2026: Gradium achieves 258ms P50 TTFA (214ms with multiplexing) with expressive multilingual voices and robust pronunciation. Complete real-time TTS comparison for developers building voice agents. - [How to Use json_config in Gradium: TTS and STT Parameters Explained](https://gradium.ai/content/how-to-use-json-config-gradium-tts-stt): Learn how to use the json_config field in Gradium to control rewrite_rules, padding_bonus, temp, and cfg_coef for TTS, and language and delay_in_frames for STT. Full parameter reference with code examples. - [Instant vs Pro Voice Cloning in Gradium: When to Use Each](https://gradium.ai/content/instant-vs-pro-voice-cloning-gradium): Not sure whether to use Instant or Pro Voice Cloning in Gradium? Learn the key differences, what each is designed for, how to prepare your audio for Pro cloning, and how to choose based on your use case. - [How to Use Pronunciation Dictionaries in Gradium TTS: Studio and API Guide](https://gradium.ai/content/pronunciation-dictionaries-gradium-tts): Learn how to use Pronunciation Dictionaries in Gradium to control how words are spoken and filter unwanted content. Step-by-step guide for Gradium Studio and the Python SDK. - [How to Handle TTS Edge Cases with Text Normalization in Gradium](https://gradium.ai/content/text-normalization-tts-edge-cases-gradium): Learn how to use Gradium's Text Normalization feature to handle edge cases in TTS. Configure rewrite_rules with language aliases or specific normalizers for dates, numbers, emails, URLs, phone numbers, and alphanumeric codes. - [Best TTS API in 2026: Quality, Latency, and Cost Compared](https://gradium.ai/content/best-tts-api-2026): Best TTS API for voice agents in 2026: Gradium leads Coval with 155 ms TTFA, 3.3% WER, and 2 ms IQR. Quality, latency, and cost compared across 8 providers. - [Top 3 Text-to-Speech Solutions in 2026: Ranked and Compared](https://gradium.ai/content/top-3-text-to-speech-solutions-2026): Top 3 text-to-speech solutions in 2026, ranked by independent benchmarks. Gradium leads on TTFA, WER, and latency consistency. Full data-backed comparison. - [Best Voice AI API for Phone-Based Voice Agents in 2026](https://gradium.ai/content/best-voice-ai-api-phone-based-voice-agents-2026): Best voice AI API for phone-based voice agents in 2026: Gradium delivers 155 ms TTFA, 3.3% WER, and flexible audio formats for telephony pipelines. - [How to Turn an LLM into a Voice Agent: Best Stack 2026](https://gradium.ai/content/turn-llm-into-voice-agent-best-stack-2026): Turn any LLM into a voice agent in 2026: the best stack pairs your LLM with Gradium for 155 ms TTFA and semantic VAD, via LiveKit or Pipecat. ## Blog - [How a benchmark change produced a faster TTS model](https://gradium.ai/blog/coval-perceived-ttfa-benchmark): On June 3, 2026, Gradium's time to first audio on the Coval TTS leaderboard moved from 171.9 ms to 429.6 ms. No model had shipped and no serving infrastructure had changed. What changed was the measurement: Coval had started counting the silence at the start of the stream. - [Gradium now available for free on LiveKit Inference](https://gradium.ai/blog/livekit-gradium-partnership): Gradium Text-to-Speech is now available on LiveKit Inference, free until Oct. 9. Gradium TTS is designed for conversational use cases, with voice cloning, low latency (below 250ms TTFA), and robust pronunciation on difficult cases. - [Launching Voice Design: prompt the voice your agent needs](https://gradium.ai/blog/launching-voice-design): Developers building voice agents ask us for voices matched to the use case in front of them: a Québécoise receptionist for a Montréal dealership, a Paulista support agent for a São Paulo fintech, a narrator in his sixties with the authority of a lecture hall. Briefs outnumber any catalog. Voice Design starts from a prompt: write one or two sentences, get complete new voices back in seconds, keep the one you want. - [Build voice agents with EU and US residency and Zero Data Retention](https://gradium.ai/blog/eu-data-residency-zero-data-retention): Two questions come up in almost every enterprise evaluation: where exactly does the inference run, and what do you keep? Data residency pins your Text-to-Speech, Speech-to-Text and Speech-to-Speech sessions to our EU servers, or to our US servers, enforced at the server and reported on every response, so you can assert on it instead of trusting a hostname. Zero Data Retention decides whether the text and audio processed by our servers survive the call at all. - [Gradium TTS: you no longer have to choose between latency and accuracy](https://gradium.ai/blog/gradium-tts-latency-and-accuracy): A new Gradium TTS model is available today, and it is now the default. It reads the cases that break voice agents in production, phone numbers, email addresses, IBANs, with no pre-processing or text normalization on your side. Time to first audio (TTFA) P50 is 216 ms on Coval, 170 ms faster than the model it replaces, with a 30 ms p75 - p25 spread, the tightest of the five models tested. Here is the hard-case evaluation set, the head-to-head samples, and how to hear the difference yourself. - [Gradium voices are now live on AudioStack](https://gradium.ai/blog/audiostack-gradium-partnership): AudioStack, the agentic audio production platform for media and advertising, has added Gradium to its roster of voice providers. Our voices are live on the platform now, across multiple languages and regional accents. - [New voices: how we pick the one that wins](https://gradium.ai/blog/new-voices-how-we-pick-the-one-that-wins): Picking a voice works like casting: one candidate in a thousand has a hit factor you cannot explain. Here is the methodology behind every flagship voice in the Gradium catalog, from generating hundreds of candidates per category to ranking them head-to-head against the incumbents on an ELO benchmark, why the same prompt does not travel across locales, and how to write the target yourself. - [Gradium TTS public beta: open for your hardest cases](https://gradium.ai/blog/gradium-tts-public-beta): A new Gradium TTS model is available today in public beta. It keeps improving on the current production model, handling the complex cases natively, phone numbers, email addresses, IBAN numbers, time expressions and more, with no pre-processing or text normalization on your side. Anyone can test it through the API, and everyone who sends us feedback during the beta receives 1M Gradium credits. - [Keyword boosting: teaching Speech-To-Text your vocabulary](https://gradium.ai/blog/stt-keyword-boosting): Speech-To-Text models rarely see your product's vocabulary, brand names, drugs, athletes, and places, so they fall back on a similar-sounding common word. Keyword boosting raises the probability of the terms you care about while decoding, in real time, with no retraining. Here is how the boost parameter behaves, what a sweep across hard clips shows, and how to add it in a few lines. - [Smaller, clearer, better: Phonon Multilingual beats Magpie and NeuTTS Nano in French, German, and Spanish](https://gradium.ai/blog/phonon-multilingual-smaller-clearer-better): Phonon, our 100M-parameter on-device Text-To-Speech model, produces up to 3.5x fewer word errors than NVIDIA's Magpie TTS in French, German, and Spanish at 3.6x fewer parameters, and leads NeuTTS Nano on both word error rate and speaker similarity. It clones a reference voice across five languages: English, French, German, Spanish, and Portuguese. - [Natural conversation, live retrieval: Keenable search with Gradium](https://gradium.ai/blog/keenable-gradium-partnership): Gradium and Keenable are partnering to bring realtime web search into natural voice agents. Gradbot, Gradium's open source voice agent framework, is adding Keenable search so agents can retrieve fresh answers from the live web without breaking the flow of conversation. - [Gradium Extends Funding to $100 Million and Expands to Silicon Valley](https://gradium.ai/blog/gradium-100-million-funding-nvidia): Gradium extends its seed funding to $100 million, welcoming new investors including NVIDIA, and opens a San Francisco Bay Area office to scale its real-time voice AI. - [Gradium Powers RMC BFM Drive: AI-Generated Personalized Radio in Renault Vehicles](https://gradium.ai/blog/rmc-bfm-drive-gradium): Gradium powers RMC BFM Drive, the first fully AI-generated, personalized radio experience built into Renault connected vehicles. Our Text-To-Speech models synthesize transitions in real time using Pro Voice Clones of RMC BFM journalists, turning an AI-curated program into something that sounds and feels like a live broadcast. - [Launching Gradium Translate: the best accuracy-latency tradeoff against gemini-3.5-live-translate and gpt-realtime-translate](https://gradium.ai/blog/launching-gradium-translate): Gradium launches stt-translate and s2s-translate: real-time Speech-To-Text and Speech-To-Speech translation across English, French, German, Spanish, and Portuguese, with a better accuracy-latency tradeoff than gpt-realtime-translate and gemini-3.5-live-translate and full control of the output voice, including cloning. - [Gradium TTS, upgraded: more accurate Text-To-Speech](https://gradium.ai/blog/gradium-tts-upgrade): Gradium TTS now runs on a new model: more natural prosody and substantially more accurate pronunciation on the cases that break voice agents in production, including spelling, acronyms, emails, phone numbers, and codes. It wins head-to-head against our previous model in all five languages, and leads real-time competitors (Cartesia Sonic 3.5, Inworld TTS 1.5 Max, ElevenLabs Flash v2.5 and Multilingual v2) on the hardest pronunciation cases. Available now as the new default, with custom voices carried over. - [Semantic VAD: turn detection that uses meaning, not silence](https://gradium.ai/blog/semantic-vad): Acoustic VAD answers "is there a voice right now?" Semantic VAD answers "is the user done talking?" Here's why the distinction decides whether a voice agent cuts users off, how Gradium STT emits multi-horizon turn-completion predictions every 80 ms, and how to tune delay_in_frames, horizon, and flushing for your use case. - [Phonon update: 1.00% WER on Seed-TTS, smaller than every model we beat](https://gradium.ai/blog/phonon-update-may-2026): Phonon, our 100M-parameter on-device Text-To-Speech model, reaches 1.00% WER on the Seed-TTS English benchmark, outperforming NeuTTS Air, KaniTTS2, and NeuTTS Nano. With a fixed voice, it drops to 0.83% WER, ahead of Kokoro and Magpie. - [Gradium #1 on Coval TTS Benchmarks](https://gradium.ai/blog/coval-tts-benchmarks-may-2026): Independent Coval TTS benchmark (May 2026): Gradium ranks first on P50 TTFA (158ms), latency IQR (2ms), and provides SOTA WER (3.7%) against ElevenLabs Turbo v2.5, Flash v2.5, Multilingual v2, Cartesia Sonic-3, Deepgram Aura-2, Rime Mist-v3, Arcana, and OpenAI TTS-1-HD. - [Gradium Voice Launches on AWS as a SaaS Subscription and a SageMaker Model Image](https://gradium.ai/blog/gradium-aws-launch): Gradium is now available on AWS through two paths: a fully managed SaaS subscription via AWS Marketplace, and a deployable model image via Amazon SageMaker for teams that need in-VPC inference. - [The most accurate multilingual text-to-speech, by the numbers](https://gradium.ai/blog/word-error-rate-evaluations): How we measure WER for TTS at Gradium: text normalization, jiwer alignment, results on the MiniMax Multilingual benchmark across English, French, Spanish, Portuguese and German — and why the standard metric is starting to saturate. - [Gradbot: Vibe code voice agents in 50 lines of code](https://gradium.ai/blog/gradbot): Gradbot is our open-source framework for prototyping voice agents in minutes. Built on a Rust orchestration core, it handles turn-taking, interruptions, silence, and async tool calls so you can ship a working voice experience in around 50 lines of code. - [Evaluating Phonon: how we made the best TTS model for edge devices](https://gradium.ai/blog/evaluating-phonon): An evaluation of Gradium Phonon, our on-device text-to-speech model. Despite its small size, it significantly outperforms larger models. - [Gradium Phonon: On-Device TTS for Consumer Apps, NPCs, and Offline Products](https://gradium.ai/blog/gradium-phonon): Announcing Gradium Phonon, our new on-device text-to-speech model designed for consumer apps, NPCs, and offline products. - [Time to First Audio: Measuring and Reducing TTS Latency in Voice Agents](https://gradium.ai/blog/time-to-first-audio): In natural conversation, the gap between one person finishing a sentence and the other starting to respond averages around 200 milliseconds. For voice agents this is the target to match. - [InteractionLabs (Ongo) and Gradium Partner to Redefine Human-Robot Interaction](https://gradium.ai/blog/interactionlabs-gradium-partnership): InteractionLabs, the company behind the Ongo living lamp robot, and Gradium announce a partnership to bring expressive, real-time voice AI to robotics. - [Optimizing Quality vs. Latency in Real-Time Text-to-Speech AI Models](https://gradium.ai/blog/optimizing-quality-vs-latency): Explore strategies for balancing quality and latency in real-time TTS AI models. Learn how Gradium achieves low-latency, high-quality speech synthesis for voice applications. - [Building Voice Agents From the Ground Up: The Gradium Startup Program](https://gradium.ai/blog/gradium-startup-program): Get 6 months free access to Gradium's voice AI platform. 9M monthly credits, voice cloning, STT/TTS APIs for seed-funded startups building voice-first products. - [Acolad and Gradium Partner to Advance Enterprise-Ready AI Interpreting](https://gradium.ai/blog/acolad-gradium-partnership): Acolad, the global leader in language and content solutions, and Gradium just announced a strategic partnership. The partnership reflects Acolad’s commitment to delivering secure, scalable, and governed AI-powered interpreting solutions, designed for enterprise and public-sector environments. - [Invincible Voice: How Gradium's Real-Time Voice AI Helps ALS Patients Speak Again](https://gradium.ai/blog/invincible-voice): Gradium's voice AI technology powers Invincible Voice, an open-source assistive system helping people with ALS and speech loss communicate in real-time. - [Why Your Voice Cloning Sounds Fake (And How to Fix It)](https://gradium.ai/blog/voice-cloning-sounds-fake): Discover how Gradium's instant voice cloning achieves superior speaker similarity to ElevenLabs. Benchmark results across 4 languages with 3,220 human evaluations. - [Powering Wonderful's Voice Agents](https://gradium.ai/blog/wonderful): We're proud to power real-time voice agents on Wonderful's platform, bringing cutting-edge voice AI from experimental to deployable. - [Gradium: Solving voice](https://gradium.ai/blog/gradium): Today we're excited to launch Gradium, the core engine powering the next generation of voice products and interactions. ## Company Information - Product: Voice AI platform (TTS, STT, Speech-to-Speech) - Key differentiator: Ultra-low latency TTS (155 ms TTFA P50 on the independent Coval benchmark, May 2026), robust pronunciation on hard cases (phone numbers, emails, codes), Instant Voice Cloning from 10 seconds of audio - Languages: English, French, German, Spanish, Portuguese, with native fluency and mid-sentence code-switching - STT: streaming Speech-to-Text with native semantic VAD, same credit pool as TTS - On-device: Gradium Phonon, ~100M-parameter on-device TTS (private beta) - Deployment options: Cloud API, dedicated instances, self-hosted, on-premises, on-device - Free tier: Available with no credit card required