NewOur fastest Text-to-Speech model yet ・ Sub-50ms latency

Comparison

Best Speech APIs 2026: TTS and STT From One Vendor

7 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • This page is about buying both halves of the speech stack, Text-to-Speech and Speech-to-Text, from one vendor. For Text-to-Speech alone, start at best Text-to-Speech APIs 2026.
  • Four vendors ship both halves as of September 2026: Gradium, Deepgram, ElevenLabs and Inworld. Cartesia added Speech-to-Text with Ink-2 on July 9, 2026. OpenAI sells them as separate products.
  • Gradium is the only one of the four that draws both from a single credit pool: 1 credit per Text-to-Speech character and 3 credits per Speech-to-Text second, so one hour of transcription costs 10,800 credits.
  • On the Coval TTS leaderboard read September 8, 2026, Gradium TTS recorded 214 ms median perceived time to first audio with 0 ms leading silence, rank 1 of 27 on that metric on the September 10, 2026 board.
  • Deepgram is the cheapest one-vendor stack for transcription-heavy workloads at $0.0043 per minute for Nova-3, or about $0.26 per hour, against Gradium's $0.39 to $0.62 per hour depending on plan.

Why buy both halves from one vendor?

Three reasons that actually show up in production, and one that does not.

One credit pool absorbs a lopsided month. A voice agent's Text-to-Speech and Speech-to-Text volumes rarely move together. Separate contracts mean paying for unused capacity on one side while topping up the other. Gradium's plans draw both from one allocation, so a month heavy on transcription costs the same as a month heavy on synthesis.

Turn detection sits between the two. Deciding when the caller has stopped talking is a Speech-to-Text-side function whose output gates the Text-to-Speech call. When the two are one product, that boundary is inside the vendor rather than inside your orchestration code. Gradium's Speech-to-Text carries native semantic voice activity detection. See turn-taking and VAD and semantic VAD configuration.

One vendor review instead of two. Data residency, retention, subprocessors and an SLA get negotiated once.

The reason that does not hold up: latency. Nothing about sharing a vendor makes the round trip shorter. The two calls are separate network operations either way, and the stage that actually costs time is turn detection, not vendor affinity.

Who ships both halves?

Source: vendor documentation and pricing pages, read September 8, 2026.

Vendor Text-to-Speech Speech-to-Text Billing
Gradium Gradium TTS Streaming, with native semantic VAD One credit pool for both
Deepgram Aura-2 Nova-3 Separate line items, one platform
ElevenLabs Flash v2.5, Eleven v3 Conversational, Eleven v3, Multilingual v2 Scribe v2, Scribe v2 Realtime Separate products
Inworld Realtime TTS-2, TTS-2 Flash Bundled Speech-to-Text Bundled with Text-to-Speech tiers
Cartesia Sonic 3.6, Sonic 3.5 Ink-2, released July 9, 2026 Credit-based
OpenAI tts-1, tts-1-hd, GPT-4o mini TTS Transcription models Separate products

What does a one-vendor stack cost?

Source: vendor pricing pages, read September 8, 2026. Text-to-Speech per 1M characters, Speech-to-Text per hour of audio. Gradium figures are the effective rates of each plan's credit allocation at 1 credit per character and 3 credits per second.

Vendor and tier Text-to-Speech per 1M characters Speech-to-Text per hour
Deepgram $30.00 (Aura-2) $0.26 (Nova-3, $0.0043/min)
Inworld $25.00 down to $5.00 (Realtime TTS-2) $0.15 down to $0.10, bundled
Gradium L $35.90 effective $0.39 effective
Gradium M $37.80 effective $0.41 effective
Gradium S $47.80 effective $0.52 effective
Gradium XS $57.80 effective $0.62 effective
ElevenLabs $50.00 (Flash v2.5) $0.22 (Scribe v2), $0.39 (Scribe v2 Realtime)

Read it against your own ratio. An agent that listens far more than it speaks is dominated by the Speech-to-Text column, where Inworld and Deepgram are cheapest. An agent that reads long confirmations back is dominated by the Text-to-Speech column. Gradium's single pool is the option that does not require predicting the ratio in advance, since unused Text-to-Speech credits are spendable on Speech-to-Text and the reverse.

Two things the table cannot show: Gradium plan credits do not roll over month to month (pay-as-you-go credits do), and ElevenLabs was running a 50%-off-for-life API promotion through September 11, 2026.

The plan allocations behind the Gradium column: Free 45,000 credits (about 1 Text-to-Speech hour or 4 Speech-to-Text hours), XS 225,000, S 900,000, M 9M and L 45M. Concurrency runs 2, 5, 5, 10 and 15 on Text-to-Speech and 3, 20, 20, 40 and 60 on Speech-to-Text.

How good is each half?

Text-to-Speech. On the Coval leaderboard read September 8, 2026 (1-day window, 480 samples per model, 26 models), median perceived time to first audio was 185 ms for ElevenLabs Flash v2.5, 214 ms for Gradium TTS, 269 ms for Cartesia Sonic 3.5, 290 ms for Deepgram Aura-2, 320 ms for ElevenLabs Eleven v3 Conversational and 440 ms for Cartesia Sonic 3.6. Gradium's figure carries 0 ms leading silence, rank 1 of 27 on the September 10, 2026 board and the only model at zero.

Note that the metric changed on June 3, 2026: Coval now counts leading silence inside the stream, which accounts for 28% of time to first audio across the board. Any figure dated earlier measures round trip only. See How a benchmark change produced a faster TTS model.

On pronunciation, Gradium TTS passed 81.0% of a 500-sentence hard-case set rated by independent native speakers in August 2026, against 75.1% for Cartesia Sonic 3.6, 65.4% for ElevenLabs Eleven v3 Conversational, 61.5% for Inworld Realtime TTS-2 and 49.5% for Fish Audio S2.1 Pro. On the MiniMax Multilingual TTS Test Set in April 2026, Gradium averaged 1.11% word error rate across five languages.

Speech-to-Text. There is no equivalent public cross-vendor leaderboard, which means the comparison has to be run on your own audio. Test with your accents, your background noise and your domain vocabulary, and test keyword boosting explicitly if names or product terms matter. The measurement approach is in STT API benchmark 2026.

What to check before consolidating

  1. Language coverage on both halves. Gradium covers English, French, Spanish, Portuguese and German on both. If your Speech-to-Text needs a language your Text-to-Speech does not, one vendor may not cover it.
  2. Turn detection. Ask whether voice activity detection is semantic or a silence timer. A fixed timer cuts off users who pause mid-thought.
  3. Streaming on both halves. Batch transcription and streaming transcription are different products at most vendors.
  4. Concurrency on the plan you would buy, not the top tier.
  5. Data residency and retention. Gradium offers eu.api.gradium.ai and us.api.gradium.ai; EU pinning covers inference and account data, US pinning covers inference only, with account data, custom voices and pronunciation dictionaries staying in the EU. Zero Data Retention has been self-serve on all paid plans since September 2, 2026.
  6. Session length. Gradium allows 3,000 seconds per session on both halves.

Which one-vendor stack, in one line each

Gradium if you want one credit pool across both halves, native semantic turn detection, and on-premise or on-device as an option.

Deepgram if transcription volume dominates and you want the lowest per-hour Speech-to-Text price with a competent Text-to-Speech attached.

ElevenLabs if the voice catalogue or a language outside Gradium's five is the requirement, and separate billing for Scribe v2 is acceptable.

Inworld if you want the lowest bundled price at volume and the highest preference-rated voice of these four is not required.

Glossary

Speech API. Any API that converts between text and speech. In practice, a Text-to-Speech endpoint, a Speech-to-Text endpoint, or both from one vendor.

Credit pool. A single monthly allocation spendable across products. Gradium charges 1 credit per Text-to-Speech character and 3 credits per Speech-to-Text second, so an hour of transcription costs 10,800 credits.

Semantic voice activity detection. Turn detection that uses the meaning of what was said, not only how long the silence lasted, to decide the speaker has finished.

Perceived time to first audio. Round-trip time to first audio plus the leading silence inside the stream before the first audible sample. The Coval headline metric since June 3, 2026.

Hard-case pass rate. The share of adversarial sentences where a native-speaker rater hears every element pronounced correctly and completely.

References

  1. Coval TTS leaderboard: benchmarks.coval.ai/tts
  2. Coval and Gradium, "How a benchmark change produced a faster TTS model", September 9, 2026: gradium.ai/blog/coval-perceived-ttfa-benchmark
  3. Gradium, "Gradium TTS: latency and accuracy", August 31, 2026: gradium.ai/blog/gradium-tts-latency-and-accuracy
  4. Gradium pricing: gradium.ai/pricing
  5. Gradium data residency: docs.gradium.ai/guides/data-residency
  6. Gradium, "EU data residency and Zero Data Retention", September 2, 2026: gradium.ai/blog/eu-data-residency-zero-data-retention

This is the one-vendor stack page in the Text-to-Speech selection cluster:

Beyond this topic

A one-vendor stack still needs orchestration around it: best API to build an AI voice agent covers the frameworks, turn-taking and VAD covers the stage between the two halves, and cascaded voice agents vs speech-to-speech covers whether you need two halves at all.

Gradium's free tier includes 45,000 credits per month with no credit card. API reference and quickstarts are at docs.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions