NewOur fastest Text-to-Speech model yet ・ Sub-50ms latency

Speech-to-text

Speech-to-textbuilt for live conversation.

Streaming transcription that knows when your user is done talking, and learns your vocabulary. Semantic turn detection, keyword boosting and tunable latency over one WebSocket.

Your transcription will appear here...
Tap to record

The numbers that make an agent listen like a human

  • every 80ms

    Knows when you've finished

    Listens to your pauses and your words, so the agent replies right when you're done

  • 0.56s to 3.84s

    Speed you can tune

    Shorter for a fast live agent, longer for the most accurate transcript

  • up to 500 words

    Custom vocabulary

    Add product names, people and jargon. No retraining, and each word's odds can rise ~20x

No credit card required

A pause is not the end of a turn

A pause to think sounds exactly like the end of a sentence, so systems that only listen for silence cut people off. Ours listens to what is being said, so your agent waits for people to finish, then answers without an awkward gap.

Semantic VAD

Do you like any other… uh… other animals or anything?

I gotta figure out this… uh… packing tape.

No marketing spin,just code and hard numbers

Paste one of these prompts into your coding agent and it will read the guide, install the SDK and open the stream. Nothing else to do, you're ready to go.

Add Gradium real-time speech-to-text to this project.

Read https://docs.gradium.ai/guides/speech-to-text first and follow it over
anything you assume. Use the official Gradium SDK for this codebase's
language. Read the API key from an environment variable, never hardcode it.

Then implement the flow, in this order:

1. Open an stt_realtime context manager with the input format and a
   json_config holding the language and delay_in_frames. Use 16 frames unless
   I ask for something more reactive.
2. Push 80 ms PCM frames from the audio source in a producer task and call
   send_eos when it ends.
3. Iterate the session in a consumer task. Accumulate text messages into the
   current turn, and close the turn when a step message shows the inactivity
   probability crossing 0.5 on the 1 s, 2 s or 3 s horizon. Make the horizon
   and the threshold configurable.
4. When the app decides a turn is over before VAD does, call send_flush with
   an ID and wait for the matching flushed message before treating the
   transcript as final.

Keep it small: one module, typed, with clear errors when frames are the wrong
size, when the socket drops, or when the key is missing. Do not add a UI
unless I ask for one.

Push-based, for microphones and telephony

No distortion under load

Latency stays flat as concurrency climbs, flexible deployment that fits your setup.

Real-time API

Bidirectional WebSocket streaming. Start sending before ready returns; first audio in ~200 ms. Telephony formats built in.

in · pcm 24 kHz

CLIENT/AGENT

WEBSOCKET

Gradium API

wss://api.gradium.ai/
api/speech/s2s

BidirectionalTelephony ulaw / alaw~200 ms

out · pcm 48 kHz

VOICE ENGINE

Predictable & secure

Flat latency under load, enterprise SLAs, zero data retention, ISO 27001 certified.

1 conversationQuiet hour

I’d like to move my appointment.

50ms

Of course. Which day works?

10 000 conversationsPeak

I’d like to move my appointment.

50ms

Of course. Which day works?

Runs where you need it

Cloud API, dedicated instances, self-hosted, on-prem for data sovereignty, and cloud-provider marketplaces.

Predictableand scalable pricing

Native fluencyacross languages

Every Gradium model speaks every language we support, with the same quality. Regional variations and accents included, with seamless mid-sentence code-switching without latency.

Questions ?We’re here to help

The questions teams ask before putting live transcription in front of real users.

Startbuilding

No credit card required