NewOur fastest Text-to-Speech model yet ・ Sub-50ms latency

Live translation

Hear them in your language,while they are still speaking.

Hold a conversation in a language you don’t speak.

Translate to

No credit card required

The numbers behind live translation

  • 20 pairs

    Every direction covered

    Any of five languages into any other, in any direction.

  • 400+ voices

    Any voice, or your own

    Every output voice in the catalogue, or a clone of your own.

  • 3.0 s

    End-to-end latency

    Average end-to-end latency across all language pairs.

No credit card required

Fewer models, faster translation

The standard stack transcribes, then translates the text, then speaks it. Ours translates as it listens so one whole model and its handoff leave the critical path.

Launching Gradium Translate

Models in the path

2

Average end to end

3.0 s

Intermediate transcript

None

No marketing spin,just code and hard numbers

Paste one of these prompts into your coding agent and it will read the guide, install the SDK and open the stream. Nothing else to do, you're ready to go.

Add Gradium live speech-to-speech translation to this project.

Read https://docs.gradium.ai/guides/speech-to-speech-overview first and
follow it over anything you assume. Use the official Gradium SDK for this
codebase's language. Read the API key from an environment variable, never
hardcode it.

Then implement the flow, in this order:

1. Open one s2s_realtime session with model_name "s2s-translate",
   stt_model_name "stt-translate", tts_model_name "default", the target
   language in json_config, and a voice_id that exists in that target
   language. voice_id is required, not optional.
2. Stream source audio in as PCM frames in one task and call send_eos when it
   ends.
3. Read the socket in a second task. Play audio messages as they arrive and
   collect text messages as a running translated transcript. Stop on
   end_of_stream.
4. Make the target language and the voice ID configurable per session, not
   per build, so a call can be re-pointed without a redeploy.

Keep it small: one module, typed, with clear errors when the voice ID does
not match the target language, when the socket drops, or when the key is
missing. Do not add a UI unless I ask for one.

Speech in, speech out

No distortion under load

Latency stays flat as concurrency climbs, flexible deployment that fits your setup.

Real-time API

Bidirectional WebSocket streaming. Start sending before ready returns; first audio in ~200 ms. Telephony formats built in.

in · pcm 24 kHz

CLIENT/AGENT

WEBSOCKET

Gradium API

wss://api.gradium.ai/
api/speech/s2s

BidirectionalTelephony ulaw / alaw~200 ms

out · pcm 48 kHz

VOICE ENGINE

Predictable & secure

Flat latency under load, enterprise SLAs, zero data retention, ISO 27001 certified.

1 conversationQuiet hour

I’d like to move my appointment.

50ms

Of course. Which day works?

10 000 conversationsPeak

I’d like to move my appointment.

50ms

Of course. Which day works?

Runs where you need it

Cloud API, dedicated instances, self-hosted, on-prem for data sovereignty, and cloud-provider marketplaces.

Predictableand scalable pricing

Native fluencyacross languages

Every Gradium model speaks every language we support, with the same quality. Regional variations and accents included, with seamless mid-sentence code-switching without latency.

Questions ?We’re here to help

The questions teams ask before putting live translation in front of real users.

Startbuilding

No credit card required