
Free Tier
Run up to 1,000 requests per month for free
Free/month
- 45,000 credits
- 4hrs
- 3 concurrency
Speech-to-text
Streaming transcription that knows when your user is done talking, and learns your vocabulary. Semantic turn detection, keyword boosting and tunable latency over one WebSocket.
Knows when you've finished
Listens to your pauses and your words, so the agent replies right when you're done
Speed you can tune
Shorter for a fast live agent, longer for the most accurate transcript
Custom vocabulary
Add product names, people and jargon. No retraining, and each word's odds can rise ~20x
No credit card required
A pause to think sounds exactly like the end of a sentence, so systems that only listen for silence cut people off. Ours listens to what is being said, so your agent waits for people to finish, then answers without an awkward gap.
Do you like any other… uh… other animals or anything?
I gotta figure out this… uh… packing tape.
Paste one of these prompts into your coding agent and it will read the guide, install the SDK and open the stream. Nothing else to do, you're ready to go.
Add Gradium real-time speech-to-text to this project.
Read https://docs.gradium.ai/guides/speech-to-text first and follow it over
anything you assume. Use the official Gradium SDK for this codebase's
language. Read the API key from an environment variable, never hardcode it.
Then implement the flow, in this order:
1. Open an stt_realtime context manager with the input format and a
json_config holding the language and delay_in_frames. Use 16 frames unless
I ask for something more reactive.
2. Push 80 ms PCM frames from the audio source in a producer task and call
send_eos when it ends.
3. Iterate the session in a consumer task. Accumulate text messages into the
current turn, and close the turn when a step message shows the inactivity
probability crossing 0.5 on the 1 s, 2 s or 3 s horizon. Make the horizon
and the threshold configurable.
4. When the app decides a turn is over before VAD does, call send_flush with
an ID and wait for the matching flushed message before treating the
transcript as final.
Keep it small: one module, typed, with clear errors when frames are the wrong
size, when the socket drops, or when the key is missing. Do not add a UI
unless I ask for one.
Push-based, for microphones and telephony
Latency stays flat as concurrency climbs, flexible deployment that fits your setup.
Real-time API
Bidirectional WebSocket streaming. Start sending before ready returns; first audio in ~200 ms. Telephony formats built in.

in · pcm 24 kHz
CLIENT/AGENT

Gradium API
wss://api.gradium.ai/
api/speech/s2s
out · pcm 48 kHz
VOICE ENGINE
Predictable & secure
Flat latency under load, enterprise SLAs, zero data retention, ISO 27001 certified.
I’d like to move my appointment.
Of course. Which day works?

I’d like to move my appointment.
Of course. Which day works?

Runs where you need it
Cloud API, dedicated instances, self-hosted, on-prem for data sovereignty, and cloud-provider marketplaces.
Every Gradium model speaks every language we support, with the same quality. Regional variations and accents included, with seamless mid-sentence code-switching without latency.
The questions teams ask before putting live transcription in front of real users.
No credit card required