New

Voice Design. Live today.

Voice Design

Get a brand-new voice
from a short text prompt

Imagine who is speaking and hear them in seconds.

A calm woman in her forties, clearly audibleor

4.06 / 5

Accent adherence

Best of four voice design systems tested.

83.4%

Prompt adherence

On the InstructTTSEval English benchmark, outperforming all previously reported systems.

100+

Accents

From Colombian Spanish to Québécois French or Swiss German, we cover all the regions.

Any speaker you can describe, in every detail, faster than everywhere else.

A voice from anywhere, of any age

Worldwide regional accents and perceived age are both fully controllable.

0:00

Before you ship,
the practical part

Voice Design is free and comes with every plan. Open the studio and design your first voice today.

See pricing

One prompt,
From a sentence to a voice

Paste this into your coding agent. It describes the voice, listens to the candidates, keeps the one you want, then uses its voice ID with the same TTS endpoint you already call.

Add Gradium voice design to this project.

Read https://docs.gradium.ai/guides/voices/voice-design first and
follow it over anything you assume. Use the official Gradium SDK for this
codebase's language. Read the API key from an environment variable, never
hardcode it.

Then implement the flow, in this order:

1. Take a description of a speaker (age, gender, accent, texture, mood) and
   a language code. Ask for three candidates and wait until they are ready.
2. Render each candidate reading one short line, so I can listen and pick.
   Save the previews as WAV files named after the candidate.
3. Keep the candidate I choose and give the voice a name. Store the returned
   voice ID where the rest of the app can read it.
4. Synthesise text with that voice ID through the standard TTS endpoint,
   exactly as you would with any voice from the catalog.

Keep it small: one module, typed, with clear errors when a candidate is not
ready yet or the key is missing. Do not add a UI unless I ask for one.

Start from a working example

Three builds that take a designed voice into a product. Each one runs today, and the code is public.

  • 0:09 / 0:24

    Turn a portrait into a talking avatar

    Design the voice to match the face, render the line with Gradium TTS, and Pruna's video-avatar model animates the still. A finished clip, not a live call.

  • A voice, taking shape

    Ready for a description.

    Tap to talk

    Design a voice by talking to one

    A speech-first Gradbot agent asks what the voice should sound like, designs a candidate and switches the live call to it. The next question is the audition.

  • LIVE

    A live avatar agent you can call

    Gradium for the voice, the ears and the speech, LiveKit for the realtime session, LemonSlice for the animated face. An image, a voice description and a role is all it takes.

Discover +400 voices in our studio

Not every project needs a new voice

The catalog covers five languages and thirteen regional variants, and every voice in it cleared listener preference testing rather than internal taste.

Questions?
We're here to help