Voice Design
Imagine who is speaking and hear them in seconds.
4.06 / 5
Accent adherence
Best of four voice design systems tested.
83.4%
Prompt adherence
On the InstructTTSEval English benchmark, outperforming all previously reported systems.
100+
Accents
From Colombian Spanish to Québécois French or Swiss German, we cover all the regions.
Worldwide regional accents and perceived age are both fully controllable.
Voice Design is free and comes with every plan. Open the studio and design your first voice today.
See pricingDescribe the speaker: age, gender, accent, texture, mood. Every prompt goes through moderation, and use is governed by our Terms of Service, in particular Section 3.4, Usage Restrictions.
Read the Terms of ServiceYou retain rights to your inputs and to the audio you generate, and warrant that you hold the necessary rights to any content you provide and generate. Gradium retains all rights in and to its models. Any voice that resembles an identifiable person is subject to our Terms in their entirety, and in particular to Section 3.4.
Read who owns whatFree Tier
Run up to 1,000 requests
per month
Free /month
45k credits
~1hr TTS
5 custom voices
Start for freeDescribe the voice you want
A British narrator in his fifties, warm and unhurried, with a dry wit… or Get inspired
Abigail
Adult · Female · General American
A warm and airy American adult voice that adds a touch of empathy to any story.
KRo-uwfno-KcEgBM
Adam
Adult · Male · General American
A joyful and smooth American adult voice that greets the morning with radio-ready energy.
EbIA5CIcQoa6NNd2
Alex
Adult · Male · General American
A joyful high-pitched American adult voice that grabs attention in advertisements.
91EdXxJDbWICDBgz
Alexander
Adult · Male · General American
A warm low-pitched American adult voice that motivates with the resonance of a fitness coach.
8sWSyTC7byLsbHkr
Alexandra
Adult · Female · General American
A joyful and smooth American adult voice ideal for reading and hosting duties.
4nAcNUlNhEA_Kyjo
Alexis
Adult · Female · General American
A joyful American adult voice that delivers customer service scripts with a bright tone.
74asmf7CXzjfopIX
Paste this into your coding agent. It describes the voice, listens to the candidates, keeps the one you want, then uses its voice ID with the same TTS endpoint you already call.
Add Gradium voice design to this project.
Read https://docs.gradium.ai/guides/voices/voice-design first and
follow it over anything you assume. Use the official Gradium SDK for this
codebase's language. Read the API key from an environment variable, never
hardcode it.
Then implement the flow, in this order:
1. Take a description of a speaker (age, gender, accent, texture, mood) and
a language code. Ask for three candidates and wait until they are ready.
2. Render each candidate reading one short line, so I can listen and pick.
Save the previews as WAV files named after the candidate.
3. Keep the candidate I choose and give the voice a name. Store the returned
voice ID where the rest of the app can read it.
4. Synthesise text with that voice ID through the standard TTS endpoint,
exactly as you would with any voice from the catalog.
Keep it small: one module, typed, with clear errors when a candidate is not
ready yet or the key is missing. Do not add a UI unless I ask for one.Three builds that take a designed voice into a product. Each one runs today, and the code is public.

Design the voice to match the face, render the line with Gradium TTS, and Pruna's video-avatar model animates the still. A finished clip, not a live call.
A voice, taking shape
Ready for a description.
A speech-first Gradbot agent asks what the voice should sound like, designs a candidate and switches the live call to it. The next question is the audition.
LIVEGradium for the voice, the ears and the speech, LiveKit for the realtime session, LemonSlice for the animated face. An image, a voice description and a role is all it takes.
The catalog covers five languages and thirteen regional variants, and every voice in it cleared listener preference testing rather than internal taste.
No. Each run gives you a few candidates to listen to, and they may differ each time. But once you like one and decide to keep it, that voice is fixed and will return the same speaker every time. So if you need the same voice again later, keep it rather than running the prompt again.
Cloning reproduces a speaker who exists, from about 10 seconds of their audio, and needs their consent. Design invents a speaker from a description, so there is no recording to provide and no speaker to clear. The same Terms of Service apply to both.
English, French, German, Spanish and Portuguese, with regional accents in each. Accents are where the model is most precise, so it pays to be specific: try asking for Colombian rather than Latin American.
You retain the intellectual property rights over your input and the output you generate, and Gradium retains all rights in its models.
Usage remains governed by our Terms of Service, in particular Section 3.4 (Usage Restrictions), which prohibits generating content that infringes a third party's publicity or personality rights, or that is used to impersonate someone without their explicit authorization.
Read the Terms of ServiceYes. The licence Gradium grants allows commercial use of the audio you generate on all paid plans. The free plan and beta services are for internal, non-commercial use only (Terms of Service, Section 6.2).
For more information, the pricing page lists commercial use per plan.
See the pricing page