Describe the accent, gender, age and delivery in one sentence. Gradium returns a voice id you call from the API or the SDK, with no recording and no actor licensing. The prompt below is the one that produced Desmond, the voice in the sample.
American English, US English, Standard American, GenAm and general american all describe the same voice. Write it as American English in your prompt.
Hear it
“Imagine a problem so stubborn that it went unsolved for decades, and then, one afternoon, a graduate student notices something nobody else had ever thought to question.”
An American English male voice, 55 to 65: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.
This exact wording produced Desmond, a English · US voice now in the Gradium library, which is live in the library as bwRhQrJel4IuvxLF.
A prompt is built from parts: the accent, the age, then pitch, pace, energy and what the voice is for. Change one part and you get a variation. These four are starting points to adapt, not finished prompts.
The prompt as written. Low pitch, slow pacing and resonant timbre is the documentary register, and the gentle low-to-high flow is what stops it sounding like a recorded announcement.
An American English male voice, 55 to 65: clean, deliberate and precise, with low pitch, slow pacing, low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.
Copy this as it stands.
Everything to the middle. An agent interrupts and gets interrupted, so extremes on any axis become obvious across hundreds of turns in a way they never are in a single demo line.
An American English male voice, 30 to 40, for a real-time voice agent: clear and easy, with mid pitch, steady natural pacing, medium energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.
In colour: the 6 parts that make this style.
American tolerates high energy far better than British does, because its intonation range is naturally wider. This is the one accent where a 5 on energy still sounds like a person.
An American English male voice, 25 to 35: bright and punchy, with mid-high pitch, fast pacing, high energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority.
In colour: the 5 parts that make this style.
Slow it further than feels right and let the resonance carry the weight. Asking for more energy here is the usual mistake: trailer voices are quiet and close, not loud.
An American English male voice, 40 to 55: gravelled and deliberate, with low pitch, slow pacing, low-to-mid energy, deep chest resonance with light vocal fry and a gentle low-to-high flow. Ideal for projecting academic authority.
In colour: the 4 parts that make this style.
Each feature below is something the prompt has to encode, or something a careless prompt will accidentally undo.
Ready to call with no design step. Hear a few here, then open any in the studio or write a prompt of your own.
Adult · Female
A warm and airy American adult voice that adds a touch of magic and empathy to any story.
KRo-uwfno-KcEgBMAdult · Female
A joyful and smooth American adult voice ideal for reading and hosting duties.
4nAcNUlNhEA_KyjoAdult · Female
A joyful American adult voice that delivers customer service scripts with a bright distinct tone.
74asmf7CXzjfopIXAdult · Female
A joyful high-pitched American adult voice that brings high energy to training and teaching.
yU6yxQ3e8LKRwU84Another 74 General American voices are in the catalogue. Sign in to hear them all, or design one to your own spec from a prompt.
The method behind the prompt: what each clause does, what breaks it, and what it was tested on.
Each row is one part of the prompt above and what it is set to here. Change a single row and you get a variation on this voice; change several and you get a different one.
| Clause | This prompt |
|---|---|
| Accent / locale | American English |
| Gender | male |
| Age range | 55 to 65 |
| Character | clean, deliberate and precise |
| Pitch | low pitch |
| Pace | slow pacing |
| Energy | low-to-mid energy |
| Resonance / timbre | resonant timbre |
| Modifiers | a gentle low-to-high flow |
| Applications | projecting academic authority |
4 failure modes we hit repeatedly, and what fixes each one.
Descriptor words do not transfer between locales. Energetic and bright travel intact, but most descriptors resolve to a different sound depending on the language they are read in. A prompt that produces a warm American voice will not produce the equivalent warmth translated into French, because the brief has to be rewritten for the locale rather than translated.
Writing your first prompt? How to write a voice prompt covers every part a prompt can carry, and what each one changes.
"An American English male voice, 55 to 65: clean, deliberate and precise, with low pitch, slow pacing and low-to-mid energy, resonant timbre and a gentle low-to-high flow. Ideal for projecting academic authority." That is the exact prompt that produced Desmond, the voice in the sample above.
Yes. Searches for an American voice almost always mean General American, the broadcast-neutral variety with no regional marking. Southern American, New York and Californian are separate mapped labels and you get them by naming them.
79, spread across conversational, customer service, podcast and advertising work. They need no design step: pick a voice id and call the API. The American voice library page lists all of them with their ids.
Almost always because the prompt anchors to a British reference for some other quality, such as warmth or authority. American is rhotic and British is not, so a British reference pulls the r out mid-line. Keep every reference inside the target accent.
Both. Gender is its own part of the prompt, so it changes on its own: swap male for female and everything else holds. Same accent, same age, same delivery, different voice.
A designed voice is a normal voice id: same endpoints, same SDK, same latency as any library voice. Every voice we ship clears blind listener comparison and an ELO ranking first, because roughly one candidate in a thousand has it.
Open a voice, type your own script, and listen. Then write a prompt of your own to design a variant. No recording, no casting, no licensing.