Voice Design
You know how the voice in your head sounds. Voice Design turns that into a sentence and the sentence into a voice id, in seconds. Here is what a prompt can say, what each part changes, and how to get from close to right.
A prompt that shipped
A British female voice, 20 to 30, glossy and confident, with girly chatter, high pitch, fast pacing, high energy and bright sparkling resonance. Ideal for a friendly receptionist or assistant.
This one produced Freya, a voice in the library today.
What it produced
Eleven things, and no prompt needs all of them. Accent, age range and the three dials carry most of the result; the rest is tuning once the voice is close.
Freya's prompt with each clause named, so you can see which one to change. Every clause below is doing something you can hear.
| Clause | This prompt |
|---|---|
| Accent / locale | British |
| Gender | female |
| Age range | 20 to 30 |
| Character | glossy and confident |
| Pitch | high pitch |
| Pace | fast pacing |
| Energy | high energy |
| Resonance / timbre | bright sparkling resonance |
| Modifiers | girly chatter |
| Applications | a friendly receptionist or assistant |
Read it as a sentence and it is one line of description. Read it as clauses and it is a set of dials. That is the whole method: the sentence is how you write it, the clauses are how you change it.
The first voice is rarely the one you ship, and that is the normal way round. Four habits that make the second one better than the first.
Each accent has its own page: what makes it sound like itself, the words that expose it, and prompts that produce it. Start from the one you need rather than from a blank box.
One or two sentences. The limit is 500 characters and most good prompts use half of it. Length is not what makes a prompt work: naming the accent, an age range and the three dials in one sentence beats a paragraph of adjectives, because adjectives past the fourth or fifth start pulling against each other.
Yes, and give it as a numeric range rather than a word. "20 to 30" and "60 to 75" both work, where "young" and "old" underperform because they leave the boundary open. Age and pitch tend to move together, so a prompt asking for gravity without moving the age range usually returns a young voice performing gravity.
Yes. A designed voice is a normal voice id with the same endpoints, the same SDK and the same latency as any library voice, so nothing about your pipeline changes once you have it. If a library voice runs in your agent today, a designed one will too.
Almost always because more than one clause changed at once. Move pitch, or pace, or the stated purpose, on its own and listen again. Change three together and the result is a new voice rather than a version of the one you had, which is useful when you are exploring and frustrating when you were close.
No. Write the prompt in English and name the language and accent inside it. "A Carioca Brazilian Portuguese female voice, 30 to 40" is an English prompt that produces a Portuguese voice.
Between one and five, and more than one is usually worth it. The same prompt describes a space rather than a point, so a batch of three gives you something to choose between and shows you which parts of your prompt are doing the work.
One that exposes the accent rather than a neutral sentence any accent could read. Every accent page here lists the words and vowels that give that accent away, and a line built from two or three of them tells you in five seconds what a paragraph of general copy will not.
Yes. There is no actor to license and no recording session behind it, which is the practical reason teams design a voice rather than cast one: the voice is yours to ship, and you can make a second one that matches it next week.
One sentence, a few seconds, and a voice id you call from the API or the SDK. No recording session, no actor licensing, and a second voice that matches it whenever you need one.