BetaOur fastest Text-to-Speech model yet ・ Sub-50ms latency

Giving voice agents a voice that adapts to context

Three classical marble statues against a dark, speckled background, one of them winking and sticking out its tongue, surrounded by floating voice prompt chips reading "stern and professional", "calming concern", "sad and hopeless" and "unapologetically gen-z".
9 min read

Most voice APIs come either with a library of predefined voice presets or the ability to create a custom voice from an audio sample.

With Voice Design, you can now create natural, realistic voices from a text description (or "prompt"). Voice characteristics become output your code computes, from the same context it already uses to decide what the agent says.

Reading that context has also become inexpensive. A new class of decision models, such as Jev from TypeSafe AI, classifies a message against labels you define for about $0.00002 per request (in the following demo), with no training. This opens use cases that a fixed catalog of voices cannot serve well:

  • a game that gives every procedurally generated character a voice of its own,
  • an interactive story whose narrator changes with the scene,
  • a support agent that answers inquiries with an appropriate tone.

Consider this latter case. You are building an automated voice agent for the customer support line. Most requests are routine, like asking for an address change. Now one ticket arrives: "I have been charged twice and nobody has answered my emails."

A single voice answers both in the same tone, whether the caller wants a quick confirmation or needs reassurance. You could map each type of request to a prebuilt voice, but you would have to anticipate every case in advance, for every language the line supports, and find a preset with the right delivery for each. With Voice Design, you describe the delivery you need, such as "calm and reassuring, warm rounded resonance", and the agent creates the voices it is missing in the background as new cases appear.

In this article we build an end-to-end pipeline for that agent:

  1. Extract a few signals from an incoming message with Jev
  2. Map those signals to a Voice Design description
  3. Resolve the description to a voice the agent can speak with now, and design new voices in the background.
  4. Synthesize the reply with Gradium TTS.

You can follow along in the live demo (bring your own Gradium API key), or run it yourself from the GitHub code repository.

Prerequisites:

  • Node.js 18+
  • Access to Jev, either directly through TypeSafe AI or via OpenRouter (without a Jev key, a simple keyword classifier stands in)
  • A Gradium API key. Each voice converted from Voice Design uses one custom-voice slot on your account, shared with voice clones, so check your allowance before generating a library.

Note: the live demo is a simplified version of the pipeline described below. It creates each voice the first time it's needed and reuses it afterwards, so the first reply in a new context takes a few extra seconds. The full pipeline avoids that wait by preparing voices ahead of time.

The pipeline

The context-to-voice pipeline An incoming message goes through context extraction to a voice spec, then to a decision: is there an approved voice for this spec? Yes goes straight to TTS with a voice_id. No goes to the nearest approved voice, which also reaches TTS, and a dotted branch enqueues the spec for Voice Design, which generates, auditions and converts a voice, adding it to the voice library that feeds the decision. yes no enqueue adds voice BACKGROUND WORKER Incoming message Context extraction Voice spec Approved voice for this spec? Nearest approved voice TTS with voice_id Voice Design: generate, audition, convert Voice library

The solid path runs on every turn and never calls Voice Design. The dotted path runs outside the conversation and grows the library over time.

Step 1: Extract context with a decision model

The first step turns an incoming message into a few structured signals. This is a classification task, and decision models such as Jev are particularly well suited to it: they answer typed questions about an input and return a probability for each option and a confidence score. Three properties make it a good fit for a voice pipeline:

  • Fast. In our tests, a request took between 240 and 850 milliseconds including the network round trip. That is short enough to run on every incoming message.
  • Cheap. Jev costs $0.042 per million input tokens on OpenRouter, and output is free.
  • No training required. The labels and their descriptions are written in plain language in the request. Changing a label set only means editing an object in your code.

Jev is available through OpenRouter as typesafe/jev-1.13. Because the request carries an API key, this code runs on the server. Tone and energy become two choice questions sent in a single request:

javascript
const QUESTIONS = {
  tone: {
    type: "choice",
    instructions: "How does the writer of this message feel?",
    criteria: {
      frustrated: "Annoyed, upset, or dissatisfied.",
      neutral: "Matter-of-fact, no strong emotion.",
      enthusiastic: "Excited, happy, or eager.",
    },
  },
  energy: {
    type: "choice",
    instructions: "How urgent is this request?",
    criteria: {
      relaxed: "Can be handled at a normal pace.",
      urgent: "Needs prompt attention.",
    },
  },
};

const DEFAULTS = { tone: "neutral", energy: "relaxed" };
const MIN_CONFIDENCE = 0.6;

export async function extractContext(message, { language, register }) {
  const res = await fetch("https://openrouter.ai/api/alpha/decisions", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ model: "typesafe/jev-1.13", state: message, questions: QUESTIONS }),
  });
  if (!res.ok) throw new Error(`Jev request failed: ${res.status}`);
  const { answers } = await res.json();

  const context = { language, register };
  for (const [field, answer] of Object.entries(answers)) {
    context[field] = answer.confidence >= MIN_CONFIDENCE ? answer.choice : DEFAULTS[field];
  }
  return context;
}

For a support line, the message "I have been charged twice and nobody has answered my emails." comes back as:

json
{
  "language": "en",
  "register": "formal",
  "tone": "frustrated",
  "energy": "urgent"
}

The three example messages resolve as expected, with the same results over repeated runs (values in parentheses are the probabilities of the selected options):

Message tone energy
"I have been charged twice and nobody has answered my emails." frustrated (1.00) urgent (1.00)
"Hello, I moved last month and would like to update the delivery address on my account." neutral (1.00) relaxed (0.99)
"Thank you so much for your support!!!" enthusiastic (1.00) relaxed (1.00)

Three rules to keep the extraction reliable:

  • Keep the label sets small and closed. Every value has to map to a voice later, so each new label multiplies the number of voices you need.
  • Fall back to a default below a confidence threshold. Jev reports a confidence score separate from the option probabilities. When it is low, a safe default such as "neutral" is better than a guess.
  • Prefer known properties over inferred ones. language and register are passed in rather than classified: the language comes from the user's account settings, and the register from the product, since a support line answers formally even when the customer writes casually.

Note that the classifier here could be replaced by local models if messages cannot leave the browser or your infrastructure.

Step 2: Turn signals into a voice description

We recommend describing a voice with concrete attributes: gender, age range, accent, pitch, pace, energy, timbre, manner, and intended use. Some of those should vary with context, but not all. Split the description in two:

  • Persona: who the agent is. Gender, age, accent. Chosen by the product, fixed per language.
  • Delivery: how the agent sounds in this interaction. Pace, energy, manner, resonance. Derived from context.
javascript
const PERSONA = {
  en: "A female voice, 30 to 40, neutral American accent, mid pitch",
  fr: "A female voice, 30 to 40, Parisian accent, mid pitch",
  de: "A female voice, 30 to 40, standard German accent, mid pitch",
};

const DELIVERY = {
  tone: {
    frustrated: "calm and reassuring, warm rounded resonance",
    neutral: "clear and direct, balanced resonance",
    enthusiastic: "bright and upbeat, lively resonance",
  },
  energy: {
    relaxed: "unhurried pacing, medium energy",
    urgent: "steady, efficient pacing, focused energy",
  },
  register: {
    formal: "Ideal for a professional support line.",
    casual: "Ideal for a friendly in-app assistant.",
  },
};

export function voicePrompt({ language, register, tone, energy }) {
  const persona = PERSONA[language];
  if (!persona) throw new Error(`No persona for language "${language}"`);

  const prompt =
    `${persona}: ${DELIVERY.tone[tone]}, ${DELIVERY.energy[energy]}. ` +
    DELIVERY.register[register];

  if (prompt.length > 500) throw new Error("Voice description exceeds 500 characters");
  return prompt;
}

For the double-charge message, this produces:

"A female voice, 30 to 40, neutral American accent, mid pitch: calm and reassuring, warm rounded resonance, steady, efficient pacing, focused energy. Ideal for a professional support line."

Notice how the agent's voice responds to the user's tone instead of mirroring it: a frustrated customer gets a calm voice, and an urgent request gets efficient pacing rather than rushed delivery. The mapping encodes what the agent should do about the user's state, which is a product decision you can review in code.

Another important property is that the user's message never reaches the prompt; only values from closed lists do. That makes every possible prompt enumerable, and guarantees that nothing a user types can steer the voice beyond the options you allowed.

Step 3: Resolve a voice without waiting for one

Voice Design doesn't return a production voice directly. It returns candidates you audition before converting one to a permanent voice_id that can be used for TTS.

So the context-to-voice pipeline needs to create a library of approved voices keyed by spec, and a resolver that always returns one of them immediately.

This is why the spec space is deliberately left small: 3 personas x 2 registers x 3 tones x 2 energy levels = 36 combinations (60 if you cover all five languages Voice Design supports), and in practice you'd approve only the ones your product uses. That's few enough to design and audition by hand before launch.

For combinations you didn't anticipate, the resolver picks the closest approved voice in the same language and flags the miss:

javascript
const WEIGHTS = { tone: 2, energy: 1, register: 1 };

export const specKey = (s) => `${s.language}:${s.register}:${s.tone}:${s.energy}`;

export function resolveVoice(spec, library) {
  const exact = library.get(specKey(spec));
  if (exact) return { voiceId: exact.voiceId, exact: true };

  let best = null;
  for (const entry of library.values()) {
    if (entry.spec.language !== spec.language) continue;
    const distance = Object.entries(WEIGHTS)
      .reduce((d, [field, w]) => d + (entry.spec[field] === spec[field] ? 0 : w), 0);
    if (!best || distance < best.distance) best = { ...entry, distance };
  }
  if (!best) throw new Error(`No approved voice for "${spec.language}"`);
  return { voiceId: best.voiceId, exact: false };
}

Tone carries more weight than energy or register because getting it wrong is the most audible mistake: a bright voice answering a complaint is worse than a slightly slow one answering an urgent question.

The per-turn handler never awaits Voice Design:

javascript
export function voiceForTurn(spec, library, designQueue) {
  const { voiceId, exact } = resolveVoice(spec, library);
  if (!exact) designQueue.add(spec); // designed later
  return voiceId;
}

Designing voices in the background

The specs added to designQueue are processed by a worker: a background job, separate from the conversation, that takes one spec at a time and turns it into an approved voice. For each spec, the worker generates candidate voices from the spec's prompt, waits until they are ready, then auditions them and converts the one that passes. The first two steps look like this:

javascript
const BASE = "https://api.gradium.ai/api";
const headers = {
  "x-api-key": process.env.GRADIUM_API_KEY,
  "Content-Type": "application/json",
};

async function generateCandidates(spec, n = 3) {
  const res = await fetch(`${BASE}/voice-generator/generate`, {
    method: "POST",
    headers,
    body: JSON.stringify({ prompt: voicePrompt(spec), language: spec.language, n_samples: n }),
  });
  if (!res.ok) throw new Error(`generate failed: ${res.status}`);
  const { embeddings } = await res.json();
  return embeddings.map((e) => e.embedding_id);
}

async function waitUntilReady(id, timeoutMs = 120_000, intervalMs = 500) {
  const deadline = Date.now() + timeoutMs;
  while (Date.now() < deadline) {
    const res = await fetch(`${BASE}/voice-generator/embeddings?embedding_id=${id}`, { headers });
    const { embeddings } = await res.json();
    if (embeddings[0]?.ready) return;
    await new Promise((r) => setTimeout(r, intervalMs));
  }
  throw new Error(`Candidate ${id} not ready after ${timeoutMs} ms`);
}

Generation typically takes about 3 seconds, so polling every 500 ms adds at most half a second to it. The timeout is required: a request with an out-of-range generation setting still returns 201, but its candidates never become ready. The two-minute bound follows the Voice Design guide.

Once candidates are ready, audition them on a short line (100 characters max) through the regular TTS endpoint, then convert the one you approve:

javascript
async function convert(candidateId, spec) {
  const res = await fetch(`${BASE}/voices/from-embedding`, {
    method: "POST",
    headers,
    body: JSON.stringify({
      voxium_embedding_id: candidateId,
      name: specKey(spec),
      description: voicePrompt(spec),
    }),
  });
  if (!res.ok) throw new Error(`convert failed: ${res.status}`);
  const { uid } = await res.json();
  return uid; // permanent voice_id, usable over REST and WebSocket
}

Whether a human approves the candidate or the worker converts the first one automatically is a product choice. For a support line, listen first, because the generated voice may not match the prompt exactly. For a game spawning hundreds of background characters, auto-conversion may be acceptable. Keep the custom-voice allowance in mind either way.

When generation latency is fine

Generation latency is only a problem inside a conversational turn. There are moments where a few seconds go unnoticed:

  • Before launch, designing the whole spec space.
  • At session start, when context is known before anyone speaks: a CRM record, a selected language, a character-creation screen. The agent greets in a default voice while the tailored one is prepared.
  • Between sessions, when the queue of cache misses is processed.

The design is the same in all three cases: voice selection runs on the hot path, voice creation runs wherever a few seconds are free.

Step 4: Speak

With a permanent voice_id, synthesis is an ordinary TTS request:

javascript
async function speak(text, voiceId) {
  const res = await fetch(`${BASE}/post/speech/tts`, {
    method: "POST",
    headers,
    body: JSON.stringify({ text, voice_id: voiceId, output_format: "wav", only_audio: true }),
  });
  if (!res.ok) throw new Error(`TTS failed: ${res.status}`);
  return res.arrayBuffer();
}

A live agent would use the WebSocket endpoint with the same voice_id.

Keep voice changes deliberate

A few rules to keep in mind for a clean context-aware voice:

  • Pin the voice per session. Resolve once at the start of a conversation, or when something meaningful changes: the user switches language, sets a preference, or the conversation moves from small talk to a billing dispute. An agent that changes its voice every turn sounds like a relay of different people.
  • Never infer a persona from the user. The classifier only decides how the agent delivers its replies. Don't derive the agent's gender, age, or accent from what you guess about the user. Persona belongs to the product and, where offered, to the user's explicit settings.
  • Let explicit preferences win. If a user has chosen a voice, that choice overrides anything the classifier says.

Voice is now a function of context

The pipeline in this article is deliberately simple: one message associated with four signals, which delivers a clean mapping. Real conversations are messier, and a production agent needs more context than a single message and firmer rules about when to change voice. But the core takeaway holds: once a voice can be described in text, your application can decide how it should sound.

A common complaint about voice agents is that they sound robotic. Part of that comes from the voice model, and part comes from using one tone for everything: to confirm an address change and to answer someone who was charged twice.

Voice Design moves that choice into your code. The agent's voice becomes a mapping from context to voice, built from ordinary application logic you can test, review, and change. Adding a language or a new kind of conversation means adding a description, not recording a voice or searching a preset catalog.