NewOur fastest Text-to-Speech model yet with 1000+ voices ・ Sub-50ms latency

Tutorial

Instant vs Pro Voice Cloning in Gradium: When to Use Each

4 min read
Updated

By Gradium. Data as of September 2026.

Key takeaways

  • Instant Voice Clone: 10 seconds of clean speech, a voice ID in seconds. 5 clones on the free tier restricted to non-commercial use, 1,000 per month on paid plans.
  • Pro Voice Clone: 30 minutes of clean audio minimum, 2 hours recommended, available from the M plan up. One-time training fee of 1M credits plus 1.2 credits per character thereafter.
  • The tradeoff is not similarity, it is range. Instant gets close to the source; Pro holds identity across long passages and across emotional registers.
  • Both are called the same way from the same API, so moving from one to the other is a change of voice ID, not a rewrite.
  • In a blinded study over 3,220 voice pairs across English, French, Spanish and German published January 22, 2026, Gradium held the highest live Elo in every language from a 10-second source sample, against ElevenLabs Flash.

What is Instant Voice Cloning?

Upload 10 seconds of clean speech, get a voice ID back within seconds, start generating. No training step, no queue.

The output sounds very close to the source. What you trade for the speed is range: emotional variation and stability over long passages are more limited than a Pro clone. For a ten-second turn in a support conversation, that is invisible. For a forty-minute chapter, it is not.

Where Instant is the right answer:

  • Prototyping and demos, where the question is whether the concept works
  • Internal tools
  • Generating many voices quickly to compare styles before committing to one
  • Per-user voices at a volume where training each one is not viable
  • Any product where the voice speaks in short turns

What is Pro Voice Cloning?

Pro finetunes a model on your speaker. That is why it needs more audio and more time, and why the result holds together where Instant drifts.

Where Pro is the right answer:

  • A branded voice that is part of the product identity
  • Scripts that swing between registers, where one delivery has to cover all of them
  • Anything where prosody and emotional range are the point
  • Podcasts, audiobooks and long-form narration
  • Customer-facing agents where voice quality carries trust

Preparing audio for a Pro clone

The minimum is 30 minutes. Gradium recommends one to two hours, and the difference between the two is audible.

Recording quality matters more than quantity past the minimum. A quiet room, a consistent microphone position, no compression artifacts, no background music. Thirty clean minutes beats two noisy hours.

The content of the script is irrelevant. What is being captured is how the voice behaves: tone, prosody, rhythm, pitch, and how it expresses surprise, fear, happiness and empathy. So the speaker should deliberately range across styles and emotions rather than reading one flat passage. A recording of a single register produces a clone that can only do that register.

How do they compare?

Dimension Instant Voice Clone Pro Voice Clone
Audio required 10 seconds of clean speech 30 minutes minimum, 1 to 2 hours recommended
Time to a usable voice Seconds Finetuning run
Similarity to source Very close Very close, and holds under pressure
Emotional range Limited Wide
Long-form stability Can drift Maintains identity over long durations
Availability Free tier (5, non-commercial) and all paid plans (1,000 per month) M plan (5 clones) and L plan (20 clones)
Cost Included in the plan 1M credits one-time training, then 1.2 credits per character
Best for Prototypes, per-user voices, short turns Branded voices, audiobooks, customer-facing agents
API Same endpoint, same parameters Same endpoint, same parameters

Source: gradium.ai/pricing and the pricing FAQ, read September 8, 2026.

Note the cost line, because it changes the arithmetic at volume. A Pro clone bills at 1.2 credits per character against 1 credit for a standard or Instant voice, so a Pro voice costs 20% more per character to run, on top of the one-time training fee. At audiobook volumes that is worth modelling before committing. See how to compare TTS pricing across providers.

How do you use either one?

Identically. The clone ID goes in voice_id, and nothing else about the call changes:

import asyncio
import gradium

async def main():
    client = gradium.client.GradiumClient()

    result = await client.tts(
        setup={
            # An Instant clone ID and a Pro clone ID are interchangeable here.
            "voice_id": "<CLONE_ID>",
            "model_name": "default",
            "output_format": "wav",
            "json_config": {
                "rewrite_rules": "en",
                # Similarity to the target voice. 2.0 to 3.0 covers most
                # cloning work; higher tracks the source more tightly and,
                # past a point, introduces artifacts.
                "cfg_coef": 2.5,
            },
        },
        text="Thanks for calling. Your order ships on Friday.",
    )

    with open("output.wav", "wb") as f:
        f.write(result.raw_data)

asyncio.run(main())

That interchangeability is the practical argument for starting on Instant. Prototype with an Instant clone, validate the experience, and swap in a Pro clone ID when the product is worth the recording session. Nothing else in the integration moves.

cfg_coef is the parameter to reach for if a clone is not tracking the source closely enough. Range 1.0 to 4.0, default 2.0, with 2.0 to 3.0 covering most cloning work. Higher values increase similarity and eventually introduce artifacts, so raise it in small steps and listen. Full parameter reference in json_config.

Which should you choose?

Three questions settle it.

How long does the voice speak without a break? Short turns favour Instant. Anything over a few minutes of continuous speech is where Pro's stability earns its cost.

How many registers does one voice have to cover? A single calm register is within Instant's range. A voice that must be warm, then apologetic, then brisk in the same script needs Pro.

How many distinct voices do you need? Pro clones are capped by plan: 5 on M, 20 on L, unlimited on Enterprise. A product giving every user their own voice is an Instant product by construction, at 1,000 per month on paid plans.

The common path is both: Instant for experimentation and to shortlist voices, Pro for the one that ships. Iterating with Instant clones is also the cheapest way to audition a range of speakers before booking a recording session for the winner.

One thing to settle before either: licensing. Gradium's free-tier Instant clones are restricted to non-commercial use. Confirm you have the speaker's consent and the rights you need before a cloned voice reaches production.

What the blind study actually measured

The cloning comparison Gradium published on January 22, 2026 is worth reading precisely, because "highest Elo in every language" is a claim whose scope matters.

The setup: 3,220 voice pairs, blinded A/B with a live Elo score, across English, French, Spanish and German. 890 sentences per language at three complexity levels, 20 voices per language, each clone built from a 10-second source sample. Gradium held the highest Elo in every one of the four languages.

Two limits to carry with the number. The comparison was against ElevenLabs Flash only, not against a field, so it says nothing about Cartesia or Inworld. And it tested Instant clones: 10 seconds in, which is the harder condition and the one that makes the result interesting, but not a measurement of Pro output.

The architecture behind it is described in Why cloned voices sound fake: cross-attention layers attending to the speaker recording rather than prefix conditioning, which is what preserves micro-traits such as vocal fry, rasp, breathiness and pitch dynamics, even when switching languages or scripts.

Glossary

Instant Voice Clone. A voice created from about 10 seconds of audio with no training step, usable within seconds.

Pro Voice Clone. A voice created by finetuning on 30 minutes or more of clean audio, with wider emotional range and better long-form stability.

Voice ID (voice_id). The identifier of a stored voice, passed on the setup message. Instant and Pro clone IDs are interchangeable at the call site.

cfg_coef. The json_config parameter controlling similarity to the target voice. Range 1.0 to 4.0, default 2.0.

Speaker similarity. How close a clone is to its source, usually measured as cosine distance between speaker embeddings.

Emotional range. How far a voice can move between registers while still sounding like the same speaker. The main axis on which Pro exceeds Instant.

Elo (voice cloning). A preference score from blind pairwise votes comparing cloned samples. Preference, not similarity measured numerically.

Part of the voice cloning cluster, hub at best voice cloning APIs 2026. Its siblings:

Beyond this topic

A cloned voice still has to read the text correctly, which is a separate problem: pronunciation dictionaries and text normalization edge cases cover it, and TTS WER benchmark 2026 covers how it is measured.

Gradium's free tier includes 45,000 credits per month with no credit card. Try cloning in the browser at studio.gradium.ai. For enterprise evaluations, use our contact form.

Frequently Asked Questions