Service

AI voice cloning: a familiar voice over a native base

Cloning does not replace the voice that speaks the language — it shapes it. The base voice supplies the pronunciation, the clone supplies the familiarity, and there is a limit to how far that goes.

Cloning is often sold as "your voice, speaking any language". That is not what happens, and the real mechanism is worth understanding before you spend credits on it.

A dub starts from a base voice that already speaks the target language properly — native pronunciation, correct stress, the right rhythm for that language. Cloning is applied on top of that base: it pulls the result toward something warmer and more familiar, closer to a person than to a narrator. What it does not do is transplant your identity wholesale into Hindi or Arabic.

It costs 8 credits a minute on top of whichever dubbing tier you choose. That premium is worth paying in some projects and wasted in others, and the rest of this page is about telling those apart — including the cases where the base voice alone is the better result.

What cloning actually does, and what it does not

What it does: shapes an already-native-sounding voice toward something more human and more familiar. The base voice handles the language; the clone handles the character. The result is a track that sounds less like a competent narrator and more like a person the audience has met before.

What it does not do — three things, and all three matter:

That third point is why cloning is not a substitute for reviewing the translation. A cloned voice reading a stiff, over-literal line sounds exactly like a person having a bad day, which is worse than a neutral base voice reading the same line.

Worth saying plainly: on many projects the base voice on its own is the better result. A well-cast native voice with no cloning is clean and unremarkable, which is exactly what you want for narration. Cloning is for when unremarkable is not the goal.

Cloning a voice is not a technical decision, it is a permissions decision. Every speaker in a project carries their own consent record, stored per speaker rather than per account, and the pipeline will not generate a cloned track for a speaker who does not have one.

This matters more than it sounds in three common situations:

You can also mark a speaker as skipped, which leaves them on a library voice while the rest of the project is cloned. That is the normal case for anything with more than one person on the microphone.

When cloning earns its credits

SituationClone?Why
A creator whose face and voice are the channelYesThe voice is the brand; a stock voice reads as a different channel
A founder or spokesperson in a company videoYesRecognition carries across markets, and re-recording is not an option
Course or training material in one voiceUsuallyConsistency across dozens of modules matters more than any single line
Narrated documentary or explainer, no on-screen hostRarelyA well-cast library voice does the same job for 8 fewer credits a minute
Product demos and screen recordingsNoNobody is listening for whose voice it is
Multi-speaker interviewsSelectivelyClone the host, cast the guests — consent and cost both point the same way

The pattern: clone when the audience already knows the voice. If they do not, you are paying a premium to reproduce a voice nobody recognises, and a well-chosen voice from the library will serve you better.

The accent question cloning does not answer

A cloned voice speaking Spanish still has to pick a Spanish. Cloning carries your timbre; it does not decide whether your audience in Bogotá hears a Castilian or a Colombian register, and that decision affects reception more than the timbre does. The Spanish guide covers the nine regional options, and the case-study data covers what getting it right is worth.

The same applies inside Arabic, where thirteen dialects behave as different registers rather than different accents, and inside Portuguese, where Brazil and Portugal are two decisions rather than one.

What a cloned dub costs

Cloning is an addition to the dubbing rate, not a replacement for it. Everything runs off one credit balance:

A ten-minute video dubbed at Standard with cloning is 330 credits — the 250 for the dub plus 80 for the clone. On the $50 plan that is roughly $14 against $11 without it.

The economics improve sharply with a second language. Transcription and the master timeline happen once, so the second market costs the dubbing pass plus the clone premium over the same timecodes. Cloning a channel into four languages is closer in price to two than to four. The pipeline breakdown has the full table.

Credits top up from $10 with no subscription, which is enough to clone one short video and hear the result before deciding whether it is worth doing at scale.

What makes a clone sound wrong

1. A bad source recording. Cloning reproduces what it is given. Room echo, a noisy air conditioner and a microphone worn too far from the mouth all survive into the clone, and they are more noticeable in a language where the listener is working slightly harder to follow.

2. Length mismatch. Translations rarely run the same length as the original. If a line is generated without being held inside its timecode, the cloned voice either rushes or leaves a gap — and because the voice is recognisably yours, the artificiality lands harder than it would on a stock voice.

3. Loudness drift. A cloned track generated at one level and published next to your original audio gives the whole thing away. Every export is mastered to the destination's loudness spec so both tracks sit at the same level.

4. Skipping the review. The pipeline stops after transcription and again after translation, before any credit goes on voice. With cloning that checkpoint is worth more, not less: regenerating a cloned track costs the dubbing rate and the clone premium again.

Projects with more than one voice

Most guidance on cloning assumes one speaker. Real footage rarely is, and the multi-speaker case is where the decisions get interesting.

A two-host podcast. Both hosts are known to the audience, both need consent, and the value of cloning is high because listeners recognise the pair as a pair. This is the clearest case for cloning everyone.

An interview. The host is the constant; the guest appears once. Cloning the host and casting a library voice for the guest is usually right — it solves the consent question, halves the premium, and the audience is not expecting to recognise the guest's voice in another language anyway.

A panel or a roundtable. Four voices, four consents, four clone premiums. Here the honest advice is usually to clone nobody and cast four distinct library voices instead: what the viewer needs is to tell the speakers apart, and distinct casting achieves that for 8 fewer credits a minute per speaker.

Corporate video with a narrator and a spokesperson. Clone the spokesperson, whose face is on screen and whose identity carries; leave the narrator on a library voice, because nobody knows what the narrator sounds like.

The common thread is worth stating plainly: cloning buys recognition, and recognition only has value where it already exists. Speaker labels from the transcription stage carry through the whole pipeline, so these choices are made per speaker rather than per project.

Cloning against the alternatives

There are three ways to put your content into another language with your identity attached, and they are not equally sensible.

ApproachWhat it costsWhere it breaks
Learn the language and re-recordMonths per languageDoes not scale past one or two, and accent gives it away
Hire a voice actor per marketStudio rates, per language, per revisionScheduling; and the voice is not yours in any market
A native base voice, optionally cloned+8 credits a minute for the cloneResemblance rather than substitution; needs a clean sample

The comparison people actually make, though, is cloning against a good library voice — and there the answer depends entirely on whether the audience knows you. A channel with a face and a name loses something real when the voice changes. A tutorial channel with no on-camera presence loses nothing at all.

Worth noting what cloning does not replace: it does not lip-sync. If your video shows a face talking to camera and mouth movement matters to you, that is a different product and the HeyGen comparison is honest about where that line falls.

What the source recording needs to be

A clone is only as good as the audio it was built from, and the requirements are less demanding than people expect in one way and more demanding in another.

Less demanding: you do not need a studio, a treated room or an expensive microphone. A decent USB or lapel microphone in a normally furnished room is enough.

More demanding: it has to be consistent. A sample assembled from three recordings made at different distances, in different rooms, on different days produces a clone that inherits all three and settles on none of them. One continuous recording beats a longer one stitched together.

What actually helps, in order:

If the only footage you have is a finished video with music under the voice, the music and effects can be separated first and the voice used from the isolated stem — which is the same separation the dubbing stage uses to keep your soundtrack under the new audio.

How to judge the result in a language you do not speak

This is the awkward part nobody writes about. You cannot hear whether a Hindi track is saying the right thing, and you cannot fully judge whether it sounds native either.

Three checks work without knowing the language:

1. Compare it against the base voice, not against yourself. Generate a short sample both ways — base voice alone, and base voice with the clone applied — and listen back to back. The question is not "does this sound like me", because it will not, exactly. The question is whether the cloned version sounds warmer and more human than the base. If it does not, you are paying 8 credits a minute for nothing on that project.

2. Watch the timing against the picture. Lines that end early leave the speaker gesturing in silence; lines that run long collide with the next shot. Both are visible without understanding a word, and both are fixable at the translation stage rather than the audio stage.

3. Check the names. Your product name, your company name and your own name should be recognisable in the audio even when everything around them is not. If they have been translated or mangled, the glossary needs fixing before you generate anything else.

What none of that covers is register — whether the translation addresses your audience as a peer or a stranger. That needs someone who speaks the language, and it needs twenty minutes of their time reading the translation, not the finished audio. Which is the whole reason the pipeline stops before generating voice: the cheapest moment to involve them is before the credits are spent, not after.

What happens to your voice sample

A reasonable question to ask before uploading a recording of yourself, and one most pages skip.

The sample exists to generate audio for your projects. It is tied to the speaker record it was created against, alongside the consent file for that speaker, which is why consent is stored per speaker rather than per account — the permission and the voice it authorises live together.

Two practical consequences. First, a speaker marked as skipped keeps their consent record and simply is not cloned, so the decision is reversible without re-collecting anything. Second, because the record is per speaker rather than per project, a host who appears across fifty episodes is cloned once rather than fifty times.

If you are dubbing on behalf of a client, this is the part worth showing them. "We have a consent record for each voice we cloned" is a materially different conversation from "we assumed it was fine", and it is the kind of thing that only becomes important after the work has shipped.

Where to go next

The other services on the same balance: transcription, translation, subtitles and SRT export and review and approval. The pipeline breakdown shows how they connect and what each stage costs.

Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.

Questions people ask

How much does AI voice cloning cost?

Voice cloning adds 8 credits per minute on top of the dubbing tier you choose. A ten-minute video at Standard with cloning is 330 credits — 250 for the dub and 80 for the clone — which is roughly $14 on the $50 plan. Credits can be topped up from $10 with no subscription.

Do I need permission to clone someone else’s voice?

Yes, and the platform enforces it. Consent is recorded per speaker rather than per account, and a cloned track will not be generated for a speaker without one. In an interview you can clone the host and use a library voice for the guest in the same project.

Does cloning make the voice sound exactly like me?

No, and it is worth being clear about that. The dub starts from a base voice that already pronounces the target language properly, and cloning shapes that voice toward something warmer and more familiar. A listener who knows you will hear a resemblance rather than a substitution, and generated speech still carries an audible AI quality in its stress and intonation that cloning reduces but does not remove.

Is cloning always better than a library voice?

No. Clone when the audience already knows the voice — a creator, a founder, a spokesperson. For narration, product demos or screen recordings nobody is listening for whose voice it is, and a well-cast native base voice does the same job for 8 fewer credits a minute. On plain narration the base voice is often the cleaner result.

What makes a cloned voice sound artificial?

Usually the source recording rather than the clone: room echo and microphone distance survive into the output. After that, length mismatch between the translation and the original timecode, and loudness drift between the dubbed track and the original audio. Some AI quality in the stress and intonation remains regardless — cloning reduces it, it does not eliminate it.

Run the whole pipeline in one session.

Transcription, translation and dubbing on a single timeline, with one credit balance across all of them.

Get Started