Cloning is often sold as "your voice, speaking any language". That is not what happens, and the real mechanism is worth understanding before you spend credits on it.
A dub starts from a base voice that already speaks the target language properly — native pronunciation, correct stress, the right rhythm for that language. Cloning is applied on top of that base: it pulls the result toward something warmer and more familiar, closer to a person than to a narrator. What it does not do is transplant your identity wholesale into Hindi or Arabic.
It costs 8 credits a minute on top of whichever dubbing tier you choose. That premium is worth paying in some projects and wasted in others, and the rest of this page is about telling those apart — including the cases where the base voice alone is the better result.
What cloning actually does, and what it does not
What it does: shapes an already-native-sounding voice toward something more human and more familiar. The base voice handles the language; the clone handles the character. The result is a track that sounds less like a competent narrator and more like a person the audience has met before.
What it does not do — three things, and all three matter:
- It does not carry your identity intact across languages. A listener who knows you will hear a resemblance, not a substitution. The further the target language sits from your own, the more the base voice's characteristics show through.
- It does not remove the AI accent. This is the honest limit. Cloning makes a voice warmer; it does not make it indistinguishable from a recording of a person. Stress patterns and intonation still carry the signature of generated speech, and a listener paying attention will hear it.
- It does not carry your performance. Where you pause for effect, which word you lean on, when you speed up because you are excited — those come from the translated script and the timing, not from the clone.
That third point is why cloning is not a substitute for reviewing the translation. A cloned voice reading a stiff, over-literal line sounds exactly like a person having a bad day, which is worse than a neutral base voice reading the same line.
Worth saying plainly: on many projects the base voice on its own is the better result. A well-cast native voice with no cloning is clean and unremarkable, which is exactly what you want for narration. Cloning is for when unremarkable is not the goal.
Consent, and why it is built into the workflow
Cloning a voice is not a technical decision, it is a permissions decision. Every speaker in a project carries their own consent record, stored per speaker rather than per account, and the pipeline will not generate a cloned track for a speaker who does not have one.
This matters more than it sounds in three common situations:
- Interviews and podcasts. Your guest's voice is not yours to clone. Per-speaker consent means you can clone the host and use a library voice for the guest in the same episode, rather than facing an all-or-nothing choice.
- Client work. An agency dubbing a client's spokesperson needs a record that the permission existed, not an assurance that it did.
- Staff turnover. A voice recorded by an employee who has since left is a question worth having answered in writing before the track ships to six markets.
You can also mark a speaker as skipped, which leaves them on a library voice while the rest of the project is cloned. That is the normal case for anything with more than one person on the microphone.
When cloning earns its credits
| Situation | Clone? | Why |
|---|---|---|
| A creator whose face and voice are the channel | Yes | The voice is the brand; a stock voice reads as a different channel |
| A founder or spokesperson in a company video | Yes | Recognition carries across markets, and re-recording is not an option |
| Course or training material in one voice | Usually | Consistency across dozens of modules matters more than any single line |
| Narrated documentary or explainer, no on-screen host | Rarely | A well-cast library voice does the same job for 8 fewer credits a minute |
| Product demos and screen recordings | No | Nobody is listening for whose voice it is |
| Multi-speaker interviews | Selectively | Clone the host, cast the guests — consent and cost both point the same way |
The pattern: clone when the audience already knows the voice. If they do not, you are paying a premium to reproduce a voice nobody recognises, and a well-chosen voice from the library will serve you better.
The accent question cloning does not answer
A cloned voice speaking Spanish still has to pick a Spanish. Cloning carries your timbre; it does not decide whether your audience in Bogotá hears a Castilian or a Colombian register, and that decision affects reception more than the timbre does. The Spanish guide covers the nine regional options, and the case-study data covers what getting it right is worth.
The same applies inside Arabic, where thirteen dialects behave as different registers rather than different accents, and inside Portuguese, where Brazil and Portugal are two decisions rather than one.
What a cloned dub costs
Cloning is an addition to the dubbing rate, not a replacement for it. Everything runs off one credit balance:
- Transcription — 7 credits a minute, with speaker identification.
- Translation Deluxe — 18, or 14 if you supply the transcript.
- AI Dubbing Standard — 25, or 10 with transcript and translation.
- AI Dubbing Deluxe — 37, or 20 on the same basis.
- Voice cloning — adds 8 to whichever tier you picked.
A ten-minute video dubbed at Standard with cloning is 330 credits — the 250 for the dub plus 80 for the clone. On the $50 plan that is roughly $14 against $11 without it.
The economics improve sharply with a second language. Transcription and the master timeline happen once, so the second market costs the dubbing pass plus the clone premium over the same timecodes. Cloning a channel into four languages is closer in price to two than to four. The pipeline breakdown has the full table.
Credits top up from $10 with no subscription, which is enough to clone one short video and hear the result before deciding whether it is worth doing at scale.
What makes a clone sound wrong
1. A bad source recording. Cloning reproduces what it is given. Room echo, a noisy air conditioner and a microphone worn too far from the mouth all survive into the clone, and they are more noticeable in a language where the listener is working slightly harder to follow.
2. Length mismatch. Translations rarely run the same length as the original. If a line is generated without being held inside its timecode, the cloned voice either rushes or leaves a gap — and because the voice is recognisably yours, the artificiality lands harder than it would on a stock voice.
3. Loudness drift. A cloned track generated at one level and published next to your original audio gives the whole thing away. Every export is mastered to the destination's loudness spec so both tracks sit at the same level.
4. Skipping the review. The pipeline stops after transcription and again after translation, before any credit goes on voice. With cloning that checkpoint is worth more, not less: regenerating a cloned track costs the dubbing rate and the clone premium again.
Projects with more than one voice
Most guidance on cloning assumes one speaker. Real footage rarely is, and the multi-speaker case is where the decisions get interesting.
A two-host podcast. Both hosts are known to the audience, both need consent, and the value of cloning is high because listeners recognise the pair as a pair. This is the clearest case for cloning everyone.
An interview. The host is the constant; the guest appears once. Cloning the host and casting a library voice for the guest is usually right — it solves the consent question, halves the premium, and the audience is not expecting to recognise the guest's voice in another language anyway.
A panel or a roundtable. Four voices, four consents, four clone premiums. Here the honest advice is usually to clone nobody and cast four distinct library voices instead: what the viewer needs is to tell the speakers apart, and distinct casting achieves that for 8 fewer credits a minute per speaker.
Corporate video with a narrator and a spokesperson. Clone the spokesperson, whose face is on screen and whose identity carries; leave the narrator on a library voice, because nobody knows what the narrator sounds like.
The common thread is worth stating plainly: cloning buys recognition, and recognition only has value where it already exists. Speaker labels from the transcription stage carry through the whole pipeline, so these choices are made per speaker rather than per project.
Cloning against the alternatives
There are three ways to put your content into another language with your identity attached, and they are not equally sensible.
| Approach | What it costs | Where it breaks |
|---|---|---|
| Learn the language and re-record | Months per language | Does not scale past one or two, and accent gives it away |
| Hire a voice actor per market | Studio rates, per language, per revision | Scheduling; and the voice is not yours in any market |
| A native base voice, optionally cloned | +8 credits a minute for the clone | Resemblance rather than substitution; needs a clean sample |
The comparison people actually make, though, is cloning against a good library voice — and there the answer depends entirely on whether the audience knows you. A channel with a face and a name loses something real when the voice changes. A tutorial channel with no on-camera presence loses nothing at all.
Worth noting what cloning does not replace: it does not lip-sync. If your video shows a face talking to camera and mouth movement matters to you, that is a different product and the HeyGen comparison is honest about where that line falls.
What the source recording needs to be
A clone is only as good as the audio it was built from, and the requirements are less demanding than people expect in one way and more demanding in another.
Less demanding: you do not need a studio, a treated room or an expensive microphone. A decent USB or lapel microphone in a normally furnished room is enough.
More demanding: it has to be consistent. A sample assembled from three recordings made at different distances, in different rooms, on different days produces a clone that inherits all three and settles on none of them. One continuous recording beats a longer one stitched together.
What actually helps, in order:
- Speak the way you normally speak. A sample recorded in a careful, announcer voice produces a clone that sounds like you doing an impression of a newsreader, and it will do that in every language.
- Cover your range. Include a question, a moment of emphasis, and a sentence delivered flat. A sample that is uniformly enthusiastic gives the model nothing to work with when the script goes quiet.
- Avoid music and effects underneath. They get absorbed into the voice rather than ignored.
- Keep the distance constant. Moving toward and away from the microphone during the sample teaches the clone an inconsistency it will reproduce.
If the only footage you have is a finished video with music under the voice, the music and effects can be separated first and the voice used from the isolated stem — which is the same separation the dubbing stage uses to keep your soundtrack under the new audio.
How to judge the result in a language you do not speak
This is the awkward part nobody writes about. You cannot hear whether a Hindi track is saying the right thing, and you cannot fully judge whether it sounds native either.
Three checks work without knowing the language:
1. Compare it against the base voice, not against yourself. Generate a short sample both ways — base voice alone, and base voice with the clone applied — and listen back to back. The question is not "does this sound like me", because it will not, exactly. The question is whether the cloned version sounds warmer and more human than the base. If it does not, you are paying 8 credits a minute for nothing on that project.
2. Watch the timing against the picture. Lines that end early leave the speaker gesturing in silence; lines that run long collide with the next shot. Both are visible without understanding a word, and both are fixable at the translation stage rather than the audio stage.
3. Check the names. Your product name, your company name and your own name should be recognisable in the audio even when everything around them is not. If they have been translated or mangled, the glossary needs fixing before you generate anything else.
What none of that covers is register — whether the translation addresses your audience as a peer or a stranger. That needs someone who speaks the language, and it needs twenty minutes of their time reading the translation, not the finished audio. Which is the whole reason the pipeline stops before generating voice: the cheapest moment to involve them is before the credits are spent, not after.
What happens to your voice sample
A reasonable question to ask before uploading a recording of yourself, and one most pages skip.
The sample exists to generate audio for your projects. It is tied to the speaker record it was created against, alongside the consent file for that speaker, which is why consent is stored per speaker rather than per account — the permission and the voice it authorises live together.
Two practical consequences. First, a speaker marked as skipped keeps their consent record and simply is not cloned, so the decision is reversible without re-collecting anything. Second, because the record is per speaker rather than per project, a host who appears across fifty episodes is cloned once rather than fifty times.
If you are dubbing on behalf of a client, this is the part worth showing them. "We have a consent record for each voice we cloned" is a materially different conversation from "we assumed it was fine", and it is the kind of thing that only becomes important after the work has shipped.
Where to go next
The other services on the same balance: transcription, translation, subtitles and SRT export and review and approval. The pipeline breakdown shows how they connect and what each stage costs.
Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.
Questions people ask
How much does AI voice cloning cost?
Voice cloning adds 8 credits per minute on top of the dubbing tier you choose. A ten-minute video at Standard with cloning is 330 credits — 250 for the dub and 80 for the clone — which is roughly $14 on the $50 plan. Credits can be topped up from $10 with no subscription.
Do I need permission to clone someone else’s voice?
Yes, and the platform enforces it. Consent is recorded per speaker rather than per account, and a cloned track will not be generated for a speaker without one. In an interview you can clone the host and use a library voice for the guest in the same project.
Does cloning make the voice sound exactly like me?
No, and it is worth being clear about that. The dub starts from a base voice that already pronounces the target language properly, and cloning shapes that voice toward something warmer and more familiar. A listener who knows you will hear a resemblance rather than a substitution, and generated speech still carries an audible AI quality in its stress and intonation that cloning reduces but does not remove.
Is cloning always better than a library voice?
No. Clone when the audience already knows the voice — a creator, a founder, a spokesperson. For narration, product demos or screen recordings nobody is listening for whose voice it is, and a well-cast native base voice does the same job for 8 fewer credits a minute. On plain narration the base voice is often the cleaner result.
What makes a cloned voice sound artificial?
Usually the source recording rather than the clone: room echo and microphone distance survive into the output. After that, length mismatch between the translation and the original timecode, and loudness drift between the dubbed track and the original audio. Some AI quality in the stress and intonation remains regardless — cloning reduces it, it does not eliminate it.