Service

Video translation, reviewed before it becomes a voice

Translating a video is not translating a document with timestamps attached. The line has to fit a shot, match a speaker, survive being spoken aloud, and stay consistent across an hour of footage. Here is what changes because of that.

Translation sits in the middle of the pipeline: it inherits the segments and speakers from transcription, and everything after it — the subtitle files, the dubbed audio, the cloned voice — is built on what it produces.

It costs 18 credits a minute, or 14 if you supply the transcript yourself. And it is the stage where reviewing costs nothing and not reviewing costs the most.

Why translating for video is a different job

A document translator optimises for accuracy and readability. A video translator has four extra constraints, and they frequently pull against each other.

1. The line has to fit the time it has. A translation that is perfect and 40% longer than the original does not fit the shot. Something has to give: either the sentence gets shorter or the voice gets faster, and only one of those sounds natural.

2. It has to survive being spoken. Written translations can use subordinate clauses and dense constructions that a reader handles comfortably and a listener does not. Anything that needs re-reading has already failed out loud.

3. It has to know who is speaking. The same sentence is translated differently depending on whether it is a host addressing an audience or a guest answering a question. Speaker labels carry through from transcription for exactly this reason.

4. It has to stay consistent for an hour. A document is translated in one sitting by one mind. A video is a sequence of segments, and the risk is that your product name is rendered three different ways across forty minutes without anyone noticing until it is in the audio.

How much languages expand, and what to do about it

Translations rarely land at the same length as the source, and the direction is fairly predictable:

From English into…TypicallyWhat it causes
GermanLongerCompound nouns and verb-final order; the voice speeds up to fit
ArabicLongerSame problem, worsened by formal register
Spanish, Portuguese, FrenchModerately longerManageable, but accumulates across a long video
ChineseShorterSilence where the speaker is still visibly talking
JapaneseLongerParticles and polite forms add syllables
TurkishUnpredictableAgglutination can compress a clause into one word

Short lines are as much of a problem as long ones — a hole in the audio while the speaker is mid-gesture reads as a glitch. The practical answer is to see the speech-rate figure beside each segment before generating audio, and to shorten or extend the line rather than letting the delivery absorb the difference.

The glossary is the part that pays for itself

Three decisions run through every sentence of a video, and if nobody makes them the translation engine makes them silently and sticks with whatever it picked.

Formal or informal. Tu or vous, du or Sie, tum or aap. Your video addresses the viewer either as a peer or as a stranger. For a creator video the informal register is usually right; for finance or B2B in Germany it is usually wrong.

What stays in English. In South Asian languages especially, everyday speech mixes English constantly — technology, business and brand vocabulary. Translating every term out produces something correct on paper and stilted in the ear.

Names that are not words. Product names, feature names, brand names. These are exactly what a general model has least reason to leave alone, and exactly what your audience needs to hear unchanged so they can search for it afterwards.

Setting these once at the translation stage costs nothing. Discovering them after the audio exists costs the whole dubbing stage again — and, if you cloned a voice, the clone premium with it.

The checkpoint, and why it exists

The pipeline stops here on purpose. You read the translation against the original, segment by segment, and correct it while correcting it is free. Only then does a single credit go on voice.

The reason is arithmetic rather than philosophy. An error in the translation hardens into the generated audio, so catching it afterwards means paying to regenerate the dubbing stage — 25 credits a minute at Standard, 37 at Deluxe, plus 8 more if the voice was cloned. It was free to fix here and expensive to fix later.

This is also the stage where a bilingual colleague is worth more than any tool. You do not need them to translate; you need them to read forty lines and tell you whether the register is right. That is twenty minutes of somebody's time against a regeneration bill.

What video translation costs

Everything runs from one credit balance:

A ten-minute video transcribed and translated is 250 credits — roughly $11 on the $50 plan — and that already gives you subtitle files in both languages, with no audio generated. If you stop there, you have a video that is searchable and watchable in a second language for a fraction of a dub.

Credits top up from $10 with no subscription. The pipeline breakdown has every rate in one table.

Machine translation and translation for dubbing

Free translation tools are genuinely good at what they were built for: rendering a sentence accurately, in isolation, for a reader. Video breaks three of those assumptions at once.

Isolation. A subtitle is not read in isolation — it follows the one before it and precedes the one after. Pronouns, callbacks and running jokes all depend on segments the translator of a single line cannot see.

For a reader. Written and spoken registers differ in every language. A construction that reads elegantly can be almost unsayable, and the difference only appears when a voice tries to deliver it.

Accuracy as the goal. In dubbing, a slightly less literal line that fits the shot and sounds natural out loud beats a perfectly accurate one that does not. The constraint is the timeline, and a translator that cannot see the timing cannot honour it.

This is why the translation stage works on segments with timecodes and speaker labels rather than on a block of text — and why the speech-rate figure appears beside each line before any audio is generated.

What does not survive translation, and what to do

Some things simply do not cross, and knowing which is the difference between a dub that works and one that is technically correct and quietly dead.

Wordplay. A pun is a coincidence of one language. The options are to replace it with a different joke that lands in the target language, or to drop it and let the line be plain. What does not work is translating it literally, which produces a sentence that is clearly meant to be funny and is not.

Cultural references. A comparison to a television programme nobody in the target market watched is worse than no comparison. Usually the fix is to generalise: naming the category instead of the example.

Idioms. Every language has its own, and they map badly. The literal version is the single most common reason a dub reads as machine-made.

Register jokes. Humour built on suddenly being very formal or very casual depends on the target language having the same registers in the same places. Often it does not.

Units, dates and currency. Not a translation problem but a localisation one, and easy to miss: 03/04 is two different days depending on the market, and a price in dollars means nothing to a viewer who cannot buy in dollars.

The practical approach is to mark these at the review stage rather than expect any engine to solve them. They are a handful of lines in a typical video, and they are the lines that decide whether the result sounds like it was made for the audience or aimed at them.

Working with a reviewer who is not a translator

You do not need a professional translator to get most of the value from the review checkpoint. You need someone who speaks the language and will answer four questions.

  1. Does this address the viewer the way we want? Formal or informal — the single highest-impact question, and answerable in thirty seconds by any native speaker.
  2. Does anything sound translated? Not wrong, just foreign in construction. Native speakers spot this instantly even when they cannot explain why.
  3. Are the names right? Product, company, people. These should survive unchanged so viewers can search for them.
  4. Is anything accidentally rude or odd? Rare, but it is the one that causes real damage, and it is invisible to everyone who does not speak the language.

Twenty minutes of that, on the translation, before any credit goes on voice, is worth more than any amount of checking the finished audio — because at this stage every fix is free, and afterwards each one costs the dubbing stage again.

Where the difficulty concentrates

Not all language pairs are equally hard, and the difficulty is not where people expect.

Closely related languages — English to German, Spanish to Portuguese — are easier for meaning and treacherous for false friends: words that look identical and are not. These produce errors a reviewer catches instantly and a model does not flag.

Distant pairs — English to Japanese, English to Arabic — restructure the sentence so completely that literal errors are less likely, because nothing is close enough to be copied. The risk shifts to register and length instead.

Languages with heavy honorifics — Japanese, Korean, Hindi — carry the speaker-listener relationship in the grammar itself. There is no neutral option; the translation picks one and lives with it for the whole video.

Languages with major regional splits — Spanish, Arabic, Portuguese — need a market decision before the translation runs, not after. Translating into "Spanish" and choosing the accent afterwards produces a Colombian voice reading Castilian vocabulary, which is more jarring than either choice made consistently.

Translating for subtitles and translating for dubbing

The same source sentence produces two different translations depending on where it is going, and treating them as one job is a common and visible mistake.

For subtitlesFor dubbing
ConstraintReading speed — how fast a person reads while watchingSpeaking time — how long the line takes out loud
CompressionExpected; subtitles routinely condenseAvoided; a condensed line leaves silence
RegisterCan stay closer to written normsMust be sayable, in spoken register
RepetitionUsually trimmedOften kept — people repeat themselves out loud
Failure modeToo dense to read in timeRushed delivery or a gap on screen

Because both are produced from the same stored translation, the practical approach is to translate for the spoken version and let the subtitle inherit it. A line written to be said aloud reads perfectly well; the reverse is not reliably true, and a subtitle-first translation delivered by a voice is where "technically correct but oddly stiff" comes from.

Where the two genuinely diverge — a line that is comfortable to hear and too long to read at that timecode — it is the caption that gets adjusted, because that adjustment costs nothing and regenerating audio does not.

Translating a library, not a video

Everything above assumes one video. Translating a back catalogue changes which decisions matter.

The glossary stops being optional. Across forty videos, terminology drift is guaranteed unless the terms are fixed once. Your product name rendered three ways across a library is worse than any single bad translation, because it breaks search for the viewer who heard it in one video and looked for it later.

Register has to be decided at library level. If half your videos address the viewer formally and half informally, the channel reads as inconsistent even to a viewer who never watches two in a row.

Order by evidence, not by intuition. Your analytics already show which markets watch you without understanding you. Translating the three videos that carry the most traffic into one language tells you more than translating one video into six.

The transcript becomes the asset. Once a video is transcribed and corrected, every additional language costs the translation and dubbing stages only. A library that is fully transcribed is a library where adding a market is a per-minute cost rather than a project — which is the same reason supplying your own transcript drops translation to 14 credits a minute and Standard dubbing to 10.

The YouTube guide covers how to read the analytics that decide the order, and the case-study data covers what the accent choice inside each language is worth once the videos are live.

The short version of this whole page: translation is the stage where a mistake is still free, and the last one where that is true. Everything after it — the audio, the voice, the master — inherits whatever you approve here, which is why the twenty minutes spent reading it back is the best-value time in the entire pipeline.

Where to go next

The other services on the same balance: transcription, voice cloning, subtitles and SRT export and review and approval. The pipeline breakdown shows how they connect and what each stage costs.

Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.

Questions people ask

How much does it cost to translate a video?

Translation Deluxe is 18 credits a minute, or 14 if you supply the transcript yourself. A ten-minute video transcribed and translated is 250 credits — roughly $11 on the $50 plan — and that includes subtitle files in both languages with no audio generated.

Do I have to dub the video after translating it?

No. Translation produces subtitle files in the target language on its own, and dubbing is a separate stage you can skip. A video translated but not dubbed is still searchable and watchable in the second language.

Why does the pipeline stop after translation?

Because an error in the translation hardens into the generated voice. Catching it afterwards means paying to regenerate the dubbing stage — 25 credits a minute at Standard, 37 at Deluxe, plus 8 more if the voice was cloned. Reviewing at this stage costs nothing.

How do you handle brand and product names?

Through a glossary set at the translation stage. Product names, feature names and brand names are exactly what a general model is most likely to translate and exactly what your audience needs to hear unchanged so they can search for it later.

Does translation work from languages other than English?

Yes. Translation runs from the source language directly across 80+ languages rather than routing through English, which avoids losing a layer of meaning on the way.

Run the whole pipeline in one session.

Transcription, translation and dubbing on a single timeline, with one credit balance across all of them.

Get Started