Service

Video transcription: the timeline everything else inherits

Transcription is where a project becomes editable — and where a mistake is cheapest to fix. Every later stage inherits its segments, its speakers and its timings, which is why it is worth getting right before anything else runs.

Transcription is the entry point of the pipeline and costs 7 credits a minute. What comes out is not a wall of text: it is a segmented timeline, with speaker labels and timecodes, that translation and dubbing both build on top of.

That structure is the reason the price of everything downstream depends on this stage.

What you get back

What accuracy to expect, honestly

On clear audio in a well-supported language, expect 90 to 95% before any correction. That is a useful number rather than a flattering one, and the range matters more than the midpoint, because what moves you inside it is not the model — it is the recording:

Which is why the transcript arrives editable and the pipeline stops there. Fixing your product name in eight places takes a minute at this stage; discovering it in a finished dub in six languages does not.

Why the transcript decides the bill

An error here does not stay here. It propagates into the translation and hardens into the generated voice, so catching it at the end means paying to regenerate all three stages — including the two that were correct.

It was free to fix at step one and expensive at step three.

So the pipeline stops between stages by design. You review the segments, the speakers and the timings while correcting them costs nothing. Only then does translation run, on a transcript you approved — and it stops the same way before a single credit goes on audio.

There is a second, more direct effect on price. Because you can supply your own transcript, the stages that follow get cheaper: translation drops from 18 credits a minute to 14, and Standard dubbing drops from 25 to 10. Uploading a transcript and a translation as a CSV is the single biggest lever on a bill — it cuts the dubbing rate by 60%.

What people actually do with a transcript

UseWhat it needsCost
Subtitles in the original languageTranscription only7 cr/min
Searchable text for SEOTranscription only7 cr/min
Subtitles in another language+ translation25 cr/min
A dubbed track+ translation + dubbingfrom 25 cr/min
Repurposing into articles or clipsTranscription only7 cr/min

The first two are worth noting because they need no dubbing at all. A transcript gives a search index text where it previously had only audio, and it gives you the raw material for written formats — which is the cheapest thing on this list and the one most often skipped.

Languages and speakers

Transcription covers the same 80+ languages as the rest of the platform, and the language of the source does not have to be English. A Spanish interview transcribes in Spanish and translates from there, rather than routing through English and losing a layer of meaning on the way.

For multi-speaker recordings the speaker labels carry through the whole pipeline, so a two-host podcast dubbed into German keeps two distinct voices rather than collapsing into one. The language index lists what is available, and voice cloning can be applied per speaker where consent exists.

What transcription costs

Seven credits a minute, out of the same balance as everything else. A ten-minute video is 70 credits — roughly $3 on the $50 plan — and that includes the speaker labels, the timecodes and an exportable subtitle file.

Credits top up from $10 with no subscription, which transcribes well over an hour of video. The pipeline breakdown has the rates for every stage.

How to get more than 95% out of it

The model is the same for everyone. What separates a transcript that needs ten corrections from one that needs a hundred is almost entirely the recording, and most of it is decided before you hit record.

None of this is exotic. It is the same advice that makes a video sound better to a human listener, which is the point: transcription accuracy is mostly a proxy for how intelligible the recording already was.

Correcting a transcript without losing the timeline

The transcript is not a text document with timestamps attached — it is a set of segments, each owning a start and an end, and every later stage inherits those boundaries. That changes what "editing" means.

Correcting text inside a segment is free and has no side effects. Fixing a misheard product name changes nothing about the timing.

Fixing speaker attribution is worth doing carefully, because those labels carry all the way through: a line attributed to the wrong speaker will be generated in the wrong voice, in every language you produce.

Adjusting segment boundaries is the one that has consequences downstream. Segments are what a dubbed line is held inside, so a segment that is too long gives the voice room to drift and one that is too short forces it to rush. If a sentence was split across two segments in a way that reads oddly, it is better to fix it here than to discover it in the audio.

Overlapping speech should stay overlapping. The temptation is to tidy cross-talk into a clean sequence. Resist it: the overlapping timecodes are what let two dubbed voices talk over each other the way the original did, and flattening them into a queue is the most common way an interview dub ends up sounding like a reading.

All of this happens before translation runs, which is the point of stopping there. The pipeline breakdown covers what each subsequent stage inherits.

Transcription is not the same as captions

These are used as synonyms and they are not, which causes real confusion when quoting work.

A transcript is the text of what was said, segmented and timed, with speakers labelled. It is a working document: editable, searchable, and the input to everything else.

Captions are that text formatted to be read on screen while the video plays — constrained by reading speed, line length and minimum duration. Every caption comes from a transcript; not every transcript makes a good caption without adjustment.

Subtitles are captions in a different language, which means a translation stage between the two.

The distinction matters because the transcript is useful on its own, and often more useful than the caption file. It is the thing you search, quote, repurpose and hand to a colleague — and it is what makes the later stages cheaper, because supplying it drops translation from 18 credits a minute to 14. The subtitle page covers the caption side in detail.

What a transcript is worth beyond subtitles

Most people transcribe to caption or to dub, and then leave the most valuable artefact sitting in the project.

Search visibility. A video page with no text is opaque to a search index. The transcript is the only text a video-first page has, and publishing it — even as a collapsed section under the player — gives an index something to read in a language it previously had none of.

Repurposing. A forty-minute interview contains a newsletter, three or four short-form clips and a written article. Finding them in a transcript takes minutes; finding them by scrubbing the video takes an afternoon. The timecodes mean that once you have found the passage you also know exactly where to cut.

Internal search. A library of transcribed recordings is searchable. A library of video files is a list of filenames.

Compliance and records. For interviews, webinars and training material, a timestamped record of what was said is frequently the reason the transcription was worth 7 credits a minute in the first place.

All four need transcription only — no translation, no dubbing. That makes them the cheapest thing available here and the most commonly skipped.

Recordings that are genuinely hard

Accuracy is a range because material varies. These are the cases where the low end of 90 to 95% is realistic and worth planning for.

None of these make transcription unusable — they change how much correction to budget for. And because correction happens before translation and dubbing run, the cost of that correction is your time rather than credits. On difficult material the order matters more than usual: transcribe, correct thoroughly, and only then let the expensive stages inherit it.

Speaker identification, and where it needs help

Speaker labels are the part of transcription with the largest downstream effect, because they decide which voice says which line in every language you produce afterwards.

Where it works well: two to four speakers with distinct voices, recorded at reasonable quality, taking turns. This is most interviews, most podcasts and nearly all corporate video.

Where it needs correcting:

Fixing these takes a couple of minutes at this stage. Left alone, a mislabelled line is generated in the wrong voice in every target language — and if you are also cloning voices, in the wrong person's cloned voice, which is the version of this problem that actually causes a re-render.

Where transcription sits against the alternatives

Transcription is a commodity in the sense that many tools do it. What differs is what happens next.

A standalone transcription tool gives you text and stops. If all you need is a record of a meeting, that is the right tool and probably cheaper.

A captioning tool gives you text formatted for the screen, usually in the source language, and stops there too.

Transcription inside a dubbing pipeline produces a timeline rather than a document: segments that later hold a dubbed line inside their own duration, speaker labels that select a voice per person, and a transcript whose corrections make the next two stages cheaper — 18 credits a minute down to 14 for translation, 25 down to 10 for Standard dubbing.

Which of those you want depends entirely on whether the video is going to travel. If it never leaves its original language, a cheaper standalone tool is a sensible choice. If there is any chance of a second market, transcribing inside the pipeline means that market costs a dubbing pass rather than starting over. The pipeline breakdown shows how the stages connect.

Transcribe once, reuse across every video

This is the part of the pricing that most people never use, and it is the one that changes a bill rather than trimming it.

A transcript is not locked to the project that produced it. From the New file modal you can import an existing transcript — as CSV or SRT — instead of paying to transcribe from scratch. The new project starts with the segments already in place.

That matters in more situations than it first appears:

Export CSV, not SRT, when you plan to reuse it

Both import. They do not carry the same information, and the difference is the one thing you most want to keep.

CSVSRT
TimecodesYesYes
TextYesYes
Assigned speakersYesNo
Best forReusing inside the platformUploading captions to a player

The CSV carries a speaker column alongside the timings and the text. An SRT has nowhere to put it — the format simply has no field for who is talking.

So a reused SRT arrives as one undifferentiated voice. You get the words and the timings back, and then you reassign every speaker by hand before dubbing can cast them. On a two-host podcast that is an annoyance; on a panel or a drama it is most of an afternoon, repeated on every episode.

The rule is short: export SRT for players, export CSV for the platform. If there is any chance the transcript will be reused, export both — they cost nothing and only one of them survives the round trip intact.

What reuse actually saves

Importing a transcript removes the transcription charge and moves the later stages onto their reduced rates, because you are supplying work the platform would otherwise do:

StageFrom scratchWith your imported transcript
Transcription7 cr/minnot charged
Translation Deluxe18 cr/min14
AI Dubbing Standard25 cr/min10 with the translation too
AI Dubbing Deluxe37 cr/min20 on the same basis

On a ten-minute video dubbed at Standard, going in with a transcript and a translation takes the job from 250 credits to 100 — roughly $4.40 instead of $11 on the $50 plan. That is a 60% cut on the dubbing rate, and it is the single largest lever available on a bill.

For a series it compounds. Transcribe episode one, correct it properly, export the CSV, and every later episode that shares dialogue starts from corrected text with its speakers already attached — which also means the corrections you made once do not have to be made again in every language.

Where to go next

The other services on the same balance: translation, voice cloning, subtitles and SRT export and review and approval. The pipeline breakdown shows how they connect and what each stage costs.

Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.

Questions people ask

How accurate is AI video transcription?

On clear audio in a well-supported language, expect 90 to 95% before correction. What moves you within that range is the recording rather than the model: microphone distance is the biggest factor, followed by cross-talk, proper nouns and jargon, and music under speech. The transcript arrives editable so those can be fixed before anything downstream runs.

How much does video transcription cost?

Seven credits a minute. A ten-minute video is 70 credits, roughly $3 on the $50 plan, and that includes speaker identification, timecodes and an exportable subtitle file. Credits top up from $10 with no subscription.

Does it identify who is speaking?

Yes. Speakers are labelled and those labels carry through translation and dubbing, so a two-host podcast keeps two distinct voices rather than collapsing into one. Overlapping speech stays as separate lines with overlapping timecodes instead of being flattened into a queue.

Can I upload my own transcript instead?

Yes, and it makes everything after it cheaper. Supplying a transcript drops translation from 18 credits a minute to 14, and supplying both a transcript and a translation drops Standard dubbing from 25 to 10 — a 60% cut on the dubbing rate.

Does the source video have to be in English?

No. Transcription covers the same 80+ languages as the rest of the platform, and translation runs from the source language directly rather than routing through English.

Run the whole pipeline in one session.

Transcription, translation and dubbing on a single timeline, with one credit balance across all of them.

Get Started