Transcription is the entry point of the pipeline and costs 7 credits a minute. What comes out is not a wall of text: it is a segmented timeline, with speaker labels and timecodes, that translation and dubbing both build on top of.
That structure is the reason the price of everything downstream depends on this stage.
What you get back
- Segments with timecodes, not a paragraph. Each line owns a start and an end, which is what later lets a dubbed line be held inside its own duration instead of drifting.
- Speaker identification. Who said what, labelled, so an interview does not arrive as one undifferentiated monologue.
- Overlapping speech kept separate. Two people talking over each other stay two lines with overlapping timecodes rather than being flattened into a queue.
- An editable transcript, correctable before anything else runs.
- A subtitle file in the original language, as a direct export.
What accuracy to expect, honestly
On clear audio in a well-supported language, expect 90 to 95% before any correction. That is a useful number rather than a flattering one, and the range matters more than the midpoint, because what moves you inside it is not the model — it is the recording:
- Microphone distance is the single biggest factor. A lapel mic and a laptop mic in the same room produce noticeably different transcripts.
- Cross-talk costs accuracy on both speakers, not just the quieter one.
- Proper nouns, brand names and jargon are where errors concentrate, because they are exactly the words a general model has least reason to expect.
- Music under speech degrades it more than steady background noise does.
Which is why the transcript arrives editable and the pipeline stops there. Fixing your product name in eight places takes a minute at this stage; discovering it in a finished dub in six languages does not.
Why the transcript decides the bill
An error here does not stay here. It propagates into the translation and hardens into the generated voice, so catching it at the end means paying to regenerate all three stages — including the two that were correct.
It was free to fix at step one and expensive at step three.
So the pipeline stops between stages by design. You review the segments, the speakers and the timings while correcting them costs nothing. Only then does translation run, on a transcript you approved — and it stops the same way before a single credit goes on audio.
There is a second, more direct effect on price. Because you can supply your own transcript, the stages that follow get cheaper: translation drops from 18 credits a minute to 14, and Standard dubbing drops from 25 to 10. Uploading a transcript and a translation as a CSV is the single biggest lever on a bill — it cuts the dubbing rate by 60%.
What people actually do with a transcript
| Use | What it needs | Cost |
|---|---|---|
| Subtitles in the original language | Transcription only | 7 cr/min |
| Searchable text for SEO | Transcription only | 7 cr/min |
| Subtitles in another language | + translation | 25 cr/min |
| A dubbed track | + translation + dubbing | from 25 cr/min |
| Repurposing into articles or clips | Transcription only | 7 cr/min |
The first two are worth noting because they need no dubbing at all. A transcript gives a search index text where it previously had only audio, and it gives you the raw material for written formats — which is the cheapest thing on this list and the one most often skipped.
Languages and speakers
Transcription covers the same 80+ languages as the rest of the platform, and the language of the source does not have to be English. A Spanish interview transcribes in Spanish and translates from there, rather than routing through English and losing a layer of meaning on the way.
For multi-speaker recordings the speaker labels carry through the whole pipeline, so a two-host podcast dubbed into German keeps two distinct voices rather than collapsing into one. The language index lists what is available, and voice cloning can be applied per speaker where consent exists.
What transcription costs
Seven credits a minute, out of the same balance as everything else. A ten-minute video is 70 credits — roughly $3 on the $50 plan — and that includes the speaker labels, the timecodes and an exportable subtitle file.
Credits top up from $10 with no subscription, which transcribes well over an hour of video. The pipeline breakdown has the rates for every stage.
How to get more than 95% out of it
The model is the same for everyone. What separates a transcript that needs ten corrections from one that needs a hundred is almost entirely the recording, and most of it is decided before you hit record.
- Get the microphone closer. A lapel or a boom at 20cm beats a laptop microphone across a desk by a wider margin than any other single change. This matters more than the price of the microphone.
- Record speakers on separate tracks when you can. Cross-talk is the second-largest source of error, and separate tracks eliminate it rather than mitigating it.
- Kill the room, not the noise. Steady background hum is handled reasonably well; reverberation is not, because it smears the boundaries between words.
- Say proper nouns clearly the first time. Brand names, product names and people's names are where errors concentrate — and where they are most expensive, because they repeat throughout the video and propagate into every language.
- Avoid music under speech. It degrades accuracy more than constant noise at the same level, and it is usually added in the edit rather than being unavoidable.
None of this is exotic. It is the same advice that makes a video sound better to a human listener, which is the point: transcription accuracy is mostly a proxy for how intelligible the recording already was.
Correcting a transcript without losing the timeline
The transcript is not a text document with timestamps attached — it is a set of segments, each owning a start and an end, and every later stage inherits those boundaries. That changes what "editing" means.
Correcting text inside a segment is free and has no side effects. Fixing a misheard product name changes nothing about the timing.
Fixing speaker attribution is worth doing carefully, because those labels carry all the way through: a line attributed to the wrong speaker will be generated in the wrong voice, in every language you produce.
Adjusting segment boundaries is the one that has consequences downstream. Segments are what a dubbed line is held inside, so a segment that is too long gives the voice room to drift and one that is too short forces it to rush. If a sentence was split across two segments in a way that reads oddly, it is better to fix it here than to discover it in the audio.
Overlapping speech should stay overlapping. The temptation is to tidy cross-talk into a clean sequence. Resist it: the overlapping timecodes are what let two dubbed voices talk over each other the way the original did, and flattening them into a queue is the most common way an interview dub ends up sounding like a reading.
All of this happens before translation runs, which is the point of stopping there. The pipeline breakdown covers what each subsequent stage inherits.
Transcription is not the same as captions
These are used as synonyms and they are not, which causes real confusion when quoting work.
A transcript is the text of what was said, segmented and timed, with speakers labelled. It is a working document: editable, searchable, and the input to everything else.
Captions are that text formatted to be read on screen while the video plays — constrained by reading speed, line length and minimum duration. Every caption comes from a transcript; not every transcript makes a good caption without adjustment.
Subtitles are captions in a different language, which means a translation stage between the two.
The distinction matters because the transcript is useful on its own, and often more useful than the caption file. It is the thing you search, quote, repurpose and hand to a colleague — and it is what makes the later stages cheaper, because supplying it drops translation from 18 credits a minute to 14. The subtitle page covers the caption side in detail.
What a transcript is worth beyond subtitles
Most people transcribe to caption or to dub, and then leave the most valuable artefact sitting in the project.
Search visibility. A video page with no text is opaque to a search index. The transcript is the only text a video-first page has, and publishing it — even as a collapsed section under the player — gives an index something to read in a language it previously had none of.
Repurposing. A forty-minute interview contains a newsletter, three or four short-form clips and a written article. Finding them in a transcript takes minutes; finding them by scrubbing the video takes an afternoon. The timecodes mean that once you have found the passage you also know exactly where to cut.
Internal search. A library of transcribed recordings is searchable. A library of video files is a list of filenames.
Compliance and records. For interviews, webinars and training material, a timestamped record of what was said is frequently the reason the transcription was worth 7 credits a minute in the first place.
All four need transcription only — no translation, no dubbing. That makes them the cheapest thing available here and the most commonly skipped.
Recordings that are genuinely hard
Accuracy is a range because material varies. These are the cases where the low end of 90 to 95% is realistic and worth planning for.
- Conference and panel audio. Room reverberation, distant microphones and several people talking. Expect to correct names heavily.
- Phone and video-call recordings. Compressed, band-limited audio loses the high frequencies that distinguish consonants — which is why "s" and "f" confusions cluster here.
- Heavy technical or medical vocabulary. Not because the terms are difficult but because they are rare, and a general model weights toward common words.
- Code-switching mid-sentence. A speaker moving between languages inside one sentence is the hardest case in this list, and common in South Asian and Arabic content.
- Strong regional accents against a general model. Improving, but still the difference between the top and the bottom of the range.
None of these make transcription unusable — they change how much correction to budget for. And because correction happens before translation and dubbing run, the cost of that correction is your time rather than credits. On difficult material the order matters more than usual: transcribe, correct thoroughly, and only then let the expensive stages inherit it.
Speaker identification, and where it needs help
Speaker labels are the part of transcription with the largest downstream effect, because they decide which voice says which line in every language you produce afterwards.
Where it works well: two to four speakers with distinct voices, recorded at reasonable quality, taking turns. This is most interviews, most podcasts and nearly all corporate video.
Where it needs correcting:
- Similar voices. Two speakers of the same gender with similar pitch and pace are the classic case, and the errors cluster at turn boundaries rather than being spread evenly.
- Short interjections. A one-word "right" or "exactly" from the other speaker is easy to attribute to whoever was already talking.
- A speaker who joins late. Someone appearing thirty minutes in sometimes gets folded into an existing label instead of getting their own.
- Heavy cross-talk. Overlapping segments are kept as separate lines, but which line belongs to whom is exactly what is hardest to determine in that moment.
Fixing these takes a couple of minutes at this stage. Left alone, a mislabelled line is generated in the wrong voice in every target language — and if you are also cloning voices, in the wrong person's cloned voice, which is the version of this problem that actually causes a re-render.
Where transcription sits against the alternatives
Transcription is a commodity in the sense that many tools do it. What differs is what happens next.
A standalone transcription tool gives you text and stops. If all you need is a record of a meeting, that is the right tool and probably cheaper.
A captioning tool gives you text formatted for the screen, usually in the source language, and stops there too.
Transcription inside a dubbing pipeline produces a timeline rather than a document: segments that later hold a dubbed line inside their own duration, speaker labels that select a voice per person, and a transcript whose corrections make the next two stages cheaper — 18 credits a minute down to 14 for translation, 25 down to 10 for Standard dubbing.
Which of those you want depends entirely on whether the video is going to travel. If it never leaves its original language, a cheaper standalone tool is a sensible choice. If there is any chance of a second market, transcribing inside the pipeline means that market costs a dubbing pass rather than starting over. The pipeline breakdown shows how the stages connect.
Transcribe once, reuse across every video
This is the part of the pricing that most people never use, and it is the one that changes a bill rather than trimming it.
A transcript is not locked to the project that produced it. From the New file modal you can import an existing transcript — as CSV or SRT — instead of paying to transcribe from scratch. The new project starts with the segments already in place.
That matters in more situations than it first appears:
- The same video, re-cut. A long-form upload and its trailer share dialogue. Transcribe the master once and import into the cut.
- Re-uploads and versions. A corrected export, a version with a new intro, a re-render at a different resolution — same words, new file.
- Series and recurring formats. Recurring intros, outros and legal read-outs are identical across episodes.
- A transcript you already own. If a client or a subtitling vendor already produced one, there is no reason to buy it twice.
Export CSV, not SRT, when you plan to reuse it
Both import. They do not carry the same information, and the difference is the one thing you most want to keep.
| CSV | SRT | |
|---|---|---|
| Timecodes | Yes | Yes |
| Text | Yes | Yes |
| Assigned speakers | Yes | No |
| Best for | Reusing inside the platform | Uploading captions to a player |
The CSV carries a speaker column alongside the timings and the text. An SRT has nowhere to put it — the format simply has no field for who is talking.
So a reused SRT arrives as one undifferentiated voice. You get the words and the timings back, and then you reassign every speaker by hand before dubbing can cast them. On a two-host podcast that is an annoyance; on a panel or a drama it is most of an afternoon, repeated on every episode.
The rule is short: export SRT for players, export CSV for the platform. If there is any chance the transcript will be reused, export both — they cost nothing and only one of them survives the round trip intact.
What reuse actually saves
Importing a transcript removes the transcription charge and moves the later stages onto their reduced rates, because you are supplying work the platform would otherwise do:
| Stage | From scratch | With your imported transcript |
|---|---|---|
| Transcription | 7 cr/min | not charged |
| Translation Deluxe | 18 cr/min | 14 |
| AI Dubbing Standard | 25 cr/min | 10 with the translation too |
| AI Dubbing Deluxe | 37 cr/min | 20 on the same basis |
On a ten-minute video dubbed at Standard, going in with a transcript and a translation takes the job from 250 credits to 100 — roughly $4.40 instead of $11 on the $50 plan. That is a 60% cut on the dubbing rate, and it is the single largest lever available on a bill.
For a series it compounds. Transcribe episode one, correct it properly, export the CSV, and every later episode that shares dialogue starts from corrected text with its speakers already attached — which also means the corrections you made once do not have to be made again in every language.
Where to go next
The other services on the same balance: translation, voice cloning, subtitles and SRT export and review and approval. The pipeline breakdown shows how they connect and what each stage costs.
Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.
Questions people ask
How accurate is AI video transcription?
On clear audio in a well-supported language, expect 90 to 95% before correction. What moves you within that range is the recording rather than the model: microphone distance is the biggest factor, followed by cross-talk, proper nouns and jargon, and music under speech. The transcript arrives editable so those can be fixed before anything downstream runs.
How much does video transcription cost?
Seven credits a minute. A ten-minute video is 70 credits, roughly $3 on the $50 plan, and that includes speaker identification, timecodes and an exportable subtitle file. Credits top up from $10 with no subscription.
Does it identify who is speaking?
Yes. Speakers are labelled and those labels carry through translation and dubbing, so a two-host podcast keeps two distinct voices rather than collapsing into one. Overlapping speech stays as separate lines with overlapping timecodes instead of being flattened into a queue.
Can I upload my own transcript instead?
Yes, and it makes everything after it cheaper. Supplying a transcript drops translation from 18 credits a minute to 14, and supplying both a transcript and a translation drops Standard dubbing from 25 to 10 — a 60% cut on the dubbing rate.
Does the source video have to be in English?
No. Transcription covers the same 80+ languages as the rest of the platform, and translation runs from the source language directly rather than routing through English.