Most platforms treat subtitles and dubbing as two products with two prices. Inside a single pipeline they are the same work seen twice: the transcript carries the timings, the translation carries the text, and an SRT is what you get when you export those two together instead of sending them to a voice.
Which means the honest answer to "what do subtitles cost" is: nothing beyond the stages you already ran.
Subtitles and dubbing serve two different people
This is worth stating because the choice is often framed as either/or, and it is not.
The dub serves the viewer who wants to watch in their language, hands free, eyes on the picture — and the viewer whose reading speed does not match the pace of your edit.
The subtitle file serves the viewer watching muted on a commute or in an open-plan office, the viewer who is deaf or hard of hearing, the viewer who prefers the original performance with help following it — and the search index, which reads text and cannot hear audio.
That last one is the reason a caption file has value even when you have already dubbed. A video with no text attached is opaque to search; a video with an accurate transcript in two languages is not.
What separates a professional SRT from a machine dump
Any tool can put words next to timecodes. A caption file that a broadcaster or a subtitler would accept has to respect constraints the raw transcript does not:
- Reading speed. There is an upper limit to how fast a person reads a line while also watching a picture. A caption that is technically synchronised but too dense to read is a failed caption.
- Line length and line breaks. Two lines maximum, broken at a grammatical boundary rather than wherever the character count runs out.
- Minimum duration. A caption that flashes for a third of a second registers as a flicker, not as text.
- Shot changes. A caption that straddles a cut makes the edit feel broken.
Exported SRTs carry these parameters, which is what makes them usable directly — uploaded to YouTube's subtitle uploader or dropped into an edit suite — instead of needing to be re-timed by hand first.
Where subtitles get harder than audio
Right-to-left scripts. Hebrew, Persian and Urdu are written right to left and the dub does not care — but the caption file does. Mixed content, like a Latin-script brand name or a number inside an Urdu sentence, is where files get mangled, and it is invisible unless you open the export rather than trusting the preview.
Languages without spaces. Chinese and Japanese do not separate words with spaces, so line breaking cannot follow the rules an English caption uses. The East Asian guide covers the Traditional and Simplified split, which changes the subtitle file while leaving the audio identical.
Spelling varieties. Colour or color, organise or organize. The dub is unaffected; the subtitle is not, and a British audience reading American spelling notices. Decide the variety once at the transcript stage and the export inherits it.
Expansion. German and Arabic translations tend to run longer than the English they came from, which collides directly with the reading-speed limit. This is the most common reason a translated caption file needs editing rather than just exporting.
One pass, every language
Because the transcript and translation are stored against the same segments, adding a market adds a caption file rather than a project. Ten languages of subtitles come off the same timeline as the first one, with the same timings and the same speaker attribution.
This is also why subtitles and dubbing compound rather than compete: a channel that dubs into four languages and exports captions in all four is serving eight distinct audiences from one transcription pass.
What subtitles cost
There is no separate subtitle price. You pay for the stages that produce the text:
- Transcription — 7 credits a minute. This alone gives you a subtitle file in the original language.
- Translation Deluxe — 18 a minute, or 14 with your own transcript. This gives you a subtitle file in the target language.
- Dubbing — only if you also want audio. Subtitles do not require it.
So captions in one language cost 7 credits a minute, and captions in a second language cost 25 — transcription plus translation. A ten-minute video captioned in English and Spanish is 250 credits, roughly $11 on the $50 plan, with no audio generated at all.
Credits top up from $10 with no subscription, which covers a short video end to end. The pipeline breakdown shows where each stage sits.
SRT, and when you need something else
SRT is the format almost everything accepts — YouTube, Vimeo, the major edit suites, most players. It carries sequence numbers, start and end timecodes, and the text. That is all, and for the large majority of work that is exactly enough.
What it deliberately does not carry is styling: position on screen, colour, italics for off-screen speech. Broadcast subtitling formats do carry those, and if you are delivering to a broadcaster you will be told which one they want. For a channel, a course, a marketing video or a client deliverable, SRT is the right answer and adding complexity buys nothing.
The practical consequence is that a caption file should be judged on its timing and its line breaking rather than its format, because the format is rarely the constraint.
Where captions fit in a real workflow
Two orders of operations are common, and they are not equivalent.
Captions first, dub later. Transcribe, correct the transcript, export the original-language SRT, and publish. The video is now searchable and accessible for 7 credits a minute, with no audio generated. If the analytics later show a market worth serving, the translation and dubbing stages run on a transcript you have already corrected — so the expensive stages start from clean input.
Dub first, captions as a by-product. Run the whole pipeline and export caption files alongside the audio. Nothing wrong with it, and it is the right order when you already know which markets you want.
The first order is underrated. It is the cheapest thing on the menu, it produces two assets a search index can read, and it front-loads the correction work at the point where correction is free. A channel that captions everything and dubs selectively is spending its credits in roughly the right places.
Four ways caption files go wrong
1. Trusting the preview instead of the file. A caption that renders correctly in a web preview can still be malformed in the exported SRT, especially with right-to-left scripts or mixed Latin content. Open the file.
2. Translating captions separately from the audio. If the subtitle text and the dubbed audio come from two different translation passes, they will disagree — and viewers who have both on notice immediately. Running both off the same stored translation keeps them consistent.
3. Ignoring reading speed after translation. A caption that fit comfortably in English can become unreadable in German or Arabic at the same timecode. Expansion is a subtitle problem before it is an audio problem.
4. Uploading the original-language file to every market. More common than it sounds on channels that add languages gradually. Each language needs its own file, generated from its own translation, off the same timeline.
Captions, subtitles and SDH are three different things
The words get used interchangeably and they do not mean the same thing, which matters because clients and platforms ask for them by name.
| Term | What it is | Who it is for |
|---|---|---|
| Captions | Text in the same language as the audio | Viewers watching muted, or in noisy places |
| Subtitles | Text in a different language from the audio | Viewers who do not speak the source language |
| SDH | Captions plus non-speech information — speaker labels, sounds, music cues | Deaf and hard-of-hearing viewers |
The practical consequence: a transcription pass gives you captions. A transcription plus a translation gives you subtitles. SDH is captions with extra information added, and the speaker labels that come out of transcription are most of what that requires — which is why the labels are worth correcting even on a project you do not intend to dub.
If a client asks for "subtitles" and means captions in the original language, you have been quoted for a translation you do not need. It is worth asking which one they mean.
Where the file goes, platform by platform
SRT is accepted essentially everywhere, but what each platform does with it differs enough to matter.
YouTube accepts SRT per language and lets viewers switch. It will also auto-generate captions if you upload nothing, and those are noticeably worse than an SRT you corrected — particularly on proper nouns, which is exactly where a viewer searching for your product will be looking. Uploading your own file replaces the automatic one.
Instagram, TikTok and Shorts lean toward burned-in text rather than sidecar files, because most viewing happens with sound off and without the interface controls a long-form player has. For these, the caption file is the source you burn from rather than the thing you upload.
LinkedIn and Facebook accept SRT, and both are worth captioning, because a large share of feed video plays muted by default.
Edit suites import SRT to burn or to restyle. This is where the timing quality shows: a file with clean line breaks needs no adjustment, and a machine dump needs re-timing line by line.
Burned-in or sidecar?
A sidecar file — an SRT alongside the video — can be switched off, switched between languages, and read by a search index. A burned-in caption is part of the picture: it always shows, in one language, and no index can read it.
Use sidecar for anything on a platform with a subtitle track: YouTube, Vimeo, a course platform, a client deliverable. It preserves the viewer's choice and it gives you the search benefit.
Use burned-in for vertical social video, where the player often has no subtitle control and the viewing context is silent by default.
Do not use both at once in the same language. Two sets of text on screen is the most common self-inflicted caption problem, and it happens most often when a burned-in video is later uploaded somewhere that adds automatic captions on top.
Since the caption file comes out of the same timeline as the dub, producing both from one project is normal: a sidecar SRT for the long-form upload and a burned-in vertical cut from the same corrected text, with the transcript as the shared source.
How long captioning actually takes
Worth setting expectations, because the generation step is not the slow part.
Generation is minutes. A ten-minute video transcribes and exports an SRT without you doing anything.
Correction is where the time goes, and it scales with the recording rather than the length. Clean audio with one speaker and no jargon might need a couple of minutes of fixes across ten minutes of video. A panel discussion full of names and technical terms can take as long as the video itself the first time, and much less afterwards once the names are in the glossary.
Translation review adds a similar pass per language, though it is faster than the first correction because the segments and timings are already right — you are reading for register and phrasing, not for accuracy of transcription.
The order that saves the most time is to correct once, thoroughly, in the source language. Every error left in the transcript is an error that gets translated into each target language and then has to be fixed once per language rather than once in total.
Reading a caption file critically
Four things to look at when judging whether an exported file is ready to ship:
Does any caption exceed two lines? Three lines cover too much of the picture and push the viewer's eyes down and away from the action.
Do the line breaks fall at sensible places? A break between an article and its noun, or between a preposition and its object, forces the reader to hold an incomplete phrase across a line. Breaking at a clause boundary reads without effort.
Does any caption appear for less than about a second? Below that, the text registers as a flicker rather than as words, and short interjections are the usual culprits.
Does any caption straddle a cut? Text that persists across a shot change makes the edit feel wrong even to viewers who could not say why.
Exports are built to these conventions, so most files pass without adjustment. It is the translated versions that need the closest look, because expansion into German or Arabic pushes against the reading-speed limit that the English original had room for.
One last thing worth saying plainly: captions are the cheapest thing on this platform and the most frequently postponed. A video that is captioned in its own language is already searchable, already accessible, and already halfway to a second market — because the transcript a caption comes from is the same transcript the translation stage starts from.
Where to go next
The other services on the same balance: transcription, translation, voice cloning and review and approval. The pipeline breakdown shows how they connect and what each stage costs.
Choosing a platform? The side-by-sides are honest about where the other one wins: vs HeyGen, vs ElevenLabs, vs DittoDub, vs Maestra, vs Dubverse, vs CAMB.AI and vs Rask AI, or all of them side by side.
Questions people ask
How do I generate an SRT file from a video?
Upload the video and run transcription. That produces a segmented transcript with timecodes, which exports directly as an SRT in the original language for 7 credits a minute. To get a subtitle file in another language, run the translation stage on top for 18 a minute, or 14 if you supply the transcript yourself.
Do I have to pay for dubbing to get subtitles?
No. Subtitles come from the transcript and the translation, and the dubbing stage is optional. Captions in one language cost 7 credits a minute and captions in a second language cost 25, with no audio generated at all.
What makes an exported SRT professional rather than raw?
Reading speed limits, a maximum of two lines broken at grammatical boundaries, minimum caption duration, and captions that do not straddle shot changes. Exports carry these parameters so the file can be uploaded to YouTube or dropped into an edit suite without being re-timed by hand.
Can I get subtitles in several languages from one upload?
Yes. The transcript and translations are stored against the same segments, so each additional language is a caption file off the same timeline rather than a new project — the same timings and the same speaker attribution.
Do subtitles help with search?
They give a search index text where it otherwise has only audio, in a language the video previously had none in. That is why a caption file has value even on a video you have already dubbed.