ClipSpark AI

Blog

Why a clip needs its captions burned in

A platform's caption track is not part of the video file, so it does not survive a download or a re-upload. What burning subtitles in fixes, what it costs, and why you still have to read them.

There are two ways to put words on a video, and they are not variations of the same thing. One writes the text into the picture, frame by frame, so it is part of the image data. The other ships the text alongside the video as a separate track that a player reads and draws on top.

The second way is better in almost every respect — it is editable, it is selectable, it can be translated, a screen reader can read it, and a viewer can turn it off. It has one weakness, and the weakness decides this argument: it is not part of the file.

What a caption track is attached to

When you upload a video to a platform and it generates captions, those captions live in the platform's database next to your video. They are rendered by the platform's player. That is why they can be corrected after publishing, and why they appear in search on that platform.

Now download that video. What arrives is an MP4 with a video stream and an audio stream. The captions were never in it. Upload that file somewhere else and there is nothing to display — the new platform will either generate its own from scratch, or show nothing.

This is the whole reason burned-in captions exist. A clip is a file that gets moved: downloaded, sent to somebody in a message, reposted, scheduled through a tool, re-uploaded a month later to a different account. Anything that is not in the picture does not survive that journey.

Why the words matter at all

Two reasons, and only one of them is about sound. A feed plays clips silently until somebody decides to listen, so a talking-head clip with no visible text is a person mouthing at a stranger. And a viewer who does have sound on is often in a room with other noise, on a phone speaker, listening to a recording made on a microphone that was not in a studio.

There is also the plain accessibility case: some of your audience cannot hear the audio at all, and a clip without text excludes them.

What it costs, honestly

Burning subtitles in is irreversible. Once the text is in the picture there is no way to remove it, restyle it, or fix a misspelled name without rendering the clip again — and rendering it again spends credits again. That is the real price, and it is why the choice of look is made before the render rather than after.

Subtitles are a paid option in ClipSpark: a few extra credits on top of the clip. If the render finishes and no subtitles were actually delivered — no speech found in the segment, or the caption step could not produce anything — those credits go back automatically. You are charged for subtitles you got, not for subtitles you asked for.

Three looks, and why they are opinionated

There are three styles rather than a font picker, because a font picker on this particular feature is a way to make an unfixable mistake. All three are designed to stay readable on a phone held at arm's length over a moving background:

  • Clean — plain, unobtrusive, the safe default. It gets out of the way of the picture.
  • Bold — heavier weight and stronger contrast, for footage that is busy or brightly lit behind the text.
  • Punch — the loud one, sized and weighted like the captions on a native short. It is the most visible and the least subtle.

Pick by what is behind the text, not by taste. A style that looks elegant on a dark studio background can be unreadable over a white slide.

You still have to read them

The captions are generated from the recording's own audio, which means they are as good as the audio and no better. The failures are predictable: proper nouns, people's names, product names, jargon, acronyms, and anything said over crosstalk. A clip that confidently misspells your guest's name in fifty-point type is worse than a clip with no captions at all, and it is in the picture, so it cannot be quietly corrected later.

Watch the clip before you post it. That is thirty seconds of your time and it is the last point at which the text is free to change.

What burning in does not give you

Being in the picture is a trade, not a pure win, and the same property that makes the text survive a re-upload makes it invisible to software:

  • A screen reader cannot read it. Burned-in text is pixels, so it is not an accessibility caption track in the sense a standard requires.
  • A platform's search cannot index it, and a viewer cannot copy a quote out of it.
  • It cannot be translated or switched off by the viewer.

So the two are not simply additive. If a platform offers a real caption track and you upload a clip that already has the words in the picture, a viewer who turns that track on sees both — the platform's text drawn over yours. If the clip is only ever going to live on one platform, and that platform generates a track you can review and correct, that is the better option and burning in buys you nothing.

Burn them in when the file is going to travel: when you will download it, send it, schedule it through something, post it to more than one place, or keep it and use it again. That is most clips, which is why the option is there — but ticking it is a decision about where the file is going, not a rule.