Captions — The Complete Guide
Captions is a mobile-first AI video editing app built around talking-head content — subtitles, eye contact correction, and AI presenters.
>
Skip to contentCaptions is a mobile-first AI video editing app built around talking-head content — subtitles, eye contact correction, and AI presenters.
Captions records or imports talking-head video and handles the editing automatically: accurate animated subtitles, removal of filler words and pauses, background replacement, and colour correction.
Its best-known feature corrects eye contact, making it appear the speaker is looking at the camera when they were reading a script off-screen.
Creators and marketers producing frequent talking-head video on a phone — social content, updates, ads, and short explainers.
It is not a general video editor and does not attempt to be.
Reading from a script while appearing to address the camera is genuinely difficult. The correction adjusts gaze direction so the speaker appears to look at the lens.
It works well in good lighting with a steady frame and produces visible artefacts in poor conditions. Review the output before publishing — when it fails, it fails conspicuously.
Captions also offers AI presenters — synthetic speakers delivering a script without filming — and translation with lip-synced dubbing into other languages.
Disclosure matters here. Synthetic presenters and dubbed voices should be identified as such, both ethically and increasingly as a regulatory expectation.
The pattern that works: record on a phone with decent light, import, let the tool cut filler and generate captions, then manually fix caption errors and trim the opening.
Caption accuracy is good but not perfect, and an error in burned-in subtitles cannot be fixed after posting.
Descript is stronger for long-form and transcript-based editing. CapCut is stronger for general mobile editing. Captions is narrower and better at the specific talking-head case.
If most of your output is one person speaking to camera, the specialisation is worth it. Otherwise a general editor covers more ground.
The same features that make this efficient — synthetic presenters, gaze correction, voice cloning for dubbing — sit on a spectrum toward misrepresentation.
Correcting eye contact on your own face is cosmetic. Publishing a synthetic presenter as a real person is not. Decide where your line is before the tools make it easy to drift past it.
Automated video tools produce a competent first cut quickly, and the value depends entirely on what happens next.
They are strong at the mechanical work: finding the segments, cutting to length, burning captions, reformatting aspect ratio. This is genuinely tedious and genuinely automatable.
They are weak at judgement — which moment is actually the interesting one, where a cut lands emotionally, when a pause should be kept. A tool optimising for engagement patterns produces clips that look like every other clip.
So the workable pattern is automate the cut, decide the selection. Let the tool propose, and have a person choose which of the proposals to publish. The time saved is real; the selection is where your channel sounds like yours rather than like the tool.
Check the captions rather than trusting them, particularly for names, figures and anything technical. Burned-in captions are permanent and a wrong figure on screen is worse than no caption.
And keep the source. Automated output is disposable; the recording is the asset, and a workflow that discards the original in favour of the export loses the ability to recut later.
Vendor demonstrations are built on material the tool handles well, so the only evaluation that predicts anything is one run on your own inputs.
Assemble twenty real examples before the trial starts, including the awkward ones — the messy input, the edge case, the one that went wrong last month. A set of clean examples measures a situation you do not have.
Define what good looks like in writing, before you see any output. Deciding afterwards is choosing the answer rather than measuring it, and it is what makes most tool trials inconclusive.
Time the whole task, not the tool. A tool that halves the generation step and adds a verification step has not saved anything. Measure the end-to-end time including checking and correction, because that is the number your team experiences.
Have two people run the same examples. Tolerance for a given failure varies more between people than between tools, and a decision made by one enthusiast rarely survives contact with the team.
And price the failure, not just the licence. What does a wrong output cost here — a correction, an apology, a customer? That number decides how much checking you need, which is usually the real cost of adoption.
Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.
Nothing is sent from here. The link carries only this page’s title and address.