Caption Animation
Turn a transcript or voiceover into word-timed, scroll-stopping captions: each word pops in on the syllable, the active word highlights, and the text stays glued to the narration. Build the kind of captions that drive watch-time on TikTok, Reels, and Shorts — readable, on-beat, and inside the safe area.
When to use
- Word-by-word or karaoke captions on short-form vertical video (9:16).
- Burning open captions onto a voiceover/narration track from a transcript.
- Turning a Whisper/SRT transcript into animated text in Remotion or on the web.
- Restyling auto-generated subtitles into a branded, "Hormozi-style" reveal.
The pipeline
Two non-negotiable rules
- Word-level timing, not line-level. Karaoke reads as magic only when each word lands on the syllable. Always transcribe to word timestamps; never fake them by splitting a line evenly over its duration — drift is instantly visible.
- Readability beats style. Captions are read on a phone, in sunlight, muted. Bold sans-serif, heavy stroke or shadow, high contrast first; decoration second. A caption nobody can read is decoration, not a caption.
Word timing from a transcript
Whisper emits word-level timestamps directly. Normalize every source into one flat token shape so the renderer never cares where the words came from:
If only sentence-level cues exist, re-transcribe for word timing — do not interpolate. Whisper's medium.en is ~2x faster than large and accurate enough for clean voice audio.
Per-word pop-in (Remotion)
Drive each word's scale and opacity off a spring anchored to its own startMs. The micro-overshoot is what makes it feel "alive".
Compute active by testing frame against each token's [startMs, endMs]. Keep the highlight a single accent color or a filled "box" behind the live word — never animate every word's color at once.
Pages, not walls of text
Show 1–4 words at a time. Use Remotion's createTikTokStyleCaptions() (or group manually) and tune combineTokensWithinMilliseconds: ~200–500ms for true word-by-word, ~1000–1500ms for short readable phrases.
Web/CSS pop (no Remotion)
For a DOM/live player, give each word its own animation-delay equal to startMs:
Readable type and safe-area placement (9:16, 1080×1920)
Place captions in the center band — not pinned to the very bottom, where the platform UI (username, audio, buttons) lives. Pad text to a max of ~2 lines.
Sync to a voiceover / narration
The transcript already carries the timing of the exact audio. Lay that same audio under the composition and keep timestamps in the audio's timebase — captions and voice stay locked with zero manual nudging. If the VO is re-recorded, re-transcribe; never hand-shift offsets.
Burn-in vs sidecar SRT/VTT
Emit both: burn the animated captions for social, and write a plain SRT/VTT from the same tokens for accessible/SEO playback.
Output checklist
- Word-level timestamps (real, from transcription) — no even-split fakery.
- Per-word pop with micro-overshoot; one accent color for the active word.
- 1–4 words per page; never a wall of text.
- Bold sans, 56–80px, 2–6px stroke + shadow, high contrast.
- Center band, clear of bottom ~280px and side rails.
- Audio laid under composition; timestamps in the audio timebase.
- Ship burn-in for social and a sidecar SRT/VTT for accessibility.
Deliver & verify (rendered stills → MP4)
Packaged helper (
scripts/): tile your stills withscripts/contact-sheet.sh sheet.png f-hook.png f-mid.png f-end.png, then assert the encode withscripts/probe-mp4.sh out.mp4 [WxH] [fps]. Seescripts/README.md.
Captions ship as a Remotion composition (<Composition> + zod schema + defaultProps) — all word motion frame-driven off useCurrentFrame(), never Date.now() / Math.random() / timers. Deliverable = out/*.mp4 (burned-in) + the project + the sidecar SRT/VTT. 9:16 vertical (1080×1920) is the default.
Verify loop — render stills → inspect → encode. Word timing is the thing that breaks; check it at exact frames before you spend an encode.
Use npx remotion compositions to read durationInFrames/fps and pick the active-word + end frames.
Before you finish:
- Stills render cleanly at frame 0, a mid active-word frame, and last — no missing font/audio.
- The correct word is highlighted at the sampled frame (frame/fps lands inside its token); no even-split fakery.
- Burn-in is legible (stroke+shadow) and the caption is fully inside the 9:16 safe area at every checked frame.
- Frame-driven only — no
Date.now()/Math.random()/ timers. - Shipped props (real tokens, not just
defaultProps) render correctly; MP4 + sidecar SRT/VTT emitted, GIF optional.
Reference files
references/word-timed-captions.md— end-to-end build: Whisper transcription and theCaptiontype, a full SRT parser, Remotion word-timed component with active-highlight, manual paging, an SRT/VTT emitter, per-platform safe-area maps, and a readable-type spec sheet.

