Agent skill

Embedded Video Captions

by heygen-com in heygen-com/hyperframes

Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.

Apache-2.0Auto-check passedMedia & Creative

Install Embedded Video Captions

skills CLI
$ npx skills add heygen-com/hyperframes --skill embedded-captions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install heygen-com/hyperframes embedded-captions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/embedded-captions .claude/skills/embedded-captions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
embedded-captions
GitHub stars
59k
Used in
3 other repos
Token cost
~8.6k tokens
SKILL.md length
3,708 words
Files
160 (incl. scripts, references, assets)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.

  • Adding plain verbatim subtitles to a talking-head video
  • SKILL.md covers Runtime prerequisites, Operational flow (TL;DR), Caption model — rail + embed and Step 0 — pick ONE identity…, plus 9 more sections
  • Calls node, ffmpeg and npx
  • Making cinematic captions that sit behind the speaker

What it does

The skill adds subtitles to an existing talking-head video and picks a look from a catalog of 35 styles. Standard, the default, builds a quiet rail of verbatim lower-third captions plus one embedded caption composited behind the subject at the peak. Cinematic embeds every caption behind the subject, and Theme is a full themed treatment for effects-heavy requests, with themes named ordnance, terminal, neonsign, stardust and stomp. The skill warns that embedding every word is a common mistake.

The whole workflow runs locally, including transcription and subject matting, and rendering waits for the HyperFrames CLI to exit before compositing. Footage with several shots must be split first. Setup installs Sharp, Puppeteer with its Chromium browser and GSAP in the caption project, Bash and FFmpeg or ffprobe must be on the PATH, and matting and transcription may download their own models on first use.

The folder ships a style catalog, preset files for looks such as documentary, editorial, glitch and neon, and bundled font and stroke-font assets. The skill also suggests refreshing itself with npx hyperframes skills update embedded-captions, after confirming with you.

When your agent uses it

  • Adding plain verbatim subtitles to a talking-head video
  • Making cinematic captions that sit behind the speaker
  • Applying a named caption style to explainer or voiceover footage
  • Requests for flashy, effects-heavy subtitles, including ones phrased in Chinese

Example prompts

  • “Add clean subtitles to ./clips/intro.mp4 without changing the footage.”
  • “Put cinematic captions behind the speaker in my product demo video.”
  • “Give this explainer video the documentary caption style from the catalog.”
  • “Make the captions on talk.mp4 look like VFX, with the key line sitting behind my head.”

Requirements

  • Bash and FFmpeg or ffprobe on the PATH
  • npm, to install Sharp, Puppeteer and GSAP in the caption project
  • The HyperFrames CLI (bundled with plugin installs)

What it can do on your machine

Read from SKILL.md and the folder at commit 0c76e52. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • ffmpeg
    • npx
    • bash
    • npm
    • ffprobe
    • python3
    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, npm and uvx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Embedded Video Captions loads about 8.6k tokens when it runs, and up to ~53k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 3,708 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~8.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~53k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from heygen-com/hyperframes at commit 0c76e52, republished under its Apache-2.0 licence (© heygen-com). 3,708 words, ~8,558 tokens.

Download SKILL.mdSave it as .claude/skills/embedded-captions/SKILL.md (or your agent's skills folder). This skill also uses 159 other files; get the full folder from GitHub.
name
embedded-captions
description
Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.

Plugin installs: Before setup or freshness commands, follow plugin execution rules when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below.

First, keep this skill fresh — confirm with the user before running: npx hyperframes skills update embedded-captions. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them.

Embedded Captions

One catalog, picked up front (CATALOG.md — 35 identities; the engines behind it are backend detail). Standard (default) builds a clean verbatim rail (lower-third subtitle carrying most text) + an embed climax composited into the scene behind the subject at the peak. Cinematic is pure embed — no rail, every caption composited behind the subject (hero typography, accumulation, occlusion as the effect). Theme is a complete themed constitution — body paradigm × hero setpiece × front fx × plate reaction, composed from registries (themes/README.md): ordnance terminal neonsign stardust stomp. Most explainer / voiceover is Standard; embed is the scarce, earned peak — embedding every word is the common mistake; Theme is for VFX-grade asks ("炸", "特效", "像 AE 做的").


Runtime prerequisites

Plugin installs use the bundled, manifest-pinned CLI for matting, transcription, and rendering; no source checkout is required. The local preview and caption measurement helpers also need Sharp, Puppeteer (with its Chromium browser), and GSAP. Install these in the caption project, not inside the read-only plugin:

bash
npm install --prefix <project> --save-dev --save-exact sharp@0.35.3 puppeteer@25.8.0 gsap@3.15.0

Keep the project's lockfile. If these dependencies already exist, use its locked versions instead of overwriting them. Bash and FFmpeg/ffprobe must be on PATH. Matting and transcription may download their own models on first use.

Rendering waits for the CLI to exit successfully before compositing. The old HF_TIMEOUT_S shell watchdog is no longer used: a large partial file is not proof that rendering finished. An explicit built-checkout argument or HYPERFRAMES_ROOT selects the contributor CLI instead of the plugin pin. Cancel a stalled render normally through the CLI/terminal; the caption helper does not force-kill or recover a render from a process snapshot.

Operational flow (TL;DR)

Routed through /hyperframes, the intent layer confirms only the input (which clip) and announces the identity pick as a deferred ask — the shortlist needs the probed clip, so it stays at step 1 below; the layer's run-shape questions don't apply (the footage is untouched, there is no storyboard to review). A BRIEF.md, when present, carries the confirmed input and any user notes — read it first.

The craft prose below is long; the pipeline itself is short — and everything deterministic is computed or compiled, never hand-written:

  1. Decision gate (refuse bad clips) → pick ONE identity from CATALOG.md (35 identities; engine/compiler derived by lookup — never surface a mode/category question)
  2. hyperframes init (skip it if the project dir already exists with the video inside — matte.cjs/transcribe.cjs adopt any video in the dir as source.mp4) → bash scripts/prepare.sh <project> (matte ∥ transcribe ∥ audio-envelope in parallel, then safe-zones v2 with scene palette/optics/lighting — one command, nothing forgotten)
  3. author a small JSON of creative choices (read safe-zones.json first): Cinematic → cinematic.json → make-cinematic.cjs (derives plan.json and compiles it); Theme → theme.json → make-theme.cjs (rail/panel/poem/takeover paradigms; anchor is the quiet rail default)
  4. Visual QA: node scripts/preview-frames.cjs <project> → faithful composite previews in ~2s/frame (no render). Check § Visual QA before paying for a render.
  5. render-and-composite.sh → gates (timing / occlusion+hero / overflow / hand-off) → final.mp4

Load-bearing rules people miss:

  • rail (default) + embed (promotion). drop (filler, not shown) / rail (verbatim lower-third subtitle, in front, carries most text) / embed (a peak word composited behind the subject). Standard mode does both, embedding only the peak(s). See § Caption model.
  • The video is delivered UNTOUCHED (Standard/Cinematic; Theme mode's PLATE budget is the one sanctioned exception — register-gated reaction beats (charge-dim, punch, shake, grain) defined per theme DNA and applied AFTER the matte composite so subject+text+plate move as one frame) — captions are the only thing added; the matte just lets the subject occlude the embed track. Never grade/recolor/scanline the footage.
  • Two rulebooks: rail → references/rail.md (thin), embed craft → references/composition-craft.md (rich, embed-only). Skim by need.

Caption model — rail + embed

Every spoken phrase is one of three things:

WhatHow it's shown
dropfiller — um/uh, stutters, self-correctionsnot shown
railthe default — ordinary spoken content (verbatim)clean lower-third subtitle, in front, readable. A punch word can get an inline emphasis highlight (accent colour / active-word pop) — it stays on the rail.
embeda promoted peak — the headline beatone big word composited behind the subject (matte occlusion), designed entrance + exit

The rail carries most of the text; embed is the scarce, earned peak. Scarcity is per beat/block, not per clip: ≤1 hero per block (thought), never two co-visible, ≥ a beat of air between hero windows (the compiler warns under 0.6s). A short clip → usually 1–2; a long explainer → ~one per section. Among multiple heroes, the largest authored one is the APEX (it alone gets the full lockup embed + width-fit raise); smaller ones are MINOR peaks that ride their column as oversized emphasis lines (fg, damped motion) — not every beat needs the matte showcase, which is exactly what keeps the apex an event. Embedding every word is still the common mistake.

Rail-surface identities build exactly this (rail = rail.html, embed = the climax in index.html). Column-flow identities drop the rail and make everything embed-style — recommend them only for mood-over-verbatim asks, never for explainer / voiceover where the words must read (CATALOG.md encodes this per identity).


Step 0 — pick ONE identity from the CATALOG

One front-end, three engines behind. The user picks an IDENTITY from CATALOG.md (35 entries: 10 classic + 25 themed); the engine, compiler and authoring file are derived by lookup from the catalog row. Never surface "Standard vs Cinematic vs Theme" as a question — those are backend names (a product has one UX even with several engines). The catalog encodes everything routing needs: reading surface, voice, recommend-for, scene needs, adjacency notes for the genuinely-close pairs (loud↔ordnance, neon↔neonsign, cream↔stardust).

The identity pick is a preference gate (../hyperframes/references/brief-contract.md § 1): in autonomous mode ("surprise me" / "decide for me"), pick from your shortlist yourself and state the one-line why instead of asking.

Procedure: probe the clip → shortlist 2–3 identities from the catalog → recommend ONE with a one-line why → the user picks (autonomous mode: you pick, stating the why) → author that identity's file. Identities are engine-locked (no cross combos; opening one is a validation event — see dna/README.md).

Always present your recommendation and let the user pick before you author. Don't silently default.

(The full identity table lives in CATALOG.md — single source of truth for routing. The engine docs below describe each backend's authoring contract.)

CATALOG.md is the whole answer space here: this workflow does not search the HyperFrames component registry. The composition workflows run npx hyperframes catalog before authoring a named look; this one must not. Its engines are locked compilers that consume cinematic.json / theme.json and emit the composition themselves, so a registry item — the caption-* blocks included — has nothing to mount into. A registry block styles text on a designed canvas; this skill burns captions into somebody's footage through a matte. When no identity fits the ask, say so and pick the nearest, rather than reaching outside the catalog.

Recommendation heuristic: use the "Shortlisting heuristics" in CATALOG.md — they are identity-level (e.g. "炸" shortlists ordnance/stomp/terminal/loud and picks by WHAT should explode), never category-level. Unsure → anchor.

  • Cinematic → write cinematic.json for a locked template, compiled by make-cinematic.cjs.
  • Theme → read themes/README.md, author theme.json, run scripts/render-theme.sh (compiles + renders + plate reaction → final_fx.mp4).

Decision gate — RUN FIRST

Probe the video and classify the scene before either mode.

bash
ffprobe <video.mp4>                    # specs
ffmpeg -ss <t> -i <video.mp4> -vframes 1 sample.png   # at 20/50/80%

Read the samples. Refuse if:

  • Multiple speakers / hard cuts (split & render each shot, or refuse)
  • No human subject (this skill is for talking-head)
  • Under 3 seconds, no speech, or face never clearly visible — transcribe.cjs warns when audio is near-silent (Whisper hallucinates words like "Thank you." over silence); heed it and refuse rather than caption fabricated words
  • Source already has burned-in captions / subtitles / heavy text graphics — adding a second caption system conflicts and the footage ships untouched (no covering/inpainting). Burned text often appears only mid-clip: sample a 1fps contact sheet (ffmpeg -i in.mp4 -vf "fps=1,scale=160:-1,tile=10x5" sheet.png), don't trust 3 spot frames.
  • Transcript is garbage — non-native/heavy-accent speech can transcribe into confident gibberish. Sanity-read transcript.json before authoring; if it doesn't parse as language, try WHISPER_MODEL=medium once, else refuse (a verbatim rail of fabricated words is worse than no captions).
  • Busy handheld with fast motion (matte flickers)
Pre-flight probes (cost nothing, prevent the worst failures)
  1. Shot-cut probe. Sample frames at 20%, 50%, 80%. If a different subject/scene appears, trim the clip before the cut.
  2. Letterbox / pillarbox probe. Black bars on the first frame? Compute safe content rect and constrain caption placement inside it.
  3. Luminance probe. Sample the caption region's average luminance — under 60 → light text reads as-is, 60-180 → add the glyph scrim, 180+ → opaque text + scrim (never bare light text). Cinematic templates are cream+screen and LOCKED — use this probe to pick a fitting identity (bright scenes → ink, or the opaque-rail anchor theme), never to recolour one.
  4. Identity recommendation by tone (you recommend; the user picks — see Step 0 + CATALOG.md). explainer / interview / must-read words → rail/panel-surface identities; poetic / social / "cinematic" → column-flow identities by register; "炸 / 特效 / VFX" / named worlds → themed identities. When unsure → anchor (words read, scene safe) — but present a shortlist and let the user choose.

Pipeline — 5 steps

1. hyperframes init <project> --non-interactive --video <video.mp4> --skill=embedded-captions
2. bash scripts/prepare.sh <project>       # matte ∥ transcribe (parallel) → safe-zones. One command.
                                           #   → frames_fg/ transcript.json safe-zones.json
3. [AGENT STEP — the only creative step] author a small JSON; see below by mode
   Cinematic: author cinematic.json → node scripts/make-cinematic.cjs <project>
   Theme:     author theme.json → bash scripts/render-theme.sh <project>   (compiles + renders + plate fx)
4. node scripts/preview-frames.cjs <project>   # ~2s/frame composite previews → § Visual QA (BEFORE the render)
5. bash scripts/render-and-composite.sh <project>  # gates → final.mp4 + history/ snapshot
   (Theme mode: SKIP steps 3b/5 — render-theme.sh already runs compile + render-and-composite
    + _postfx.sh; the deliverable is final_fx.mp4, final.mp4 is pre-plate-reaction)

Step 1's init checks the installed skills against the latest on GitHub and updates the global set if any are out of date.

Step 3 differs by mode:

Step 3 — Cinematic mode (pure embed)
  1. Read safe-zones.json first. Narration planes go in zones.hugLeft/hugRight — clean strips ABUTTING the silhouette (text far from the body reads as floating, not embedded; far corners are the fallback, not the default). The hero defaults to heroAnchor/heroBands.best (centered ON the subject, ~30–55% occluded). recommendation:"fg" moves NARRATION in front for legibility; the hero stays embedded whenever heroBands.feasible — hero-fg is the last resort.
  2. The DNA is the identity you picked in Step 0 (CATALOG.md) — do not re-open the choice here. Sanity-check it against the scene (bright hero band luma > 150 wants ink; full pick guidance lives in the catalog, covering all ten incl. neon / glitch / chrome / velocity). State your pick + why; the user decides. The DNA locks type/palette/blend/motion + hero three-act; safe-zones v2 (palette/optics/lighting) parameterizes it to THIS scene automatically.
  3. Author <project>/cinematic.json — "dna": "<name>" + thought-BLOCKS, not raw groups: each block = lines of words (grouped 2–5 at clause boundaries) + the plane it stacks in + per-line css (size/weight/style only — no positions) + at most ONE line marked "hero": true (the promoted word; "text" for display form). Schema: scripts/make-cinematic.cjs header.
  4. Compile: node scripts/make-cinematic.cjs <project> — lowers blocks → plan.json → index.html. Generated for you: transcript-sequenced timings, accumulate-within-block, page-flip-between-blocks, the hero LOCKUP (a hero block's pre-context, HERO and post-context stack as ONE bonded composition centered on the subject — reading order top→bottom = spoken order by construction; context floats in FRONT while the hero embeds BEHIND = the depth sandwich; a mass rule keeps the hero dominating its context), apex/minor hero split, reading order by construction, fg fallback per safe-zones. Then the gates run as usual. (Hand-authoring plan.json directly remains possible for designs blocks can't express — then run fill-timings.cjs + fit-fonts.cjs + make-composition.cjs yourself.)
Step 3 — Theme mode (themed constitution)

Read themes/README.md FIRST — paradigm/setpiece registries, linkages, hard rules, and the exact theme.json schema.

  1. Pick a theme DNA by content register (each themes/<name>.json has voice + when). State your pick + why; the user decides.
  2. Author <project>/theme.json — dna, lines (verbatim, transcript order; 1–5 words each — for takeover each line is one CARD), minors (emphasis words), hero:{match} (the climax word/phrase; leave it OUT of lines for embed setpieces, keep it IN for inline setpieces and panel+redact).
  3. Render: bash scripts/render-theme.sh <project> — compiles (verbatim-completeness gate at compile time), renders both layers, composites, applies the plate reaction → final_fx.mp4. Use preview-frames.cjs between compile and render for Visual QA.

Visual QA — preview BEFORE you render

node scripts/preview-frames.cjs <project> [t…] composites faithful preview frames in ~2s each (caption layers screenshotted at seek-time + real video frame + matte occlusion + rail overlay = what the final composite will look like at that moment). Default samples = each group/climax window. A full render costs minutes — never use it to discover layout problems.

Check the previews (<project>/preview/sheet.png) against this list — these are the failures the geometric gates cannot catch:

  1. Washout — light text over a bright region (window/sign/sky): unreadable → move the plane or change DNA/mode (bright scene → ink).
  2. Text-on-text — captions over the scene's own text/graphics, or two caption groups colliding.
  3. Reading order — on-screen vertical order must match spoken order; the hero must not sit below later words.
  4. Hero presence — the climax should be BIG and visibly behind the subject (~30–55% occluded), not a floating label in a margin.
  5. Balance — one coherent column/band, not scattered fragments; margins breathing; nothing clipped.

Then the 5 positive checks in references/reference-bar.md (poster test · timid test · one-glance hierarchy · scene handshake · dead-air audit) — the failure list keeps a render from being broken; the positive list is what makes it designed. Ship when both pass.

Fresh-eyes review (recommended for anything user-facing): you have confirmation bias about your own layout. If you can spawn a subagent, give it ONLY the preview sheet + this checklist and ask for PASS/FIX verdicts per frame ("review these caption previews against the 5-point checklist; answer PASS or the specific fix per frame"). Apply fixes in cinematic.json / theme.json, recompile, re-preview — each loop costs seconds. Render once, when the previews pass.


Show full SKILL.md (1,509 more words)Show less

The DNA registry — ten visual languages (replaces the template catalog)

Both modes draw from dna/ — ten art-directed visual languages that parameterize per scene (accent sampled from the footage, contact shadow along the measured light direction, depth-match blur, RMS-coupled hero amplitude):

DNARegisterScene fitVoice
creampremium-warmdark/mid warm scenesInter + warm cream + screen; glowing emergence hero (successor of cinematic-cream)
inkpremiumbright scenes (luma > 150)near-black multiply — type printed ON the wall; the bright-scene answer
editorialeditorial-luxeintrospective / fashion / poeticBodoni Moda, lowercase-italic hero — magazine elegance
keynotetech-premiumproduct / launchopaque white Inter 800, dead-center stillness
documentaryformalinterview / seriousburn-in reveals, no hero — gravitas IS the style
loudloudhype / sport / socialAnton + scene-sampled accent, single-unit slam + ripple; body ANNOUNCES in front (bodyLayer: fg)
neonloud-neonneon-noir / nightlife / tech-noir (dark scenes)electric-cyan signage, ignition flicker, the hero powers ON like a sign
glitchloud-neondigital / hacker / AIRGB-split echoes snap together on landing; machine-percussive timing
chromeloud-luxeY2K / fashion-tech / musicliquid-metal gradient hero + one sheen sweep during the hold
velocityloud-sportsport / auto / fitnessevery word arrives along its motion vector (streak+skew), hero passes with speed trails

Pick by safe-zones.json (heroAnchor.bandLuma, palette.temperature) × content register — dna/README.md has the decision rule. Authoring: cinematic.json takes "dna": "<name>".

The engine generates the hero three-act from the DNA (no authoring needed): co-visible captions dim (setup) → per-letter entrance with amplitude ∝ spoken loudness (impact) → breathe + glow until exit (afterglow).

(Legacy: plan.template:"cinematic-cream" maps to dna:"cream" automatically. The retired 54-template library is archived outside this repo and is not distributed with the skill; _motion.md remains in-skill as the motion-verb reference catalog.)


Aesthetic decision — tone × shot × platform (input to the catalog shortlist, NOT a second router)

Classify the clip on 3 axes and feed the result into CATALOG.md's shortlisting — this section never picks a mode/engine by itself:

Tone (what feel does the content have?)

  • documentary | conversational | energetic | poetic | keynote | investigative | music-video

Shot (what's the framing?)

  • close-up (head + shoulders) | mid-shot (torso+) | wide (full body+) | cut-montage (mixed shots)

Platform (where will it play?)

  • 9:16 portrait (TikTok/IG/Shorts) | 16:9 landscape (YouTube/web) | 1:1 square | broadcast export

Cross-reference in references/direction-catalog.md § Classification matrix for direction language — then return to CATALOG.md to shortlist identities (this matrix informs the shortlist; the catalog is the only routing surface).

Composition craft (embed track) — read before embedding

The full embed-track playbook lives in references/composition-craft.md: transcript role-annotation, phrase grouping, planes & clean-zone anchoring, zone coherence, climax pop & readability, edge-breathing, the occlusion 3-step judgement, and accumulation/persistence. It governs how a promoted phrase sits INTO the scene — read it before authoring any embed (Cinematic cinematic.json or Standard index.html). The default rail track has its own, much simpler spec → references/rail.md.


Shared knowledge

DocWhat
references/rail.mdThe rail track — standard lower-third subtitle spec (the default; carries most text).
references/composition-craft.mdThe embed-track playbook — grouping, planes, climax pop, occlusion judgement, accumulation/persistence. Read before embedding.
dna/README.mdThe DNA registry — ten scene-parameterized visual languages; how to pick.
references/reference-bar.mdThe taste bar — per-register world-class references + the 5 positive checks.
references/aesthetic-principles.mdThe 18 rules. Beat Veed AI on taste. Read first.
references/motion-vocabulary.md10 named motion primitives + tone→timing lookup
references/direction-catalog.md10 ship-ready aesthetics + tone×shot×platform matrix
references/anti-patterns.mdBugs already locked out (CoreML, letter-spacing reflow, etc.)
references/scene-types.mdWhen a wall surface is usable (4 conditions)
references/layout-heuristics.mdPlane positioning, clean-zone selection, crown 3 conditions, pillarbox math
references/typography-presets.mdFont-size × column-width matrix (starting points)
references/caption-grouping.mdWord → group rules (pauses, sentence boundaries)
references/failure-modes.mdLong tail of dev gotchas
references/bespoke-vs-presets.mdWhy presets fail sometimes; clone-and-tweak pattern

Read the aesthetic principles and direction catalog FIRST. Everything else is implementation detail.


Non-negotiables

  • Face must never be 100%-covered continuously — every 0.3s window, face bbox ≥30% uncovered.
  • WCAG contrast — final render lints; fix palette if it fails.
  • Deterministic — no Math.random(), no Date.now(), no repeat:-1.
  • Never grade/recolor the video. The footage ships untouched — captions are the only addition. No full-frame scanlines / duotone / darken / vignette over the a-roll. neon-noir/CRT texture belongs inside a caption element, not over the whole frame.
  • Rail-first for talking-head / explainer. Don't embed the whole transcript — most text is the rail; embed only peaks. Embedding everything is the default mistake.
  • Embed is scarce + spaced. ≤1 embed per sentence/beat, never two adjacent or co-visible, ≥ a beat apart, at most one apex. climax = per-beat peak, not "the single payoff of the entire clip."
  • Matte = the PERSON (hyperframes remove-background, u2net_human_seg, Apache-2.0). Human segmentation by intent, but not surgically: thin offset furniture (mic boom arms) is usually excluded — captions render over it, behind the person — while large salient objects NEAR the subject (a telescope, a desk rig) can still leak into the matte and occlude captions. Objects HELD by the subject (products, phones) may drop out intermittently, letting captions pass in front. NEVER assume: sample frames_fg/ at 2-3 timestamps before placing the hero, and prefer hero positions clear of any leaked furniture (heroAnchor can be skewed by leaks — cross-check against frames_bg).
  • safe-zones is PROP-BLIND — eyeball every band you use. Zones/heroBands score subject occlusion + luma only: a mic, telescope, or screen sitting inside a "clean" zone is invisible to them (and a prop leaking INTO the matte skews heroAnchor.centerXPct off the person). Before authoring, extract ONE frame of each band you intend to use; if a prop lives there, measure its bbox and move/shrink the plane. Two real cases shipped clean only because the agent did exactly this. (Auto prop-saliency is a known gap; zones' peakLuma only catches moving bright objects.)
  • Captions stay on-frame. Cinematic mode hard-gates frame-overflow; Standard mode runs check-overflow.cjs as a WARNING (intentional bleed is the only exception — read the warning).
  • Each caption ≥ 0.5s on screen — shorter = unreadable.
  • Word timings must match transcript.json within 80ms — a caption firing 500ms off-beat destroys the scene illusion. Cinematic runs check-timing.cjs --strict before rendering (via render-and-composite.sh); THEME mode enforces the same timings at compile time instead (make-theme's sequential transcript matcher + verbatim completeness gate — drift is a compile error). Never pack multiple transcript words into one entry (e.g. "FUTURE OF" or an IT + line-break + ALL stack with one start/end) — the second word inherits the first's timestamp and fires early. Split them into separate word entries with their own timings, even if you want them on the same visual line (use CSS white-space / natural wrap instead of <br>). Creative substitutions where caption text ≠ transcript (e.g. "15%" replacing "fifteen percent") are supported — register them in CREATIVE_SUBS inside check-timing.cjs.
  • Group windows must envelop their words — group.in ≤ min(word.start) and group.out ≥ max(word.end) for every group. If group.in is later than a word's start, the word is silently delayed until the container mounts (we've shipped 800ms lag bugs from this). The validator enforces this.
  • No two caption groups may overlap in both time AND screen region — overlapping-in-time captions create text-on-text pileups. Options: (a) spatial separation — place each group in a non-overlapping vertical band so they can coexist (memory-wall cascade style); (b) handoff — set the earlier group's out ≤ the next group's in so only one is on screen; (c) deliberate layered typography — add "allow_overlap": true on one of the groups to silence the validator. The validator estimates each group's vertical bbox from its CSS and flags collisions. Pick (a) by default — it's what makes cinematic-cream feel like a poem accumulating, not a subtitle track replacing itself.
  • Screen-blend fails on bright backgrounds (>180 luminance). Cinematic templates are cream + screen and that DNA is locked (the plan can't recolour them) → on a bright backdrop they wash out, so pick ink (letterpress built FOR bright surfaces) or the anchor theme (opaque rail surface) rather than overriding a look.
  • Don't animate letter-spacing or filter:blur on word entrance — inline-block reflow causes line-jumps.
  • CoreML banned for matting — the onnxruntime CoreML EP's mixed-precision partitioning corrupted face alpha (observed with the previous RVM engine; don't re-try it). Matting is CPU-only (~2 fps @1080p ≈ 2-3 min per 10s clip; budget for it on long clips).

Dependencies

  • HyperFrames CLI: plugin installs use the bundled manifest-pinned launcher. Source contributors can use a built checkout (packages/cli/dist/cli.js) via HYPERFRAMES_ROOT, the skill’s source tree, or ~/Downloads/hyperframes.
  • Node-first; two Python touchpoints via uvx (no manual installs): transcription runs WhisperX through uvx (word-level timings; falls back to an existing word-level transcript.json), and Theme's drawon setpiece shells python3 scripts/gen-stroke-path.py at compile time. Everything else runs on the toolchain hyperframes already ships: matting via the hyperframes CLI's remove-background (u2net_human_seg; weights auto-download once, ~168 MB, to ~/.cache/hyperframes/), image/alpha math via sharp, layout/occlusion/overflow via puppeteer, plus ffmpeg. Install Sharp, Puppeteer, and GSAP in the caption project as described in Runtime prerequisites above. The helpers check that project first and retain checkout dependency lookup for source contributors.
  • Transcription = WhisperX via uvx (word-level timings + alignment; no manual install — transcribe.cjs drives uvx whisperx). Falls back to an existing word-level transcript.json if present.
  • Source video — matte.cjs / transcribe.cjs auto-resolve source.mp4 (or glob the clip / read hyperframes.json), so hyperframes init --video X.mp4 needs no manual rename.
  • fps — matte.cjs extracts at the source's native rate and records matte.fps; render-and-composite.sh uses that so the matte stays frame-aligned.
  • Matting weights are NOT bundled: matte.cjs shells the hyperframes CLI's remove-background, which downloads u2net_human_seg (~168 MB, Apache-2.0) once to ~/.cache/hyperframes/background-removal/models/. First prepare on a fresh machine needs network for that one download.

If a hard dependency is missing, STOP and ask the user — don't silently skip steps.

© heygen-com, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 159 other files (scripts, references, assets) in skills/embedded-captions of heygen-com/hyperframes.

  • SKILL.md
  • .gitignore
  • CATALOG.md
  • assets/fonts/char-widths.json
  • assets/strokefonts/HersheyScript1.svg
  • assets/strokefonts/HersheyScriptMed.svg
  • dna/README.md
  • dna/chrome.json
  • dna/cream.json
  • dna/documentary.json
  • dna/editorial.json
  • dna/glitch.json
  • dna/ink.json
  • dna/keynote.json
  • dna/loud.json
  • dna/neon.json
  • dna/velocity.json
  • … and 143 more

Open the folder on GitHubat commit 0c76e52

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in heygen-com/hyperframes, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Embedded Video Captions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Embedded Video Captions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Embedded Video Captions this skillheygen-com/hyperframes59k3 repos~8.6kAutomated safety check: PassApache-2.0
Noti Tiktok Vnnotivn/AIEV126—~5.4kAutomated safety check: PassMIT
Talking-Head Video Pipelinenaive-kun/naive-video-skill132—~3.3kAutomated safety check: PassMIT
Content To Videoarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0
KinocutKyaniteLabs/kinocut192—~5.7kAutomated safety check: PassApache-2.0
Yuv Viral Videohoodini/ai-agents-skills281—~7.5kAutomated safety check: NotesNone

Similar skills

  • Noti Tiktok Vn

    notivn/AIEV

    Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…

    126 GitHub stars~5.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Talking-Head Video Pipeline

    naive-kun/naive-video-skill

    Turns raw or rough-cut talking-head footage into a captioned, animated final video through a staged, resumable production pipeline.

    132 GitHub stars~3.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Kinocut

    KyaniteLabs/kinocut

    Use Kinocut for guarded video editing, source-backed planning, FFmpeg operations, media analysis, subtitles, audio workflows, Hyperframes or Revideo rendering, repurposing packages, and release…

    192 GitHub stars~5.7k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Yuv Viral Video

    hoodini/ai-agents-skills

    Edit any selfie or screen-share footage into a viral short-form video in YUV.AI's signature style — Apple-style liquid-glass cards (real CSS backdrop-filter), dark-mode polish, MrBeast-paced cuts…

    281 GitHub stars~7.5k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.

    126 GitHub stars~4.8k tokensUpdated today
    Media & CreativeAuto-check passed

More from heygen-com/hyperframes

All 32 skills in this repo
  • HyperFrames Animation

    heygen-com/hyperframes

    Collects motion rules, scene blueprints, transitions and runtime adapters for HyperFrames video compositions, with GSAP as the default animation runtime.

    59k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Weekly Changelog Video

    heygen-com/hyperframes

    Turns a weekly changelog markdown file into a branded HyperFrames video with voiceover, animated mock-UI scenes and captions, using fonts, background and scripts bundled in the skill.

    59k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Faceless Explainer Video

    heygen-com/hyperframes

    Turns an article, notes or a topic brief into an explainer video whose visuals are invented per scene, built frame by frame in HyperFrames with no footage.

    59k GitHub starsUsed in 3 repos~7.7k tokens
    Auto-check: notes
  • Figma to HyperFrames

    heygen-com/hyperframes

    Imports Figma assets, brand tokens, components and motion into a HyperFrames video composition, using the Figma REST API with a connector or native export for shaders.

    59k GitHub starsUsed in 3 repos~4.5k tokens
    Auto-check: notes
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Auto-check passed

Questions about Embedded Video Captions

What does Embedded Video Captions do?

Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker. The skill adds subtitles to an existing talking-head video and picks a look from a catalog of 35 styles. Standard, the default, builds a quiet rail of verbatim lower-third captions plus one embedded caption composited behind the subject at the peak.

When should I use Embedded Video Captions?

Embedded Video Captions fits situations like: adding plain verbatim subtitles to a talking-head video; making cinematic captions that sit behind the speaker; applying a named caption style to explainer or voiceover footage; requests for flashy, effects-heavy subtitles, including ones phrased in Chinese.

How do I install Embedded Video Captions in Claude Code?

Run `npx skills add heygen-com/hyperframes --skill embedded-captions -a claude-code`. Or copy the skill folder (skills/embedded-captions in heygen-com/hyperframes) into .claude/skills/embedded-captions in your project. Claude Code loads it when a task matches its description.

How do I install Embedded Video Captions in Codex?

Run `npx skills add heygen-com/hyperframes --skill embedded-captions -a codex`. Or copy the skill folder (skills/embedded-captions in heygen-com/hyperframes) into .agents/skills/embedded-captions in your project. Codex loads it when a task matches its description.

Can I use Embedded Video Captions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add heygen-com/hyperframes --skill embedded-captions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/embedded-captions, .gemini/skills/embedded-captions, .github/skills/embedded-captions and .opencode/skills/embedded-captions in your project.

What does Embedded Video Captions need to run?

Going by SKILL.md and its folder, Embedded Video Captions needs the command-line tools its instructions call (node, ffmpeg, npx, bash, npm and ffprobe). Our summary lists: Bash and FFmpeg or ffprobe on the PATH; npm, to install Sharp, Puppeteer and GSAP in the caption project; The HyperFrames CLI (bundled with plugin installs).

Does Embedded Video Captions access the network?

SKILL.md contains no URLs. Its commands use npx, npm and uvx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Embedded Video Captions safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Embedded Video Captions use?

Embedded Video Captions is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Embedded Video Captions use?

About 8.6k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 44k tokens, read only when the agent opens those files.

What are the alternatives to Embedded Video Captions?

Skills that share tags, products or a category with Embedded Video Captions: Noti Tiktok Vn (notivn/AIEV, 126 stars), Talking-Head Video Pipeline (naive-kun/naive-video-skill, 132 stars), Content To Video (architectds/modeldock, 117 stars) and Kinocut (KyaniteLabs/kinocut, 192 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Embedded Video Captions?

heygen-com (a GitHub organization) maintains it in heygen-com/hyperframes, which has 58,669 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 8, 2026.

Source: heygen-com/hyperframes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.