Agent skill

Caption Burn

by gooseworks-ai in gooseworks-ai/goose-skills

Burned-in captions for a finished vertical video, three kinds.

MITAuto-check passedMedia & Creative

Install Caption Burn

skills CLI
$ npx skills add gooseworks-ai/goose-skills --skill caption-burn -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gooseworks-ai/goose-skills caption-burn --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gooseworks-ai/goose-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ads/capabilities/caption-burn .claude/skills/caption-burn && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
caption-burn
GitHub stars
1.2k
Token cost
~1.2k tokens
SKILL.md length
582 words
Files
9 (incl. scripts)
Skills in repo
273
Repo updated
First seen
Licence
MIT

At a glance

Burned-in captions for a finished vertical video, three kinds.

  • Works in 5 steps: Transcribe the FINISHED audio. Joining… → The last caption holds to the last… → A word Whisper writes differently ("200"… → …
  • Tasks that involve Transcription
  • SKILL.md covers Run, Placement and style, Captions for a format with no… and Rules, plus 1 more section
  • Runs Python scripts from its folder; calls python

What it does

Caption Burn is an agent skill from gooseworks-ai/goose-skills. Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette…

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `scripts/_fonts.py`, `scripts/captions.py` and `scripts/footprint.py`).

It sits in Media & Creative, covering Transcription, Speech recognition and synthesis and Video production. It works with FFmpeg. The repository describes itself as: Library of Growth & GTM skills + data APIs for Claude Code, Codex, Cursor to run ads, social, content, lead gen, seo and data scraping. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Video production

Example prompts

  • “/caption-burn”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Transcribe the FINISHED audio. Joining takes and aligning lines shifts timing;
  2. The last caption holds to the last frame. It is the call to action.
  3. A word Whisper writes differently ("200" for "two hundred") is interpolated between
  4. Without --words timing is estimated from syllables. Use that to judge placement,
  5. Fonts: a bold sans is found on macOS, Linux or Windows; if none is present, Roboto

What it can do on your machine

Read from SKILL.md and the folder at commit c650c6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Caption Burn loads about 1.2k tokens when it runs. Until then it costs about 169 tokens; SKILL.md has 582 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~169
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gooseworks-ai/goose-skills at commit c650c6d, republished under its MIT licence (© gooseworks-ai). 582 words, ~1,197 tokens.

Download SKILL.mdSave it as .claude/skills/caption-burn/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
caption-burn
description
Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.
status
superseded
superseded_by
captions-layer@1.1.0

Superseded: the video kit now does this with the captions-layer part, version 1.1.0, in the parts folder of this repository. It burns the same plate captions from the speech timings, holds the last caption to the end, transcribes only when there are no timings, and returns a WebVTT of what it drew. This atom stays, unchanged in behaviour, for skills outside the kit until they move; its scripts still run.

caption-burn

Captions timed to what is actually said in the finished video, not to the script's estimate.

Run

bash
python transcribe.py --media reel.mp4 --out reel.words.json          # paid, cents
python captions.py --video reel.mp4 --beats cutlist.aligned.json \
    --words reel.words.json --out final.mp4 [--style plate|outline] [--highlight Brand]

--beats supplies the lines (vo), their timing and, for split layouts, seam, size and each beat's state. Any file with beats: [{start, end, vo}] works.

Placement and style

  • --anchor seam (the default when the beats have a seam): on split beats the plate is pinned to the seam, positioned by the plate, not the text: 25% of the plate above the line, 75% below. Full-frame beats use --full-y.

  • --anchor fixed --y 0.62: every caption's plate centred at that fraction of the height.

  • plate (default): white bold on a dark grey rounded plate, 1–2 words, cap ~0.019 H.

  • outline: white bold with a dark outline, no plate, 1–3 words, cap ~0.034 H.

  • --highlight WORD colours that word yellow (the CTA keyword). Repeatable.

  • serif-word: ONE word at a time, heavy serif (Georgia Bold), white with a black outline, on a fixed baseline at 0.77 H (the screen-insert look). The highlight word is quoted.

  • --card "LINE ONE|LINE TWO" --card-until 4.7: a white rounded hook card with two lines of heavy red capitals near the top, for the opening seconds. ~14 characters a line.

Captions for a format with no voice (plates.py)

bash
python plates.py --video walk.mp4 --beats cutlist.json --out captioned.mp4 [--logo logo.png]

Each beat's caption (a string or list of lines) shows for the whole beat on ONE black block (square rectangles unioned, then rounded as a single silhouette: rounding each line leaves seams), lines left-aligned, the block centred on its widest line. It goes in the emptiest band of that beat's frame unless the beat pins cap_y. logo: true on a beat hangs the logo tile under the block. Write lines a person would type: the same short "fragment. fragment." shape three times reads as AI-written. No emoji twice.

Show full SKILL.md (218 more words)Show less

Rules

  1. Transcribe the FINISHED audio. Joining takes and aligning lines shifts timing; only the final video's audio gives correct cues.
  2. The last caption holds to the last frame. It is the call to action.
  3. A word Whisper writes differently ("200" for "two hundred") is interpolated between its neighbours rather than dropped or stretched over the whole line.
  4. Without --words timing is estimated from syllables. Use that to judge placement, never to ship.
  5. Fonts: a bold sans is found on macOS, Linux or Windows; if none is present, Roboto Bold is fetched once into ~/.cache/gooseworks/fonts. --font or GW_CAPTION_FONT overrides.

Footprint before composition

The bundled footprint helper uses this renderer's cue grouping, selected font, stroke, plate padding and anchor. Export it from the approved beat list and pass it to the footage-cutlist preview. The JSON includes every group and the union of its rendered bounds per beat. Use the same style, anchor, font and fixed-position settings for the final burn; regenerate after alignment or copy changes. Caption coverage must leave claim qualifications readable. Inspect final captioned frames as well as this planned coverage.

Pass the same highlight terms to footprint and burn, including serif-word quotes. Preview rejects a changed cut list until its footprint is rebuilt. An explicitly supplied missing font fails in both commands.

© gooseworks-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/ads/capabilities/caption-burn of gooseworks-ai/goose-skills.

  • SKILL.md
  • scripts/_fonts.py
  • scripts/captions.py
  • scripts/footprint.py
  • scripts/media_proxy.py
  • scripts/plates.py
  • scripts/transcribe.py
  • skill.meta.json
  • tests/smoke-test.md

Open the folder on GitHubat commit c650c6d

Compare with similar skills

Caption Burn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Caption Burn compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Caption Burn this skillgooseworks-ai/goose-skills1.2k—~1.2kAutomated safety check: PassMIT
Karaoke CaptionsAI-Builder-Club/skills1.3k—~850Automated safety check: PassNone
AutoshortsUpload-Post/skill-autoshorts151—~5.3kAutomated safety check: NotesMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Video Clip Extractorlinzzzzzz/openclip569—~2.8kAutomated safety check: WarnMIT

Similar skills

  • Karaoke Captions

    AI-Builder-Club/skills

    Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.

    1.3k GitHub stars~850 tokensUpdated 25 days ago
    Media & CreativeAuto-check passed
  • Autoshorts

    Upload-Post/skill-autoshorts

    Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…

    151 GitHub stars~5.3k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Bggg Tiktok Readvideo

    binggandata/bggg-skills

    把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。

    605 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from gooseworks-ai/goose-skills

All 273 skills in this repo
  • Reddit Post Finder

    gooseworks-ai/goose-skills

    Scrape and search Reddit posts using Apify. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Image Fal

    gooseworks-ai/goose-skills

    Generate or edit an image via any FAL image model (nano-banana edit, gpt-image, flux, ...), ROUTED THROUGH THE fal-proxy so it bills the Ads agent.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Render Hook Replacement

    gooseworks-ai/goose-skills

    Replace an existing video's opening with a supplied clip or free kinetic text hook while retaining and verifying every original body frame, audio, captions and ending.

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Blog Feed Monitor

    gooseworks-ai/goose-skills

    Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites.

    1.2k GitHub starsUsed in 1 repo~578 tokens
    Auto-check passed
  • Competitor Post Engagers

    gooseworks-ai/goose-skills

    Find leads by scraping engagers from a competitor's top LinkedIn posts.

    1.2k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Render Chatgpt Chat

    gooseworks-ai/goose-skills

    Assemble a ChatGPT chat-reveal video ad from a thread + timeline JSON — one continuous Playwright recording of a ChatGPT mobile chat (user types with the iOS keyboard up → taps send → keyboard…

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Caption Burn

What does Caption Burn do?

Burned-in captions for a finished vertical video, three kinds. Caption Burn is an agent skill from gooseworks-ai/goose-skills. Burned-in captions for a finished vertical video, three kinds.

When should I use Caption Burn?

Caption Burn fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis; tasks that involve Video production.

How do I install Caption Burn in Claude Code?

Run `npx skills add gooseworks-ai/goose-skills --skill caption-burn -a claude-code`. Or copy the skill folder (skills/ads/capabilities/caption-burn in gooseworks-ai/goose-skills) into .claude/skills/caption-burn in your project. Claude Code loads it when a task matches its description.

How do I install Caption Burn in Codex?

Run `npx skills add gooseworks-ai/goose-skills --skill caption-burn -a codex`. Or copy the skill folder (skills/ads/capabilities/caption-burn in gooseworks-ai/goose-skills) into .agents/skills/caption-burn in your project. Codex loads it when a task matches its description.

Can I use Caption Burn in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gooseworks-ai/goose-skills --skill caption-burn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/caption-burn, .gemini/skills/caption-burn, .github/skills/caption-burn and .opencode/skills/caption-burn in your project.

What does Caption Burn need to run?

Going by SKILL.md and its folder, Caption Burn needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Caption Burn access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Caption Burn safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Caption Burn use?

Caption Burn is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Caption Burn use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Caption Burn?

Skills that share tags, products or a category with Caption Burn: Karaoke Captions (AI-Builder-Club/skills, 1.3k stars), Autoshorts (Upload-Post/skill-autoshorts, 151 stars), Watch (mathiaschu/watch, 142 stars) and Watch Video (coreyhaines31/makerskills, 851 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Caption Burn?

gooseworks-ai (a GitHub organization) maintains it in gooseworks-ai/goose-skills, which has 1,240 GitHub stars. The repository holds 273 skills in this directory. The repository was last updated on October 8, 2026.

Source: gooseworks-ai/goose-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.