Agent skill

Scenario Caption Studio

by scenario-labs in scenario-labs/skills

A skill your agent uses when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube…

MITAuto-check passedMedia & Creative

Install Scenario Caption Studio

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-caption-studio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-caption-studio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-caption-studio .claude/skills/scenario-caption-studio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-caption-studio
GitHub stars
931
Token cost
~3.2k tokens
SKILL.md length
1,643 words
Files
1
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube…

  • Works in 7 steps: asset_get the assembled master: confirm… → model_schema_get on… → Price it: model_run with dry_run: true… → …
  • A video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok
  • SKILL.md covers Overview, Ask once, then caption, The style ladder and Getting the words right, plus 2 more sections
  • Calls npx and ffmpeg

What it does

Scenario Caption Studio is an agent skill from scenario-labs/skills. Use when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript. Keywords: captions, subtitles, SRT, transcribe, karaoke, word-by-word, burn in, closed captions, translate video, caption style.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. It works with Model Context Protocol, TikTok and YouTube. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • A video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok
  • Ad captions for sound-off feeds
  • YouTube subtitles
  • Transcription of a clips audio

Example prompts

  • “/scenario-caption-studio”

Requirements

  • Node.js

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. asset_get the assembled master: confirm duration and that it is the finished cut, since the price tracks the footage (video carries…
  2. model_schema_get on model_scenario-caption-studio.
  3. Price it: model_run with dry_run: true and parameters={"video": "", "stylePreset": "tiktok-bouncy", "maxSegmentWords": 4, "textPosition"…
  4. Run it with wait: false, then jobs_wait with the job_id; on timeout re-call with the returned pending_job_ids, never job_get in a loop.
  5. Review before deriving, since a successful job proves nothing about what rendered: download the SRT sidecar with asset_download (no…
  6. Spanish variant: add targetLanguage: "es" to the same parameters, transcriptionPrompt included, and dry_run again (it moves the price)…
  7. File the master, variants, and SRT assets in a collection (scenario skill) so the delivery set stays findable.

What it can do on your machine

Read from SKILL.md and the folder at commit 91caa01. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Caption Studio loads about 3.2k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 1,643 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit 91caa01, republished under its MIT licence (© scenario-labs). 1,643 words, ~3,167 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-caption-studio/SKILL.md (or your agent's skills folder).
name
scenario-caption-studio
description
Use when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript. Keywords: captions, subtitles, SRT, transcribe, karaoke, word-by-word, burn in, closed captions, translate video, caption style.
license
MIT

Scenario Caption Studio

Overview

Caption Studio is one tool model, model_scenario-caption-studio: a video in, its speech transcribed (Whisper) or an existing SRT applied, styled captions out, burned into the picture or delivered as a soft track and an .srt sidecar. It translates into 18 languages and styles captions three ways. Running that one member is this skill's whole purpose, so the id is named rather than discovered and model_schema_get starts the flow directly; availability differs per team, so a member the team lacks is a gap to flag, not a cue to substitute. Connection and the core loop: see the scenario skill.

Captioning is the last pass on a finished cut: assemble first (scenario-video-assembly), then caption the master once. Captions carry the transcript only; text that must appear letter-perfect without being spoken (CTAs, prices, legal supers) is scenario-text-overlay territory. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Ask once, then caption

The destination decides nearly every parameter, so collect one round of answers before touching the schema: where the video ships (a sound-off mobile feed, a paid placement, a seated long-form viewer), whether the spoken language stays or translates (targetLanguage, auto keeps it), brand colors if any, and which deliverable the platform wants. The deliverable is three switches: burned-in pixels are outputSubtitles: "video_image" (the default), the toggleable track is "video_data", the sidecar file is outputSrt: true, and an SRT-only pass is that plus outputVideo: false, the first pass when the words must be letter-perfect: it priced the same as a burn-in at authoring time, so it buys certainty rather than savings, proving the words before any pixels are paid for. Then map the answers:

DestinationStyleSegmentationPositionOutput
Social mobile short (9:16 Shorts, Reels)tiktok-bouncy or word-pop; karaoke-fill when music drivesmaxSegmentWords 3 to 5; 1 with a karaoke presetmiddle: platform UI and native auto-captions own the bottomBurn in
Ad short, performance cutmodern-chip or minimal-underline, accents set to brand color3 to 7 words per cuebottom, or top when an end card or overlay sits belowBurn in for sound-off feeds; add outputSrt for the platform's caption upload
YouTube long-form, tutorial, interviewDefault look or cinematic-fade; restraint reads as professionalismmaxLines 2, maxSegmentChars 84 (two 42-character lines, the broadcast convention)bottomoutputSrt for the platform's closed captions; burn in only for re-embeds
Cinematic piece, trailer, festival cutcinematic-fadeSentence-length cues, maxSegmentDuration about 6bottomBurn in

Rows are authoring-time starting points to confirm with the user, not platform contracts; unattended, the task's own instructions answer the interview and the matching row's defaults fill what they leave unsaid. The per-placement safe zones behind the position column live in scenario-formats. middle renders at frame center, which on a centered talking head is the mouth: caption a selfie cut only once its face sits in the upper third (reframe in assembly), because the schema offers no lower-third position. An uploaded track beats burn-in for long-form because viewers toggle and restyle it, assistive tech reads it, and platforms index it for search; burn-in wins wherever the style is the point or a track cannot travel with the file.

The style ladder

Three tiers: stylePreset picks a ready-made look (7 presets, empty for the default); stylePrompt describes a look in plain words and builds a matching style (it carries cost_impact); themeTsx supplies a full custom theme that replaces the preset, with stylePrompt then refining that theme. There is no font parameter: type rides inside the tiers, and it is the strongest signal a style sends, so when the type itself must carry the mood or the brand, put the intent into stylePrompt in plain words (the weight, the letterform class, the feeling: "heavy condensed sans, high-energy", "light geometric sans, quiet and premium"); an exact brand face is themeTsx territory, and scenario-text-overlay chooses faces by meaning for the text cards around the captions. Presets also restyle the words themselves: an authoring-time run of tiktok-bouncy uppercased every caption, and the other presets are unverified for casing, so when exact casing matters (a product name, "LoRA") steer with stylePrompt or themeTsx and verify a frame before delivering. fontColor sets the body text (contrast beats aesthetics: white body text survives every backdrop the presets put behind it), and accentColorStart/accentColorEnd drive the highlight animation (karaoke fills, pops): spend the accent on one thing, usually the brand color, with equal values for a solid and different values for a gradient. Auto sizing (fontSizePx empty) rendered words about 30 pixels tall on a 1080x1920 frame at authoring time, unreadable on a phone feed: for vertical social set fontSizePx explicitly (60 to 90 on a 1920-tall frame is the authoring-time starting point) and judge a frame at phone scale. outputTsx: true returns the theme a run used, so a look that landed can be replayed exactly on the next video.

Getting the words right

  • transcriptionPrompt is a spelling hint, not a style field: list the names, brands, and jargon the audio contains. The hint raises the odds without guaranteeing them (an authoring-time run misspelled a hinted name twice), so check the transcript for every required name before trusting a burn-in, and on a miss retry with large-v3 or a sharper hint.
  • modelSize trades accuracy for speed and cost (cost_impact): the medium default is fine for clean voiceover; step up to large-v3 for noisy audio, accents, or dense terminology; .en variants are English-only.
  • The model transcribes whatever audio the master carries, so balance the music bed against dialogue before captioning (scenario-video-assembly), never after.
  • Corrections and restyles ride the subtitles input, whose contract is inline content, not a reference (authoring-time): pass the SRT text itself, base64-encoded, as the value. An asset_... id is not dereferenced there; the id string is base64-decoded as if it were content, and the run still reports success, bills, and renders zero captions (segment_count: 0 in the job record is the tell). Reuse is therefore: outputSrt: true returns the transcript as an asset, asset_download it, correct spellings locally if needed, and feed the edited text back base64-encoded, which also covers upload_asset having no text kind (authoring-time fact). The SRT inherits the run's segmentation (a karaoke run returns word-per-cue), so re-chunk it locally before an ads upload or a calmer restyle. Never deliver a subtitles run on job status alone: sweep the output as the worked example reviews it, and when captions are missing, asset_get the subtitles asset the job consumed to see what it received. The schema does not say whether segmentation caps re-chunk a supplied SRT, so set segmentation when transcribing and omit the caps alongside subtitles.
Show full SKILL.md (526 more words)Show less

Worked example: a vertical social short, then a Spanish variant

  1. asset_get the assembled master: confirm duration and that it is the finished cut, since the price tracks the footage (video carries cost_impact; trim first, scenario-video-editing).
  2. model_schema_get on model_scenario-caption-studio.
  3. Price it: model_run with dry_run: true and parameters={"video": "<asset_id>", "stylePreset": "tiktok-bouncy", "maxSegmentWords": 4, "textPosition": "middle", "transcriptionPrompt": "Scenario, LoRA, Flux", "outputSrt": true}. Re-estimate after changing targetLanguage, modelSize, stylePrompt, or outputSubtitles: all carry cost_impact.
  4. Run it with wait: false, then jobs_wait with the job_id; on timeout re-call with the returned pending_job_ids, never job_get in a loop.
  5. Review before deriving, since a successful job proves nothing about what rendered: download the SRT sidecar with asset_download (no format: it converts images only; a run that returns both video and SRT lists the two asset ids in no fixed order, and asset_get tells them apart by mimeType) and check its text for every name the transcriptionPrompt carries; then download the video and sweep it into contact sheets (ffmpeg -vf "fps=2"), reading the burned captions for spelling, casing, and placement, because a defect that appears mid-cue survives a spot check. Pop presets reveal words within a cue, so also sample the last frames of each cue: a flash under 100 ms survives an fps=2 sweep.
  6. Spanish variant: add targetLanguage: "es" to the same parameters, transcriptionPrompt included, and dry_run again (it moves the price) before running. Translation happens inside the run, so one master yields a variant per market. Latin-script segmentation caps do not transfer to Chinese, Japanese, or Korean (streaming style guides run them at a third of the characters per line), so revisit maxSegmentChars per target language.
  7. File the master, variants, and SRT assets in a collection (scenario skill) so the delivery set stays findable.

Common mistakes

  • Captioning each clip before assembly: caption the finished master once, or cues drift across cuts and every edit orphans its captions.
  • Expecting the preset look on the soft track: video_data renders as plain text the viewer toggles; styling survives only when burned in.
  • Styling legal lines or CTAs as captions: disclosures carry locked wording, size, and dwell time that caption logic would re-chunk and retime, and a toggleable track fails "visible without viewer action" rules outright; exact unspoken text is a scenario-text-overlay card composited in assembly.
  • Setting maxSegmentWords: 1 without a karaoke or pop preset: one-word cues flash as a slideshow unless the style animates them. With a pop preset (the ones that reveal words inside a cue: tiktok-bouncy, word-pop, and the karaoke pair at authoring time), never end a cue on a one-letter word: pop-in timing follows character count, so a trailing "I" or "a" showed for about 70 ms at authoring time. Re-chunking a supplied SRT means owning its timings too: a supplied SRT carries cue times only, so a pop preset spreads its words evenly across each cue and the highlight drifts from the voice wherever a cue outlasts the speech; keep each cue's end on the spoken span, and extend only the last cue to the clip's end.
  • Treating a jobs_wait timeout as failure: re-call with pending_job_ids; video tool jobs outlast the wait window routinely.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-caption-studio of scenario-labs/skills.

Open the folder on GitHubat commit 91caa01

Compare with similar skills

Scenario Caption Studio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Caption Studio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Caption Studio this skillscenario-labs/skills931—~3.2kAutomated safety check: PassMIT
Watching Videosoxbshw/watch-skill469—~599Automated safety check: NotesMIT
WatchTheCraigHewitt/skills157—~1.6kAutomated safety check: NotesMIT
Video Transcribewendy7756/AI-Video-Transcriber3.3k—~937Automated safety check: NotesApache-2.0
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
ShowtimeFavioVazquez/showtime206—~3kAutomated safety check: PassMIT

Similar skills

  • Watching Videos

    oxbshw/watch-skill

    The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…

    469 GitHub stars~599 tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes
  • Watch

    TheCraigHewitt/skills

    When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports.

    157 GitHub stars~1.6k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: notes
  • Video Transcribe

    wendy7756/AI-Video-Transcriber

    Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

    3.3k GitHub stars~937 tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    206 GitHub stars~3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Hotclip

    xixihhhh/hotclip

    Turn long videos & livestream VODs into viral vertical shorts, 100% locally — on-device transcription, LLM highlight detection, 9:16 reframe with karaoke captions, and a per-clip render-QA report.

    304 GitHub stars~1.1k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from scenario-labs/skills

All 143 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    931 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    931 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Multiplayer

    scenario-labs/skills

    A skill your agent uses when a Godot 4 game goes online or co-op: host and join with ENet, WebSocket for a web build, RPCs (@rpc, rpcid, anypeer), MultiplayerSpawner and MultiplayerSynchronizer…

    931 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Caption Studio

What does Scenario Caption Studio do?

A skill your agent uses when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube…. Scenario Caption Studio is an agent skill from scenario-labs/skills. Use when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript.

When should I use Scenario Caption Studio?

Scenario Caption Studio fits situations like: A video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok; ad captions for sound-off feeds; youTube subtitles; transcription of a clips audio.

How do I install Scenario Caption Studio in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-caption-studio -a claude-code`. Or copy the skill folder (skills/scenario-caption-studio in scenario-labs/skills) into .claude/skills/scenario-caption-studio in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Caption Studio in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-caption-studio -a codex`. Or copy the skill folder (skills/scenario-caption-studio in scenario-labs/skills) into .agents/skills/scenario-caption-studio in your project. Codex loads it when a task matches its description.

Can I use Scenario Caption Studio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-caption-studio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-caption-studio, .gemini/skills/scenario-caption-studio, .github/skills/scenario-caption-studio and .opencode/skills/scenario-caption-studio in your project.

What does Scenario Caption Studio need to run?

Going by SKILL.md and its folder, Scenario Caption Studio needs the command-line tools its instructions call (npx and ffmpeg). Our summary lists: Node.js.

Does Scenario Caption Studio access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Caption Studio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Caption Studio use?

Scenario Caption Studio is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Caption Studio use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Caption Studio?

Skills that share tags, products or a category with Scenario Caption Studio: Watching Videos (oxbshw/watch-skill, 469 stars), Watch (TheCraigHewitt/skills, 157 stars), Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars) and Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Caption Studio?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 931 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.