---
name: scenario-caption-studio
description: "Use when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript. Keywords: captions, subtitles, SRT, transcribe, karaoke, word-by-word, burn in, closed captions, translate video, caption style."
license: MIT
---

# Scenario Caption Studio

## Overview

Caption Studio is one tool model, `model_scenario-caption-studio`: a video in, its speech transcribed (Whisper) or an existing SRT applied, styled captions out, burned into the picture or delivered as a soft track and an `.srt` sidecar. It translates into 18 languages and styles captions three ways. Running that one member is this skill's whole purpose, so the id is named rather than discovered and `model_schema_get` starts the flow directly; availability differs per team, so a member the team lacks is a gap to flag, not a cue to substitute. Connection and the core loop: see the `scenario` skill.

Captioning is the last pass on a finished cut: assemble first (`scenario-video-assembly`), then caption the master once. Captions carry the transcript only; text that must appear letter-perfect without being spoken (CTAs, prices, legal supers) is `scenario-text-overlay` territory. If a sibling skill named here is missing from your available skills, ask the user to install it (`npx skills add scenario-labs/skills --skill <name>`); unattended, proceed from tool schemas and flag the gap.

## Ask once, then caption

The destination decides nearly every parameter, so collect one round of answers before touching the schema: where the video ships (a sound-off mobile feed, a paid placement, a seated long-form viewer), whether the spoken language stays or translates (`targetLanguage`, `auto` keeps it), brand colors if any, and which deliverable the platform wants. The deliverable is three switches: burned-in pixels are `outputSubtitles: "video_image"` (the default), the toggleable track is `"video_data"`, the sidecar file is `outputSrt: true`, and an SRT-only pass is that plus `outputVideo: false`, the first pass when the words must be letter-perfect: it priced the same as a burn-in at authoring time, so it buys certainty rather than savings, proving the words before any pixels are paid for. Then map the answers:

| Destination                              | Style                                                                | Segmentation                                                                          | Position                                                      | Output                                                                         |
| ---------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| Social mobile short (9:16 Shorts, Reels) | `tiktok-bouncy` or `word-pop`; `karaoke-fill` when music drives      | `maxSegmentWords` 3 to 5; 1 with a karaoke preset                                     | `middle`: platform UI and native auto-captions own the bottom | Burn in                                                                        |
| Ad short, performance cut                | `modern-chip` or `minimal-underline`, accents set to brand color     | 3 to 7 words per cue                                                                  | `bottom`, or `top` when an end card or overlay sits below     | Burn in for sound-off feeds; add `outputSrt` for the platform's caption upload |
| YouTube long-form, tutorial, interview   | Default look or `cinematic-fade`; restraint reads as professionalism | `maxLines` 2, `maxSegmentChars` 84 (two 42-character lines, the broadcast convention) | `bottom`                                                      | `outputSrt` for the platform's closed captions; burn in only for re-embeds     |
| Cinematic piece, trailer, festival cut   | `cinematic-fade`                                                     | Sentence-length cues, `maxSegmentDuration` about 6                                    | `bottom`                                                      | Burn in                                                                        |

Rows are authoring-time starting points to confirm with the user, not platform contracts; unattended, the task's own instructions answer the interview and the matching row's defaults fill what they leave unsaid. The per-placement safe zones behind the position column live in `scenario-formats`. `middle` renders at frame center, which on a centered talking head is the mouth: caption a selfie cut only once its face sits in the upper third (reframe in assembly), because the schema offers no lower-third position. An uploaded track beats burn-in for long-form because viewers toggle and restyle it, assistive tech reads it, and platforms index it for search; burn-in wins wherever the style is the point or a track cannot travel with the file.

## The style ladder

Three tiers: `stylePreset` picks a ready-made look (7 presets, empty for the default); `stylePrompt` describes a look in plain words and builds a matching style (it carries `cost_impact`); `themeTsx` supplies a full custom theme that replaces the preset, with `stylePrompt` then refining that theme. There is no font parameter: type rides inside the tiers, and it is the strongest signal a style sends, so when the type itself must carry the mood or the brand, put the intent into `stylePrompt` in plain words (the weight, the letterform class, the feeling: "heavy condensed sans, high-energy", "light geometric sans, quiet and premium"); an exact brand face is `themeTsx` territory, and `scenario-text-overlay` chooses faces by meaning for the text cards around the captions. Presets also restyle the words themselves: an authoring-time run of `tiktok-bouncy` uppercased every caption, and the other presets are unverified for casing, so when exact casing matters (a product name, "LoRA") steer with `stylePrompt` or `themeTsx` and verify a frame before delivering. `fontColor` sets the body text (contrast beats aesthetics: white body text survives every backdrop the presets put behind it), and `accentColorStart`/`accentColorEnd` drive the highlight animation (karaoke fills, pops): spend the accent on one thing, usually the brand color, with equal values for a solid and different values for a gradient. Auto sizing (`fontSizePx` empty) rendered words about 30 pixels tall on a 1080x1920 frame at authoring time, unreadable on a phone feed: for vertical social set `fontSizePx` explicitly (60 to 90 on a 1920-tall frame is the authoring-time starting point) and judge a frame at phone scale. `outputTsx: true` returns the theme a run used, so a look that landed can be replayed exactly on the next video.

## Getting the words right

- When narration was authored, retain that script as the wording reference. Compare the transcript against it for missing, repeated or altered words before burning; a spelling hint is not that comparison. Derive cue times from the actual recording, not word count or the planned shot duration. Correct the SRT through the supported input below; do not invent a script-alignment parameter.
- `transcriptionPrompt` is a spelling hint, not a style field: list the names, brands, and jargon the audio contains. The hint raises the odds without guaranteeing them (an authoring-time run misspelled a hinted name twice), so check the transcript for every required name before trusting a burn-in, and on a miss retry with `large-v3` or a sharper hint.
- `modelSize` trades accuracy for speed and cost (`cost_impact`): the `medium` default is fine for clean voiceover; step up to `large-v3` for noisy audio, accents, or dense terminology; `.en` variants are English-only.
- The model transcribes whatever audio the master carries, so balance the music bed against dialogue before captioning (`scenario-video-assembly`), never after.
- Corrections and restyles ride the `subtitles` input, whose contract is inline content, not a reference (authoring-time): pass the SRT text itself, base64-encoded, as the value. An `asset_...` id is not dereferenced there; the id string is base64-decoded as if it were content, and the run still reports success, bills, and renders zero captions (`segment_count: 0` in the job record is the tell). Reuse is therefore: `outputSrt: true` returns the transcript as an asset, `asset_download` it, correct spellings locally if needed, and feed the edited text back base64-encoded, which also covers `upload_asset` having no text kind (authoring-time fact). The SRT inherits the run's segmentation (a karaoke run returns word-per-cue), so re-chunk it locally before an ads upload or a calmer restyle. Never deliver a `subtitles` run on job status alone: sweep the output as the worked example reviews it, and when captions are missing, `asset_get` the subtitles asset the job consumed to see what it received. The schema does not say whether segmentation caps re-chunk a supplied SRT, so set segmentation when transcribing and omit the caps alongside `subtitles`.

## Worked example: a vertical social short, then a Spanish variant

1. `asset_get` the assembled master: confirm duration and that it is the finished cut, since the price tracks the footage (`video` carries `cost_impact`; trim first, `scenario-video-editing`).
2. `model_schema_get` on `model_scenario-caption-studio`.
3. Price it: `model_run` with `dry_run: true` and `parameters={"video": "<asset_id>", "stylePreset": "tiktok-bouncy", "maxSegmentWords": 4, "textPosition": "middle", "transcriptionPrompt": "Scenario, LoRA, Flux", "outputSrt": true}`. Re-estimate after changing `targetLanguage`, `modelSize`, `stylePrompt`, or `outputSubtitles`: all carry `cost_impact`.
4. Run it with `wait: false`, then `jobs_wait` with the `job_id`; on timeout re-call with the returned `pending_job_ids`, never `job_get` in a loop.
5. Review before deriving, since a successful job proves nothing about what rendered: download the SRT sidecar with `asset_download` (no `format`: it converts images only; a run that returns both video and SRT lists the two asset ids in no fixed order, and `asset_get` tells them apart by `mimeType`) and check its text against the authored script and required names; then download the video and sweep it into contact sheets (`ffmpeg -vf "fps=2"`), reading the burned captions for spelling, casing, and placement. Inspect at the intended display size, including backing panels: keep faces, hands, demonstrated actions and essential UI clear, and keep placement stable within a cue. Listen against cue starts and ends; a correct transcript does not prove synchronization. Sample frames immediately before and after cue boundaries as well as within them: reveals and fades can exceed supplied SRT times, and a flash under 100 ms survives an `fps=2` sweep. If strict boundaries are required, verify the actual theme or report the limitation; a correct sidecar alone is not a pass. Check the final spoken word survives, comparing audio and video streams separately; a shorter audio stream containing the complete narration may simply omit trailing silence.
6. Spanish variant: add `targetLanguage: "es"` to the same parameters, `transcriptionPrompt` included, and `dry_run` again (it moves the price) before running. Translation happens inside the run, so one master yields a variant per market. Latin-script segmentation caps do not transfer to Chinese, Japanese, or Korean (streaming style guides run them at a third of the characters per line), so revisit `maxSegmentChars` per target language.
7. File the master, variants, and SRT assets in a collection (`scenario` skill) so the delivery set stays findable.

## Common mistakes

- Captioning each clip before assembly: caption the finished master once, or cues drift across cuts and every edit orphans its captions.
- Expecting the preset look on the soft track: `video_data` renders as plain text the viewer toggles; styling survives only when burned in.
- Styling legal lines or CTAs as captions: disclosures carry locked wording, size, and dwell time that caption logic would re-chunk and retime, and a toggleable track fails "visible without viewer action" rules outright; exact unspoken text is a `scenario-text-overlay` card composited in assembly.
- Setting `maxSegmentWords: 1` without a karaoke or pop preset: one-word cues flash as a slideshow unless the style animates them. With a pop preset (the ones that reveal words inside a cue: `tiktok-bouncy`, `word-pop`, and the karaoke pair at authoring time), never end a cue on a one-letter word: pop-in timing follows character count, so a trailing "I" or "a" showed for about 70 ms at authoring time. Re-chunking a supplied SRT means owning its timings too: a supplied SRT carries cue times only, so a pop preset spreads its words evenly across each cue and the highlight drifts from the voice wherever a cue outlasts the speech; keep each cue's end on the spoken span, and extend only the last cue to the clip's end.
- Treating a `jobs_wait` timeout as failure: re-call with `pending_job_ids`; video tool jobs outlast the wait window routinely.
