Agent skill

Scenario Video Assembly

by scenario-labs in scenario-labs/skills

A skill your agent uses when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or…

MITAuto-check passedMedia & Creative

Install Scenario Video Assembly

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-video-assembly -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-video-assembly --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-video-assembly .claude/skills/scenario-video-assembly && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-video-assembly
GitHub stars
946
Token cost
~2.3k tokens
SKILL.md length
1,329 words
Files
1
Skills in repo
146
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or…

  • Works in 5 steps: For every source, asset_get and read… → Add them up to get each clip's… → model_run it with layers holding the… → …
  • Generated clips must become a finished video on Scenario via MCP: cutting a shot list together
  • SKILL.md covers Overview, Quick reference: pick the…, The Video Studio contract and Worked example: three clips, a…, plus 1 more section
  • Calls npx

What it does

Scenario Video Assembly is an agent skill from scenario-labs/skills. Use when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or title card, adding a music bed or voiceover, burning in or exporting captions and subtitles, trimming and splitting, or producing vertical and square variants of one master. Keywords: edit, assemble, stitch, concat, compositor, timeline, layers, overlay, captions, SRT, ad, UGC, social cut.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, Influencer and creator marketing and Text to speech and voice. It works with Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Generated clips must become a finished video on Scenario via MCP: cutting a shot list together
  • Laying a timeline
  • Concatenating with transitions
  • Overlaying a logo

Example prompts

  • “/scenario-video-assembly”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. For every source, asset_get and read properties: the real duration (audio carries it too), frameRate on video, and width/height on video…
  2. Add them up to get each clip's startTime, then model_schema_get on model_scenario-compose-video.
  3. model_run it with layers holding the three clips at their computed startTimes (zIndex: 0), a logo PNG pinned by x/y/anchor at zIndex: 1…
  4. jobs_wait, then caption the finished cut rather than each clip: model_scenario-caption-studio (an authoring-time pick, since Scenario…
  5. asset_display to review, then asset_download with format left unset (it is an image conversion target) for the file.

What it can do on your machine

Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Video Assembly loads about 2.3k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 1,329 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 1,329 words, ~2,344 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-video-assembly/SKILL.md (or your agent's skills folder).
name
scenario-video-assembly
description
Use when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or title card, adding a music bed or voiceover, burning in or exporting captions and subtitles, trimming and splitting, or producing vertical and square variants of one master. Keywords: edit, assemble, stitch, concat, compositor, timeline, layers, overlay, captions, SRT, ad, UGC, social cut.
license
MIT

Scenario Video Assembly

Overview

Scenario has no editing tools on the MCP surface. Every compositor, concatenator, trimmer and captioner is a deterministic model run through model_run. The compositor, concatenator and trimmer ids are fixed and named below; Scenario ships two captioners, so that one is discovered with recommend. Connection and the core loop: see the scenario skill; the clips themselves: see scenario-video. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference: pick the backend first

NeedModel
Clips end to end, optional transitionsmodel_scenario-video-concat
Anything overlapping: overlay, bed, titlesmodel_scenario-compose-video
Still layout (thumbnails, sheets, key art)model_scenario-compose-image

These ids are constants: each is Scenario's single deterministic tool for its operation, so discovery would only re-derive it.

Concat is sequential: videos takes 2 to 50 files (never 1), preserveAudio (default true) carries each clip's own track into the cut, and the optional transitions array's "length must be number of videos - 1". Each entry is an object {type, duration}, never a bare string: type is one of 18 names read off the schema (fade, the default, is the plain crossfade; dissolve is a different effect; none a hard cut) and duration runs 0.1 to 5 seconds, default 0.5. Omit the array for hard cuts throughout, the right call whenever the brief does not ask one shot to flow into the next.

preserveAudio is not a guarantee of sync: a concatenated cut has come back with its audio drifting against the picture, with and without transitions (a confirmed defect at authoring time). A concat of clips carrying dialogue or effects is verified, never trusted: asset_get the output and compare properties.duration with the inputs' sum (each transition shortens the cut by up to its duration), then download it and check that the audio and video streams run the same length (a local probe of the download is inspection, not editing) and that a visible cue (a cut, a mouth opening, an impact) lands on its sound. When speech must stay locked to the picture, assemble in Video Studio instead: each clip layer carries its own track at its startTime, so a drift introduced by joining has nowhere to arise, at the price of hard cuts or per-layer fades in place of the transition set. Concat stays the tool for silent clips and for cuts whose bed is laid afterwards.

Video Studio is an absolute timeline: layers holds 1 to 50 image, video or audio sources placed by startTime and stacked by zIndex (higher in front), with per-layer fadeIn/fadeOut rather than a between-clip array, so a cross-fade between two clips belongs in concat. At least one layer must be a video, so a music bed over a still is rejected.

The Video Studio contract

Per layer, three timing ideas have confusable names:

  • startTime and endTime place the layer on the timeline.
  • trimStart and trimEnd are amounts shaved off the head and tail of the source, not in and out points: to play the first 20 seconds of a 90-second track, set trimEnd: 70.
  • duration overrides the layer's length. The top-level duration is a different field, setting the whole composition, and applies only when durationMode is "custom".

Everything is in seconds at step: 0.1, so do not promise frame-accurate cuts. x and y are strings taking pixels, percentages or words, and anchor says which point of the layer they address: top-left, top-center, top-right, center-left, center, center-right, bottom-left, bottom-center, bottom-right (the middle one is bare center), defaulting to "top-left". A bottom-right logo is x: "right", y: "bottom", anchor: "bottom-right", flush to the edge unless you inset it with a percentage.

canvasMode and durationMode both default to "auto", computed from the layers: a custom frame is canvasMode: "custom" plus numeric canvasWidth and canvasHeight (there are no top-level width/height, and unlike the layers' string width/height these are numbers), and a custom length is durationMode: "custom" plus the top-level duration. fit only acts on a layer that also sets the layer's own width and height, strings like x and y: to reframe a clip, set both to the canvas size as bare pixel numbers in strings (width: "1920", height: "1080", the form of the schema's own "0" default; a percent string also works) and pass "cover" (crops) or "contain" (letterboxes), since the default "fill" stretches. Give an overlay layer an explicit width and height even at its native size: leaving them empty is documented as preserving the original but does not reliably, and a 560x80 strip came back at 1792x256 on a 1920x1080 canvas. The output field is videoOutputFormat here, imageOutputFormat in Image Studio, and outputFormat in concat and the trim utilities.

A layer's type (image, video, audio, or mask) is optional and otherwise inferred from source. There is no text layer, so titles, lower thirds and end cards arrive as rendered PNGs composited as image layers; scenario-text-overlay renders them letter-perfect.

Show full SKILL.md (502 more words)Show less

Worked example: three clips, a music bed, and captions

  1. For every source, asset_get and read properties: the real duration (audio carries it too), frameRate on video, and width/height on video and images alike. A model asked for eight seconds does not always return exactly eight.
  2. Add them up to get each clip's startTime, then model_schema_get on model_scenario-compose-video.
  3. model_run it with layers holding the three clips at their computed startTimes (zIndex: 0), a logo PNG pinned by x/y/anchor at zIndex: 1, and the music as an audio layer trimmed to the clips' total with trimEnd, since auto durationMode computes from every layer, the bed included. volume (0 to 2) and mute are per-layer and apply to video layers too, so balance the bed from either side. Set canvasMode: "custom" with canvasWidth and canvasHeight; fps defaults to 30 and takes whole numbers from 1 to 120, so set it to the clips' measured frameRate when they agree (rounded), which is what step 1 read it for.
  4. jobs_wait, then caption the finished cut rather than each clip: model_scenario-caption-studio (an authoring-time pick, since Scenario ships two captioners: re-discover with recommend) takes it as video and transcribes whatever audio the master actually carries, so lower the bed's volume rather than muting the dialogue. For Caption Studio, follow scenario-caption-studio for corrected-SRT input and rendered-caption verification rather than inferring an upload route from a file-typed schema: the exposed upload_asset kinds do not include text. It burns captions in by default, returns a sidecar when outputSrt is true, and positions them with textPosition (top, middle or bottom, default bottom, so move it off a bottom-edge overlay). Other captioners use their own live input contracts.
  5. asset_display to review, then asset_download with format left unset (it is an image conversion target) for the file.

For a revision handoff, retain the clean master, source assets, actual submitted layer payload, measured timings, narration text and voice settings, title PNGs with their editable layout settings, and any corrected SRT and returned caption theme. Keep output ids linked to the recipe. Revise the requested layer and recaption from the clean master; do not stack new captions over a burned export. These records support reconstruction, not a promise of a native editor project or a tested replay in a new environment.

Trim, split and resize have their own tool models. Note model_scenario-video-cut's startTime/endTime are source in and out points, not timeline placement.

Common mistakes

  • Reaching for an MCP tool like video_compose, or for local ffmpeg to do the edit itself. The editing surface is model_run on tool models; probing a download is fine.
  • Handing concat one clip, or one transition per clip: the minimum is 2 videos, and transitions run one fewer.
  • Trusting dry_run here as validation: it returns a cost estimate only, and assembly payloads are the most structurally complex in the catalog.
  • Captioning a concatenated master before checking its sync: the captioner times its lines to the audio it hears, so drifted audio ships drifted captions on top of it.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-video-assembly of scenario-labs/skills.

Open the folder on GitHubat commit f6f8ab7

Compare with similar skills

Scenario Video Assembly next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Video Assembly compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Video Assembly this skillscenario-labs/skills946—~2.3kAutomated safety check: PassMIT
Proofreadsonilo-ai/skills115—~4.4kAutomated safety check: NotesMIT
ShowtimeFavioVazquez/showtime220—~3kAutomated safety check: PassMIT
Azure AImicrosoft/GitHub-Copilot-for-Azure2551 repos~852Automated safety check: PassMIT
Audio And Videoglifxyz/glif-mcp-server213—~1.6kAutomated safety check: PassMIT
OrchestrationOrkas-AI/Orkas-VideoStudio499—~3.4kAutomated safety check: PassMIT

Similar skills

  • Proofread

    sonilo-ai/skills

    Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before…

    115 GitHub stars~4.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    220 GitHub stars~3k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed
  • Audio And Video

    glifxyz/glif-mcp-server

    Make or edit audio and video with Glif, from a text brief or from a reference image, video or audio file.

    213 GitHub stars~1.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Orchestration

    Orkas-AI/Orkas-VideoStudio

    The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps.

    499 GitHub stars~3.4k tokensUpdated 19 days ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes

More from scenario-labs/skills

All 146 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    946 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    946 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Chatgpt Pet Create

    scenario-labs/skills

    A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…

    946 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Video Assembly

What does Scenario Video Assembly do?

A skill your agent uses when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or…. Scenario Video Assembly is an agent skill from scenario-labs/skills. Use when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or title card, adding a music bed or voiceover, burning in or exporting captions and subtitles, trimming and splitting, or producing vertical and square variants of one master.

When should I use Scenario Video Assembly?

Scenario Video Assembly fits situations like: generated clips must become a finished video on Scenario via MCP: cutting a shot list together; laying a timeline; concatenating with transitions; overlaying a logo.

How do I install Scenario Video Assembly in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-video-assembly -a claude-code`. Or copy the skill folder (skills/scenario-video-assembly in scenario-labs/skills) into .claude/skills/scenario-video-assembly in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Video Assembly in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-video-assembly -a codex`. Or copy the skill folder (skills/scenario-video-assembly in scenario-labs/skills) into .agents/skills/scenario-video-assembly in your project. Codex loads it when a task matches its description.

Can I use Scenario Video Assembly in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-video-assembly -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-video-assembly, .gemini/skills/scenario-video-assembly, .github/skills/scenario-video-assembly and .opencode/skills/scenario-video-assembly in your project.

What does Scenario Video Assembly need to run?

Going by SKILL.md and its folder, Scenario Video Assembly needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Scenario Video Assembly access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Video Assembly safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Video Assembly use?

Scenario Video Assembly is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Video Assembly use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Video Assembly?

Skills that share tags, products or a category with Scenario Video Assembly: Proofread (sonilo-ai/skills, 115 stars), Showtime (FavioVazquez/showtime, 220 stars), Azure AI (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Audio And Video (glifxyz/glif-mcp-server, 213 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Video Assembly?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.