Agent skill

Stage Assemble

by Orkas-AI in Orkas-AI/Orkas-VideoStudio

Deterministically assemble an approved cross-modal EDL (plan.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg…

MITAuto-check passedMedia & Creative

Install Stage Assemble

skills CLI
$ npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-assemble -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orkas-AI/Orkas-VideoStudio stage-assemble --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orkas-AI/Orkas-VideoStudio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/stage-assemble .claude/skills/stage-assemble && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stage-assemble
GitHub stars
499
Token cost
~3.4k tokens
SKILL.md length
1,816 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Deterministically assemble an approved cross-modal EDL (plan.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg…

  • Works in 4 steps: Produce each segment (delegate by source) → Assemble in ffmpeg tiers (the default… → Idempotent resume → …
  • After a plan is validated and approved
  • SKILL.md covers Step 1 — Produce each segment…, Step 2 — Assemble in ffmpeg…, Director judgment (end-to-end… and Step 3 — Idempotent resume, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stage Assemble is an agent skill from Orkas-AI/Orkas-VideoStudio. Deterministically assemble an approved cross-modal EDL (plan.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg tiers: concat primaries → overlay composed layers → mix narration once (coverage check) → burn captions → verify loudness; idempotent-resumable, with QA before the draft gate. Trigger after a plan is validated and approved; do NOT trigger to ingest or re-author the plan (stage-plan).

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Video production, Text to speech and voice and AI video generation. It works with FFmpeg. The repository describes itself as: Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file. The licence is MIT.

When your agent uses it

  • After a plan is validated and approved
  • Do NOT trigger to ingest
  • Re-author the plan (stage-plan)

Example prompts

  • “/stage-assemble”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Produce each segment (delegate by source)
  2. Assemble in ffmpeg tiers (the default path)
  3. Idempotent resume
  4. QA report, then gate D

What it can do on your machine

Read from SKILL.md and the folder at commit c0c3c3e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stage Assemble loads about 3.4k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 1,816 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orkas-AI/Orkas-VideoStudio at commit c0c3c3e, republished under its MIT licence (© Orkas-AI). 1,816 words, ~3,379 tokens.

Download SKILL.mdSave it as .claude/skills/stage-assemble/SKILL.md (or your agent's skills folder).
name
stage-assemble
description
Deterministically assemble an approved cross-modal EDL (plan.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg tiers: concat primaries → overlay composed layers → mix narration once (coverage check) → burn captions → verify loudness; idempotent-resumable, with QA before the draft gate. Trigger after a plan is validated and approved; do NOT trigger to ingest or re-author the plan (stage-plan).

stage-assemble

How to execute a validated project/plan.json into one finished file. By the time you are here the plan passed ovs plan validate and the user approved it at gate B. Walk it; do not re-plan. The producers are ovs edit, ovs draft, ovs video/ovs image, ovs speak, and the assembler is ovs edit (or the equivalent MCP tools).

Step 1 — Produce each segment (delegate by source)

Iterate segments in order. For each, produce its produced_path according to source, then write that path + status:"done" back into the segment so a resume never re-produces it:

  • edit → stage-edit: ovs edit trim the input_id to [in_sec, out_sec] → project/cuts/<id>.mp4.
  • compose → stage-compose: build a small visual-only manifest-owned composition for spec.kind (title card, lower-third, stat card, captions) under project/compositions/<id>/ → run ovs draft project/compositions/<id> --out project/parts/<id>.mp4 --quality draft --report project/reports/<id>-compose-report.json. This keeps compose segments on the same manifest/source/check/video-QA path as standalone COMPOSE while still letting the assembler own narration and loudness.
  • generate → stage-generate (+ stage-consistency for recurring characters): only AFTER gate C. ovs video/ovs image → project/assets/<id>.mp4. For operation:"edit", pass the exact original reference video and obey top-level references plus edit_strategy; never widen it into regeneration. A failed/unknown paid attempt is not an automatic retry. Preserve completed siblings and require a new output path for any later authorized attempt.
  • provided → use spec.asset_id as-is (probe it first; conform aspect/fps if needed).

Billable generate segments must not run before gate C has confirmed the count from cost_estimate. Produce cheap/free segments (edit, compose, provided) freely.

Step 2 — Assemble in ffmpeg tiers (the default path)

Assemble deterministically, bottom-up. This tiered order is the default; it is predictable and cheap, and keeps each clip's real audio intact:

  1. Primary track — ovs edit concat the primary-layer produced_paths in order → project/render/primary.mp4. Conform aspect/fps on the way in if sources differ; read the returned conformance report and verify it still matches the approved canvas.
  2. Overlays / bg — Do not place a full-frame opaque composed video over source footage: it erases the base. Use genuinely transparent or bounded overlays, or amend the plan to make it a primary beat. Do not bypass the opaque-overlay refusal with resizing tricks. for each overlay/bg segment, ovs edit overlay its part onto the primary over the window of the segment named in over (title cards, lower-thirds, logos). Composed layers are VISUAL-ONLY — they must not carry their own narration audio. This includes a compose segment that IS the primary track (a full-video composition): render it SILENT — do not put a narration <audio> in its index.html. The assembler owns narration (tier 3), so a composition that bakes it in would mean narration is added TWICE (the "two voices" defect).
  3. Narration — added EXACTLY ONCE, here. If active, require the Gate-B-signed tracks.narration.synthesis profile. Run ovs narration fit for each timed line before ovs speak; shorten over-budget text in the plan without changing the approved meaning. Synthesize with the exact signed voice/model/format/speed, probe the result, rerun measured fit, and write each line's produced_path. If measured timing misses, revise once using the suggested unit budget rather than repeatedly billing or forcing speed. Add all produced lines in ONE ovs edit mix call at their start_sec. The default existing-audio rejection catches compose segments that accidentally baked narration; re-render those SILENT. Read voicedRatio, interiorGaps, maxOverlapSec and status: the last line reaching the end does not mean the whole track carries speech. Fix unintended dead air and collisions, and disclose any intentional silent tail at Gate D. Preserve a generated talking head's built-in lip-synced audio instead of adding a second voice.
  4. Music — add tracks.music ducked under narration by the planned amount.
  5. Captions — turn tracks.captions.lines ({text, start_sec, target_sec}) into a .srt, then ovs edit burnsubs. Avoid duplicating caption text already visibly present in that scene. Captions are DATA in the plan — burned ONLY here at assemble — so a later typo fix reuses the clean pre-caption primary/generated source and burns the revised file once. Never burn onto the already-captioned output and never regenerate an unchanged provider clip. If burnsubs fails because the runtime ffmpeg lacks subtitle filter support, stop and report that blocker; do not hand-write a fallback ffmpeg graph, PNG subtitle overlay, or drawtext pipeline.
  6. Loudness — run ovs edit normalize-loudness project/render/draft.mp4 --out project/render/video.mp4. It normalizes to the video-craft §7 targets (~−14 LUFS integrated, true-peak ≤ ~−1 dBTP) and returns measured loudness; use ovs edit loudness only for diagnosis without writing an output.

Apply the plan's style_kit for cohesion: composed layers (titles/captions/cards) use its palette + fonts. A single lut graded across all clips is what unifies tonally mixed sources — until a grade op is available, keep mixed sources close at capture/trim and lean on the shared palette + consistent captions for cohesion rather than promising a uniform grade.

Output project/render/video.mp4 as the deliverable; project/render/draft.mp4 is the pre-normalized intermediate.

Director judgment (end-to-end assembly)

The craft of making mixed sources feel like one video, on top of the shared craft (video-craft). The seams between footage / generated / composed are where multi-source assembly falls apart — engineer continuity across them:

  • One look across every source. Apply the style_kit so a cut from real footage → a generated shot → a composed card does not read as three videos: one type system + palette on every composed layer, one caption style throughout, matched aspect / fps, tonal proximity (a shared LUT is the unifier when available; video-craft §4).
  • Audio is the through-line that hides the visual seam. One narration voice; a continuous music bed UNDER the cuts (do not restart it per segment); duck consistently (video-craft §7). The ear's continuity carries the eye across a source change — a reveal may drop music, but the bed bridges the cut.
  • Rhythm over a mixed cut. Alternate motion vs. static and source types for momentum — do not stack three composed cards or three talking-head shots in a row (that is the repetition / slideshow smell, video-craft §3, §12). Vary holds.
  • Cut on a content change, not just plan order. A hard cut on a beat / word change is invisible and professional; a crossfade signals a gentle topic shift (video-craft §5).
  • Don't bury the hero. On a source_led piece, composed lower-thirds and captions FRAME the footage — they never cover its subject / face (video-craft §6).
  • Apply the editing cut craft ACROSS the seams. The cut mechanics live in stage-edit → "Cut craft" (best sub-window, ≤ 4 transitions, L/J-cut sound bridges, handles / no freeze-frame, adjacent-diversity, a reason per cut) — apply them at every junction between sources, since the footage → generated → composed seams are exactly where a mixed cut betrays itself.
Show full SKILL.md (746 more words)Show less

Step 3 — Idempotent resume

The plan is the checkpoint. On a re-run, skip any segment already status:"done" with a present produced_path, and skip assembly tiers whose output already exists and is newer than its inputs. Never re-run a billable generate segment that is already produced.

For a partial child failure or later revision, reset only the affected child and assembly tiers derived from it. Preserve every unaffected completed child and its evidence. When the child becomes valid, rebuild the complete parent draft and run parent QA in the same workflow; do not stop at a child-only recovery artifact.

Step 4 — QA report, then gate D

Before showing the draft, run the QA pass and write project/render_report.json with these sections:

  • technical_probe — ovs edit probe the draft/final (real duration / resolution / fps / audio present); confirm it matches the plan's aspect + total.
  • promise_preservation — ovs plan promise-check project/plan.json --probe-produced. At gate D this probes each primary segment's produced_path and computes the REAL primary-track motion ratio vs. motion_min_ratio plus the source_required invariant; missing/unreadable produced media or a fail means "slideshow / promise broken" — do not deliver. Send it back (below). Do not eyeball this; let the numbers decide.
  • visual_spotcheck — extract ~4 frames across the draft (ovs edit extract-frame) and read them for upside-down / garbled-caption / empty / wrong-product frames. Read them yourself if you are multimodal; if you cannot see images, record the spot-check as unverified and proceed — do not invent what the frames show.
  • audio_spotcheck — the normalize-loudness measured loudness numbers + the narration coverage result from step 2 (uncovered tail / overshoot / silent lead-in).
  • transcript_comparison (when there is narration) — optionally ovs transcribe the draft and confirm the spoken words match the planned narration lines.

Each section carries pass / warn / fail + a one-line reason. Then present the draft + headline findings at Gate D and resolve the user's decision through ovs gate transition.

On approve → finalize project/render/video.mp4 (loudness / captions only; never re-synthesize a talking-head voice). On revise → redo only the affected segment(s) and re-assemble.

Send-back (self-correction on a QA fail)

A QA fail does not go to the user as "here's a broken video". Diagnose which segment(s) caused it and redo ONLY those, then re-assemble and re-run QA:

  • promise_preservation fail (slideshow) → the static composed segments are too long / the motion segments too short. Rebalance segment durations or convert a static beat to footage, re-assemble.
  • visual_spotcheck fail (bad frame) → re-produce that one segment (re-trim / re-compose / re-generate), not the whole video.
  • audio fail (uncovered tail) → re-time or extend the narration / trim the tail.

Bound repetition, not recovery: allow at most 2 send-back rounds for the same failing check and unchanged strategy. Then preserve the current artifact and show concrete user directions through gate-control before starting another cycle. Create no technical confirmation. Only a required signed-plan change or new billable attempt returns through its normal authorization boundary.

Rules

  • Walk the approved plan; if assembly reveals the plan is wrong, surface it and re-gate — do not silently re-plan.
  • Write produced_path + status back per segment as you go (resumability + the QA pass depend on it).
  • One output file is the deliverable; cuts/ and parts/ are intermediates.
  • Narration is added exactly ONCE — in the mix tier, never baked into a compose render. Compose segments (including a full-video composition used as the primary track) render SILENT (no narration <audio>); the assembler mixes narration via ovs edit mix with segments placed per line. The mix's default --on-existing-audio reject enforces this — a "base already has audio" mix rejection is the signal a segment wrongly baked audio in; re-render it silent, then re-mix.
  • No ad-hoc ffmpeg fallbacks for captions. Caption burn-in is a low-freedom operation owned by ovs edit burnsubs; a failed burnsubs call is a tool/runtime blocker, not permission to invent a custom subtitles/drawtext/PNG-overlay command.

Boundary / non-goals

This skill assembles an already-approved plan. It does not ingest or decide the plan (stage-plan), and it delegates the actual production of each segment to the compose / generate / edit / consistency skills rather than re-deriving their craft here.

Before showing any plan-backed final, regardless of assembly route, run ovs plan promise-check project/plan.json --probe-produced --video project/render/video.mp4. Repair duration/aspect mismatch, overlapping/truncated narration and unverifiable lines. Explain caption warnings from the actual container/sidecar check; burned captions need visual evidence. Planned duration and a successful mix alone do not prove delivery.

For narration, preserve signed line windows; target_sec is duration. Measure each produced file and voice span. Keep intentional silence, shorten overlong text within authorized scope, and never slide later lines to hide an overlap. A failed/unknown provider request does not authorize automatic repeat billing.

© Orkas-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills/stage-assemble of Orkas-AI/Orkas-VideoStudio.

Open the folder on GitHubat commit c0c3c3e

Compare with similar skills

Stage Assemble next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stage Assemble compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stage Assemble this skillOrkas-AI/Orkas-VideoStudio499—~3.4kAutomated safety check: PassMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
AI Presenter VideoNousResearch/hermes-agent253k—~2.3kAutomated safety check: PassMIT
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Hearyourvoicekillernay/HearYourVOICE140—~10kAutomated safety check: NotesMIT
Muapi DirectorAnil-matcha/vox-ai-motion-graphics-generator246—~679Automated safety check: PassNone

Similar skills

  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • AI Presenter Video

    NousResearch/hermes-agent

    Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports.

    253k GitHub stars~2.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Hearyourvoice

    killernay/HearYourVOICE

    The repeatable workflow for short Thai documentary/explainer videos — one topic in, one finished MP4 out.

    140 GitHub stars~10k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Muapi Director

    Anil-matcha/vox-ai-motion-graphics-generator

    Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…

    246 GitHub stars~679 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    190 GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed

More from Orkas-AI/Orkas-VideoStudio

All 14 skills in this repo
  • Gate Control

    Orkas-AI/Orkas-VideoStudio

    Canonical VideoStudio review authorization and state-transition policy.

    499 GitHub stars~2.7k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Compose

    Orkas-AI/Orkas-VideoStudio

    Authoring knowledge for Orkas/OVS HTML video compositions -- write an index.html, drive animation from a paused timeline, declare canvas + duration, then run the VideoStudio draft gate to render an…

    499 GitHub stars~4.2k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    499 GitHub stars~2.4k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Consistency

    Orkas-AI/Orkas-VideoStudio

    Multi-shot narrative & character consistency — a character bible with a locked front-portrait anchor, view-matched reference selection, recent-frame carry-forward, Cameo (a user photo as the lead)…

    499 GitHub stars~1.8k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Plan

    Orkas-AI/Orkas-VideoStudio

    The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

    499 GitHub stars~4k tokensUpdated 19 days ago
    Auto-check passed
  • Frontend Design

    Orkas-AI/Orkas-VideoStudio

    Aesthetic direction for OrkasVideoStudio HTML and motion-graphics compositions.

    499 GitHub stars~4.3k tokensUpdated 19 days ago
    Auto-check passed

Works with

Questions about Stage Assemble

What does Stage Assemble do?

Deterministically assemble an approved cross-modal EDL (plan.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg…. Stage Assemble is an agent skill from Orkas-AI/Orkas-VideoStudio.json) into a finished video — produce each segment (edit/compose/generate/provided, delegated to its line), then assemble in ffmpeg tiers: concat primaries → overlay composed layers → mix narration once (coverage check) → burn captions → verify loudness; idempotent-resumable, with QA before the draft gate.

When should I use Stage Assemble?

Stage Assemble fits situations like: after a plan is validated and approved; do NOT trigger to ingest; re-author the plan (stage-plan).

How do I install Stage Assemble in Claude Code?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-assemble -a claude-code`. Or copy the skill folder (packages/skills/stage-assemble in Orkas-AI/Orkas-VideoStudio) into .claude/skills/stage-assemble in your project. Claude Code loads it when a task matches its description.

How do I install Stage Assemble in Codex?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-assemble -a codex`. Or copy the skill folder (packages/skills/stage-assemble in Orkas-AI/Orkas-VideoStudio) into .agents/skills/stage-assemble in your project. Codex loads it when a task matches its description.

Can I use Stage Assemble in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-assemble -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stage-assemble, .gemini/skills/stage-assemble, .github/skills/stage-assemble and .opencode/skills/stage-assemble in your project.

What does Stage Assemble need to run?

SKILL.md names no scripts, command-line tools or credentials: Stage Assemble is instructions for the agent only.

Does Stage Assemble access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stage Assemble safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stage Assemble use?

Stage Assemble is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stage Assemble use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stage Assemble?

Skills that share tags, products or a category with Stage Assemble: Vox Director (Alisa0808/vox-director, 2.2k stars), AI Presenter Video (NousResearch/hermes-agent, 253k stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars) and Hearyourvoice (killernay/HearYourVOICE, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stage Assemble?

Orkas-AI (a GitHub user) maintains it in Orkas-AI/Orkas-VideoStudio, which has 499 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 22, 2026.

Source: Orkas-AI/Orkas-VideoStudio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.