Agent skill

Orchestration

by Orkas-AI in Orkas-AI/Orkas-VideoStudio

The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps.

MITAuto-check passedMedia & Creative

Install Orchestration

skills CLI
$ npx skills add Orkas-AI/Orkas-VideoStudio --skill orchestration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orkas-AI/Orkas-VideoStudio orchestration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orkas-AI/Orkas-VideoStudio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/orchestration .claude/skills/orchestration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
orchestration
GitHub stars
499
Token cost
~3.4k tokens
SKILL.md length
1,764 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps.

  • Works in 2 steps: Route + lock (read video-router) → GATE A — Proposal (all lines)
  • Make / edit / cut / caption / dub / animate a video
  • SKILL.md covers Checkpoint protocol — how…, 1. Route + lock (read…, 2. GATE A — Proposal (all lines) and 2.5 Craft standard (ALL lines…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Orchestration is an agent skill from Orkas-AI/Orkas-VideoStudio. The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps. Trigger for "make / edit / cut / caption / dub / animate a video"; it sequences the compose / generate / edit lines and the approval gates. Do NOT trigger for a single low-level operation (just transcribe a file, just probe a clip) — call that operation directly.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, AI video generation and Text to speech and voice. It works with Model Context Protocol. The repository describes itself as: Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file. The licence is MIT.

When your agent uses it

  • Make / edit / cut / caption / dub / animate a video
  • It sequences the compose / generate / edit lines and the approval gates
  • A single low-level operation (just transcribe a file
  • Just probe a clip) — call that operation directly

Example prompts

  • “make / edit / cut / caption / dub / animate a video”
  • “/orchestration”

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Route + lock (read video-router)
  2. GATE A — Proposal (all lines)

What it can do on your machine

Read from SKILL.md and the folder at commit c0c3c3e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Orchestration loads about 3.4k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,764 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orkas-AI/Orkas-VideoStudio at commit c0c3c3e, republished under its MIT licence (© Orkas-AI). 1,764 words, ~3,361 tokens.

Download SKILL.mdSave it as .claude/skills/orchestration/SKILL.md (or your agent's skills folder).
name
orchestration
description
The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps. Trigger for "make / edit / cut / caption / dub / animate a video"; it sequences the compose / generate / edit lines and the approval gates. Do NOT trigger for a single low-level operation (just transcribe a file, just probe a clip) — call that operation directly.

orchestration

You are producing a short video. Run this as a TIGHT program — lean turns — but STOP at every GATE so the creative decisions stay the user's. Only the technical/assembly steps between gates run unattended. The CLI surface is ovs ... (an MCP server mirrors it 1:1; use whichever your host exposes). All work for one deliverable lives under a single project dir, e.g. project/.

Read gate-control once before the first user gate. It is the single authorization policy across all lines. After every gate reply, resumed approval, post-gate revision, or exhausted visual-QA result, run ovs gate transition and follow its one returned action; line sections below define artifacts and production steps, not a competing approval state machine.

Checkpoint protocol — how every GATE works (there is no special form UI)

  1. Show the artifact in chat so the user can actually see it — script/plan as markdown, images inline, a draft video as its output file path — plus one line of "what I'll do next" and any cost/QA note.
  2. State the options for that gate and WAIT for the user to reply. Do not run the next production step in the same turn as the gate.
  3. On reply, resolve the choice with ovs gate transition: approve → continue only with the returned operation; revise → redo only the authorized scope, re-show, and re-gate only when the resolver says so; abort → stop. Never pass a gate without explicit user confirmation, and never ask again for an unchanged artifact whose approval is already recorded.

1. Route + lock (read video-router)

Classify and LOCK the line (no silent switching):

  • COMPOSE — explain / teach / animate / motion-graphics / kinetic text, no source footage → ovs draft (+ optional ovs image / ovs video imagery, optional ovs speak narration).
  • GENERATE — "footage of / a scene of / cinematic / a presenter or avatar speaking / talking-head" → AI footage via ovs video (+ ovs image for the subject, ovs speak for voice), assembled with ovs edit.
  • EDIT — the user supplied real clips to cut / join / subtitle / localize → ovs edit (+ ovs transcribe for transcript-driven work).
  • AUTO (end-to-end) — the primary timeline weaves MORE THAN ONE axis; adding audio/captions to an existing video remains EDIT. Run the cross-modal orchestration (read stage-plan, then stage-assemble); the lock is the plan's delivery_promise.

2. GATE A — Proposal (all lines)

Show: the brief you inferred (line, aspect, duration, language) for the user to correct, plus 1–3 differentiated concepts (each: hook + look + rough length; for GENERATE add the shot count and that each clip is a billable call; for AUTO also state the proposed delivery promise — source_led / motion_led / compose_led / hybrid — and the rough segment mix). Options: pick a concept / adjust the brief / new direction. STOP.

2.5 Craft standard (ALL lines — read video-craft)

Before scripting / storyboarding / composing / generating, hold the output to video-craft: a hook in the first seconds, one idea per beat, readable type inside safe zones, restrained easing, muted-friendly captions, ducked audio, the right aspect for the platform. Bake these into the script/shotlist/composition — don't leave them to chance.

For COMPOSE or AUTO compose segments, also apply frontend-design before writing HTML. If the user supplied a DESIGN.md, brand guide, screenshot, existing app UI, Figma notes, or an explicit named style, apply design-system-importer. After a rendered draft, use composition-design-review only when its trigger applies and treat non-blocking findings as Gate D notes.

2.6 Narration voice

When the piece has voiceover, run ovs speech-capabilities and copy its executable route/model/voice/format into the Gate B plan together with the BCP-47 video language and a natural speed. Do not invent a voice id. Before Gate B, run ovs narration fit --text ... --target ...; revise over/under text internally before any paid synthesis. After ovs speak, probe the produced audio and run the same fit with --measured; retime scenes from the measured duration without silently shortening the approved target. If no TTS provider is configured, tell the user and explicitly choose silent delivery or wait for configuration.

TALKING-HEAD note: if a GENERATE clip already returned lip-synced built-in speech, THAT is the voice — do NOT synthesize a narration over it (a fresh TTS track desyncs from the mouth). Use ovs speak only for a silent clip, or for COMPOSE / EDIT / off-screen voiceover.


COMPOSE line

3C. Write project/composition/composition-manifest.json as the single plan: timeline, exact copy/narration, audio intent, language and art direction. Run free narration fit before presenting it. Do not require a duplicate script or shotlist. 4C. GATE B — Production plan confirmation. Show the canonical manifest as a readable timed plan, including exact copy and voice. Options: approve / revise / change direction. STOP. 5C. (optional) Narration: after the free fit passes, ovs speak once → project/composition/assets/narration.mp3, probe/measure it, retime the composition within the approved target, and add it as an <audio> track (see stage-compose). For a STANDALONE compose deliverable only; in AUTO the assembler mixes narration and compose segments render SILENT. 6C. (optional) Visual assets via ovs image / ovs video → project/assets/. Skip for pure typographic explainers. If any asset is billable: GATE C first — state the count + that they're billable; options approve & generate / adjust / skip. STOP, then generate. 7C. Compose (stage-compose) → project/composition/composition-manifest.json v2 and project/composition/index.html; prepare once and reconcile after manifest timing/audio changes. Use the optional HTML Preview Gate from stage-compose only when render rework is likely expensive. 8C. QA + draft: ovs draft project/composition --out project/render/draft.mp4 --quality draft --report project/render/draft-report.json --findings project/composition/qa/check.json; repair only concrete blockers within the bounded repair budget. 9C. GATE D — Draft review. Show the draft path + report/check/craft/design-review findings. Options: approve → high export / revise. STOP, then run ovs draft project/composition --out project/render/video.mp4 --quality high --report project/render/final-report.json --findings project/composition/qa/final-check.json once.

GENERATE line (follow stage-generate; for recurring characters / a story, ALSO stage-consistency)

3G. Plan shots → project/shotlist.json (each: prompt, motion, duration, which characters are visible + camera angle). Keep the count tight (cost). For a STORY / long script: first build the global character bible + scene plan per stage-consistency. 4G. GATE B — Script + shotlist sign-off. Options: approve / revise / change direction. STOP. 5G. Character anchors (talking-head / any recurring character) — per stage-consistency: generate ONE locked front portrait per character with ovs image → project/characters/ (or, for a cameo, use the user-uploaded photo AS the front portrait); record in project/characters/bible.json. NEVER regenerate a locked portrait. 6G. GATE C — Pre-generation confirm. State "N shots + M portraits ≈ X billable generations"; show the locked portrait(s) for approval. Options: approve & generate / adjust portrait / reduce scope. STOP. 7G. Generate each shot — per stage-consistency: pick references (angle-matched portrait of each visible character + the most recent prior frame), write an explicit ref→element prompt, ovs video image-to-video → project/assets/shot-N.mp4 (KEEP the clip's built-in lip-synced audio; call ovs speak ONLY when the clip came back silent). After each shot, ovs edit extract-frame its last frame → project/frames/ to carry forward. 8G. Assemble: ovs edit concat the shots → project/render/draft.mp4; add captions/lower-thirds via a small composition (stage-compose) + ovs edit overlay / ovs edit burnsubs. Captions/overlays are VISUAL-ONLY — preserve each clip's built-in audio through assembly. 9G. GATE D — Draft review. Show the assembled draft + craft findings. Options: approve / revise. STOP, then finalize → project/render/video.mp4 (carry talking-head audio through UNCHANGED; captions/loudness only).

Show full SKILL.md (616 more words)Show less

EDIT line (read stage-edit)

3E. Ingest — ovs edit probe each clip for durations/resolution. 4E. Plan → project/plan.json (the segments EDL + a tracks.narration track of timed lines, so each narration line / caption stays separately re-editable). GROUND the plan on the clip's ACTUAL content BEFORE writing any narration:

  • spoken audio → ovs transcribe raw/clip.mp4 --out project/transcripts/clip.json (--model large-v3 for non-English), pick timecodes from the transcript.
  • SILENT / screen-recording / slideshow → extract representative frames across the clip with ovs edit extract-frame and inspect those images with the current multimodal model. Then write each narration segment to match the visible slide in its window. Do NOT narrate from prior knowledge of the topic.
  • If frame reading is unavailable, STOP and ask the user for the on-screen beats. Never invent narration. 5E. GATE B — Edit plan sign-off. Show the plan (segments / order / subtitles / localization; for highlights, the chosen moments + timecodes). Options: approve / revise / change selection. STOP. 6E. Execute: ovs edit trim → project/cuts/; ovs edit concat → project/render/draft.mp4; optional burnsubs / overlay. Adding narration to footage that ALREADY has audio: ovs edit mix — choose on purpose: keep the original under the voice (--on-existing-audio mix) or drop it (--on-existing-audio replace). It rejects by default when the base already has audio, so decide explicitly. Localization/dubbing: transcribe → translate → ovs speak → ovs edit mix --on-existing-audio replace + burn captions. 7E. GATE D — Draft review. Show the edited draft + mix coverage (when narration was added) + the loudness check (ovs edit loudness vs ~−14 LUFS / −1 dBTP). Options: approve / revise. STOP, then finalize with ovs edit normalize-loudness when needed → project/render/edited.mp4.

AUTO end-to-end line (read stage-plan, then stage-assemble)

Ingest every supplied clip from evidence (probe + transcript and/or inspected extracted frames), author ONE cross-modal project/plan.json, ovs plan validate and fix every error, GATE B on the timeline (ovs plan summarize), GATE C only if the plan has billable generate segments with the exact count and exact media_kind/duration/ratio/resolution/audio/reference settings, then assemble per stage-assemble (produce each segment via its line, mix narration ONCE, music ducked, burnsubs, normalize loudness). At GATE D run ovs plan promise-check project/plan.json --probe-produced --video project/render/video.mp4 plus the production-method review defined by video-router, then finalize.


plan.json as the editable record (all lines) — keep follow-up edits cheap

Once a draft exists, keep project/plan.json faithful so a later tweak only re-touches one piece (never the whole video): (1) every produced segment carries its real output under produced_path + status:"done"; (2) narration is tracks.narration whose lines each carry their own produced_path, so one line can be re-voiced alone; (3) captions are DATA in tracks.captions.lines ({text, start_sec, target_sec}) — NOT burned into the picture — so a typo is a one-line edit re-burned at assemble; (4) set top-level "draft": "render/draft.mp4".

Local follow-up edits — make the minimal targeted change; never redo the whole video. Once a plan.json is present and the user asks to change ONE local thing (a segment's narration / caption / text, a trim, volume / speed, a single shot swap), edit ONLY the matching entry in plan.json and re-produce ONLY what it touched (ovs speak for that one line, ovs draft for that one compose segment, or ovs edit for that one cut), then re-assemble. For caption-only changes to a generated video, reuse the original clean generated segment and burn the revised captions once; never burn onto the already-captioned output or regenerate the source clip. Do NOT re-author the whole EDL and DO NOT regenerate a segment whose status is done that the user did not touch. Fall back to a full re-plan only when the request genuinely restructures the timeline.

Deliver (all lines)

Run the video-craft pre-publish review on the FINAL file (readable / timed / on-message / consistent / audio [talking-head: confirm the built-in lip-synced voice is intact] / platform; fix any blocker first). Present the final output file path + a one-line summary. End the turn.

© Orkas-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills/orchestration of Orkas-AI/Orkas-VideoStudio.

Open the folder on GitHubat commit c0c3c3e

Compare with similar skills

Orchestration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Orchestration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Orchestration this skillOrkas-AI/Orkas-VideoStudio499—~3.4kAutomated safety check: PassMIT
AI Video Scriptzrt-ai-lab/opencode-skills287—~1.7kAutomated safety check: PassNone
Pollinationssundial-org/awesome-openclaw-skills663—~1.7kAutomated safety check: PassNone
Scenario Videoscenario-labs/skills946—~4.3kAutomated safety check: PassMIT
Paw Cra Agent Video Producerpawbytes/skill-suites113—~2.3kAutomated safety check: PassMIT
Anything2explainerVincentwei1021/anything2explainer2.4k—~2.7kAutomated safety check: PassCustom licence

Similar skills

  • AI Video Script

    zrt-ai-lab/opencode-skills

    A skill your agent uses when a request asks for a Chinese-first AI video script with shot plans, image prompts, narration, subtitles, or handoff contracts for scene generation, image generation…

    287 GitHub stars~1.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Pollinations

    sundial-org/awesome-openclaw-skills

    Pollinations.ai API for AI generation - text, images, videos, audio, and analysis.

    663 GitHub stars~1.7k tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed
  • Scenario Video

    scenario-labs/skills

    A skill your agent uses when generating or editing video on Scenario via MCP: text-to-video, a music video from a song, image-to-video (still animation, frame anchors), motion prompting, lipsync and…

    946 GitHub stars~4.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Paw Cra Agent Video Producer

    pawbytes/skill-suites

    Video production specialist for short-form, long-form, episodic, and motion graphics video.

    113 GitHub stars~2.3k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Anything2explainer

    Vincentwei1021/anything2explainer

    给一个主题,产出一条黑底 MG 风格(幕底可选星点或点阵波)、有配音字幕章节进度条的科普讲解视频(中文或英文;Remotion 代码动画;时长由用户定,常用 3–5 分钟)。内含可编译模板、图元库、配音/分镜/渲染工具、风格与动效规范、多 agent 分工协议与 QC 判据,以及一条完整样片(《RAG 与知识库》)作为质量标尺。Turn any topic into a narrated…

    2.4k GitHub stars~2.7k tokensUpdated 22 days ago
    Media & CreativeAuto-check passed
  • Edit Timeline Studio

    MartinDelophy/ai-video-editor

    Analyze images, video, speech, motion, products, and websites; route local vision, audio, depth, tracking, matting, identity, and restoration models; auto-edit, replicate, enhance, caption, voice…

    905 GitHub stars~7k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from Orkas-AI/Orkas-VideoStudio

All 14 skills in this repo
  • Gate Control

    Orkas-AI/Orkas-VideoStudio

    Canonical VideoStudio review authorization and state-transition policy.

    499 GitHub stars~2.7k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Compose

    Orkas-AI/Orkas-VideoStudio

    Authoring knowledge for Orkas/OVS HTML video compositions -- write an index.html, drive animation from a paused timeline, declare canvas + duration, then run the VideoStudio draft gate to render an…

    499 GitHub stars~4.2k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    499 GitHub stars~2.4k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Consistency

    Orkas-AI/Orkas-VideoStudio

    Multi-shot narrative & character consistency — a character bible with a locked front-portrait anchor, view-matched reference selection, recent-frame carry-forward, Cameo (a user photo as the lead)…

    499 GitHub stars~1.8k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Plan

    Orkas-AI/Orkas-VideoStudio

    The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

    499 GitHub stars~4k tokensUpdated 19 days ago
    Auto-check passed
  • Frontend Design

    Orkas-AI/Orkas-VideoStudio

    Aesthetic direction for OrkasVideoStudio HTML and motion-graphics compositions.

    499 GitHub stars~4.3k tokensUpdated 19 days ago
    Auto-check passed

Questions about Orchestration

What does Orchestration do?

The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps. Orchestration is an agent skill from Orkas-AI/Orkas-VideoStudio. The master program for producing or editing a video end to end — read this at the START of any video task (after video-router), then follow the gates and the per-line steps.

When should I use Orchestration?

Orchestration fits situations like: make / edit / cut / caption / dub / animate a video; it sequences the compose / generate / edit lines and the approval gates; A single low-level operation (just transcribe a file; just probe a clip) — call that operation directly.

How do I install Orchestration in Claude Code?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill orchestration -a claude-code`. Or copy the skill folder (packages/skills/orchestration in Orkas-AI/Orkas-VideoStudio) into .claude/skills/orchestration in your project. Claude Code loads it when a task matches its description.

How do I install Orchestration in Codex?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill orchestration -a codex`. Or copy the skill folder (packages/skills/orchestration in Orkas-AI/Orkas-VideoStudio) into .agents/skills/orchestration in your project. Codex loads it when a task matches its description.

Can I use Orchestration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orkas-AI/Orkas-VideoStudio --skill orchestration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/orchestration, .gemini/skills/orchestration, .github/skills/orchestration and .opencode/skills/orchestration in your project.

What does Orchestration need to run?

SKILL.md names no scripts, command-line tools or credentials: Orchestration is instructions for the agent only.

Does Orchestration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Orchestration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Orchestration use?

Orchestration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Orchestration use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Orchestration?

Skills that share tags, products or a category with Orchestration: AI Video Script (zrt-ai-lab/opencode-skills, 287 stars), Pollinations (sundial-org/awesome-openclaw-skills, 663 stars), Scenario Video (scenario-labs/skills, 946 stars) and Paw Cra Agent Video Producer (pawbytes/skill-suites, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Orchestration?

Orkas-AI (a GitHub user) maintains it in Orkas-AI/Orkas-VideoStudio, which has 499 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 22, 2026.

Source: Orkas-AI/Orkas-VideoStudio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.