Agent skill

Stage Plan

by Orkas-AI in Orkas-AI/Orkas-VideoStudio

The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

MITAuto-check passedMedia & Creative

Install Stage Plan

skills CLI
$ npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orkas-AI/Orkas-VideoStudio stage-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orkas-AI/Orkas-VideoStudio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/stage-plan .claude/skills/stage-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stage-plan
GitHub stars
498
Token cost
~4k tokens
SKILL.md length
1,909 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

  • Works in 4 steps: Ingest from evidence, never from… → Choose the delivery promise → Decompose into a cross-modal EDL → …
  • The deliverable spans more than one axis (AUTO line)
  • SKILL.md covers Step 1 — Ingest from evidence,…, Step 2 — Choose the delivery…, Step 3 — Decompose into a… and Step 4 — Validate, then gate B, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stage Plan is an agent skill from Orkas-AI/Orkas-VideoStudio. The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided segments + narration/music/caption tracks + a delivery promise), validate it with ovs plan validate. Trigger when the deliverable spans more than one axis (AUTO line); do NOT trigger for a pure single-axis job (route to that line). The assembler half is stage-assemble.

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and AI video generation. The repository describes itself as: Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file. The licence is MIT.

When your agent uses it

  • The deliverable spans more than one axis (AUTO line)
  • Do NOT trigger for a pure single-axis job (route to that line)

Example prompts

  • “ingest + plan”
  • “/stage-plan”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Ingest from evidence, never from assumption
  2. Choose the delivery promise
  3. Decompose into a cross-modal EDL
  4. Validate, then gate B

What it can do on your machine

Read from SKILL.md and the folder at commit c0c3c3e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stage Plan loads about 4k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,909 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orkas-AI/Orkas-VideoStudio at commit c0c3c3e, republished under its MIT licence (© Orkas-AI). 1,909 words, ~3,983 tokens.

Download SKILL.mdSave it as .claude/skills/stage-plan/SKILL.md (or your agent's skills folder).
name
stage-plan
description
The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided segments + narration/music/caption tracks + a delivery promise), validate it with `ovs plan validate`. Trigger when the deliverable spans more than one axis (AUTO line); do NOT trigger for a pure single-axis job (route to that line). The assembler half is stage-assemble.

stage-plan

How to turn "here is my material + here's the video I want" into a single, inspectable plan spanning reference images/videos, deterministic or semantic editing, generation, and composition. The output is project/plan.json—a cross-modal EDL that the assembler walks. Ingest with the public ovs analysis commands, validate with ovs plan, and execute through the compose/generate/edit lines.

Use this line when the deliverable is NOT cleanly one axis — e.g. "trim my clip, add a title card and captions, and a voiceover", or "my footage for the middle, generate an opener, compose the stats". For a pure single-axis job, route to that single line instead (see video-router).

Step 1 — Ingest from evidence, never from assumption

You cannot plan against material you have not looked at. For EVERY supplied clip, before writing any segment:

  1. Probe it (ovs edit probe) for real duration / resolution / fps / audio presence. A plan that cuts past the real duration breaks.
  2. Read its content the cheapest way that fits:
    • spoken audio → ovs transcribe raw/clip.mp4 --out project/transcripts/clip.json (pass --model large-v3 for non-English) → you now have timecoded words to cut on and a reusable transcript file for edit ops.
    • silent / screen-recording / slideshow → extract representative frames with ovs edit extract-frame and inspect them with the current multimodal model. The audio being empty does NOT mean the screen is. If you cannot inspect images, ask the user for the on-screen beats. Never infer slide/screen content from the topic alone.
    • need to judge what a moment LOOKS like (is the hero shot usable? is the product right-side up?) → read frames: ovs edit extract-frame then look at them. If you cannot see images, say so and plan on probe/transcript evidence alone — mark those judgments unverified, do not invent them.
  3. Record what each input is good for in project/ingest.json: {input_id, duration, has_audio, content_summary, quality_risks:[...], usable_for:[...], planning_implications:[...]}. This is the factual basis the plan cites — segments reference input_ids from here. Rules:
    • content_summary is specific and from observation: "45 s of interview, no b-roll, mono audio" — never "user provided footage". An entry is only "reviewed" if a real probe, transcript, or inspected extracted frame supports it; never claim you looked at a clip you did not.
    • Usability heuristics: video > 10 s → hero footage; > 3 s → b-roll; has speech → dialogue source; audio-only → narration/music source, production must supply the visuals; image-only → motion must come from animation or generation.
    • Quality risks to flag: width < 720 / height < 480 (will look soft), clip < 3 s (limited use), mono audio, a still where the brief wants motion. A flagged risk the plan ignores is a planning bug — resolve it during direction confirmation.

For every supplied image/video that constrains the result, lock its requested relationship as reproduce, edit, or guide. Record roles, protected attributes, allowed changes, and target segments. Video reproduction/editing or motion/timing guidance requires source-time-to-target-segment anchors.

Step 2 — Choose the delivery promise

Pick ONE delivery_promise.type and make the whole plan keep it:

  • source_led — the user's footage is the hero (repurpose / highlight / localize). source_required: true.
  • motion_led — real motion (footage or generated video) dominates; composed cards are accents.
  • compose_led — designed HTML is the spine (explainer / data); footage/generation are accents.
  • hybrid — a deliberate mix (e.g. source hero + composed framing + generated opener).

Set motion_min_ratio to the minimum share of runtime that must be real footage/generated motion rather than composed HTML. For compose_led, use exactly 0; HTML animation quality is enforced by composition motion QA instead of this source-mix ratio. If you cannot hit a nonzero promise from the available material, say so during direction confirmation instead of quietly shipping a slideshow. If source_required is true, at least one PRIMARY segment must be real footage (source: edit | provided) or a semantic video edit that consumes a required top-level original video reference with intent edit and matching target segment and the supplied footage must play in the rendered timeline, not merely appear as a still reference frame.

Step 3 — Decompose into a cross-modal EDL

Write project/plan.json. Every segment declares HOW it is produced (source) and WHERE it sits (layer):

  • source: edit (trim a real clip — needs input_id + in_sec/out_sec), generate (needs prompt + explicit media_kind; billable; video signs supported provider settings such as generation_duration_sec, ratio, resolution, generate_audio, and reference_image_urls; omit optional settings the provider should choose; for recurring subjects also set characters, refs, and variation_type), compose (designed HTML — needs a kind), provided (needs asset_id and explicit kind: video|image; still images never count as real motion/source footage).
  • layer: primary (the main timeline), overlay (sits over a primary via over: <segment id> — captions, lower-thirds, title cards), bg (behind).
  • role: MUST be exactly one of hook / body / proof / cta / transition — the schema rejects any other value (E_SEG_ROLE) and the plan fails validation. Narrative BEAT names from the arc ("payoff", "establishing", "climax", ...) are NOT roles: map a payoff / closing / CTA beat to cta, an establishing / evidence beat to proof. Front-load the hook.

Top-level references is mandatory for every edit/provided source and generation reference. Each entry uses {id,media_type,source,intent,intent_basis,roles,required,preserve,may_change,target_segment_ids,temporal_anchors?}. Omit the field when no references exist; never emit an empty array. Roles are exactly content|identity|composition|structure|style|motion|timing|audio. User requirements win; otherwise guide/inferred is safe. Preserve and may-change must not overlap.

When OVS intelligently selects or transforms content, add edit_strategy with {mode,objectives,decision_signals,preserve,may_change}. Use deterministic for evidence-driven cuts, semantic for AI pixel edits, and mixed for both. A semantic video edit stays an EDIT/AUTO workflow but is a billable source:"generate" video segment with operation:"edit" and at least one declared original in reference_video_paths or reference_video_urls. It requires semantic_model, a matching top-level edit reference and temporal anchor, and inclusion in the paid-generation count.

tracks is required even when the project has no audio or captions; use {} for the empty case. Tracks are separate from the visual timeline. For narration, run ovs speech-capabilities and sign its executable route/model/voice/format together with the BCP-47 video language and speed under tracks.narration.synthesis; raw voice is legacy recovery only. Timed lines are {text,start_sec,target_sec}; target_sec is duration, never end time, and windows must not overlap and each receives a produced_path. Music is path + ducking. Captions stay editable DATA under tracks.captions.lines. Put the exact number of generate segments in cost_estimate.billable_generations; a mismatch is a validation error because paid-generation confirmation reviews that count.

Fit narration in the plan before any TTS call: use natural cadence (about 2.2-2.7 English words/sec or 4-5 Chinese chars/sec), shorten over-budget lines here, and do not rely on repeated synthesis to discover timing.

Author plan.json in EXACTLY this shape (copy the field names — ovs plan validate rejects any other shape):

json
{
  "aspect": "9:16",
  "total_target_sec": 30,
  "language": "en",
  "delivery_promise": { "type": "hybrid", "source_required": true, "motion_min_ratio": 0.6 },
  "segments": [
    { "id": "s1_hook", "order": 1, "role": "hook", "layer": "primary", "source": "edit",
      "target_sec": 6, "spec": { "input_id": "clipA", "in_sec": 12, "out_sec": 18 } },
    { "id": "s2_body", "order": 2, "role": "body", "layer": "primary", "source": "compose",
      "target_sec": 8, "spec": { "kind": "stat-card" } },
    { "id": "s2_cap", "order": 3, "role": "body", "layer": "overlay", "over": "s2_body",
      "source": "compose", "target_sec": 3, "spec": { "kind": "lower-third" } }
  ],
  "references": [
    {
      "id": "clip-a-source", "media_type": "video", "source": "raw/clipA.mp4",
      "intent": "edit", "intent_basis": "user", "roles": ["content", "timing", "audio"],
      "required": true, "preserve": ["approved content", "audio sync"],
      "may_change": ["signed timeline cuts"], "target_segment_ids": ["s1_hook"],
      "temporal_anchors": [{ "source_start_sec": 12, "source_end_sec": 18, "target_segment_id": "s1_hook" }]
    }
  ],
  "tracks": {
    "narration": { "synthesis": { "route_ref": "openai-compatible", "voice": "nova", "model": "tts-1", "format": "mp3", "language": "en-US", "speed": 1 },
      "segments": [ { "text": "one line of narration", "start_sec": 0, "target_sec": 6 } ] },
    "music": { "path": "assets/bed.mp3", "duck": true },
    "captions": { "style": "bold-bottom", "lines": [ { "text": "one caption line", "start_sec": 0, "target_sec": 3 } ] }
  },
  "cost_estimate": { "billable_generations": 0 }
}

Field gotchas the validator enforces (these are the common breakers):

  • source is the production-method enum edit | generate | compose | provided — NOT a file path. The actual clip/asset goes in spec.input_id (edit) or spec.asset_id (provided).
  • Every segment needs order + layer + spec; use target_sec (not target_duration_sec/duration). At least one segment must be layer:"primary".
  • Every edit/provided source and generation reference needs a matching top-level reference declaration; spec.input_id does not express intent or preservation boundaries.
  • operation:"edit" requires reference video input plus edit_strategy.mode:"semantic"|"mixed" and may change only the declared axes.
  • tracks is a required object {narration, music, captions} — NOT an array or null. Use {} when no tracks are needed.
  • delivery_promise must MATCH this deliverable (Step 2) — do NOT copy the example's hybrid/source_required:true/0.6. A designed-HTML explainer is type:"compose_led", source_required:false, motion_min_ratio:0; set source_required:true ONLY when the user's real footage must star; for other promise types, motion_min_ratio is the real-motion floor you are actually committing to.

Plan to the craft bar (video-craft): a hook in the first seconds, one idea per beat, readable type in safe zones, ducked audio, the right aspect.

Show full SKILL.md (673 more words)Show less

Step 4 — Validate, then gate B

  1. ovs plan validate on project/plan.json. Fix EVERY error before going further — errors mean the plan cannot be executed or it breaks its own promise (e.g. source_required but no source segment). It also checks every generate video segment against the configured video.provider (MuAPI's Kling endpoints accept only 16:9/9:16/1:1 and 5 or 10 s); fix the plan or switch provider now rather than letting an approved generation fail. Reconsider warnings.
  2. ovs plan promise-check on the PLAN, before producing anything. It computes the planned motion ratio vs. the promise — a fail means the plan is already a slideshow / breaks its promise. Fixing the plan now is free; re-assembling later is not. Rebalance durations or convert a static beat to footage until it passes (gate D re-checks against the real cut).
  3. ovs plan summarize → present that timeline for production plan confirmation, including the exact narrator and generation settings. After the user's reply, run ovs gate transition; do not infer approval or request it again for an unchanged already-approved plan.

Director judgment (end-to-end planning)

The craft of weaving ONE good video across sources, on top of the shared craft (video-craft). This is where a multi-source plan becomes a video instead of a tour of clips:

  • Decide the spine before the sources. Write the beat arc (hook → gap → core → proof → payoff/CTA, video-craft §2) source-agnostic FIRST, then assign each beat its cheapest sufficient source. Letting the material on hand dictate the structure is how end-to-end videos turn into a disjointed reel.
  • Assign each beat to the source that earns it. Real footage (edit / provided) carries proof / authenticity / the actual product or result — make it the hero of a source_led piece, not a cameo. generate is a last resort for a beat you can neither film nor compose (an impossible / expensive establishing shot, missing b-roll) — it is billable and reads synthetic if overused. compose is the connective tissue — titles, stats, definitions, transitions, the CTA card — cheapest and crispest for anything textual.
  • Treat the promise as an editorial commitment, not a ratio to satisfy. source_led means the user's material genuinely stars (the hero beats + real screen time), not 6 s buried under composed cards. Set motion_min_ratio to the feel you are promising.
  • Pace the plan in target_sec to video-craft §3: front-load the first payoff, one idea per beat, don't plan three equal-length beats in a row.
  • Cost-aware craft. Reach ~90% of the result with zero billable generation — reuse the user's footage, compose instead of generate, pull b-roll from existing frames. Generation is the exception you justify, not the default.
  • Plan the moment, not the whole clip. Set each edit segment's in_sec/out_sec to the one ~3 s window that earns its slot (the cut craft itself is in stage-edit). Every beat must earn a purpose (establish / proof / reaction); a beat you can't justify shouldn't be in the plan.
  • Write each visual beat as a concrete photograph, not an emotion — subject, action, environment, lighting (the rule + examples are in video-craft §11). If you can't picture a specific frame from the spec, neither can the generator.

Rules

  • The plan is the source of production intent. Gate-B signing ignores known execution-only fields such as status/produced_path and provider catalog labels, while stable content/references/settings and unknown new fields remain approval-bearing.
  • Reference real input_ids from ingest.json; never cite a clip you have not probed.
  • Keep the billable count honest in cost_estimate — gate C depends on it.

Boundary / non-goals

This skill ingests and PLANS. It does not produce or assemble — that is stage-assemble, which walks the validated plan and delegates each segment to the compose / generate / edit lines.

Generate specs use a closed media-specific field set. Put edit strategy at the plan top level; do not invent nested provider options or aliases such as duration_sec or audio. Run the validator and use its allowed fields. Source audio kept by an edit needs no extra narration track. Plan narration for new explainers/promos without existing speech unless the user requests silent/music-only delivery; preserve existing or lip-synced speech instead of adding a second voice.

© Orkas-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills/stage-plan of Orkas-AI/Orkas-VideoStudio.

Open the folder on GitHubat commit c0c3c3e

Compare with similar skills

Stage Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stage Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stage Plan this skillOrkas-AI/Orkas-VideoStudio498—~4kAutomated safety check: PassMIT
Narrator AI CLINarratorAI-Studio/narrator-ai-cli-skill3k—~4.5kAutomated safety check: PassMIT
Ergo Remotion Videoitwanger/toBeBetterJavaer18k—~1.1kAutomated safety check: PassNone
Wedding Video Guided Wizardaaronyi97/wedding-video-guided-wizard306—~1kAutomated safety check: PassMIT
RunninghubHM-RunningHub/OpenClaw_RH_Skills141—~1.6kAutomated safety check: PassApache-2.0
AI Video Production Assistantwanghui2323/ai-video-maker101—~924Automated safety check: PassMIT

Similar skills

  • Narrator AI CLI

    NarratorAI-Studio/narrator-ai-cli-skill

    AI 电影/短剧解说视频自动生成(AI 解说大师 CLI Skill)。当用户需要创建电影解说视频、短剧解说、影视二创、AI 配音旁白视频、film commentary、video narration、drama dubbing、movie narration 时触发。内置电影素材库、BGM、多语种配音、解说模板。通过 narrator-ai-cli 命令行实现:搜片→选模板→选…

    3k GitHub stars~4.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Ergo Remotion Video

    itwanger/toBeBetterJavaer

    把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在…

    18k GitHub stars~1.1k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Wedding Video Guided Wizard

    aaronyi97/wedding-video-guided-wizard

    Guide a creator through a real couple's custom wedding video, from a shareable story intake card and Kimi writing pack through narration, external GPT image prompts, image-to-video packs, music and…

    306 GitHub stars~1k tokensUpdated 29 days ago
    Media & CreativeAuto-check passed
  • Runninghub

    HM-RunningHub/OpenClaw_RH_Skills

    Generate images, videos, audio, and 3D models via RunningHub API (420 endpoints) and run any RunningHub AI Application (custom ComfyUI workflow) by webappId.

    141 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Video Production Assistant

    wanghui2323/ai-video-maker

    Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice.

    101 GitHub stars~924 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Media Gen

    clacky-ai/openclacky

    Generate or edit images, videos, or audio in the current task.

    1.2k GitHub stars~7.3k tokensUpdated today
    Media & CreativeAuto-check passed

More from Orkas-AI/Orkas-VideoStudio

All 14 skills in this repo
  • Gate Control

    Orkas-AI/Orkas-VideoStudio

    Canonical VideoStudio review authorization and state-transition policy.

    498 GitHub stars~2.7k tokensUpdated 17 days ago
    Auto-check passed
  • Stage Compose

    Orkas-AI/Orkas-VideoStudio

    Authoring knowledge for Orkas/OVS HTML video compositions -- write an index.html, drive animation from a paused timeline, declare canvas + duration, then run the VideoStudio draft gate to render an…

    498 GitHub stars~4.2k tokensUpdated 17 days ago
    Auto-check passed
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    498 GitHub stars~2.4k tokensUpdated 17 days ago
    Auto-check passed
  • Stage Consistency

    Orkas-AI/Orkas-VideoStudio

    Multi-shot narrative & character consistency — a character bible with a locked front-portrait anchor, view-matched reference selection, recent-frame carry-forward, Cameo (a user photo as the lead)…

    498 GitHub stars~1.8k tokensUpdated 17 days ago
    Auto-check passed
  • Frontend Design

    Orkas-AI/Orkas-VideoStudio

    Aesthetic direction for OrkasVideoStudio HTML and motion-graphics compositions.

    498 GitHub stars~4.3k tokensUpdated 17 days ago
    Auto-check passed
  • Video Craft

    Orkas-AI/Orkas-VideoStudio

    The craft standard that makes a video GOOD, not just rendered — hooks, story pacing, visual hierarchy, type & safe zones, motion/easing, captions, audio mix, platform conventions, shot language…

    498 GitHub stars~3.3k tokensUpdated 17 days ago
    Auto-check passed

Questions about Stage Plan

What does Stage Plan do?

The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…. Stage Plan is an agent skill from Orkas-AI/Orkas-VideoStudio.json: edit/generate/compose/provided segments + narration/music/caption tracks + a delivery promise), validate it with ovs plan validate.

When should I use Stage Plan?

Stage Plan fits situations like: the deliverable spans more than one axis (AUTO line); do NOT trigger for a pure single-axis job (route to that line).

How do I install Stage Plan in Claude Code?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-plan -a claude-code`. Or copy the skill folder (packages/skills/stage-plan in Orkas-AI/Orkas-VideoStudio) into .claude/skills/stage-plan in your project. Claude Code loads it when a task matches its description.

How do I install Stage Plan in Codex?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-plan -a codex`. Or copy the skill folder (packages/skills/stage-plan in Orkas-AI/Orkas-VideoStudio) into .agents/skills/stage-plan in your project. Codex loads it when a task matches its description.

Can I use Stage Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stage-plan, .gemini/skills/stage-plan, .github/skills/stage-plan and .opencode/skills/stage-plan in your project.

What does Stage Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Stage Plan is instructions for the agent only.

Does Stage Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stage Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stage Plan use?

Stage Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stage Plan use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stage Plan?

Skills that share tags, products or a category with Stage Plan: Narrator AI CLI (NarratorAI-Studio/narrator-ai-cli-skill, 3k stars), Ergo Remotion Video (itwanger/toBeBetterJavaer, 18k stars), Wedding Video Guided Wizard (aaronyi97/wedding-video-guided-wizard, 306 stars) and Runninghub (HM-RunningHub/OpenClaw_RH_Skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stage Plan?

Orkas-AI (a GitHub user) maintains it in Orkas-AI/Orkas-VideoStudio, which has 498 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 22, 2026.

Source: Orkas-AI/Orkas-VideoStudio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.