Agent skill

Stage Generate

by Orkas-AI in Orkas-AI/Orkas-VideoStudio

AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO.

MITAuto-check passedMedia & Creative

Install Stage Generate

skills CLI
$ npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-generate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orkas-AI/Orkas-VideoStudio stage-generate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orkas-AI/Orkas-VideoStudio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/stage-generate .claude/skills/stage-generate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stage-generate
GitHub stars
499
Token cost
~1.7k tokens
SKILL.md length
918 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO.

  • Works in 3 steps: Character still: generate one image of… → Bring it to life: generate a video from… → Polish: add captions / a lower-third / a…
  • Tasks that involve AI video generation
  • SKILL.md covers Pattern A — talking-head /…, Pattern B — cinematic / AI…, Consistency (basic — deep… and Director judgment (generation…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stage Generate is an agent skill from Orkas-AI/Orkas-VideoStudio. AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO. Generate per shot with ovs video, preserve signed reference/settings, then assemble. For recurring characters also read stage-consistency.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation and Text to speech and voice. The repository describes itself as: Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file. The licence is MIT.

When your agent uses it

  • Tasks that involve AI video generation
  • Tasks that involve Text to speech and voice

Example prompts

  • “/stage-generate”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Character still: generate one image of the presenter / avatar with the intended look. Keep this reference image and reuse it for every…
  2. Bring it to life: generate a video from that image (image-to-video). When the provider returns speech + built-in audio, that audio is the…
  3. Polish: add captions / a lower-third / a hook by authoring a small composition (stage-compose) and overlaying it onto the clip…

What it can do on your machine

Read from SKILL.md and the folder at commit c0c3c3e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stage Generate loads about 1.7k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 918 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orkas-AI/Orkas-VideoStudio at commit c0c3c3e, republished under its MIT licence (© Orkas-AI). 918 words, ~1,707 tokens.

Download SKILL.mdSave it as .claude/skills/stage-generate/SKILL.md (or your agent's skills folder).
name
stage-generate
description
AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO. Generate per shot with `ovs video`, preserve signed reference/settings, then assemble. For recurring characters also read stage-consistency.

stage-generate

How to produce AI-generated footage and execute a bounded semantic video edit already planned by EDIT/AUTO. Designed HTML remains composition work; deterministic cutting remains stage-edit work. The generation tools are ovs image, ovs video, and ovs speak, followed by ovs edit assembly. Every billable call belongs to a signed generate segment.

Pattern A — talking-head / spokesperson

  1. Character still: generate one image of the presenter / avatar with the intended look. Keep this reference image and reuse it for every shot of the same character.
  2. Bring it to life: generate a video from that image (image-to-video). When the provider returns speech + built-in audio, that audio is the deliverable voice — it is lip-synced to the mouth in the clip — so keep it and do NOT synthesize a separate narration. Only when the clip comes back silent do you synthesize the narration (ovs speak) and add it as the audio track. Synthesizing a fresh TTS track over a clip that already speaks is the #1 talking-head defect: the new audio has different wording/timing/length, so the voice no longer matches the lips.
  3. Polish: add captions / a lower-third / a hook by authoring a small composition (stage-compose) and overlaying it onto the clip — visual-only. Preserve the clip's own (lip-synced) audio through assembly; a captions composition must not carry a narration <audio> track that would replace the clip's voice.

Pattern B — cinematic / AI b-roll montage

  1. Storyboard the shots (each: prompt, camera motion, duration).
  2. Generate each shot (one ovs video call per shot; reuse a shared reference image / consistent style prompt for visual continuity). Pass the Gate-C-approved --image-urls, --ratio, --duration, --resolution, and --generate-audio values exactly; do not let provider defaults silently replace the signed plan.
  3. Assemble: concatenate the shots in order, add transitions, and overlay a title / captions from a composition.

Consistency (basic — deep consistency is a separate skill)

  • Within a shot: drive the clip from a reference image (image-to-video) to lock the subject.
  • Across shots: reuse the same reference image / style prompt.
  • Full multi-shot character consistency, Cameo (upload-a-photo-as-the-lead), and long-narrative planning live in stage-consistency.

Director judgment (generation line)

Craft calls specific to AI-generated footage, on top of the shared craft reference (video-craft):

Talking-head / spokesperson

  • Understand what's said before placing overlays; time graphics to the spoken words.
  • 3–6 overlays/min, varied types; keep them in speaker-safe zones — never over the face.
  • Cut silences and filler; for vertical, keep subtitles low so they don't cover the face.

Cinematic / AI b-roll

  • Open on a hero frame; keep a small transition palette (cut / fade-to-black / slow dissolve / restrained push-in).
  • Protect earned moments — don't over-cut a held look or a deliberate silence.
  • Design each shot first-frame → last-frame and let audio dynamics carry momentum (shot language: video-craft §10; identity across shots: stage-consistency).
  • Frames are static snapshots, never an action in progress (video-craft §10); in motion / last-frame text, name characters by visible features, not names (the model conditions on pixels, not labels).
  • Spend keyframes by how much the shot changes (cost gate). If a shot's start and end look nearly the same — a talking head, a small pose/expression change, a gentle pan — it needs only ONE keyframe and motion fills the rest (variation_type small). Only a shot that ends somewhere visually different — a new subject enters, a wide→close transition, a big camera move — needs TWO keyframes for the model to interpolate (medium / large). Don't pay to generate an end-frame you don't need.
Show full SKILL.md (357 more words)Show less

Default scope caps (cost control — do not exceed without explicit user request)

Generated clips and images are billable hosted calls, so bound the run by default:

  • Shots / clips: ≤ 6 per video.
  • Characters: ≤ 3 per video.
  • One aspect ratio per run.

If the brief seems to need more (a long story, many scenes, many characters), DO NOT silently fan out — state the larger count + the rough number of billable generations in the pre-generation confirmation and let the user opt in first. Treat anything above these caps as requiring explicit confirmation.

Rules

  • Cost/time discipline: every generated clip is a hosted, billable, multi-second call. State the exact shot/character count in the pre-generation confirmation; never start generating before the user has approved the count.
  • Audio (talking-head): a generated talking-head clip's built-in audio is lip-synced to its mouth. Treat it as the final voice: keep it through assembly and finalize. NEVER add or mux a separately-synthesized narration over a clip that already speaks — it desyncs from the lips. Synthesize narration ONLY for a silent clip, or for b-roll / off-screen voiceover where no mouth is visible.
  • Brief drives params: record and pass the exact reference images, aspect ratio, per-shot duration, resolution, and audio choice to each generation call.
  • One authorization, one output set: every billable segment in the approved plan must have a matching result. A retry that creates a new hosted task or output set needs fresh Gate-C authority; recovery may inspect or download an already-authorized result without asking again. Use ovs gate transition instead of inventing line-local confirmation rules.
  • Semantic edit boundary: operation:"edit" requires a top-level video reference with intent:"edit", target/temporal anchors, plus edit_strategy.mode:"semantic"|"mixed". The prompt may change only may_change; protected identity/content/motion/timing/audio axes remain explicit. Never turn an edit into unconstrained regeneration.
  • Partial batches: preserve completed sibling segments. A revision or failure invalidates only that shot plus downstream assembly evidence; never redispatch completed shots merely to rebuild the montage.

Boundary / non-goals

Generation-primary work uses this skill directly. EDIT/AUTO may invoke it only as executor for an approved semantic-edit segment; source analysis, preservation boundaries, and final assembly stay with the edit plan. Designed HTML goes to stage-compose and deterministic cuts/joins to stage-edit.

© Orkas-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills/stage-generate of Orkas-AI/Orkas-VideoStudio.

Open the folder on GitHubat commit c0c3c3e

Compare with similar skills

Stage Generate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stage Generate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stage Generate this skillOrkas-AI/Orkas-VideoStudio499—~1.7kAutomated safety check: PassMIT
Narrator AI CLINarratorAI-Studio/narrator-ai-cli-skill3k—~4.5kAutomated safety check: PassMIT
Ergo Remotion Videoitwanger/toBeBetterJavaer18k—~1.1kAutomated safety check: PassNone
Wedding Video Guided Wizardaaronyi97/wedding-video-guided-wizard310—~1kAutomated safety check: PassMIT
RunninghubHM-RunningHub/OpenClaw_RH_Skills142—~1.6kAutomated safety check: PassApache-2.0
AI Video Production Assistantwanghui2323/ai-video-maker101—~924Automated safety check: PassMIT

Similar skills

  • Narrator AI CLI

    NarratorAI-Studio/narrator-ai-cli-skill

    AI 电影/短剧解说视频自动生成(AI 解说大师 CLI Skill)。当用户需要创建电影解说视频、短剧解说、影视二创、AI 配音旁白视频、film commentary、video narration、drama dubbing、movie narration 时触发。内置电影素材库、BGM、多语种配音、解说模板。通过 narrator-ai-cli 命令行实现:搜片→选模板→选…

    3k GitHub stars~4.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Ergo Remotion Video

    itwanger/toBeBetterJavaer

    把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在…

    18k GitHub stars~1.1k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Wedding Video Guided Wizard

    aaronyi97/wedding-video-guided-wizard

    Guide a creator through a real couple's custom wedding video, from a shareable story intake card and Kimi writing pack through narration, external GPT image prompts, image-to-video packs, music and…

    310 GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Runninghub

    HM-RunningHub/OpenClaw_RH_Skills

    Generate images, videos, audio, and 3D models via RunningHub API (420 endpoints) and run any RunningHub AI Application (custom ComfyUI workflow) by webappId.

    142 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Video Production Assistant

    wanghui2323/ai-video-maker

    Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice.

    101 GitHub stars~924 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Media Gen

    clacky-ai/openclacky

    Generate or edit images, videos, or audio in the current task.

    1.2k GitHub stars~7.5k tokensUpdated today
    Media & CreativeAuto-check passed

More from Orkas-AI/Orkas-VideoStudio

All 14 skills in this repo
  • Gate Control

    Orkas-AI/Orkas-VideoStudio

    Canonical VideoStudio review authorization and state-transition policy.

    499 GitHub stars~2.7k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Compose

    Orkas-AI/Orkas-VideoStudio

    Authoring knowledge for Orkas/OVS HTML video compositions -- write an index.html, drive animation from a paused timeline, declare canvas + duration, then run the VideoStudio draft gate to render an…

    499 GitHub stars~4.2k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    499 GitHub stars~2.4k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Consistency

    Orkas-AI/Orkas-VideoStudio

    Multi-shot narrative & character consistency — a character bible with a locked front-portrait anchor, view-matched reference selection, recent-frame carry-forward, Cameo (a user photo as the lead)…

    499 GitHub stars~1.8k tokensUpdated 19 days ago
    Auto-check passed
  • Stage Plan

    Orkas-AI/Orkas-VideoStudio

    The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

    499 GitHub stars~4k tokensUpdated 19 days ago
    Auto-check passed
  • Frontend Design

    Orkas-AI/Orkas-VideoStudio

    Aesthetic direction for OrkasVideoStudio HTML and motion-graphics compositions.

    499 GitHub stars~4.3k tokensUpdated 19 days ago
    Auto-check passed

Questions about Stage Generate

What does Stage Generate do?

AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO. Stage Generate is an agent skill from Orkas-AI/Orkas-VideoStudio. AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO.

When should I use Stage Generate?

Stage Generate fits situations like: tasks that involve AI video generation; tasks that involve Text to speech and voice.

How do I install Stage Generate in Claude Code?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-generate -a claude-code`. Or copy the skill folder (packages/skills/stage-generate in Orkas-AI/Orkas-VideoStudio) into .claude/skills/stage-generate in your project. Claude Code loads it when a task matches its description.

How do I install Stage Generate in Codex?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-generate -a codex`. Or copy the skill folder (packages/skills/stage-generate in Orkas-AI/Orkas-VideoStudio) into .agents/skills/stage-generate in your project. Codex loads it when a task matches its description.

Can I use Stage Generate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-generate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stage-generate, .gemini/skills/stage-generate, .github/skills/stage-generate and .opencode/skills/stage-generate in your project.

What does Stage Generate need to run?

SKILL.md names no scripts, command-line tools or credentials: Stage Generate is instructions for the agent only.

Does Stage Generate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stage Generate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stage Generate use?

Stage Generate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stage Generate use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stage Generate?

Skills that share tags, products or a category with Stage Generate: Narrator AI CLI (NarratorAI-Studio/narrator-ai-cli-skill, 3k stars), Ergo Remotion Video (itwanger/toBeBetterJavaer, 18k stars), Wedding Video Guided Wizard (aaronyi97/wedding-video-guided-wizard, 310 stars) and Runninghub (HM-RunningHub/OpenClaw_RH_Skills, 142 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stage Generate?

Orkas-AI (a GitHub user) maintains it in Orkas-AI/Orkas-VideoStudio, which has 499 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 22, 2026.

Source: Orkas-AI/Orkas-VideoStudio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.