Agent skill

Stage Edit

by Orkas-AI in Orkas-AI/Orkas-VideoStudio

Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

MITAuto-check passedMedia & Creative

Install Stage Edit

skills CLI
$ npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-edit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orkas-AI/Orkas-VideoStudio stage-edit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orkas-AI/Orkas-VideoStudio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/stage-edit .claude/skills/stage-edit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stage-edit
GitHub stars
499
Token cost
~2.4k tokens
SKILL.md length
1,265 words
Files
2 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

  • Works in 4 steps: Ingest — always probe first. For every… → Plan — write an edit_decisions timeline.… → Execute in order. → …
  • Local content changes
  • SKILL.md covers Intelligent Edit Contract, The deterministic editing loop, Director judgment (editing line) and Rules, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stage Edit is an agent skill from Orkas-AI/Orkas-VideoStudio. Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained semantic AI edit. Trigger for repurpose, montage, cleanup, localization, narration, or local content changes.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/transcript-and-screen-grounded-editing.md`).

It sits in Media & Creative, covering Text to speech and voice, Internationalization and Content repurposing. It works with FFmpeg. The repository describes itself as: Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file. The licence is MIT.

When your agent uses it

  • Local content changes
  • Tasks that involve Text to speech and voice
  • Tasks that involve Internationalization

Example prompts

  • “/stage-edit”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Ingest — always probe first. For every input clip, read its metadata (duration, resolution, fps, codecs) with ovs edit probe. Never plan a…
  2. Plan — write an edit_decisions timeline. From the user's intent + the probe results, decide the exact segments and order, and write them…
  3. Execute in order.
  4. Publish the final file.

What it can do on your machine

Read from SKILL.md and the folder at commit c0c3c3e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stage Edit loads about 2.4k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 1,265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orkas-AI/Orkas-VideoStudio at commit c0c3c3e, republished under its MIT licence (© Orkas-AI). 1,265 words, ~2,447 tokens.

Download SKILL.mdSave it as .claude/skills/stage-edit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
stage-edit
description
Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained semantic AI edit. Trigger for repurpose, montage, cleanup, localization, narration, or local content changes.

stage-edit

How to intelligently edit real user-supplied footage while keeping source and decisions auditable. Deterministic operations handle cuts, joins, captions, overlays, reframes, and audio. When the request changes pixels semantically—remove an object, alter a background, relight, or make another content-aware local change—keep the EDIT route and execute only that bounded segment as a signed ovs video --operation edit call after paid-generation approval.

Subtitle safety hard rule. Burn captions only through ovs edit burnsubs; do not hand-write ffmpeg subtitle, drawtext, or PNG-overlay fallback commands. If burnsubs fails because the runtime ffmpeg lacks subtitle filter support, stop and report the blocker instead of improvising a custom ffmpeg graph.

If the task is to FIND / SELECT / REDUCE / CLEAN rather than run a known timecode edit — remove dead air, drop fillers, pick highlights, cut a long recording down — read stage-decide first: it covers understanding the footage and producing an evidence-bearing rough cut (the deterministic auto-cuts ovs edit trim-silence / remove-fillers, plus ovs scenes / ovs quality / ovs plan rank-takes). This skill is for executing cuts you have already chosen.

Two assembly paths — pick by whether the result needs to stay re-editable:

  • Plan-backed (anything the user may later adjust: narration, multi-shot, segmented edits). Author project/plan.json (the segments EDL — see stage-plan) carrying ONLY the operations the user asked for (the deltas) — everything else is the source, passed through untouched. Keep each editable concern SEPARATE in the plan: each narration line in tracks.narration.segments with its own produced_path, each caption in tracks.captions.lines as data, each segment carrying status/produced_path. Assemble with ovs edit (trim → concat → mix → burnsubs). Because plan.json holds every piece separately, a later "fix one caption / re-voice one line" is a one-entry edit + one re-render — do NOT pre-bake (e.g. one big narration file), which destroys that separability.
  • One-shot deterministic (a plain trim or concat the user just wants done). Use ovs edit directly; write no plan.json.

Intelligent Edit Contract

Write project/plan.json#edit_strategy whenever OVS decides what to change rather than merely executing user-provided timecodes. Declare its mode (deterministic, semantic, or mixed), exact objectives, only the evidence signals actually used, and non-empty, non-overlapping preserve/may_change.

Declare every source/reference in top-level references with media type, reproduce/edit/guide intent and basis, roles, required state, preservation boundary, target segments, and video temporal anchors. A semantic video edit is source:"generate", media_kind:"video", operation:"edit" with its original in reference_video_paths or reference_video_urls; it remains owned by EDIT and counts as billable.

For a plan-backed follow-up, invalidate only the changed entry and its derived assembly. Preserve source probes, transcripts, inspected-frame evidence, unaffected cuts, narration, and sibling outputs. A content-identical source moved to a new path is an implementation locator change, not new creative intent.

The deterministic editing loop

  1. Ingest — always probe first. For every input clip, read its metadata (duration, resolution, fps, codecs) with ovs edit probe. Never plan a cut blind; a trim past the real duration produces an empty or broken clip.
  2. Plan — write an edit_decisions timeline. From the user's intent + the probe results, decide the exact segments and order, and write them to project/edit_plan.json so the plan is inspectable and re-runnable. Shape:
    json
    {
      "segments": [
        { "input": "raw/clipA.mp4", "start": 12.0, "duration": 8.0 },
        { "input": "raw/clipB.mp4", "start": 0.0,  "duration": 5.5 }
      ],
      "subtitles": "raw/captions.srt",
      "overlay": { "media": "assets/logo.png", "x": 40, "y": 40 }
    }
    Every start/duration must be inside the probed duration of its input.
  3. Execute in order.
    • ovs edit trim each segment to its own file (project/cuts/seg-1.mp4, ...).
    • ovs edit concat the cut files (in plan order) into one (project/render/edited.mp4).
    • If subtitles: ovs edit burnsubs the .srt/.ass onto the concatenated video.
    • If an overlay (logo / lower-third image / PiP): ovs edit overlay it at the planned position.
  4. Publish the final file.

Read transcript and screen-grounded editing only for topic-based selection, captions, localization or narration added to existing footage. It owns evidence-first analysis, per-line speech and extracted-frame inspection.

Show full SKILL.md (665 more words)Show less

Director judgment (editing line)

Craft calls per repurpose/montage line, on top of the shared craft reference (video-craft).

Cut craft (every editing job — this is the canonical set; the assembly line references it). On top of video-craft (pacing §3, transitions §5, audio §7):

  • Cut the moment, not the clip. A 12 s clip usually holds one ~3 s moment that earns its slot — trim to that window. End the cut on a held look, not on the action moving off; leave a few frames of handle at each end so a dissolve doesn't clip the moment, and never freeze on a static last frame (reads as a glitch).
  • A restrained transition vocabulary for cut-driven pieces: ≤ 4 types across the whole piece — hard cut (default, most invisible), dissolve (emotional siblings / time passage), fade-to-black (act breaks), fade bookends. In a documentary/montage register, wipes / push-slide / zoom-blur / glitch read as social-media language — avoid (this is stricter than the explainer norm in video-craft §5, where a wipe can mark a step).
  • Bridge the hardest cuts with sound — carry the outgoing clip's ambient under the incoming for ~0.5–1.5 s (L-cut), or start the next audio early (J-cut); audio continuity hides a visual seam. Plus the one held silence from video-craft §7.
  • Adjacent-diversity + a reason per cut. Don't place the same subject at the same shot size, or the same palette, back-to-back — break the pattern at least every ~4 cuts. If you can't write a one-line reason for a cut, it's arbitrary — reconsider it.

Per repurpose/montage line:

  • Social clip / clip-factory — per clip = hook (0–2 s) → sustain → clean outro; optimize the first 2 frames; start on motion/face/result; lock a batch style (caption / hook position / watermark) so a series feels cohesive; don't crowd frame 1 with hook + caption + watermark + lower-third at once.
  • Podcast-repurpose — audio is the hero; pick quotable moments; speaker video if it exists, else a simple audiogram / quote card; keep the visual system simple and repeatable; preserve attribution + CTA.
  • Screen-demo — zoom only for legibility/orientation, steady while the viewer reads; reset to wide context between phases; ≤ 2 attention cues at once; label sped-up sections; keep UI text sharp (higher bitrate), don't force an unreadable vertical crop.
  • Localization — treat each language as its own deliverable; dubbed audio won't match source timing, so plan holds to flex; re-render or cover any baked-in text per language; subtitle line lengths differ by language; lip-sync only where a close-up mouth mismatch would distract.
  • Documentary-montage — concrete sensory shot descriptions, not abstract themes; one grade/LUT across all clips is what unifies mixed sources; budget 2–3 hero slots longer holds; a music bed + an end-tag.
  • Before publishing, normalize the mix against the targets in video-craft §7 (~−14 LUFS integrated, true-peak ≤ ~−1 dBTP) with ovs edit normalize-loudness; use ovs edit loudness for diagnosis.

Rules

  • Timecodes come from the user, from probe, from a transcript, or from inspected extracted frames — never guessed. If the target moment cannot be located from evidence, ask the user for the timestamp.
  • Layer composition over footage when the brief needs designed elements (animated lower-thirds, kinetic captions, hooks): produce those with stage-compose as an overlay/element and ovs edit overlay them, rather than trying to draw them in ffmpeg.
  • One output file at the end; intermediate cuts live under project/cuts/ and are not the deliverable.

Boundary / non-goals

This skill owns the EDIT workflow. It executes deterministic EDL operations directly and delegates only bounded semantic pixel changes to the signed video-edit provider path. It does not author HTML compositions.

Before showing any plan-backed final, regardless of assembly route, run ovs plan promise-check project/plan.json --probe-produced --video project/render/video.mp4. Repair duration/aspect mismatch, overlapping/truncated narration and unverifiable lines. Explain caption warnings from the actual container/sidecar check; burned captions need visual evidence. Planned duration and a successful mix alone do not prove delivery.

For narration, preserve signed line windows; target_sec is duration. Measure each produced file and voice span. Keep intentional silence, shorten overlong text within authorized scope, and never slide later lines to hide an overlap. A failed/unknown provider request does not authorize automatic repeat billing.

© Orkas-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in packages/skills/stage-edit of Orkas-AI/Orkas-VideoStudio.

  • SKILL.md
  • references/transcript-and-screen-grounded-editing.md

Open the folder on GitHubat commit c0c3c3e

Compare with similar skills

Stage Edit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stage Edit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stage Edit this skillOrkas-AI/Orkas-VideoStudio499—~2.4kAutomated safety check: PassMIT
Media ProductionWrongStack/WrongStack371—~1kAutomated safety check: PassMIT
AutoshortsUpload-Post/skill-autoshorts151—~5.3kAutomated safety check: NotesMIT
Proof Videoopenclaw/openclaw392k—~2.4kAutomated safety check: PassMIT
Video Productionffroliva/gflow-cli269—~7kAutomated safety check: PassMIT
Paw Cra Agent Video Producerpawbytes/skill-suites113—~2.3kAutomated safety check: PassMIT

Similar skills

  • Media Production

    WrongStack/WrongStack

    Create and process finished videos with Remotion, Motion Canvas, Manim, FFmpeg or an available AI video provider.

    371 GitHub stars~1k tokensUpdated today
    Media & CreativeAuto-check passed
  • Autoshorts

    Upload-Post/skill-autoshorts

    Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…

    151 GitHub stars~5.3k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Proof Video

    openclaw/openclaw

    Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

    392k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Production

    ffroliva/gflow-cli

    A skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…

    269 GitHub stars~7k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Paw Cra Agent Video Producer

    pawbytes/skill-suites

    Video production specialist for short-form, long-form, episodic, and motion graphics video.

    113 GitHub stars~2.3k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed

More from Orkas-AI/Orkas-VideoStudio

All 14 skills in this repo
  • Gate Control

    Orkas-AI/Orkas-VideoStudio

    Canonical VideoStudio review authorization and state-transition policy.

    499 GitHub stars~2.7k tokensUpdated 18 days ago
    Auto-check passed
  • Stage Compose

    Orkas-AI/Orkas-VideoStudio

    Authoring knowledge for Orkas/OVS HTML video compositions -- write an index.html, drive animation from a paused timeline, declare canvas + duration, then run the VideoStudio draft gate to render an…

    499 GitHub stars~4.2k tokensUpdated 18 days ago
    Auto-check passed
  • Stage Consistency

    Orkas-AI/Orkas-VideoStudio

    Multi-shot narrative & character consistency — a character bible with a locked front-portrait anchor, view-matched reference selection, recent-frame carry-forward, Cameo (a user photo as the lead)…

    499 GitHub stars~1.8k tokensUpdated 18 days ago
    Auto-check passed
  • Stage Plan

    Orkas-AI/Orkas-VideoStudio

    The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided…

    499 GitHub stars~4k tokensUpdated 18 days ago
    Auto-check passed
  • Frontend Design

    Orkas-AI/Orkas-VideoStudio

    Aesthetic direction for OrkasVideoStudio HTML and motion-graphics compositions.

    499 GitHub stars~4.3k tokensUpdated 18 days ago
    Auto-check passed
  • Video Craft

    Orkas-AI/Orkas-VideoStudio

    The craft standard that makes a video GOOD, not just rendered — hooks, story pacing, visual hierarchy, type & safe zones, motion/easing, captions, audio mix, platform conventions, shot language…

    499 GitHub stars~3.3k tokensUpdated 18 days ago
    Auto-check passed

Works with

Questions about Stage Edit

What does Stage Edit do?

Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…. Stage Edit is an agent skill from Orkas-AI/Orkas-VideoStudio. Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained semantic AI edit.

When should I use Stage Edit?

Stage Edit fits situations like: local content changes; tasks that involve Text to speech and voice; tasks that involve Internationalization.

How do I install Stage Edit in Claude Code?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-edit -a claude-code`. Or copy the skill folder (packages/skills/stage-edit in Orkas-AI/Orkas-VideoStudio) into .claude/skills/stage-edit in your project. Claude Code loads it when a task matches its description.

How do I install Stage Edit in Codex?

Run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-edit -a codex`. Or copy the skill folder (packages/skills/stage-edit in Orkas-AI/Orkas-VideoStudio) into .agents/skills/stage-edit in your project. Codex loads it when a task matches its description.

Can I use Stage Edit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orkas-AI/Orkas-VideoStudio --skill stage-edit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stage-edit, .gemini/skills/stage-edit, .github/skills/stage-edit and .opencode/skills/stage-edit in your project.

What does Stage Edit need to run?

SKILL.md names no scripts, command-line tools or credentials: Stage Edit is instructions for the agent only.

Does Stage Edit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stage Edit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stage Edit use?

Stage Edit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stage Edit use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Stage Edit?

Skills that share tags, products or a category with Stage Edit: Media Production (WrongStack/WrongStack, 371 stars), Autoshorts (Upload-Post/skill-autoshorts, 151 stars), Proof Video (openclaw/openclaw, 392k stars) and Video Production (ffroliva/gflow-cli, 269 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stage Edit?

Orkas-AI (a GitHub user) maintains it in Orkas-AI/Orkas-VideoStudio, which has 499 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 22, 2026.

Source: Orkas-AI/Orkas-VideoStudio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.