Agent skill

Open Video

by agent-next in agent-next/video-agent

Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).

Apache-2.0Auto-check passedMedia & Creative

Install Open Video

skills CLI
$ npx skills add agent-next/video-agent --skill open-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agent-next/video-agent open-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill/open-video .claude/skills/open-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
open-video
GitHub stars
120
Token cost
~3.2k tokens
SKILL.md length
1,338 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).

  • Works in 6 steps: What open-video does → Agentic procedure — run these steps in… → Key commands → …
  • The user wants to turn a concept
  • SKILL.md covers 1. What open-video does, 2. Agentic procedure — run…, 3. Key commands and 4. Constraints (H3 baseline —…, plus 2 more sections
  • Calls python and curl; needs OPEN_VIDEO_VLM_KEY

What it does

Open Video is an agent skill from agent-next/video-agent. Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Use when the user wants to turn a concept, script, or reference image into a finished video or multi-shot film — single shots, image-to-video, first-last-frame interpolation, reference-video/audio styling, or stitched long films beyond the 15s model ceiling. Covers prompt crafting, hard validation, ComfyUI-driven generation, vision judging, refine loop, and ffmpeg stitching.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation, Diffusion and image models and Video production. It works with ComfyUI, MiniMax and FFmpeg. The repository describes itself as: Open-source video generation — Ollama for MiniMax H3. Local director on ComfyUI. The licence is Apache-2.0.

When your agent uses it

  • The user wants to turn a concept
  • Reference image into a finished video
  • Multi-shot film — single shots
  • First-last-frame interpolation

Example prompts

  • “/open-video”

Requirements

  • Python 3
  • A credential in OPEN_VIDEO_VLM_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. What open-video does
  2. Agentic procedure — run these steps in order
  3. Key commands
  4. Constraints (H3 baseline — hard, evidence-based)
  5. Docs — read these for depth (do not guess from memory)
  6. Adding a model backend (Wan2.2, LTX, …)

What it can do on your machine

Read from SKILL.md and the folder at commit 12a10e9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPEN_VIDEO_VLM_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Open Video loads about 3.2k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,338 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agent-next/video-agent at commit 12a10e9, republished under its Apache-2.0 licence (© agent-next). 1,338 words, ~3,165 tokens.

Download SKILL.mdSave it as .claude/skills/open-video/SKILL.md (or your agent's skills folder).
name
open-video
description
Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Use when the user wants to turn a concept, script, or reference image into a finished video or multi-shot film — single shots, image-to-video, first-last-frame interpolation, reference-video/audio styling, or stitched long films beyond the 15s model ceiling. Covers prompt crafting, hard validation, ComfyUI-driven generation, vision judging, refine loop, and ffmpeg stitching.

open-video — autonomous director skill

v0.0.1 default for high-quality H3 clips is skill/h3-video
(Ollama for H3 + harness). Use this skill when the user wants multi-shot / long-film director behavior (plan → judge → stitch) beyond a single strong clip.

1. What open-video does

open-video is the autonomous director layer on top of open video models. It turns a concept into a finished film by running the loop no open engine ships natively:

plan → craft → validate → generate → judge → refine → stitch → deliver

  • Model-agnostic core (core/): planner (coherence bible), crafter, validator, judge-loop, stitcher, selector, ref-pack builder.
  • Pluggable backends (backends/<model>/): baseline = MiniMax H3 (#1 open model, Arena parity with closed, native stereo audio). Future: Wan2.2 (physics), LTX-2.3 (speed).
  • Engine adapter (engines/comfyui/): drives ComfyUI via its HTTP API. open-video is the brain; ComfyUI is the hands.

It is NOT a video engine (ComfyUI is the engine) and NOT a model (H3/Wan/LTX are backends). It is the agent brain ComfyUI lacks: judge→refine + multi-shot stitch + coherence planning as they land.

v0.0.1 honesty: prefer skill/h3-video for reliable high-quality single clips. Multi-minute film is the flagship design. The judge is REAL when env-wired: set OPEN_VIDEO_VLM_URL + OPEN_VIDEO_VLM_MODEL (+ OPEN_VIDEO_VLM_KEY) to any OpenAI-compatible vision endpoint and every shot is scored + diagnosed automatically; with the env unset it is an honest PASS stub — then manual frame review is mandatory. Single open models still cap ~15s/shot; longer output needs multi-shot orchestration (partial).

2. Agentic procedure — run these steps in order

Step 1 — Understand the request. Classify the job: single shot vs multi-shot film; target duration; did the user supply reference image(s) / video / audio; desired aspect; quality bar. If the request is ambiguous and the target is a film >30s, ask ONE focused question (subject + mood + length). Do not generate a long film on guesswork.

Step 2 — Pick the mode. Mode is auto-derived from inputs (see backends/h3/backend.py and scripts/validate_prompt.py detect_mode):

  • T2V / T2VA — text only, no image. Native audio + video.
  • I2V / I2VA — one image (first frame). Instruction line: "For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced."
  • FL2VA — two images (first + last frame). Continuous interpolation. This is the multi-shot chain mode: previous shot's last frame → next shot's first frame for continuous handoff.
  • L2VA — one image referenced at the final timestamp.
  • R2V / Ref2VA — reference video/audio for identity / style / motion / voice. Tag refs <Picture 1>, <Video 1>, <Audio 1> and explicitly assign each a role.

Step 3 — Craft the 3-field prompt (per backends/h3/PROMPT_GRAMMAR.md). Exact structure:

[<instruction line>     # ONLY for I2V/FL2VA/L2VA — first line, then a blank line]

integrated_multimodal_description: [Shot 1] <style first>, <composition/subjects/scene>. <camera type + amplitude + speed>. [Shot 2] At 00:0X.XXX, the camera cuts to <next beat>.
overall_soundscape: <1–4 sentences: ambient / physical / non-verbal human sound>
non_diegetic_music: <1–3 sentences: instrumentation / tempo / rhythm / dynamics — NO mood words>

Hard rules: state style first in Shot 1 (Cinematic / live-action / 2D-animated / 3D CG / claymation / watercolor / vintage film); don't timestamp Shot 1; later shots use strictly increasing cut times within the duration; keep identity/wardrobe/color/objects/spatial relations consistent across shots; camera motion = natural prose combining type + amplitude + speed (Push In / Pull Out / Pan / Truck / Tilt / Arc / Tracking / Static / POV / Roll — omit amplitude/speed when medium/normal); dialogue as <d>[lang] verbatim words</d> with stable (S1)/(S2) speaker IDs (first appearance gives age/gender/on-screen/pitch/timbre/rate/accent); on-screen text in English double quotes, verbatim; every detail must be visible or audible — no abstract mood/emotion words; prefer camera motion over a cut for mere distance/angle changes.

Step 4 — Validate before generate (core/validator.py, hard gate — never skip). Checks: all 3 required fields present; duration within 4–15s; mode matches the instruction line; image/ref counts match the mode; cut times strictly increasing and within duration; <d>[lang]…</d> well-formed. Fix every issue before spending GPU. A shot that fails validation will fail the judge.

Step 5 — Generate via ComfyUI (backends/h3/backend.py + engines/comfyui/adapter.py). Confirm the server is up first (§3). H3 defaults: 1344×768, 20 steps, res_multistep / simple scheduler, shift_video=12.0 / shift_audio=3.0, INT8 ConvRot quants, engine flags --lowvram --use-sage-attention. length snaps to the 17k+5 grid.

Step 6 — Judge the output (core/judge.py). Activate the real judge with env: OPEN_VIDEO_VLM_URL + OPEN_VIDEO_VLM_MODEL (+ OPEN_VIDEO_VLM_KEY) — the pipeline then extracts frames and assesses vs prompt intent + quality bar automatically (QualityJudge.from_env() is the entry point; explicit QualityJudge(vision_fn=…) also works). Verdict: PASS / REFINE / FAIL, with score + issues in the receipt (run --json exposes them). With the env unset the judge auto-PASSes — then you must manually review frames.

Step 7 — Refine if REFINE or FAIL. Diagnose the specific issue, apply a targeted fix (prompt tweak / +steps / different mode / ref-pack for identity lock / different seed), regenerate. Strategy is refine-primary, not best-of-N: H3 raw quality is already at parity — the loop fixes adherence, length, and consistency, which is where the gap actually lives. Best-of-N is an optional escape hatch, not the default.

Step 8 — Stitch multi-shot (core/pipeline.py LongFilmPipeline.make_film). Each subsequent shot's first_frame = the previous shot's last frame (ffmpeg-extracted at -sseof -0.1); a t2v shot is auto-upgraded to i2v when a handoff frame exists. Then ffmpeg concat (-f concat -c copy)

  • cross-shot audio continuity (music theme / dialogue language / ambient crossfade). Optional 2K upscale via API as a final step.

Step 9 — Deliver. One coherent film + per-shot receipts (prompt, seed, settings, judge verdict, extracted frames). Persist receipts under artifacts/verify/.

Show full SKILL.md (541 more words)Show less

3. Key commands

Server — start ComfyUI first (the engine), on http://127.0.0.1:8188:

bash
cd /path/to/ComfyUI && python main.py --lowvram --use-sage-attention   # H3 default flags

Health check: curl -s http://127.0.0.1:8188/system_stats (returns JSON when up). In Python: ComfyUIAdapter(server="http://127.0.0.1:8188").health().

Generate one shot — open-video Python API (the contract; lives in backends/ + engines/):

python
from open_video.backends.h3.backend import H3Backend
from open_video.core.backend import ShotRequest
from open_video.engines.comfyui.adapter import ComfyUIAdapter

engine = ComfyUIAdapter(server="http://127.0.0.1:8188")
backend = H3Backend()
req = ShotRequest(prompt=<3-field prompt string>, mode="t2v",
                  width=1344, height=768, duration_s=10.0, seed=0)
result = backend.generate(req, engine=engine)   # → ShotResult(ok, video_path, receipt)
  • User-facing CLI (shipped): open-video "<concept>" --duration 10 --model h3 --json (durations above ~15s trigger the multi-shot v0 template path — single-shot ≤15s uses your prompt verbatim and is the reliable v0.0.1 route).
  • Proven baseline scripts (same workflows, ported from the early lab):
    • Single shot: scripts/h3_agent.py --request "<simple NL>" --duration 5 --width 1344 --height 768 (use --prompt "<full 3-field prompt>" for best quality; --first-frame/--last-frame for I2V/FL2VA).
    • Validate: scripts/validate_prompt.py (exit 0 = clean, 1 = issues).

Multishot / long film — core/pipeline.py LongFilmPipeline.make_film(plan, out_path): plan is a list[Shot(scene_id, prompt, mode, duration_s, seed, …)]. The pipeline generates each shot → judges → extracts last frame → chains (FL2VA handoff) → stitches → writes output/film.mp4 and returns (film_path, plan_with_receipts).

  • Proven baseline: scripts/h3_multishot.py --plan library/plans/multishot_demo.json --out output/long_demo.mp4 (a ready example plan ships at library/plans/multishot_demo.json; plan JSON = {"shots": [{"prompt_file": "...", "duration": 10, "first_frame": null}, …]}).

4. Constraints (H3 baseline — hard, evidence-based)

  • Duration: 4–15s per shot (backends/h3/backend.py constraints() → duration_range_s: (4, 15)). No single shot >15s. For longer content, use multishot — that is the entire point of the stitcher.
  • Frame grid: num_frames snaps to 17k+5 @ 24fps (video-VAE temporal constraint; _snap_17k5). Valid lengths: 5, 22, 39, 56, … Do not pass arbitrary frame counts.
  • Resolution: local ceiling = 768 short edge (native canvas 768×1344), multiples of 32 (resolution_multiple: 32). 2K = API upscale only — never attempt 2K locally.
  • Refs: ≤9 images, ≤3 videos, ≤3 audios, ≤12 total (max_refs).
  • Quant — use INT8 ConvRot (minimax_h3_fl2va_pruned_int8_convrot.safetensors, ~21 GB, proven on RTX 5090). Avoid NVFP4 on RTX 5090 — ComfyUI issue #14157 bug (recorded in backends/h3/backend.py default_settings()["known_issues"] and docs/h3_ecosystem.md). Other quants: NF4 (~8 GB, lowest VRAM), W4 ConvRot (~10 GB, stock-compatible), BF16 (~62 GB, multi-GPU only).
  • Audio: 32 kHz stereo native; CFG-distilled (no negative prompt, no guidance scale).
  • Known issues to guard against (docs/h3_ecosystem.md "Known issues"): wide-shot face corruption (Comfy-Org #30), 2K upscale fails in ref2va (#19), ref2va persistent noise (HF #50), AMD/Apple Silicon partial support (#17/#24/#33), prompt metadata embedded in output files (#13).

5. Docs — read these for depth (do not guess from memory)

  • README.md — what/why, three interfaces (App / CLI / Skill), the thesis.
  • ARCHITECTURE.md — core / backends / engines / library layers; the quality loop; the long-film pipeline diagram.
  • PLAN.md — phased roadmap, open decisions, the make-or-break success metric.
  • backends/h3/PROMPT_GRAMMAR.md — the official H3 3-field prompt guide (condensed). Read before crafting any H3 prompt.
  • docs/h3_ecosystem.md — quants, multishot tools, speed nodes, known issues, build-on targets.
  • templates/model_backend.py — the plugin template (§6).
  • CONTRIBUTING.md — plugin points + good-first-issues.

6. Adding a model backend (Wan2.2, LTX, …)

The core never changes — add a plugin. Copy templates/model_backend.py → backends/<model>/backend.py and implement the ModelBackend ABC (core/backend.py):

  1. capabilities (Capabilities(...)) — which modes (t2v/i2v/flf2v/r2v), native_audio, max_duration_s, max_short_edge_px, strengths (the selector keys on these).
  2. prompt_guide() + craft_prompt(intent, mode) — your model's prompt grammar.
  3. constraints() — duration range, frame grid (or None), max_refs, resolution multiple.
  4. generate(req, engine) + _build_workflow(req) — build the engine workflow (e.g. ComfyUI JSON loaded from workflows/), run via the engine adapter, return ShotResult(ok, video_path, receipt).
  5. default_settings() — steps / sampler / scheduler / quant (evidence-based via bench/).
  6. duration_to_length() / resolution_for() — model-specific frame/fps and resolution-grid math.

Then add backends/<model>/workflows/ (engine JSON), PROMPT_GRAMMAR.md, and __init__.py. See backends/h3/backend.py for a complete working example — the H3 backend is the reference implementation. PR it; CONTRIBUTING.md lists plugin points as good-first-issues.

© agent-next, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skill/open-video of agent-next/video-agent.

Open the folder on GitHubat commit 12a10e9

Compare with similar skills

Open Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Open Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Open Video this skillagent-next/video-agent120—~3.2kAutomated safety check: PassApache-2.0
VRGDG H3 Short Film Pipelinevrgamegirl19/comfyui-vrgamedevgirl760—~4.2kAutomated safety check: PassCustom licence
ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit105—~12kAutomated safety check: PassApache-2.0
Cs Digital Human Product Video PipelineChenShuo2004/cs-skills193—~1.3kAutomated safety check: PassMIT
H3 Video Prompt Enhancerbenjiyaya/Calliope241—~4.6kAutomated safety check: PassMIT
MiniMax H3 Video DirectorTFboy1/oh-my-minimaxh3-director140—~2.4kAutomated safety check: PassMIT

Similar skills

  • VRGDG H3 Short Film Pipeline

    vrgamegirl19/comfyui-vrgamedevgirl

    Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.

    760 GitHub stars~4.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • 当用户要把产品事实、口播、数字人、产品界面和 CTA 制作成可验收的数字人产品介绍视频时使用:统一预检 ChatCut、FFmpeg、ComfyUI、Fish/TTS 与 Remotion,按 plan、sample、batch 三种模式编排,先完成可审批样片再批量或出成片。用于有明确产品包的横版/竖版产品视频流水线;不要用于通用自动剪辑、电商短视频复刻、纯视频选题策划或未经确认的批量生成与发布。

    193 GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • H3 Video Prompt Enhancer

    benjiyaya/Calliope

    A skill your agent uses when making MiniMax H3 video prompts from media + ideas.

    241 GitHub stars~4.6k tokensUpdated today
    Media & CreativeAuto-check passed
  • MiniMax H3 Video Director

    TFboy1/oh-my-minimaxh3-director

    Turns a script into a storyboard, assigns MiniMax H3 workflows in ComfyUI, monitors batch generation and builds a Jianying draft of the finished video.

    140 GitHub stars~2.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Long-Form H3 Director

    karuvanan/MiniMax-H3-Director-Cut-Studio

    Plans long-form MiniMax H3 video productions as Sequence, Shot, and Segment contracts that stay continuous across generation windows.

    131 GitHub stars~1.5k tokensUpdated 22 days ago
    Media & CreativeAuto-check passed

More from agent-next/video-agent

  • H3 Video

    agent-next/video-agent

    OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

    120 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Open Video

What does Open Video do?

Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Open Video is an agent skill from agent-next/video-agent.2 / LTX future).

When should I use Open Video?

Open Video fits situations like: the user wants to turn a concept; reference image into a finished video; multi-shot film — single shots; first-last-frame interpolation.

How do I install Open Video in Claude Code?

Run `npx skills add agent-next/video-agent --skill open-video -a claude-code`. Or copy the skill folder (skill/open-video in agent-next/video-agent) into .claude/skills/open-video in your project. Claude Code loads it when a task matches its description.

How do I install Open Video in Codex?

Run `npx skills add agent-next/video-agent --skill open-video -a codex`. Or copy the skill folder (skill/open-video in agent-next/video-agent) into .agents/skills/open-video in your project. Codex loads it when a task matches its description.

Can I use Open Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agent-next/video-agent --skill open-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-video, .gemini/skills/open-video, .github/skills/open-video and .opencode/skills/open-video in your project.

What does Open Video need to run?

Going by SKILL.md and its folder, Open Video needs the command-line tools its instructions call (python and curl) and credentials named OPEN_VIDEO_VLM_KEY. Our summary lists: Python 3; A credential in OPEN_VIDEO_VLM_KEY.

Does Open Video access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Open Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Open Video use?

Open Video is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Open Video use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Open Video?

Skills that share tags, products or a category with Open Video: VRGDG H3 Short Film Pipeline (vrgamegirl19/comfyui-vrgamedevgirl, 760 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars), Cs Digital Human Product Video Pipeline (ChenShuo2004/cs-skills, 193 stars) and H3 Video Prompt Enhancer (benjiyaya/Calliope, 241 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Open Video?

agent-next (a GitHub organization) maintains it in agent-next/video-agent, which has 120 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.

Source: agent-next/video-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.