VRGDG H3 Short Film Pipeline
vrgamegirl19/comfyui-vrgamedevgirl
Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.
Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).
$ npx skills add agent-next/video-agent --skill open-video -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agent-next/video-agent open-video --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill/open-video .claude/skills/open-video && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .claude/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agent-next/video-agent/tree/master/skill/open-videoType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agent-next/video-agent --skill open-video -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agent-next/video-agent open-video --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skill/open-video .agents/skills/open-video && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .agents/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agent-next/video-agent --skill open-video -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agent-next/video-agent open-video --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skill/open-video .cursor/skills/open-video && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .cursor/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agent-next/video-agent.git --path skill/open-video--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agent-next/video-agent --skill open-video -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agent-next/video-agent open-video --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skill/open-video .gemini/skills/open-video && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .gemini/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agent-next/video-agent open-videoInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agent-next/video-agent --skill open-video -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skill/open-video .github/skills/open-video && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .github/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agent-next/video-agent --skill open-video -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agent-next/video-agent open-video --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skill/open-video .opencode/skills/open-video && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "open-video" agent skill from https://github.com/agent-next/video-agent/tree/master/skill/open-video into .opencode/skills/open-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-video", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
open-videoGenerate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).
Open Video is an agent skill from agent-next/video-agent. Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Use when the user wants to turn a concept, script, or reference image into a finished video or multi-shot film — single shots, image-to-video, first-last-frame interpolation, reference-video/audio styling, or stitched long films beyond the 15s model ceiling. Covers prompt crafting, hard validation, ComfyUI-driven generation, vision judging, refine loop, and ffmpeg stitching.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering AI video generation, Diffusion and image models and Video production. It works with ComfyUI, MiniMax and FFmpeg. The repository describes itself as: Open-source video generation — Ollama for MiniMax H3. Local director on ComfyUI. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 12a10e9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythoncurlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPEN_VIDEO_VLM_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Open Video loads about 3.2k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,338 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agent-next/video-agent at commit 12a10e9, republished under its Apache-2.0 licence (© agent-next). 1,338 words, ~3,165 tokens.
.claude/skills/open-video/SKILL.md (or your agent's skills folder).v0.0.1 default for high-quality H3 clips is
skill/h3-video
(Ollama for H3 + harness). Use this skill when the user wants multi-shot / long-film director behavior (plan → judge → stitch) beyond a single strong clip.
open-video is the autonomous director layer on top of open video models. It turns a concept into a finished film by running the loop no open engine ships natively:
plan → craft → validate → generate → judge → refine → stitch → deliver
core/): planner (coherence bible), crafter, validator, judge-loop,
stitcher, selector, ref-pack builder.backends/<model>/): baseline = MiniMax H3 (#1 open model, Arena
parity with closed, native stereo audio). Future: Wan2.2 (physics), LTX-2.3 (speed).engines/comfyui/): drives ComfyUI via its HTTP API. open-video is the brain;
ComfyUI is the hands.It is NOT a video engine (ComfyUI is the engine) and NOT a model (H3/Wan/LTX are backends). It is the agent brain ComfyUI lacks: judge→refine + multi-shot stitch + coherence planning as they land.
v0.0.1 honesty: prefer skill/h3-video for reliable high-quality
single clips. Multi-minute film is the flagship design. The judge is REAL when env-wired:
set OPEN_VIDEO_VLM_URL + OPEN_VIDEO_VLM_MODEL (+ OPEN_VIDEO_VLM_KEY) to any
OpenAI-compatible vision endpoint and every shot is scored + diagnosed automatically;
with the env unset it is an honest PASS stub — then manual frame review is mandatory.
Single open models still cap ~15s/shot; longer output needs multi-shot orchestration (partial).
Step 1 — Understand the request. Classify the job: single shot vs multi-shot film; target duration; did the user supply reference image(s) / video / audio; desired aspect; quality bar. If the request is ambiguous and the target is a film >30s, ask ONE focused question (subject + mood + length). Do not generate a long film on guesswork.
Step 2 — Pick the mode. Mode is auto-derived from inputs (see backends/h3/backend.py and
scripts/validate_prompt.py detect_mode):
"For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced."<Picture 1>, <Video 1>, <Audio 1> and explicitly assign each a role.Step 3 — Craft the 3-field prompt (per backends/h3/PROMPT_GRAMMAR.md). Exact structure:
[<instruction line> # ONLY for I2V/FL2VA/L2VA — first line, then a blank line]
integrated_multimodal_description: [Shot 1] <style first>, <composition/subjects/scene>. <camera type + amplitude + speed>. [Shot 2] At 00:0X.XXX, the camera cuts to <next beat>.
overall_soundscape: <1–4 sentences: ambient / physical / non-verbal human sound>
non_diegetic_music: <1–3 sentences: instrumentation / tempo / rhythm / dynamics — NO mood words>Hard rules: state style first in Shot 1 (Cinematic / live-action / 2D-animated / 3D CG / claymation / watercolor / vintage film); don't timestamp Shot 1; later shots use strictly increasing cut times within the duration; keep identity/wardrobe/color/objects/spatial relations consistent across shots; camera motion = natural prose combining type + amplitude + speed (Push In / Pull Out / Pan / Truck / Tilt / Arc / Tracking / Static / POV / Roll — omit amplitude/speed when medium/normal); dialogue as <d>[lang] verbatim words</d> with stable (S1)/(S2) speaker IDs (first appearance gives age/gender/on-screen/pitch/timbre/rate/accent); on-screen text in English double quotes, verbatim; every detail must be visible or audible — no abstract mood/emotion words; prefer camera motion over a cut for mere distance/angle changes.
Step 4 — Validate before generate (core/validator.py, hard gate — never skip). Checks: all 3
required fields present; duration within 4–15s; mode matches the instruction line; image/ref counts
match the mode; cut times strictly increasing and within duration; <d>[lang]…</d> well-formed.
Fix every issue before spending GPU. A shot that fails validation will fail the judge.
Step 5 — Generate via ComfyUI (backends/h3/backend.py + engines/comfyui/adapter.py).
Confirm the server is up first (§3). H3 defaults: 1344×768, 20 steps, res_multistep / simple
scheduler, shift_video=12.0 / shift_audio=3.0, INT8 ConvRot quants, engine flags
--lowvram --use-sage-attention. length snaps to the 17k+5 grid.
Step 6 — Judge the output (core/judge.py). Activate the real judge with env:
OPEN_VIDEO_VLM_URL + OPEN_VIDEO_VLM_MODEL (+ OPEN_VIDEO_VLM_KEY) — the pipeline then
extracts frames and assesses vs prompt intent + quality bar automatically
(QualityJudge.from_env() is the entry point; explicit QualityJudge(vision_fn=…) also works).
Verdict: PASS / REFINE / FAIL, with score + issues in the receipt (run --json exposes them).
With the env unset the judge auto-PASSes — then you must manually review frames.
Step 7 — Refine if REFINE or FAIL. Diagnose the specific issue, apply a targeted fix (prompt tweak / +steps / different mode / ref-pack for identity lock / different seed), regenerate. Strategy is refine-primary, not best-of-N: H3 raw quality is already at parity — the loop fixes adherence, length, and consistency, which is where the gap actually lives. Best-of-N is an optional escape hatch, not the default.
Step 8 — Stitch multi-shot (core/pipeline.py LongFilmPipeline.make_film). Each subsequent
shot's first_frame = the previous shot's last frame (ffmpeg-extracted at -sseof -0.1); a t2v
shot is auto-upgraded to i2v when a handoff frame exists. Then ffmpeg concat (-f concat -c copy)
Step 9 — Deliver. One coherent film + per-shot receipts (prompt, seed, settings, judge verdict,
extracted frames). Persist receipts under artifacts/verify/.
Server — start ComfyUI first (the engine), on http://127.0.0.1:8188:
cd /path/to/ComfyUI && python main.py --lowvram --use-sage-attention # H3 default flagsHealth check: curl -s http://127.0.0.1:8188/system_stats (returns JSON when up). In Python:
ComfyUIAdapter(server="http://127.0.0.1:8188").health().
Generate one shot — open-video Python API (the contract; lives in backends/ + engines/):
from open_video.backends.h3.backend import H3Backend
from open_video.core.backend import ShotRequest
from open_video.engines.comfyui.adapter import ComfyUIAdapter
engine = ComfyUIAdapter(server="http://127.0.0.1:8188")
backend = H3Backend()
req = ShotRequest(prompt=<3-field prompt string>, mode="t2v",
width=1344, height=768, duration_s=10.0, seed=0)
result = backend.generate(req, engine=engine) # → ShotResult(ok, video_path, receipt)open-video "<concept>" --duration 10 --model h3 --json
(durations above ~15s trigger the multi-shot v0 template path — single-shot ≤15s uses your
prompt verbatim and is the reliable v0.0.1 route).scripts/h3_agent.py --request "<simple NL>" --duration 5 --width 1344 --height 768
(use --prompt "<full 3-field prompt>" for best quality; --first-frame/--last-frame for I2V/FL2VA).scripts/validate_prompt.py (exit 0 = clean, 1 = issues).Multishot / long film — core/pipeline.py LongFilmPipeline.make_film(plan, out_path):
plan is a list[Shot(scene_id, prompt, mode, duration_s, seed, …)]. The pipeline generates each
shot → judges → extracts last frame → chains (FL2VA handoff) → stitches → writes output/film.mp4
and returns (film_path, plan_with_receipts).
scripts/h3_multishot.py --plan library/plans/multishot_demo.json --out output/long_demo.mp4
(a ready example plan ships at library/plans/multishot_demo.json; plan JSON =
{"shots": [{"prompt_file": "...", "duration": 10, "first_frame": null}, …]}).backends/h3/backend.py constraints() → duration_range_s: (4, 15)).
No single shot >15s. For longer content, use multishot — that is the entire point of the stitcher._snap_17k5).
Valid lengths: 5, 22, 39, 56, … Do not pass arbitrary frame counts.resolution_multiple: 32). 2K = API upscale only — never attempt 2K locally.max_refs).minimax_h3_fl2va_pruned_int8_convrot.safetensors, ~21 GB, proven on
RTX 5090). Avoid NVFP4 on RTX 5090 — ComfyUI issue #14157 bug (recorded in
backends/h3/backend.py default_settings()["known_issues"] and docs/h3_ecosystem.md).
Other quants: NF4 (~8 GB, lowest VRAM), W4 ConvRot (~10 GB, stock-compatible), BF16 (~62 GB, multi-GPU only).docs/h3_ecosystem.md "Known issues"): wide-shot face
corruption (Comfy-Org #30), 2K upscale fails in ref2va (#19), ref2va persistent noise (HF #50),
AMD/Apple Silicon partial support (#17/#24/#33), prompt metadata embedded in output files (#13).README.md — what/why, three interfaces (App / CLI / Skill), the thesis.ARCHITECTURE.md — core / backends / engines / library layers; the quality loop; the long-film pipeline diagram.PLAN.md — phased roadmap, open decisions, the make-or-break success metric.backends/h3/PROMPT_GRAMMAR.md — the official H3 3-field prompt guide (condensed). Read before crafting any H3 prompt.docs/h3_ecosystem.md — quants, multishot tools, speed nodes, known issues, build-on targets.templates/model_backend.py — the plugin template (§6).CONTRIBUTING.md — plugin points + good-first-issues.The core never changes — add a plugin. Copy templates/model_backend.py →
backends/<model>/backend.py and implement the ModelBackend ABC (core/backend.py):
capabilities (Capabilities(...)) — which modes (t2v/i2v/flf2v/r2v),
native_audio, max_duration_s, max_short_edge_px, strengths (the selector keys on these).prompt_guide() + craft_prompt(intent, mode) — your model's prompt grammar.constraints() — duration range, frame grid (or None), max_refs, resolution multiple.generate(req, engine) + _build_workflow(req) — build the engine workflow (e.g. ComfyUI
JSON loaded from workflows/), run via the engine adapter, return ShotResult(ok, video_path, receipt).default_settings() — steps / sampler / scheduler / quant (evidence-based via bench/).duration_to_length() / resolution_for() — model-specific frame/fps and resolution-grid math.Then add backends/<model>/workflows/ (engine JSON), PROMPT_GRAMMAR.md, and __init__.py.
See backends/h3/backend.py for a complete working example — the H3 backend is the reference
implementation. PR it; CONTRIBUTING.md lists plugin points as good-first-issues.
© agent-next, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skill/open-video of agent-next/video-agent.
Open the folder on GitHubat commit 12a10e9
Open Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Open Video this skillagent-next/video-agent | 120 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| VRGDG H3 Short Film Pipelinevrgamegirl19/comfyui-vrgamedevgirl | 760 | — | ~4.2k | Automated safety check: Pass | Custom licence | |
| ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit | 105 | — | ~12k | Automated safety check: Pass | Apache-2.0 | |
| Cs Digital Human Product Video PipelineChenShuo2004/cs-skills | 193 | — | ~1.3k | Automated safety check: Pass | MIT | |
| H3 Video Prompt Enhancerbenjiyaya/Calliope | 241 | — | ~4.6k | Automated safety check: Pass | MIT | |
| MiniMax H3 Video DirectorTFboy1/oh-my-minimaxh3-director | 140 | — | ~2.4k | Automated safety check: Pass | MIT |
vrgamegirl19/comfyui-vrgamedevgirl
Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.
SlavaSexton/ComfyUI-Agent-Kit
Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.
ChenShuo2004/cs-skills
当用户要把产品事实、口播、数字人、产品界面和 CTA 制作成可验收的数字人产品介绍视频时使用:统一预检 ChatCut、FFmpeg、ComfyUI、Fish/TTS 与 Remotion,按 plan、sample、batch 三种模式编排,先完成可审批样片再批量或出成片。用于有明确产品包的横版/竖版产品视频流水线;不要用于通用自动剪辑、电商短视频复刻、纯视频选题策划或未经确认的批量生成与发布。
benjiyaya/Calliope
A skill your agent uses when making MiniMax H3 video prompts from media + ideas.
TFboy1/oh-my-minimaxh3-director
Turns a script into a storyboard, assigns MiniMax H3 workflows in ComfyUI, monitors batch generation and builds a Jianying draft of the finished video.
karuvanan/MiniMax-H3-Director-Cut-Studio
Plans long-form MiniMax H3 video productions as Sequence, Shot, and Segment contracts that stay continuous across generation windows.
agent-next/video-agent
OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).
Categories
Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future). Open Video is an agent skill from agent-next/video-agent.2 / LTX future).
Open Video fits situations like: the user wants to turn a concept; reference image into a finished video; multi-shot film — single shots; first-last-frame interpolation.
Run `npx skills add agent-next/video-agent --skill open-video -a claude-code`. Or copy the skill folder (skill/open-video in agent-next/video-agent) into .claude/skills/open-video in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agent-next/video-agent --skill open-video -a codex`. Or copy the skill folder (skill/open-video in agent-next/video-agent) into .agents/skills/open-video in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agent-next/video-agent --skill open-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-video, .gemini/skills/open-video, .github/skills/open-video and .opencode/skills/open-video in your project.
Going by SKILL.md and its folder, Open Video needs the command-line tools its instructions call (python and curl) and credentials named OPEN_VIDEO_VLM_KEY. Our summary lists: Python 3; A credential in OPEN_VIDEO_VLM_KEY.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Open Video is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Open Video: VRGDG H3 Short Film Pipeline (vrgamegirl19/comfyui-vrgamedevgirl, 760 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars), Cs Digital Human Product Video Pipeline (ChenShuo2004/cs-skills, 193 stars) and H3 Video Prompt Enhancer (benjiyaya/Calliope, 241 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agent-next (a GitHub organization) maintains it in agent-next/video-agent, which has 120 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.
Source: agent-next/video-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.