Agent skill

H3 Video

by agent-next in agent-next/video-agent

OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

Apache-2.0Auto-check passedMedia & Creative

Install H3 Video

skills CLI
$ npx skills add agent-next/video-agent --skill h3-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agent-next/video-agent h3-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agent-next/video-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill/h3-video .claude/skills/h3-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
h3-video
GitHub stars
120
Token cost
~2.9k tokens
SKILL.md length
854 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

  • Works in 10 steps: Paths (env-first) → Quality hierarchy (what actually moves… → Best-quality recipe (local max on RTX… → …
  • OpenVideo install/pull/status/run
  • SKILL.md covers 0. Paths (env-first), 1. Quality hierarchy (what…, 2. Best-quality recipe (local… and 3. Agent procedure (always…, plus 6 more sections
  • Calls python, curl and ffprobe; needs OPEN_VIDEO_VLM_KEY

What it does

H3 Video is an agent skill from agent-next/video-agent. OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend). Use for OpenVideo install/pull/status/run, official 3-field prompts, T2V/I2V/FL2VA, agent-driven video. Brand is OpenVideo — not a bare ComfyUI workflow. Triggers: OpenVideo, open-video, H3, generate video, T2V, I2V, FL2VA.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation and Diffusion and image models. It works with ComfyUI and MiniMax. The repository describes itself as: Open-source video generation — Ollama for MiniMax H3. Local director on ComfyUI. The licence is Apache-2.0.

When your agent uses it

  • OpenVideo install/pull/status/run
  • Official 3-field prompts
  • Agent-driven video

Example prompts

  • “/h3-video”

Requirements

  • Python 3
  • A credential in OPEN_VIDEO_VLM_KEY

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Paths (env-first)
  2. Quality hierarchy (what actually moves the needle)
  3. Best-quality recipe (local max on RTX 5090)
  4. Agent procedure (always this order)
  5. Ollama-shaped command cheat sheet
  6. Prompt micro-templates (copy & fill)
  7. Hard constraints
  8. When to use which skill
  9. Docs (read these, don't invent)
  10. Commit / path hygiene

What it can do on your machine

Read from SKILL.md and the folder at commit 12a10e9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • curl
    • ffprobe
    • git
    • bash
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPEN_VIDEO_VLM_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

H3 Video loads about 2.9k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 854 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agent-next/video-agent at commit 12a10e9, republished under its Apache-2.0 licence (© agent-next). 854 words, ~2,949 tokens.

Download SKILL.mdSave it as .claude/skills/h3-video/SKILL.md (or your agent's skills folder).
name
h3-video
description
OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend). Use for OpenVideo install/pull/status/run, official 3-field prompts, T2V/I2V/FL2VA, agent-driven video. Brand is OpenVideo — not a bare ComfyUI workflow. Triggers: OpenVideo, open-video, H3, generate video, T2V, I2V, FL2VA.

OpenVideo skill · v0.1.0

Brand: OpenVideo (always lead with this). MiniMax H3 is the model OpenVideo drives.

Job: use OpenVideo so an agent produces good product video, not a random diffusion click. Quality comes from prompt craft + correct mode + validated settings, not secret samplers.

0. Paths (env-first)

Recommended sibling layout (any parent dir name; no machine-absolute paths):

text
parent/
├── open-video/       # THIS product — CLI, backends, skills  ← work here
└── lab/              # ComfyUI + weights (not git)
WhatPath / env
Product root (CLI, skills)$OPEN_VIDEO_ROOT (this checkout)
Install / pull / runFrom product root: ./scripts/install.sh, python -m open_video …
Lab engine + weights$OPEN_VIDEO_LAB / $H3_LAB (sibling lab/ with ComfyUI + h3_models/)
Weights envexport OPEN_VIDEO_MODELS=$H3_LAB/h3_models
ComfyUI URLexport OPEN_VIDEO_COMFYUI=http://127.0.0.1:8188 (default)

Resolve product root (never hardcode a machine path):

bash
# Prefer env, else directory of this skill → repo root
export OPEN_VIDEO_ROOT="${OPEN_VIDEO_ROOT:-$(cd "$(dirname "$0")/../.." 2>/dev/null && pwd)}"
cd "$OPEN_VIDEO_ROOT"

Lab harness (same ComfyUI) when you keep a sibling runtime:

bash
export OPEN_VIDEO_LAB="${OPEN_VIDEO_LAB:-$OPEN_VIDEO_ROOT/../lab}"
export H3_LAB="${H3_LAB:-$OPEN_VIDEO_LAB}"
export OPEN_VIDEO_MODELS="${OPEN_VIDEO_MODELS:-$H3_LAB/h3_models}"

Product = open-video/ checkout. lab/ is runtime only — never the product root.


1. Quality hierarchy (what actually moves the needle)

RankLeverRule
13-field promptOfficial structure; concrete visible/audible detail; camera type+amplitude+speed; dialogue in <d>[lang]…</d>
2ModeT2V / I2V / FL2VA chosen correctly from inputs
3Resolution / steps1344×768, 20 steps, res_multistep + simple (defaults in harness)
4QuantINT8 on 5090-class (the only tier pull installs); lower VRAM → recommend-quant may suggest nf4/w4, which are manual + experimental — never auto-installed
5Duration5–10 s sweet spot (max 15 s / shot); multi-shot via cut times or open-video director skill
6ReviewWatch / extract frames; fix prompt; re-run. Dual vision gate only when shipping

Anti-patterns (low quality): bare one-liner prompts; abstract mood words only; wrong mode; NVFP4 on 5090; forcing 2K locally; skipping validation.


2. Best-quality recipe (local max on RTX 5090)

text
Diffusion:  fl2va_pruned_int8_convrot
Text enc:   qwen3vl_32b_int8_convrot
VAE:        video fp16 + audio fp32
ComfyUI:    --lowvram --use-sage-attention   # lowvram optional if VRAM ≥ ~22 GiB free policy
Canvas:     1344×768 (16:9), multiple of 32
Sampler:    res_multistep | Scheduler: simple | Steps: 20
Duration:   5–10 s (snaps to 17k+5 frames @ 24 fps)
Audio:      native 32 kHz stereo — describe in prompt fields
Prompt:     official 3-field only

Order-of-magnitude on a 32 GB class NVIDIA GPU: ~10–15 min / 10 s clip @ 1344×768, peak VRAM ~22–30 GB (measure on your machine).


3. Agent procedure (always this order)

A. Host ready
bash
cd "${OPEN_VIDEO_ROOT:?set OPEN_VIDEO_ROOT to this product checkout}"
python -m open_video status          # or: open-video status / ps
python -m open_video recommend-quant
  • Weights incomplete → python -m open_video pull h3 (or OPEN_VIDEO_MODELS=… pull)
  • ComfyUI down → start lab server:
bash
cd "${H3_LAB:-$OPEN_VIDEO_ROOT/../lab}"
curl -sf http://127.0.0.1:8188/system_stats || (
  mkdir -p logs && cd ComfyUI && nohup ../venv/bin/python main.py \
    --listen 127.0.0.1 --port 8188 --lowvram --use-sage-attention \
    > ../logs/comfy_server.log 2>&1 &
)
B. Mode
InputsMode
Text onlyT2V
+ 1 imageI2V (+ instruction line)
+ first & last imageFL2VA (+ alignment line)
Multi-ref identity/styleR2V (needs ref2va weights — not default)

CLI --mode tokens: t2v · i2v · flf2v (FL2VA ↔ --mode flf2v). R2V has no CLI token yet.

C. Craft the 3-field prompt (quality lever #1)
text
[<instruction line — I2V / FL2VA only, then blank line>]

integrated_multimodal_description: [Shot 1] <style first: Live-action, cinematic, …>,
  <subjects, clothing, lighting, space>. <camera: type + amplitude + speed + action>.
  [Shot 2] At 00:0X.XXX, the camera cuts to <new info only>.
overall_soundscape: <1–4 sentences: ambient / physical / non-verbal — no duplicate dialogue>
non_diegetic_music: <1–3 sentences: instruments, tempo, rhythm — no vague mood words>

Hard rules (from official guide):

  • Every claim must be visible or audible (no pure emotion adjectives).
  • Style first in Shot 1.
  • Camera prose: e.g. pushes in with small amplitude at slow speed.
  • Dialogue: speaker IDs + <d>[English] verbatim words</d>.
  • Later shots: strictly increasing cut times inside duration.
  • Full grammar: backends/h3/PROMPT_GRAMMAR.md (product) or lab docs/PROMPT_GUIDE.md.

Expand casual NL into 3-field; never ship a bare phrase as the only prompt for “high quality.”

D. Dry-run (cheap)
bash
cd "$OPEN_VIDEO_ROOT"
python -m open_video run "$(cat prompts/my_shot.txt)" --duration 8 --dry-run
# optional, only if a lab tree exists at $H3_LAB:
#   cd "$H3_LAB" && ./venv/bin/python "$OPEN_VIDEO_ROOT/scripts/h3_agent.py" --prompt "$(cat …)" --dry-run

Fix validator issues before GPU spend.

E. Generate (product CLI or lab agent — both use H3 workflows)

Preferred product path (OpenVideo harness):

bash
cd "$OPEN_VIDEO_ROOT"
python -m open_video run "$(cat prompts/my_shot.txt)" \
  --duration 8 --model h3 --output output/film.mp4
# aliases: open-video "…", open-video run "…"

I2V / FL2VA (product CLI — supply frames; FL2VA is --mode flf2v):

bash
python -m open_video run "$(cat prompts/my_shot.txt)" \
  --mode i2v --first-frame inputs/start.png --duration 8
python -m open_video run "$(cat prompts/my_shot.txt)" \
  --mode flf2v --first-frame inputs/start.png --last-frame inputs/end.png --duration 8

Optional lab path (only when a lab tree exists — $H3_LAB set, with its own venv): cd "$H3_LAB" && ./venv/bin/python "$OPEN_VIDEO_ROOT/scripts/h3_agent.py" --prompt … --width 1344 --height 768 --duration 8 --seed 42 (mp4 → output/, receipt → artifacts/verify/agent_*.json).

Show full SKILL.md (372 more words)Show less
F. Review & iterate
  1. Automatic VLM judge (opt-in): set OPEN_VIDEO_VLM_URL + OPEN_VIDEO_VLM_MODEL (+ OPEN_VIDEO_VLM_KEY) to any OpenAI-compatible vision endpoint and the pipeline judges every shot for real (score + issues in the --json output). Without these env vars the verdict is honestly SKIPPED (score 0) — never a fake PASS — so the manual review below is mandatory, not optional.
  2. Play the mp4 (native audio matters).
  3. Extract a contact sheet:
    ffmpeg -y -i out.mp4 -vf "fps=1,scale=320:-1,tile=4x2" contact.png
  4. If weak: fix prompt first (specificity, camera, continuity), then seed, then duration.
  5. Shipping / public claim: dual visual review per org rules if required.
G. Deliver

Verify before claiming done: exit code 0 AND DONE -> <path> printed, then

bash
[ -f "$path" ] && ffprobe -v error -show_entries format=duration -of csv=p=0 "$path"

(or run with --json and read film + per-shot verdict/judge_score from the final stdout line). Only then tell the user: path · duration · resolution · seed · mode · wall time (from receipt).


4. Ollama-shaped command cheat sheet

bash
# Install (once) — v0.1.0 from the GitHub tag
git clone --depth 1 --branch v0.1.0 https://github.com/agent-next/video-agent
cd video-agent && bash scripts/install.sh   # Windows: run inside WSL2
# The website installer is updated separately and may serve an older version.

cd "$OPEN_VIDEO_ROOT"
python -m open_video pull h3              # download / verify weights
python -m open_video pull h3 --check-only
python -m open_video status               # = ps
python -m open_video recommend-quant

python -m open_video run "<3-field or concept>" --duration 8
python -m open_video "concept" --dry-run
python -m open_video list-models
EnvMeaning
OPEN_VIDEO_ROOTProduct checkout
OPEN_VIDEO_MODELSWeights root (h3_models or ComfyUI/models)
OPEN_VIDEO_COMFYUIComfyUI base URL
OPEN_VIDEO_COMFYUI_INPUTI2V/FL2V staging dir — must belong to that ComfyUI server (unique filenames per run)
OPEN_VIDEO_MODELDefault backend (h3)
H3_LABOptional lab tree with ComfyUI + h3_agent
OPEN_VIDEO_VLM_URLOpenAI-compatible vision endpoint → real judge
OPEN_VIDEO_VLM_MODELVision model id for the judge
OPEN_VIDEO_VLM_KEYBearer token for the judge endpoint (optional)
OPEN_VIDEO_JUDGE_RETRIESExtra takes on REFINE verdicts (default 1; best score kept)

5. Prompt micro-templates (copy & fill)

T2V cinematic single beat

text
integrated_multimodal_description: [Shot 1] Live-action, cinematic, <wide/medium> shot frames <subject + clothing + age/gender if speaking>. <lighting + location>. The camera <push/pull/pan/track/arc> with <small|medium|large> amplitude at <slow|medium|fast> speed as <concrete action>.
overall_soundscape: <ambient>… <physical>…
non_diegetic_music: <instruments + tempo + dynamics>…

T2V multi-shot (cut adds new info)

text
integrated_multimodal_description: [Shot 1] … [Shot 2] At 00:05.000, the camera cuts to …
overall_soundscape: …
non_diegetic_music: …

I2V

text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, starting from Picture 1, …
overall_soundscape: …
non_diegetic_music: …

6. Hard constraints

ConstraintLimit
Duration / shot4–15 s
Frame grid17k+5 @ 24 fps
Local resShort edge ≤ 768; multiple of 32
2KAPI only — cannot upscale local 768p with open weights
NVFP4 on 5090Forbidden (ComfyUI #14157)
LicenseMiniMax weight terms (region/commercial) + code Apache-2.0

7. When to use which skill

IntentSkill
High-quality H3 clip, agent generate videoh3-video (this file) — default
Multi-minute film, judge→refine→stitchskill/open-video (experimental; not full product)
Website / Pagesseparate open-video-web repo

8. Docs (read these, don't invent)

DocLocation
Prompt grammar$OPEN_VIDEO_ROOT/backends/h3/PROMPT_GRAMMAR.md
Prompt guide$OPEN_VIDEO_ROOT/docs/h3/PROMPT_GUIDE.md
Best practices$OPEN_VIDEO_ROOT/docs/h3/BEST_PRACTICES.md
Quants / issues$OPEN_VIDEO_ROOT/docs/h3_ecosystem.md
Quickstart$OPEN_VIDEO_ROOT/docs/QUICKSTART.md

9. Commit / path hygiene

  • Product commits: inside the open-video git checkout only.
  • Do not hardcode host-absolute lab paths as the product root; use Path(__file__).resolve() or OPEN_VIDEO_ROOT / H3_LAB.
  • Lab scripts should resolve ROOT from __file__ (relative), not a frozen absolute path.

© agent-next, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skill/h3-video of agent-next/video-agent.

Open the folder on GitHubat commit 12a10e9

Compare with similar skills

H3 Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

H3 Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
H3 Video this skillagent-next/video-agent120—~2.9kAutomated safety check: PassApache-2.0
ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit105—~12kAutomated safety check: PassApache-2.0
H3 Video Prompt Enhancerbenjiyaya/Calliope241—~4.6kAutomated safety check: PassMIT
Minimax H3calesthio/OpenMontage66k—~580Automated safety check: PassAGPL-3.0
High Density Fight PromptJGRFW/comfyui-AICG3D166—~590Automated safety check: PassGPL-3.0
9Router Image Generationdecolua/9router30k—~830Automated safety check: PassMIT

Similar skills

  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • H3 Video Prompt Enhancer

    benjiyaya/Calliope

    A skill your agent uses when making MiniMax H3 video prompts from media + ideas.

    241 GitHub stars~4.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Minimax H3

    calesthio/OpenMontage

    Generate MiniMax H3 (Hailuo 3.0) video through the official MiniMax v2 API, fal.ai, Runway, ComfyUI Partner Nodes, or local open weights in ComfyUI.

    66k GitHub stars~580 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • High Density Fight Prompt

    JGRFW/comfyui-AICG3D

    根据参考图和用户设定创作高密度、连续因果的电影级打斗视频提示词,并输出保持同一时间线的中文导演稿与 MiniMax H3 Ref2VA 英文六段稿。适用于 15 秒动作设计、武器战、徒手战、巨物战和参考图驱动的连续攻防。

    166 GitHub stars~590 tokensUpdated today
    Media & CreativeAuto-check passed
  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    30k GitHub stars~830 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • MiniMax H3 Video Director

    TFboy1/oh-my-minimaxh3-director

    Turns a script into a storyboard, assigns MiniMax H3 workflows in ComfyUI, monitors batch generation and builds a Jianying draft of the finished video.

    141 GitHub stars~2.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from agent-next/video-agent

  • Open Video

    agent-next/video-agent

    Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).

    120 GitHub stars~3.2k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about H3 Video

What does H3 Video do?

OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend). H3 Video is an agent skill from agent-next/video-agent.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

When should I use H3 Video?

H3 Video fits situations like: openVideo install/pull/status/run; official 3-field prompts; agent-driven video.

How do I install H3 Video in Claude Code?

Run `npx skills add agent-next/video-agent --skill h3-video -a claude-code`. Or copy the skill folder (skill/h3-video in agent-next/video-agent) into .claude/skills/h3-video in your project. Claude Code loads it when a task matches its description.

How do I install H3 Video in Codex?

Run `npx skills add agent-next/video-agent --skill h3-video -a codex`. Or copy the skill folder (skill/h3-video in agent-next/video-agent) into .agents/skills/h3-video in your project. Codex loads it when a task matches its description.

Can I use H3 Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agent-next/video-agent --skill h3-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/h3-video, .gemini/skills/h3-video, .github/skills/h3-video and .opencode/skills/h3-video in your project.

What does H3 Video need to run?

Going by SKILL.md and its folder, H3 Video needs the command-line tools its instructions call (python, curl, ffprobe, git, bash and ffmpeg) and credentials named OPEN_VIDEO_VLM_KEY. Our summary lists: Python 3; A credential in OPEN_VIDEO_VLM_KEY.

Does H3 Video access the network?

SKILL.md contains no URLs. Its commands use curl and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is H3 Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does H3 Video use?

H3 Video is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does H3 Video use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to H3 Video?

Skills that share tags, products or a category with H3 Video: ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars), H3 Video Prompt Enhancer (benjiyaya/Calliope, 241 stars), Minimax H3 (calesthio/OpenMontage, 66k stars) and High Density Fight Prompt (JGRFW/comfyui-AICG3D, 166 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains H3 Video?

agent-next (a GitHub organization) maintains it in agent-next/video-agent, which has 120 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.

Source: agent-next/video-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.