H3 Video
agent-next/video-agent
OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).
A skill your agent uses when making MiniMax H3 video prompts from media + ideas.
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .claude/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .claude/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .claude/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .agents/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .agents/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .agents/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .cursor/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .cursor/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benjiyaya/Calliope.git --path calliope-backend/skills_builtin/h3-video-prompt-enhancer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .gemini/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .gemini/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .github/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .github/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .github/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .opencode/skills/h3-video-prompt-enhancer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "h3-video-prompt-enhancer" agent skill from https://github.com/benjiyaya/Calliope/tree/main/calliope-backend/skills_builtin/h3-video-prompt-enhancer into .opencode/skills/h3-video-prompt-enhancer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "h3-video-prompt-enhancer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
h3-video-prompt-enhancerA skill your agent uses when making MiniMax H3 video prompts from media + ideas.
H3 Video Prompt Enhancer is an agent skill from benjiyaya/Calliope. Use when making MiniMax H3 video prompts from media + ideas.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `references/base-multishot-format.md` and `references/creative-showcase.md`).
It sits in Media & Creative, covering AI video generation and Diffusion and image models. It works with MiniMax, ComfyUI and SvelteKit. The repository describes itself as: Local-first AI Idea-to-video studio — FastAPI + SvelteKit + ComfyUI. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0a75919. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
H3 Video Prompt Enhancer loads about 4.6k tokens when it runs, and up to ~30k if it reads all its reference files. Until then it costs about 21 tokens; SKILL.md has 2,384 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from benjiyaya/Calliope at commit 0a75919, republished under its MIT licence (© benjiyaya). 2,384 words, ~4,566 tokens.
.claude/skills/h3-video-prompt-enhancer/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Transform a user's rough video idea + attached assets into a production-grade MiniMax H3 video generation prompt. The skill handles two H3 generation modes:
references/ref2va-format.md for the full spec.references/base-multishot-format.md for the full spec.The skill's distinctive value: it doesn't just format-comply — it creatively enhances the brief with professional-grade cinematic detail (camera aesthetics, visual texture, lighting design, pacing arcs, spatial choreography, continuity tracking) before mapping everything into the exact H3 output format. Load references/creative-showcase.md for quality benchmarks and pattern examples from advanced long-form prompts.
Trigger when the user:
Don't use for:
Determine the H3 mode from what the user attached and stated:
| User provides | Mode | Format reference |
|---|---|---|
| Reference images/videos/audio (character sheets, style refs, voice clips) | Ref2VA | references/ref2va-format.md |
| Nothing — just a text idea | T2VA (text-to-video) | references/base-multishot-format.md |
| 1 image as the first frame | I2VA (image-to-video) | references/base-multishot-format.md |
| 2 images (first frame + last frame) | FL2VA (first-last-to-video) | references/base-multishot-format.md |
| 1 image as the last frame only | L2VA (last-frame-to-video) | references/base-multishot-format.md |
Key distinction: An image used as a first/last frame anchor = I2VA/FL2VA/L2VA. An image used as a character/style reference (not a frame position) = Ref2VA. When ambiguous, ask the user: "Is this image a frame anchor (first/last frame of the video) or a reference (character/style/scene)?"
In DSH you must actually see or transcribe each reference asset before you can describe it faithfully. Two requirements apply:
model ... does not declare image input, the current model is not vision-capable. Switch to an image-capable model (or delegate the visual read to a vision-capable partner) before describing reference images. Do NOT guess a character's appearance from a filename.subject_definitions.(Sx) speaker IDs. If speech transcription is unavailable in the harness, note the delivery style from any user-provided description and keep (Sx) IDs generic.Confirm these before enhancing (ask if missing, but proceed if the idea is clear enough):
<Picture k> or <Subject N>; Video k → <Video k>; Audio k → <Audio k>.This is the core value-add. Take the user's idea and enrich it across seven dimensions. The goal: produce a prompt with the depth and cinematic intelligence of a professional storyboard. Consult references/creative-showcase.md for full examples of each pattern.
1. Camera Identity — Assign a distinct camera aesthetic that matches the scene's tone:
2. Visual Texture (LOOK) — Define the image quality and color science:
3. Pacing Arc — Plan the energy progression across the full duration:
4. Character Detail — Flesh out every on-screen person with specificity:
5. Spatial Geography — For action sequences or multi-location videos:
6. Continuity Progression — Track what changes across shots so the video feels coherent:
7. Sound Design Plan — Map the full audio landscape:
non_diegetic_music field)Every shot in the storyboard must specify:
Load the appropriate format reference and produce the final H3 prompt.
For Ref2VA → Load references/ref2va-format.md. Output exactly 6 sections in order:
subject_definitions: — one line per tracked itemsummary: — task-type prefix + one paragraphretention_analysis: — fidelity markers per labeldetailed_description: — 350–500 words, opens with style, then [Shot N] timelineoverall_soundscape: — 1–4 sentencesnon_diegetic_music: — 1–3 sentences or N/AFor Base MultiShot → Load references/base-multishot-format.md. Output:
integrated_multimodal_description: — timed multi-shot timelineoverall_soundscape: — 1–4 sentencesnon_diegetic_music: — 1–3 sentences or N/A<d> tags and visible on-screen text keep their original language verbatim[Shot 1] has NO timestamp and opens with style + initial composition. Later shots: [Shot N] At MM:SS.mmm, the camera cuts to ... with strictly increasing times within duration<d>; inside <d> ONLY language tag + exact words: <d>[English] Wait for us!</d>Types: Zoom In/Out, Push In/Pull Out, Pan L/R, Truck L/R, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, Roll CW/CCW. Amplitude: "with small amplitude" / "with large amplitude" (omit when medium). Speed: "at slow speed" / "at fast speed" (omit when normal).
After generating the H3 prompt:
The user maintains advanced long-form prompt examples representing the quality bar. Load references/creative-showcase.md for full examples. Key transferable patterns:
Long-form Storytelling (montage, day-in-the-life, narrative):
Action Choreography (combat, chase, sports):
"Slop" in H3 prompts means shots that look flat, static, or lifeless in the generated video. The most common cause is too many wide shots, static camera holds, and vague action descriptions. Before finalizing any prompt, verify every shot against these rules:
| Slop | Fix |
|---|---|
| Wide establishing shot during action | Cut to close-up of feet/hands/face at action onset |
| "Camera holds on the standoff" | Camera pushes in, orbits, or whip-pans to maintain motion |
| "She ducks, grabs, pivots, and throws" in one shot | Split into: duck (close-up) → grab (hand close-up) → pivot+throw (tracking) |
| "The camera follows them" (vague) | "The camera tracks at fast speed, whip-panning between striker and target" |
| Characters standing/walking with no action | Every beat needs a physical micro-action: eyes narrowing, fingers tightening, weight shifting |
| Static aftermath shot | Even aftermath needs camera movement: pull-back, tilt-up, slow orbit |
When a prompt feels flat, identify the weakest shots and rewrite with this pattern:
Treating Ref2VA reference images as frame anchors. A character sheet or style reference is a <Subject>, not a <Picture>. Only use <Picture N> standalone when the image IS a concrete frame position (first frame, keyframe, last frame). When in doubt, cite the image inside the relevant <Subject N> line.
Cramming multiple actions into one shot. One dominant action per shot — this is a hard H3 constraint. If the brief describes sequential actions, split them across multiple shots with cuts.
Forgetting camera motion. Every shot needs an explicit camera movement specification — even "Static Shot" if the camera doesn't move. Omitting it leaves the model guessing.
Inconsistent character identity across shots. Repeat identity anchors (hair, clothing, key props) in every shot, phrased freshly but consistently. If hair gets messy in shot 3, it stays messy in shot 4.
Wrong shot count for duration. Respect the budget: 4–6s → 1–2 shots; 7–10s → 2–3 shots; 11–15s → 3–5 shots. Don't plan 5 shots for a 5-second video.
Mixing diegetic and non-diegetic music. Music audible to characters (radio, live performance, phone speaker) goes in the shot description. Background score the characters can't hear goes in non_diegetic_music. Never put the same music in both.
Inventing reference labels (Ref2VA). Never create labels beyond those defined in subject_definitions. <Subject 3> means the same thing in every section where it appears. Different indices for different modalities are numbered independently.
Translating or rewriting dialogue. Preserve the user's exact words inside <d> tags — including punctuation, hesitations, and language. Never translate to English if the user wrote in another language.
Skipping the style opener. The detailed_description (Ref2VA) or integrated_multimodal_description (Base) MUST open with 1–2 sentences of overall style BEFORE [Shot 1]. Don't jump straight into the first shot.
Flat timestamps. [Shot 1] never has a timestamp. Every subsequent shot needs At MM:SS.mmm with strictly increasing times. Convert duration to seconds with proper formatting (8s → 8.00 for the instruction line).
<d>[Language] ...</d> format[Shot 1]non_diegetic_music© benjiyaya, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in calliope-backend/skills_builtin/h3-video-prompt-enhancer of benjiyaya/Calliope.
Open the folder on GitHubat commit 0a75919
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in benjiyaya/Calliope, which our catalogue first saw on October 7, 2026.
H3 Video Prompt Enhancer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| H3 Video Prompt Enhancer this skillbenjiyaya/Calliope | 242 | — | ~4.6k | Automated safety check: Pass | MIT | |
| H3 Videoagent-next/video-agent | 120 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Open Videoagent-next/video-agent | 120 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit | 105 | — | ~12k | Automated safety check: Pass | Apache-2.0 | |
| Seedance 2 5calesthio/OpenMontage | 66k | — | ~3.2k | Automated safety check: Pass | AGPL-3.0 | |
| Minimax H3calesthio/OpenMontage | 66k | — | ~580 | Automated safety check: Pass | AGPL-3.0 |
agent-next/video-agent
OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).
agent-next/video-agent
Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).
SlavaSexton/ComfyUI-Agent-Kit
Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.
calesthio/OpenMontage
Generate 4-30 second cinematic video with ByteDance Seedance 2.5 through fal.ai, Volcengine Ark, Runway, or ComfyUI Partner Nodes.
calesthio/OpenMontage
Generate MiniMax H3 (Hailuo 3.0) video through the official MiniMax v2 API, fal.ai, Runway, ComfyUI Partner Nodes, or local open weights in ComfyUI.
JGRFW/comfyui-AICG3D
根据参考图和用户设定创作高密度、连续因果的电影级打斗视频提示词,并输出保持同一时间线的中文导演稿与 MiniMax H3 Ref2VA 英文六段稿。适用于 15 秒动作设计、武器战、徒手战、巨物战和参考图驱动的连续攻防。
benjiyaya/Calliope
A skill your agent uses when the user asks to build, pose, frame, or ANIMATE a 3D scene in Build Scene — characters, primitives, shot framing, keyframe motion, or exporting a blockout as a…
benjiyaya/Calliope
Write finished Calliope story content — beats, cast, screenplay scenes, shot clips, continuity requirements — into an existing Calliope project through the calliope-cli bridge.
benjiyaya/Calliope
A skill your agent uses when writing or improving character image prompts — sheets, portraits, and reference images that stay consistent across scenes.
benjiyaya/Calliope
A skill your agent uses when turning a Calliope scene into a video job — chaining from a previous clip, ordering reference images, or picking duration settings.
Categories
A skill your agent uses when making MiniMax H3 video prompts from media + ideas. H3 Video Prompt Enhancer is an agent skill from benjiyaya/Calliope. Use when making MiniMax H3 video prompts from media + ideas.
H3 Video Prompt Enhancer fits situations like: making MiniMax H3 video prompts from media + ideas; tasks that involve AI video generation; tasks that involve Diffusion and image models.
Run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a claude-code`. Or copy the skill folder (calliope-backend/skills_builtin/h3-video-prompt-enhancer in benjiyaya/Calliope) into .claude/skills/h3-video-prompt-enhancer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a codex`. Or copy the skill folder (calliope-backend/skills_builtin/h3-video-prompt-enhancer in benjiyaya/Calliope) into .agents/skills/h3-video-prompt-enhancer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/h3-video-prompt-enhancer, .gemini/skills/h3-video-prompt-enhancer, .github/skills/h3-video-prompt-enhancer and .opencode/skills/h3-video-prompt-enhancer in your project.
SKILL.md names no scripts, command-line tools or credentials: H3 Video Prompt Enhancer is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
H3 Video Prompt Enhancer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with H3 Video Prompt Enhancer: H3 Video (agent-next/video-agent, 120 stars), Open Video (agent-next/video-agent, 120 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars) and Seedance 2 5 (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benjiyaya (a GitHub user) maintains it in benjiyaya/Calliope, which has 242 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.
Source: benjiyaya/Calliope on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.