Agent skill

MiniMax H3 Video Prompting

by SlavaSexton in SlavaSexton/ComfyUI-Agent-Kit

Guides prompting and local ComfyUI use of MiniMax H3 (Hailuo 3), which generates video with synchronized audio, including prompt format, quants and troubleshooting.

Apache-2.0Auto-check passedMedia & Creative

Install MiniMax H3 Video Prompting

skills CLI
$ npx skills add SlavaSexton/ComfyUI-Agent-Kit --skill minimax-h3 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SlavaSexton/ComfyUI-Agent-Kit minimax-h3 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SlavaSexton/ComfyUI-Agent-Kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/shared/minimax-h3 .claude/skills/minimax-h3 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-h3
GitHub stars
105
Token cost
~2.5k tokens
SKILL.md length
1,188 words
Files
2
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides prompting and local ComfyUI use of MiniMax H3 (Hailuo 3), which generates video with synchronized audio, including prompt format, quants and troubleshooting.

  • Writing prompts for MiniMax H3 text-to-video with audio
  • SKILL.md covers The prompt format is not…, Dialogue: the single most…, Camera is a controlled… and Reference-to-video: label…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Running the open H3 weights locally in ComfyUI

What it does

H3 generates video and stereo audio together from text, images, reference video or reference audio, in any mix. The skill separates two paths that are easy to confuse: the hosted API through partner nodes, with 2K output and per-second pricing, and the open weights run locally through core nodes such as MiniMaxH3ImageToVideo and MiniMaxH3ReferenceToVideo, at 768p and free. Most of the guidance is about the local path.

Local prompts are not free-form prose. The base model expects the output shape of a hosted refiner that is not in the open release, so you write an optional instruction line followed by integrated_multimodal_description, overall_soundscape and non_diegetic_music fields, and image modes need a fixed first line. A reference.md file covers weights, quant sizes and acceleration packs. Listed fixes include gibberish speech, identity drift, garbled audio after a latent upscale and builds that refuse to run.

When your agent uses it

  • Writing prompts for MiniMax H3 text-to-video with audio
  • Running the open H3 weights locally in ComfyUI
  • Choosing a quant or acceleration LoRA that fits your VRAM
  • Fixing gibberish speech, identity drift or garbled audio in a generated clip

Example prompts

  • “Write an H3 prompt for a five-second shot of a bicycle mechanic talking while fixing a wheel.”
  • “Pick an H3 quant and acceleration LoRA for my graphics card with limited VRAM.”
  • “My H3 clip has gibberish speech; rework the prompt so the dialogue is clear.”
  • “Wire a reference-to-video graph in ComfyUI with one image and one audio clip.”

Requirements

  • ComfyUI with the MiniMax H3 core nodes for local runs
  • Enough VRAM for the chosen quant of the open weights

What it can do on your machine

Read from SKILL.md and the folder at commit 74f5b0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

MiniMax H3 Video Prompting loads about 2.5k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 1,188 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SlavaSexton/ComfyUI-Agent-Kit at commit 74f5b0b, republished under its Apache-2.0 licence (© SlavaSexton). 1,188 words, ~2,466 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-h3/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
minimax-h3
description
Use when writing or debugging prompts for MiniMax H3 (Hailuo 3) video-with-audio generation, running the open weights locally in ComfyUI, choosing a quant or an acceleration LoRA for the VRAM you have, wiring reference-to-video with images, video or audio, or when a generated clip produces gibberish speech, drifts off a reference identity, garbles audio after a latent upscale, or refuses to run on a build that looks current.

MiniMax H3 (Hailuo 3)

H3 generates video and synchronised stereo audio jointly, from text, images, reference video, reference audio, or any mix. It is the same model family whether you hit the hosted API or run the open weights, but the two paths are wired completely differently and only one of them is free.

Two paths, do not confuse them.

  • Hosted API, partner nodes MinimaxHailuo03TextToVideoNode / ...FirstLastFrameNode / ...ReferenceNode, category partner/video/MiniMax. 2K output, priced per second, no local weights.
  • Local open weights, core nodes MiniMaxH3ImageToVideo and MiniMaxH3ReferenceToVideo from comfy_extras/nodes_minimax_h3.py. 768p, free, and everything below is about this path.

Who owns what, so you open one file, not three. THIS file owns the prompt format and the operating rules. reference.md next to it owns weights, quant sizes, acceleration packs and their wiring. The MiniMax entry in the kit's MODELS.md owns the node-level graph (every node and socket) and the licence. When they disagree, the node code wins and the discrepancy is a bug worth reporting.

The prompt format is not free-form prose

H3-Base consumes the output of a hosted prompt refiner (H3-Context-IR) that is not in the open release. So locally you write in the refiner's output shape yourself. From MiniMax's own VIDEO_PROMPT_WRITING_GUIDE_base_en.md, that is an optional instruction line, a blank line, then three fields:

integrated_multimodal_description: [Shot 1] ... [Shot 2] At 00:04.500, ...

overall_soundscape: ...

non_diegetic_music: ...
  • integrated_multimodal_description carries visuals, action, shots, speakers, dialogue and diegetic sound along the timeline. overall_soundscape sums ambience and physical-action sound. non_diegetic_music is score the characters cannot hear.
  • Image modes need a fixed first line. I2V: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. First-and-last-frame uses the alignment sentence naming both pictures and the second each lands on, to two decimals.
A complete prompt, end to end

Nothing above is usable until you have seen one whole. This is a 5 s text-to-video brief in the official shape:

integrated_multimodal_description: [Shot 1] Cinematic medium-wide shot, Push In slowly. A bicycle mechanic in
a navy work coat lowers a metal shutter in a narrow workshop at dusk; warm tungsten light spills across
scattered tools and rain-dark pavement outside. He pauses, looks toward the street. At 00:03.200 he switches
off the bench lamp and the frame drops to ambient blue. The mechanic (weathered voice, mid-fifties, speaks
English only) says quietly: <d>[English] That's enough for today.</d>

overall_soundscape: Steady rain on a metal awning, the rolling clatter of the shutter, one soft click of the
lamp switch, distant tyres on wet asphalt. No music from within the scene.

non_diegetic_music: Sparse solo piano, slow, minor key, entering after the shutter closes and fading to
silence on the lamp click.

For image-to-video the same block is preceded by the fixed line and one blank line:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] ...

Note what is doing the work: one physical action the camera and the sound can both follow, a camera move from the fixed vocabulary, the spoken line wrapped in <d> so the words are exact, and the two audio fields kept separate so diegetic sound and score do not fight.

Dialogue: the single most common cause of "the speech is gibberish"

Speakers get stable IDs (S1), (S2), joint (S1,S2). The words go inside <d> with a language tag, and everything about who says it and how stays outside. Copy the line verbatim, do not paraphrase or translate:

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

Without <d> the model is never told the exact words and improvises phonetics. Voiceover needs the exact phrase says in an off-screen voiceover plus a statement that the lips stay closed. Use <scenetrans> when a line crosses a cut, <cutoff> when speech is truncated by the end. On-screen text goes in double quotes, verbatim.

Camera is a controlled vocabulary

Motion type plus amplitude plus speed, and medium amplitude at normal speed is the default you simply omit: Zoom In/Out, Push In/Pull Out, Pan Left/Right, Truck Left/Right, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot.

Reference-to-video: label every file with a job

<Subject N> is the one that does the real work, because it binds sources: "<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>." Then <Picture N> is a frame or composition anchor, <Video N> an editing or temporal source, <Audio N> a copied signal.

Reference-to-video has its OWN output contract, and it is six sections, not the three above. MiniMax ships two prompt guides, and until 2026-08-09 this skill knew only the first. VIDEO_PROMPT_WRITING_GUIDE_base_en.md governs T2VA / I2VA / FL2VA / L2VA and gives the three fields at the top of this file. VIDEO_PROMPT_WRITING_GUIDE_ref_en.md governs full-reference mode and opens: "A complete rewrite output consists of six sections in the following order":

subject_definitions: <Subject 1> is ... whose appearance comes from <Picture 1> ...
summary: ...
retention_analysis: ...
detailed_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...

subject_definitions declares the references and their labels; summary states the task type, the target video and the main relationships to the references; retention_analysis says how each reference is retained, transferred or reused; detailed_description carries visuals, action, shots, sound and dialogue in playback order; the last two match their base-mode meanings. The order is fixed, and the field is retention_analysis, not retention. Note this is six named sections of prose in one string, not a JSON array, whatever a third-party node's docs call it. Confirmed by reading the guide on MiniMaxAI/MiniMax-H3 (23 553 bytes, 2026-08-09); the core node itself validates none of this, comfy_extras/nodes_minimax_h3.py takes a plain multiline string, so nothing will tell you when you get it wrong except the result.

Limits from the official card: 9 images, 3 video clips, 3 audio clips, 12 files total, clips 2 to 15 s each and 15 s total, and audio can never be the only reference. The local MiniMaxH3ReferenceToVideo node has four Autogrow families reaching those ceilings: ref_images (max 9), ref_videos (3), ref_video_audios (3, the soundtrack of the same-numbered reference video) and ref_audios (3, standalone). A template showing three image sockets is not the limit.

Show full SKILL.md (354 more words)Show less

Ten production jobs to write against

Krea's 2026-08-05 guide frames H3 prompts as production paperwork, one job per clip: director's single-shot brief · timed three-beat teaser · first-to-last-frame passage · influencer identity-and-voice lock · reference-motion performance · native-audio product reveal · brand-title reveal with required text and negatives · protected-frame object swap · UI walkthrough · omni-reference director brief. Pick the job first, then decide which control has to survive: camera, timing, identity, motion source, or the sound event line.

Note their examples use a [0-3s] beat style, which is a readable shorthand rather than MiniMax's own [Shot N] plus At 00:04.500 convention. Both work; the official form is the safer default locally.

Operating facts that bite

  • Frame count sits on a grid. length must leave remainder 5 modulo 17 (124 frames = ~5 s, 73 = ~3 s, 362 = ~15 s). The node calls it the "17k+5 grid" and the trained range is ~124 to 362 frames.
  • Native size is ~1 MP, template default 1344 x 768 at 24 fps. Higher costs time and VRAM without more real detail; the official 2K route is a hosted regenerate pass that is not open.
  • Duration: trained and tested 5 to 15 s. Longer runs but is untrained territory.
  • Speed: the open release ships full attention only; sparse attention is promised later. This is why it is slow, and why the community acceleration below matters.
  • Licence: open weights, not open source, and the territory clause is unusually strict. See MODELS.md.

When it goes wrong

SymptomMost likely cause
Speech is fluent-sounding nonsenseThe line is not inside <d>[Language] ... </d>
Face drifts across the clipNo <Subject N> binding the identity source, or refs at the wrong scale after an upscale
Audio garbles after a latent upscaleaudio_denoise left at its default 1.0 (inferred cause, not measured); audio settles late, so run more of the schedule in pass 1
Wrong face animated in a crowdReference sizing set to match when identity needed max
Acceleration node refuses to loadThe build is older than the pack requires; a tagged release is not automatically new enough
Output length is not what you askedlength fell off the 17k+5 grid

© SlavaSexton, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in shared/minimax-h3 of SlavaSexton/ComfyUI-Agent-Kit.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 74f5b0b

Compare with similar skills

MiniMax H3 Video Prompting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

MiniMax H3 Video Prompting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
MiniMax H3 Video Prompting this skillSlavaSexton/ComfyUI-Agent-Kit105—~2.5kAutomated safety check: PassApache-2.0
Hong Kong Comic Fighter for H3karuvanan/MiniMax-H3-Director-Cut-Studio132—~4.4kAutomated safety check: PassCustom licence
VRGDG H3 Short Film Pipelinevrgamegirl19/comfyui-vrgamedevgirl765—~4.2kAutomated safety check: PassCustom licence
MiniMax H3 Video DirectorTFboy1/oh-my-minimaxh3-director143—~2.4kAutomated safety check: PassMIT
H3 Videoagent-next/video-agent120—~2.9kAutomated safety check: PassApache-2.0
Open Videoagent-next/video-agent120—~3.2kAutomated safety check: PassApache-2.0

Similar skills

  • Hong Kong Comic Fighter for H3

    karuvanan/MiniMax-H3-Director-Cut-Studio

    Special skill that turns loaded Hong Kong comic panels into photoreal MiniMax H3 martial-arts sequences with readable attack and defence action.

    132 GitHub stars~4.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • VRGDG H3 Short Film Pipeline

    vrgamegirl19/comfyui-vrgamedevgirl

    Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.

    765 GitHub stars~4.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • MiniMax H3 Video Director

    TFboy1/oh-my-minimaxh3-director

    Turns a script into a storyboard, assigns MiniMax H3 workflows in ComfyUI, monitors batch generation and builds a Jianying draft of the finished video.

    143 GitHub stars~2.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • H3 Video

    agent-next/video-agent

    OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

    120 GitHub stars~2.9k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Open Video

    agent-next/video-agent

    Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).

    120 GitHub stars~3.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • H3 Video Prompt Enhancer

    benjiyaya/Calliope

    A skill your agent uses when making MiniMax H3 video prompts from media + ideas.

    242 GitHub stars~4.6k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from SlavaSexton/ComfyUI-Agent-Kit

  • Krea in ComfyUI

    SlavaSexton/ComfyUI-Agent-Kit

    Helps choose and wire Krea models in ComfyUI: the hosted Krea 2 API nodes versus local open weights, FLUX.1 Krea Dev, and add-on packs for ControlNet and editing.

    105 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Auto-check passed
  • Seedance Video Prompting

    SlavaSexton/ComfyUI-Agent-Kit

    Helps write and debug prompts for ByteDance Seedance video models, with labelled multimodal references and fixes for drift, subtitles and duplicated characters.

    105 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about MiniMax H3 Video Prompting

What does MiniMax H3 Video Prompting do?

Guides prompting and local ComfyUI use of MiniMax H3 (Hailuo 3), which generates video with synchronized audio, including prompt format, quants and troubleshooting. H3 generates video and stereo audio together from text, images, reference video or reference audio, in any mix. The skill separates two paths that are easy to confuse: the hosted API through partner nodes, with 2K output and per-second pricing, and the open weights run locally through core nodes such as MiniMaxH3ImageToVideo and MiniMaxH3ReferenceToVideo, at 768p and free.

When should I use MiniMax H3 Video Prompting?

MiniMax H3 Video Prompting fits situations like: writing prompts for MiniMax H3 text-to-video with audio; running the open H3 weights locally in ComfyUI; choosing a quant or acceleration LoRA that fits your VRAM; fixing gibberish speech, identity drift or garbled audio in a generated clip.

How do I install MiniMax H3 Video Prompting in Claude Code?

Run `npx skills add SlavaSexton/ComfyUI-Agent-Kit --skill minimax-h3 -a claude-code`. Or copy the skill folder (shared/minimax-h3 in SlavaSexton/ComfyUI-Agent-Kit) into .claude/skills/minimax-h3 in your project. Claude Code loads it when a task matches its description.

How do I install MiniMax H3 Video Prompting in Codex?

Run `npx skills add SlavaSexton/ComfyUI-Agent-Kit --skill minimax-h3 -a codex`. Or copy the skill folder (shared/minimax-h3 in SlavaSexton/ComfyUI-Agent-Kit) into .agents/skills/minimax-h3 in your project. Codex loads it when a task matches its description.

Can I use MiniMax H3 Video Prompting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SlavaSexton/ComfyUI-Agent-Kit --skill minimax-h3 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-h3, .gemini/skills/minimax-h3, .github/skills/minimax-h3 and .opencode/skills/minimax-h3 in your project.

What does MiniMax H3 Video Prompting need to run?

SKILL.md names no scripts, command-line tools or credentials: MiniMax H3 Video Prompting is instructions for the agent only. Our summary lists: ComfyUI with the MiniMax H3 core nodes for local runs; Enough VRAM for the chosen quant of the open weights.

Does MiniMax H3 Video Prompting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is MiniMax H3 Video Prompting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does MiniMax H3 Video Prompting use?

MiniMax H3 Video Prompting is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does MiniMax H3 Video Prompting use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to MiniMax H3 Video Prompting?

Skills that share tags, products or a category with MiniMax H3 Video Prompting: Hong Kong Comic Fighter for H3 (karuvanan/MiniMax-H3-Director-Cut-Studio, 132 stars), VRGDG H3 Short Film Pipeline (vrgamegirl19/comfyui-vrgamedevgirl, 765 stars), MiniMax H3 Video Director (TFboy1/oh-my-minimaxh3-director, 143 stars) and H3 Video (agent-next/video-agent, 120 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains MiniMax H3 Video Prompting?

SlavaSexton (a GitHub user) maintains it in SlavaSexton/ComfyUI-Agent-Kit, which has 105 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 3, 2026.

Source: SlavaSexton/ComfyUI-Agent-Kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.