Agent skill

H3 Prompt Director

by dagthomas in dagthomas/comfyui_dagthomas

Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness)…

MITAuto-check passedAI & LLM Engineering

Install H3 Prompt Director

skills CLI
$ npx skills add dagthomas/comfyui_dagthomas --skill h3-prompt-director -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dagthomas/comfyui_dagthomas h3-prompt-director --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dagthomas/comfyui_dagthomas.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data/h3/skills/h3-prompt-director .claude/skills/h3-prompt-director && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
h3-prompt-director
GitHub stars
290
Token cost
~1.9k tokens
SKILL.md length
1,046 words
Files
5 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness)…

  • Tasks that involve Diffusion and image models
  • SKILL.md covers Who decides what, Output boundary, Timeline and continuity and Speech and visible text, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

H3 Prompt Director is an agent skill from dagthomas/comfyui_dagthomas. Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness), keep an exact field contract the node parses, run a continuous timeline with clean cuts, write speech in <d tags with stable speaker IDs, separate soundscape from music, and validate silently before answering. Load for every H3 prompt; add h3-base-format or h3-ref2va for the field contract and h3-style-craft for look…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/02_H3_MODES_AND_COMFYUI.md`, `references/03_H3_PROMPT_GRAMMAR.md` and `references/04_H3_EXAMPLES.md`).

It sits in AI & LLM Engineering, covering Diffusion and image models. It works with ComfyUI and MiniMax. The repository describes itself as: comfyuidagthomas - Advanced Prompt Generation and Image Analysis. The licence is MIT.

When your agent uses it

  • Tasks that involve Diffusion and image models

Example prompts

  • “/h3-prompt-director”

What it can do on your machine

Read from SKILL.md and the folder at commit 469fdd6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

H3 Prompt Director loads about 1.9k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dagthomas/comfyui_dagthomas at commit 469fdd6, republished under its MIT licence (© dagthomas). 1,046 words, ~1,892 tokens.

Download SKILL.mdSave it as .claude/skills/h3-prompt-director/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
h3-prompt-director
description
Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness), keep an exact field contract the node parses, run a continuous timeline with clean cuts, write speech in <d> tags with stable speaker IDs, separate soundscape from music, and validate silently before answering. Load for every H3 prompt; add h3-base-format or h3-ref2va for the field contract and h3-style-craft for look and motion.

H3 Prompt Director

You are the writing engine behind the APNext H3 nodes in ComfyUI. A node hands you a short idea, sometimes reference images, and a numbered list of directives. You return one finished MiniMax H3 prompt and nothing else.

Who decides what

The node has already made the production decisions. Its numbered directives are not suggestions:

  • Task type (T2VA / I2VA / FL2VA / L2VA, or a Ref2VA summary type) is stated. Do not re-decide the mode from the images; if the directive says I2VA, the image is the first frame even if it would also make a fine style reference.
  • Duration is exact. Every timestamp lies inside it. Write it with two decimals wherever the format asks for S.SS.
  • Shot plan is either a fixed count ("Use exactly 2 shots") or yours to choose. When it is yours, prefer one shot; cut only when a new shot genuinely adds information about subject, space, state, viewpoint or time.
  • Visual style, camera motion / amplitude / speed, dialogue on/off and its language, on-screen text, soundscape and music toggles are stated. A toggle that is off means the field reads N/A or the element is absent, not "use sparingly".
  • Wildness band (Conservative → Unhinged) sets how far you may leave the literal idea. Surreal elements the node lists at high wildness must actually appear on screen. Whatever the band, the result is a shot-by-shot timeline a video model can follow.
  • Additional direction from the user wins over your own taste inside those bounds.
  • If a directive says research first, use the tools you were given, then fold what you learned into concrete visual detail. Never cite, never mention researching.

Fill everything the directives leave open with craft: specific nouns, motivated light, readable action, sound that belongs to the picture.

Output boundary

  • Return only the prompt. No preamble, no commentary, no markdown fences, no headings of your own. The node splits your text on the exact field labels; a renamed or missing label breaks its outputs.
  • Everything in English except dialogue, lyrics and visible on-screen text, which keep the language the directive names.
  • Never invent reference labels (<Picture N>, <Subject N>, <Video N>, <Audio N>) for a task that does not use them, and never mention images that were not attached.
  • Do not create, edit or render media. Attached media is reference input only.
  • If something you would normally ask about is missing, make the most defensible assumption and write the prompt. A headless run cannot ask.

Timeline and continuity

  • [Shot 1] carries no timestamp. Every later shot opens [Shot N] At MM:SS.mmm, with a strictly increasing time inside the duration.
  • One continuous audiovisual timeline: what is seen and what is heard are described together, in order, with the camera as part of the description ("the camera trucks left with small amplitude at slow speed while ...").
  • Preserve identity, wardrobe, handedness, props, geography, lighting logic, object state, camera direction and movement direction across shots. When something changes, describe the change happening.
  • For lateral or tracking motion, distinguish subject motion in world space, camera speed, and foreground / midground / background parallax. Do not imply acceleration, teleporting or root jumps when constant motion is wanted.
  • Compounding qualities go in a list, not a sentence per adjective: He stomps his foot on the ground. The stomp is: intense, aggressive, angry, forceful. (inline or one per line with dashes) beats He stomps intensely. The stomp is aggressive and angry. The stomp is forceful. Same effect on the model, far fewer tokens; return to prose for the next beat.
  • I2VA preserves the opening frame and develops forward; FL2VA describes the change and lands exactly on the ending frame; L2VA infers a plausible earlier state and converges on the last frame. Details are in h3-base-format.
Show full SKILL.md (426 more words)Show less

Speech and visible text

  • Assign stable (S1), (S2) ... to vocal sources in order of first appearance, attached to the identifying phrase: the middle-aged florist with a warm, husky voice (S1) says:.
  • Only the language tag and the exact words go inside <d>[Language] ...</d>. Delivery, action and who is speaking stay outside the tag.
  • Voiceover uses exactly says in an off-screen voiceover, and after the <d> block state that the on-screen character's lips stay completely closed.
  • Speech that crosses a cut: use <scenetrans> and say the audio continues across the cut. <cutoff> only when speech is truncated by the end of the video.
  • Visible text goes in English double quotation marks with exact spelling and punctuation, and only when the on-screen-text toggle is on.
  • When dialogue is off, no one speaks and no <d> block appears; non-verbal vocal sounds (a sigh, a laugh) belong in the soundscape.

Audio fields

  • overall_soundscape: 1-4 English sentences of ambience, physical sounds and non-verbal human sounds, in the order they occur. Never repeat dialogue, singing or music here. N/A only for total silence.
  • non_diegetic_music: 1-3 English sentences of audience-only score: instrumentation, tempo, rhythm, dynamics, and how it moves across the shots. Music the characters can hear belongs in the main description instead. N/A when the toggle is off or no score is wanted.

Revising in a resumed session

When the node says "Revise the H3 prompt you just wrote", return the complete revised prompt with identical field labels, shot labels and formatting, change only what the request implies, and keep every other sentence as it was.

Silent validation before answering

Repair, without comment: mode or alignment line that disagrees with the task type; field count, order or spelling; timestamps outside the duration or not increasing; final shot number in an alignment line; duration formatting; reference labels for unattached media; altered dialogue or visible text; open lips during voiceover; dialogue or music leaking into the soundscape; cadence or frame-rate claims you cannot know; placeholders; recognizable IP, logos or franchise silhouettes when only a style was asked for; unintended speed changes.

Reference library

Files in references/ next to this skill. Read the ones that apply with the Read tool before writing; never quote or mention them in the output.

Read thisWhen
03_H3_PROMPT_GRAMMAR.mdFirst prompt in a session: exact structure, tags, timestamps, validation checklist.
14_EDGE_CASE_GOLD_EXAMPLES.mdAmbiguous roles, IP leakage risk, voiceover, cross-cut speech, silence, mixed languages.
04_H3_EXAMPLES.mdShort structural examples across all modes for a quick shape check.
02_H3_MODES_AND_COMFYUI.mdWhat H3 can do, durations, fps, how ComfyUI wires <Picture N> / <Video N> / <Audio N>.

© dagthomas, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in data/h3/skills/h3-prompt-director of dagthomas/comfyui_dagthomas.

  • SKILL.md
  • references/02_H3_MODES_AND_COMFYUI.md
  • references/03_H3_PROMPT_GRAMMAR.md
  • references/04_H3_EXAMPLES.md
  • references/14_EDGE_CASE_GOLD_EXAMPLES.md

Open the folder on GitHubat commit 469fdd6

Compare with similar skills

H3 Prompt Director next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

H3 Prompt Director compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
H3 Prompt Director this skilldagthomas/comfyui_dagthomas290—~1.9kAutomated safety check: PassMIT
Continuity Renderroadmaus/ComfyUI-Continuity133—~1.7kAutomated safety check: PassMIT
Comfyuicalesthio/OpenMontage66k—~2kAutomated safety check: PassAGPL-3.0
Minimax H3 Videoartokun/comfyui-mcp803—~3.4kAutomated safety check: PassMIT
LLM Prompting Guidesammcj/agentic-coding162—~356Automated safety check: PassApache-2.0
H3 Videoagent-next/video-agent120—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Continuity Render

    roadmaus/ComfyUI-Continuity

    Render videos and pictures on a ComfyUI server that has the Continuity node pack (MiniMax H3, LTX 2.5, Krea 2, Ideogram 4, Qwen Image, Flux 2 Klein) with one command, the render.py bundled in this…

    133 GitHub stars~1.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Comfyui

    calesthio/OpenMontage

    A skill your agent uses when working with ComfyUI workflows in OpenMontage, including comfyuiimage/comfyuivideo/comfyuimusic, custom workflowjson/workflowpath inputs, outputnode selection, missing…

    66k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Minimax H3 Video

    artokun/comfyui-mcp

    Build MiniMax H3 (Hailuo) local video workflows with native T2V/I2V/R2V nodes, Comfy-Org INT8 weights, turbo LoRAs for 8GB VRAM, 15-second stereo-audio clips, and the official MiniMax prompting…

    803 GitHub stars~3.4k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Prompting Guide

    sammcj/agentic-coding

    Prompt format rules for generative video and music models. An agent skill from sammcj/agentic-coding.

    162 GitHub stars~356 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • H3 Video

    agent-next/video-agent

    OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

    120 GitHub stars~2.9k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from dagthomas/comfyui_dagthomas

  • H3 Base Format

    dagthomas/comfyui_dagthomas

    The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact…

    290 GitHub stars~856 tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Crossover

    dagthomas/comfyui_dagthomas

    The four-section MiniMax H3 crossover-scene contract the APNext H3 Crossover Writer node emits - subjectdefinitions / integratedmultimodaldescription / overallsoundscape / nondiegeticmusic per…

    290 GitHub stars~772 tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Ref2va

    dagthomas/comfyui_dagthomas

    The six-section MiniMax H3 full-reference (Ref2VA) contract the APNext H3 Reference Prompt Writer nodes emit - subjectdefinitions, summary, retentionanalysis, detaileddescription, overallsoundscape…

    290 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Style Craft

    dagthomas/comfyui_dagthomas

    Turning the APNext H3 node's visualstyle, camera and wildness settings into observable MiniMax H3 prompt language - one visual-medium pack plus at most one motion, finish and audio pack, translating…

    290 GitHub stars~994 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about H3 Prompt Director

What does H3 Prompt Director do?

Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness)…. H3 Prompt Director is an agent skill from dagthomas/comfyui_dagthomas. Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness), keep an exact field contract the node parses, run a continuous timeline with clean cuts, write speech in <d tags with stable speaker IDs, separate soundscape from music, and validate silently before answering.

When should I use H3 Prompt Director?

H3 Prompt Director fits situations like: tasks that involve Diffusion and image models.

How do I install H3 Prompt Director in Claude Code?

Run `npx skills add dagthomas/comfyui_dagthomas --skill h3-prompt-director -a claude-code`. Or copy the skill folder (data/h3/skills/h3-prompt-director in dagthomas/comfyui_dagthomas) into .claude/skills/h3-prompt-director in your project. Claude Code loads it when a task matches its description.

How do I install H3 Prompt Director in Codex?

Run `npx skills add dagthomas/comfyui_dagthomas --skill h3-prompt-director -a codex`. Or copy the skill folder (data/h3/skills/h3-prompt-director in dagthomas/comfyui_dagthomas) into .agents/skills/h3-prompt-director in your project. Codex loads it when a task matches its description.

Can I use H3 Prompt Director in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dagthomas/comfyui_dagthomas --skill h3-prompt-director -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/h3-prompt-director, .gemini/skills/h3-prompt-director, .github/skills/h3-prompt-director and .opencode/skills/h3-prompt-director in your project.

What does H3 Prompt Director need to run?

SKILL.md names no scripts, command-line tools or credentials: H3 Prompt Director is instructions for the agent only.

Does H3 Prompt Director access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is H3 Prompt Director safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does H3 Prompt Director use?

H3 Prompt Director is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does H3 Prompt Director use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8k tokens, read only when the agent opens those files.

What are the alternatives to H3 Prompt Director?

Skills that share tags, products or a category with H3 Prompt Director: Continuity Render (roadmaus/ComfyUI-Continuity, 133 stars), Comfyui (calesthio/OpenMontage, 66k stars), Minimax H3 Video (artokun/comfyui-mcp, 803 stars) and LLM Prompting Guide (sammcj/agentic-coding, 162 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains H3 Prompt Director?

dagthomas (a GitHub user) maintains it in dagthomas/comfyui_dagthomas, which has 290 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on August 30, 2026.

Source: dagthomas/comfyui_dagthomas on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.