Agent skill

H3 Base Format

by dagthomas in dagthomas/comfyui_dagthomas

The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact…

MITAuto-check passed

Install H3 Base Format

skills CLI
$ npx skills add dagthomas/comfyui_dagthomas --skill h3-base-format -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dagthomas/comfyui_dagthomas h3-base-format --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dagthomas/comfyui_dagthomas.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data/h3/skills/h3-base-format .claude/skills/h3-base-format && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
h3-base-format
GitHub stars
290
Token cost
~856 tokens
SKILL.md length
437 words
Files
4 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact…

  • SKILL.md covers Alignment lines, How the node's images map and Reference library
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

H3 Base Format is an agent skill from dagthomas/comfyui_dagthomas. The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact first-frame, first-and-last-frame and last-frame alignment lines, how the node's image batch maps to firstframe and lastframe, and how each mode develops from or converges onto its keyframes. Load with h3-prompt-director whenever the task is not Ref2VA.

Its SKILL.md is about 860 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/06_OFFICIAL_GUIDE_CLEAN_COPY.md`, `references/11_T2VA_GOLD_EXAMPLES.md` and `references/12_KEYFRAME_GOLD_EXAMPLES.md`).

It works with MiniMax. The repository describes itself as: comfyuidagthomas - Advanced Prompt Generation and Image Analysis. The licence is MIT.

Example prompts

  • “/h3-base-format”

What it can do on your machine

Read from SKILL.md and the folder at commit 469fdd6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

H3 Base Format loads about 856 tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 437 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~856
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dagthomas/comfyui_dagthomas at commit 469fdd6, republished under its MIT licence (© dagthomas). 437 words, ~856 tokens.

Download SKILL.mdSave it as .claude/skills/h3-base-format/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
h3-base-format
description
The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integrated_multimodal_description / overall_soundscape / non_diegetic_music, the exact first-frame, first-and-last-frame and last-frame alignment lines, how the node's image batch maps to first_frame and last_frame, and how each mode develops from or converges onto its keyframes. Load with h3-prompt-director whenever the task is not Ref2VA.

H3 base format (T2VA / I2VA / FL2VA / L2VA)

This is what the APNext H3 Prompt Writer and APNext H3 Claude Code Writer nodes expect back. The node splits the text on these three labels, so they must appear exactly once, in this order, each followed by one blank line:

integrated_multimodal_description: [Shot 1] <style>, <one continuous timeline> ...

overall_soundscape: ...

non_diegetic_music: ...

The stated visual style opens [Shot 1] (the node names it, or asks you to choose one).

Alignment lines

T2VA begins directly with integrated_multimodal_description:. The other modes put one alignment line first, then exactly one blank line, then the three fields. Copy the wording character for character; only N (the number of the final shot) and S.SS (the duration, two decimals) change.

I2VA: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

FL2VA: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

L2VA: How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.

The node's own directive repeats the right line for the chosen task type; if it and this file ever differ, follow the node.

Show full SKILL.md (222 more words)Show less

How the node's images map

The node sends whatever was connected to its image socket as an ordered batch and hands the same frames back out as first_frame and last_frame for the video node:

  • I2VA - the image is <Picture 1>, the exact opening frame. Preserve its identity, clothing, colours, key objects and geography, then develop forward from it.
  • L2VA - the image is <Picture 1>, the exact final frame, and belongs to the last shot. Infer a plausible earlier state, then a transition path, then a gradual convergence, then land on the image.
  • FL2VA - frame 0 of the batch is <Picture 1> (opening), the last frame is <Picture 2> (ending). Describe the change between them and land exactly on the ending.
  • T2VA with an image attached - the image is visual context to describe from, not a keyframe. No alignment line, no <Picture N> labels.

Describe what you actually see in the frames: subject, wardrobe, colour, light direction, objects, spatial layout. Do not paraphrase the user's idea back when the picture already answers the question.

Reference library

Read thisWhen
11_T2VA_GOLD_EXAMPLES.mdTask is T2VA - match its density, shot pacing and audio balance.
12_KEYFRAME_GOLD_EXAMPLES.mdTask is I2VA, FL2VA or L2VA - alignment grammar and how to develop from or converge onto frames.
06_OFFICIAL_GUIDE_CLEAN_COPY.mdCondensed official guide, if the full guide is not already in your context.

© dagthomas, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in data/h3/skills/h3-base-format of dagthomas/comfyui_dagthomas.

  • SKILL.md
  • references/06_OFFICIAL_GUIDE_CLEAN_COPY.md
  • references/11_T2VA_GOLD_EXAMPLES.md
  • references/12_KEYFRAME_GOLD_EXAMPLES.md

Open the folder on GitHubat commit 469fdd6

Compare with similar skills

H3 Base Format next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

H3 Base Format compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
H3 Base Format this skilldagthomas/comfyui_dagthomas290—~856Automated safety check: PassMIT
Minimax DOCXpoco-ai/poco-claw1.4k7 repos~3.9kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Minimax XLSXpoco-ai/poco-claw1.4k6 repos~2.1kAutomated safety check: PassMIT
Minimax PDFpoco-ai/poco-claw1.4k6 repos~2.1kAutomated safety check: PassMIT
Evals Contextzgsm-ai/costrict4.4k1 repos~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Minimax DOCX

    poco-ai/poco-claw

    Professional DOCX document creation, editing, and formatting using OpenXML SDK (.NET).

    1.4k GitHub starsUsed in 7 repos~3.9k tokens
    Documents & OfficeAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Minimax XLSX

    poco-ai/poco-claw

    Open, create, read, analyze, edit, or validate Excel/spreadsheet files (.xlsx, .xlsm, .csv, .tsv).

    1.4k GitHub starsUsed in 6 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • Minimax PDF

    poco-ai/poco-claw

    A skill your agent uses when visual quality and design identity matter for a PDF.

    1.4k GitHub starsUsed in 6 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • Evals Context

    zgsm-ai/costrict

    Provides context about the CoStrict evals system structure in this monorepo.

    4.4k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • KrillinAI CLI Operator

    krillinai/OpenCreator

    Routes agents to the right KrillinAI command for subtitles, dubbing, video rendering, covers and speech, and explains how to read its JSON and manifest output.

    13k GitHub stars~869 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from dagthomas/comfyui_dagthomas

  • H3 Crossover

    dagthomas/comfyui_dagthomas

    The four-section MiniMax H3 crossover-scene contract the APNext H3 Crossover Writer node emits - subjectdefinitions / integratedmultimodaldescription / overallsoundscape / nondiegeticmusic per…

    290 GitHub stars~772 tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Ref2va

    dagthomas/comfyui_dagthomas

    The six-section MiniMax H3 full-reference (Ref2VA) contract the APNext H3 Reference Prompt Writer nodes emit - subjectdefinitions, summary, retentionanalysis, detaileddescription, overallsoundscape…

    290 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Prompt Director

    dagthomas/comfyui_dagthomas

    Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness)…

    290 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • H3 Style Craft

    dagthomas/comfyui_dagthomas

    Turning the APNext H3 node's visualstyle, camera and wildness settings into observable MiniMax H3 prompt language - one visual-medium pack plus at most one motion, finish and audio pack, translating…

    290 GitHub stars~994 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about H3 Base Format

What does H3 Base Format do?

The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact…. H3 Base Format is an agent skill from dagthomas/comfyui_dagthomas. The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integratedmultimodaldescription / overallsoundscape / nondiegeticmusic, the exact first-frame, first-and-last-frame and last-frame alignment lines, how the node's image batch maps to firstframe and lastframe, and how each mode develops from or converges onto its keyframes.

How do I install H3 Base Format in Claude Code?

Run `npx skills add dagthomas/comfyui_dagthomas --skill h3-base-format -a claude-code`. Or copy the skill folder (data/h3/skills/h3-base-format in dagthomas/comfyui_dagthomas) into .claude/skills/h3-base-format in your project. Claude Code loads it when a task matches its description.

How do I install H3 Base Format in Codex?

Run `npx skills add dagthomas/comfyui_dagthomas --skill h3-base-format -a codex`. Or copy the skill folder (data/h3/skills/h3-base-format in dagthomas/comfyui_dagthomas) into .agents/skills/h3-base-format in your project. Codex loads it when a task matches its description.

Can I use H3 Base Format in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dagthomas/comfyui_dagthomas --skill h3-base-format -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/h3-base-format, .gemini/skills/h3-base-format, .github/skills/h3-base-format and .opencode/skills/h3-base-format in your project.

What does H3 Base Format need to run?

SKILL.md names no scripts, command-line tools or credentials: H3 Base Format is instructions for the agent only.

Does H3 Base Format access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is H3 Base Format safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does H3 Base Format use?

H3 Base Format is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does H3 Base Format use?

About 856 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.3k tokens, read only when the agent opens those files.

What are the alternatives to H3 Base Format?

Skills that share tags, products or a category with H3 Base Format: Minimax DOCX (poco-ai/poco-claw, 1.4k stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Minimax XLSX (poco-ai/poco-claw, 1.4k stars) and Minimax PDF (poco-ai/poco-claw, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains H3 Base Format?

dagthomas (a GitHub user) maintains it in dagthomas/comfyui_dagthomas, which has 290 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on August 30, 2026.

Source: dagthomas/comfyui_dagthomas on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.