Agent skill

Minimax H3 Reference Video Prompt

by unknowlei in unknowlei/minimax-h3-opencode-skills

Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role.

MITAuto-check passedMedia & Creative

Install Minimax H3 Reference Video Prompt

skills CLI
$ npx skills add unknowlei/minimax-h3-opencode-skills --skill minimax-h3-reference-video-prompt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install unknowlei/minimax-h3-opencode-skills minimax-h3-reference-video-prompt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/unknowlei/minimax-h3-opencode-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/minimax-h3-reference-video-prompt .claude/skills/minimax-h3-reference-video-prompt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-h3-reference-video-prompt
GitHub stars
122
Token cost
~3k tokens
SKILL.md length
1,545 words
Files
3 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role.

  • Works in 8 steps: Inventory every supplied image, video,… → Map what each asset contributes:… → Distinguish reusable visible content… → …
  • Tasks that involve Video production
  • SKILL.md covers Routing Contract, Official Format Authority, Confirmed Multishot Handoff and Official Input Envelope, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Minimax H3 Reference Video Prompt is an agent skill from unknowlei/minimax-h3-opencode-skills. Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role. Use the official six-section full-reference format for character/person/object consistency, scene/style/action/camera/storyboard/voice/audio reference, source-video editing or continuation, ambiguous image roles, ordinary image animation, and prompt-level first/key/last-frame anchoring.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/reference-rules.md`).

It sits in Media & Creative, covering Video production and Comics and storyboards. It works with MiniMax. The repository describes itself as: OpenCode skill suite for MiniMax H3 directing, routing, multishot planning, prompt generation, and review. The licence is MIT.

When your agent uses it

  • Tasks that involve Video production
  • Tasks that involve Comics and storyboards

Example prompts

  • “/minimax-h3-reference-video-prompt”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Inventory every supplied image, video, audio asset, upload order, and text requirement; validate the official input envelope, then run the…
  2. Map what each asset contributes: identity, appearance, clothing, object, environment, style, pose, action, camera, storyboard, edit…
  3. Distinguish reusable visible content from source files. Assign stable , , , and labels only where their defined roles apply.
  4. Determine all applicable task types: keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio…
  5. Determine the retention relationship for every label.
  6. Build the shot and sound timeline in playback order. Prove where each important reference first appears or takes effect.
  7. Validate label consistency, source provenance, temporal feasibility, speaker identity, audio continuity, and any keyframe landing.
  8. Return the required bilingual output.

What it can do on your machine

Read from SKILL.md and the folder at commit 2e2096f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Minimax H3 Reference Video Prompt loads about 3k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 1,545 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from unknowlei/minimax-h3-opencode-skills at commit 2e2096f, republished under its MIT licence (© unknowlei). 1,545 words, ~3,031 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-h3-reference-video-prompt/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
minimax-h3-reference-video-prompt
description
Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role. Use the official six-section full-reference format for character/person/object consistency, scene/style/action/camera/storyboard/voice/audio reference, source-video editing or continuation, ambiguous image roles, ordinary image animation, and prompt-level first/key/last-frame anchoring.

MiniMax H3 Full-Reference Video Prompt

Convert mixed reference assets and user intent into a traceable six-section full-reference prompt. Make every asset role explicit and write the target video as a detailed audiovisual timeline.

Routing Contract

New MiniMax H3 requests should enter through minimax-h3-creative-director. This is the default specialist for any request containing images. Route away only when the user explicitly declares pure first/last boundary roles and no image also carries identity, character, person, object, costume, scene, style, action, camera, composition, or another reusable reference responsibility. Ambiguity stays here. Full-reference may use prompt-level first/key/last-frame anchoring while maintaining consistency.

Official Format Authority

Before drafting, read ../h3-prompt-writing/SKILL.md and ../h3-prompt-writing/references/ref-en.txt. Consult ../h3-prompt-writing/references/base-en.txt for the shared shot, camera, dialogue, visible-text, and sound rules when needed. Treat these official MiniMax files as the canonical prompt-format specification. If this skill or its local references conflict with them, follow the official files.

Read references/reference-rules.md before drafting.

Confirmed Multishot Handoff

When the director supplies a confirmed multishot_plan, treat its shot count, timing, content, framing, performance, camera, transitions, sound, active references, and continuity decisions as already answered. Do not ask those questions again. Map every confirmed shot into detailed_description, keep reference labels and responsibilities stable, and carry the continuity ledger into retention and preservation instructions. Reopen a decision only when the plan is physically impossible, conflicts with a source asset, or exceeds the effective duration.

Official Input Envelope

For the official online product/API, keep output within 4-15 seconds at 24 FPS and the prompt within 7000 characters. Accept at most 9 images, 3 videos, 3 audio files, and 12 mixed files total. Each video or audio file must be 2-15 seconds; combined video duration and combined audio duration must each be at most 15 seconds. Audio cannot be the only reference type. Respect the documented per-file and API-body limits. Treat local ComfyUI constraints separately when they are narrower.

Mandatory Format

Always return the official six sections in their required order. Never return a free-form natural-language paragraph, keyword list, abbreviated prompt, base three-field prompt, or alternate schema as the final English prompt for a full-reference task.

Interactive Direction Check

Before assigning labels or drafting, use the host's structured-choice tool whenever the brief is sparse; gives only an image plus a generic motion request; omits two or more of action progression, scene treatment, preservation priorities, visual style, camera/editing, dialogue/voice, sound/music, or endpoint; leaves a decisive reference responsibility unclear; or asks the AI to improvise.

In OpenCode, call the built-in question tool rather than displaying Markdown options. In Codex, use the available structured user-input tool; if unavailable, ask one concise plain-text question.

Ask as many related questions as the current decision stage requires; normally 1-5, but three is not a per-call, per-turn, or per-session ceiling. A strict yes/no question may contain exactly two options; every other choice question must contain at least five materially different, feasible options. Put the context-specific recommendation first and append (Recommended) to its label. Explain the consequence of each choice, omit Other, and use multiple selection only when roles can truly be combined. Before every question tool call, count the options in the actual payload. If any non-binary choice has fewer than five, do not submit it; expand it with meaningful alternatives or make it open-ended. Do not create near-duplicates merely to reach five.

Prioritize:

  • Task relationship: reference generation, source-video editing, continuation, or a combination
  • Asset responsibility: identity/appearance, clothing/scene, action/camera, storyboard/keyframe, audio/voice/music
  • Fidelity policy: strict preservation, balanced adaptation, attribute transfer, broad stylistic reference
  • Prompt-level picture anchor: no concrete anchor, first frame, intermediate keyframe, last frame
  • Audio policy: direct reuse, partial reuse, timbre/style reference, newly designed sound

For sparse or delegated briefs, ask enough high-impact questions in one batch to establish the current stage, normally 3-5, even when the asset itself is visually clear. Do not ask the user to classify every asset; default images to character/object/scene consistency and ask about the desired creative outcome. Skip questions only when the user explicitly prohibits them. A request to improvise still requires questions.

Workflow

  1. Inventory every supplied image, video, audio asset, upload order, and text requirement; validate the official input envelope, then run the interactive direction check when required.
  2. Map what each asset contributes: identity, appearance, clothing, object, environment, style, pose, action, camera, storyboard, edit source, continuation point, voice, sound, music, or rhythm.
  3. Distinguish reusable visible content from source files. Assign stable <Subject N>, <Picture N>, <Video N>, and <Audio N> labels only where their defined roles apply.
  4. Determine all applicable task types: keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference.
  5. Determine the retention relationship for every label.
  6. Build the shot and sound timeline in playback order. Prove where each important reference first appears or takes effect.
  7. Validate label consistency, source provenance, temporal feasibility, speaker identity, audio continuity, and any keyframe landing.
  8. Return the required bilingual output.

Prompt-Level Keyframe Anchoring

Allow a full-reference task to designate a referenced image as a concrete first frame, keyframe, edited keyframe, last frame, or composition anchor. Define it as a standalone <Picture N> and state its exact role, for example:

text
<Picture 2> is the last frame of [Shot 3], defining the final pose, object placement, camera angle, lighting, and composition.

Combine keyframe completion with other task types in summary when appropriate. In detailed_description, describe a continuous path into or out of the anchored picture rather than repeating static image descriptions.

When a picture both anchors a boundary frame and preserves a person, character, object, costume, scene, style, or composition, keep the task in full-reference. Define both responsibilities explicitly instead of moving to the pure keyframe specialist.

Observe the current API distinction: prompt-level keyframe semantics are supported by the full-reference format, but an API request must not mix first_frame/last_frame roles with any reference_* role. When mixed reference assets are required, pass images as reference_image and express the concrete frame relationship in the prompt. Do not claim that the API role itself is a hard keyframe role.

Show full SKILL.md (573 more words)Show less

Required Prompt Structure

Produce exactly these six sections in order:

text
subject_definitions:

summary:

retention_analysis:

detailed_description:

overall_soundscape:

non_diegetic_music:

Write all six sections in English, preserving the original language only for dialogue, lyrics, and text visibly present in the scene.

Drafting Standard

  • Define content units, not every source file mechanically.
  • Assign every important uploaded asset an explicit role, including whether audio is reused fully, reused by track/time segment, or referenced only for timbre/style.
  • Describe each shot's current composition, subject appearance and position, environment, lighting, actions, state changes, camera movement, sound, and active reference relationship.
  • Keep reference definitions stable across all six sections.
  • Do not reduce detailed_description to a plot summary or a list of reference links.
  • Scale detail to the task's complexity while keeping the executable English prompt within the official 7000-character maximum.
  • Never invent dialogue, lyrics, visible text, or unintelligible source words. Use [unclear] for unintelligible spans.
  • Preserve exact transcripts or lyrics when speech, singing, or source audio must remain. Fit dialogue to shot duration and state cross-cut continuation explicitly.
  • For source-video editing, maintain a preservation ledger for all content that must remain unchanged.

Output Contract

Return:

  1. ### English Prompt
  2. One text code block containing only the complete six-section English prompt
  3. ### 中文翻译
  4. A complete Chinese translation outside code blocks

Language Policy

Write all six sections in English except exact literal content that must remain in its source form:

  • Dialogue, voiceover, speech, singing, and lyrics inside <d>[Language] ...</d>
  • Text visibly present in the scene, including signs, captions, subtitles, labels, interfaces, titles, logos, and messages
  • Exact personal names, place names, brands, organizations, product/work titles, account names, and handles
  • Filenames, URLs, asset names, model names, reference labels, and control tokens

Write subject definitions, reference relationships, visual description, actions, camera, editing, sound design, and music in English. Do not retain non-English prose merely because the user's request was written in that language.

After completing all six English sections, translate every prompt component into Chinese, including the meaning of dialogue, lyrics, visible text, task types, retention relationships, and sound descriptions. Keep field names, task/relationship markers, <Subject N>, <Picture N>, <Video N>, <Audio N>, [Shot N], timestamps, (Sx), and control tags visible alongside Chinese explanations. Preserve exact proper nouns and source literals, adding Chinese meaning where helpful. The Chinese section is explanatory and not an executable prompt.

Add ### 采用的假设 before the English prompt only if materially useful.

Quality Gate

  • Every important asset has one clear role; no asset is referenced vaguely as part of "all materials."
  • A <Subject N> denotes reusable visible content; <Picture N> denotes a concrete frame or planning anchor; <Video N> denotes a whole-video source or temporal structure; <Audio N> denotes an audio signal or audio reference.
  • Task types match actual use rather than mere asset presence.
  • Each retention marker belongs to the correct visual or audio vocabulary.
  • Every label used later is defined once and retains the same meaning.
  • Keyframe anchors name the shot and frame role, and the timeline reaches them coherently.
  • Actual speakers have stable (Sx) IDs; audio-only verbal cues do not create fictitious speakers.
  • Dialogue fits its shot duration; full and partial audio reuse identify exact source scope.
  • Required visible text, logos, captions, slogans, and interface labels are quoted exactly.
  • Sound layers are placed in the correct section.
  • The prompt does not contradict one-take versus cuts or music versus N/A.
  • The final English prompt is no more than 7000 characters.
  • The Chinese translation covers every section and literal verbal or visible element without introducing new creative content.

© unknowlei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/minimax-h3-reference-video-prompt of unknowlei/minimax-h3-opencode-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/reference-rules.md

Open the folder on GitHubat commit 2e2096f

Compare with similar skills

Minimax H3 Reference Video Prompt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Minimax H3 Reference Video Prompt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Minimax H3 Reference Video Prompt this skillunknowlei/minimax-h3-opencode-skills122—~3kAutomated safety check: PassMIT
MiniMax H3 Video DirectorTFboy1/oh-my-minimaxh3-director143—~2.4kAutomated safety check: PassMIT
Article Explainer Videowwwzhouhui/skills_collection283—~2.2kAutomated safety check: PassNone
H3 Cinematic Directorjtydhr88/ComfyTV1.1k—~2.7kAutomated safety check: PassMIT
Scroll Promo Site Builderkangarooking/kangarooking-skills662—~2.6kAutomated safety check: PassMIT
Novel Storyboardeternityspring/shuohao-skills4.3k—~1.8kAutomated safety check: NotesApache-2.0

Similar skills

  • MiniMax H3 Video Director

    TFboy1/oh-my-minimaxh3-director

    Turns a script into a storyboard, assigns MiniMax H3 workflows in ComfyUI, monitors batch generation and builds a Jianying draft of the finished video.

    143 GitHub stars~2.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Article Explainer Video

    wwwzhouhui/skills_collection

    把一篇技术长文/论文解读自动做成章节式解说视频(1080p, 5-8 分钟)。双主题:warm(奶油底+珊瑚红+cozy-handdrawn 透明插图,亲和感)和 midnight(深蓝黑底+琥珀金+宋体标题+executive-tech 插图,AI 科技感),storyboard 一个 theme 字段切换。每章三种 layout 混排:illustration(左文右图+Ken…

    283 GitHub stars~2.2k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • H3 Cinematic Director

    jtydhr88/ComfyTV

    Convert approved scripts, shot briefs, storyboards, keyframes, character sheets, scene assets, prop sheets, and sound briefs into director-level storyboards and production-ready MiniMax H3 prompts.

    1.1k GitHub stars~2.7k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Scroll Promo Site Builder

    kangarooking/kangarooking-skills

    Create a scroll-controlled cinematic product website with rich motion (动效网站) from product materials, reference pages or videos, and brand assets.

    662 GitHub stars~2.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Novel Storyboard

    eternityspring/shuohao-skills

    给 AI 短剧出分镜:三层结构——段(一次视频生成,≤15 秒)→ 分镜(段内 2–5 秒的剪切,认领剧本节拍) → 分镜图(每切一张关键帧:主分镜图钉 0.00 秒,子分镜图钉各自切点)。

    4.3k GitHub stars~1.8k tokensUpdated 3 days ago
    Media & CreativeAuto-check: notes
  • Paper Collage Explainer Generator

    tl2012tl/comfyUI-llama-TE

    For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.

    241 GitHub starsUsed in 4 repos~5.2k tokens
    Media & CreativeAuto-check passed

More from unknowlei/minimax-h3-opencode-skills

  • Minimax H3 Multishot Planner

    unknowlei/minimax-h3-opencode-skills

    Non-skippable planning-only MiniMax H3 subskill used by minimax-h3-creative-director before final prompt formatting.

    122 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax H3 Keyframe Video Prompt

    unknowlei/minimax-h3-opencode-skills

    Narrow downstream MiniMax H3 specialist for pure I2VA, FL2VA, and L2VA boundary-frame prompts.

    122 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax H3 Text Video Prompt

    unknowlei/minimax-h3-opencode-skills

    Downstream MiniMax H3 specialist for professional text-to-video (T2VA) prompts using the official three-field format.

    122 GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax H3 Creative Director

    unknowlei/minimax-h3-opencode-skills

    Primary mandatory entrypoint for every MiniMax H3 video-generation request.

    122 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax H3 Prompt Reviewer

    unknowlei/minimax-h3-opencode-skills

    Downstream MiniMax H3 specialist that audits, repairs, and rewrites T2VA, I2VA, FL2VA, L2VA, and full-reference prompts into an official structured format.

    122 GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Minimax H3 Reference Video Prompt

What does Minimax H3 Reference Video Prompt do?

Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role. Minimax H3 Reference Video Prompt is an agent skill from unknowlei/minimax-h3-opencode-skills. Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role.

When should I use Minimax H3 Reference Video Prompt?

Minimax H3 Reference Video Prompt fits situations like: tasks that involve Video production; tasks that involve Comics and storyboards.

How do I install Minimax H3 Reference Video Prompt in Claude Code?

Run `npx skills add unknowlei/minimax-h3-opencode-skills --skill minimax-h3-reference-video-prompt -a claude-code`. Or copy the skill folder (skills/minimax-h3-reference-video-prompt in unknowlei/minimax-h3-opencode-skills) into .claude/skills/minimax-h3-reference-video-prompt in your project. Claude Code loads it when a task matches its description.

How do I install Minimax H3 Reference Video Prompt in Codex?

Run `npx skills add unknowlei/minimax-h3-opencode-skills --skill minimax-h3-reference-video-prompt -a codex`. Or copy the skill folder (skills/minimax-h3-reference-video-prompt in unknowlei/minimax-h3-opencode-skills) into .agents/skills/minimax-h3-reference-video-prompt in your project. Codex loads it when a task matches its description.

Can I use Minimax H3 Reference Video Prompt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add unknowlei/minimax-h3-opencode-skills --skill minimax-h3-reference-video-prompt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-h3-reference-video-prompt, .gemini/skills/minimax-h3-reference-video-prompt, .github/skills/minimax-h3-reference-video-prompt and .opencode/skills/minimax-h3-reference-video-prompt in your project.

What does Minimax H3 Reference Video Prompt need to run?

SKILL.md names no scripts, command-line tools or credentials: Minimax H3 Reference Video Prompt is instructions for the agent only.

Does Minimax H3 Reference Video Prompt access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Minimax H3 Reference Video Prompt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Minimax H3 Reference Video Prompt use?

Minimax H3 Reference Video Prompt is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Minimax H3 Reference Video Prompt use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Minimax H3 Reference Video Prompt?

Skills that share tags, products or a category with Minimax H3 Reference Video Prompt: MiniMax H3 Video Director (TFboy1/oh-my-minimaxh3-director, 143 stars), Article Explainer Video (wwwzhouhui/skills_collection, 283 stars), H3 Cinematic Director (jtydhr88/ComfyTV, 1.1k stars) and Scroll Promo Site Builder (kangarooking/kangarooking-skills, 662 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Minimax H3 Reference Video Prompt?

unknowlei (a GitHub user) maintains it in unknowlei/minimax-h3-opencode-skills, which has 122 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on August 9, 2026.

Source: unknowlei/minimax-h3-opencode-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.