Agent skill

H3 Video Prompt Enhancer

by benjiyaya in benjiyaya/Calliope

A skill your agent uses when making MiniMax H3 video prompts from media + ideas.

MITAuto-check passedMedia & Creative

Install H3 Video Prompt Enhancer

skills CLI
$ npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benjiyaya/Calliope h3-video-prompt-enhancer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benjiyaya/Calliope.git skills-src && mkdir -p .claude/skills && cp -r skills-src/calliope-backend/skills_builtin/h3-video-prompt-enhancer .claude/skills/h3-video-prompt-enhancer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
h3-video-prompt-enhancer
GitHub stars
242
Token cost
~4.6k tokens
SKILL.md length
2,384 words
Files
5 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when making MiniMax H3 video prompts from media + ideas.

  • Works in 5 steps: Classify the Mode → Gather Parameters → Creative Enhancement → …
  • Making MiniMax H3 video prompts from media + ideas
  • SKILL.md covers Overview, When to Use, Step 0: Classify the Mode and Step 1: Gather Parameters, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

H3 Video Prompt Enhancer is an agent skill from benjiyaya/Calliope. Use when making MiniMax H3 video prompts from media + ideas.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `references/base-multishot-format.md` and `references/creative-showcase.md`).

It sits in Media & Creative, covering AI video generation and Diffusion and image models. It works with MiniMax, ComfyUI and SvelteKit. The repository describes itself as: Local-first AI Idea-to-video studio — FastAPI + SvelteKit + ComfyUI. The licence is MIT.

When your agent uses it

  • Making MiniMax H3 video prompts from media + ideas
  • Tasks that involve AI video generation
  • Tasks that involve Diffusion and image models

Example prompts

  • “/h3-video-prompt-enhancer”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Classify the Mode
  2. Gather Parameters
  3. Creative Enhancement
  4. Format and Output
  5. Present to User

What it can do on your machine

Read from SKILL.md and the folder at commit 0a75919. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

H3 Video Prompt Enhancer loads about 4.6k tokens when it runs, and up to ~30k if it reads all its reference files. Until then it costs about 21 tokens; SKILL.md has 2,384 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~30k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benjiyaya/Calliope at commit 0a75919, republished under its MIT licence (© benjiyaya). 2,384 words, ~4,566 tokens.

Download SKILL.mdSave it as .claude/skills/h3-video-prompt-enhancer/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
h3-video-prompt-enhancer
description
Use when making MiniMax H3 video prompts from media + ideas.
version
1.1.0
license
MIT
metadata.author
Benjamin Law (Muse-AI)
metadata.tags
video, prompt-engineering, minimax-h3, comfyui, text-to-video, image-to-video, creative
metadata.related_skills
comfyui

H3 Video Prompt Enhancer

Overview

Transform a user's rough video idea + attached assets into a production-grade MiniMax H3 video generation prompt. The skill handles two H3 generation modes:

  • Ref2VA — when the user provides reference assets (character sheets, style images, reference videos, voice samples). Outputs a 6-section structured prompt. Load references/ref2va-format.md for the full spec.
  • Base MultiShot (T2VA / I2VA / FL2VA / L2VA) — for text-to-video or image-anchored generation. Outputs a timed multi-shot prompt. Load references/base-multishot-format.md for the full spec.

The skill's distinctive value: it doesn't just format-comply — it creatively enhances the brief with professional-grade cinematic detail (camera aesthetics, visual texture, lighting design, pacing arcs, spatial choreography, continuity tracking) before mapping everything into the exact H3 output format. Load references/creative-showcase.md for quality benchmarks and pattern examples from advanced long-form prompts.

When to Use

Trigger when the user:

  • Attaches 1+ images/videos/audio files and describes a video they want to create
  • Says "make a video prompt" / "enhance this for H3" / "write a Ref2VA prompt" / "img2vid" / "txt2vid"
  • Provides a video concept and wants it structured for MiniMax H3 generation
  • Mentions H3, MiniMax, Ref2VA, ComfyUI video generation
  • Pastes a creative brief (CAMERA/LOOK/STYLE format) and wants it converted to H3 format

Don't use for:

  • Non-H3 video models (Sora, Runway, Kling, etc.) — the output format is H3-specific
  • Pure image generation prompts
  • Video editing tasks that don't involve new generation

Step 0: Classify the Mode

Determine the H3 mode from what the user attached and stated:

User providesModeFormat reference
Reference images/videos/audio (character sheets, style refs, voice clips)Ref2VAreferences/ref2va-format.md
Nothing — just a text ideaT2VA (text-to-video)references/base-multishot-format.md
1 image as the first frameI2VA (image-to-video)references/base-multishot-format.md
2 images (first frame + last frame)FL2VA (first-last-to-video)references/base-multishot-format.md
1 image as the last frame onlyL2VA (last-frame-to-video)references/base-multishot-format.md

Key distinction: An image used as a first/last frame anchor = I2VA/FL2VA/L2VA. An image used as a character/style reference (not a frame position) = Ref2VA. When ambiguous, ask the user: "Is this image a frame anchor (first/last frame of the video) or a reference (character/style/scene)?"

Reading the Assets (DSH)

In DSH you must actually see or transcribe each reference asset before you can describe it faithfully. Two requirements apply:

  • Vision capability. The active model must accept image input. Use the harness image tool on each sheet — if the tool reports model ... does not declare image input, the current model is not vision-capable. Switch to an image-capable model (or delegate the visual read to a vision-capable partner) before describing reference images. Do NOT guess a character's appearance from a filename.
  • Asset description depth. Character sheets usually show multiple views (front/side/back, several poses). Extract per view: identity (gender, age, build), face (hair, skin, eyes, markings), full outfit head-to-toe with exact colors/materials, weapons/props, and any on-screen labels. This feeds subject_definitions.
  • Voice/audio assets. Transcribe or listen to each clip to capture the speaker's timbre, pitch, rate, and accent for the (Sx) speaker IDs. If speech transcription is unavailable in the harness, note the delivery style from any user-provided description and keep (Sx) IDs generic.

Step 1: Gather Parameters

Confirm these before enhancing (ask if missing, but proceed if the idea is clear enough):

  • duration_s: 4–15 seconds (integer). Default to 8 if unspecified.
  • aspect ratio: 16:9, 9:16, 1:1, 4:3, 21:9. Default to 16:9.
  • shot count: Let the planner decide, or respect user's explicit count. Budget: 4–6s → 1–2 shots; 7–10s → 2–3 shots; 11–15s → 3–5 shots.
  • asset inventory: What each attached file is and its role. Numbering per modality defines labels: Image k → <Picture k> or <Subject N>; Video k → <Video k>; Audio k → <Audio k>.

Step 2: Creative Enhancement

This is the core value-add. Take the user's idea and enrich it across seven dimensions. The goal: produce a prompt with the depth and cinematic intelligence of a professional storyboard. Consult references/creative-showcase.md for full examples of each pattern.

Enhancement Dimensions

1. Camera Identity — Assign a distinct camera aesthetic that matches the scene's tone:

  • Physical type: handheld, tripod, drone, steadicam, dolly, security cam, dashcam, POV, arc shot
  • Imperfections to preserve (when stylistically appropriate): hand tremor, autofocus hunting, exposure fluctuation, lens flare, motion blur, awkward zoom
  • Format/aesthetic hint: 16mm film, DV tape, digital clean, anamorphic, vintage camcorder, broadcast

2. Visual Texture (LOOK) — Define the image quality and color science:

  • Grain/noise: film grain, electronic noise, clean digital, VHS tracking, subtle blur
  • Color palette: warm/cool, saturated/desaturated, high/low contrast, natural skin tones
  • Lighting design: natural, studio, neon, golden hour, mixed sources, practical lights
  • Lighting transitions: if locations change across shots, describe how the light shifts

3. Pacing Arc — Plan the energy progression across the full duration:

  • Build patterns: quiet→energetic, tense→release, slow build→explosive peak→settle, steady rhythm
  • Cut rhythm: accelerating cuts toward climax, slow contemplative holds, musical cutting on beats
  • Match the pacing arc to the emotional intent of the brief

4. Character Detail — Flesh out every on-screen person with specificity:

  • Physical: age range, build, hair color/style, skin tone, distinctive features (scars, freckles, heterochromia)
  • Wardrobe: specific garments with colors, materials, textures, accessories — note changes across locations or time
  • Visual signature: a recurring color or visual element that makes the character instantly recognizable in every shot (e.g., "electric purple energy trails", "glossy teal jacket reflections", "always wears a red scarf")
  • Coverage note: describe outfits fully — avoid implying revealing clothing unless the user explicitly requests it

5. Spatial Geography — For action sequences or multi-location videos:

  • Screen direction: who enters from where, movement vectors (Left→Right, Deep→Front, foreground↔background)
  • Key action moments: the 2–3 critical motion beats that define the sequence
  • Environmental layout: what's in the space, how it's lit, reflective surfaces, depth

6. Continuity Progression — Track what changes across shots so the video feels coherent:

  • Physical state: damage accumulates, hair gets messier, clothes get wet/torn/dusty
  • Environmental: props move, lights flicker, weather shifts, debris scatters
  • Emotional: expressions and body language evolve naturally across the timeline

7. Sound Design Plan — Map the full audio landscape:

  • Ambience: room tone, environmental atmosphere, background chatter, traffic, wind
  • Physical action sounds: footsteps, impacts, fabric rustling, door creaks, liquid pouring
  • Non-diegetic score: instrumentation, tempo, rhythm, dynamic changes (goes in non_diegetic_music field)
  • Diegetic music: source-visible music — radio, speaker, live performance (goes in shot description)
  • Dialogue vs. voiceover: clearly mark which is spoken on-camera vs. off-screen narration
Per-Shot Quality Bar

Every shot in the storyboard must specify:

  • Composition: framing (wide, medium, close-up, extreme close-up, macro), angle (eye-level, low, high, overhead, Dutch)
  • Camera motion: type + amplitude + speed in natural English (e.g., "The camera pushes in with small amplitude at slow speed")
  • Subject action: exactly ONE dominant action per shot — never cram multiple actions
  • Environment/lighting: what's visible, how it's lit, time of day cues
  • Sound cue: what's audible in this specific moment
  • Reference labels (Ref2VA only): where referenced content appears or takes effect

Step 3: Format and Output

Load the appropriate format reference and produce the final H3 prompt.

For Ref2VA → Load references/ref2va-format.md. Output exactly 6 sections in order:

  1. subject_definitions: — one line per tracked item
  2. summary: — task-type prefix + one paragraph
  3. retention_analysis: — fidelity markers per label
  4. detailed_description: — 350–500 words, opens with style, then [Shot N] timeline
  5. overall_soundscape: — 1–4 sentences
  6. non_diegetic_music: — 1–3 sentences or N/A

For Base MultiShot → Load references/base-multishot-format.md. Output:

  1. Instruction line (I2VA/FL2VA/L2VA only — exact template from reference)
  2. integrated_multimodal_description: — timed multi-shot timeline
  3. overall_soundscape: — 1–4 sentences
  4. non_diegetic_music: — 1–3 sentences or N/A
Critical Format Rules (both modes)
  • Output ONLY the specified fields — no preamble, no explanation, no markdown fences, no commentary
  • Write everything in English. Exceptions: dialogue/lyrics inside <d> tags and visible on-screen text keep their original language verbatim
  • Timestamps: [Shot 1] has NO timestamp and opens with style + initial composition. Later shots: [Shot N] At MM:SS.mmm, the camera cuts to ... with strictly increasing times within duration
  • A cut must add NEW information (subject, space, state, viewpoint, time). If only distance/angle changes, use camera motion instead of a cut
  • Camera motion = motion type + amplitude + speed as natural English action, not stacked labels
  • Dialogue format: identifying phrase + speaker ID + delivery OUTSIDE <d>; inside <d> ONLY language tag + exact words: <d>[English] Wait for us!</d>
  • Preserve the user's dialogue verbatim — never translate, rewrite, or paraphrase
  • Voiceover: exact phrase "says in an off-screen voiceover" + immediately state "while his/her lips remain completely closed"
  • On-screen text (signs, neon, subtitles): English double quotation marks, verbatim, no translation
  • Avoid named third-party IP, real celebrities, trademarked characters — describe them generically
Camera Motion Vocabulary

Types: Zoom In/Out, Push In/Pull Out, Pan L/R, Truck L/R, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, Roll CW/CCW. Amplitude: "with small amplitude" / "with large amplitude" (omit when medium). Speed: "at slow speed" / "at fast speed" (omit when normal).

Show full SKILL.md (973 more words)Show less

Step 4: Present to User

After generating the H3 prompt:

  1. Present the full prompt in a code block so it can be copied directly
  2. State which mode was detected and why
  3. Flag any assumptions made (e.g., "Assumed 8s duration and 16:9 — adjust if needed")
  4. Offer to refine specific aspects: camera aesthetic, pacing, character detail, shot count, sound design

Creative Patterns from Showcase

The user maintains advanced long-form prompt examples representing the quality bar. Load references/creative-showcase.md for full examples. Key transferable patterns:

Long-form Storytelling (montage, day-in-the-life, narrative):

  • Explicit CAMERA/LOOK/STYLE blocks defining the aesthetic before the storyboard
  • Outfit and location progression across time, with lighting transitions
  • Voiceover carrying emotional narrative over non-synchronized visuals
  • Pace building from quiet opening to high-energy finale
  • Per-shot storyboard with timing, action, and VO lines

Action Choreography (combat, chase, sports):

  • Per-character color lock — each character has a signature color for their effects/trails/reflections
  • Explicit spatial layout with screen directions (Deep→Front, Left→Right)
  • Action vectors — the key motion moment that defines each shot
  • Progressive continuity — damage accumulates, hair gets windblown, dust/cracks appear
  • Strict shot count enforcement — state the count and hold it
  • Environmental reactivity — holograms flicker, floors reflect, walls crack in response to action

Anti-Slop: Motion Quality Checklist

"Slop" in H3 prompts means shots that look flat, static, or lifeless in the generated video. The most common cause is too many wide shots, static camera holds, and vague action descriptions. Before finalizing any prompt, verify every shot against these rules:

Every Shot Must Have:
  1. A close-up or tight framing — never a wide establishing shot during action beats. Extreme close-ups of hands, faces, weapons, impact points.
  2. ONE clear physical action broken down at body-part level: fingers gripping, eyes tracking, feet grinding, hips rotating, jaw clenching. Not "she fights him" — describe the exact limb movement.
  3. Camera movement — every shot needs explicit camera motion (type + amplitude + speed). No "camera holds" or "static shot" during action. Whip-pans, tracking, orbits, push-ins.
  4. No crammed actions — if the brief describes 3 sequential actions, split across 3 shots with cuts. One action per shot is a hard rule.
Slop Patterns to Avoid:
SlopFix
Wide establishing shot during actionCut to close-up of feet/hands/face at action onset
"Camera holds on the standoff"Camera pushes in, orbits, or whip-pans to maintain motion
"She ducks, grabs, pivots, and throws" in one shotSplit into: duck (close-up) → grab (hand close-up) → pivot+throw (tracking)
"The camera follows them" (vague)"The camera tracks at fast speed, whip-panning between striker and target"
Characters standing/walking with no actionEvery beat needs a physical micro-action: eyes narrowing, fingers tightening, weight shifting
Static aftermath shotEven aftermath needs camera movement: pull-back, tilt-up, slow orbit
Rewriting Slop:

When a prompt feels flat, identify the weakest shots and rewrite with this pattern:

  1. Zoom in — change wide/medium to close-up or extreme close-up
  2. Add body-part detail — "his fist" → "his white-knuckle fist, tendons visible"
  3. Add camera motion — "the shot cuts to" → "the camera whip-pans to" or "tracks at fast speed"
  4. Split crammed actions — extract one action per shot, add a cut between them
  5. Replace static moments — "holds on the aftermath" → "pulls back slowly as..." or "orbits around..."

Common Pitfalls

  1. Treating Ref2VA reference images as frame anchors. A character sheet or style reference is a <Subject>, not a <Picture>. Only use <Picture N> standalone when the image IS a concrete frame position (first frame, keyframe, last frame). When in doubt, cite the image inside the relevant <Subject N> line.

  2. Cramming multiple actions into one shot. One dominant action per shot — this is a hard H3 constraint. If the brief describes sequential actions, split them across multiple shots with cuts.

  3. Forgetting camera motion. Every shot needs an explicit camera movement specification — even "Static Shot" if the camera doesn't move. Omitting it leaves the model guessing.

  4. Inconsistent character identity across shots. Repeat identity anchors (hair, clothing, key props) in every shot, phrased freshly but consistently. If hair gets messy in shot 3, it stays messy in shot 4.

  5. Wrong shot count for duration. Respect the budget: 4–6s → 1–2 shots; 7–10s → 2–3 shots; 11–15s → 3–5 shots. Don't plan 5 shots for a 5-second video.

  6. Mixing diegetic and non-diegetic music. Music audible to characters (radio, live performance, phone speaker) goes in the shot description. Background score the characters can't hear goes in non_diegetic_music. Never put the same music in both.

  7. Inventing reference labels (Ref2VA). Never create labels beyond those defined in subject_definitions. <Subject 3> means the same thing in every section where it appears. Different indices for different modalities are numbered independently.

  8. Translating or rewriting dialogue. Preserve the user's exact words inside <d> tags — including punctuation, hesitations, and language. Never translate to English if the user wrote in another language.

  9. Skipping the style opener. The detailed_description (Ref2VA) or integrated_multimodal_description (Base) MUST open with 1–2 sentences of overall style BEFORE [Shot 1]. Don't jump straight into the first shot.

  10. Flat timestamps. [Shot 1] never has a timestamp. Every subsequent shot needs At MM:SS.mmm with strictly increasing times. Convert duration to seconds with proper formatting (8s → 8.00 for the instruction line).

Verification Checklist

  • Mode correctly detected from attached assets and user intent
  • All seven enhancement dimensions applied (camera, look, pacing, character, spatial, continuity, sound)
  • Output contains ONLY the required H3 fields — no preamble, no markdown fences, no extra commentary
  • Timestamps are strictly increasing and fall within duration_s
  • Exactly ONE dominant action per shot
  • Camera motion specified for every shot (type + amplitude + speed)
  • Character identity consistent across all shots (repeat anchors every shot)
  • Dialogue preserved verbatim in <d>[Language] ...</d> format
  • Style opener present before [Shot 1]
  • Reference labels (Ref2VA only) used consistently and never invented beyond subject_definitions
  • Duration (4–15s) and aspect ratio specified or assumed-and-flagged
  • No named IP, celebrities, or trademarked character names
  • Diegetic music in shot descriptions, non-diegetic score in non_diegetic_music

© benjiyaya, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in calliope-backend/skills_builtin/h3-video-prompt-enhancer of benjiyaya/Calliope.

  • SKILL.md
  • README.md
  • references/base-multishot-format.md
  • references/creative-showcase.md
  • references/ref2va-format.md

Open the folder on GitHubat commit 0a75919

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in benjiyaya/Calliope, which our catalogue first saw on October 7, 2026.

Compare with similar skills

H3 Video Prompt Enhancer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

H3 Video Prompt Enhancer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
H3 Video Prompt Enhancer this skillbenjiyaya/Calliope242—~4.6kAutomated safety check: PassMIT
H3 Videoagent-next/video-agent120—~2.9kAutomated safety check: PassApache-2.0
Open Videoagent-next/video-agent120—~3.2kAutomated safety check: PassApache-2.0
ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit105—~12kAutomated safety check: PassApache-2.0
Seedance 2 5calesthio/OpenMontage66k—~3.2kAutomated safety check: PassAGPL-3.0
Minimax H3calesthio/OpenMontage66k—~580Automated safety check: PassAGPL-3.0

Similar skills

  • H3 Video

    agent-next/video-agent

    OpenVideo skill (v0.1.0): generate high-quality local video with the OpenVideo product (MiniMax H3 backend).

    120 GitHub stars~2.9k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Open Video

    agent-next/video-agent

    Generate, edit, or direct videos via open-source models (MiniMax H3 baseline; Wan2.2 / LTX future).

    120 GitHub stars~3.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Seedance 2 5

    calesthio/OpenMontage

    Generate 4-30 second cinematic video with ByteDance Seedance 2.5 through fal.ai, Volcengine Ark, Runway, or ComfyUI Partner Nodes.

    66k GitHub stars~3.2k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Minimax H3

    calesthio/OpenMontage

    Generate MiniMax H3 (Hailuo 3.0) video through the official MiniMax v2 API, fal.ai, Runway, ComfyUI Partner Nodes, or local open weights in ComfyUI.

    66k GitHub stars~580 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • High Density Fight Prompt

    JGRFW/comfyui-AICG3D

    根据参考图和用户设定创作高密度、连续因果的电影级打斗视频提示词,并输出保持同一时间线的中文导演稿与 MiniMax H3 Ref2VA 英文六段稿。适用于 15 秒动作设计、武器战、徒手战、巨物战和参考图驱动的连续攻防。

    187 GitHub stars~590 tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from benjiyaya/Calliope

  • Shot Composer Blockout

    benjiyaya/Calliope

    A skill your agent uses when the user asks to build, pose, frame, or ANIMATE a 3D scene in Build Scene — characters, primitives, shot framing, keyframe motion, or exporting a blockout as a…

    242 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Calliope

    benjiyaya/Calliope

    Write finished Calliope story content — beats, cast, screenplay scenes, shot clips, continuity requirements — into an existing Calliope project through the calliope-cli bridge.

    242 GitHub stars~5.5k tokensUpdated 3 days ago
    Auto-check passed
  • Character Consistency

    benjiyaya/Calliope

    A skill your agent uses when writing or improving character image prompts — sheets, portraits, and reference images that stay consistent across scenes.

    242 GitHub stars~510 tokensUpdated 3 days ago
    Auto-check passed
  • Scene To Video

    benjiyaya/Calliope

    A skill your agent uses when turning a Calliope scene into a video job — chaining from a previous clip, ordering reference images, or picking duration settings.

    242 GitHub stars~552 tokensUpdated 3 days ago
    Auto-check passed

Questions about H3 Video Prompt Enhancer

What does H3 Video Prompt Enhancer do?

A skill your agent uses when making MiniMax H3 video prompts from media + ideas. H3 Video Prompt Enhancer is an agent skill from benjiyaya/Calliope. Use when making MiniMax H3 video prompts from media + ideas.

When should I use H3 Video Prompt Enhancer?

H3 Video Prompt Enhancer fits situations like: making MiniMax H3 video prompts from media + ideas; tasks that involve AI video generation; tasks that involve Diffusion and image models.

How do I install H3 Video Prompt Enhancer in Claude Code?

Run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a claude-code`. Or copy the skill folder (calliope-backend/skills_builtin/h3-video-prompt-enhancer in benjiyaya/Calliope) into .claude/skills/h3-video-prompt-enhancer in your project. Claude Code loads it when a task matches its description.

How do I install H3 Video Prompt Enhancer in Codex?

Run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a codex`. Or copy the skill folder (calliope-backend/skills_builtin/h3-video-prompt-enhancer in benjiyaya/Calliope) into .agents/skills/h3-video-prompt-enhancer in your project. Codex loads it when a task matches its description.

Can I use H3 Video Prompt Enhancer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benjiyaya/Calliope --skill h3-video-prompt-enhancer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/h3-video-prompt-enhancer, .gemini/skills/h3-video-prompt-enhancer, .github/skills/h3-video-prompt-enhancer and .opencode/skills/h3-video-prompt-enhancer in your project.

What does H3 Video Prompt Enhancer need to run?

SKILL.md names no scripts, command-line tools or credentials: H3 Video Prompt Enhancer is instructions for the agent only.

Does H3 Video Prompt Enhancer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is H3 Video Prompt Enhancer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does H3 Video Prompt Enhancer use?

H3 Video Prompt Enhancer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does H3 Video Prompt Enhancer use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.

What are the alternatives to H3 Video Prompt Enhancer?

Skills that share tags, products or a category with H3 Video Prompt Enhancer: H3 Video (agent-next/video-agent, 120 stars), Open Video (agent-next/video-agent, 120 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars) and Seedance 2 5 (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains H3 Video Prompt Enhancer?

benjiyaya (a GitHub user) maintains it in benjiyaya/Calliope, which has 242 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: benjiyaya/Calliope on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.