Agent skill

Scenario Minimax Video

by scenario-labs in scenario-labs/skills

A skill your agent uses when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or…

MITAuto-check passedMedia & Creative

Install Scenario Minimax Video

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-minimax-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-minimax-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-minimax-video .claude/skills/scenario-minimax-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-minimax-video
GitHub stars
946
Token cost
~2.6k tokens
SKILL.md length
1,338 words
Files
1
Skills in repo
146
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or…

  • Works in 6 steps: search with target="models",… → model_schema_get with that id: modes,… → upload_asset two clean, evenly lit… → …
  • Editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video
  • SKILL.md covers Overview, Quick reference, Picking the member and Editing a clip you already have, plus 3 more sections
  • Calls npx

What it does

Scenario Minimax Video is an agent skill from scenario-labs/skills. Use when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or audio, native stereo audio, bracketed camera commands, extending a clip, inserting a new scene into one, or recasting the people in it from photos. Keywords: MiniMax, Hailuo 3.0, H3, H3 Max, Turbo, Extend Video, Insert Video, Recast, Hailuo 2.3, T2V, I2V, V2V, 2K.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation. It works with MiniMax and Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video
  • First and last frame anchors
  • Reference images
  • Native stereo audio

Example prompts

  • “/scenario-minimax-video”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. search with target="models", query="minimax hailuo", public=true. Prefer the newest non-deprecated hit, e.g. model_minimax-h3 (a live hit…
  2. model_schema_get with that id: modes, caps, and allowed values before anything else.
  3. upload_asset two clean, evenly lit character stills (see the scenario skill) to get asset ids.
  4. model_run with that model_id, dry_run=true, and the exact parameters={"prompt": "[Tracking shot] The scout from the reference images…
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second…
  6. asset_display the output and review it with sound on: dialogue and SFX sync are part of what you paid for.

What it can do on your machine

Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Minimax Video loads about 2.6k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 1,338 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 1,338 words, ~2,649 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-minimax-video/SKILL.md (or your agent's skills folder).
name
scenario-minimax-video
description
Use when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or audio, native stereo audio, bracketed camera commands, extending a clip, inserting a new scene into one, or recasting the people in it from photos. Keywords: MiniMax, Hailuo 3.0, H3, H3 Max, Turbo, Extend Video, Insert Video, Recast, Hailuo 2.3, T2V, I2V, V2V, 2K.
license
MIT

Scenario MiniMax Video

Overview

MiniMax's Hailuo video family on Scenario has three kinds of member. H3 (Hailuo 3.0) folds text, keyframe, and reference conditioning into one model and generates stereo audio in the same pass. The H3 Max line splits that into single-mode members (text, image, reference, lip sync), each with a faster Turbo twin on the first two. Three Max members edit a clip you already have instead of generating one: Extend Video (with a Turbo twin), Insert Video, and Recast. The older Hailuo 2.3 pair carried a deprecated tag naming H3 as successor at authoring time. Discover members with search and treat model_schema_get as the contract: they agree on almost nothing, not even the spelling of a resolution.

Connection and the core loop: see the scenario skill in this repo; model-agnostic video work: the scenario-video skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

H3's mode follows from the inputs (names from the live schema):

ModeInputsBehavior
TextpromptaspectRatio honored (21:9 through 9:16, default adaptive)
First framefirstFrameImage (+ prompt)shape follows the image; aspectRatio ignored
First + lastfirstFrameImage + lastFrameImagelastFrameImage is valid only alongside firstFrameImage
ReferencereferenceImages, referenceVideos, referenceAudioguides subject, style, motion, and voice, not forced frames; aspectRatio still applies

Keyframes and references are mutually exclusive: neither frame combines with any reference array. referenceAudio never rides alone; it requires at least one image or video reference. Reference parameters are arrays even for one asset. At authoring time H3 took 9 reference images, 3 videos, and 3 audio files (videos and audio each 2 to 15 seconds, and 2 to 15 seconds in total), 5 to 15 seconds of output, at 768P or 2K.

The 2.3 members take only prompt and firstFrameImage, plus a coupled pair: at authoring time 10 second duration was available only at 768p, and 1080p only at 6 seconds. Their promptOptimizer (default true) rewrites the prompt before generation: leave it on for thin prompts, switch it off when engineered wording must survive verbatim. H3 has no such switch, and neither H3 nor 2.3 takes a seed.

Picking the member

H3 is the one member that mixes modes in a single call, and, lip sync aside, the one generator that reaches 2K; it is also slow and expensive, and reference media adds more (reference videos bill per second of uploaded footage). The Max text and image members cap at 768P and take a seed and a promptExpansionMode (disabled, balanced, quality); set it to disabled when engineered wording must survive verbatim. Their Turbo twins share the same inputs and are the iteration tier: draft at 480P on Turbo, then spend the expensive member on the take you keep. A draft is a composition and motion check, not a preview of the final pixels: a different member or resolution is a new generation even with the same seed. dry_run the draft and the final payloads before a batch. Re-discover a member before naming 2.3 for new work: deprecated members can disappear from the catalog.

Editing a clip you already have

These members take the existing video as video (any video asset id, generated or uploaded) and bill only what they make or touch, so price each with dry_run on the real source.

MemberWhat it doesKey inputs and caps at authoring time
Extend (+Turbo)adds 1 to 15 s after the sourcesource 1.625 to 60 s, up to 50 MB, aspect 0.4 to 2.5; output extended returns source plus new footage, continuation the new footage behind a lead-in copied from the source (see Common mistakes); referenceAudio replaces the source soundtrack as the audio guide; aspectRatio other than auto crops; 480P to 2K
Insert Videoa new 5 to 13 s scene, then the original resumesstartTime and resumeTime in source seconds (both required); duration is the new scene only; up to 9 referenceImages and 3 referenceVideos, billed by reference tokens; colorMatch on by default; 480p or 768p, lowercase
Recastswaps up to 4 people for new ones from photossource 5 to 30 s, no single shot over 15 s, billed per source second; referenceImages one photo per person, required, mapped left to right by default; prompt optional, to say who becomes whom; motion, camera, cuts, and audio are kept; 768P or 1080P

Write an Extend prompt about what happens next, never a description of the source: the model already sees it, and enablePromptExpansion (on by default) rewrites the prompt from the source to keep the continuation consistent. Recast is the lane for re-shooting the same performance with a different person; it keeps everything else, so it is the wrong tool for a new setting or style.

Show full SKILL.md (537 more words)Show less

Camera in brackets, motion in moderation

The whole family reads bracketed camera commands inline in the prompt: [Push in], [Pull out], [Pan left], [Tilt up], [Truck right], [Pedestal up], [Zoom out], [Shake], [Tracking shot], [Static shot]. Up to three moves combine in one bracket ([Pan left, Pedestal up]); separate brackets sequence them. Keep 2 or 3 motion cues per shot in total: piling on moves invites background wobble and texture flicker. Image-to-video tends to drift even unprompted, so lock statics explicitly ("Static camera, locked shot, tripod mounted").

Write the rest as natural prose, ordered camera, subject, action, scene, lighting and mood, style. On H3, sound lives in the prompt too: describe dialogue lines, SFX, and ambience inline so they sync to the action; no audio parameter exists to switch instead.

Worked example: a character clip from reference stills

  1. search with target="models", query="minimax hailuo", public=true. Prefer the newest non-deprecated hit, e.g. model_minimax-h3 (a live hit at authoring time: re-discover each session).
  2. model_schema_get with that id: modes, caps, and allowed values before anything else.
  3. upload_asset two clean, evenly lit character stills (see the scenario skill) to get asset ids.
  4. model_run with that model_id, dry_run=true, and the exact parameters={"prompt": "[Tracking shot] The scout from the reference images sprints across a rooftop at dusk, coat snapping in the wind, warm rim light, footsteps and distant traffic in the audio, cinematic realism.", "referenceImages": ["asset_a", "asset_b"], "duration": 8, "resolution": "2K", "aspectRatio": "16:9"} for the cost estimate; re-estimate after changing duration, resolution, or the reference count.
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second model_run; H3 at 2K commonly runs several minutes.
  6. asset_display the output and review it with sound on: dialogue and SFX sync are part of what you paid for.

Common mistakes

  • Combining firstFrameImage with a reference array: keyframes and references never mix.
  • Passing referenceAudio alone: it requires at least one image or video reference.
  • Sending lastFrameImage without firstFrameImage: the pair anchors both endpoints or neither.
  • Expecting aspectRatio to win over a first frame: the shape follows the image.
  • Carrying one member's values to another: H3 spells resolutions 768P and 2K, the Max generators other than lip sync stop at 768P, Insert spells 480p and 768p in lowercase, and Recast offers only 768P and 1080P.
  • Sending Recast a clip with a single shot over 15 seconds, or one under 5: cut it first (the scenario-video-editing skill).
  • Pricing Recast on the output: it bills every second of the source, so trim the source to the part that needs the new people.
  • Expecting either Extend output to be new footage only: extended returns the source plus the continuation, and continuation opens on a copy of the source's last 1.625 s (39 frames at 24 fps in two authoring-time runs), then new footage that ran a little past the request (2 s gave 2.1 s, 3 s gave 3.5 s, each billed as requested). For pieces assembled later (the scenario-video-assembly skill), take continuation and cut from startTime 1.7: the cut tool steps by 0.1 s and 1.6 keeps a source frame (the scenario-video-editing skill).
  • Stacking camera moves: past 2 or 3 cues the background wobbles and textures flicker.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-minimax-video of scenario-labs/skills.

Open the folder on GitHubat commit f6f8ab7

Compare with similar skills

Scenario Minimax Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Minimax Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Minimax Video this skillscenario-labs/skills946—~2.6kAutomated safety check: PassMIT
Miora Video Studiokangarooking/director-skills181—~4.6kAutomated safety check: PassMIT
ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit105—~12kAutomated safety check: PassApache-2.0
Videoguaardvark/guaardvark258—~1.2kAutomated safety check: PassMIT
Avatar Videocalesthio/OpenMontage66k—~1.6kAutomated safety check: PassAGPL-3.0
Create Videocalesthio/OpenMontage66k—~1.3kAutomated safety check: PassAGPL-3.0

Similar skills

  • Miora Video Studio

    kangarooking/director-skills

    miora 视频生成的可靠性与规格层——补内置技能未覆盖的部分:四必填参数闸门(模式/分辨率/时长/画幅缺一不可提交)、时长区间闸门、完成信号的判定与轮询等待、官方规格口径与本通道实测差异(时长/分辨率/参考图上限)、文生/首尾帧/参考生三种模式的参数事实、提示词的自动修正边界、生成状态向用户的同步。当用户要生成视频、动效、短视频,或提到参考生视频/首尾帧/图生视频、miora…

    181 GitHub stars~4.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video

    guaardvark/guaardvark

    Generate video clips on the user's own GPU through Guaardvark: text-to-video, image-to-video, first+last frame animation, clips with their own soundtrack and dialogue (MiniMax H3), short looping…

    258 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Cassette Video Edit

    Cassette-Editor/oh-my-cassette

    Edit, trim, cut, caption, subtitle, reframe, combine, add background music to, or export video, audio, and image files through Cassette.

    119 GitHub starsUsed in 1 repo~3.4k tokens
    Media & CreativeAuto-check passed

More from scenario-labs/skills

All 146 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    946 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    946 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Chatgpt Pet Create

    scenario-labs/skills

    A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…

    946 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Minimax Video

What does Scenario Minimax Video do?

A skill your agent uses when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or…. Scenario Minimax Video is an agent skill from scenario-labs/skills. Use when generating or editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video, image-to-video, first and last frame anchors, reference images, videos, or audio, native stereo audio, bracketed camera commands, extending a clip, inserting a new scene into one, or recasting the people in it from photos.

When should I use Scenario Minimax Video?

Scenario Minimax Video fits situations like: editing video with MiniMax Hailuo models on Scenario via MCP: text-to-video; first and last frame anchors; reference images; native stereo audio.

How do I install Scenario Minimax Video in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-minimax-video -a claude-code`. Or copy the skill folder (skills/scenario-minimax-video in scenario-labs/skills) into .claude/skills/scenario-minimax-video in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Minimax Video in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-minimax-video -a codex`. Or copy the skill folder (skills/scenario-minimax-video in scenario-labs/skills) into .agents/skills/scenario-minimax-video in your project. Codex loads it when a task matches its description.

Can I use Scenario Minimax Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-minimax-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-minimax-video, .gemini/skills/scenario-minimax-video, .github/skills/scenario-minimax-video and .opencode/skills/scenario-minimax-video in your project.

What does Scenario Minimax Video need to run?

Going by SKILL.md and its folder, Scenario Minimax Video needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Scenario Minimax Video access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Minimax Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Minimax Video use?

Scenario Minimax Video is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Minimax Video use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Minimax Video?

Skills that share tags, products or a category with Scenario Minimax Video: Miora Video Studio (kangarooking/director-skills, 181 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars), Video (guaardvark/guaardvark, 258 stars) and Avatar Video (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Minimax Video?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.