Agent skill

MiniMax Multimodal Toolkit

by madebyaris in madebyaris/advance-minimax-m3-cursor-rules

Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media.

MITAuto-check passedMedia & Creative

Install MiniMax Multimodal Toolkit

skills CLI
$ npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-multimodal-toolkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install madebyaris/advance-minimax-m3-cursor-rules minimax-multimodal-toolkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/madebyaris/advance-minimax-m3-cursor-rules.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/minimax-multimodal-toolkit .claude/skills/minimax-multimodal-toolkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-multimodal-toolkit
GitHub stars
126
Token cost
~1.6k tokens
SKILL.md length
684 words
Files
2
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media.

  • Works in 3 steps: Determine the Real Goal → Route to the Right Path → Inspect Before Generating
  • Generating an image, clip, narration or music asset with MiniMax
  • SKILL.md covers When to Use, Step 0: Determine the Real Goal, Step 1: Route to the Right Path and Step 2: Inspect Before…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill sends media tasks down the smallest path that can honestly produce the requested artifact, which keeps the always-on prompt light. Step 0 classifies the request as direct asset generation, product integration code, pipeline work on existing files, or capability research, and warns against jumping into API integration when direct generation is enough. A routing table maps one-off images to the runtime's own image tool, video, TTS, voice and music to current MiniMax docs and the repo's tools, edits to local tooling such as FFmpeg, and app features to a small tested integration slice.

Before generating, the agent inspects the repo for existing media patterns, asset folders, environment handling and helpers, and checks auth and output paths first. The skill also notes MiniMax M3's native multimodal input, where an attached image or video frame can be fed straight to the model for design parity or error UI triage, with a separate minimax-m3-multimodal-input skill covering that workflow. A reference.md file adds routing and output conventions.

When your agent uses it

  • Generating an image, clip, narration or music asset with MiniMax
  • Integrating MiniMax media APIs into an application
  • Converting, trimming or joining generated media with FFmpeg
  • Using an attached screenshot as ground truth through M3 multimodal input

Example prompts

  • “Generate a narration for this product script with MiniMax text to speech and save it under assets/audio.”
  • “Trim the generated clip and join it with the intro video using FFmpeg.”
  • “Add an endpoint to our app that calls the MiniMax image API, with a focused test.”

Requirements

  • MiniMax credentials for the media APIs
  • FFmpeg for local media processing (optional)

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Determine the Real Goal
  2. Route to the Right Path
  3. Inspect Before Generating

What it can do on your machine

Read from SKILL.md and the folder at commit 4d6c552. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

MiniMax Multimodal Toolkit loads about 1.6k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 684 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from madebyaris/advance-minimax-m3-cursor-rules at commit 4d6c552, republished under its MIT licence (© madebyaris). 684 words, ~1,618 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-multimodal-toolkit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
minimax-multimodal-toolkit
description
MiniMax-native multimodal workflow for image, video, voice, music, and media-processing tasks. Use when the user asks to generate image/video/audio assets, wants MiniMax-specific media APIs, needs TTS or voice workflows, wants reproducible local media outputs, or needs FFmpeg-style processing around generated media. M3's native multimodal input means image/video inputs can be fed directly to the model for grounded decisions in coding work.
license
MIT
metadata.version
1.1.0
metadata.category
media-generation
metadata.sources
MiniMax platform media capabilities, Current runtime tool surface, FFmpeg documentation

MiniMax Multimodal Toolkit

Use MiniMax-native media workflows without bloating the always-on prompt. Route the task to the smallest path that can honestly produce the requested artifact.

When to Use

  • The user asks for image, video, voice, speech, music, or multimodal asset generation
  • The user explicitly mentions MiniMax media capabilities or wants MiniMax API integration
  • The user wants reproducible local media outputs rather than only in-chat prose
  • The task involves media conversion, trimming, concatenation, or extraction around generated assets

For deeper routing notes, output conventions, and implementation details, also read reference.md in this skill directory.


Step 0: Determine the Real Goal

Classify the task before acting:

  1. Direct asset generation: user wants an image, clip, narration, or music artifact
  2. Product integration: user wants app code that calls MiniMax media APIs
  3. Media pipeline work: user already has files and needs processing, conversion, or stitching
  4. Capability research: user wants comparison, planning, or API guidance before building

Do not jump into API integration when a direct generation path is enough.


Step 1: Route to the Right Path

User needPrimary pathNotes
One-off image assetUse the runtime's direct image-generation tool if availableFastest path for explicit image requests
Video, TTS, voice, music, or MiniMax-specific generationUse current MiniMax docs and the repo/runtime tool surfaceCheck auth and output path first
Existing media needs editingUse local tooling such as FFmpeg when availableAvoid re-generation unless needed
App feature using MiniMax media APIsImplement integration code and verify with a focused request or fixturePrefer smallest vertical slice
Planning or research onlyGather current docs and synthesizeDo not implement prematurely
M3 input path: read an attached image / video frame as ground truthFeed the file/frame into the model directly via the runtime's multimodal input — no separate "describe the image" stepUse for design parity, error UI triage, screenshot-driven dev. See the minimax-m3-multimodal-input skill for the full workflow.

Step 2: Inspect Before Generating

Before any implementation or generation:

  1. Inspect the repo for existing media patterns, asset folders, env handling, and helper utilities
  2. Check the current runtime for direct generation tools before inventing scripts
  3. Check whether required MiniMax credentials or host configuration already exist
  4. Clarify only if the missing answer changes the route:
    • output medium
    • target format
    • duration or size constraints
    • whether the user wants direct generation or product integration

Show full SKILL.md (292 more words)Show less

M3 Native Multimodal Input

On M3, image and video inputs can be fed to the model directly. This collapses the older "read the file, write a text description, then reason about the description" loop into a single grounded step:

  • The user attaches an image, screenshot, mock, or short clip; the runtime passes it to M3 as native input.
  • Ground decisions in what the image actually shows. Quote visible text, cite regions, name the file path.
  • For the full input-handling workflow (region citations, before/after diffing, multi-frame video, design parity), load the minimax-m3-multimodal-input skill.

This skill (minimax-multimodal-toolkit) remains the source of truth for generation paths — calling MiniMax media APIs, FFmpeg pipelines, and reproducible local outputs. The two skills are complementary: this one for output, minimax-m3-multimodal-input for input.

Core Rules

  • Prefer the smallest path that produces the requested artifact honestly
  • Use direct generation tools for explicit image requests when available
  • Use MiniMax-specific API flows when the user asks for MiniMax integration, reproducibility, video, TTS, voice, or music
  • Keep generated outputs in a predictable project-local folder rather than scattering temp files
  • Never hardcode secrets; use environment variables and document the missing configuration
  • Do not claim a generated asset exists until you have verified the file or response
  • For integration work, verify one focused happy-path request before broadening the feature

Verification Expectations

Match proof to the task:

  • Asset generation: verify the output file exists or the tool returned a concrete artifact
  • API integration: verify one focused request, script, or runtime flow
  • Media processing: verify the output file was created and matches the requested format or duration
  • UI integration: verify at the user surface, not only by build success

If the artifact was designed but not generated, report it as changed and unverified, not complete.


Workflow

text
1. CLASSIFY -> direct asset, integration, processing, or research
2. ROUTE -> choose direct tool, MiniMax API path, or local media tooling
3. INSPECT -> repo patterns, runtime surface, env/auth, output constraints
4. EXECUTE -> make the smallest honest slice
5. VERIFY -> prove the artifact or integration at the relevant surface

Quick Reference

text
IMAGE      -> direct image tool first when available
VIDEO/TTS  -> MiniMax-specific workflow or integration path
MUSIC      -> MiniMax-specific workflow or integration path
PROCESSING -> local media tooling, usually FFmpeg
INTEGRATE  -> smallest API slice + focused verification

ALWAYS     -> inspect runtime first, use env vars for secrets, verify outputs
NEVER      -> hardcode keys, promise files that were not produced, skip surface proof

© madebyaris, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .cursor/skills/minimax-multimodal-toolkit of madebyaris/advance-minimax-m3-cursor-rules.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 4d6c552

Compare with similar skills

MiniMax Multimodal Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

MiniMax Multimodal Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
MiniMax Multimodal Toolkit this skillmadebyaris/advance-minimax-m3-cursor-rules126—~1.6kAutomated safety check: PassMIT
Media Productionleon-ai/leon18k—~999Automated safety check: PassMIT
Qiaomu Cutjoeseesun/qiaomu-cut-skill372—~6.8kAutomated safety check: NotesMIT
Video Productionspeechlab0210/video-production-skill105—~4.1kAutomated safety check: NotesMIT
Whiteboard Videognipbao/codex-whiteboard-video-skill327—~7.2kAutomated safety check: NotesMIT
Content To Videoarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Media Production

    leon-ai/leon

    Generates images, audio and video through Leon's media tools, joins them with FFmpeg, checks the output and attaches playable files for the owner.

    18k GitHub stars~999 tokensUpdated today
    Media & CreativeAuto-check passed
  • Qiaomu Cut

    joeseesun/qiaomu-cut-skill

    把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…

    372 GitHub stars~6.8k tokensUpdated 13 days ago
    Media & CreativeAuto-check: notes
  • Video Production

    speechlab0210/video-production-skill

    AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.

    105 GitHub stars~4.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    327 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Vox Explainer

    CK42BB/vox-explainer-skill

    End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

    109 GitHub stars~2.6k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed

More from madebyaris/advance-minimax-m3-cursor-rules

  • 3D Web Experiences

    madebyaris/advance-minimax-m3-cursor-rules

    Builds 3D web scenes with Three.js and React Three Fiber, with attention to art direction, a performance budget, mobile behavior and fallbacks.

    126 GitHub stars~3k tokensUpdated 3 mo ago
    Auto-check passed
  • Deep Research Loop

    madebyaris/advance-minimax-m3-cursor-rules

    Runs multi-step research with a loop of search, compress, reflect and synthesize, scaling effort from a quick sourced answer to an exhaustive cited report.

    126 GitHub stars~2.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Incident Triage Harness

    madebyaris/advance-minimax-m3-cursor-rules

    Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

    126 GitHub stars~984 tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax M3 Long-Context Discipline

    madebyaris/advance-minimax-m3-cursor-rules

    Teaches how to work within MiniMax M3's 1M-token context: decide per source what to keep, summarize or drop, plan the loading, and cap raw blocks across iterations.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax M3 Multimodal Input

    madebyaris/advance-minimax-m3-cursor-rules

    Teaches an agent on MiniMax M3 to ground visual claims in attached images, screenshots, mockups and clips, and to re-read results after a visual change.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about MiniMax Multimodal Toolkit

What does MiniMax Multimodal Toolkit do?

Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media. This skill sends media tasks down the smallest path that can honestly produce the requested artifact, which keeps the always-on prompt light. Step 0 classifies the request as direct asset generation, product integration code, pipeline work on existing files, or capability research, and warns against jumping into API integration when direct generation is enough.

When should I use MiniMax Multimodal Toolkit?

MiniMax Multimodal Toolkit fits situations like: generating an image, clip, narration or music asset with MiniMax; integrating MiniMax media APIs into an application; converting, trimming or joining generated media with FFmpeg; using an attached screenshot as ground truth through M3 multimodal input.

How do I install MiniMax Multimodal Toolkit in Claude Code?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-multimodal-toolkit -a claude-code`. Or copy the skill folder (.cursor/skills/minimax-multimodal-toolkit in madebyaris/advance-minimax-m3-cursor-rules) into .claude/skills/minimax-multimodal-toolkit in your project. Claude Code loads it when a task matches its description.

How do I install MiniMax Multimodal Toolkit in Codex?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-multimodal-toolkit -a codex`. Or copy the skill folder (.cursor/skills/minimax-multimodal-toolkit in madebyaris/advance-minimax-m3-cursor-rules) into .agents/skills/minimax-multimodal-toolkit in your project. Codex loads it when a task matches its description.

Can I use MiniMax Multimodal Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-multimodal-toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-multimodal-toolkit, .gemini/skills/minimax-multimodal-toolkit, .github/skills/minimax-multimodal-toolkit and .opencode/skills/minimax-multimodal-toolkit in your project.

What does MiniMax Multimodal Toolkit need to run?

SKILL.md names no scripts, command-line tools or credentials: MiniMax Multimodal Toolkit is instructions for the agent only. Our summary lists: MiniMax credentials for the media APIs; FFmpeg for local media processing (optional).

Does MiniMax Multimodal Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is MiniMax Multimodal Toolkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does MiniMax Multimodal Toolkit use?

MiniMax Multimodal Toolkit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does MiniMax Multimodal Toolkit use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to MiniMax Multimodal Toolkit?

Skills that share tags, products or a category with MiniMax Multimodal Toolkit: Media Production (leon-ai/leon, 18k stars), Qiaomu Cut (joeseesun/qiaomu-cut-skill, 372 stars), Video Production (speechlab0210/video-production-skill, 105 stars) and Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 327 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains MiniMax Multimodal Toolkit?

madebyaris (a GitHub user) maintains it in madebyaris/advance-minimax-m3-cursor-rules, which has 126 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 16, 2026.

Source: madebyaris/advance-minimax-m3-cursor-rules on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.