Agent skill

MiniMax M3 Multimodal Input

by madebyaris in madebyaris/advance-minimax-m3-cursor-rules

Teaches an agent on MiniMax M3 to ground visual claims in attached images, screenshots, mockups and clips, and to re-read results after a visual change.

MITAuto-check passedFrontend & Design

Install MiniMax M3 Multimodal Input

skills CLI
$ npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install madebyaris/advance-minimax-m3-cursor-rules minimax-m3-multimodal-input --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/madebyaris/advance-minimax-m3-cursor-rules.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/minimax-m3-multimodal-input .claude/skills/minimax-m3-multimodal-input && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-m3-multimodal-input
GitHub stars
126
Token cost
~1.6k tokens
SKILL.md length
776 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Teaches an agent on MiniMax M3 to ground visual claims in attached images, screenshots, mockups and clips, and to re-read results after a visual change.

  • Works in 6 steps: Inventory The Input → Ground In The File, Not The Prose → Visual-Fidelity Claims → …
  • Attaching a screenshot of a UI bug and asking why it looks wrong
  • SKILL.md covers When to Use, Step 0: Inventory The Input, Step 1: Ground In The File,… and Step 2: Visual-Fidelity Claims, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

MiniMax M3 accepts text, image and video as native input, and this skill makes the agent use that honestly. It inventories what you attached, reads the actual file, cites its path and the region involved, and quotes visible text such as error messages and button labels exactly. Your own description of an image is treated as a hint, and if the agent cannot see a file it must say so rather than invent contents.

Video or screen recordings are handled through a few representative frames, such as start, middle and end, with the motion between them reasoned about explicitly. Sets such as desktop, tablet and mobile mockups, or before and after screenshots, are each re-read before the agent claims a responsive match or a fix. Any claim that something looks right or matches a design needs a grounded check. For creating images, video or audio, the separate `minimax-multimodal-toolkit` skill applies instead.

When your agent uses it

  • Attaching a screenshot of a UI bug and asking why it looks wrong
  • Matching a page to a design mockup or reference clip
  • Reading an error message from a screenshot
  • Confirming a styling fix by re-reading the result

Example prompts

  • “Make the pricing page look like this mockup. The PNG is attached.”
  • “Why does the header look broken in the attached screenshot?”
  • “Compare the before and after screenshots and tell me if the layout bug is fixed.”
  • “Read the error in the attached screen recording and tell me which step fails.”

Requirements

  • A MiniMax M3 model that accepts image or video input
  • Image, screenshot or clip files the agent can open

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inventory The Input
  2. Ground In The File, Not The Prose
  3. Visual-Fidelity Claims
  4. Design Parity (Mock → Code)
  5. Error UI / Bug Reports
  6. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit 4d6c552. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

MiniMax M3 Multimodal Input loads about 1.6k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 776 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from madebyaris/advance-minimax-m3-cursor-rules at commit 4d6c552, republished under its MIT licence (© madebyaris). 776 words, ~1,615 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-m3-multimodal-input/SKILL.md (or your agent's skills folder).
name
minimax-m3-multimodal-input
description
How to use MiniMax M3's native multimodal input (image, video) for grounded decisions in coding work. Covers reading attached images/frames, treating them as ground truth for visual claims, screenshot diffing, design parity from mockups, and routing visual evidence through reports and PRs. Load when the user attaches an image, screenshot, mockup, frame, or short clip — or when the task involves "make it look like this", "match this design", "why does this UI look wrong", or "read this error screenshot".
license
MIT
metadata.version
1.0.0
metadata.category
workflow
metadata.sources
MiniMax M3 multimodal input documentation, Screenshot-driven development practice

M3 Multimodal Input

M3 accepts text + image + video as native input. The point of this skill is to use that capability honestly: ground every visual claim in the actual file the model can read, and re-read the post-change state before declaring a visual fix done.

When to Use

  • The user attaches an image, screenshot, mockup, frame, or short clip.
  • The task involves "make it look like this", "match this design", "why does this UI look wrong", "read this error screenshot", or "match this reference video".
  • You are about to claim a visual or styling result. Any visual claim without a multimodal-grounded read is a guess.
  • A bug report or feature request references a UI state the user can show you (rather than describe in words).

If the user is only talking about generating media (creating new images, video, TTS, music), the minimax-multimodal-toolkit skill is the right one — that skill is for output; this one is for input.

Step 0: Inventory The Input

Before reasoning, identify what the user attached and what you can actually read:

  • Static image (PNG / JPG / WebP / SVG) — Read the file path the user provided or that the runtime surfaces.
  • Multi-frame video or screen recording — pick a small number of representative frames (start, mid, end) and reason about the in-between motion explicitly.
  • Inline attachment that the runtime renders into the chat — the model sees it directly; cite it as "the attached image" with the visible region.
  • Multiple images in a set (desktop + tablet + mobile mockups, before + after screenshots) — re-read each before claiming a responsive match or a fix.

If you cannot actually see the file, say so and ask. Do not invent the contents of an image you did not open.

Step 1: Ground In The File, Not The Prose

Always Read the file/frame and reference exact paths. Do not paraphrase a guessed description.

  • Cite the file path (/path/to/screenshot.png) and, when relevant, the region (top-right nav, hero block, error toast).
  • Quote visible text directly — error messages, button labels, empty-state copy. Do not paraphrase.
  • If the user described the image in prose, treat their prose as a hint, not as evidence. The image is the evidence.

Step 2: Visual-Fidelity Claims

Any time the user (or you) say "looks right", "matches", "fixed", or "as designed", you owe a multimodal-grounded proof per the always-on status rule.

Shape:

text
Visual claim: [what you say it looks like]
Reference (pre / target): [file path]
Actual (post / current):   [file path]
Region inspected:           [where in the frame you looked]
Verdict:                    [matches / does not match / partial — say which and why]

If you cannot re-read the post-change state, downgrade the claim to unverified — do not say "fixed" without seeing the fix.

Show full SKILL.md (369 more words)Show less

Step 3: Design Parity (Mock → Code)

For "match this mock" work, use the image as the contract:

  • Identify the regions of the mock: hero, nav, primary CTA, footer, repeating components. Reason per region, not per pixel.
  • Match the mock's intent, not its exact pixel values, when the project has its own design tokens. Override tokens only when the mock is high-fidelity and the user wants a pixel match.
  • For multi-resolution mockups, hold the layout intent constant; let the responsive behavior do the work.
  • For dark/light mode parity, read both the dark and light mock frames before declaring a mode-aware fix.

Cite the mock's file path in the PR / commit message so the next reviewer can re-read the contract.

Step 4: Error UI / Bug Reports

When the user attaches a screenshot or clip of a broken UI:

  • Read it. Quote the visible error or broken text directly in the bug writeup.
  • Name the file path of the screenshot/clip in the report.
  • If the bug is interactive, ask the user to attach a short screen recording (or trigger the interaction and screenshot the result yourself in a headless browser when possible).
  • Do not invent text that is not in the image. "Likely says 'Connection failed'" without reading it is an inference, not evidence.

Step 5: Verification

After a UI change:

  1. Render or capture the post-change state (screenshot, frame, or browser snapshot).
  2. Read the post-change file in the current session.
  3. Compare to the reference (pre-change state, mock, or expected behavior).
  4. State the verdict with the four fields from Step 2.
  5. If the verdict is "does not match" or "partial", say what specifically diverged and the smallest next change.

This is the loop the always-on minimax-m3-status-verification rule calls multimodal-grounded. Skipping the post-change read is the most common visual-fidelity failure.

Anti-Patterns

  • Describing an image you did not actually open. If the runtime could not surface the file, say so.
  • "Looks right" without re-reading the post-change state.
  • Paraphrasing visible text instead of quoting it.
  • Treating a single frame as evidence for a multi-frame or animated UI.
  • Claiming a responsive match when you only re-read the desktop frame.
  • Inventing colors / spacing / typography instead of reading them from the mock.

Quick Reference

text
INVENTORY  -> image, multi-frame video, inline attachment, multi-image set?
GROUND     -> Read the file; cite the path; quote visible text
CLAIM      -> any visual claim needs a reference path + actual path + region + verdict
PARITY     -> mock is the contract; match intent by region; cite the path in the PR
ERROR UI   -> read the screenshot/clip; quote visible text; name the path
VERIFY     -> render/capture post-change; re-read; compare; state verdict

© madebyaris, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/minimax-m3-multimodal-input of madebyaris/advance-minimax-m3-cursor-rules.

Open the folder on GitHubat commit 4d6c552

Compare with similar skills

MiniMax M3 Multimodal Input next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

MiniMax M3 Multimodal Input compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
MiniMax M3 Multimodal Input this skillmadebyaris/advance-minimax-m3-cursor-rules126—~1.6kAutomated safety check: PassMIT
Pixel Perfectbitovi/ai-enablement-prompts121—~2.6kAutomated safety check: PassMIT
Pixel Perfect ReplicationYu-369/VibeCurb979—~8.7kAutomated safety check: PassMIT
Figma To StaticLeoYeAI/openclaw-master-skills2.2k—~7.1kAutomated safety check: PassMIT
Figma Design Handoffyonatangross/orchestkit292—~2.7kAutomated safety check: PassMIT
Replicaadobe/skills197—~8.9kAutomated safety check: PassApache-2.0

Similar skills

  • Pixel Perfect

    bitovi/ai-enablement-prompts

    Orchestrate a complete visual QA workflow to make a dev/Storybook page match a baseline URL pixel-for-pixel.

    121 GitHub stars~2.6k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Image-to-code replication pipeline. An agent skill from Yu-369/VibeCurb.

    979 GitHub stars~8.7k tokensUpdated 2 mo ago
    DatabasesAuto-check passed
  • Figma To Static

    LeoYeAI/openclaw-master-skills

    Convert Figma design files to pixel-level mobile-first static HTML/CSS pages.

    2.2k GitHub stars~7.1k tokensUpdated 2 mo ago
    Frontend & DesignAuto-check passed
  • Figma Design Handoff

    yonatangross/orchestkit

    Figma-to-code design handoff patterns including Figma Variables to design tokens pipeline, component spec extraction, Dev Mode inspection, Auto Layout to CSS Flexbox/Grid mapping, and visual…

    292 GitHub stars~2.7k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Replica

    adobe/skills

    Same-design migration — re-platform a site to AEM Edge Delivery (or any clean front end) keeping its current design near pixel-perfect.

    197 GitHub stars~8.9k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Control UI

    cursor/plugins

    Official

    Build or adapt a local browser/CDP harness to drive and inspect a web, IDE, or Electron UI.

    11k GitHub starsUsed in 2 repos~1.2k tokens
    Frontend & DesignAuto-check passed

More from madebyaris/advance-minimax-m3-cursor-rules

  • 3D Web Experiences

    madebyaris/advance-minimax-m3-cursor-rules

    Builds 3D web scenes with Three.js and React Three Fiber, with attention to art direction, a performance budget, mobile behavior and fallbacks.

    126 GitHub stars~3k tokensUpdated 3 mo ago
    Auto-check passed
  • Deep Research Loop

    madebyaris/advance-minimax-m3-cursor-rules

    Runs multi-step research with a loop of search, compress, reflect and synthesize, scaling effort from a quick sourced answer to an exhaustive cited report.

    126 GitHub stars~2.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Incident Triage Harness

    madebyaris/advance-minimax-m3-cursor-rules

    Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

    126 GitHub stars~984 tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax M3 Long-Context Discipline

    madebyaris/advance-minimax-m3-cursor-rules

    Teaches how to work within MiniMax M3's 1M-token context: decide per source what to keep, summarize or drop, plan the loading, and cap raw blocks across iterations.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • MiniMax Multimodal Toolkit

    madebyaris/advance-minimax-m3-cursor-rules

    Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media.

    126 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about MiniMax M3 Multimodal Input

What does MiniMax M3 Multimodal Input do?

Teaches an agent on MiniMax M3 to ground visual claims in attached images, screenshots, mockups and clips, and to re-read results after a visual change. MiniMax M3 accepts text, image and video as native input, and this skill makes the agent use that honestly. It inventories what you attached, reads the actual file, cites its path and the region involved, and quotes visible text such as error messages and button labels exactly.

When should I use MiniMax M3 Multimodal Input?

MiniMax M3 Multimodal Input fits situations like: attaching a screenshot of a UI bug and asking why it looks wrong; matching a page to a design mockup or reference clip; reading an error message from a screenshot; confirming a styling fix by re-reading the result.

How do I install MiniMax M3 Multimodal Input in Claude Code?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input -a claude-code`. Or copy the skill folder (.cursor/skills/minimax-m3-multimodal-input in madebyaris/advance-minimax-m3-cursor-rules) into .claude/skills/minimax-m3-multimodal-input in your project. Claude Code loads it when a task matches its description.

How do I install MiniMax M3 Multimodal Input in Codex?

Run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input -a codex`. Or copy the skill folder (.cursor/skills/minimax-m3-multimodal-input in madebyaris/advance-minimax-m3-cursor-rules) into .agents/skills/minimax-m3-multimodal-input in your project. Codex loads it when a task matches its description.

Can I use MiniMax M3 Multimodal Input in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-m3-multimodal-input, .gemini/skills/minimax-m3-multimodal-input, .github/skills/minimax-m3-multimodal-input and .opencode/skills/minimax-m3-multimodal-input in your project.

What does MiniMax M3 Multimodal Input need to run?

SKILL.md names no scripts, command-line tools or credentials: MiniMax M3 Multimodal Input is instructions for the agent only. Our summary lists: A MiniMax M3 model that accepts image or video input; Image, screenshot or clip files the agent can open.

Does MiniMax M3 Multimodal Input access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is MiniMax M3 Multimodal Input safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does MiniMax M3 Multimodal Input use?

MiniMax M3 Multimodal Input is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does MiniMax M3 Multimodal Input use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to MiniMax M3 Multimodal Input?

Skills that share tags, products or a category with MiniMax M3 Multimodal Input: Pixel Perfect (bitovi/ai-enablement-prompts, 121 stars), Pixel Perfect Replication (Yu-369/VibeCurb, 979 stars), Figma To Static (LeoYeAI/openclaw-master-skills, 2.2k stars) and Figma Design Handoff (yonatangross/orchestkit, 292 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains MiniMax M3 Multimodal Input?

madebyaris (a GitHub user) maintains it in madebyaris/advance-minimax-m3-cursor-rules, which has 126 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 16, 2026.

Source: madebyaris/advance-minimax-m3-cursor-rules on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.