Agent skill

Minimax Music Gen

by aiskillstore in aiskillstore/marketplace

A skill your agent uses when user wants to generate music, songs, or audio tracks.

MITAuto-check passedMedia & Creative

Install Minimax Music Gen

skills CLI
$ npx skills add aiskillstore/marketplace --skill minimax-music-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiskillstore/marketplace minimax-music-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiskillstore/marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/minimax-ai/minimax-music-gen .claude/skills/minimax-music-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-music-gen
GitHub stars
433
Token cost
~3.5k tokens
SKILL.md length
1,392 words
Files
3 (incl. references)
Skills in repo
1,044
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when user wants to generate music, songs, or audio tracks.

  • Works in 6 steps: Detect Intent → Basic Mode → Advanced Control Mode → …
  • User wants to generate music
  • SKILL.md covers Prerequisites, CLI Tool, Storage and Language & Interaction, plus 5 more sections
  • Calls npm

What it does

Minimax Music Gen is an agent skill from aiskillstore/marketplace. Use when user wants to generate music, songs, or audio tracks. Triggers on any request involving music creation, song writing, lyrics generation, audio production, or covers. Also triggers when user provides lyrics and wants them turned into a song, or describes a mood/scene and wants background music. Supports multilingual triggers — match equivalent phrases in any language. Do NOT use for music playback of existing files, music theory questions, or music recommendation without generation.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/prompt_guide.md` and `skill-report.json`).

It sits in Media & Creative, covering Music and audio generation. It works with MiniMax. The repository describes itself as: Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified. The licence is MIT.

When your agent uses it

  • User wants to generate music
  • Any request involving music creation
  • Lyrics generation
  • Audio production

Example prompts

  • “/minimax-music-gen”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Detect Intent
  2. Basic Mode
  3. Advanced Control Mode
  4. Call mmx
  5. Playback
  6. Feedback & Iteration

What it can do on your machine

Read from SKILL.md and the folder at commit 44923f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.minimaxi.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Minimax Music Gen loads about 3.5k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 1,392 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aiskillstore/marketplace at commit 44923f3, republished under its MIT licence (© aiskillstore). 1,392 words, ~3,508 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-music-gen/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
minimax-music-gen
description
Use when user wants to generate music, songs, or audio tracks. Triggers on any request involving music creation, song writing, lyrics generation, audio production, or covers. Also triggers when user provides lyrics and wants them turned into a song, or describes a mood/scene and wants background music. Supports multilingual triggers — match equivalent phrases in any language. Do NOT use for music playback of existing files, music theory questions, or music recommendation without generation.
license
MIT
metadata.version
1.1
metadata.category
creative

MiniMax Music Generation Skill

Generate songs (vocal or instrumental) using the MiniMax Music API. Supports two creation modes: Basic (one-sentence-in, song-out) and Advanced Control (edit lyrics, refine prompt, plan before generating).

Prerequisites

  • mmx CLI (required): Music generation uses the mmx command-line tool.

    Check if installed:

    bash
    command -v mmx && mmx --version || echo "mmx not found"

    Install (requires Node.js):

    bash
    npm install -g mmx-cli

    Authenticate (first time only):

    bash
    mmx auth login --api-key <your-minimax-api-key>

    The API key can be obtained from MiniMax Platform. Credentials are saved to ~/.mmx/credentials.json and persist across sessions.

    Verify:

    bash
    mmx quota show
  • Audio player (recommended): mpv, ffplay, or afplay (macOS built-in) for local playback. mpv is preferred for its interactive controls.

CLI Tool

This skill uses the mmx CLI for all music generation:

  • Music Generation: mmx music generate — model: music-2.6-free

    • Supports --lyrics-optimizer to auto-generate lyrics from prompt
    • Supports --instrumental for instrumental tracks
    • Supports --lyrics for user-provided lyrics
    • Structured params: --genre, --mood, --vocals, --instruments, --bpm, --key, --tempo, --structure, --references
  • Cover: mmx music cover — model: music-cover-free

    • Takes reference audio via --audio-file <path> or --audio <url>
    • --prompt describes the target cover style

Agent flags: Always add --quiet --non-interactive when calling mmx from agents.

Pipeline:

  • Vocal: User description -> mmx music generate --lyrics-optimizer -> MP3
  • Instrumental: User description -> mmx music generate --instrumental -> MP3
  • Cover: Source audio + style -> mmx music cover -> MP3

Storage

All generated music is saved to ~/Music/minimax-gen/. Create the directory if it doesn't exist. Files are named with a timestamp and a short slug derived from the prompt: YYYYMMDD_HHMMSS_<slug>.mp3


Language & Interaction

Detect the user's language from their first message and respond in that language for the entire session. This applies to all interaction text, questions, confirmations, and feedback prompts.

User-facing text localization rule:

  • ALL text shown to the user — including preview labels, field names, confirmations, status messages, playback info, feedback prompts, and the prompt/description preview — MUST be fully translated into the user's language.
  • The API prompt sent to the model should always be written in English for best generation quality. However, when previewing the prompt to the user, show a localized description in the user's language instead of the raw English prompt. The English prompt is an internal implementation detail — the user does not need to see it.
  • The templates below are written in English as reference. At runtime, translate every label and message into the user's detected language.

Lyrics language rule:

  • Default lyrics language = the user's language. A Chinese-speaking user gets Chinese lyrics; an English-speaking user gets English lyrics.
  • Only generate lyrics in a different language if the user explicitly requests it.
  • When a different lyrics language is needed, embed it naturally into the vocal or genre description in the prompt. For example, instead of appending "with Korean lyrics", use "featuring a Korean female vocalist" or specify a genre that implies the language (e.g., "K-pop", "J-rock", "Mandopop", "Latin pop").

Workflow

Step 0: Detect Intent

Parse the user's message to determine:

  1. Song category: vocal (with lyrics), instrumental (no vocals), or cover
  2. Creation mode preference: did they provide detailed requirements (Advanced) or a casual one-liner (Basic)?

If ambiguous, ask using this decision tree:

Q1: What type of music?
  - Vocal (with lyrics)
  - Instrumental (no vocals)
  - Cover

Q2: Creation mode?
  - Basic — one-line description, auto-generate
  - Advanced — edit lyrics, refine prompt, plan

If the user gives a clear one-liner like "make me a sad piano piece", skip the questions — infer instrumental + basic mode and proceed.


Step 1: Basic Mode

Goal: User provides a short description, the skill auto-generates everything, then calls the API.

  1. Expand the description into a prompt: Take the user's one-liner and expand it into a rich music prompt. Refer to the Prompt Writing Guide appendix at the end of this document for style vocabulary, genre/instrument references, and prompt structure. The API prompt should always be written in English for best generation quality, regardless of the user's language.

    Follow this pattern:

    A [mood] [BPM optional] [genre] song, featuring [vocal description],
    about [narrative/theme], [atmosphere], [key instruments and production].
  2. Show the user a preview before generating. Translate all labels AND the prompt description into the user's language. The English prompt is only used internally when calling the API — the user should never see it. Example template (English reference — localize everything at runtime):

    About to generate:
    Type: Vocal / Instrumental
    Description: indie folk, melancholy, acoustic guitar, gentle female voice
    Lyrics: Auto-generated (--lyrics-optimizer)
    
    Confirm? (press enter to confirm, or tell me what to change)
  3. Call mmx: Generate the music directly.


Step 2: Advanced Control Mode

Goal: User has full control over every parameter before generation.

  1. Lyrics phase:

    • If user provided lyrics: display them formatted with section markers, ask for edits. The final lyrics will be passed via --lyrics to mmx.
    • If user has a theme but no lyrics: will use --lyrics-optimizer to auto-generate.
    • Support iterative editing: "change the second chorus" -> only rewrite that section.
    • User can also write lyrics themselves and pass via --lyrics.
  2. Prompt phase:

    • Generate a recommended prompt based on the lyrics' mood and content.
    • Present it as editable tags the user can add/remove/modify.
    • Refer to the Prompt Writing Guide appendix for the full vocabulary.
  3. Advanced planning (optional, offer but don't force):

    • Song structure: verse-chorus-verse-chorus-bridge-chorus or custom
    • BPM suggestion (encode in prompt as tempo descriptor)
    • Reference style: "something like X style" -> map to prompt tags
    • Vocal character description
  4. Final confirmation: Show complete parameter summary, then generate.


Show full SKILL.md (598 more words)Show less
Step 3: Call mmx

Generate music using the mmx CLI:

Vocal with auto-generated lyrics:

bash
mmx music generate \
  --prompt "<prompt>" \
  --lyrics-optimizer \
  --genre "<genre>" --mood "<mood>" --vocals "<vocal style>" \
  --instruments "<instruments>" --bpm <bpm> \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

Vocal with user-provided lyrics:

bash
mmx music generate \
  --prompt "<prompt>" \
  --lyrics "<lyrics with section markers>" \
  --genre "<genre>" --mood "<mood>" --vocals "<vocal style>" \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

Instrumental (no vocal):

bash
mmx music generate \
  --prompt "<prompt>" \
  --instrumental \
  --genre "<genre>" --mood "<mood>" --instruments "<instruments>" \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

Use structured flags (--genre, --mood, --vocals, --instruments, --bpm, --key, --tempo, --structure, --references, --avoid, --use-case) to give the API fine-grained control instead of cramming everything into --prompt.

Display a progress indicator while waiting. Typical generation takes 30-120 seconds.


Step 4: Playback

After generation, detect an available audio player and play the file.

Detect player:

bash
command -v mpv || command -v ffplay || command -v afplay

Play based on detected player (in priority order):

PlayerCommandControls
mpv (preferred)mpv --no-video ~/Music/minimax-gen/<filename>.mp3space = pause/resume, q = quit, left/right = seek
ffplayffplay -nodisp -autoexit ~/Music/minimax-gen/<filename>.mp3q = quit
afplay (macOS)afplay ~/Music/minimax-gen/<filename>.mp3Ctrl+C = stop
None foundDo not attempt playbackShow file path only

After starting playback, tell the user (localize all text):

Now playing: <filename>.mp3
Saved to: ~/Music/minimax-gen/<filename>.mp3

Do NOT show playback controls (e.g. keyboard shortcuts) — they don't work in this environment since the player runs in the background.

If no player is found (localize all text):

No audio player detected.
File saved to: ~/Music/minimax-gen/<filename>.mp3
Tip: Install mpv for the best playback experience (brew install mpv).

Step 5: Feedback & Iteration

After playback, ask for feedback:

How was this song?
  1. Love it, keep it!
  2. Not quite, adjust and regenerate
  3. Fine-tune lyrics/style then regenerate
  4. Don't want it, start over

Based on feedback:

  • Satisfied: Done. Mention the file path again.
  • Adjust & regenerate: Ask what to change (prompt? lyrics? style?), apply edits, re-run generation. Keep the old file with a _v1 suffix for comparison.
  • Fine-tune: Enter Advanced Control Mode with the current parameters pre-filled.
  • Delete & restart: Remove the file, go back to Step 0.

Cover Mode

Generate a cover version of a song based on reference audio. Model: music-cover-free.

Reference audio requirements: mp3, wav, flac — duration 6s to 6min, max 50MB. If no lyrics are provided, the original lyrics are extracted via ASR automatically.

Workflow

When the user selects Cover mode:

  1. Ask for the source audio — a local file path or URL
  2. Ask for the target cover style (e.g., "acoustic cover, stripped-down, intimate vocal")
  3. Optionally ask for custom lyrics or lyrics file
Commands

Cover from local file:

bash
mmx music cover \
  --prompt "<cover style description>" \
  --audio-file <source.mp3> \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

Cover from URL:

bash
mmx music cover \
  --prompt "<cover style description>" \
  --audio <source_url> \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

With custom lyrics (text):

bash
mmx music cover \
  --prompt "<style>" \
  --audio-file <source.mp3> \
  --lyrics "<custom lyrics>" \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive

With custom lyrics (file):

bash
mmx music cover \
  --prompt "<style>" \
  --audio-file <source.mp3> \
  --lyrics-file <lyrics.txt> \
  --out ~/Music/minimax-gen/<filename>.mp3 \
  --quiet --non-interactive
Optional flags
FlagDescription
--seed <number>Random seed 0-1000000 for reproducible results
--channel <n>1 (mono) or 2 (stereo, default)
--format <fmt>mp3 (default), wav, pcm
--sample-rate <hz>Sample rate (default: 44100)
--bitrate <bps>Bitrate (default: 256000)
After generation

Proceed with normal playback and feedback flow (Step 4 & 5).


Error Handling

ErrorAction
mmx not foundnpm install -g mmx-cli
mmx auth error (exit code 3)mmx auth login
Quota exceeded (exit code 4)Report quota limit, suggest waiting or upgrading
API timeout (exit code 5)Retry once, then report failure
Content filter (exit code 10)Adjust prompt to avoid filtered content
Invalid lyrics formatAuto-fix section markers, warn user
No audio player foundSave file and tell user the path, suggest installing mpv
Network errorShow error detail, suggest checking connection

Important Notes

  • Never reproduce copyrighted lyrics. When doing covers, always write original lyrics inspired by the song's theme. Explain this to the user.
  • Prompt language: The API prompt works best with English tags. Chinese tags are also acceptable. Mixing is OK.
  • Section markers in lyrics: The API recognizes [verse], [chorus], [bridge], [outro], [intro]. Always include them when providing --lyrics.
  • File management: If ~/Music/minimax-gen/ has more than 50 files, suggest cleanup when starting a new session.
  • Structured params: Prefer using --genre, --mood, --vocals, --instruments, --bpm etc. over embedding everything in --prompt. This gives the API better control.
  • Lyrics language via style: When the user wants lyrics in a specific language, express it through the vocal description or genre (e.g., "Japanese female vocalist", "Mandopop ballad") rather than appending a language directive to the prompt.

Appendix: Prompt Writing Guide

See references/prompt_guide.md for the complete prompt writing guide, including genre/vocal/instrument references and BPM tables.

© aiskillstore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/minimax-ai/minimax-music-gen of aiskillstore/marketplace.

  • SKILL.md
  • references/prompt_guide.md
  • skill-report.json

Open the folder on GitHubat commit 44923f3

Compare with similar skills

Minimax Music Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Minimax Music Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Minimax Music Gen this skillaiskillstore/marketplace433—~3.5kAutomated safety check: PassMIT
Music Caption RewriterT8mars/T8-penguin-canvas615—~2.2kAutomated safety check: PassMIT
Mm Music ExpertLeoYeAI/openclaw-master-skills2.2k—~4.6kAutomated safety check: PassMIT
Venice Audio Musicveniceai/skills144—~3.1kAutomated safety check: PassMIT
Videoguaardvark/guaardvark258—~1.2kAutomated safety check: PassMIT
Audio Jinglesanqiufong/slides-from-anything1321 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Music Caption Rewriter

    T8mars/T8-penguin-canvas

    Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.

    615 GitHub stars~2.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Mm Music Expert

    LeoYeAI/openclaw-master-skills

    Create music with MiniMax music models (music-2.5+, music-2.5).

    2.2k GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Video

    guaardvark/guaardvark

    Generate video clips on the user's own GPU through Guaardvark: text-to-video, image-to-video, first+last frame animation, clips with their own soundtrack and dialogue (MiniMax H3), short looping…

    258 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Audio Jingle

    sanqiufong/slides-from-anything

    Audio generation skill — jingles, beds, voiceover, and sound effects.

    132 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Video Audio Continuity

    nodetool-ai/nodetool

    Keep sound continuous across a multi-scene piece cut from generated video — why one clip per scene hard-cuts the audio at every boundary, when to write all the scenes into a single generation…

    560 GitHub stars~1.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from aiskillstore/marketplace

All 1,040 skills in this repo
  • Code Stats

    aiskillstore/marketplace

    Analyze codebase with tokei (fast line counts by language) and difft (semantic AST-aware diffs).

    433 GitHub starsUsed in 1 repo~697 tokens
    Auto-check: notes
  • Data Processing

    aiskillstore/marketplace

    Process JSON with jq and YAML/TOML with yq. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 1 repo~720 tokens
    Auto-check: notes
  • Doc Scanner

    aiskillstore/marketplace

    Scans for project documentation files (AGENTS.md, CLAUDE.md, GEMINI.md, COPILOT.md, CURSOR.md, WARP.md, and 15+ other formats) and synthesizes guidance.

    433 GitHub starsUsed in 1 repo~644 tokens
    Auto-check: notes
  • File Search

    aiskillstore/marketplace

    Modern file and content search using fd, ripgrep (rg), and fzf.

    433 GitHub starsUsed in 1 repo~598 tokens
    Auto-check: notes
  • Find Replace

    aiskillstore/marketplace

    Modern find-and-replace using sd (simpler than sed) and batch replacement patterns.

    433 GitHub starsUsed in 1 repo~527 tokens
    Auto-check: notes
  • Project Planner

    aiskillstore/marketplace

    Detects stale project plans and suggests session commands. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 1 repo~504 tokens
    Auto-check passed

Works with

Questions about Minimax Music Gen

What does Minimax Music Gen do?

A skill your agent uses when user wants to generate music, songs, or audio tracks. Minimax Music Gen is an agent skill from aiskillstore/marketplace. Use when user wants to generate music, songs, or audio tracks.

When should I use Minimax Music Gen?

Minimax Music Gen fits situations like: user wants to generate music; any request involving music creation; lyrics generation; audio production.

How do I install Minimax Music Gen in Claude Code?

Run `npx skills add aiskillstore/marketplace --skill minimax-music-gen -a claude-code`. Or copy the skill folder (skills/minimax-ai/minimax-music-gen in aiskillstore/marketplace) into .claude/skills/minimax-music-gen in your project. Claude Code loads it when a task matches its description.

How do I install Minimax Music Gen in Codex?

Run `npx skills add aiskillstore/marketplace --skill minimax-music-gen -a codex`. Or copy the skill folder (skills/minimax-ai/minimax-music-gen in aiskillstore/marketplace) into .agents/skills/minimax-music-gen in your project. Codex loads it when a task matches its description.

Can I use Minimax Music Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiskillstore/marketplace --skill minimax-music-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-music-gen, .gemini/skills/minimax-music-gen, .github/skills/minimax-music-gen and .opencode/skills/minimax-music-gen in your project.

What does Minimax Music Gen need to run?

Going by SKILL.md and its folder, Minimax Music Gen needs the command-line tools its instructions call (npm). Our summary lists: Node.js.

Does Minimax Music Gen access the network?

SKILL.md names 1 domain. As links in the text: platform.minimaxi.com. This is read from the text; nothing was executed.

Is Minimax Music Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Minimax Music Gen use?

Minimax Music Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Minimax Music Gen use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Minimax Music Gen?

Skills that share tags, products or a category with Minimax Music Gen: Music Caption Rewriter (T8mars/T8-penguin-canvas, 615 stars), Mm Music Expert (LeoYeAI/openclaw-master-skills, 2.2k stars), Venice Audio Music (veniceai/skills, 144 stars) and Video (guaardvark/guaardvark, 258 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Minimax Music Gen?

aiskillstore (a GitHub organization) maintains it in aiskillstore/marketplace, which has 433 GitHub stars. The repository holds 1,044 skills in this directory. The repository was last updated on October 10, 2026.

Source: aiskillstore/marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.