Agent skill

Music Caption Rewriter

by T8mars in T8mars/T8-penguin-canvas

Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.

MITAuto-check passedMedia & Creative

Install Music Caption Rewriter

skills CLI
$ npx skills add T8mars/T8-penguin-canvas --skill music-caption-rewriter -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install T8mars/T8-penguin-canvas music-caption-rewriter --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/T8mars/T8-penguin-canvas.git skills-src && mkdir -p .claude/skills && cp -r skills-src/public/official-skills/minimax-music3/music-caption-rewriter .claude/skills/music-caption-rewriter && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
music-caption-rewriter
GitHub stars
615
Token cost
~2.2k tokens
SKILL.md length
1,179 words
Files
1,023 (incl. references)
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.

  • Works in 8 steps: Build a private Music Brief from the… → Resolve explicit constraints and… → Read references/genre-router.md. → …
  • Users ask to enhance a music-generation prompt
  • SKILL.md covers Inputs, Workflow, Build the Music Brief and Resolve Constraints, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Music Caption Rewriter is an agent skill from T8mars/T8-penguin-canvas. Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement. Use when users ask to enhance a music-generation prompt, preserve lyric-section directives, retrieve a similar style from bundled templates, fuse styles, or produce JSON or JSONL caption output.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1024 other files, including reference files (for example `README.md`, `agents/openai.yaml` and `references/genre-router.md`).

It sits in Media & Creative, covering Music and audio generation. It works with MiniMax. The licence is MIT.

When your agent uses it

  • Users ask to enhance a music-generation prompt
  • Preserve lyric-section directives
  • Retrieve a similar style from bundled templates
  • JSONL caption output

Example prompts

  • “/music-caption-rewriter”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Build a private Music Brief from the inputs.
  2. Resolve explicit constraints and section-local directives.
  3. Read references/genre-router.md.
  4. Read one primary family index and, only when useful, one secondary family index.
  5. Select up to three references with distinct roles.
  6. Read only the complete template files named by those cards.
  7. Design a coherent section-by-section timeline.
  8. Render and validate the new caption.

What it can do on your machine

Read from SKILL.md and the folder at commit 7ba4ede. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Music Caption Rewriter loads about 2.2k tokens when it runs, and up to ~125k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 1,179 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~125k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from T8mars/T8-penguin-canvas at commit 7ba4ede, republished under its MIT licence (© T8mars). 1,179 words, ~2,240 tokens.

Download SKILL.mdSave it as .claude/skills/music-caption-rewriter/SKILL.md (or your agent's skills folder). This skill also uses 1022 other files; get the full folder from GitHub.
name
music-caption-rewriter
description
Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement. Use when users ask to enhance a music-generation prompt, preserve lyric-section directives, retrieve a similar style from bundled templates, fuse styles, or produce JSON or JSONL caption output.

Music Caption Rewriter

Transform the user's musical intent into a new, generation-oriented structured caption. Find useful references through progressive disclosure: route to a small style family, compare compact cards, then read only the selected complete templates.

Use natural-language reasoning and local text files only. Do not execute scripts, build a database, calculate embeddings, call external APIs, or scan all 1,000 templates.

Inputs

Accept:

  • Caption: required natural-language music description.
  • Lyrics: optional lyrics containing bracketed section or control tags.
  • Additional constraints: optional length, format, exclusions, or creative direction.

Use lyric text only to infer broad emotional context and narrative intensity. Never quote, paraphrase, summarize, or reproduce it. Treat only bracketed tags as executable structural, musical, vocal, or production directives.

Workflow

Follow these stages in order:

  1. Build a private Music Brief from the inputs.
  2. Resolve explicit constraints and section-local directives.
  3. Read references/genre-router.md.
  4. Read one primary family index and, only when useful, one secondary family index.
  5. Select up to three references with distinct roles.
  6. Read only the complete template files named by those cards.
  7. Design a coherent section-by-section timeline.
  8. Render and validate the new caption.

Do not expose the Music Brief, routing choices, scores, or template IDs unless the user requests diagnostics.

Build the Music Brief

Extract only supported or reasonably inferred values:

  • macro genre, subgenres, and cultural or market style
  • mood and emotional arc
  • approximate tempo, meter, and groove
  • vocal presence, gender, register, timbre, and delivery
  • core instruments and production texture
  • section structure and section-specific changes
  • spatial character and explicit exclusions

Classify each value internally as explicit, tagged, inferred, or unspecified.

Do not invent a precise key, BPM, vocal gender, melodic interval, or production technique when a broader description is sufficient.

Preserve an explicit instrumental request. Do not add vocals. If vocal presence is unspecified, choose a conservative treatment supported by the user's description and the closest style family.

Resolve Constraints

Apply this precedence:

  1. Explicit user requirements and exclusions.
  2. Section-local directives from lyric tags, within that section.
  3. Strong implications from the user's Caption.
  4. Selected reference characteristics.
  5. Conservative musical defaults.

A section tag may change its local arrangement without replacing the song's global genre. Preserve a hard user exclusion when a tag conflicts with it.

When two explicit instructions conflict, prefer the more specific and later instruction if the intent remains clear. Otherwise make the smallest musically coherent compromise.

Never silently reverse an explicit vocal gender, instrumental requirement, tempo limit, required instrument, or prohibited element.

Route by Progressive Disclosure

Read the genre router first. Choose:

  • one primary family for a clear genre request
  • one primary and one secondary family for an explicit fusion
  • at most two plausible families for an ambiguous genre
  • the general pop and ballad family when only mood or imagery is available

Use genre, groove, instrumentation, and cultural context as stronger routing signals than generic adjectives such as emotional, epic, dark, or modern.

Read only the family indexes selected by the router. Do not inspect every family index, reconstruct a global catalog, or scan every template filename.

Select References

Compare cards in the selected family indexes using this priority:

  1. Genre and subgenre compatibility.
  2. Explicit requirements and exclusions.
  3. Groove and tempo compatibility, including plausible half-time or double-time relationships.
  4. Vocal configuration.
  5. Instrumentation.
  6. Mood and emotional arc.
  7. Production character.

Apply a strong penalty to direct conflicts. Prefer a close musical family over a card that merely shares mood vocabulary.

Select up to three references with different responsibilities:

  • Foundation: closest overall identity, groove, and songwriting language.
  • Modifier: best source for a requested secondary genre, vocal character, cultural color, or production texture.
  • Arrangement: best source for section development, energy contour, transitions, and instrument lifecycle.

Use one or two references when the request is simple. Do not select a weak match merely to reach three.

Use Templates Safely

Use the Foundation for broad musical identity, the Modifier only for its matched dimension, and the Arrangement reference only for timeline logic.

Do not inherit unsupported details such as a template's exact key, BPM, vocalist, instruments, emotional story, or section order.

Do not copy sentences, distinctive phrases, or a template's complete structure. Synthesize a new caption around the user's brief.

Show full SKILL.md (478 more words)Show less

Plan the Timeline

Build around the user's section tags when present. Otherwise choose only sections appropriate to the style, for example:

Intro → Verse → Pre-Chorus → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro

For every included section, state what enters, exits, changes, or intensifies. Keep instrument behavior continuous and make transitions musically plausible.

Create a readable energy arc rather than a static equipment list or a stack of production terminology.

Output Contract

Write the final caption in English unless the user explicitly requests another language.

Return exactly these three top-level headings in this order:

Global Metadata

Include genre and subgenres, tempo, emotional progression, and overall sonic and production profile. Use an exact BPM only when explicit or strongly justified; otherwise use a range or qualitative tempo. Include key and scale only when explicit or musically useful.

Vocal Details

For vocal music, describe the lead configuration, timbre, register, delivery, harmony or backing vocals, and restrained vocal effects.

For instrumental music, state that the piece is instrumental and identify the instrument or texture carrying the lead melodic role.

Do not invent lyrical subject matter or reproduce lyrics.

Arrangement

Describe the song as a section-by-section timeline. Explain primary and secondary instrument lifecycles, groove development, transitions, embellishments, texture, and spatial effects only where relevant.

Prefer concrete musical changes over decorative prose. Default to approximately 250–450 English words unless the user requests another length.

Do not include a song title, track ID, selected template ID, reasoning trace, or copied lyric line.

Machine-Readable Output

Return JSON or JSONL only when explicitly requested. Include original inputs and rewritten_caption. Include routing diagnostics or selected template IDs only when explicitly requested.

Never include complete template contents in machine-readable output unless the user specifically asks for them.

Validate Before Returning

Verify that:

  • every explicit user constraint is preserved
  • every actionable section tag appears in the matching section
  • no quoted, paraphrased, or summarized lyric content, title, or track ID appears
  • an instrumental request remains instrumental
  • vocal gender is not contradicted
  • genre and local modifiers coexist coherently
  • the three required headings are present
  • the arrangement follows a readable timeline
  • instruments have coherent entrances, changes, and exits
  • exact BPM, key, and technical details are not fabricated
  • no template sentence or complete template structure is copied
  • the caption is specific enough to guide generation without becoming an essay

Revise once when any check fails, then return only the corrected result.

Static Library Maintenance

Keep the library entirely text-based. When adding a template:

  1. Add one complete Caption file under templates/.
  2. Add one compact card to exactly one family index linked from the genre router.
  3. Record compatible secondary families in that card instead of duplicating it.
  4. Confirm that the card ID matches the template filename and that its path exists.
  5. Update the family count in that index.

Do not add scripts, generated catalogs, embeddings, vector stores, databases, or external service configuration.

© T8mars, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1,022 other files (references) in public/official-skills/minimax-music3/music-caption-rewriter of T8mars/T8-penguin-canvas.

  • SKILL.md
  • README.md
  • agents/openai.yaml
  • references/genre-router.md
  • references/index-cinematic-orchestral-epic.md
  • references/index-cinematic-pop-ballad.md
  • references/index-club-edm-house-trance.md
  • references/index-contemporary-folk-acoustic.md
  • references/index-country-americana.md
  • references/index-dance-pop-disco-funk.md
  • references/index-east-asian-ballad-heritage.md
  • references/index-east-asian-modern.md
  • references/index-electronic-synth-ambient-pop.md
  • references/index-general-pop-ballad.md
  • references/index-hip-hop-rap.md
  • references/index-jazz-swing-big-band.md
  • references/index-metal-heavy-rock.md
  • references/index-modern-rnb-neo-soul.md
  • references/index-pop-alternative-rock.md
  • … and 1,004 more

Open the folder on GitHubat commit 7ba4ede

Compare with similar skills

Music Caption Rewriter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Music Caption Rewriter compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Music Caption Rewriter this skillT8mars/T8-penguin-canvas615—~2.2kAutomated safety check: PassMIT
Mm Music ExpertLeoYeAI/openclaw-master-skills2.2k—~4.6kAutomated safety check: PassMIT
Minimax Music Genaiskillstore/marketplace433—~3.5kAutomated safety check: PassMIT
Venice Audio Musicveniceai/skills144—~3.1kAutomated safety check: PassMIT
Videoguaardvark/guaardvark257—~1.2kAutomated safety check: PassMIT
Audio Jinglesanqiufong/slides-from-anything1321 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Mm Music Expert

    LeoYeAI/openclaw-master-skills

    Create music with MiniMax music models (music-2.5+, music-2.5).

    2.2k GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Minimax Music Gen

    aiskillstore/marketplace

    A skill your agent uses when user wants to generate music, songs, or audio tracks.

    433 GitHub stars~3.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Video

    guaardvark/guaardvark

    Generate video clips on the user's own GPU through Guaardvark: text-to-video, image-to-video, first+last frame animation, clips with their own soundtrack and dialogue (MiniMax H3), short looping…

    257 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Audio Jingle

    sanqiufong/slides-from-anything

    Audio generation skill — jingles, beds, voiceover, and sound effects.

    132 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Video Audio Continuity

    nodetool-ai/nodetool

    Keep sound continuous across a multi-scene piece cut from generated video — why one clip per scene hard-cuts the audio at every boundary, when to write all the scenes into a single generation…

    560 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed

More from T8mars/T8-penguin-canvas

  • Product Image Direction

    T8mars/T8-penguin-canvas

    Turn one supplied product image and a short brief into one product-led advertising image.

    615 GitHub stars~268 tokensUpdated yesterday
    Auto-check passed
  • Product Motion Shot

    T8mars/T8-penguin-canvas

    Plan and generate one continuous product video shot from one reference image, with a coherent camera move and product identity checks.

    615 GitHub stars~255 tokensUpdated yesterday
    Auto-check passed
  • Product Prompt Brief

    T8mars/T8-penguin-canvas

    Turn product facts and a loose image brief into a complete, editable image prompt with clear constraints.

    615 GitHub stars~229 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Music Caption Rewriter

What does Music Caption Rewriter do?

Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement. Music Caption Rewriter is an agent skill from T8mars/T8-penguin-canvas. Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.

When should I use Music Caption Rewriter?

Music Caption Rewriter fits situations like: users ask to enhance a music-generation prompt; preserve lyric-section directives; retrieve a similar style from bundled templates; JSONL caption output.

How do I install Music Caption Rewriter in Claude Code?

Run `npx skills add T8mars/T8-penguin-canvas --skill music-caption-rewriter -a claude-code`. Or copy the skill folder (public/official-skills/minimax-music3/music-caption-rewriter in T8mars/T8-penguin-canvas) into .claude/skills/music-caption-rewriter in your project. Claude Code loads it when a task matches its description.

How do I install Music Caption Rewriter in Codex?

Run `npx skills add T8mars/T8-penguin-canvas --skill music-caption-rewriter -a codex`. Or copy the skill folder (public/official-skills/minimax-music3/music-caption-rewriter in T8mars/T8-penguin-canvas) into .agents/skills/music-caption-rewriter in your project. Codex loads it when a task matches its description.

Can I use Music Caption Rewriter in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add T8mars/T8-penguin-canvas --skill music-caption-rewriter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/music-caption-rewriter, .gemini/skills/music-caption-rewriter, .github/skills/music-caption-rewriter and .opencode/skills/music-caption-rewriter in your project.

What does Music Caption Rewriter need to run?

SKILL.md names no scripts, command-line tools or credentials: Music Caption Rewriter is instructions for the agent only.

Does Music Caption Rewriter access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Music Caption Rewriter safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Music Caption Rewriter use?

Music Caption Rewriter is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Music Caption Rewriter use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 123k tokens, read only when the agent opens those files.

What are the alternatives to Music Caption Rewriter?

Skills that share tags, products or a category with Music Caption Rewriter: Mm Music Expert (LeoYeAI/openclaw-master-skills, 2.2k stars), Minimax Music Gen (aiskillstore/marketplace, 433 stars), Venice Audio Music (veniceai/skills, 144 stars) and Video (guaardvark/guaardvark, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Music Caption Rewriter?

T8mars (a GitHub user) maintains it in T8mars/T8-penguin-canvas, which has 615 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.

Source: T8mars/T8-penguin-canvas on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.