Roo Translation
zgsm-ai/costrict
Provides comprehensive guidelines for translating and localizing CoStrict extension strings.
Explains MiniMax-Music3's Structured Caption format (Global Metadata / Vocal Details / Arrangement) and the vendored upstream music-caption-rewriter reference library.
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/mm3-captioning .claude/skills/mm3-captioning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .claude/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/mm3-captioning .agents/skills/mm3-captioning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .agents/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/mm3-captioning .cursor/skills/mm3-captioning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .cursor/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scragnog/HOT-Step-CPP.git --path .claude/skills/mm3-captioning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/mm3-captioning .gemini/skills/mm3-captioning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .gemini/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/mm3-captioning .github/skills/mm3-captioning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .github/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scragnog/HOT-Step-CPP mm3-captioning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/mm3-captioning .opencode/skills/mm3-captioning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "mm3-captioning" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/mm3-captioning into .opencode/skills/mm3-captioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm3-captioning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
mm3-captioningExplains MiniMax-Music3's Structured Caption format (Global Metadata / Vocal Details / Arrangement) and the vendored upstream music-caption-rewriter reference library.
Mm3 Captioning is an agent skill from scragnog/HOT-Step-CPP. Explains MiniMax-Music3's Structured Caption format (Global Metadata / Vocal Details / Arrangement) and the vendored upstream music-caption-rewriter reference library. Use when formatting a caption/prompt for the MiniMax-Music3 backend (mm3- engine, /api/generate with backend=minimax), building the prompt-assembly / request-translator increment for MM3, wiring Lyric Studio output toward an MM3 caption, or debugging genre drift / low adherence in MM3 generations.
Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Writing & Content, covering Translation. It works with MiniMax. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 91e92a8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Mm3 Captioning loads about 6.2k tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 3,018 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scragnog/HOT-Step-CPP at commit 91e92a8, republished under its MIT licence (© scragnog). 3,018 words, ~6,200 tokens.
.claude/skills/mm3-captioning/SKILL.md (or your agent's skills folder).MiniMax-Music3 is the second local generation backend (alongside the native ACE-Step
LM→DiT→VAE pipeline) — see engine/src/minimax/mm3-*.h, engine/tools/convert-mm3.py,
server/src/services/backends/types.ts, ui/src/stores/backendStore.ts. Unlike
ACE-Step's caption field, MM3 was trained on a specific three-section Structured
Caption format, and adherence to that format is the main lever for controlling the
output (see "Empirical context" below).
This skill vendors MiniMax's own official caption-authoring skill
(upstream/, fetched verbatim from MiniMax-AI/MiniMax-Music3) and summarizes it for
our agents. Read upstream/SKILL.md in full before writing or reviewing any
MM3 caption-assembly code — this file is a map and a set of house notes, not a
replacement.
mm3-*
backend, or the instructions field the reference HTTP API expects).<|caption_start|>...<|caption_end|>.server/src/services/lireek/) output toward an MM3-targeted
caption instead of (or alongside) the ACE-Step caption it already produces.upstream/SKILL.md)MM3 captions have exactly three top-level sections, in this order, as plain text labels (not markdown headings — see "Skill vs. pipeline-template discrepancy" below):
Other hard rules worth internalizing:
[Verse], [Chorus],
[Bridge], [Instrumental], …) in the lyric text act as musical directives for
that section's local arrangement — they must be honored in the Arrangement section
— but the lyric words themselves are never quoted, paraphrased, or summarized into
the caption.upstream/references/ (genre router → family index → template)This is a progressive-disclosure retrieval system — do not scan all 1,000 templates. Three layers:
upstream/references/genre-router.md — entry point. Maps genre/mood/cultural
cues (with a CN/EN alias table) to one of 18 style families, e.g.
east-asian-modern, hip-hop-rap, metal-heavy-rock, cinematic-orchestral-epic,
contemporary-folk-acoustic, general-pop-ballad (fallback when only mood/imagery
is given). Read this first, pick at most one primary + one secondary family.upstream/references/index-<family>.md (18 files) — compact style cards per
family. Read only the 1–2 files the router pointed at.upstream/templates/<slug>_NNNN.txt (1,000 files) — full example captions in
the exact target format. Select up to three by distinct role — Foundation
(overall identity/groove), Modifier (one specific requested dimension:
secondary genre, vocal character, cultural color, production texture),
Arrangement (section timeline/energy-contour logic only) — then synthesize a
new caption; never copy a template's sentences, exact key/BPM, or full section
order verbatim.Example templates worth opening as calibration references (plain-text label format, not markdown):
upstream/templates/acoustic-blues-folk_0001.txt — sparse solo-instrument
arrangement, good minimal-instrumentation example.upstream/templates/index-east-asian-modern.md cards → templates therein, if
MM3 output needs to match HOT-Step's existing East-Asian-heavy adapter/dataset mix.index-hip-hop-rap.md or index-metal-heavy-rock.md card, for genres where
our current ACE-Step captions already lean on strong groove/production language —
good starting point for a side-by-side format comparison.Cross-checked upstream/SKILL.md's Output Contract against
D:\Ace-Step-Latest\mm3-weights\fixtures\tok_prompt_template.txt, the literal
prompt-assembly template the reference pipeline builds:
<|im_start|><|caption_start|>Energetic synthwave with driving bass, retro drums, and soaring lead synths. 120 BPM, A minor.<|caption_end|><|lyrics_start|>[start]
[verse]
Neon lights across the bay
...
<|lyrics_end|><|im_end|><|audio_start|>Findings:
upstream/SKILL.md presents the three required sections as markdown ###
headings in its own instructions (### Global Metadata, etc.), which could be
misread as "the output caption text should literally contain ### markdown".
The actual reference templates (e.g. templates/acoustic-blues-folk_0001.txt)
use plain text section-name lines (Global Metadata / Vocal Details /
Arrangement, no #, no blank-line separation requirement) — confirmed by
reading a template directly. The request-translator should emit plain-text
labels, not markdown, to match what MM3 was actually trained/templated on.tok_prompt_template.txt's <|caption_start|>...<|caption_end|> slot is a fully
opaque string — the tokenizer doesn't parse or require internal structure, it
just wraps whatever text is handed to it. The fixture's own example caption is a
minimal one-liner ("Energetic synthwave... 120 BPM, A minor."), not a
three-section Structured Caption — i.e. the reference pipeline was smoke-tested
with minimal captions, not the skill's prescribed structured form. This is
consistent with the skill's own README calling the plain description "used
directly" as one valid mode, with the Structured Caption as the richer,
optional upgrade path — both are legal inputs to the same
<|caption_start|> slot.<|caption_start|>.Minimal one-line captions (the fixture's own smoke-test style) produce high
take-variance across seeds — genre drift for MM3: the same short prompt lands in
noticeably different genre/mood territory seed to seed. Detailed Structured
Captions (the three-section form this skill describes) are the adherence lever —
more explicit Global Metadata / Vocal Details / Arrangement content reduces that
drift. This directly motivates building the request-translator to always emit a
full Structured Caption rather than passing a short user-typed description straight
through to <|caption_start|>.
The format is necessary and not sufficient. A caption can pass
validateMm3Caption cleanly, name a specific genre, and still render a
different one — because MM3 reads the ~560 words of description, not the two
words after "scale is minor.".
Worked case: a caption whose Basic Attributes line ended Hardcore Punk.
rendered as southern rock on every seed. Grepping its distinctive vocabulary
against the 1,000 templates says why:
| phrase used | template families that use it |
|---|---|
live-room | country-americana, blues-rock-southern-rock, blues-rock-indie-soul |
baritone | indie-folk-acoustic-pop, blues-rock-soul, traditional-pop |
close-miked | indie-folk-acoustic-pop, soul-blues-ballad, dark-folk-americana |
galloping | power-metal / symphonic-metal |
garage, minimal polish | zero templates — out of distribution |
And the corpus's own punk/hardcore templates say the opposite on every axis: "heavily distorted" not mild distortion; "wide soundstage, panned hard L/R, wall of sound" not "tight midrange focus"; "heavily compressed, modern rock radio" not "live-room honesty"; "clear youthful tenor with a nasal edge" not "chest-voice baritone"; "palm-muted chugging / power chords" not "ringing open-chord texture".
Two rules follow, both checkable:
Hardcore Punk. appears in none of the
1,000 templates. Every punk/hardcore entry is paired — Pop Punk / Alternative Rock, Metalcore / Post-Hardcore, Alternative Rock / Post-Hardcore, J-Rock / Pop Punk — and the corpus-wide mode is two
slash-joined genres. An unpaired, unattested term is a genre the model was
never taught.This supersedes nothing above; it is the layer under "THE GENRE MUST BE
SPECIFIC" in MM3_CAPTION_SYSTEM_PROMPT. Specific and attested and consistent
with the prose. The structural fix is retrieval — few-shot the caption writer
with 2-3 real templates from the routed family — designed in
docs/plans/2026-08-21-mm3-prompt-translator.md.
Confound to control first: the caption is consumed ONLY by the LM (the flow
DiT's uncond branch is zeros_like(condition), not an empty prompt —
engine/src/minimax/mm3-dit-graph.h:459). Never judge caption adherence on a
quantised LM; q8_0/NVFP4/MXFP4 degrade exactly the stage being measured.
A controlled A/B settled this for training data. One track (the album-A artist,
albumA), 30 s, identical lyrics, 5 seeds per arm, f16/f16, no adapter:
Basic Attributes synthesised from the sidecar's bpm/key/signature.Rob's verdict by ear: B better "by a massive margin"; most A takes were not even the right genre (1 of 5 was), while B was on-genre and sounded far better.
Consequences, both load-bearing:
Do not judge this class of change by spectral proxies. In the same run,
flatness (0.080 → 0.200) and centroid (1847 → 3187 Hz) both rose sharply, which
reads as "noisier/harsher" — and by ear that was simply the genre arriving (crash
cymbals and wall-of-sound distortion ARE spectrally flat and bright). An
intro_ratio (opening RMS ÷ body RMS) also favoured B, while B was the arm with
no piano at all: it measures a dynamic envelope, not instrumentation. Cross-seed
spread rose for B rather than falling, contradicting the "informative caption →
tighter clustering" hypothesis. Every proxy either misled or measured something
adjacent. Ears decided it; keep it that way.
Artifacts: M:\HOT-Step-CPP\_experiments\caption-ab-2026-08-14\ (10 WAVs, both
captions, results.json, and the runner).
ace-caption --mode mm3The two sections below describe a problem that is now solved. Read them for the traps, not for the recommendation.
MOSS-Music-8B, ported natively by the other agent, writes MM3 Structured
Captions from the audio. ace-caption --models <gguf dir> --src-audio <file> --mode mm3 --ffmpeg <path>. Rob's ear test — same track, same lyrics, same 5
seeds, both arms declaring IDENTICAL Basic Attributes so the only variable was
the prose:
| arm | verdict |
|---|---|
mm3-caption-restructure.py | "all rock… more plain rock, not particularly punk/emo" |
ace-caption --mode mm3 | "WAY better"; 2 of 5 seeds "sounds like the album-A artist already" |
Two of five seeds landed on the target artist with NO ADAPTER LOADED, from the caption alone. Caption quality is that dominant.
Two things to carry forward:
Basic Attributes with Essentia. MOSS is unreliable on the two
facts the sidecar already knows exactly — on 03-burn it said ~102 BPM /
C# minor where Essentia has 90 / E major. Tempo is a documented MOSS blind
spot and key behaves the same. engine/tools/mm3-caption-hybrid.py keeps
every section MOSS heard and rebuilds only that line from
bpm/keyscale/timesignature.Basic Attributes line.mm3-caption-restructure.py is superseded and carries a header saying so.
Follow-up to the above: engine/tools/mm3-caption-restructure.py converts the
existing Gemini caption corpus into the format mechanically. Five ear-judged
rounds on albumA, 5 seeds each, same track/lyrics/model:
| caption | on-genre |
|---|---|
| raw ACE caption | 1/5 |
| hand-written Structured Caption | 4–5/5 |
| scripted restructure v1 | 1–2/5 |
| scripted v2 (genre statement moved to lead Global Emotional Progression) | worse — country, "big band" |
| scripted, on a typical track (no piano) | heavy rock / old-school punk, still not pop-punk |
Only hand-written prose reached the target. Restructuring preserves the source's emphasis faithfully — which is the problem, because the source was written for a different model and weights things differently. It is a bridge for an existing corpus, not a solution. The real fix is to caption in MM3 format in the first place, from the audio, via the Training Studio Gemini prompt; that also turns the ~22 % boilerplate (Application Scenarios & Imagery, Vocal Style, Harmony/Backing Vocals, Vocal FX — identical across every track a script emits) into real per-track observation.
Three traps, each of which cost a wrong conclusion:
intro_ratio favoured the arm with no piano (it measures a dynamic envelope,
not instrumentation). Scripted v2 was statistically identical to the
hand-written winner (0.192 ± 0.070 vs 0.189 ± 0.075) and sounded like big
band. Cross-seed spread rose where the hypothesis said it should fall. Judge
caption changes BY EAR; use the numbers only to spot a gross regression.Scope note. All of the above judges the BASE model's genre fidelity from a caption. For training conditioning that is a proxy, not the objective: the conditioning rollout is hallucinated and content-misaligned with the target audio by construction, so what the adapter can learn is the target's timbre/production marginal, and the trigger word is what binds style to invocation. Do not spend unbounded effort perfecting base-model genre before a training run has ever been judged.
Lyric Studio's Generate Lyrics now produces two captions per song. They are separate DB fields and neither is a reformatting of the other:
| field | backend | shape |
|---|---|---|
generations.caption | ACE-Step 1.5 | 2-4 sentences of flowing description |
generations.caption_mm3 | MiniMax-Music3 | the three-heading Structured Caption above (3 headings + 13 labels) |
Where it lives — all prompt text is single-sourced in
server/src/services/lireek/prompts.ts (MM3_CAPTION_SYSTEM_PROMPT,
buildMm3CaptionPrompt, normalizeMm3Caption, validateMm3Caption,
MM3_CAPTION_FIELDS), consumed by both the in-app pipeline
(llm/orchestration.ts::writeMm3Caption) and the MCP server
(prepare_mm3_caption tool → save_generation's caption_mm3 param).
Four design points, each load-bearing:
Basic Attributes: is rebuilt deterministically from the stored
bpm/key/signature after the model answers — the same correction
mm3-caption-hybrid.py applies to MOSS. The genre clause the model wrote
is preserved, because genre is the one part of that line a model judges
better than we do.MM3_CAPTION_FIELDS.
validateMm3Caption passes 998 of those 1,000 (the two misses run
* Primary: inline). It is advisory: it drives ONE retry and never rejects,
because a partly-malformed Structured Caption still beats an ACE caption.ui/src/utils/captionForBackend.ts.
Read it live from backendStore, never from a queued param snapshot:
routes/generate.ts routes on getActiveBackendId() and ignores the
request's backend field entirely.Open — no backfill. The ~1,485 generations written before this all have
caption_mm3 = '' and fall back to their ACE caption on MM3. A bulk backfill
would look like the existing CAPTION_REPLAN machinery, but note the ceiling
established above: mechanically restructuring an ACE caption never reached the
target genre in five ear-judged rounds. A backfill should re-run the real MM3
caption call per song, not restructure.
Measured on a album B adapter (GOODCAPS: quality recipe, per-track captions), plan-only, same six seeds, candidates off:
| lyrics | caption (a training track's own, verbatim) | natural endings |
|---|---|---|
| 5528 | castaway | 1/6 |
| 5528 | warning | 4/6 |
| 5530 | warning | 4/6 |
| 5530 | castaway | 4/6 |
_experiments/2026-09-07-mm3-eos-probe/caption_stage2b.py); if they
underperform training captions, the caption prompt is the lever.mm3-captioning/
├── SKILL.md This file — house notes for HOT-Step agents
└── upstream/ Verbatim vendor copy, MiniMax-AI/MiniMax-Music3 @ main, 2026-08-13
├── PROVENANCE.md Source, fetch method, license note (read before redistributing further)
├── SKILL.md MiniMax's own skill instructions — the authoritative workflow
├── README.md MiniMax's skill-level README (usage, output contract, layout)
├── agents/openai.yaml Agent metadata (display name, default prompt)
├── references/
│ ├── genre-router.md Entry point: 18-family routing table + CN/EN aliases + fusion rules
│ └── index-*.md 18 family indexes (compact style cards), e.g. index-hip-hop-rap.md
└── templates/ 1,000 full example Structured Captions, `<slug>_NNNN.txt`华语流行/国风流行/氛围 R&B etc.) is directly reusable if Lyric Studio ever
needs to normalize CN genre input the way it already handles artist-profile
vocabulary — see project-vocal-pacing / project-section-tag-vocabulary memory
entries for our existing tag-vocabulary discipline.upstream/references/genre-router.md's "modifier vs. genre" discipline (treat
ballad, emotional, epic, modern, dark, cinematic as modifiers, never
primary-genre evidence) is a good sanity check to borrow for any future MM3-side
genre-tag validation, mirroring the OOD-tag lesson in the
project-section-tag-vocabulary memory entry (don't infer out-of-distribution
tags from our own datasets — check the authoritative doc first).© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/mm3-captioning of scragnog/HOT-Step-CPP.
Open the folder on GitHubat commit 91e92a8
Mm3 Captioning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Mm3 Captioning this skillscragnog/HOT-Step-CPP | 173 | — | ~6.2k | Automated safety check: Pass | MIT | |
| Roo Translationzgsm-ai/costrict | 4.4k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| H3 Style Craftdagthomas/comfyui_dagthomas | 290 | — | ~994 | Automated safety check: Pass | MIT | |
| Minimax H3 Prompt Reviewerunknowlei/minimax-h3-opencode-skills | 122 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Street Fighter Live Action H3karuvanan/MiniMax-H3-Director-Cut-Studio | 131 | — | ~12k | Automated safety check: Pass | Custom licence | |
| Create Release Blogweb-infra-dev/rstest | 505 | — | ~4.7k | Automated safety check: Pass | MIT |
zgsm-ai/costrict
Provides comprehensive guidelines for translating and localizing CoStrict extension strings.
dagthomas/comfyui_dagthomas
Turning the APNext H3 node's visualstyle, camera and wildness settings into observable MiniMax H3 prompt language - one visual-medium pack plus at most one motion, finish and audio pack, translating…
unknowlei/minimax-h3-opencode-skills
Downstream MiniMax H3 specialist that audits, repairs, and rewrites T2VA, I2VA, FL2VA, L2VA, and full-reference prompts into an official structured format.
karuvanan/MiniMax-H3-Director-Cut-Studio
Design 15-45 second MiniMax H3 live-action arcade martial-arts movie scenes with two readable fighters, grounded striking and MMA ground-game choreography, non-repeating multi-Segment action and a…
web-infra-dev/rstest
Generate a narrative version release blog post from commits within a tag range.
manycoretech/aholo-viewer
Guides writing and maintaining Aholo Viewer documentation: README, AGENTS.md, architecture notes, bilingual manual pages and AI collaboration guides.
scragnog/HOT-Step-CPP
The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…
scragnog/HOT-Step-CPP
Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
scragnog/HOT-Step-CPP
Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.
scragnog/HOT-Step-CPP
The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.
scragnog/HOT-Step-CPP
Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.
scragnog/HOT-Step-CPP
Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.
Works with
Categories
Explains MiniMax-Music3's Structured Caption format (Global Metadata / Vocal Details / Arrangement) and the vendored upstream music-caption-rewriter reference library. Mm3 Captioning is an agent skill from scragnog/HOT-Step-CPP. Explains MiniMax-Music3's Structured Caption format (Global Metadata / Vocal Details / Arrangement) and the vendored upstream music-caption-rewriter reference library.
Mm3 Captioning fits situations like: formatting a caption/prompt for the MiniMax-Music3 backend (mm3- engine; /api/generate with backend=minimax); building the prompt-assembly / request-translator increment for MM3; wiring Lyric Studio output toward an MM3 caption.
Run `npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a claude-code`. Or copy the skill folder (.claude/skills/mm3-captioning in scragnog/HOT-Step-CPP) into .claude/skills/mm3-captioning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a codex`. Or copy the skill folder (.claude/skills/mm3-captioning in scragnog/HOT-Step-CPP) into .agents/skills/mm3-captioning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill mm3-captioning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mm3-captioning, .gemini/skills/mm3-captioning, .github/skills/mm3-captioning and .opencode/skills/mm3-captioning in your project.
SKILL.md names no scripts, command-line tools or credentials: Mm3 Captioning is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Mm3 Captioning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Mm3 Captioning: Roo Translation (zgsm-ai/costrict, 4.4k stars), H3 Style Craft (dagthomas/comfyui_dagthomas, 290 stars), Minimax H3 Prompt Reviewer (unknowlei/minimax-h3-opencode-skills, 122 stars) and Street Fighter Live Action H3 (karuvanan/MiniMax-H3-Director-Cut-Studio, 131 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 173 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 7, 2026.
Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.