Agent skill

MUSA GPU Training Optimizer

by open-infra-skills in open-infra-skills/infra-skills

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

Apache-2.0Auto-check passedAI & LLM Engineering

Install MUSA GPU Training Optimizer

skills CLI
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .claude/skills/optimize-musa-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimize-musa-training
GitHub stars
141
Token cost
~1.7k tokens
SKILL.md length
765 words
Files
10 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

  • Works in 7 steps: Define The Metric And Invariants → Capture The Environment → Establish A Steady Baseline → …
  • Raising training throughput on Moore Threads MUSA GPUs
  • SKILL.md covers Guardrails, Route The Task, Workflow and Artifact Contract
  • Runs Python scripts from its folder; calls python

What it does

The skill sets guardrails first: keep the architecture, data semantics, optimizer math, precision policy and checkpoint compatibility unless you authorize a change, record a versioned baseline, and compare outputs, loss, gradients, memory and steady-state throughput after every retained change. It keeps profiler overhead out of throughput numbers, separates useful model FLOPs from executed and profiler-attributed FLOPs before anything is called MFU, and treats MUSA behavior as something to detect, not infer from CUDA.

Task routing sends the agent to reference files on environment and preflight problems, measurement and profiling, an optimization playbook (FA2, GEMM, compile, launch, memory, dataloader, FSDP, MCCL), correctness experiments and a low-batch S5000 case study. Scripts compute MFU, report the MUSA environment and summarize step timings. Credentials, internal hostnames and private paths stay out of public artifacts.

When your agent uses it

  • Raising training throughput on Moore Threads MUSA GPUs
  • Computing MFU or HFU and profiling a training step
  • Debugging distributed hangs or MCCL problems
  • Migrating CUDA training performance work to MUSA

Example prompts

  • “Profile our fine-tuning job on the MTT GPUs and tell me where the time goes.”
  • “Compute MFU for this training run and explain which FLOP definition you used.”
  • “Training hangs on multi-GPU with MCCL; help me find the cause.”
  • “Our low-batch utilization is poor on MUSA; look for kernel launch overhead.”

Requirements

  • Moore Threads MUSA GPUs with Torch MUSA
  • Python for the bundled scripts

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Define The Metric And Invariants
  2. Capture The Environment
  3. Establish A Steady Baseline
  4. Quantify Utilization
  5. Descend Through Three Profiling Layers
  6. Form One Evidence-Backed Hypothesis
  7. Validate And Decide

What it can do on your machine

Read from SKILL.md and the folder at commit 72fd3e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

MUSA GPU Training Optimizer loads about 1.7k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 765 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from open-infra-skills/infra-skills at commit 72fd3e6, republished under its Apache-2.0 licence (© open-infra-skills). 765 words, ~1,739 tokens.

Download SKILL.mdSave it as .claude/skills/optimize-musa-training/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
optimize-musa-training
description
Profile, benchmark, debug, and optimize AI training workloads on Moore Threads MUSA GPUs while preserving numerical behavior. Use for Torch MUSA, MTT GPUs, mthreads-gmi, Moore Perf System (msys), Moore Perf Compute (mcu), MFU/HFU, FlashAttention/FA2, SDPA, torch.compile, FSDP/FSDP2, MCCL, distributed hangs, low-batch utilization, kernel launch or transfer overhead, memory pressure, and CUDA-to-MUSA performance migration.

Optimize MUSA Training

Use a measurement-first workflow to improve MUSA training throughput without changing model semantics. Treat framework timing, system traces, and kernel counters as different layers of evidence.

Guardrails

  • Preserve model architecture, data semantics, optimizer math, precision policy, and checkpoint compatibility unless the user explicitly authorizes a change.
  • Establish a versioned baseline before editing code. Compare forward outputs, loss, gradients, memory, and steady-state throughput after every retained change.
  • Keep profiler overhead out of the throughput denominator. Measure FLOPs in a profiled run and steady step time in an otherwise equivalent non-profiled run.
  • Never label profiler-attributed executed FLOPs as model MFU without stating the FLOP definition and coverage. Distinguish useful model FLOPs, executed hardware FLOPs, and profiler-attributed FLOPs.
  • Do not infer MUSA behavior from CUDA behavior. Feature-detect the installed driver, SDK, Torch MUSA, muDNN, muBLAS, MCCL, attention backend, and profiler versions.
  • Keep cluster transport separate from profiling logic. Do not require PowerShell, VS Code, a jump host, a specific scheduler, or a particular client operating system.
  • Keep credentials, internal hostnames, private image registries, dataset paths, and proprietary reports out of public artifacts.

Route The Task

  1. For environment, import, device-selection, or container failures, read environment-and-preflight.md.
  2. For MFU/HFU, PyTorch profiling, timeline analysis, transfers, or kernel counters, read measurement-and-profiling.md.
  3. For FA2, GEMM, compile, launch, memory, dataloader, FSDP, or MCCL optimization, read optimization-playbook.md.
  4. Before retaining any optimization, read correctness-and-experiments.md.
  5. For a real low-batch S5000 case and negative results worth avoiding, read s5000-case-study.md.

Workflow

1. Define The Metric And Invariants

Record:

  • workload, model revision, dataset/sample bucket, precision, sequence shape, batch per device, accumulation, device count, and distributed strategy;
  • useful-model FLOPs, executed FLOPs, or profiler-attributed FLOPs;
  • peak denominator by SKU and precision;
  • model behaviors that must remain unchanged;
  • target metric such as samples/s, tokens/s, MFU, HFU, memory, or time-to-train.

Do not optimize a mixed workload as though every sample has the longest shape. Benchmark each meaningful bucket and the actual weighted mixture.

2. Capture The Environment

Run:

bash
python scripts/musa_env_report.py --output <run-dir>/environment.json

Also save the container image digest, source commit, working-tree diff, launch command, and relevant environment switches. Verify physical device visibility from inside the process rather than trusting shell variables alone.

3. Establish A Steady Baseline
  • Warm up imports, allocator state, compilation, autotuning, dataloader workers, and collectives.
  • Exclude compile steps, epoch boundaries, checkpoint saves, validation, and profiler steps.
  • Capture enough steady steps to expose variance. Reverse A/B order and repeat when the expected gain is below 3%.
  • Keep all ranks on the same shape bucket for a distributed step.

Summarize logs with:

bash
python scripts/summarize_steps.py train.log --skip-first 2 --json
4. Quantify Utilization

For one shape:

bash
python scripts/compute_mfu.py \
  --flops-per-device-step-tflop <F> \
  --step-seconds <T> \
  --peak-tflops-per-device <C> \
  --flops-kind profiler-attributed

For mixed buckets, provide a JSON config with per-bucket FLOPs, time, and either step weight or sample count plus global batch. Use the aggregate total-FLOPs / total-time result, not a naive arithmetic mean.

Show full SKILL.md (310 more words)Show less
5. Descend Through Three Profiling Layers

Use the cheapest layer that answers the current question:

  1. Framework layer: separate input pipeline, forward, backward/recompute, optimizer, clipping, and collectives.
  2. System layer: use Moore Perf System to inspect CPU/GPU overlap, launches, copies, synchronization, queues, streams, and rank skew.
  3. Kernel layer: use Moore Perf Compute only on a small reproducible segment to inspect LaunchStats, MemoryWorkloadAnalysis, SpeedOfLight, occupancy, registers, memory pipelines, and Roofline position.

Do not run full end-to-end training under MCU unless the capture is tightly filtered. MCU replays and serializes kernels to collect counters; its duration is not an end-to-end throughput measurement.

6. Form One Evidence-Backed Hypothesis

Examples:

  • FA2 is graph-breaking or poorly tiled for the observed head/sequence shape.
  • GEMM is using a vendor tensor-core kernel but surrounding cast/reduction traffic dominates.
  • repeated static metadata construction creates fill and launch overhead;
  • host reads such as .item(), .cpu(), logging, or metric synchronization serialize the step;
  • FSDP queue time reflects rank skew rather than slow collective kernels;
  • an activation larger than LLC is repeatedly materialized by cat/split/copy operations;
  • short and long buckets need different static compiled policies.

Change one variable at a time. Keep a decision log for both positive and negative experiments.

7. Validate And Decide

Require all of the following before retaining a change:

  • numerical differences are characterized and acceptable for the target dtype;
  • the same checkpoint and data produce finite forward, backward, and optimizer behavior;
  • the gain survives repeated non-profiled A/B runs and exceeds normal variance;
  • peak memory and all supported buckets remain acceptable;
  • the end-to-end result agrees with the microbenchmark direction;
  • distributed scaling and checkpoint load/save still work when affected.

Prefer small stable gains that compose, but keep experimental paths disabled by default until their full-training benefit is repeatable.

Artifact Contract

Create a self-contained run directory with:

text
run/
  environment.json
  command.txt
  source.txt
  baseline.json
  correctness.json
  profiles/
    framework/
    system/
    compute/
  decisions.md

Record exact versions and commands, but sanitize machine-specific and secret values before sharing.

© open-infra-skills, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in skills/accelerators/optimize-musa-training of open-infra-skills/infra-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/correctness-and-experiments.md
  • references/environment-and-preflight.md
  • references/measurement-and-profiling.md
  • references/optimization-playbook.md
  • references/s5000-case-study.md
  • scripts/compute_mfu.py
  • scripts/musa_env_report.py
  • scripts/summarize_steps.py

Open the folder on GitHubat commit 72fd3e6

Compare with similar skills

MUSA GPU Training Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

MUSA GPU Training Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
MUSA GPU Training Optimizer this skillopen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT
Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT
GPU OptimizerMathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Mamba State-Space Models

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

    13k GitHub starsUsed in 2 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    329 GitHub stars~3.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 5 days ago
    DevelopmentAuto-check: notes

Works with

Questions about MUSA GPU Training Optimizer

What does MUSA GPU Training Optimizer do?

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. The skill sets guardrails first: keep the architecture, data semantics, optimizer math, precision policy and checkpoint compatibility unless you authorize a change, record a versioned baseline, and compare outputs, loss, gradients, memory and steady-state throughput after every retained change. It keeps profiler overhead out of throughput numbers, separates useful model FLOPs from executed and profiler-attributed FLOPs before anything is called MFU, and treats MUSA behavior as something to detect, not infer from CUDA.

When should I use MUSA GPU Training Optimizer?

MUSA GPU Training Optimizer fits situations like: raising training throughput on Moore Threads MUSA GPUs; computing MFU or HFU and profiling a training step; debugging distributed hangs or MCCL problems; migrating CUDA training performance work to MUSA.

How do I install MUSA GPU Training Optimizer in Claude Code?

Run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a claude-code`. Or copy the skill folder (skills/accelerators/optimize-musa-training in open-infra-skills/infra-skills) into .claude/skills/optimize-musa-training in your project. Claude Code loads it when a task matches its description.

How do I install MUSA GPU Training Optimizer in Codex?

Run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a codex`. Or copy the skill folder (skills/accelerators/optimize-musa-training in open-infra-skills/infra-skills) into .agents/skills/optimize-musa-training in your project. Codex loads it when a task matches its description.

Can I use MUSA GPU Training Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-musa-training, .gemini/skills/optimize-musa-training, .github/skills/optimize-musa-training and .opencode/skills/optimize-musa-training in your project.

What does MUSA GPU Training Optimizer need to run?

Going by SKILL.md and its folder, MUSA GPU Training Optimizer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Moore Threads MUSA GPUs with Torch MUSA; Python for the bundled scripts.

Does MUSA GPU Training Optimizer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is MUSA GPU Training Optimizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does MUSA GPU Training Optimizer use?

MUSA GPU Training Optimizer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does MUSA GPU Training Optimizer use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.

What are the alternatives to MUSA GPU Training Optimizer?

Skills that share tags, products or a category with MUSA GPU Training Optimizer: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 408 stars), Mamba State-Space Models (Orchestra-Research/AI-Research-SKILLs, 13k stars) and GPU Optimizer (Mathews-Tom/armory, 329 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains MUSA GPU Training Optimizer?

open-infra-skills (a GitHub organization) maintains it in open-infra-skills/infra-skills, which has 141 GitHub stars. The repository was last updated on July 10, 2026.

Source: open-infra-skills/infra-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.