Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .claude/skills/optimize-musa-training && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .claude/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-trainingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .agents/skills/optimize-musa-training && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .agents/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .cursor/skills/optimize-musa-training && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .cursor/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-infra-skills/infra-skills.git --path skills/accelerators/optimize-musa-training--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .gemini/skills/optimize-musa-training && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .gemini/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-infra-skills/infra-skills optimize-musa-trainingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .github/skills/optimize-musa-training && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .github/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-infra-skills/infra-skills optimize-musa-training --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-infra-skills/infra-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/accelerators/optimize-musa-training .opencode/skills/optimize-musa-training && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "optimize-musa-training" agent skill from https://github.com/open-infra-skills/infra-skills/tree/main/skills/accelerators/optimize-musa-training into .opencode/skills/optimize-musa-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-musa-training", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
optimize-musa-trainingProfiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
The skill sets guardrails first: keep the architecture, data semantics, optimizer math, precision policy and checkpoint compatibility unless you authorize a change, record a versioned baseline, and compare outputs, loss, gradients, memory and steady-state throughput after every retained change. It keeps profiler overhead out of throughput numbers, separates useful model FLOPs from executed and profiler-attributed FLOPs before anything is called MFU, and treats MUSA behavior as something to detect, not infer from CUDA.
Task routing sends the agent to reference files on environment and preflight problems, measurement and profiling, an optimization playbook (FA2, GEMM, compile, launch, memory, dataloader, FSDP, MCCL), correctness experiments and a low-batch S5000 case study. Scripts compute MFU, report the MUSA environment and summarize step timings. Credentials, internal hostnames and private paths stay out of public artifacts.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 72fd3e6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
MUSA GPU Training Optimizer loads about 1.7k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 765 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from open-infra-skills/infra-skills at commit 72fd3e6, republished under its Apache-2.0 licence (© open-infra-skills). 765 words, ~1,739 tokens.
.claude/skills/optimize-musa-training/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Use a measurement-first workflow to improve MUSA training throughput without changing model semantics. Treat framework timing, system traces, and kernel counters as different layers of evidence.
Record:
Do not optimize a mixed workload as though every sample has the longest shape. Benchmark each meaningful bucket and the actual weighted mixture.
Run:
python scripts/musa_env_report.py --output <run-dir>/environment.jsonAlso save the container image digest, source commit, working-tree diff, launch command, and relevant environment switches. Verify physical device visibility from inside the process rather than trusting shell variables alone.
Summarize logs with:
python scripts/summarize_steps.py train.log --skip-first 2 --jsonFor one shape:
python scripts/compute_mfu.py \
--flops-per-device-step-tflop <F> \
--step-seconds <T> \
--peak-tflops-per-device <C> \
--flops-kind profiler-attributedFor mixed buckets, provide a JSON config with per-bucket FLOPs, time, and either step weight or sample count plus global batch. Use the aggregate total-FLOPs / total-time result, not a naive arithmetic mean.
Use the cheapest layer that answers the current question:
Do not run full end-to-end training under MCU unless the capture is tightly filtered. MCU replays and serializes kernels to collect counters; its duration is not an end-to-end throughput measurement.
Examples:
.item(), .cpu(), logging, or metric synchronization serialize the step;Change one variable at a time. Keep a decision log for both positive and negative experiments.
Require all of the following before retaining a change:
Prefer small stable gains that compose, but keep experimental paths disabled by default until their full-training benefit is repeatable.
Create a self-contained run directory with:
run/
environment.json
command.txt
source.txt
baseline.json
correctness.json
profiles/
framework/
system/
compute/
decisions.mdRecord exact versions and commands, but sanitize machine-specific and secret values before sharing.
© open-infra-skills, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (scripts, references) in skills/accelerators/optimize-musa-training of open-infra-skills/infra-skills.
Open the folder on GitHubat commit 72fd3e6
MUSA GPU Training Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| MUSA GPU Training Optimizer this skillopen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Magpie Kernel Evaluatoramd/skills | 408 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | |
| GPU OptimizerMathews-Tom/armory | 329 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Orchestra-Research/AI-Research-SKILLs
Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
Categories
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. The skill sets guardrails first: keep the architecture, data semantics, optimizer math, precision policy and checkpoint compatibility unless you authorize a change, record a versioned baseline, and compare outputs, loss, gradients, memory and steady-state throughput after every retained change. It keeps profiler overhead out of throughput numbers, separates useful model FLOPs from executed and profiler-attributed FLOPs before anything is called MFU, and treats MUSA behavior as something to detect, not infer from CUDA.
MUSA GPU Training Optimizer fits situations like: raising training throughput on Moore Threads MUSA GPUs; computing MFU or HFU and profiling a training step; debugging distributed hangs or MCCL problems; migrating CUDA training performance work to MUSA.
Run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a claude-code`. Or copy the skill folder (skills/accelerators/optimize-musa-training in open-infra-skills/infra-skills) into .claude/skills/optimize-musa-training in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a codex`. Or copy the skill folder (skills/accelerators/optimize-musa-training in open-infra-skills/infra-skills) into .agents/skills/optimize-musa-training in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-infra-skills/infra-skills --skill optimize-musa-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-musa-training, .gemini/skills/optimize-musa-training, .github/skills/optimize-musa-training and .opencode/skills/optimize-musa-training in your project.
Going by SKILL.md and its folder, MUSA GPU Training Optimizer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Moore Threads MUSA GPUs with Torch MUSA; Python for the bundled scripts.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
MUSA GPU Training Optimizer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with MUSA GPU Training Optimizer: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 408 stars), Mamba State-Space Models (Orchestra-Research/AI-Research-SKILLs, 13k stars) and GPU Optimizer (Mathews-Tom/armory, 329 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-infra-skills (a GitHub organization) maintains it in open-infra-skills/infra-skills, which has 141 GitHub stars. The repository was last updated on July 10, 2026.
Source: open-infra-skills/infra-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.