MUSA GPU Training Optimizer
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .claude/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .agents/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .cursor/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .gemini/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .github/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ncu-report-skill" agent skill from https://github.com/mit-han-lab/ncu-report-skill/tree/main into .opencode/skills/ncu-report-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-report-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ncu-report-skillProfile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.
Ncu Report Skill is an agent skill from mit-han-lab/ncu-report-skill. Profile CUDA kernels with Nsight Compute on B200 / sm100. Use when the user asks to profile a kernel, analyze its performance, diagnose bottlenecks, read an ncu report, or write an optimization plan — including variants in Chinese ("profile 一下", "为什么慢", "ncu 报告").
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files (for example `README.md`, `blackwell-cuda-programming.md` and `helpers/README.md`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 74a1291. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ncu Report Skill loads about 2k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 847 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mit-han-lab/ncu-report-skill at commit 74a1291, republished under its MIT licence (© mit-han-lab). 847 words, ~2,050 tokens.
.claude/skills/ncu-report-skill/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.When to use: user asks to profile a CUDA kernel, analyze its performance, find its bottlenecks, or write an optimization plan based on Nsight Compute data. Triggers include: "profile X", "为什么这个 kernel 慢", "ncu report 说...", "下一步怎么优化", "帮我看一下这份 ncu 报告".
Target hardware (this repo): NVIDIA B200 (sm_100, CC 10.0, 148 SMs, 192 GB HBM3e). Most advice below is generic; B200-specific notes are explicitly marked.
Profile → Diagnose → Plan, in that order. Never guess.
Most under-performing CUDA kernels are under-performing for exactly one reason that ncu can tell you in 10 seconds. Don't invent hypotheses before you have the report. Don't start coding a fix before you've matched the observed pattern to a known diagnosis. Don't write a wall of suggestions — rank them by evidence and expected impact.
Create a new run directory first under profile/<run_name>/ at the repo root — one directory per run, never reuse an existing one. Each run contains its own harness/, reports/, analysis/, and REPORT.md. This rule is mandatory in this repo. See reference/00-directory-layout.md.
Decide what you're profiling. What inputs? Which dispatch path? What question do you want answered? If the kernel takes variable-sized inputs (variable seq lengths, variable batch sizes), you must pick specific representative shapes from the user's workload — don't profile with arbitrary inputs.
Build a standalone harness unless the user is profiling through their existing binary. Harnesses compile in seconds, run the kernel in isolation, and let you use -lineinfo cleanly so ncu can map SASS back to source. Compile into profile/<run_name>/harness/. See reference/02-harness-guide.md and the template in helpers/harness_template.cu.
Run two profiles: --set full (with PmSampling sections) for the overview, and --set source --section SourceCounters for per-line stall attribution. Write outputs to profile/<run_name>/reports/. See reference/03-collection.md.
Parse with ncu_report Python module — not by eye-balling the CLI. Write analysis outputs to profile/<run_name>/analysis/. Use the helpers in helpers/. See reference/04-python-api.md.
Work through the six analysis dimensions. See reference/05-analysis-dimensions.md. Every one matters, but on any given kernel only 1–2 will dominate.
Match patterns to the diagnosis playbook. See reference/06-diagnosis-playbook.md. It maps NCU signal → likely cause → concrete fix, with example counts for "how big is this".
Write the report at profile/<run_name>/REPORT.md with evidence-backed recommendations, ranked by expected impact. See reference/07-report-template.md.
| File | Purpose |
|---|---|
reference/00-directory-layout.md | Read first. Directory / naming conventions — one run = one subdirectory, no cross-contamination |
reference/01-workflow.md | End-to-end checklist from "user request" to "final report" |
reference/02-harness-guide.md | When and how to build a standalone harness (mandatory for TVM-FFI, PyTorch kernels, JIT-compiled code) |
reference/03-collection.md | ncu command recipes: full, source-level, PM sampling, custom sections |
reference/04-python-api.md | ncu_report Python API patterns with copy-pasteable code |
reference/05-analysis-dimensions.md | Six analysis dimensions: occupancy, balance, stalls, tensor core, timeline, memory |
reference/06-diagnosis-playbook.md | Pattern → diagnosis → fix. Merges Blackwell programming principles with NCU signals |
reference/07-report-template.md | How to structure the final report |
reference/08-b200-metric-names.md | sm_100 metric names vs older GPUs — many common names are different |
reference/09-common-issues.md | Permissions, PM sampling gaps, TVM-FFI / PyTorch gotchas |
| File | Purpose |
|---|---|
helpers/harness_template.cu | Standalone harness template — paste your kernel, fill in input allocation, done |
helpers/safetensors_loader.h | Header-only safetensors reader (no external deps) for loading real workload tensors |
helpers/analyze_reports.py | Extract key metrics, produce side-by-side comparisons |
helpers/extract_stall_hotspots.py | Per-line stall aggregation via action.source_info(pc) |
helpers/plot_timeline.py | ASCII PM-sampling timeline plotter — makes tail effect visible |
helpers/list_flashinfer_workloads.py | Browse a flashinfer-trace dataset — shape histograms, filter by axis, resolve safetensors paths for specific UUIDs |
helpers/ncu_utils.py | Shared Python helpers: safe metric access, per-instance extraction, report loading |
The stock ncu_profile_skill.md metric names don't all work on B200. Names like smsp__inst_executed_op_global_ld.sum, dram__bytes.sum, l1tex__average_t_sectors_per_request*.ratio return None on sm_100. Use the sm_100 names in reference/08-b200-metric-names.md or enumerate via action.metric_names().
Always compile with -lineinfo. Without it, ncu's source view is blank and you cannot do per-line stall analysis. If you can't add -lineinfo to the build system (TVM-FFI, PyTorch inline, JIT), build a standalone harness — that's the whole point.
PM sampling is the only way to see tail effects. Static metrics average over the whole kernel; only the time-series (either pmsampling: metrics or the ASCII plotter in helpers/) shows the shape of utilization over time.
Load-imbalance on variable-length inputs is often the #1 bottleneck. If the user's workload has sequences of varying length, per-SM active-cycle variance will often dwarf every other effect. Always check the input distribution.
NCU's rule engine (--page details) already does half the work. Each rule comes with Est. Speedup: X%. Read them first — they often point straight at the answer.
Don't delegate understanding. Run the profiles yourself, open the reports, cite specific metric values. Never write "the profile shows it's memory-bound" — instead, name the two or three metric values that back your conclusion (e.g., "dram__bytes_read.sum.pct_of_peak_sustained_elapsed well under 10%, and long_scoreboard stalls dominate the pcsamp histogram, so the kernel is latency-bound on L1, not DRAM-bandwidth-bound"). Fill in the actual numbers from your report. Specificity is the deliverable.
blackwell-cuda-programming.md — Blackwell-specific programming principles and checklists, preserved as a companion reference. Use it when proposing new kernel designs; use this skill when diagnosing existing kernels.© mit-han-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files in the repository root of mit-han-lab/ncu-report-skill.
Open the folder on GitHubat commit 74a1291
Ncu Report Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ncu Report Skill this skillmit-han-lab/ncu-report-skill | 251 | — | ~2k | Automated safety check: Pass | MIT | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill | 214 | — | ~4.3k | Automated safety check: Pass | MIT | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| DGX Spark Training Gotchaswshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT |
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
KernelFlow-ops/cuda-optimized-skill
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
wshobson/agents
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
Orchestra-Research/AI-Research-SKILLs
Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.
Orchestra-Research/AI-Research-SKILLs
Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.
Works with
Categories
Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. Ncu Report Skill is an agent skill from mit-han-lab/ncu-report-skill. Profile CUDA kernels with Nsight Compute on B200 / sm100.
Ncu Report Skill fits situations like: the user asks to profile a kernel; analyze its performance; diagnose bottlenecks; read an ncu report.
Run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a claude-code`. Or copy the skill folder (the mit-han-lab/ncu-report-skill repository) into .claude/skills/ncu-report-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a codex`. Or copy the skill folder (the mit-han-lab/ncu-report-skill repository) into .agents/skills/ncu-report-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncu-report-skill, .gemini/skills/ncu-report-skill, .github/skills/ncu-report-skill and .opencode/skills/ncu-report-skill in your project.
Going by SKILL.md and its folder, Ncu Report Skill needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ncu Report Skill is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ncu Report Skill: MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 214 stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and DGX Spark Training Gotchas (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mit-han-lab (a GitHub organization) maintains it in mit-han-lab/ncu-report-skill, which has 251 GitHub stars. The repository was last updated on August 26, 2026.
Source: mit-han-lab/ncu-report-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.