Agent skill

Ncu Report Skill

by mit-han-lab in mit-han-lab/ncu-report-skill

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

MITAuto-check passedAI & LLM Engineering

Install Ncu Report Skill

skills CLI
$ npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mit-han-lab/ncu-report-skill ncu-report-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ncu-report-skill
GitHub stars
251
Token cost
~2k tokens
SKILL.md length
847 words
Files
23
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

  • Works in 8 steps: Create a new run directory first under… → Decide what you're profiling. What… → Build a standalone harness unless the… → …
  • The user asks to profile a kernel
  • SKILL.md covers Golden rule, Quickstart (what to do when…, File index and Critical lessons (don't skip), plus 1 more section
  • Runs Python scripts from its folder

What it does

Ncu Report Skill is an agent skill from mit-han-lab/ncu-report-skill. Profile CUDA kernels with Nsight Compute on B200 / sm100. Use when the user asks to profile a kernel, analyze its performance, diagnose bottlenecks, read an ncu report, or write an optimization plan — including variants in Chinese ("profile 一下", "为什么慢", "ncu 报告").

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files (for example `README.md`, `blackwell-cuda-programming.md` and `helpers/README.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA. The licence is MIT.

When your agent uses it

  • The user asks to profile a kernel
  • Analyze its performance
  • Diagnose bottlenecks
  • Read an ncu report

Example prompts

  • “profile 一下”
  • “ncu 报告”
  • “/ncu-report-skill”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Create a new run directory first under profile// at the repo root — one directory per run, never reuse an existing one. Each run contains…
  2. Decide what you're profiling. What inputs? Which dispatch path? What question do you want answered? If the kernel takes variable-sized…
  3. Build a standalone harness unless the user is profiling through their existing binary. Harnesses compile in seconds, run the kernel in…
  4. Run two profiles: --set full (with PmSampling sections) for the overview, and --set source --section SourceCounters for per-line stall…
  5. Parse with ncu_report Python module — not by eye-balling the CLI. Write analysis outputs to profile//analysis/. Use the helpers in…
  6. Work through the six analysis dimensions. See reference/05-analysis-dimensions.md. Every one matters, but on any given kernel only 1–2…
  7. Match patterns to the diagnosis playbook. See reference/06-diagnosis-playbook.md. It maps NCU signal → likely cause → concrete fix, with…
  8. Write the report at profile//REPORT.md with evidence-backed recommendations, ranked by expected impact. See reference/07-report-template.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 74a1291. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ncu Report Skill loads about 2k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 847 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mit-han-lab/ncu-report-skill at commit 74a1291, republished under its MIT licence (© mit-han-lab). 847 words, ~2,050 tokens.

Download SKILL.mdSave it as .claude/skills/ncu-report-skill/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.
name
ncu-report-skill
description
Profile CUDA kernels with Nsight Compute on B200 / sm_100. Use when the user asks to profile a kernel, analyze its performance, diagnose bottlenecks, read an ncu report, or write an optimization plan — including variants in Chinese ("profile 一下", "为什么慢", "ncu 报告").

Skill: CUDA Kernel Profiling (B200 / Nsight Compute)

When to use: user asks to profile a CUDA kernel, analyze its performance, find its bottlenecks, or write an optimization plan based on Nsight Compute data. Triggers include: "profile X", "为什么这个 kernel 慢", "ncu report 说...", "下一步怎么优化", "帮我看一下这份 ncu 报告".

Target hardware (this repo): NVIDIA B200 (sm_100, CC 10.0, 148 SMs, 192 GB HBM3e). Most advice below is generic; B200-specific notes are explicitly marked.


Golden rule

Profile → Diagnose → Plan, in that order. Never guess.

Most under-performing CUDA kernels are under-performing for exactly one reason that ncu can tell you in 10 seconds. Don't invent hypotheses before you have the report. Don't start coding a fix before you've matched the observed pattern to a known diagnosis. Don't write a wall of suggestions — rank them by evidence and expected impact.


Quickstart (what to do when someone says "profile this kernel")

  1. Create a new run directory first under profile/<run_name>/ at the repo root — one directory per run, never reuse an existing one. Each run contains its own harness/, reports/, analysis/, and REPORT.md. This rule is mandatory in this repo. See reference/00-directory-layout.md.

  2. Decide what you're profiling. What inputs? Which dispatch path? What question do you want answered? If the kernel takes variable-sized inputs (variable seq lengths, variable batch sizes), you must pick specific representative shapes from the user's workload — don't profile with arbitrary inputs.

  3. Build a standalone harness unless the user is profiling through their existing binary. Harnesses compile in seconds, run the kernel in isolation, and let you use -lineinfo cleanly so ncu can map SASS back to source. Compile into profile/<run_name>/harness/. See reference/02-harness-guide.md and the template in helpers/harness_template.cu.

  4. Run two profiles: --set full (with PmSampling sections) for the overview, and --set source --section SourceCounters for per-line stall attribution. Write outputs to profile/<run_name>/reports/. See reference/03-collection.md.

  5. Parse with ncu_report Python module — not by eye-balling the CLI. Write analysis outputs to profile/<run_name>/analysis/. Use the helpers in helpers/. See reference/04-python-api.md.

  6. Work through the six analysis dimensions. See reference/05-analysis-dimensions.md. Every one matters, but on any given kernel only 1–2 will dominate.

  7. Match patterns to the diagnosis playbook. See reference/06-diagnosis-playbook.md. It maps NCU signal → likely cause → concrete fix, with example counts for "how big is this".

  8. Write the report at profile/<run_name>/REPORT.md with evidence-backed recommendations, ranked by expected impact. See reference/07-report-template.md.


File index

Reference docs (read these when you need details)
FilePurpose
reference/00-directory-layout.mdRead first. Directory / naming conventions — one run = one subdirectory, no cross-contamination
reference/01-workflow.mdEnd-to-end checklist from "user request" to "final report"
reference/02-harness-guide.mdWhen and how to build a standalone harness (mandatory for TVM-FFI, PyTorch kernels, JIT-compiled code)
reference/03-collection.mdncu command recipes: full, source-level, PM sampling, custom sections
reference/04-python-api.mdncu_report Python API patterns with copy-pasteable code
reference/05-analysis-dimensions.mdSix analysis dimensions: occupancy, balance, stalls, tensor core, timeline, memory
reference/06-diagnosis-playbook.mdPattern → diagnosis → fix. Merges Blackwell programming principles with NCU signals
reference/07-report-template.mdHow to structure the final report
reference/08-b200-metric-names.mdsm_100 metric names vs older GPUs — many common names are different
reference/09-common-issues.mdPermissions, PM sampling gaps, TVM-FFI / PyTorch gotchas
Show full SKILL.md (344 more words)Show less
Helpers (reusable code)
FilePurpose
helpers/harness_template.cuStandalone harness template — paste your kernel, fill in input allocation, done
helpers/safetensors_loader.hHeader-only safetensors reader (no external deps) for loading real workload tensors
helpers/analyze_reports.pyExtract key metrics, produce side-by-side comparisons
helpers/extract_stall_hotspots.pyPer-line stall aggregation via action.source_info(pc)
helpers/plot_timeline.pyASCII PM-sampling timeline plotter — makes tail effect visible
helpers/list_flashinfer_workloads.pyBrowse a flashinfer-trace dataset — shape histograms, filter by axis, resolve safetensors paths for specific UUIDs
helpers/ncu_utils.pyShared Python helpers: safe metric access, per-instance extraction, report loading

Critical lessons (don't skip)

  1. The stock ncu_profile_skill.md metric names don't all work on B200. Names like smsp__inst_executed_op_global_ld.sum, dram__bytes.sum, l1tex__average_t_sectors_per_request*.ratio return None on sm_100. Use the sm_100 names in reference/08-b200-metric-names.md or enumerate via action.metric_names().

  2. Always compile with -lineinfo. Without it, ncu's source view is blank and you cannot do per-line stall analysis. If you can't add -lineinfo to the build system (TVM-FFI, PyTorch inline, JIT), build a standalone harness — that's the whole point.

  3. PM sampling is the only way to see tail effects. Static metrics average over the whole kernel; only the time-series (either pmsampling: metrics or the ASCII plotter in helpers/) shows the shape of utilization over time.

  4. Load-imbalance on variable-length inputs is often the #1 bottleneck. If the user's workload has sequences of varying length, per-SM active-cycle variance will often dwarf every other effect. Always check the input distribution.

  5. NCU's rule engine (--page details) already does half the work. Each rule comes with Est. Speedup: X%. Read them first — they often point straight at the answer.

  6. Don't delegate understanding. Run the profiles yourself, open the reports, cite specific metric values. Never write "the profile shows it's memory-bound" — instead, name the two or three metric values that back your conclusion (e.g., "dram__bytes_read.sum.pct_of_peak_sustained_elapsed well under 10%, and long_scoreboard stalls dominate the pcsamp histogram, so the kernel is latency-bound on L1, not DRAM-bandwidth-bound"). Fill in the actual numbers from your report. Specificity is the deliverable.


  • blackwell-cuda-programming.md — Blackwell-specific programming principles and checklists, preserved as a companion reference. Use it when proposing new kernel designs; use this skill when diagnosing existing kernels.

© mit-han-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 22 other files in the repository root of mit-han-lab/ncu-report-skill.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • blackwell-cuda-programming.md
  • helpers/README.md
  • helpers/analyze_reports.py
  • helpers/extract_stall_hotspots.py
  • helpers/harness_template.cu
  • helpers/list_flashinfer_workloads.py
  • helpers/ncu_utils.py
  • helpers/plot_timeline.py
  • helpers/safetensors_loader.h
  • reference/00-directory-layout.md
  • reference/01-workflow.md
  • reference/02-harness-guide.md
  • reference/03-collection.md
  • reference/04-python-api.md
  • reference/05-analysis-dimensions.md
  • … and 4 more

Open the folder on GitHubat commit 74a1291

Compare with similar skills

Ncu Report Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ncu Report Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ncu Report Skill this skillmit-han-lab/ncu-report-skill251—~2kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
DGX Spark Training Gotchaswshobson/agents40k—~2kAutomated safety check: PassMIT
Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Mamba State-Space Models

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

    13k GitHub starsUsed in 2 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Lambda Labs GPU Cloud

    Orchestra-Research/AI-Research-SKILLs

    Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

    13k GitHub starsUsed in 4 repos~3k tokens
    AI & LLM EngineeringAuto-check: warnings

Works with

Questions about Ncu Report Skill

What does Ncu Report Skill do?

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. Ncu Report Skill is an agent skill from mit-han-lab/ncu-report-skill. Profile CUDA kernels with Nsight Compute on B200 / sm100.

When should I use Ncu Report Skill?

Ncu Report Skill fits situations like: the user asks to profile a kernel; analyze its performance; diagnose bottlenecks; read an ncu report.

How do I install Ncu Report Skill in Claude Code?

Run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a claude-code`. Or copy the skill folder (the mit-han-lab/ncu-report-skill repository) into .claude/skills/ncu-report-skill in your project. Claude Code loads it when a task matches its description.

How do I install Ncu Report Skill in Codex?

Run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a codex`. Or copy the skill folder (the mit-han-lab/ncu-report-skill repository) into .agents/skills/ncu-report-skill in your project. Codex loads it when a task matches its description.

Can I use Ncu Report Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mit-han-lab/ncu-report-skill --skill ncu-report-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncu-report-skill, .gemini/skills/ncu-report-skill, .github/skills/ncu-report-skill and .opencode/skills/ncu-report-skill in your project.

What does Ncu Report Skill need to run?

Going by SKILL.md and its folder, Ncu Report Skill needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Ncu Report Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ncu Report Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ncu Report Skill use?

Ncu Report Skill is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ncu Report Skill use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ncu Report Skill?

Skills that share tags, products or a category with Ncu Report Skill: MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 214 stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and DGX Spark Training Gotchas (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ncu Report Skill?

mit-han-lab (a GitHub organization) maintains it in mit-han-lab/ncu-report-skill, which has 251 GitHub stars. The repository was last updated on August 26, 2026.

Source: mit-han-lab/ncu-report-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.