Agent skill

Ncu Cuda Profiling

by maxiaosong1124 in maxiaosong1124/ncu-cuda-profiling-skill

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

MITAuto-check passedAI & LLM Engineering

Install Ncu Cuda Profiling

skills CLI
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ncu-cuda-profiling
GitHub stars
129
Token cost
~1.6k tokens
SKILL.md length
136 words
Files
15
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

  • Works in 3 steps: 数据获取 (优先顺序) → 数据持久化 → 自动诊断
  • AI & LLM Engineering work in your project
  • SKILL.md covers 🚀 快速开始, 📋 AI 分析流程, 📊 输出模板 and 📖 诊断规则详解, plus 4 more sections
  • Runs Shell and Python scripts from its folder; calls python

What it does

Ncu Cuda Profiling is an agent skill from maxiaosong1124/ncu-cuda-profiling-skill. Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files (for example `.github/workflows/ci.yml`, `AGENTS_COMPATIBILITY.md` and `FINAL_RELEASE.md`).

It sits in AI & LLM Engineering. It works with CUDA. The repository describes itself as: let coding agents use ncu skills analysis cuda program automatically! The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/ncu-cuda-profiling”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 数据获取 (优先顺序)
  2. 数据持久化
  3. 自动诊断

What it can do on your machine

Read from SKILL.md and the folder at commit 063727f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ncu Cuda Profiling loads about 1.6k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 136 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maxiaosong1124/ncu-cuda-profiling-skill at commit 063727f, republished under its MIT licence (© maxiaosong1124). 136 words, ~1,559 tokens.

Download SKILL.mdSave it as .claude/skills/ncu-cuda-profiling/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
ncu-cuda-profiling
description
Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage
version
1.0.0
author
maxiaosong1124
tags
cuda, profiling, ncu, performance, optimization

NCU CUDA 自动化性能分析

本 Skill 提供完整的自动化 NCU 性能分析流程,支持全量指标采集和持久化存储。


🚀 快速开始

推荐: 一键完整采集
bash
# 使用 --set full 采集所有指标,并持久化保存
ncu --set full \
    -o <report_name> \
    --target-processes all \
    ./your_kernel

# 示例
ncu --set full -o matmul_analysis --target-processes all ./matmul0_perf

# 自动生成:
# - matmul_analysis.ncu-rep    (NCU 报告文件)
# - matmul_analysis.csv        (CSV 格式指标)
指标提取 (采集后)
bash
# 从已保存的报告提取关键指标 (无需重新运行 kernel)
ncu --import matmul_analysis.ncu-rep --print-summary per-kernel

# 导出为 CSV
ncu --import matmul_analysis.ncu-rep --page raw --csv > metrics.csv

📋 AI 分析流程

当用户提供 NCU 数据时,AI 按以下流程处理:

Phase 1: 数据获取 (优先顺序)

情况 A: 用户提供了 .ncu-rep 文件

bash
# 直接导入已有报告
ncu --import <file.ncu-rep> --print-summary per-kernel

情况 B: 用户需要新分析

bash
# 完整采集并持久化
ncu --set full -o <report_name> --target-processes all ./kernel

情况 C: 用户提供了截图/文本

  • 直接提取其中的数值进行分析
Phase 2: 数据持久化

AI 会自动保存分析数据到项目目录:

project_root/
├── ncu_reports/                    # NCU 报告目录
│   ├── matmul_analysis.ncu-rep    # 完整报告
│   ├── matmul_analysis.csv        # CSV 指标
│   └── matmul_analysis.md         # AI 分析报告
└── ...
Phase 3: 自动诊断

使用决策引擎自动分析:

python
def auto_diagnose(metrics):
    roofline = metrics.get('roofline_ratio', 0)
    dram = metrics.get('dram_throughput', 0)
    l1tex = metrics.get('l1tex_throughput', 0)
    sm_busy = metrics.get('sm_busy', 0)
    occupancy = metrics.get('occupancy', 0)
    
    if roofline < 30:
        if dram > 70:
            return "DRAM_MEMORY_BOUND"
        elif l1tex > 80 and dram < 30:
            return "L1_PRESSURE_BOUND"
        else:
            return "LATENCY_BOUND"
    elif roofline > 60:
        if sm_busy > 80:
            return "COMPUTE_BOUND"
        else:
            return "OCCUPANCY_BOUND"
    else:
        return "MIXED_BOUND"

📊 输出模板

markdown
# NCU 性能分析报告

## 📁 报告信息
- **Kernel**: {kernel_name}
- **采集时间**: {timestamp}
- **报告文件**: {report_file}
- **原始数据**: {csv_file}

## 📈 执行摘要

| 项目 | 数值 |
|------|------|
| **主要瓶颈** | {bottleneck_type} |
| **置信度** | {confidence} |
| **性能** | {performance} GFLOPS |
| **优化潜力** | {potential}x |

## 📊 关键指标

### 性能指标
| 指标 | 数值 | 健康阈值 | 状态 |
|------|------|----------|------|
| Roofline 性能比 | {roofline}% | > 60% | {status} |
| SM Busy | {sm_busy}% | > 70% | {status} |
| Occupancy | {occupancy}% | > 50% | {status} |

### 内存指标
| 指标 | 数值 | 健康阈值 | 状态 |
|------|------|----------|------|
| DRAM Throughput | {dram}% | < 50% | {status} |
| L1/TEX Throughput | {l1tex}% | < 80% | {status} |
| L2 Throughput | {l2}% | < 80% | {status} |

## 🔍 诊断详情

**瓶颈类型**: {bottleneck_type}

**判断依据**:
- {reason_1}
- {reason_2}

## 💡 优化建议

### 高优先级
{high_priority_suggestions}

## 🛠️ 下一步操作

### 建议的 NCU 命令
```bash
# 优化后重新采集
ncu --set full -o {report_name}_optimized --target-processes all ./kernel_optimized
验证清单
  • 实施建议的优化
  • 重新运行 NCU 采集
  • 对比优化前后数据

---

## 🔧 工具使用说明

### 完整采集 (推荐)

```bash
# 采集所有指标并保存
ncu --set full -o my_analysis --target-processes all ./kernel

# 参数说明:
# --set full          # 采集完整指标集
# -o my_analysis      # 输出文件名 (生成 my_analysis.ncu-rep)
# --target-processes all  # 监控所有进程
增量分析 (已有报告)
bash
# 从已有报告提取特定指标
ncu --import my_analysis.ncu-rep --print-summary per-kernel

# 导出为 CSV 便于分析
ncu --import my_analysis.ncu-rep --page raw --csv > metrics.csv
自动化脚本

使用提供的自动化脚本:

bash
cd examples/

# 全自动分析
./auto_profile.sh ./kernel report_name

# Python 分析器
python ncu_analyzer.py --import report_name.ncu-rep

📖 诊断规则详解

DRAM_MEMORY_BOUND
IF dram_throughput > 70% AND roofline < 30%:
    诊断: DRAM_MEMORY_BOUND (置信度: HIGH)
    
    优化策略:
    1. Block Tiling (共享内存缓存)
    2. Vectorized Load (float4)
    3. Prefetching (数据预取)
L1_PRESSURE_BOUND
IF l1tex_throughput > 80% AND dram_throughput < 30%:
    诊断: L1_PRESSURE_BOUND (置信度: HIGH)
    
    优化策略:
    1. Shared Memory Padding
    2. Data Transpose
    3. Fragment Caching
LATENCY_BOUND
IF sm_busy < 50% AND occupancy > 60%:
    诊断: LATENCY_BOUND (置信度: HIGH)
    
    优化策略:
    1. Double Buffering
    2. Instruction-level Parallelism
    3. Loop Unrolling
COMPUTE_BOUND
IF roofline > 60% AND sm_busy > 80%:
    诊断: COMPUTE_BOUND (置信度: HIGH)
    
    优化策略:
    1. Use FMA instructions
    2. Reduce precision (FP32 -> FP16/TF32)
    3. Tensor Cores
OCCUPANCY_BOUND
IF occupancy < 30% AND sm_busy > 70%:
    诊断: OCCUPANCY_BOUND (置信度: HIGH)
    
    优化策略:
    1. Reduce register usage
    2. Adjust block size
    3. Use __launch_bounds__

🎯 优化策略速查

瓶颈类型立即行动代码示例预期收益
DRAM_MEMORY_BOUNDBlock Tiling__shared__ float As[BM][BK];3-5x
L1_PRESSURE_BOUNDPaddingAs[BM][BK+1]1.2-2x
LATENCY_BOUNDDouble BufferAs[2][BM*BK]1.2-1.5x
COMPUTE_BOUNDFMAfmaf(a, b, c)1.1-1.3x
OCCUPANCY_BOUND调整 block size__launch_bounds__(256, 2)1.2-2x

📚 完整 NCU 命令参考

推荐采集命令
bash
# 完整采集 (推荐)
ncu --set full -o report_name --target-processes all ./kernel

# 指定 sections
ncu --section SpeedOfLight,Occupancy,LaunchStats -o report_name ./kernel

# 特定指标
ncu --metrics sm__throughput.avg.pct,dram__throughput.avg.pct -o report_name ./kernel
报告操作
bash
# 查看摘要
ncu --import report.ncu-rep --print-summary per-kernel

# 查看详情
ncu --import report.ncu-rep --page details

# 导出 CSV
ncu --import report.ncu-rep --page raw --csv > metrics.csv

# 对比两个报告
ncu --diff report1.ncu-rep report2.ncu-rep

⚠️ 常见误区

  1. 高 Throughput ≠ 高效率

    • Compute + Memory Throughput 都很高但 Roofline 很低 = GPU 在"忙碌地等待"
  2. DRAM Throughput 低可能是好事

    • 优化后 DRAM 降低说明数据在缓存中复用
  3. Occupancy 不是越高越好

    • 目标是最小足够 occupancy 隐藏延迟

🔗 相关资源


本 Skill 支持完整的自动化 NCU 性能分析工作流,包含全量采集和持久化存储

© maxiaosong1124, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files in the repository root of maxiaosong1124/ncu-cuda-profiling-skill.

  • SKILL.md
  • .github/workflows/ci.yml
  • .gitignore
  • AGENTS_COMPATIBILITY.md
  • FINAL_RELEASE.md
  • GITHUB_SETUP.md
  • LICENSE
  • README.md
  • RELEASE.md
  • check_env.sh
  • examples/README.md
  • examples/auto_profile.sh
  • examples/ncu_analyzer.py
  • install.sh
  • publish.sh

Open the folder on GitHubat commit 063727f

Compare with similar skills

Ncu Cuda Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ncu Cuda Profiling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ncu Cuda Profiling this skillmaxiaosong1124/ncu-cuda-profiling-skill129—~1.6kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Benchmark TuneMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Areno Develop Kernel

    inclusionAI/AReno

    Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

    323 GitHub stars~498 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Ncu Cuda Profiling

What does Ncu Cuda Profiling do?

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage. Ncu Cuda Profiling is an agent skill from maxiaosong1124/ncu-cuda-profiling-skill.

When should I use Ncu Cuda Profiling?

Ncu Cuda Profiling fits situations like: AI & LLM Engineering work in your project.

How do I install Ncu Cuda Profiling in Claude Code?

Run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a claude-code`. Or copy the skill folder (the maxiaosong1124/ncu-cuda-profiling-skill repository) into .claude/skills/ncu-cuda-profiling in your project. Claude Code loads it when a task matches its description.

How do I install Ncu Cuda Profiling in Codex?

Run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a codex`. Or copy the skill folder (the maxiaosong1124/ncu-cuda-profiling-skill repository) into .agents/skills/ncu-cuda-profiling in your project. Codex loads it when a task matches its description.

Can I use Ncu Cuda Profiling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncu-cuda-profiling, .gemini/skills/ncu-cuda-profiling, .github/skills/ncu-cuda-profiling and .opencode/skills/ncu-cuda-profiling in your project.

What does Ncu Cuda Profiling need to run?

Going by SKILL.md and its folder, Ncu Cuda Profiling needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.

Does Ncu Cuda Profiling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ncu Cuda Profiling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ncu Cuda Profiling use?

Ncu Cuda Profiling is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ncu Cuda Profiling use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ncu Cuda Profiling?

Skills that share tags, products or a category with Ncu Cuda Profiling: Esmfold2 (JimLiu/science-skills, 228 stars), MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Benchmark Tune (Mesh-LLM/mesh-llm, 3.5k stars) and Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 214 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ncu Cuda Profiling?

maxiaosong1124 (a GitHub user) maintains it in maxiaosong1124/ncu-cuda-profiling-skill, which has 129 GitHub stars. The repository was last updated on May 25, 2026.

Source: maxiaosong1124/ncu-cuda-profiling-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.