Esmfold2
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .claude/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .agents/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .cursor/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .gemini/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profilingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .github/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install maxiaosong1124/ncu-cuda-profiling-skill ncu-cuda-profiling --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ncu-cuda-profiling" agent skill from https://github.com/maxiaosong1124/ncu-cuda-profiling-skill/tree/main into .opencode/skills/ncu-cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ncu-cuda-profiling", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ncu-cuda-profilingAutomated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage
Ncu Cuda Profiling is an agent skill from maxiaosong1124/ncu-cuda-profiling-skill. Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files (for example `.github/workflows/ci.yml`, `AGENTS_COMPATIBILITY.md` and `FINAL_RELEASE.md`).
It sits in AI & LLM Engineering. It works with CUDA. The repository describes itself as: let coding agents use ncu skills analysis cuda program automatically! The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 063727f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell and Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ncu Cuda Profiling loads about 1.6k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 136 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from maxiaosong1124/ncu-cuda-profiling-skill at commit 063727f, republished under its MIT licence (© maxiaosong1124). 136 words, ~1,559 tokens.
.claude/skills/ncu-cuda-profiling/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.本 Skill 提供完整的自动化 NCU 性能分析流程,支持全量指标采集和持久化存储。
# 使用 --set full 采集所有指标,并持久化保存
ncu --set full \
-o <report_name> \
--target-processes all \
./your_kernel
# 示例
ncu --set full -o matmul_analysis --target-processes all ./matmul0_perf
# 自动生成:
# - matmul_analysis.ncu-rep (NCU 报告文件)
# - matmul_analysis.csv (CSV 格式指标)# 从已保存的报告提取关键指标 (无需重新运行 kernel)
ncu --import matmul_analysis.ncu-rep --print-summary per-kernel
# 导出为 CSV
ncu --import matmul_analysis.ncu-rep --page raw --csv > metrics.csv当用户提供 NCU 数据时,AI 按以下流程处理:
情况 A: 用户提供了 .ncu-rep 文件
# 直接导入已有报告
ncu --import <file.ncu-rep> --print-summary per-kernel情况 B: 用户需要新分析
# 完整采集并持久化
ncu --set full -o <report_name> --target-processes all ./kernel情况 C: 用户提供了截图/文本
AI 会自动保存分析数据到项目目录:
project_root/
├── ncu_reports/ # NCU 报告目录
│ ├── matmul_analysis.ncu-rep # 完整报告
│ ├── matmul_analysis.csv # CSV 指标
│ └── matmul_analysis.md # AI 分析报告
└── ...使用决策引擎自动分析:
def auto_diagnose(metrics):
roofline = metrics.get('roofline_ratio', 0)
dram = metrics.get('dram_throughput', 0)
l1tex = metrics.get('l1tex_throughput', 0)
sm_busy = metrics.get('sm_busy', 0)
occupancy = metrics.get('occupancy', 0)
if roofline < 30:
if dram > 70:
return "DRAM_MEMORY_BOUND"
elif l1tex > 80 and dram < 30:
return "L1_PRESSURE_BOUND"
else:
return "LATENCY_BOUND"
elif roofline > 60:
if sm_busy > 80:
return "COMPUTE_BOUND"
else:
return "OCCUPANCY_BOUND"
else:
return "MIXED_BOUND"# NCU 性能分析报告
## 📁 报告信息
- **Kernel**: {kernel_name}
- **采集时间**: {timestamp}
- **报告文件**: {report_file}
- **原始数据**: {csv_file}
## 📈 执行摘要
| 项目 | 数值 |
|------|------|
| **主要瓶颈** | {bottleneck_type} |
| **置信度** | {confidence} |
| **性能** | {performance} GFLOPS |
| **优化潜力** | {potential}x |
## 📊 关键指标
### 性能指标
| 指标 | 数值 | 健康阈值 | 状态 |
|------|------|----------|------|
| Roofline 性能比 | {roofline}% | > 60% | {status} |
| SM Busy | {sm_busy}% | > 70% | {status} |
| Occupancy | {occupancy}% | > 50% | {status} |
### 内存指标
| 指标 | 数值 | 健康阈值 | 状态 |
|------|------|----------|------|
| DRAM Throughput | {dram}% | < 50% | {status} |
| L1/TEX Throughput | {l1tex}% | < 80% | {status} |
| L2 Throughput | {l2}% | < 80% | {status} |
## 🔍 诊断详情
**瓶颈类型**: {bottleneck_type}
**判断依据**:
- {reason_1}
- {reason_2}
## 💡 优化建议
### 高优先级
{high_priority_suggestions}
## 🛠️ 下一步操作
### 建议的 NCU 命令
```bash
# 优化后重新采集
ncu --set full -o {report_name}_optimized --target-processes all ./kernel_optimized
---
## 🔧 工具使用说明
### 完整采集 (推荐)
```bash
# 采集所有指标并保存
ncu --set full -o my_analysis --target-processes all ./kernel
# 参数说明:
# --set full # 采集完整指标集
# -o my_analysis # 输出文件名 (生成 my_analysis.ncu-rep)
# --target-processes all # 监控所有进程# 从已有报告提取特定指标
ncu --import my_analysis.ncu-rep --print-summary per-kernel
# 导出为 CSV 便于分析
ncu --import my_analysis.ncu-rep --page raw --csv > metrics.csv使用提供的自动化脚本:
cd examples/
# 全自动分析
./auto_profile.sh ./kernel report_name
# Python 分析器
python ncu_analyzer.py --import report_name.ncu-repIF dram_throughput > 70% AND roofline < 30%:
诊断: DRAM_MEMORY_BOUND (置信度: HIGH)
优化策略:
1. Block Tiling (共享内存缓存)
2. Vectorized Load (float4)
3. Prefetching (数据预取)IF l1tex_throughput > 80% AND dram_throughput < 30%:
诊断: L1_PRESSURE_BOUND (置信度: HIGH)
优化策略:
1. Shared Memory Padding
2. Data Transpose
3. Fragment CachingIF sm_busy < 50% AND occupancy > 60%:
诊断: LATENCY_BOUND (置信度: HIGH)
优化策略:
1. Double Buffering
2. Instruction-level Parallelism
3. Loop UnrollingIF roofline > 60% AND sm_busy > 80%:
诊断: COMPUTE_BOUND (置信度: HIGH)
优化策略:
1. Use FMA instructions
2. Reduce precision (FP32 -> FP16/TF32)
3. Tensor CoresIF occupancy < 30% AND sm_busy > 70%:
诊断: OCCUPANCY_BOUND (置信度: HIGH)
优化策略:
1. Reduce register usage
2. Adjust block size
3. Use __launch_bounds__| 瓶颈类型 | 立即行动 | 代码示例 | 预期收益 |
|---|---|---|---|
| DRAM_MEMORY_BOUND | Block Tiling | __shared__ float As[BM][BK]; | 3-5x |
| L1_PRESSURE_BOUND | Padding | As[BM][BK+1] | 1.2-2x |
| LATENCY_BOUND | Double Buffer | As[2][BM*BK] | 1.2-1.5x |
| COMPUTE_BOUND | FMA | fmaf(a, b, c) | 1.1-1.3x |
| OCCUPANCY_BOUND | 调整 block size | __launch_bounds__(256, 2) | 1.2-2x |
# 完整采集 (推荐)
ncu --set full -o report_name --target-processes all ./kernel
# 指定 sections
ncu --section SpeedOfLight,Occupancy,LaunchStats -o report_name ./kernel
# 特定指标
ncu --metrics sm__throughput.avg.pct,dram__throughput.avg.pct -o report_name ./kernel# 查看摘要
ncu --import report.ncu-rep --print-summary per-kernel
# 查看详情
ncu --import report.ncu-rep --page details
# 导出 CSV
ncu --import report.ncu-rep --page raw --csv > metrics.csv
# 对比两个报告
ncu --diff report1.ncu-rep report2.ncu-rep高 Throughput ≠ 高效率
DRAM Throughput 低可能是好事
Occupancy 不是越高越好
examples/本 Skill 支持完整的自动化 NCU 性能分析工作流,包含全量采集和持久化存储
© maxiaosong1124, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files in the repository root of maxiaosong1124/ncu-cuda-profiling-skill.
Open the folder on GitHubat commit 063727f
Ncu Cuda Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ncu Cuda Profiling this skillmaxiaosong1124/ncu-cuda-profiling-skill | 129 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Esmfold2JimLiu/science-skills | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Benchmark TuneMesh-LLM/mesh-llm | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill | 214 | — | ~4.3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Modelshuggingface/skills | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 |
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Mesh-LLM/mesh-llm
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
KernelFlow-ops/cuda-optimized-skill
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
inclusionAI/AReno
Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.
Works with
Categories
Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage. Ncu Cuda Profiling is an agent skill from maxiaosong1124/ncu-cuda-profiling-skill.
Ncu Cuda Profiling fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a claude-code`. Or copy the skill folder (the maxiaosong1124/ncu-cuda-profiling-skill repository) into .claude/skills/ncu-cuda-profiling in your project. Claude Code loads it when a task matches its description.
Run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a codex`. Or copy the skill folder (the maxiaosong1124/ncu-cuda-profiling-skill repository) into .agents/skills/ncu-cuda-profiling in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maxiaosong1124/ncu-cuda-profiling-skill --skill ncu-cuda-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ncu-cuda-profiling, .gemini/skills/ncu-cuda-profiling, .github/skills/ncu-cuda-profiling and .opencode/skills/ncu-cuda-profiling in your project.
Going by SKILL.md and its folder, Ncu Cuda Profiling needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ncu Cuda Profiling is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ncu Cuda Profiling: Esmfold2 (JimLiu/science-skills, 228 stars), MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Benchmark Tune (Mesh-LLM/mesh-llm, 3.5k stars) and Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 214 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
maxiaosong1124 (a GitHub user) maintains it in maxiaosong1124/ncu-cuda-profiling-skill, which has 129 GitHub stars. The repository was last updated on May 25, 2026.
Source: maxiaosong1124/ncu-cuda-profiling-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.