CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

MITAuto-check: notesAI & LLM Engineering

Install Cuda Profiling

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/cuda-profiling .claude/skills/cuda-profiling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-profiling
GitHub stars
253
Token cost
~1.6k tokens
SKILL.md length
364 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

  • Works in 8 steps: Choose profiling tool → Nsight Systems — timeline profiling → NVTX range annotations → …
  • Profiling kernels with Nsight Systems
  • SKILL.md covers Purpose, When to Use, Workflow and Common Problems, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cuda Profiling is an agent skill from mohitmishra786/low-level-dev-skills. CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight, NCU, ncu CLI, GPU roofline, occupancy metrics, or CUDA profiling workflow.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Performance optimization and GPU and accelerator computing. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Profiling kernels with Nsight Systems
  • Interpreting roofline models
  • Diagnosing memory-bound vs compute-bound kernels
  • Annotating code with NVTX ranges

Example prompts

  • “/cuda-profiling”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Choose profiling tool
  2. Nsight Systems — timeline profiling
  3. NVTX range annotations
  4. Nsight Compute — kernel analysis
  5. NCU CLI metrics
  6. Memory-bound vs compute-bound diagnosis
  7. Occupancy analysis
  8. Profiling workflow checklist

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and cpp).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda Profiling loads about 1.6k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 364 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:177
    ficient profiling permissions | Run with sudo or set `NVreg_RestrictProfilingToAdminUsers=0` |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 364 words, ~1,615 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-profiling/SKILL.md (or your agent's skills folder).
name
cuda-profiling
description
CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight, NCU, ncu CLI, GPU roofline, occupancy metrics, or CUDA profiling workflow.

CUDA Profiling

Purpose

Guide agents through profiling CUDA applications with Nsight Systems (timeline-level) and Nsight Compute (kernel-level metrics), using the NCU CLI for automated metric collection, interpreting roofline models, and diagnosing whether kernels are memory-bound or compute-bound.

When to Use

  • A CUDA kernel is slower than expected and you need bottleneck identification
  • Comparing kernel variants (tiling strategies, block sizes)
  • Building CI performance regression checks with ncu metrics
  • Correlating CPU and GPU activity in multi-stream pipelines
  • Annotating application phases with NVTX for timeline visibility
  • Interpreting occupancy, memory throughput, and SM utilization metrics

Workflow

1. Choose profiling tool
What do you need?
├── System-wide timeline (CPU+GPU+CUDA API) → Nsight Systems (nsys)
├── Per-kernel deep metrics (occupancy, memory) → Nsight Compute (ncu)
└── Quick metric from CLI in CI → ncu --metrics ...
2. Nsight Systems — timeline profiling
bash
# Profile entire application
nsys profile --trace=cuda,nvtx,osrt --output=report ./my_cuda_app

# Open report
nsys-ui report.nsys-rep

# CLI summary
nsys stats report.nsys-rep

What to look for in the timeline:

  • Gaps between kernel launches (CPU bottleneck or sync points)
  • cudaDeviceSynchronize stalls
  • Overlap between H2D copies and kernel execution across streams
  • CUDA API call overhead
bash
# Capture with CUDA graph info
nsys profile --capture-range=cudaProfilerApi ./my_cuda_app
3. NVTX range annotations
cpp
#include <nvtx3/nvToolsExt.h>

void pipeline(void) {
    nvtxRangePushA("H2D copy");
    cudaMemcpyAsync(d_in, h_in, size, cudaMemcpyHostToDevice, stream);
    nvtxRangePop();

    nvtxRangePushA("kernel");
    my_kernel<<<grid, block, 0, stream>>>(d_in, d_out, n);
    nvtxRangePop();

    nvtxRangePushA("D2H copy");
    cudaMemcpyAsync(h_out, d_out, size, cudaMemcpyDeviceToHost, stream);
    nvtxRangePop();
}

Compile with -lnvToolsExt or link nvtx3 header-only. Ranges appear as colored bands in Nsight Systems.

4. Nsight Compute — kernel analysis
bash
# Profile all kernels, save report
ncu -o kernel_report ./my_cuda_app

# Profile specific kernel by name
ncu --kernel-name regex:matmul_tiled ./my_cuda_app

# Launch UI
ncu-ui kernel_report.ncu-rep

Key sections in NCU report:

  • Speed of Light: SM throughput vs memory throughput vs peak
  • Occupancy: Active warps vs hardware limit
  • Memory Workload Analysis: L1/L2 hit rates, coalescing efficiency
  • Warp State Statistics: Stall reasons (memory, barrier, dispatch)
5. NCU CLI metrics
bash
# Essential metrics set
ncu --metrics \
  sm__throughput.avg.pct_of_peak_sustained_elapsed,\
  dram__throughput.avg.pct_of_peak_sustained_elapsed,\
  sm__warps_active.avg.pct_of_peak_sustained_active,\
  l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\
  smsp__sass_thread_inst_executed_op_ffma_pred_on.sum \
  ./my_cuda_app

# CSV export for CI
ncu --csv --metrics dram__bytes_read.sum,dram__bytes_write.sum ./my_cuda_app

# Set kernel replay mode for accurate counters
ncu --kernel-replay-mode application ./my_cuda_app
6. Memory-bound vs compute-bound diagnosis
Roofline interpretation
├── dram__throughput near peak AND sm__throughput low → memory-bound
│   └── Fix: coalescing, shared mem tiling, reduce traffic
├── sm__throughput near peak AND dram low → compute-bound
│   └── Fix: tensor cores, loop unrolling, ILP
└── Both low → launch config, occupancy, or sync overhead

Roofline model (conceptual):

Performance (GFLOP/s)
    |     /\  compute roof
    |    /  \
    |   /    \____ memory roof (bandwidth-limited region)
    |  /
    +------------------ Arithmetic Intensity (FLOP/byte)

Measure arithmetic intensity: smsp__sass_thread_inst_executed_op_ffma_pred_on.sum * 2 / dram__bytes.sum

Show full SKILL.md (153 more words)Show less
7. Occupancy analysis
bash
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active,\
launch__occupancy_limit_registers,\
launch__occupancy_limit_shared_mem,\
launch__occupancy_limit_block_size \
./my_cuda_app
Limiting factorTypical fix
Registers-maxrregcount, simplify kernel
Shared memoryReduce tile size, split phases
Block sizeTry 128 or 256 instead of 512+
8. Profiling workflow checklist
bash
# 1. Build with line info (not -G unless debugging)
nvcc -lineinfo -O3 -arch=sm_80 -o app main.cu

# 2. Timeline first
nsys profile --trace=cuda,nvtx -o timeline ./app

# 3. Deep dive on hot kernel
ncu --kernel-name regex:hot_kernel --set full ./app

# 4. Compare before/after
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v1 > v1.csv
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v2 > v2.csv

Common Problems

SymptomCauseFix
ERR_NVGPUCTRPERMInsufficient profiling permissionsRun with sudo or set NVreg_RestrictProfilingToAdminUsers=0
All metrics show zeroProfiling disabled or wrong GPUCheck CUDA_VISIBLE_DEVICES; use --target-processes all
NCU report emptyKernel too short or not launchedIncrease workload; verify cudaGetLastError()
Huge profiling overheadFull metric sets on many kernelsUse --kernel-name filter; --launch-skip
Timeline shows no overlapSingle default streamCreate multiple streams; use async copies
Occupancy looks fine but kernel slowMemory latency not hiddenCheck memory coalescing; increase active warps
  • skills/gpu/cuda — kernel writing, occupancy tuning, nvcc flags
  • skills/gpu/gpu-memory-model — coalescing, bank conflicts, SIMT model
  • skills/gpu/cuda-debugging — correctness before performance tuning
  • skills/profilers/intel-vtune-amd-uprof — CPU-side roofline and hotspot analysis
  • skills/profilers/flamegraphs — CPU flamegraphs complementary to nsys timeline
  • skills/profilers/hardware-counters — general perf stat concepts

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu/cuda-profiling of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Cuda Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda Profiling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda Profiling this skillmohitmishra786/low-level-dev-skills253—~1.6kAutomated safety check: NotesMIT
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT
Cv DeployLMIXR/CV_Deployment_skill146—~547Automated safety check: PassNone

Similar skills

  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    146 GitHub stars~547 tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    911 GitHub stars~2k tokensUpdated 3 days ago
    DevelopmentAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    253 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Cuda Profiling

What does Cuda Profiling do?

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. Cuda Profiling is an agent skill from mohitmishra786/low-level-dev-skills. CUDA profiling skill for NVIDIA GPU performance analysis.

When should I use Cuda Profiling?

Cuda Profiling fits situations like: profiling kernels with Nsight Systems; interpreting roofline models; diagnosing memory-bound vs compute-bound kernels; annotating code with NVTX ranges.

How do I install Cuda Profiling in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a claude-code`. Or copy the skill folder (skills/gpu/cuda-profiling in mohitmishra786/low-level-dev-skills) into .claude/skills/cuda-profiling in your project. Claude Code loads it when a task matches its description.

How do I install Cuda Profiling in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a codex`. Or copy the skill folder (skills/gpu/cuda-profiling in mohitmishra786/low-level-dev-skills) into .agents/skills/cuda-profiling in your project. Codex loads it when a task matches its description.

Can I use Cuda Profiling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-profiling, .gemini/skills/cuda-profiling, .github/skills/cuda-profiling and .opencode/skills/cuda-profiling in your project.

What does Cuda Profiling need to run?

SKILL.md names no scripts, command-line tools or credentials: Cuda Profiling is instructions for the agent only.

Does Cuda Profiling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda Profiling safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cuda Profiling use?

Cuda Profiling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cuda Profiling use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cuda Profiling?

Skills that share tags, products or a category with Cuda Profiling: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Graphsignal (graphsignal/graphsignal, 257 stars), Cuda (technillogue/ptx-isa-markdown, 229 stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda Profiling?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.