LLM Torch Profiler Trace Analysis
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/cuda-profiling .claude/skills/cuda-profiling && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .claude/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profilingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu/cuda-profiling .agents/skills/cuda-profiling && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .agents/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu/cuda-profiling .cursor/skills/cuda-profiling && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .cursor/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitmishra786/low-level-dev-skills.git --path skills/gpu/cuda-profiling--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu/cuda-profiling .gemini/skills/cuda-profiling && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .gemini/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profilingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu/cuda-profiling .github/skills/cuda-profiling && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .github/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-profiling --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu/cuda-profiling .opencode/skills/cuda-profiling && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-profiling" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-profiling into .opencode/skills/cuda-profiling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-profiling", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-profilingCUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.
Cuda Profiling is an agent skill from mohitmishra786/low-level-dev-skills. CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight, NCU, ncu CLI, GPU roofline, occupancy metrics, or CUDA profiling workflow.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Performance optimization and GPU and accelerator computing. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and cpp).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cuda Profiling loads about 1.6k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 364 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
ficient profiling permissions | Run with sudo or set `NVreg_RestrictProfilingToAdminUsers=0` |Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 364 words, ~1,615 tokens.
.claude/skills/cuda-profiling/SKILL.md (or your agent's skills folder).Guide agents through profiling CUDA applications with Nsight Systems (timeline-level) and Nsight Compute (kernel-level metrics), using the NCU CLI for automated metric collection, interpreting roofline models, and diagnosing whether kernels are memory-bound or compute-bound.
ncu metricsWhat do you need?
├── System-wide timeline (CPU+GPU+CUDA API) → Nsight Systems (nsys)
├── Per-kernel deep metrics (occupancy, memory) → Nsight Compute (ncu)
└── Quick metric from CLI in CI → ncu --metrics ...# Profile entire application
nsys profile --trace=cuda,nvtx,osrt --output=report ./my_cuda_app
# Open report
nsys-ui report.nsys-rep
# CLI summary
nsys stats report.nsys-repWhat to look for in the timeline:
cudaDeviceSynchronize stalls# Capture with CUDA graph info
nsys profile --capture-range=cudaProfilerApi ./my_cuda_app#include <nvtx3/nvToolsExt.h>
void pipeline(void) {
nvtxRangePushA("H2D copy");
cudaMemcpyAsync(d_in, h_in, size, cudaMemcpyHostToDevice, stream);
nvtxRangePop();
nvtxRangePushA("kernel");
my_kernel<<<grid, block, 0, stream>>>(d_in, d_out, n);
nvtxRangePop();
nvtxRangePushA("D2H copy");
cudaMemcpyAsync(h_out, d_out, size, cudaMemcpyDeviceToHost, stream);
nvtxRangePop();
}Compile with -lnvToolsExt or link nvtx3 header-only. Ranges appear as colored bands in Nsight Systems.
# Profile all kernels, save report
ncu -o kernel_report ./my_cuda_app
# Profile specific kernel by name
ncu --kernel-name regex:matmul_tiled ./my_cuda_app
# Launch UI
ncu-ui kernel_report.ncu-repKey sections in NCU report:
# Essential metrics set
ncu --metrics \
sm__throughput.avg.pct_of_peak_sustained_elapsed,\
dram__throughput.avg.pct_of_peak_sustained_elapsed,\
sm__warps_active.avg.pct_of_peak_sustained_active,\
l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\
smsp__sass_thread_inst_executed_op_ffma_pred_on.sum \
./my_cuda_app
# CSV export for CI
ncu --csv --metrics dram__bytes_read.sum,dram__bytes_write.sum ./my_cuda_app
# Set kernel replay mode for accurate counters
ncu --kernel-replay-mode application ./my_cuda_appRoofline interpretation
├── dram__throughput near peak AND sm__throughput low → memory-bound
│ └── Fix: coalescing, shared mem tiling, reduce traffic
├── sm__throughput near peak AND dram low → compute-bound
│ └── Fix: tensor cores, loop unrolling, ILP
└── Both low → launch config, occupancy, or sync overheadRoofline model (conceptual):
Performance (GFLOP/s)
| /\ compute roof
| / \
| / \____ memory roof (bandwidth-limited region)
| /
+------------------ Arithmetic Intensity (FLOP/byte)Measure arithmetic intensity: smsp__sass_thread_inst_executed_op_ffma_pred_on.sum * 2 / dram__bytes.sum
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active,\
launch__occupancy_limit_registers,\
launch__occupancy_limit_shared_mem,\
launch__occupancy_limit_block_size \
./my_cuda_app| Limiting factor | Typical fix |
|---|---|
| Registers | -maxrregcount, simplify kernel |
| Shared memory | Reduce tile size, split phases |
| Block size | Try 128 or 256 instead of 512+ |
# 1. Build with line info (not -G unless debugging)
nvcc -lineinfo -O3 -arch=sm_80 -o app main.cu
# 2. Timeline first
nsys profile --trace=cuda,nvtx -o timeline ./app
# 3. Deep dive on hot kernel
ncu --kernel-name regex:hot_kernel --set full ./app
# 4. Compare before/after
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v1 > v1.csv
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v2 > v2.csv| Symptom | Cause | Fix |
|---|---|---|
ERR_NVGPUCTRPERM | Insufficient profiling permissions | Run with sudo or set NVreg_RestrictProfilingToAdminUsers=0 |
| All metrics show zero | Profiling disabled or wrong GPU | Check CUDA_VISIBLE_DEVICES; use --target-processes all |
| NCU report empty | Kernel too short or not launched | Increase workload; verify cudaGetLastError() |
| Huge profiling overhead | Full metric sets on many kernels | Use --kernel-name filter; --launch-skip |
| Timeline shows no overlap | Single default stream | Create multiple streams; use async copies |
| Occupancy looks fine but kernel slow | Memory latency not hidden | Check memory coalescing; increase active warps |
skills/gpu/cuda — kernel writing, occupancy tuning, nvcc flagsskills/gpu/gpu-memory-model — coalescing, bank conflicts, SIMT modelskills/gpu/cuda-debugging — correctness before performance tuningskills/profilers/intel-vtune-amd-uprof — CPU-side roofline and hotspot analysisskills/profilers/flamegraphs — CPU flamegraphs complementary to nsys timelineskills/profilers/hardware-counters — general perf stat concepts© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu/cuda-profiling of mohitmishra786/low-level-dev-skills.
Open the folder on GitHubat commit bdc5847
Cuda Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cuda Profiling this skillmohitmishra786/low-level-dev-skills | 253 | — | ~1.6k | Automated safety check: Notes | MIT | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Cudatechnillogue/ptx-isa-markdown | 229 | — | ~2.5k | Automated safety check: Pass | None | |
| TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Cv DeployLMIXR/CV_Deployment_skill | 146 | — | ~547 | Automated safety check: Pass | None |
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
technillogue/ptx-isa-markdown
CUDA kernel development, debugging, and performance optimization for Claude Code.
Orchestra-Research/AI-Research-SKILLs
Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
BBuf/AI-Infra-Auto-Driven-SKILLS
Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.
mohitmishra786/low-level-dev-skills
Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.
mohitmishra786/low-level-dev-skills
Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.
mohitmishra786/low-level-dev-skills
Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.
mohitmishra786/low-level-dev-skills
Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.
mohitmishra786/low-level-dev-skills
Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.
Works with
Categories
CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. Cuda Profiling is an agent skill from mohitmishra786/low-level-dev-skills. CUDA profiling skill for NVIDIA GPU performance analysis.
Cuda Profiling fits situations like: profiling kernels with Nsight Systems; interpreting roofline models; diagnosing memory-bound vs compute-bound kernels; annotating code with NVTX ranges.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a claude-code`. Or copy the skill folder (skills/gpu/cuda-profiling in mohitmishra786/low-level-dev-skills) into .claude/skills/cuda-profiling in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a codex`. Or copy the skill folder (skills/gpu/cuda-profiling in mohitmishra786/low-level-dev-skills) into .agents/skills/cuda-profiling in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-profiling, .gemini/skills/cuda-profiling, .github/skills/cuda-profiling and .opencode/skills/cuda-profiling in your project.
SKILL.md names no scripts, command-line tools or credentials: Cuda Profiling is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Cuda Profiling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cuda Profiling: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Graphsignal (graphsignal/graphsignal, 257 stars), Cuda (technillogue/ptx-isa-markdown, 229 stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.
Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.