CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

MITAuto-check passedDevelopment

Install Cuda Debugging

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/cuda-debugging .claude/skills/cuda-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-debugging
GitHub stars
253
Token cost
~1.5k tokens
SKILL.md length
307 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

  • Works in 8 steps: Build for debugging → Compute Sanitizer — automated checks → cuda-gdb interactive session → …
  • Debugging with cuda-gdb
  • SKILL.md covers Purpose, When to Use, Workflow and Common Problems, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cuda Debugging is an agent skill from mohitmishra786/low-level-dev-skills. CUDA debugging skill for GPU program correctness. Use when debugging with cuda-gdb, running NVIDIA Compute Sanitizer memcheck/racecheck, analyzing GPU core dumps, or interpreting CUDA error codes 700/702. Activates on queries about cuda-gdb, compute-sanitizer, illegal memory access, launch timeout, or device printf.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging and GPU and accelerator computing. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Debugging with cuda-gdb
  • Running NVIDIA Compute Sanitizer memcheck/racecheck
  • Analyzing GPU core dumps
  • Interpreting CUDA error codes 700/702

Example prompts

  • “/cuda-debugging”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Build for debugging
  2. Compute Sanitizer — automated checks
  3. cuda-gdb interactive session
  4. Device printf
  5. Error code triage
  6. GPU core dumps
  7. Debugging decision tree
  8. Multi-GPU and MIG notes

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash, c and gdb).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda Debugging loads about 1.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 307 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 307 words, ~1,537 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-debugging/SKILL.md (or your agent's skills folder).
name
cuda-debugging
description
CUDA debugging skill for GPU program correctness. Use when debugging with cuda-gdb, running NVIDIA Compute Sanitizer memcheck/racecheck, analyzing GPU core dumps, or interpreting CUDA error codes 700/702. Activates on queries about cuda-gdb, compute-sanitizer, illegal memory access, launch timeout, or device printf.

CUDA Debugging

Purpose

Guide agents through debugging CUDA programs with cuda-gdb for interactive GPU thread inspection, NVIDIA Compute Sanitizer for automated memory and race detection, GPU core dump analysis, device-side printf, and triaging common CUDA runtime error codes.

When to Use

  • cudaErrorIllegalAddress (700) or segmentation fault on device
  • Intermittent correctness failures in multi-threaded GPU code
  • Debugging race conditions between warps or between host and device
  • Stepping through kernel code line-by-line with cuda-gdb
  • Validating uninitialized memory reads with initcheck
  • Kernel hang or cudaErrorLaunchTimeout (702)

Workflow

1. Build for debugging
bash
# Debug build — disables optimizations, enables device debug
nvcc -G -g -O0 -arch=sm_80 -o app_debug main.cu

# Sanitizer-friendly build (lineinfo helps reports)
nvcc -lineinfo -g -O2 -arch=sm_80 -o app_san main.cu

-G is required for cuda-gdb source-level stepping. Sanitizers work with optimized builds but -G gives clearer line numbers.

2. Compute Sanitizer — automated checks
bash
# Memory errors (OOB, misaligned, leak)
compute-sanitizer --tool memcheck ./app_san

# Shared memory and global memory races
compute-sanitizer --tool racecheck ./app_san

# Uninitialized memory reads
compute-sanitizer --tool initcheck ./app_san

# Synchronization errors (missing __syncthreads)
compute-sanitizer --tool synccheck ./app_san

# Verbose with source correlation
compute-sanitizer --tool memcheck --show-reachable=yes --log-file san.log ./app_san

Typical memcheck output:

======== Invalid __global__ write of size 4
========     at 0x1a0 in vector_add(vector_add.cu:12)
========     by thread (0,0,0) in block (0,0,0)
======== Address 0x7f... is out of bounds
3. cuda-gdb interactive session
bash
# Launch under cuda-gdb
cuda-gdb ./app_debug

# Or attach to running process
cuda-gdb -p <pid>

Essential commands:

gdb
# Break at kernel entry
(cuda-gdb) break vector_add
(cuda-gdb) run

# Focus on GPU threads
(cuda-gdb) info cuda kernels
(cuda-gdb) cuda kernel 0
(cuda-gdb) cuda thread (0,0,0)   # block (x,y,z), thread (x,y,z)

# Inspect device memory
(cuda-gdb) print data[i]
(cuda-gdb) x/10f d_ptr

# Step in kernel
(cuda-gdb) cuda step
(cuda-gdb) cuda next

# All threads in block
(cuda-gdb) info cuda threads
(cuda-gdb) cuda thread (0,0,5)
4. Device printf
c
__global__ void debug_kernel(float *data, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n) {
        if (i < 5)  // limit output
            printf("thread %d: data[%d] = %f\n", i, i, data[i]);
        data[i] *= 2.0f;
    }
}
bash
# Buffer size for printf (default may truncate)
cuda-gdb) set cuda printf_buffer_size 16777216

Flush with cudaDeviceSynchronize() before checking output. Excessive printf from all threads will overwhelm the buffer.

5. Error code triage
CodeNameCommon cause
700cudaErrorIllegalAddressOOB access, use-after-free, bad pointer
701cudaErrorLaunchOutOfResourcesToo much shared mem or registers per block
702cudaErrorLaunchTimeoutInfinite loop, TDR watchdog (Windows/default Linux)
719cudaErrorLaunchFailureAssert in kernel, stack overflow
c
// Always check after launch
kernel<<<grid, block>>>(args);
cudaError_t err = cudaGetLastError();
if (err != cudaSuccess)
    fprintf(stderr, "launch: %s\n", cudaGetErrorString(err));
cudaDeviceSynchronize();
err = cudaGetLastError();
if (err != cudaSuccess)
    fprintf(stderr, "exec: %s\n", cudaGetErrorString(err));
6. GPU core dumps
bash
# Enable coredump (driver 450+)
export CUDA_ENABLE_COREDUMP_ON_EXCEPTION=1
export CUDA_COREDUMP_FILE=/tmp/cuda_coredump_%h.%p

./app_san   # crash generates dump

# Analyze with cuda-gdb
cuda-gdb ./app_san /tmp/cuda_coredump_hostname.pid
(cuda-gdb) cuda coredump load /tmp/cuda_coredump_hostname.pid
(cuda-gdb) bt
(cuda-gdb) info cuda kernels
7. Debugging decision tree
Crash or wrong results?
├── Consistent wrong values → logic bug; use printf or cuda-gdb
├── Intermittent / depends on size → OOB or race
│   ├── compute-sanitizer --tool memcheck
│   └── compute-sanitizer --tool racecheck
├── Hang / timeout 702 → infinite loop or barrier mismatch
│   └── synccheck; audit __syncthreads paths
└── Works in debug (-G), fails in release → uninitialized mem or race
    └── initcheck + racecheck on release build
8. Multi-GPU and MIG notes
bash
# Isolate GPU
CUDA_VISIBLE_DEVICES=0 compute-sanitizer --tool memcheck ./app

# MIG instances appear as separate devices
nvidia-smi -L

Common Problems

SymptomCauseFix
cuda-gdb can't break in kernelBuilt without -GRebuild with nvcc -G -g -O0
Sanitizer reports no errors but crash persistsAsync error delayedAdd cudaDeviceSynchronize() after kernel
printf shows nothingBuffer full or no syncLimit prints; increase buffer; sync
racecheck false positive on atomicsNon-atomic RMWUse atomicAdd/atomicCAS
Attach failsProcess not in CUDA contextBreak after first cudaMalloc
TDR timeout on WindowsLong-running kernelSplit kernel; cudaDeviceSetLimit or regedit TDR
  • skills/gpu/cuda — kernel patterns, memory hierarchy, launch config
  • skills/gpu/cuda-profiling — performance after correctness is verified
  • skills/gpu/gpu-memory-model — understanding races and coalescing
  • skills/debuggers/gdb — host-side GDB commands shared with cuda-gdb
  • skills/runtimes/sanitizers — ASan/TSan concepts for host code
  • skills/kernel/kernel-debugging — kgdb/kprobes for driver-level issues

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu/cuda-debugging of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Cuda Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda Debugging this skillmohitmishra786/low-level-dev-skills253—~1.5kAutomated safety check: PassMIT
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Tilelang Developeryzlnew/infra-skills149—~2.4kAutomated safety check: PassNone
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNone

Similar skills

  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • Tilelang Developer

    yzlnew/infra-skills

    Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

    149 GitHub stars~2.4k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    253 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Cuda Debugging

What does Cuda Debugging do?

CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. Cuda Debugging is an agent skill from mohitmishra786/low-level-dev-skills. CUDA debugging skill for GPU program correctness.

When should I use Cuda Debugging?

Cuda Debugging fits situations like: debugging with cuda-gdb; running NVIDIA Compute Sanitizer memcheck/racecheck; analyzing GPU core dumps; interpreting CUDA error codes 700/702.

How do I install Cuda Debugging in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a claude-code`. Or copy the skill folder (skills/gpu/cuda-debugging in mohitmishra786/low-level-dev-skills) into .claude/skills/cuda-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Cuda Debugging in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a codex`. Or copy the skill folder (skills/gpu/cuda-debugging in mohitmishra786/low-level-dev-skills) into .agents/skills/cuda-debugging in your project. Codex loads it when a task matches its description.

Can I use Cuda Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-debugging, .gemini/skills/cuda-debugging, .github/skills/cuda-debugging and .opencode/skills/cuda-debugging in your project.

What does Cuda Debugging need to run?

SKILL.md names no scripts, command-line tools or credentials: Cuda Debugging is instructions for the agent only.

Does Cuda Debugging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cuda Debugging use?

Cuda Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cuda Debugging use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cuda Debugging?

Skills that share tags, products or a category with Cuda Debugging: Debug Distributed Hang (sgl-project/sglang, 37k stars), CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars) and Tilelang Developer (yzlnew/infra-skills, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda Debugging?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.