Debug Distributed Hang
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/cuda-debugging .claude/skills/cuda-debugging && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .claude/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debuggingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu/cuda-debugging .agents/skills/cuda-debugging && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .agents/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu/cuda-debugging .cursor/skills/cuda-debugging && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .cursor/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitmishra786/low-level-dev-skills.git --path skills/gpu/cuda-debugging--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu/cuda-debugging .gemini/skills/cuda-debugging && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .gemini/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debuggingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu/cuda-debugging .github/skills/cuda-debugging && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .github/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills cuda-debugging --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu/cuda-debugging .opencode/skills/cuda-debugging && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-debugging" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/cuda-debugging into .opencode/skills/cuda-debugging/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-debugging", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-debuggingCUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.
Cuda Debugging is an agent skill from mohitmishra786/low-level-dev-skills. CUDA debugging skill for GPU program correctness. Use when debugging with cuda-gdb, running NVIDIA Compute Sanitizer memcheck/racecheck, analyzing GPU core dumps, or interpreting CUDA error codes 700/702. Activates on queries about cuda-gdb, compute-sanitizer, illegal memory access, launch timeout, or device printf.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Debugging and GPU and accelerator computing. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash, c and gdb).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cuda Debugging loads about 1.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 307 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 307 words, ~1,537 tokens.
.claude/skills/cuda-debugging/SKILL.md (or your agent's skills folder).Guide agents through debugging CUDA programs with cuda-gdb for interactive GPU thread inspection, NVIDIA Compute Sanitizer for automated memory and race detection, GPU core dump analysis, device-side printf, and triaging common CUDA runtime error codes.
cudaErrorIllegalAddress (700) or segmentation fault on devicecudaErrorLaunchTimeout (702)# Debug build — disables optimizations, enables device debug
nvcc -G -g -O0 -arch=sm_80 -o app_debug main.cu
# Sanitizer-friendly build (lineinfo helps reports)
nvcc -lineinfo -g -O2 -arch=sm_80 -o app_san main.cu-G is required for cuda-gdb source-level stepping. Sanitizers work with optimized builds but -G gives clearer line numbers.
# Memory errors (OOB, misaligned, leak)
compute-sanitizer --tool memcheck ./app_san
# Shared memory and global memory races
compute-sanitizer --tool racecheck ./app_san
# Uninitialized memory reads
compute-sanitizer --tool initcheck ./app_san
# Synchronization errors (missing __syncthreads)
compute-sanitizer --tool synccheck ./app_san
# Verbose with source correlation
compute-sanitizer --tool memcheck --show-reachable=yes --log-file san.log ./app_sanTypical memcheck output:
======== Invalid __global__ write of size 4
======== at 0x1a0 in vector_add(vector_add.cu:12)
======== by thread (0,0,0) in block (0,0,0)
======== Address 0x7f... is out of bounds# Launch under cuda-gdb
cuda-gdb ./app_debug
# Or attach to running process
cuda-gdb -p <pid>Essential commands:
# Break at kernel entry
(cuda-gdb) break vector_add
(cuda-gdb) run
# Focus on GPU threads
(cuda-gdb) info cuda kernels
(cuda-gdb) cuda kernel 0
(cuda-gdb) cuda thread (0,0,0) # block (x,y,z), thread (x,y,z)
# Inspect device memory
(cuda-gdb) print data[i]
(cuda-gdb) x/10f d_ptr
# Step in kernel
(cuda-gdb) cuda step
(cuda-gdb) cuda next
# All threads in block
(cuda-gdb) info cuda threads
(cuda-gdb) cuda thread (0,0,5)__global__ void debug_kernel(float *data, int n) {
int i = blockIdx.x * blockDim.x + threadIdx.x;
if (i < n) {
if (i < 5) // limit output
printf("thread %d: data[%d] = %f\n", i, i, data[i]);
data[i] *= 2.0f;
}
}# Buffer size for printf (default may truncate)
cuda-gdb) set cuda printf_buffer_size 16777216Flush with cudaDeviceSynchronize() before checking output. Excessive printf from all threads will overwhelm the buffer.
| Code | Name | Common cause |
|---|---|---|
| 700 | cudaErrorIllegalAddress | OOB access, use-after-free, bad pointer |
| 701 | cudaErrorLaunchOutOfResources | Too much shared mem or registers per block |
| 702 | cudaErrorLaunchTimeout | Infinite loop, TDR watchdog (Windows/default Linux) |
| 719 | cudaErrorLaunchFailure | Assert in kernel, stack overflow |
// Always check after launch
kernel<<<grid, block>>>(args);
cudaError_t err = cudaGetLastError();
if (err != cudaSuccess)
fprintf(stderr, "launch: %s\n", cudaGetErrorString(err));
cudaDeviceSynchronize();
err = cudaGetLastError();
if (err != cudaSuccess)
fprintf(stderr, "exec: %s\n", cudaGetErrorString(err));# Enable coredump (driver 450+)
export CUDA_ENABLE_COREDUMP_ON_EXCEPTION=1
export CUDA_COREDUMP_FILE=/tmp/cuda_coredump_%h.%p
./app_san # crash generates dump
# Analyze with cuda-gdb
cuda-gdb ./app_san /tmp/cuda_coredump_hostname.pid
(cuda-gdb) cuda coredump load /tmp/cuda_coredump_hostname.pid
(cuda-gdb) bt
(cuda-gdb) info cuda kernelsCrash or wrong results?
├── Consistent wrong values → logic bug; use printf or cuda-gdb
├── Intermittent / depends on size → OOB or race
│ ├── compute-sanitizer --tool memcheck
│ └── compute-sanitizer --tool racecheck
├── Hang / timeout 702 → infinite loop or barrier mismatch
│ └── synccheck; audit __syncthreads paths
└── Works in debug (-G), fails in release → uninitialized mem or race
└── initcheck + racecheck on release build# Isolate GPU
CUDA_VISIBLE_DEVICES=0 compute-sanitizer --tool memcheck ./app
# MIG instances appear as separate devices
nvidia-smi -L| Symptom | Cause | Fix |
|---|---|---|
| cuda-gdb can't break in kernel | Built without -G | Rebuild with nvcc -G -g -O0 |
| Sanitizer reports no errors but crash persists | Async error delayed | Add cudaDeviceSynchronize() after kernel |
printf shows nothing | Buffer full or no sync | Limit prints; increase buffer; sync |
| racecheck false positive on atomics | Non-atomic RMW | Use atomicAdd/atomicCAS |
| Attach fails | Process not in CUDA context | Break after first cudaMalloc |
| TDR timeout on Windows | Long-running kernel | Split kernel; cudaDeviceSetLimit or regedit TDR |
skills/gpu/cuda — kernel patterns, memory hierarchy, launch configskills/gpu/cuda-profiling — performance after correctness is verifiedskills/gpu/gpu-memory-model — understanding races and coalescingskills/debuggers/gdb — host-side GDB commands shared with cuda-gdbskills/runtimes/sanitizers — ASan/TSan concepts for host codeskills/kernel/kernel-debugging — kgdb/kprobes for driver-level issues© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu/cuda-debugging of mohitmishra786/low-level-dev-skills.
Open the folder on GitHubat commit bdc5847
Cuda Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cuda Debugging this skillmohitmishra786/low-level-dev-skills | 253 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Debug Distributed Hangsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Cudatechnillogue/ptx-isa-markdown | 229 | — | ~2.5k | Automated safety check: Pass | None | |
| Tilelang Developeryzlnew/infra-skills | 149 | — | ~2.4k | Automated safety check: Pass | None | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 900 | — | ~2.8k | Automated safety check: Pass | None |
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
technillogue/ptx-isa-markdown
CUDA kernel development, debugging, and performance optimization for Claude Code.
yzlnew/infra-skills
Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
mohitmishra786/low-level-dev-skills
Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.
mohitmishra786/low-level-dev-skills
Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.
mohitmishra786/low-level-dev-skills
Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.
mohitmishra786/low-level-dev-skills
Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.
mohitmishra786/low-level-dev-skills
Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.
Works with
Categories
CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. Cuda Debugging is an agent skill from mohitmishra786/low-level-dev-skills. CUDA debugging skill for GPU program correctness.
Cuda Debugging fits situations like: debugging with cuda-gdb; running NVIDIA Compute Sanitizer memcheck/racecheck; analyzing GPU core dumps; interpreting CUDA error codes 700/702.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a claude-code`. Or copy the skill folder (skills/gpu/cuda-debugging in mohitmishra786/low-level-dev-skills) into .claude/skills/cuda-debugging in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a codex`. Or copy the skill folder (skills/gpu/cuda-debugging in mohitmishra786/low-level-dev-skills) into .agents/skills/cuda-debugging in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill cuda-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-debugging, .gemini/skills/cuda-debugging, .github/skills/cuda-debugging and .opencode/skills/cuda-debugging in your project.
SKILL.md names no scripts, command-line tools or credentials: Cuda Debugging is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cuda Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cuda Debugging: Debug Distributed Hang (sgl-project/sglang, 37k stars), CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars) and Tilelang Developer (yzlnew/infra-skills, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.
Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.