CUTLASS FMHA Incremental Rebuild
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang debug-distributed-hang --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .claude/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .claude/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hangType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang debug-distributed-hang --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .agents/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .agents/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang debug-distributed-hang --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .cursor/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .cursor/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/debug-distributed-hang--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang debug-distributed-hang --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .gemini/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .gemini/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang debug-distributed-hangInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .github/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .github/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang debug-distributed-hang --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .opencode/skills/debug-distributed-hang && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debug-distributed-hang" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-distributed-hang into .opencode/skills/debug-distributed-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-distributed-hang", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debug-distributed-hangDebug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
Debug Distributed Hang is an agent skill from sgl-project/sglang. Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Covers identifying hang locations via py-spy/watchdog/cuda coredump, per-rank logging to find state divergence, binary-search methodology for locating the first diverge point, and fix patterns. Use when a multi-GPU SGLang run hangs, freezes, or times out during collective operations.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Debugging and GPU and accelerator computing. It works with SGLang and CUDA. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debug Distributed Hang loads about 2.4k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 972 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 972 words, ~2,369 tokens.
.claude/skills/debug-distributed-hang/SKILL.md (or your agent's skills folder).Hangs in distributed inference happen when ranks diverge in state, causing collective operations (AllGather, AllReduce, Broadcast, Barrier) to deadlock. Common causes:
pip install py-spy or system package. Requires root or CAP_SYS_PTRACE to attach to running processes.PATH.SGLang's watchdog automatically dumps py-spy traces on timeout. Look for:
Scheduler watchdog timeout (self.watchdog_timeout=300, self.soft=False)The py-spy dump shows the stack trace of each thread. The hanging thread is typically blocked in a CUDA synchronize or NCCL collective:
Thread (active): "MainThread"
cuStreamSynchronize (libcuda.so)
...
forward_extend (model_runner.py)SGLang has two watchdog modes (see python/sglang/srt/utils/watchdog.py):
soft=False, default): dumps py-spy traces then sends SIGQUIT to kill the parent process.soft=True): only logs the timeout without killing the process, giving you more time to manually attach debuggers or collect coredumps.If the watchdog doesn't trigger, manually dump:
py-spy dump --pid <scheduler_pid>export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=COLLLook for the last collective logged before the hang. Mismatched sizes show up as one rank waiting and another never entering.
When a process hangs, you can trigger a GPU coredump on demand to see which kernel is stuck. Set these env vars before launching:
export CUDA_ENABLE_USER_TRIGGERED_COREDUMP=1
export CUDA_COREDUMP_PIPE="/tmp/cuda_pipe_%h_%p"
export CUDA_COREDUMP_FILE="/tmp/cuda_coredump_%h_%p"
export CUDA_COREDUMP_SHOW_PROGRESS=1
export CUDA_COREDUMP_GENERATION_FLAGS='skip_nonrelocated_elf_images,skip_global_memory,skip_shared_memory,skip_local_memory,skip_constbank_memory'While the process is hanging, find the pipe via /proc/<pid>/fd/ and write to it to trigger the dump:
ls /proc/<pid>/fd/ -la 2>/dev/null | grep cuda_pipe
dd if=/dev/zero bs=1M count=1 > /tmp/cuda_pipe_<hostname>_<pid>Alternatively, if you don't need to keep the process alive, kill -SIGABRT <pid> also triggers a CUDA coredump (but terminates the process).
Then open with cuda-gdb --batch -ex "target cudacore <coredump_file>". On load, it immediately shows which kernel is stuck. For example:
Opening GPU coredump: <coredump_file>
[Current focus set to CUDA kernel 0, grid 622721, cluster (4,0,0), block (16,0,0), thread (64,0,0), device 0, sm 0, warp 0, lane 0]
#0 0x00007f8029b2b040 in ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<4096ul>)<<<(24,1,1),(512,1,1)>>> ()This told us the hang was in an NCCL AllGather — not a compute kernel. Combined with the py-spy stack pointing to LogitsProcessor.forward → tensor_model_parallel_all_gather, we knew it was an AllGather size mismatch between TP ranks.
From the stack traces and logs, identify:
LogitsProcessor, tensor_model_parallel_all_gather)The key technique: each rank writes its own log file so you can diff them.
import os
_debug_files = {}
def get_debug_file(rank):
key = f"rank{rank}"
if key not in _debug_files:
_debug_files[key] = open(f"/tmp/debug_rank{rank}.log", "w")
return _debug_files[key]Gate logging behind an env var to avoid overhead in production. SGLANG_DEBUG_HANG is not a built-in SGLang env var — you need to add this check yourself in the code you're instrumenting:
if os.environ.get("SGLANG_DEBUG_HANG"):
f = get_debug_file(rank)
f.write(f"EVENT_NAME key1={val1} key2={val2}\n")
f.flush()Log structured events at key state-mutation points:
f.write(f"SCHED_BATCH step={step} num_reqs={n} extend_lens={lens}\n")
f.write(f"VERIFY predict_hash={hash} accept_len={alen}\n")
f.write(f"CACHE_INSERT rid={rid} num_tokens={n}\n")Use consistent event names (uppercase prefix) for easy grep/diff.
For tensor values, compute a hash instead of dumping raw data:
import hashlib
h = hashlib.md5(tensor.cpu().numpy().tobytes()).hexdigest()[:8]
f.write(f"LOGITS logits_hash={h}\n")For token ID lists, str(list).encode() works:
h = hashlib.md5(str(tensor.tolist()).encode()).hexdigest()[:8]tensor.cpu(), tensor.tolist(), and tensor.numpy() all trigger CUDA synchronization. This can:
Prefer logging values that are already on CPU (e.g., Python ints, list lengths, request IDs). When you must hash a GPU tensor, do it at a point where the GPU is already idle (e.g., between scheduler steps, not inside a model forward pass).
# Extract specific event type
grep "^VERIFY" /tmp/debug_rank0.log > /tmp/v_r0.txt
grep "^VERIFY" /tmp/debug_rank1.log > /tmp/v_r1.txt
diff /tmp/v_r0.txt /tmp/v_r1.txt | head -20grep -c "^VERIFY" /tmp/debug_rank*.logIf counts differ, one rank executed more iterations — that's already a diverge signal.
The first diff line tells you the exact step where ranks diverge. All lines before it are identical — the root cause is at or before this step.
Once you find the diverging event, trace backwards:
For the diverging operation, list all its inputs. Add hash logging for each:
f.write(
f"OP_INPUTS input_a_hash={h_a} input_b_hash={h_b} "
f"input_c_hash={h_c} input_d_hash={h_d}\n"
)Compare the hashes. Some inputs will match, some won't. The non-matching input is where divergence entered.
For the non-matching input, trace where it was produced and repeat: hash its inputs, diff across ranks, find the divergent one. Continue until you reach the root cause.
Symptom: All "logical" inputs are identical (same logits after all-gather), but derived floating-point values (softmax, probabilities) differ across GPUs.
Example: EAGLE speculative decoding — F.softmax → top_k_renorm_prob → top_p_renorm_prob produces slightly different target_probs on each GPU. The sampling kernel then picks different tokens. These flow into output_ids → radix cache → different prefix match depths → different extend_seq_lens → AllGather size mismatch → hang.
Symptom: Operations using torch.rand produce different values on each rank.
Fix: Generate on rank 0 and broadcast, or use a shared seed.
Symptom: A condition (e.g., memory check, queue length) evaluates differently on different ranks, causing one rank to enter a collective while another skips it.
Fix: Synchronize the condition value before branching, or restructure to ensure all ranks take the same path.
Symptom: In PP setups, one stage issues a send that the next stage never recvs (or vice versa), causing both to block indefinitely. Unlike TP hangs (collective mismatches), PP hangs typically involve point-to-point operations.
Fix: Ensure all stages agree on the number of microbatches and the sequence of send/recv calls for each microbatch.
Run the failing test multiple times to confirm the fix is stable. Intermittent hangs require many runs. A test that hung ~30% of the time needs at least 10 clean passes to be confident.
| Technique | When to Use |
|---|---|
| py-spy dump | First step — see where each rank is stuck |
NCCL_DEBUG=INFO | Identify which collective and sizes |
CUDA coredump + cuda-gdb | See which GPU kernel is blocked |
| Per-rank log files | Compare rank states over time |
| Hash of tensors | Efficiently compare large tensors across ranks |
diff on extracted events | Find the exact step of divergence |
broadcast(result, src=0) | Fix floating-point or sampling non-determinism |
© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/debug-distributed-hang of sgl-project/sglang.
Open the folder on GitHubat commit 1c42ad3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Debug Distributed Hang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debug Distributed Hang this skillsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Cudatechnillogue/ptx-isa-markdown | 229 | — | ~2.5k | Automated safety check: Pass | None | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 900 | — | ~2.8k | Automated safety check: Pass | None | |
| Cuda Cpp Kernelvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Debuggingmohitmishra786/low-level-dev-skills | 253 | — | ~1.5k | Automated safety check: Pass | MIT |
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
technillogue/ptx-isa-markdown
CUDA kernel development, debugging, and performance optimization for Claude Code.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
mohitmishra786/low-level-dev-skills
CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
sgl-project/sglang
Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…
Categories
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Debug Distributed Hang is an agent skill from sgl-project/sglang. Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
Debug Distributed Hang fits situations like: A multi-GPU SGLang run hangs; times out during collective operations.
Run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a claude-code`. Or copy the skill folder (.agents/skills/debug-distributed-hang in sgl-project/sglang) into .claude/skills/debug-distributed-hang in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a codex`. Or copy the skill folder (.agents/skills/debug-distributed-hang in sgl-project/sglang) into .agents/skills/debug-distributed-hang in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-distributed-hang, .gemini/skills/debug-distributed-hang, .github/skills/debug-distributed-hang and .opencode/skills/debug-distributed-hang in your project.
Going by SKILL.md and its folder, Debug Distributed Hang needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debug Distributed Hang is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Debug Distributed Hang: CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars) and Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.