Agent skill

Debug Distributed Hang

by sgl-project in sgl-project/sglang

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

Apache-2.0Auto-check passedDevelopment

Install Debug Distributed Hang

skills CLI
$ npx skills add sgl-project/sglang --skill debug-distributed-hang -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang debug-distributed-hang --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debug-distributed-hang .claude/skills/debug-distributed-hang && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-distributed-hang
GitHub stars
37k
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
972 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

  • Works in 6 steps: Confirm and Locate the Hang → Per-Rank Logging → Diff to Find the Diverge Point → …
  • A multi-GPU SGLang run hangs
  • SKILL.md covers Overview, Prerequisites, Step 1: Confirm and Locate the… and Step 2: Per-Rank Logging, plus 5 more sections
  • Calls pip

What it does

Debug Distributed Hang is an agent skill from sgl-project/sglang. Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Covers identifying hang locations via py-spy/watchdog/cuda coredump, per-rank logging to find state divergence, binary-search methodology for locating the first diverge point, and fix patterns. Use when a multi-GPU SGLang run hangs, freezes, or times out during collective operations.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging and GPU and accelerator computing. It works with SGLang and CUDA. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • A multi-GPU SGLang run hangs
  • Times out during collective operations

Example prompts

  • “/debug-distributed-hang”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Confirm and Locate the Hang
  2. Per-Rank Logging
  3. Diff to Find the Diverge Point
  4. Binary-Search the Root Cause
  5. Common Root Causes and Fixes
  6. Verify the Fix

What it can do on your machine

Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Distributed Hang loads about 2.4k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 972 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 972 words, ~2,369 tokens.

Download SKILL.mdSave it as .claude/skills/debug-distributed-hang/SKILL.md (or your agent's skills folder).
name
debug-distributed-hang
description
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Covers identifying hang locations via py-spy/watchdog/cuda coredump, per-rank logging to find state divergence, binary-search methodology for locating the first diverge point, and fix patterns. Use when a multi-GPU SGLang run hangs, freezes, or times out during collective operations.

Debugging Distributed Hangs in SGLang

Overview

Hangs in distributed inference happen when ranks diverge in state, causing collective operations (AllGather, AllReduce, Broadcast, Barrier) to deadlock. Common causes:

  • Size mismatch: ranks pass different tensor sizes to a collective
  • Branch divergence: one rank enters a collective, another skips it
  • Cascading state drift: a small non-determinism (e.g., floating-point) propagates into different batch structures
  • Resource exhaustion: one rank OOMs or crashes, others wait forever

Prerequisites

  • py-spy: pip install py-spy or system package. Requires root or CAP_SYS_PTRACE to attach to running processes.
  • cuda-gdb: Ships with the CUDA toolkit. Ensure it's on your PATH.

Step 1: Confirm and Locate the Hang

1a. Watchdog / py-spy

SGLang's watchdog automatically dumps py-spy traces on timeout. Look for:

Scheduler watchdog timeout (self.watchdog_timeout=300, self.soft=False)

The py-spy dump shows the stack trace of each thread. The hanging thread is typically blocked in a CUDA synchronize or NCCL collective:

Thread (active): "MainThread"
    cuStreamSynchronize (libcuda.so)
    ...
    forward_extend (model_runner.py)

SGLang has two watchdog modes (see python/sglang/srt/utils/watchdog.py):

  • Hard watchdog (soft=False, default): dumps py-spy traces then sends SIGQUIT to kill the parent process.
  • Soft watchdog (soft=True): only logs the timeout without killing the process, giving you more time to manually attach debuggers or collect coredumps.

If the watchdog doesn't trigger, manually dump:

bash
py-spy dump --pid <scheduler_pid>
1b. NCCL Debug Logging
bash
export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=COLL

Look for the last collective logged before the hang. Mismatched sizes show up as one rank waiting and another never entering.

1c. CUDA Coredump

When a process hangs, you can trigger a GPU coredump on demand to see which kernel is stuck. Set these env vars before launching:

bash
export CUDA_ENABLE_USER_TRIGGERED_COREDUMP=1
export CUDA_COREDUMP_PIPE="/tmp/cuda_pipe_%h_%p"
export CUDA_COREDUMP_FILE="/tmp/cuda_coredump_%h_%p"
export CUDA_COREDUMP_SHOW_PROGRESS=1
export CUDA_COREDUMP_GENERATION_FLAGS='skip_nonrelocated_elf_images,skip_global_memory,skip_shared_memory,skip_local_memory,skip_constbank_memory'

While the process is hanging, find the pipe via /proc/<pid>/fd/ and write to it to trigger the dump:

bash
ls /proc/<pid>/fd/ -la 2>/dev/null | grep cuda_pipe
dd if=/dev/zero bs=1M count=1 > /tmp/cuda_pipe_<hostname>_<pid>

Alternatively, if you don't need to keep the process alive, kill -SIGABRT <pid> also triggers a CUDA coredump (but terminates the process).

Then open with cuda-gdb --batch -ex "target cudacore <coredump_file>". On load, it immediately shows which kernel is stuck. For example:

Opening GPU coredump: <coredump_file>
[Current focus set to CUDA kernel 0, grid 622721, cluster (4,0,0), block (16,0,0), thread (64,0,0), device 0, sm 0, warp 0, lane 0]
#0  0x00007f8029b2b040 in ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<4096ul>)<<<(24,1,1),(512,1,1)>>> ()

This told us the hang was in an NCCL AllGather — not a compute kernel. Combined with the py-spy stack pointing to LogitsProcessor.forward → tensor_model_parallel_all_gather, we knew it was an AllGather size mismatch between TP ranks.

1d. Identify the Collective

From the stack traces and logs, identify:

  • Which collective hangs (AllGather, AllReduce, Broadcast)
  • Which code path invokes it (e.g., LogitsProcessor, tensor_model_parallel_all_gather)
  • Whether it's a size mismatch or a missing participant

Step 2: Per-Rank Logging

The key technique: each rank writes its own log file so you can diff them.

Setup Pattern
python
import os

_debug_files = {}

def get_debug_file(rank):
    key = f"rank{rank}"
    if key not in _debug_files:
        _debug_files[key] = open(f"/tmp/debug_rank{rank}.log", "w")
    return _debug_files[key]

Gate logging behind an env var to avoid overhead in production. SGLANG_DEBUG_HANG is not a built-in SGLang env var — you need to add this check yourself in the code you're instrumenting:

python
if os.environ.get("SGLANG_DEBUG_HANG"):
    f = get_debug_file(rank)
    f.write(f"EVENT_NAME key1={val1} key2={val2}\n")
    f.flush()
What to Log

Log structured events at key state-mutation points:

python
f.write(f"SCHED_BATCH step={step} num_reqs={n} extend_lens={lens}\n")
f.write(f"VERIFY predict_hash={hash} accept_len={alen}\n")
f.write(f"CACHE_INSERT rid={rid} num_tokens={n}\n")

Use consistent event names (uppercase prefix) for easy grep/diff.

Hash Large Tensors

For tensor values, compute a hash instead of dumping raw data:

python
import hashlib
h = hashlib.md5(tensor.cpu().numpy().tobytes()).hexdigest()[:8]
f.write(f"LOGITS logits_hash={h}\n")

For token ID lists, str(list).encode() works:

python
h = hashlib.md5(str(tensor.tolist()).encode()).hexdigest()[:8]
Avoid Implicit Synchronization

tensor.cpu(), tensor.tolist(), and tensor.numpy() all trigger CUDA synchronization. This can:

  • Change timing and mask or move the hang
  • Deadlock if the log point is between two collectives that must run back-to-back

Prefer logging values that are already on CPU (e.g., Python ints, list lengths, request IDs). When you must hash a GPU tensor, do it at a point where the GPU is already idle (e.g., between scheduler steps, not inside a model forward pass).

Step 3: Diff to Find the Diverge Point

Basic Diff
bash
# Extract specific event type
grep "^VERIFY" /tmp/debug_rank0.log > /tmp/v_r0.txt
grep "^VERIFY" /tmp/debug_rank1.log > /tmp/v_r1.txt
diff /tmp/v_r0.txt /tmp/v_r1.txt | head -20
Count Events
bash
grep -c "^VERIFY" /tmp/debug_rank*.log

If counts differ, one rank executed more iterations — that's already a diverge signal.

Find First Diverge

The first diff line tells you the exact step where ranks diverge. All lines before it are identical — the root cause is at or before this step.

Show full SKILL.md (377 more words)Show less

Step 4: Binary-Search the Root Cause

Once you find the diverging event, trace backwards:

4a. Identify Inputs

For the diverging operation, list all its inputs. Add hash logging for each:

python
f.write(
    f"OP_INPUTS input_a_hash={h_a} input_b_hash={h_b} "
    f"input_c_hash={h_c} input_d_hash={h_d}\n"
)
4b. Diff Inputs Across Ranks

Compare the hashes. Some inputs will match, some won't. The non-matching input is where divergence entered.

4c. Recurse

For the non-matching input, trace where it was produced and repeat: hash its inputs, diff across ranks, find the divergent one. Continue until you reach the root cause.

Step 5: Common Root Causes and Fixes

Floating-Point Non-Determinism

Symptom: All "logical" inputs are identical (same logits after all-gather), but derived floating-point values (softmax, probabilities) differ across GPUs.

Example: EAGLE speculative decoding — F.softmax → top_k_renorm_prob → top_p_renorm_prob produces slightly different target_probs on each GPU. The sampling kernel then picks different tokens. These flow into output_ids → radix cache → different prefix match depths → different extend_seq_lens → AllGather size mismatch → hang.

Random Number Divergence

Symptom: Operations using torch.rand produce different values on each rank.

Fix: Generate on rank 0 and broadcast, or use a shared seed.

Conditional Code Paths

Symptom: A condition (e.g., memory check, queue length) evaluates differently on different ranks, causing one rank to enter a collective while another skips it.

Fix: Synchronize the condition value before branching, or restructure to ensure all ranks take the same path.

Pipeline Parallel (PP) Send/Recv Mismatch

Symptom: In PP setups, one stage issues a send that the next stage never recvs (or vice versa), causing both to block indefinitely. Unlike TP hangs (collective mismatches), PP hangs typically involve point-to-point operations.

Fix: Ensure all stages agree on the number of microbatches and the sequence of send/recv calls for each microbatch.

Step 6: Verify the Fix

Run the failing test multiple times to confirm the fix is stable. Intermittent hangs require many runs. A test that hung ~30% of the time needs at least 10 clean passes to be confident.

Quick Reference

TechniqueWhen to Use
py-spy dumpFirst step — see where each rank is stuck
NCCL_DEBUG=INFOIdentify which collective and sizes
CUDA coredump + cuda-gdbSee which GPU kernel is blocked
Per-rank log filesCompare rank states over time
Hash of tensorsEfficiently compare large tensors across ranks
diff on extracted eventsFind the exact step of divergence
broadcast(result, src=0)Fix floating-point or sampling non-determinism

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/debug-distributed-hang of sgl-project/sglang.

Open the folder on GitHubat commit 1c42ad3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Debug Distributed Hang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Distributed Hang compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Distributed Hang this skillsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNone
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Cuda Debuggingmohitmishra786/low-level-dev-skills253—~1.5kAutomated safety check: PassMIT

Similar skills

  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Debugging

    mohitmishra786/low-level-dev-skills

    CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Hip Rocm

    mohitmishra786/low-level-dev-skills

    HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.6k tokensUpdated 3 mo ago
    DevelopmentAuto-check: notes

More from sgl-project/sglang

All 31 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Kl Consistency Test

    sgl-project/sglang

    Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…

    37k GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check passed

Works with

Questions about Debug Distributed Hang

What does Debug Distributed Hang do?

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). Debug Distributed Hang is an agent skill from sgl-project/sglang. Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

When should I use Debug Distributed Hang?

Debug Distributed Hang fits situations like: A multi-GPU SGLang run hangs; times out during collective operations.

How do I install Debug Distributed Hang in Claude Code?

Run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a claude-code`. Or copy the skill folder (.agents/skills/debug-distributed-hang in sgl-project/sglang) into .claude/skills/debug-distributed-hang in your project. Claude Code loads it when a task matches its description.

How do I install Debug Distributed Hang in Codex?

Run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a codex`. Or copy the skill folder (.agents/skills/debug-distributed-hang in sgl-project/sglang) into .agents/skills/debug-distributed-hang in your project. Codex loads it when a task matches its description.

Can I use Debug Distributed Hang in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill debug-distributed-hang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-distributed-hang, .gemini/skills/debug-distributed-hang, .github/skills/debug-distributed-hang and .opencode/skills/debug-distributed-hang in your project.

What does Debug Distributed Hang need to run?

Going by SKILL.md and its folder, Debug Distributed Hang needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Debug Distributed Hang access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Debug Distributed Hang safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Distributed Hang use?

Debug Distributed Hang is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Distributed Hang use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Distributed Hang?

Skills that share tags, products or a category with Debug Distributed Hang: CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars) and Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Distributed Hang?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.