Debug Distributed Hang
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .claude/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .claude/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuildType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .agents/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .agents/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .cursor/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .cursor/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/onnxruntime.git --path .github/skills/cuda-cutlass-fmha-incremental-rebuild--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .gemini/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .gemini/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuildInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .github/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .github/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .opencode/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-cutlass-fmha-incremental-rebuild" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/cuda-cutlass-fmha-incremental-rebuild into .opencode/skills/cuda-cutlass-fmha-incremental-rebuild/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cutlass-fmha-incremental-rebuild", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-cutlass-fmha-incremental-rebuildExplains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
The problem is that `nvcc` depfiles do not track the headers under `onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/`, such as `kernel_forward.h` and `fmha_launch_template.h`, even though the `fmha_sm*.cu` units include them. After you edit one, an incremental `build.sh` skips those units, prints a built target and exits 0, leaving the object files and `libonnxruntime_providers_cuda.so` unchanged. Tests then exercise the old kernel, which silently invalidates any fail-to-pass or pass-to-fail check.
The fix is to `touch` the `.cu` files in that folder before the normal build so they recompile against the header change. To confirm the rebuild was real, compare the modification times of the `fmha_sm*.cu.o` files and the `.so` with your edit, not the test executable, because the CUDA provider is a shared module that the test binary loads at run time and does not relink. The skill also covers disk-space frugality on shared GPU dev boxes and points to the `ort-test` skill for general false-green cases.
Read from SKILL.md and the folder at commit 8fc9fd5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CUTLASS FMHA Incremental Rebuild loads about 1.3k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 568 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/onnxruntime at commit 8fc9fd5, republished under its MIT licence (© microsoft). 568 words, ~1,299 tokens.
.claude/skills/cuda-cutlass-fmha-incremental-rebuild/SKILL.md (or your agent's skills folder).The general false-green principles (stale binary, wrong-artifact mtime) are summarised in the
ort-testskill's "False-green taxonomy". This skill is the CUDA/CUTLASS-specific detail.
nvcc-generated depfiles do not track the CUTLASS fused-MHA headers under
onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/ (e.g. kernel_forward.h,
fmha_launch_template.h). These headers are #included by the fmha_sm*.cu
translation units, but the build system does not record that dependency.
Consequence: after you edit one of those headers, an incremental build.sh:
fmha_sm*.cu,[100%] Built target ... and exits 0,fmha_sm*.cu.o objects and the
libonnxruntime_providers_cuda.so they link into — unchanged (same mtime as
the pre-edit build).(Do not use the gtest test-exe mtime as the stale symptom: in the shared-provider
build the exe dlopens the .so and is not relinked, so its mtime stays old even
after a correct rebuild — see "How to confirm" below. The reliable diagnostic signal
is the fmha_sm*.cu.o / .so mtime.)
So your "successful" rebuild is running the old kernel. Tests that should now pass (or fail) reflect the previous code, not your edit. This silently invalidates any FAIL→PASS / PASS→FAIL verification.
Before rebuilding after editing any cutlass_fmha/*.h header:
touch onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/*.cuThen run the normal build command. This forces the fmha_sm*.cu translation units
(and downstream binaries) to recompile against your header change.
Confirm that the artifact which actually links the recompiled fmha_sm*.cu.o
is newer than your header edit.
⚠️ Do NOT just check the test EXE mtime — it can falsely flag a good build as
stale. In the shared-provider build configuration (the default here), the CUDA
execution provider is a shared module: the recompiled fmha_sm*.cu.o link into
libonnxruntime_providers_cuda.so, and the onnxruntime_provider_test executable
dlopens that .so — it is not relinked. So after a correct rebuild the
test exe mtime stays old while the .so advances. Checking the exe alone
would wrongly conclude the build was stale.
Check the right artifact for your link mode:
.so that links the recompiled .o —
build/<dir>/<cfg>/libonnxruntime_providers_cuda.soonnxruntime_provider_test)Safest check — stat both the recompiled object and the .so, and confirm BOTH
are newer than the header edit:
stat -c '%y %n' onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/kernel_forward.h
# in your build dir, e.g. build/Debug_quickbuild/Debug/:
stat -c '%y %n' libonnxruntime_providers_cuda.so
# and the actual recompiled object (path varies by build dir):
find . -name 'fmha_sm80.cu.o' -exec stat -c '%y %n' {} +If the .so (and the fmha_sm*.cu.o) timestamps are older than (or equal to) the
header edit, the build was stale — touch the .cu files and rebuild. The most
reliable signal of all is behavioral: a test that was failing now passes (a stale
binary cannot flip its result).
This is the CUDA/CUTLASS instance of false-green mode 1 (zero-match / wrong binary) —
see the ort-test skill's "False-green taxonomy" for the general principle. In short:
attention/MEA/Flash boundary gtests (e.g. FlashStructuralEmptyRows*,
Attention_Causal_NonPadKVSeqLen_MEA_*) live in onnxruntime_provider_test, which CI
runs; onnxruntime_test_all does not contain them and gives a false green. Verify the
MEA/Flash boundary fix against onnxruntime_provider_test.
Full ORT CUDA builds are large (test binaries ~1 GB each; a build dir can reach
tens of GB). On a shared box, /home filling to 100% makes builds fail in
non-obvious places — e.g. git submodule sync reporting No space left on device
or a config.lock error, not an obvious "disk full" at the compile step.
Before a big rebuild, check free space and clean only clearly-stale, regenerable build directories (old dated experiment dirs). Never delete another agent's active build dir or anything ambiguous:
df -h /home
du -sh build/* | sort -h© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .github/skills/cuda-cutlass-fmha-incremental-rebuild of microsoft/onnxruntime.
Open the folder on GitHubat commit 8fc9fd5
CUTLASS FMHA Incremental Rebuild next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| CUTLASS FMHA Incremental Rebuild this skillmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Debug Distributed Hangsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Cudatechnillogue/ptx-isa-markdown | 229 | — | ~2.5k | Automated safety check: Pass | None | |
| Cuda Cpp Kernelvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Debuggingmohitmishra786/low-level-dev-skills | 252 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Hip Rocmmohitmishra786/low-level-dev-skills | 252 | — | ~1.6k | Automated safety check: Notes | MIT |
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
technillogue/ptx-isa-markdown
CUDA kernel development, debugging, and performance optimization for Claude Code.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
mohitmishra786/low-level-dev-skills
CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
microsoft/onnxruntime
Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
microsoft/onnxruntime
Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.
microsoft/onnxruntime
Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.
microsoft/onnxruntime
Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.
microsoft/onnxruntime
Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.
Categories
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. cu` units include them.so` unchanged.
CUTLASS FMHA Incremental Rebuild fits situations like: rebuilding ONNX Runtime CUDA after editing a CUTLASS fused-MHA header; investigating a header edit that passed a build but changed no test behavior; checking that a CUDA test run used freshly compiled kernels; saving disk space on a shared GPU dev box.
Run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a claude-code`. Or copy the skill folder (.github/skills/cuda-cutlass-fmha-incremental-rebuild in microsoft/onnxruntime) into .claude/skills/cuda-cutlass-fmha-incremental-rebuild in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a codex`. Or copy the skill folder (.github/skills/cuda-cutlass-fmha-incremental-rebuild in microsoft/onnxruntime) into .agents/skills/cuda-cutlass-fmha-incremental-rebuild in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-cutlass-fmha-incremental-rebuild, .gemini/skills/cuda-cutlass-fmha-incremental-rebuild, .github/skills/cuda-cutlass-fmha-incremental-rebuild and .opencode/skills/cuda-cutlass-fmha-incremental-rebuild in your project.
Going by SKILL.md and its folder, CUTLASS FMHA Incremental Rebuild needs the command-line tools its instructions call (git). Our summary lists: An ONNX Runtime source checkout with a CUDA build; The NVIDIA CUDA toolchain including nvcc.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
CUTLASS FMHA Incremental Rebuild is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with CUTLASS FMHA Incremental Rebuild: Debug Distributed Hang (sgl-project/sglang, 37k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars), Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars) and Cuda Debugging (mohitmishra786/low-level-dev-skills, 252 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/onnxruntime, which has 22,049 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 11, 2026.
Source: microsoft/onnxruntime on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.