Official agent skill

CUTLASS FMHA Incremental Rebuild

by microsoft in microsoft/onnxruntime

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

OfficialMITAuto-check passedDevelopment

Install CUTLASS FMHA Incremental Rebuild

skills CLI
$ npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/onnxruntime cuda-cutlass-fmha-incremental-rebuild --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cuda-cutlass-fmha-incremental-rebuild .claude/skills/cuda-cutlass-fmha-incremental-rebuild && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-cutlass-fmha-incremental-rebuild
GitHub stars
22k
Token cost
~1.3k tokens
SKILL.md length
568 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

  • Rebuilding ONNX Runtime CUDA after editing a CUTLASS fused-MHA header
  • SKILL.md covers The gotcha…, The fix — force recompile the…, How to confirm the rebuild was… and Related: pick the right test…, plus 1 more section
  • Calls git
  • Investigating a header edit that passed a build but changed no test behavior

What it does

The problem is that `nvcc` depfiles do not track the headers under `onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/`, such as `kernel_forward.h` and `fmha_launch_template.h`, even though the `fmha_sm*.cu` units include them. After you edit one, an incremental `build.sh` skips those units, prints a built target and exits 0, leaving the object files and `libonnxruntime_providers_cuda.so` unchanged. Tests then exercise the old kernel, which silently invalidates any fail-to-pass or pass-to-fail check.

The fix is to `touch` the `.cu` files in that folder before the normal build so they recompile against the header change. To confirm the rebuild was real, compare the modification times of the `fmha_sm*.cu.o` files and the `.so` with your edit, not the test executable, because the CUDA provider is a shared module that the test binary loads at run time and does not relink. The skill also covers disk-space frugality on shared GPU dev boxes and points to the `ort-test` skill for general false-green cases.

When your agent uses it

  • Rebuilding ONNX Runtime CUDA after editing a CUTLASS fused-MHA header
  • Investigating a header edit that passed a build but changed no test behavior
  • Checking that a CUDA test run used freshly compiled kernels
  • Saving disk space on a shared GPU dev box

Example prompts

  • “I edited kernel_forward.h and the build passed, but the tests behave the same. Check whether the kernels are stale.”
  • “Force a real rebuild of the fused-MHA kernels after my header change.”
  • “How can I confirm libonnxruntime_providers_cuda.so was actually rebuilt?”

Requirements

  • An ONNX Runtime source checkout with a CUDA build
  • The NVIDIA CUDA toolchain including nvcc

What it can do on your machine

Read from SKILL.md and the folder at commit 8fc9fd5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CUTLASS FMHA Incremental Rebuild loads about 1.3k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 568 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/onnxruntime at commit 8fc9fd5, republished under its MIT licence (© microsoft). 568 words, ~1,299 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-cutlass-fmha-incremental-rebuild/SKILL.md (or your agent's skills folder).
name
cuda-cutlass-fmha-incremental-rebuild
description
Use when rebuilding ONNX Runtime CUDA after editing CUTLASS fused-MHA headers (onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/*.h such as kernel_forward.h or fmha_launch_template.h), or when a header edit "passed" an incremental build but test behavior did not change. Explains the nvcc depfile gotcha that produces stale Memory-Efficient-Attention (MEA) kernels and binaries, and how to force a correct recompile. Also covers disk-space frugality on shared GPU dev boxes.

Incremental rebuilds silently use STALE CUTLASS fused-MHA kernels

The general false-green principles (stale binary, wrong-artifact mtime) are summarised in the ort-test skill's "False-green taxonomy". This skill is the CUDA/CUTLASS-specific detail.

The gotcha (verification-integrity bug)

nvcc-generated depfiles do not track the CUTLASS fused-MHA headers under onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/ (e.g. kernel_forward.h, fmha_launch_template.h). These headers are #included by the fmha_sm*.cu translation units, but the build system does not record that dependency.

Consequence: after you edit one of those headers, an incremental build.sh:

  • does not recompile fmha_sm*.cu,
  • reports [100%] Built target ... and exits 0,
  • leaves the recompiled artifacts — the fmha_sm*.cu.o objects and the libonnxruntime_providers_cuda.so they link into — unchanged (same mtime as the pre-edit build).

(Do not use the gtest test-exe mtime as the stale symptom: in the shared-provider build the exe dlopens the .so and is not relinked, so its mtime stays old even after a correct rebuild — see "How to confirm" below. The reliable diagnostic signal is the fmha_sm*.cu.o / .so mtime.)

So your "successful" rebuild is running the old kernel. Tests that should now pass (or fail) reflect the previous code, not your edit. This silently invalidates any FAIL→PASS / PASS→FAIL verification.

The fix — force recompile the .cu units

Before rebuilding after editing any cutlass_fmha/*.h header:

bash
touch onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/*.cu

Then run the normal build command. This forces the fmha_sm*.cu translation units (and downstream binaries) to recompile against your header change.

How to confirm the rebuild was real (don't trust "[100%] Built")

Confirm that the artifact which actually links the recompiled fmha_sm*.cu.o is newer than your header edit.

⚠️ Do NOT just check the test EXE mtime — it can falsely flag a good build as stale. In the shared-provider build configuration (the default here), the CUDA execution provider is a shared module: the recompiled fmha_sm*.cu.o link into libonnxruntime_providers_cuda.so, and the onnxruntime_provider_test executable dlopens that .so — it is not relinked. So after a correct rebuild the test exe mtime stays old while the .so advances. Checking the exe alone would wrongly conclude the build was stale.

Check the right artifact for your link mode:

  • Shared-provider build (default): the .so that links the recompiled .o — build/<dir>/<cfg>/libonnxruntime_providers_cuda.so
  • Statically-linked provider: the test exe itself (onnxruntime_provider_test)

Safest check — stat both the recompiled object and the .so, and confirm BOTH are newer than the header edit:

bash
stat -c '%y %n' onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/kernel_forward.h
# in your build dir, e.g. build/Debug_quickbuild/Debug/:
stat -c '%y %n' libonnxruntime_providers_cuda.so
# and the actual recompiled object (path varies by build dir):
find . -name 'fmha_sm80.cu.o' -exec stat -c '%y %n' {} +

If the .so (and the fmha_sm*.cu.o) timestamps are older than (or equal to) the header edit, the build was stale — touch the .cu files and rebuild. The most reliable signal of all is behavioral: a test that was failing now passes (a stale binary cannot flip its result).

Show full SKILL.md (151 more words)Show less

This is the CUDA/CUTLASS instance of false-green mode 1 (zero-match / wrong binary) — see the ort-test skill's "False-green taxonomy" for the general principle. In short: attention/MEA/Flash boundary gtests (e.g. FlashStructuralEmptyRows*, Attention_Causal_NonPadKVSeqLen_MEA_*) live in onnxruntime_provider_test, which CI runs; onnxruntime_test_all does not contain them and gives a false green. Verify the MEA/Flash boundary fix against onnxruntime_provider_test.

Full ORT CUDA builds are large (test binaries ~1 GB each; a build dir can reach tens of GB). On a shared box, /home filling to 100% makes builds fail in non-obvious places — e.g. git submodule sync reporting No space left on device or a config.lock error, not an obvious "disk full" at the compile step.

Before a big rebuild, check free space and clean only clearly-stale, regenerable build directories (old dated experiment dirs). Never delete another agent's active build dir or anything ambiguous:

bash
df -h /home
du -sh build/* | sort -h

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/cuda-cutlass-fmha-incremental-rebuild of microsoft/onnxruntime.

Open the folder on GitHubat commit 8fc9fd5

Compare with similar skills

CUTLASS FMHA Incremental Rebuild next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CUTLASS FMHA Incremental Rebuild compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CUTLASS FMHA Incremental Rebuild this skillmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Cuda Debuggingmohitmishra786/low-level-dev-skills252—~1.5kAutomated safety check: PassMIT
Hip Rocmmohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT

Similar skills

  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Debugging

    mohitmishra786/low-level-dev-skills

    CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Hip Rocm

    mohitmishra786/low-level-dev-skills

    HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.6k tokensUpdated 3 mo ago
    DevelopmentAuto-check: notes
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed

More from microsoft/onnxruntime

All 14 skills in this repo
  • Official

    Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.

    22k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • ONNX Runtime Source Build

    microsoft/onnxruntime

    Official

    Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

    22k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • ONNX Runtime CI Management

    microsoft/onnxruntime

    Official

    Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.

    22k GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • ONNX Runtime Release Notes

    microsoft/onnxruntime

    Official

    Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.

    22k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Auto-check passed

Works with

Questions about CUTLASS FMHA Incremental Rebuild

What does CUTLASS FMHA Incremental Rebuild do?

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. cu` units include them.so` unchanged.

When should I use CUTLASS FMHA Incremental Rebuild?

CUTLASS FMHA Incremental Rebuild fits situations like: rebuilding ONNX Runtime CUDA after editing a CUTLASS fused-MHA header; investigating a header edit that passed a build but changed no test behavior; checking that a CUDA test run used freshly compiled kernels; saving disk space on a shared GPU dev box.

How do I install CUTLASS FMHA Incremental Rebuild in Claude Code?

Run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a claude-code`. Or copy the skill folder (.github/skills/cuda-cutlass-fmha-incremental-rebuild in microsoft/onnxruntime) into .claude/skills/cuda-cutlass-fmha-incremental-rebuild in your project. Claude Code loads it when a task matches its description.

How do I install CUTLASS FMHA Incremental Rebuild in Codex?

Run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a codex`. Or copy the skill folder (.github/skills/cuda-cutlass-fmha-incremental-rebuild in microsoft/onnxruntime) into .agents/skills/cuda-cutlass-fmha-incremental-rebuild in your project. Codex loads it when a task matches its description.

Can I use CUTLASS FMHA Incremental Rebuild in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-cutlass-fmha-incremental-rebuild, .gemini/skills/cuda-cutlass-fmha-incremental-rebuild, .github/skills/cuda-cutlass-fmha-incremental-rebuild and .opencode/skills/cuda-cutlass-fmha-incremental-rebuild in your project.

What does CUTLASS FMHA Incremental Rebuild need to run?

Going by SKILL.md and its folder, CUTLASS FMHA Incremental Rebuild needs the command-line tools its instructions call (git). Our summary lists: An ONNX Runtime source checkout with a CUDA build; The NVIDIA CUDA toolchain including nvcc.

Does CUTLASS FMHA Incremental Rebuild access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is CUTLASS FMHA Incremental Rebuild safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CUTLASS FMHA Incremental Rebuild use?

CUTLASS FMHA Incremental Rebuild is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CUTLASS FMHA Incremental Rebuild use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CUTLASS FMHA Incremental Rebuild?

Skills that share tags, products or a category with CUTLASS FMHA Incremental Rebuild: Debug Distributed Hang (sgl-project/sglang, 37k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars), Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars) and Cuda Debugging (mohitmishra786/low-level-dev-skills, 252 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CUTLASS FMHA Incremental Rebuild?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/onnxruntime, which has 22,049 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 11, 2026.

Source: microsoft/onnxruntime on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.