Agent skill

Debug Fa Runtime Failure

by Xilinx in Xilinx/mlir-air

A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

MITAuto-check passedDevelopment

Install Debug Fa Runtime Failure

skills CLI
$ npx skills add Xilinx/mlir-air --skill debug-fa-runtime-failure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air debug-fa-runtime-failure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-fa-runtime-failure .claude/skills/debug-fa-runtime-failure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-fa-runtime-failure
GitHub stars
150
Token cost
~1.9k tokens
SKILL.md length
863 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

  • Works in 4 steps: Classify the symptom → Verify .o flag conventions (most common… → Workaround the seq-first dk_chunks > 1… → …
  • NPU FlashAttention hangs (ERTCMDSTATETIMEOUT)
  • SKILL.md covers Purpose, Knowledge base references, Triggers and Workflow, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Debug Fa Runtime Failure is an agent skill from Xilinx/mlir-air. Use when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128. Discriminates the three known root causes (compile-flag mismatch, seq-first dkchunks bug, true L1 overflow) via a symptom-classification table and applies the documented fix.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis. The licence is MIT.

When your agent uses it

  • NPU FlashAttention hangs (ERTCMDSTATETIMEOUT)
  • Produces NaN at headdim ≥ 128

Example prompts

  • “/debug-fa-runtime-failure”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Classify the symptom
  2. Verify .o flag conventions (most common — fixes NaN/garbage)
  3. Workaround the seq-first dk_chunks > 1 upstream bug (Option C)
  4. True L1 overflow (rare)

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are makefile).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Fa Runtime Failure loads about 1.9k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 863 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 863 words, ~1,870 tokens.

Download SKILL.mdSave it as .claude/skills/debug-fa-runtime-failure/SKILL.md (or your agent's skills folder).
name
debug-fa-runtime-failure
description
Use when NPU FlashAttention hangs (`ERT_CMD_STATE_TIMEOUT`) or produces NaN at head_dim ≥ 128. Discriminates the three known root causes (compile-flag mismatch, seq-first dk_chunks bug, true L1 overflow) via a symptom-classification table and applies the documented fix.

Purpose

NPU FlashAttention failures at head_dim ≥ 128 manifest in three distinct ways with three distinct root causes. This skill walks the diagnosis efficiently instead of bisecting from scratch — head_dim=128 deployments hit each of these.

Use this as a diagnostic decision tree, not a mechanical fix-applier — match symptom to hypothesis BEFORE applying the documented remedy.

Knowledge base references

  • programming_examples/flash_attention/kernel_fusion_based/Makefile — canonical -D flag conventions (the ground truth for compile flags)
  • programming_examples/flash_attention/kernel_fusion_based/attn_npu2.py vs .../attn_npu2_seqfirst.py — the two Python builders (head-first vs seq-first); compile from the same C++ kernel but different IR
  • programming_examples/llms/llama_kernel_builder/external_kernels.py — the FA compile API that derives per-tile flags correctly (the compile_attn_npu2* helpers)

Triggers

ALL of these route here:

  • RuntimeError: Command failed to complete successfully (ERT_CMD_STATE_TIMEOUT) from cache.load_and_run("flash_attn", ...) or XRTRunner.run_test
  • All-NaN output from FA when inputs are well-formed (no NaN/Inf in)
  • Compile passes but runtime output is constant garbage (e.g., all 49.0, 844.0, growing magnitudes — typical of softmax-not-running)

Workflow

Step 1: Classify the symptom

Run a minimal repro at the failing shape. Match the symptom:

SymptomMost-likely root causeJump to
HANG (timeout) at all dk_chunks > 1 configs but PASS at dk_chunks = 1Seq-first dk_chunks > 1 upstream bugStep 3
NaN at any config (including ones the lit test passes with uniform(0,4) inputs)compile_attn_npu2* flag mismatch (per-launch sizes baked into .o)Step 2
Garbage non-NaN output (large constant values, no softmax behavior)Same as NaN — flag mismatch, just numerically different surfaceStep 2
HANG at one specific shape but PASS at smaller variantsTrue L1 overflow at the larger shapeStep 4

If the symptom doesn't fit any row, this is a new failure mode — escalate per "Failure mode" at the bottom.

Step 2: Verify .o flag conventions (most common — fixes NaN/garbage)

The attn_npu2.cc kernel's lqp/lkp/dk/dv defines are per-tile, NOT per-launch. The Makefile's convention is canonical (verify against this):

makefile
LQP_TILE := $(shell echo $$(($(LQP) / $(NUM_Q_TILES))))
... -Dlqp=$(LQP_TILE) -Dlkp=$(LKP) \
    -Ddk=$(LKP) -Ddk_full=$(DK) \
    -Ddv=$(LKP) -Ddv_full=$(DV) ...

Diagnostic: diff your compile call against the Makefile. The FA compile helper in programming_examples/llms/llama_kernel_builder/external_kernels.py derives lqp_tile = lqp // num_q_tiles internally and emits the right per-tile flags.

Remedy: if your .o was compiled with per-launch flags, delete the stale .o and the cached flash_attn.elf, then rebuild via the external-kernels FA compile helper (which emits per-tile flags). Re-run.

Step 3: Workaround the seq-first dk_chunks > 1 upstream bug (Option C)

Diagnostic: attn_npu2_seqfirst.py (the seq-first Python builder) has an untested dk_chunks > 1 shim-DMA path that hangs at runtime. Verify via bisect: every dk_chunks=2 config hangs in seq-first, regardless of (n_heads/n_kv, lq=lk). The HEAD-first kernel attn_npu2.py at the same shape PASSES (make run DK=128 DV=128 NUM_HEADS=32 NUM_KV_HEADS=8 → corr ≈ 0.996).

Remedy — Option C (head-first FA + host transposes):

  1. Route the attention call through the head-first kernel (attn_npu2.build_module(...)) instead of seq-first, and add host transposes around it: reshape seq-first [seq, n_heads*head_dim] ↔ head-first [n_heads, seq, head_dim] on the way in and out.
  2. Wrap this as a helper that intercepts the flash_attn cached call so the rest of the pipeline stays seq-first and only the FA call sees head-first layout. Build it once in the deployment's setup() / block-compile path.

Cost: a few ms/layer host transpose. Gain: NPU FA actually runs (a head_dim=128 deployment that fell back to CPU attention recovers a multi-× warm prefill speedup once NPU FA works).

Show full SKILL.md (347 more words)Show less
Step 4: True L1 overflow (rare)

If both Step 2 and Step 3 are clean (correct flags, head-first variant) and the kernel STILL hangs at one specific shape but PASSES at smaller variants, you're hitting the actual 64 KB per-core L1 limit. Per-tile budget for FA at (tile_size_q, lkp, dk_full, dv_full):

Q tile          : tile_size_q * lkp_per_dk_chunk * 2B
K tile (per dk) : lkp * lkp * 2B            (= 8 KB at lkp=64)
V tile (per dv) : lkp * lkp * 2B            (= 8 KB at lkp=64)
Gp accumulator  : tile_size_q * dv_full * 2B
misc (up,sp,r)  : ~2 KB

With lkp=64, the per-dk_chunk budget is small (~8 KB each) and shared buffers are off (lkp != dk_full at hd=128). Sum stays well under 64 KB at typical tile_size_q ≤ 64.

Remedy: drop lqp in the Python builder (which reduces tile_size_q = lqp / num_q_tiles) and recompile.

True L1 overflow is rare for the shapes LLM deployments use. If you hit it, also document the failing shape — it's a useful data point.

Reusable bisect harness

When the symptom doesn't immediately fit Step 1's table, bisect across (n_heads, n_kv_heads, lq=lk, dk) one axis at a time toward your production config. The first axis that flips PASS → HANG/NaN tells you which dimension is the offender.

The pattern is straightforward — invoke the external-kernels FA compile helper to build the .o at a given shape, build the module via attn_npu2[_seqfirst].build_module(...), generate random inputs + NumPy causal-SDPA reference, run via XRTRunner. Catch TIMEOUT → "HANG"; cosine < threshold → "FAIL_NUMERICAL". A <model>_phaseN_test.py script that exercises FA at the production shape is the worked example to mirror.

Verification

This skill is "successful" when the failing FA invocation produces correct output (cosine ≥ Phase 2's head_dim-scaled threshold) and runs without hang. Capture the resolution path in <model>/docs/development_progress/debug_log.md.

Failure mode (when this skill itself can't resolve)

If the symptom matches one of Step 1's rows but the documented remedy in Step 2/3/4 doesn't resolve, OR the symptom doesn't fit any row: this is a new failure mode not covered by current knowledge.

Escalate to the user with:

  • The bisect matrix (which axis was varied, which configs PASS/HANG/NaN)
  • The exact compile flags used
  • The kernel cache state (was .o rebuilt? was flash_attn.elf re-cached?)

Do NOT silently wrap-fix or further-bisect for hours — the 3 documented root causes are well-characterized; a real new failure mode deserves human triage.

Update protocol

On successful diagnosis, append to <model>/docs/development_progress/debug_log.md:

## debug-fa-runtime-failure recovery (YYYY-MM-DD)
- Failing shape: lq=X, lk=Y, dk=Z, hd=W, n_heads=A, n_kv=B
- Symptom: HANG / NaN / garbage
- Step matched: 2 / 3 / 4
- Fix applied: <one-line description>
- Verified: cosine ≥ <threshold> at production shape

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/debug-fa-runtime-failure of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Debug Fa Runtime Failure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Fa Runtime Failure compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Fa Runtime Failure this skillXilinx/mlir-air150—~1.9kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT

Similar skills

  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Om Auto Fix Issue

    go-musicfox/go-musicfox

    Fix or implement a tracker issue end to end from a single command — takes an issue id or a plain problem description (filed first via om-prepare-issue), classifies, then drives the bug autofix chain…

    2.6k GitHub starsUsed in 1 repo~5k tokens
    DevelopmentAuto-check: notes

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Debug Fa Runtime Failure

What does Debug Fa Runtime Failure do?

A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128. Debug Fa Runtime Failure is an agent skill from Xilinx/mlir-air. Use when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

When should I use Debug Fa Runtime Failure?

Debug Fa Runtime Failure fits situations like: NPU FlashAttention hangs (ERTCMDSTATETIMEOUT); produces NaN at headdim ≥ 128.

How do I install Debug Fa Runtime Failure in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill debug-fa-runtime-failure -a claude-code`. Or copy the skill folder (.claude/skills/debug-fa-runtime-failure in Xilinx/mlir-air) into .claude/skills/debug-fa-runtime-failure in your project. Claude Code loads it when a task matches its description.

How do I install Debug Fa Runtime Failure in Codex?

Run `npx skills add Xilinx/mlir-air --skill debug-fa-runtime-failure -a codex`. Or copy the skill folder (.claude/skills/debug-fa-runtime-failure in Xilinx/mlir-air) into .agents/skills/debug-fa-runtime-failure in your project. Codex loads it when a task matches its description.

Can I use Debug Fa Runtime Failure in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill debug-fa-runtime-failure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-fa-runtime-failure, .gemini/skills/debug-fa-runtime-failure, .github/skills/debug-fa-runtime-failure and .opencode/skills/debug-fa-runtime-failure in your project.

What does Debug Fa Runtime Failure need to run?

SKILL.md names no scripts, command-line tools or credentials: Debug Fa Runtime Failure is instructions for the agent only. Our summary lists: Python 3.

Does Debug Fa Runtime Failure access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Fa Runtime Failure safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Fa Runtime Failure use?

Debug Fa Runtime Failure is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Fa Runtime Failure use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Fa Runtime Failure?

Skills that share tags, products or a category with Debug Fa Runtime Failure: Code Design Rationale Investigator (cursor/plugins, 10k stars), OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars) and Root Cause Debugging (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Fa Runtime Failure?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.