Debug
asgeirtj/system_prompts_leaks
Enable debug logging for this session and help diagnose issues
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .claude/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .claude/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .agents/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .agents/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .cursor/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .cursor/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/debug-bo-corruption--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .gemini/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .gemini/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air debug-bo-corruptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .github/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .github/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .opencode/skills/debug-bo-corruption && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debug-bo-corruption" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-bo-corruption into .opencode/skills/debug-bo-corruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-bo-corruption", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debug-bo-corruptionA skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Debug Bo Corruption is an agent skill from Xilinx/mlir-air. Use when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. Common symptoms: correct first invocation but wrong on subsequent calls; correct in isolation but wrong when chained with other kernels.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debug Bo Corruption loads about 1.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 715 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 715 words, ~1,525 tokens.
.claude/skills/debug-bo-corruption/SKILL.md (or your agent's skills folder).Diagnose Buffer Object (BO) corruption — bugs that surface only at integration time, characterized by a kernel that passes standalone correctness but produces wrong output in a larger pipeline. The 4 hypotheses below cover every BO-corruption case observed across the 6 LLM deployments to date. Use this skill as a diagnostic checklist, not a mechanical fix-applier — understand which hypothesis fits before applying the corresponding remedy.
programming_examples/llms/llama_kernel_builder/cache.py
— static_input_indices, intermediate_indices mechanics (the
KernelCache host-optimization knobs)programming_examples/llms/llama_kernel_builder/stitching.py
— _wrap_ir_in_launch is the helper that wraps a bare herd in
air.launch + air.segment (see hypothesis 4 below); also where
airrt.herd_load vs airrt.segment_load semantics matterThis skill matches when ANY of these apply:
For each hypothesis: check the symptom-fit, then apply the listed remedy ONLY if the diagnostic matches. Don't apply remedies blindly.
Symptom-fit: Output is correct on first call, wrong on subsequent calls. Wrong values match a different invocation.
Diagnostic: Find the kernel's cache.load_and_run(...) call.
Inspect whether static_input_indices=[...] is passed. Cross-reference
the kernel's MLIR signature: which input indices are weights (written
once, reused) vs activations (overwritten each call)?
Remedy: Add the weight indices to static_input_indices so
load_and_run skips re-writing them every call. Re-test.
Symptom-fit: Output buffer has stale values from previous call; NaN on first call (uninitialized) but a "successful" cosine number later.
Diagnostic: Identify buffers in the kernel that are outputs the
kernel will fully overwrite. Check whether intermediate_indices=[...]
lists these.
Remedy: Add the output indices to intermediate_indices so
load_and_run skips the initial host write. Re-test.
Symptom-fit: Per-layer kernel call works in isolation; multiple layers chained produce wrong output. Different layers are clobbering each other's BO state.
Diagnostic: Search for bo_key strings across layers. Confirm
each layer uses a unique key (the convention is
f"kernel_L{layer_idx}").
Remedy: Parameterize on layer index. Re-test after fix in a multi-layer call (e.g., 2-layer test before going to N-layer).
air.herd missing the air.launch + air.segment wrapperSymptom-fit: Output is all-zero (or undefined) — the kernel "succeeds" silently but never actually executes the DMA path (a silent-corruption pattern: cosine ≈ 0 / NaN).
Diagnostic: Search the multi-launch builder (or the standalone
test harness) for air.herd ops. Check each is wrapped in `air.launch
. Reference: the airrt-to-npulowering pass needsairrt.segment_load(notairrt.herd_load) to attach the launch region to the aie.device` op. A bare herd without the segment wrapper
gets silently dropped.Remedy: Wrap via _wrap_ir_in_launch(mlir_text) from
programming_examples/llms/llama_kernel_builder/stitching.py. The fused multi-launch
builders in llama32_1b/multi_launch_builder/ use this wrapper around
every bare herd (e.g. the RMSNorm and Eltwise-Add herd_x=8 builders,
which emit bare herds) — mirror that pattern.
Re-test.
After applying any fix:
rtol=1e-3 — too tight for K ≥ 1024 BF16)If fix succeeded: record recovered_via=debug-bo-corruption and
which hypothesis fired in
<model>/docs/development_progress/debug_log.md.
If no hypothesis fits the symptom OR no remedy works: this is a new
failure mode. Escalate via <model>/TODO.md "Active blockers" with
the failing test, the hypotheses tried, and the unchanged failure
output. Don't apply remedies speculatively.
| Symptom doesn't fit any of 4 hypotheses | What to check |
|---|---|
| Output is wrong on iter 1 too (not just stale across calls) | Not BO corruption; this is a kernel correctness bug — re-run Phase 1 standalone test for that kernel |
| Output is wrong only at specific seq_len / layer index | Could be KV cache layout bug; invoke superpowers:systematic-debugging |
bo_key is unique per layer but still aliasing | KernelCache may be reusing artifact across bo_key — inspect cache.artifacts dict at runtime |
On success: append to <model>/docs/development_progress/debug_log.md:
## debug-bo-corruption recovery (YYYY-MM-DD)
- Failing item: <kernel/test name>
- Hypothesis fired: 1 / 2 / 3 / 4
- Fix applied: <one-line description>
- Verified: 3/3 consistent runs, cosine ≥ 0.99On escalation: append to <model>/TODO.md Active blockers with full
failure context and which hypotheses were ruled out.
© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/debug-bo-corruption of Xilinx/mlir-air.
Open the folder on GitHubat commit bca27e5
Debug Bo Corruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debug Bo Corruption this skillXilinx/mlir-air | 150 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Debugasgeirtj/system_prompts_leaks | 69k | — | ~439 | Automated safety check: Pass | CC0-1.0 | |
| Debugging Toolkitsickn33/agentic-awesome-skills | 47k | 1 repos | ~344 | Automated safety check: Pass | MIT | |
| Hypothesis-Driven Debuggingcode-yeongyu/oh-my-openagent | 70k | — | ~3.2k | Automated safety check: Pass | Custom licence | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| Agent Introspection Debuggingaffaan-m/ECC | 274k | 2 repos | ~1.4k | Automated safety check: Pass | MIT |
asgeirtj/system_prompts_leaks
Enable debug logging for this session and help diagnose issues
sickn33/agentic-awesome-skills
A skill your agent uses when working with debugging toolkit smart debug (Alias for debugging-toolkit-smart-debug)
code-yeongyu/oh-my-openagent
Runs a hypothesis-driven debugging loop for crashes, hangs and silent failures in any language, grounding every claim in runtime evidence and locking the fix with a test.
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
affaan-m/ECC
Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports.
parcadei/Continuous-Claude-v3
Debug issues by investigating logs, database state, and git history
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Xilinx/mlir-air
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. Debug Bo Corruption is an agent skill from Xilinx/mlir-air. Use when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Debug Bo Corruption fits situations like: an NPU kernel passes its standalone shape test but produces NaN; stale values when invoked as part of a larger pipeline.
Run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a claude-code`. Or copy the skill folder (.claude/skills/debug-bo-corruption in Xilinx/mlir-air) into .claude/skills/debug-bo-corruption in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a codex`. Or copy the skill folder (.claude/skills/debug-bo-corruption in Xilinx/mlir-air) into .agents/skills/debug-bo-corruption in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-bo-corruption, .gemini/skills/debug-bo-corruption, .github/skills/debug-bo-corruption and .opencode/skills/debug-bo-corruption in your project.
SKILL.md names no scripts, command-line tools or credentials: Debug Bo Corruption is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debug Bo Corruption is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Debug Bo Corruption: Debug (asgeirtj/system_prompts_leaks, 69k stars), Debugging Toolkit (sickn33/agentic-awesome-skills, 47k stars), Hypothesis-Driven Debugging (code-yeongyu/oh-my-openagent, 70k stars) and Aoti Debug (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.