Agent skill

Debug Bo Corruption

by Xilinx in Xilinx/mlir-air

A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

MITAuto-check passed

Install Debug Bo Corruption

skills CLI
$ npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air debug-bo-corruption --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-bo-corruption .claude/skills/debug-bo-corruption && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-bo-corruption
GitHub stars
150
Token cost
~1.5k tokens
SKILL.md length
715 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

  • Works in 3 steps: Re-run the failing test (the standalone… → Confirm output cosine ≥ 0.99 vs CPU… → Run the test 3 times in a loop to…
  • An NPU kernel passes its standalone shape test but produces NaN
  • SKILL.md covers Purpose, Knowledge base references, Trigger pattern and Diagnostic hypothesis tree, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Debug Bo Corruption is an agent skill from Xilinx/mlir-air. Use when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. Common symptoms: correct first invocation but wrong on subsequent calls; correct in isolation but wrong when chained with other kernels.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

When your agent uses it

  • An NPU kernel passes its standalone shape test but produces NaN
  • Stale values when invoked as part of a larger pipeline

Example prompts

  • “/debug-bo-corruption”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Re-run the failing test (the standalone XRT runner test or the
  2. Confirm output cosine ≥ 0.99 vs CPU reference (BF16 convention; do
  3. Run the test 3 times in a loop to confirm consistency across

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Bo Corruption loads about 1.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 715 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 715 words, ~1,525 tokens.

Download SKILL.mdSave it as .claude/skills/debug-bo-corruption/SKILL.md (or your agent's skills folder).
name
debug-bo-corruption
description
Use when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. Common symptoms: correct first invocation but wrong on subsequent calls; correct in isolation but wrong when chained with other kernels.

Purpose

Diagnose Buffer Object (BO) corruption — bugs that surface only at integration time, characterized by a kernel that passes standalone correctness but produces wrong output in a larger pipeline. The 4 hypotheses below cover every BO-corruption case observed across the 6 LLM deployments to date. Use this skill as a diagnostic checklist, not a mechanical fix-applier — understand which hypothesis fits before applying the corresponding remedy.

Knowledge base references

  • programming_examples/llms/llama_kernel_builder/cache.py — static_input_indices, intermediate_indices mechanics (the KernelCache host-optimization knobs)
  • programming_examples/llms/llama_kernel_builder/stitching.py — _wrap_ir_in_launch is the helper that wraps a bare herd in air.launch + air.segment (see hypothesis 4 below); also where airrt.herd_load vs airrt.segment_load semantics matter

Trigger pattern

This skill matches when ANY of these apply:

  • Output tensor contains NaN despite kernel passing standalone XRTRunner test
  • Kernel produces correct output on first invocation but wrong on second+ invocation in a loop
  • Kernel produces correct output standalone but wrong when chained with another kernel
  • Output tensor has correct shape but values match a previous layer/iteration

Diagnostic hypothesis tree

For each hypothesis: check the symptom-fit, then apply the listed remedy ONLY if the diagnostic matches. Don't apply remedies blindly.

Hypothesis 1: Stale BO state from a prior call

Symptom-fit: Output is correct on first call, wrong on subsequent calls. Wrong values match a different invocation.

Diagnostic: Find the kernel's cache.load_and_run(...) call. Inspect whether static_input_indices=[...] is passed. Cross-reference the kernel's MLIR signature: which input indices are weights (written once, reused) vs activations (overwritten each call)?

Remedy: Add the weight indices to static_input_indices so load_and_run skips re-writing them every call. Re-test.

Hypothesis 2: Intermediate buffer not marked

Symptom-fit: Output buffer has stale values from previous call; NaN on first call (uninitialized) but a "successful" cosine number later.

Diagnostic: Identify buffers in the kernel that are outputs the kernel will fully overwrite. Check whether intermediate_indices=[...] lists these.

Remedy: Add the output indices to intermediate_indices so load_and_run skips the initial host write. Re-test.

Hypothesis 3: Buffer aliasing across layers

Symptom-fit: Per-layer kernel call works in isolation; multiple layers chained produce wrong output. Different layers are clobbering each other's BO state.

Diagnostic: Search for bo_key strings across layers. Confirm each layer uses a unique key (the convention is f"kernel_L{layer_idx}").

Remedy: Parameterize on layer index. Re-test after fix in a multi-layer call (e.g., 2-layer test before going to N-layer).

Show full SKILL.md (340 more words)Show less
Hypothesis 4: Bare air.herd missing the air.launch + air.segment wrapper

Symptom-fit: Output is all-zero (or undefined) — the kernel "succeeds" silently but never actually executes the DMA path (a silent-corruption pattern: cosine ≈ 0 / NaN).

Diagnostic: Search the multi-launch builder (or the standalone test harness) for air.herd ops. Check each is wrapped in `air.launch

  • air.segment. Reference: the airrt-to-npulowering pass needsairrt.segment_load(notairrt.herd_load) to attach the launch region to the aie.device` op. A bare herd without the segment wrapper gets silently dropped.

Remedy: Wrap via _wrap_ir_in_launch(mlir_text) from programming_examples/llms/llama_kernel_builder/stitching.py. The fused multi-launch builders in llama32_1b/multi_launch_builder/ use this wrapper around every bare herd (e.g. the RMSNorm and Eltwise-Add herd_x=8 builders, which emit bare herds) — mirror that pattern.

Re-test.

Verification

After applying any fix:

  1. Re-run the failing test (the standalone XRT runner test or the integration test that surfaced the bug)
  2. Confirm output cosine ≥ 0.99 vs CPU reference (BF16 convention; do NOT use rtol=1e-3 — too tight for K ≥ 1024 BF16)
  3. Run the test 3 times in a loop to confirm consistency across invocations (NaN on iter 2+ is a different bug than failure on iter 1)

If fix succeeded: record recovered_via=debug-bo-corruption and which hypothesis fired in <model>/docs/development_progress/debug_log.md.

If no hypothesis fits the symptom OR no remedy works: this is a new failure mode. Escalate via <model>/TODO.md "Active blockers" with the failing test, the hypotheses tried, and the unchanged failure output. Don't apply remedies speculatively.

Failure modes (when this skill itself can't resolve)

Symptom doesn't fit any of 4 hypothesesWhat to check
Output is wrong on iter 1 too (not just stale across calls)Not BO corruption; this is a kernel correctness bug — re-run Phase 1 standalone test for that kernel
Output is wrong only at specific seq_len / layer indexCould be KV cache layout bug; invoke superpowers:systematic-debugging
bo_key is unique per layer but still aliasingKernelCache may be reusing artifact across bo_key — inspect cache.artifacts dict at runtime

Update protocol

On success: append to <model>/docs/development_progress/debug_log.md:

## debug-bo-corruption recovery (YYYY-MM-DD)
- Failing item: <kernel/test name>
- Hypothesis fired: 1 / 2 / 3 / 4
- Fix applied: <one-line description>
- Verified: 3/3 consistent runs, cosine ≥ 0.99

On escalation: append to <model>/TODO.md Active blockers with full failure context and which hypotheses were ruled out.

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/debug-bo-corruption of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Debug Bo Corruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Bo Corruption compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Bo Corruption this skillXilinx/mlir-air150—~1.5kAutomated safety check: PassMIT
Debugasgeirtj/system_prompts_leaks69k—~439Automated safety check: PassCC0-1.0
Debugging Toolkitsickn33/agentic-awesome-skills47k1 repos~344Automated safety check: PassMIT
Hypothesis-Driven Debuggingcode-yeongyu/oh-my-openagent70k—~3.2kAutomated safety check: PassCustom licence
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
Agent Introspection Debuggingaffaan-m/ECC274k2 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Debug

    asgeirtj/system_prompts_leaks

    Enable debug logging for this session and help diagnose issues

    69k GitHub stars~439 tokensUpdated yesterday
    Auto-check passed
  • Debugging Toolkit

    sickn33/agentic-awesome-skills

    A skill your agent uses when working with debugging toolkit smart debug (Alias for debugging-toolkit-smart-debug)

    47k GitHub starsUsed in 1 repo~344 tokens
    DevelopmentAuto-check passed
  • Hypothesis-Driven Debugging

    code-yeongyu/oh-my-openagent

    Runs a hypothesis-driven debugging loop for crashes, hangs and silent failures in any language, grounding every claim in runtime evidence and locking the fix with a test.

    70k GitHub stars~3.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports.

    274k GitHub starsUsed in 2 repos~1.4k tokens
    DevelopmentAuto-check passed
  • Debug

    parcadei/Continuous-Claude-v3

    Debug issues by investigating logs, database state, and git history

    3.9k GitHub starsUsed in 1 repo~1.3k tokens
    DevelopmentAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Debug Bo Corruption

What does Debug Bo Corruption do?

A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. Debug Bo Corruption is an agent skill from Xilinx/mlir-air. Use when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

When should I use Debug Bo Corruption?

Debug Bo Corruption fits situations like: an NPU kernel passes its standalone shape test but produces NaN; stale values when invoked as part of a larger pipeline.

How do I install Debug Bo Corruption in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a claude-code`. Or copy the skill folder (.claude/skills/debug-bo-corruption in Xilinx/mlir-air) into .claude/skills/debug-bo-corruption in your project. Claude Code loads it when a task matches its description.

How do I install Debug Bo Corruption in Codex?

Run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a codex`. Or copy the skill folder (.claude/skills/debug-bo-corruption in Xilinx/mlir-air) into .agents/skills/debug-bo-corruption in your project. Codex loads it when a task matches its description.

Can I use Debug Bo Corruption in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill debug-bo-corruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-bo-corruption, .gemini/skills/debug-bo-corruption, .github/skills/debug-bo-corruption and .opencode/skills/debug-bo-corruption in your project.

What does Debug Bo Corruption need to run?

SKILL.md names no scripts, command-line tools or credentials: Debug Bo Corruption is instructions for the agent only.

Does Debug Bo Corruption access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Bo Corruption safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Bo Corruption use?

Debug Bo Corruption is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Bo Corruption use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Bo Corruption?

Skills that share tags, products or a category with Debug Bo Corruption: Debug (asgeirtj/system_prompts_leaks, 69k stars), Debugging Toolkit (sickn33/agentic-awesome-skills, 47k stars), Hypothesis-Driven Debugging (code-yeongyu/oh-my-openagent, 70k stars) and Aoti Debug (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Bo Corruption?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.