Fla Ascend Performance
fla-org/flash-linear-attention
Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .claude/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .claude/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-mergeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .agents/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .agents/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .cursor/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .cursor/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/debug-multi-launch-merge--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .gemini/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .gemini/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air debug-multi-launch-mergeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .github/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .github/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .opencode/skills/debug-multi-launch-merge && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debug-multi-launch-merge" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/debug-multi-launch-merge into .opencode/skills/debug-multi-launch-merge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-multi-launch-merge", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debug-multi-launch-mergeA skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Debug Multi Launch Merge is an agent skill from Xilinx/mlir-air. Use when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA stride limitation). Discriminates the 6 known compile blockers via a symptom-classification table.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Security, covering Threat modeling. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debug Multi Launch Merge loads about 1.8k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 858 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 858 words, ~1,848 tokens.
.claude/skills/debug-multi-launch-merge/SKILL.md (or your agent's skills folder).When opt-merge-multi-launch-kernels produces a fused ELF and aircc.py
rejects it for hardware-resource reasons, this skill identifies which
of the 6 documented constraints was hit. Use as a diagnostic
decision tree, not a mechanical fix-applier — confirm the trigger
matches your stderr before applying the corresponding remedy.
programming_examples/kernel_registry/details/<Kernel>_bf16.md —
per-kernel constraints / placeability notes (the authority for that
kernel's hard limits): BD inner-dim 1024, GEMV K-DMA repeat, combined
channel reads, L2 cap — the constraints the merge blockers below map
back to (for GEMV-specific limits see details/GEMV_bf16.md)programming_examples/kernel_registry/supported_kernels.md
— per-kernel constraints + silent-corruption traps (the merge
constraints in context of each leaf kernel)programming_examples/llms/llama32_1b/multi_launch_builder/ — working
fused-ELF builders to diff a failing merge against (what DOES merge,
and the FA-stays-separate boundary)This skill matches when ANY of these appear in aircc.py stderr:
buffer descriptor / BD / out of buffer descriptors (BD exhaustion)channel routing / cannot route / channel allocation failedherd shape mismatch / herd dimensionaie.tile location conflictairrt.herd_load not found / undefined symbolstride must be 1 for BF16 / sub-32b typesIf your error doesn't match any of these patterns, the issue is likely
NOT a multi-launch resource collision — invoke
superpowers:systematic-debugging instead.
For each hypothesis: confirm the symptom-fit, then apply the listed remedy. Don't apply remedies speculatively.
Symptom-fit: stderr mentions buffer descriptor / BD / out of buffer descriptors. Often triggered by a non-1024-aligned dim (see the
kernel's details/<Kernel>_bf16.md placeability notes) because the DMA
auto-splits into many shim BDs that exhaust the pool.
Diagnostic: identify which model dim is non-1024-aligned (e.g.
emb_dim=1536). Check the merged ELF's launch count and whether each
launch's DMA pattern is BD-friendly (see the kernel's
details/<Kernel>_bf16.md).
Remedy (in order of preference):
phase-2-single-block-validation Step 2 + the kernel's details/<Kernel>_bf16.md placeability notes)Symptom-fit: stderr mentions channel routing / cannot route /
channel allocation failed. Adjacent launches use overlapping channel
IDs that cannot coexist physically on the AIE2P fabric.
Diagnostic: print the merged MLIR (make print or
--print-module-only); search for air.channel declarations and
check IDs across the merged launches.
Remedy: rename channels in one of the offending launches. The
_rename_all(text, prefix=...) helper in
programming_examples/llms/llama_kernel_builder/stitching.py already prefixes every SSA
name (including channels) per-kernel — confirm your stitching code
used distinct prefixes per sub-kernel. If two sub-kernels were
stitched with the same prefix, that's the bug. Re-run stitching with
distinct prefixes; recompile.
Symptom-fit: stderr mentions herd shape mismatch / herd dimension. Two launches need different herd shapes (e.g., [8,4] for
GEMM and [8,1] for RMSNorm) that the placement pass can't reconcile
in one segment.
Diagnostic: identify the herd sizes=[N, M] of each sub-kernel.
GEMM is typically [8, 4], GEMV/RMSNorm/RoPE/Eltwise are [8, 1].
Remedy: in practice these CAN coexist in one segment when both fit the chip's 8 columns (8 cols × max(M_i) = chip rows). If the compiler still rejects, the cleanest path is to keep the offending kernels as separate XRT calls (don't merge them). This isn't a regression — it's a scoping decision. Document in TODO and accept the slightly higher dispatch overhead.
air.herd missing the launch+segment wrapperSymptom-fit: stderr mentions airrt.herd_load not found /
undefined symbol / failed to legalize airrt.dma_memcpy_nd.
Diagnostic: same as debug-bo-corruption Hypothesis 4 — search
the multi-launch builder for air.herd ops not wrapped in
air.launch + air.segment.
Remedy: wrap via _wrap_ir_in_launch(mlir_text) from
programming_examples/llms/llama_kernel_builder/stitching.py. The fused builders in
llama32_1b/multi_launch_builder/ apply this wrapper around every bare
herd (e.g. the RMSNorm and Eltwise-Add herd_x=8 builders).
Symptom-fit: stderr mentions stride must be 1 for BF16 or other
sub-32b types.
Diagnostic: BF16 DMA on AIE2P requires stride=1 for the inner
dim. A producer kernel emitting an output with stride > 1 (e.g., a
transpose layout with non-contiguous BF16 elements) will hit this.
Remedy: restructure the data layout so the offending DMA has
stride=1. Often requires changing memref shape or transpose order
in the producer kernel. See
compiler_issues/weight_broadcast_dma.md for examples.
Symptom-fit: compile doesn't fail with an error — it just takes
5 minutes (per
compiler_scaling.md).
Diagnostic: not a correctness bug; a workflow blocker.
Remedy: reduce the merge scope by dropping the most-recently added launch from the merged set. Document the soft cap in TODO so future deployments don't push past it.
After applying a fix:
rtol=1e-3 — too tight for K ≥ 1024 BF16); log
max_abs/max_rel informationalIf output mismatch: invoke opt-merge-multi-launch-kernels Step 5 bisect
(un-merge the last-added kernel to localize the offender).
If your stderr doesn't match any of the 6 hypothesis trigger patterns,
this is a new compile blocker. Escalate via <model>/TODO.md "Active
blockers" with:
make print output)On success: append to <model>/docs/development_progress/debug_log.md:
## debug-multi-launch-merge recovery (YYYY-MM-DD)
- Failing merge: <group_name>
- Hypothesis fired: 1 / 2 / 3 / 4 / 5 / 6
- Fix applied: <one-line description>
- Verified: merged ELF compiles, output cosine ≥ 0.99 vs unmerged baseline© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/debug-multi-launch-merge of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Debug Multi Launch Merge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debug Multi Launch Merge this skillXilinx/mlir-air | 150 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Fla Ascend Performancefla-org/flash-linear-attention | 5.8k | — | ~6.3k | Automated safety check: Pass | MIT | |
| Forensifyalexgreensh/repo-forensics | 188 | — | ~2.5k | Automated safety check: Notes | Custom licence | |
| Create Rulecartography-cncf/cartography | 4.1k | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Commit Security Scancodexstar69/bug-hunter | 519 | — | ~629 | Automated safety check: Pass | MIT | |
| Auditing Code For Vulnerabilitiestrilwu/secskills | 156 | — | ~3.2k | Automated safety check: Pass | MIT |
fla-org/flash-linear-attention
Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.
alexgreensh/repo-forensics
Cross-agent self-inspection of your AI-agent stack. An agent skill from alexgreensh/repo-forensics.
cartography-cncf/cartography
Author a Cartography security rule (one or more Cypher Facts plus a Pydantic Finding output model) under cartography/rules/data/rules/.
codexstar69/bug-hunter
Scan code changes for security vulnerabilities using Bug Hunter-native artifacts and STRIDE context.
trilwu/secskills
Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis.
wshobson/agents
Match identified threats to preventive, detective and corrective controls across network, application, data, endpoint and process layers to plan remediation.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Xilinx/mlir-air
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
Categories
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…. Debug Multi Launch Merge is an agent skill from Xilinx/mlir-air. Use when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA stride limitation).
Debug Multi Launch Merge fits situations like: stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion; channel routing; herd shape conflict; IR validation error.
Run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a claude-code`. Or copy the skill folder (.claude/skills/debug-multi-launch-merge in Xilinx/mlir-air) into .claude/skills/debug-multi-launch-merge in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a codex`. Or copy the skill folder (.claude/skills/debug-multi-launch-merge in Xilinx/mlir-air) into .agents/skills/debug-multi-launch-merge in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-multi-launch-merge, .gemini/skills/debug-multi-launch-merge, .github/skills/debug-multi-launch-merge and .opencode/skills/debug-multi-launch-merge in your project.
Going by SKILL.md and its folder, Debug Multi Launch Merge needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debug Multi Launch Merge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Debug Multi Launch Merge: Fla Ascend Performance (fla-org/flash-linear-attention, 5.8k stars), Forensify (alexgreensh/repo-forensics, 188 stars), Create Rule (cartography-cncf/cartography, 4.1k stars) and Commit Security Scan (codexstar69/bug-hunter, 519 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.