Agent skill

Debug Multi Launch Merge

by Xilinx in Xilinx/mlir-air

A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

MITAuto-check passedSecurity

Install Debug Multi Launch Merge

skills CLI
$ npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air debug-multi-launch-merge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-multi-launch-merge .claude/skills/debug-multi-launch-merge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-multi-launch-merge
GitHub stars
150
Token cost
~1.8k tokens
SKILL.md length
858 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

  • Works in 3 steps: Pad the offending dim to 1024-aligned… → If padding is infeasible: un-merge the… → Switch to kernel-first split-ELF path…
  • Stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion
  • SKILL.md covers Purpose, Knowledge base references, Trigger pattern and Diagnostic hypothesis tree, plus 3 more sections
  • Calls make

What it does

Debug Multi Launch Merge is an agent skill from Xilinx/mlir-air. Use when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA stride limitation). Discriminates the 6 known compile blockers via a symptom-classification table.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Security, covering Threat modeling. The licence is MIT.

When your agent uses it

  • Stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion
  • Channel routing
  • Herd shape conflict
  • IR validation error

Example prompts

  • “/debug-multi-launch-merge”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Pad the offending dim to 1024-aligned via GQA-aware reindexed padding
  2. If padding is infeasible: un-merge the most recently added launch
  3. Switch to kernel-first split-ELF path (more, smaller ELFs)

What it can do on your machine

Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Multi Launch Merge loads about 1.8k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 858 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 858 words, ~1,848 tokens.

Download SKILL.mdSave it as .claude/skills/debug-multi-launch-merge/SKILL.md (or your agent's skills folder).
name
debug-multi-launch-merge
description
Use when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA stride limitation). Discriminates the 6 known compile blockers via a symptom-classification table.

Purpose

When opt-merge-multi-launch-kernels produces a fused ELF and aircc.py rejects it for hardware-resource reasons, this skill identifies which of the 6 documented constraints was hit. Use as a diagnostic decision tree, not a mechanical fix-applier — confirm the trigger matches your stderr before applying the corresponding remedy.

Knowledge base references

  • programming_examples/kernel_registry/details/<Kernel>_bf16.md — per-kernel constraints / placeability notes (the authority for that kernel's hard limits): BD inner-dim 1024, GEMV K-DMA repeat, combined channel reads, L2 cap — the constraints the merge blockers below map back to (for GEMV-specific limits see details/GEMV_bf16.md)
  • programming_examples/kernel_registry/supported_kernels.md — per-kernel constraints + silent-corruption traps (the merge constraints in context of each leaf kernel)
  • programming_examples/llms/llama32_1b/multi_launch_builder/ — working fused-ELF builders to diff a failing merge against (what DOES merge, and the FA-stays-separate boundary)

Trigger pattern

This skill matches when ANY of these appear in aircc.py stderr:

  • buffer descriptor / BD / out of buffer descriptors (BD exhaustion)
  • channel routing / cannot route / channel allocation failed
  • herd shape mismatch / herd dimension
  • aie.tile location conflict
  • airrt.herd_load not found / undefined symbol
  • stride must be 1 for BF16 / sub-32b types

If your error doesn't match any of these patterns, the issue is likely NOT a multi-launch resource collision — invoke superpowers:systematic-debugging instead.

Diagnostic hypothesis tree

For each hypothesis: confirm the symptom-fit, then apply the listed remedy. Don't apply remedies speculatively.

Hypothesis 1: BD exhaustion

Symptom-fit: stderr mentions buffer descriptor / BD / out of buffer descriptors. Often triggered by a non-1024-aligned dim (see the kernel's details/<Kernel>_bf16.md placeability notes) because the DMA auto-splits into many shim BDs that exhaust the pool.

Diagnostic: identify which model dim is non-1024-aligned (e.g. emb_dim=1536). Check the merged ELF's launch count and whether each launch's DMA pattern is BD-friendly (see the kernel's details/<Kernel>_bf16.md).

Remedy (in order of preference):

  1. Pad the offending dim to 1024-aligned via GQA-aware reindexed padding (technique in phase-2-single-block-validation Step 2 + the kernel's details/<Kernel>_bf16.md placeability notes)
  2. If padding is infeasible: un-merge the most recently added launch from the merged ELF; recompile
  3. Switch to kernel-first split-ELF path (more, smaller ELFs)
Hypothesis 2: Channel routing congestion

Symptom-fit: stderr mentions channel routing / cannot route / channel allocation failed. Adjacent launches use overlapping channel IDs that cannot coexist physically on the AIE2P fabric.

Diagnostic: print the merged MLIR (make print or --print-module-only); search for air.channel declarations and check IDs across the merged launches.

Remedy: rename channels in one of the offending launches. The _rename_all(text, prefix=...) helper in programming_examples/llms/llama_kernel_builder/stitching.py already prefixes every SSA name (including channels) per-kernel — confirm your stitching code used distinct prefixes per sub-kernel. If two sub-kernels were stitched with the same prefix, that's the bug. Re-run stitching with distinct prefixes; recompile.

Hypothesis 3: Herd shape conflict

Symptom-fit: stderr mentions herd shape mismatch / herd dimension. Two launches need different herd shapes (e.g., [8,4] for GEMM and [8,1] for RMSNorm) that the placement pass can't reconcile in one segment.

Diagnostic: identify the herd sizes=[N, M] of each sub-kernel. GEMM is typically [8, 4], GEMV/RMSNorm/RoPE/Eltwise are [8, 1].

Remedy: in practice these CAN coexist in one segment when both fit the chip's 8 columns (8 cols × max(M_i) = chip rows). If the compiler still rejects, the cleanest path is to keep the offending kernels as separate XRT calls (don't merge them). This isn't a regression — it's a scoping decision. Document in TODO and accept the slightly higher dispatch overhead.

Show full SKILL.md (323 more words)Show less
Hypothesis 4: Bare air.herd missing the launch+segment wrapper

Symptom-fit: stderr mentions airrt.herd_load not found / undefined symbol / failed to legalize airrt.dma_memcpy_nd.

Diagnostic: same as debug-bo-corruption Hypothesis 4 — search the multi-launch builder for air.herd ops not wrapped in air.launch + air.segment.

Remedy: wrap via _wrap_ir_in_launch(mlir_text) from programming_examples/llms/llama_kernel_builder/stitching.py. The fused builders in llama32_1b/multi_launch_builder/ apply this wrapper around every bare herd (e.g. the RMSNorm and Eltwise-Add herd_x=8 builders).

Hypothesis 5: DMA stride limitation (sub-32b)

Symptom-fit: stderr mentions stride must be 1 for BF16 or other sub-32b types.

Diagnostic: BF16 DMA on AIE2P requires stride=1 for the inner dim. A producer kernel emitting an output with stride > 1 (e.g., a transpose layout with non-contiguous BF16 elements) will hit this.

Remedy: restructure the data layout so the offending DMA has stride=1. Often requires changing memref shape or transpose order in the producer kernel. See compiler_issues/weight_broadcast_dma.md for examples.

Hypothesis 6: Compile-time blowup

Symptom-fit: compile doesn't fail with an error — it just takes

5 minutes (per compiler_scaling.md).

Diagnostic: not a correctness bug; a workflow blocker.

Remedy: reduce the merge scope by dropping the most-recently added launch from the merged set. Document the soft cap in TODO so future deployments don't push past it.

Verification

After applying a fix:

  1. Recompile the merged ELF — should succeed within reasonable time (< 5 min)
  2. Run the merged ELF on the same input as a single XRT call
  3. Compare output against the unmerged baseline (cosine ≥ 0.99; do NOT use rtol=1e-3 — too tight for K ≥ 1024 BF16); log max_abs/max_rel informational

If output mismatch: invoke opt-merge-multi-launch-kernels Step 5 bisect (un-merge the last-added kernel to localize the offender).

Failure mode (when this skill itself can't resolve)

If your stderr doesn't match any of the 6 hypothesis trigger patterns, this is a new compile blocker. Escalate via <model>/TODO.md "Active blockers" with:

  • Full stderr
  • The merged MLIR (make print output)
  • Which sub-kernels were being stitched
  • Hypotheses tried (and ruled out, with diagnostic evidence)

Update protocol

On success: append to <model>/docs/development_progress/debug_log.md:

## debug-multi-launch-merge recovery (YYYY-MM-DD)
- Failing merge: <group_name>
- Hypothesis fired: 1 / 2 / 3 / 4 / 5 / 6
- Fix applied: <one-line description>
- Verified: merged ELF compiles, output cosine ≥ 0.99 vs unmerged baseline

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/debug-multi-launch-merge of Xilinx/mlir-air.

Open the folder on GitHubat commit 6e81ce1

Compare with similar skills

Debug Multi Launch Merge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Multi Launch Merge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Multi Launch Merge this skillXilinx/mlir-air150—~1.8kAutomated safety check: PassMIT
Fla Ascend Performancefla-org/flash-linear-attention5.8k—~6.3kAutomated safety check: PassMIT
Forensifyalexgreensh/repo-forensics188—~2.5kAutomated safety check: NotesCustom licence
Create Rulecartography-cncf/cartography4.1k—~3kAutomated safety check: PassApache-2.0
Commit Security Scancodexstar69/bug-hunter519—~629Automated safety check: PassMIT
Auditing Code For Vulnerabilitiestrilwu/secskills156—~3.2kAutomated safety check: PassMIT

Similar skills

  • Fla Ascend Performance

    fla-org/flash-linear-attention

    Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.

    5.8k GitHub stars~6.3k tokensUpdated today
    SecurityAuto-check passed
  • Forensify

    alexgreensh/repo-forensics

    Cross-agent self-inspection of your AI-agent stack. An agent skill from alexgreensh/repo-forensics.

    188 GitHub stars~2.5k tokensUpdated 11 days ago
    SecurityAuto-check: notes
  • Create Rule

    cartography-cncf/cartography

    Author a Cartography security rule (one or more Cypher Facts plus a Pydantic Finding output model) under cartography/rules/data/rules/.

    4.1k GitHub stars~3k tokensUpdated today
    SecurityAuto-check passed
  • Commit Security Scan

    codexstar69/bug-hunter

    Scan code changes for security vulnerabilities using Bug Hunter-native artifacts and STRIDE context.

    519 GitHub stars~629 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis.

    156 GitHub stars~3.2k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Match identified threats to preventive, detective and corrective controls across network, application, data, endpoint and process layers to plan remediation.

    40k GitHub starsUsed in 8 repos~742 tokens
    SecurityAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Debug Multi Launch Merge

What does Debug Multi Launch Merge do?

A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…. Debug Multi Launch Merge is an agent skill from Xilinx/mlir-air. Use when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA stride limitation).

When should I use Debug Multi Launch Merge?

Debug Multi Launch Merge fits situations like: stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion; channel routing; herd shape conflict; IR validation error.

How do I install Debug Multi Launch Merge in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a claude-code`. Or copy the skill folder (.claude/skills/debug-multi-launch-merge in Xilinx/mlir-air) into .claude/skills/debug-multi-launch-merge in your project. Claude Code loads it when a task matches its description.

How do I install Debug Multi Launch Merge in Codex?

Run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a codex`. Or copy the skill folder (.claude/skills/debug-multi-launch-merge in Xilinx/mlir-air) into .agents/skills/debug-multi-launch-merge in your project. Codex loads it when a task matches its description.

Can I use Debug Multi Launch Merge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill debug-multi-launch-merge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-multi-launch-merge, .gemini/skills/debug-multi-launch-merge, .github/skills/debug-multi-launch-merge and .opencode/skills/debug-multi-launch-merge in your project.

What does Debug Multi Launch Merge need to run?

Going by SKILL.md and its folder, Debug Multi Launch Merge needs the command-line tools its instructions call (make).

Does Debug Multi Launch Merge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Multi Launch Merge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Multi Launch Merge use?

Debug Multi Launch Merge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Multi Launch Merge use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Multi Launch Merge?

Skills that share tags, products or a category with Debug Multi Launch Merge: Fla Ascend Performance (fla-org/flash-linear-attention, 5.8k stars), Forensify (alexgreensh/repo-forensics, 188 stars), Create Rule (cartography-cncf/cartography, 4.1k stars) and Commit Security Scan (codexstar69/bug-hunter, 519 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Multi Launch Merge?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.