Agent skill

Opt Merge Multi Launch Kernels

by Xilinx in Xilinx/mlir-air

Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

MITAuto-check passed

Install Opt Merge Multi Launch Kernels

skills CLI
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .claude/skills/opt-merge-multi-launch-kernels && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opt-merge-multi-launch-kernels
GitHub stars
150
Token cost
~1.8k tokens
SKILL.md length
795 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

  • Works in 6 steps: Identify merge candidates → Pick merge boundaries → Author the multi-launch builder → …
  • SKILL.md covers Purpose, Success criteria, Knowledge base references and Workflow, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Opt Merge Multi Launch Kernels is an agent skill from Xilinx/mlir-air. Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). Invoked by phase-4-prefill-optimization and phase-5-decode-optimization to fuse kernel groups when building NEW model-specific fused ELFs (kernel-first path). Reduces XRT dispatch overhead (~50–200 µs per call on NPU2).

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

Example prompts

  • “/opt-merge-multi-launch-kernels”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Identify merge candidates
  2. Pick merge boundaries
  3. Author the multi-launch builder
  4. Compile via KernelCache
  5. Validate output vs unmerged baseline
  6. Measure perf gain

What it can do on your machine

Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opt Merge Multi Launch Kernels loads about 1.8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 795 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 795 words, ~1,819 tokens.

Download SKILL.mdSave it as .claude/skills/opt-merge-multi-launch-kernels/SKILL.md (or your agent's skills folder).
name
opt-merge-multi-launch-kernels
description
Procedural recipe for fusing multiple `air.launch` kernels into one multi-launch ELF (single XRT invocation). Invoked by phase-4-prefill-optimization and phase-5-decode-optimization to fuse kernel groups when building NEW model-specific fused ELFs (kernel-first path). Reduces XRT dispatch overhead (~50–200 µs per call on NPU2).

Purpose

When a deployment goes the kernel-first path in Phase 4/5 (model has new ops or stitching boundaries differ from llama), build new model-specific fused ELFs by stitching multiple air.launch kernels into a single func.func. This skill is the procedural recipe.

The reference llama3 deployment went 10 launches/layer → 3 launches/layer using this recipe — the dominant prefill perf win.

Success criteria

A merge is "successful" when ALL hold:

  1. Merged ELF compiles without errors
  2. Output cosine ≥ 0.99 vs the unmerged baseline (run the same kernels as separate XRT calls, compare output). Match the BF16 tolerance convention from the kernel registry — rtol=1e-3 is too tight for K=2048 BF16; use cosine + max_abs/max_rel informational logging (see phase-1-kernel-validation PASS criteria for the convention).
  3. Wall-clock time is lower than the unmerged baseline

If (1) fails → invoke debug-multi-launch-merge. If (2) fails → bisect by un-merging the last-added kernel. If (3) fails (compile slower than unmerged) → not a correctness bug; document and accept or revert.

Knowledge base references

  • programming_examples/llms/llama_kernel_builder/stitching.py — the helpers this recipe uses (_rename_all, _fix_launch_func_args, _wrap_ir_in_launch, _rename_all_with_externs, _rename_all_gemv)
  • programming_examples/llms/llama32_1b/multi_launch_builder/rms_gemms_rope_multi.py — reference 6-launch prefill merge (RMSNorm + Q/K/V GEMM + RoPE Q/K)
  • programming_examples/llms/llama32_1b/multi_launch_builder/o_ffn_multi.py — reference 8-launch prefill merge (O + add + RMSNorm + Gate/Up + SwiGLU + Down + add)
  • programming_examples/llms/llama32_1b/multi_launch_builder/o_gemv_ffn_multi.py — decode merge with 2-K extern rename (the pattern to extend to 3-K when n_heads*head_dim != emb_dim)
  • programming_examples/kernel_registry/supported_kernels.md — per-kernel constraints; the FA section notes FlashAttention stays a separate XRT call (does NOT merge into the fused block)

Workflow

Step 1: Identify merge candidates

List the per-layer kernel sequence. Mark each as one of:

  • Mergeable — pure compute (RMSNorm, GEMM, GEMV, RoPE, eltwise)
  • Hard-stop — has data-dependent control flow or unsupported merge pattern. FlashAttention is the canonical hard-stop (see full_block.md)
  • Conditional — mergeable but check herd-shape compatibility with neighbors (e.g., GEMM uses 8×4 herd; eltwise uses 8×1; coexisting in one segment is OK but verify)
Step 2: Pick merge boundaries

Group consecutive mergeable launches between hard-stops. For a typical decoder-only LLM:

  • Group A: RMSNorm + Q/K/V proj + RoPE Q + RoPE K (6 launches — rms_gemms_rope_multi.py shape)
  • Hard-stop: FlashAttention (separate XRT call)
  • Group B: O proj + Add + RMSNorm + Gate + Up + SiLU+Mul + Down + Add (8 launches — o_ffn_multi.py shape)

For decode, replace GEMM with GEMV throughout.

Step 3: Author the multi-launch builder

For each group, create <model>/multi_launch_builder/<group_name>_multi.py that:

  1. Builds each sub-kernel's IR via @module_builder
  2. Imports stitching helpers: from llama_kernel_builder.stitching import (_rename_all, _fix_launch_func_args, _wrap_ir_in_launch, ...) (resolved against the shared programming_examples/llms/llama_kernel_builder/ via sys.path)
  3. For each sub-kernel: extract its function body via _extract_between_func_and_return(ir_text)
  4. Rename SSA values with a per-kernel prefix via _rename_all(body, prefix=...) to avoid collisions across the merged module
  5. Remap function arguments to the merged module's args via _fix_launch_func_args(body, prefix, arg_map)
  6. Concatenate the renamed bodies into a single func.func. If any sub-kernel emits a bare air.herd (RMSNorm herd_x>1, Eltwise Add at herd_x>1), wrap each via _wrap_ir_in_launch(...) BEFORE stitching — otherwise the lowering's airrt-to-npu pass drops the bare herd
  7. For multi-K GEMV in one ELF (decode kernel-first): use _rename_all_with_externs with per-launch extern allowlists to keep different mv_*.o symbols distinct (extend the 2-K rename in o_gemv_ffn_multi.py to 3-K when n_heads*head_dim != emb_dim)

The canonical reference is rms_gemms_rope_multi.py — copy from there.

Show full SKILL.md (291 more words)Show less
Step 4: Compile via KernelCache

Use KernelCache.compile_and_cache(name=<group_name>, builder=<your_builder_function>, ...).

If compile fails → invoke debug-multi-launch-merge.

Step 5: Validate output vs unmerged baseline

Run the merged ELF on a fixed input. Run the same kernels as separate XRT calls (the unmerged baseline). Compare the merged output to the unmerged output:

  • cosine ≥ 0.99 (gate)
  • log max_abs / max_rel informational

If cosine fails: bisect by un-merging the last-added kernel from Step 3. If still fails after un-merging that one, the bug is in an EARLIER merge addition — bisect further.

Step 6: Measure perf gain

Profile the merged version vs unmerged. Expected at NPU2 dispatch overhead levels: ≥ 20% reduction per merged group at moderate scale (llama3 saw 16-XRT-call/layer prefill → 3-XRT-call/layer = much larger reduction). If the gain is much smaller than expected, the per-call XRT overhead may not have been the bottleneck — the opt-buffer-object-reuse skill (static weight BOs) is often the missing piece in that case.

Record gain in <model>/docs/development_progress/phase{4,5}_*.md.

Failure modes

SymptomLikely causeWhere to look
Compile fails (BD exhaustion, channel routing, herd shape conflict, IR validation)Multi-launch hardware-resource collisionInvoke debug-multi-launch-merge for the diagnostic decision tree
Output mismatch vs unmerged baselineOne of the per-kernel stitch operations corrupted SSA names, arg mapping, or layout boundaryBisect by un-merging the last-added kernel; the first un-merge that restores correctness identifies the offender
BO corruption after merge (NaN on 2nd+ call, stale values)static_input_indices / intermediate_indices not propagated to the merged ELF's cache.load_and_run()Invoke debug-bo-corruption
Bare-herd kernel produces all-zero in merged ELFRMSNorm or Eltwise at herd_x>1 emits bare air.herd; needs _wrap_ir_in_launch BEFORE stitchingWrap in Step 3 step 6
Merged compile slower than unmerged baselineCompile-time scaling with ELF size; not a correctness bugDocument in TODO; accept or drop the most-recent merge addition

Update protocol

Append to <model>/docs/development_progress/phase{4,5}_*.md:

## Multi-launch merge: <group_name>
- Sub-kernels merged: N
- Latency before: X ms
- Latency after: Y ms
- Gain: Z%
- Cosine vs unmerged: <value>
- max_abs / max_rel: <value> / <value>

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/opt-merge-multi-launch-kernels of Xilinx/mlir-air.

Open the folder on GitHubat commit 6e81ce1

Compare with similar skills

Opt Merge Multi Launch Kernels next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opt Merge Multi Launch Kernels compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opt Merge Multi Launch Kernels this skillXilinx/mlir-air150—~1.8kAutomated safety check: PassMIT
Mergeremotion-dev/remotion62k—~508Automated safety check: PassCustom licence
Mergewithastro/astro63k—~153Automated safety check: PassCustom licence
Shipping and Launch Checklistaddyosmani/agent-skills103k1 repos~2.8kAutomated safety check: PassMIT
Merge Upsymfony/symfony31k—~4kAutomated safety check: PassMIT
Mergealirezarezvani/claude-skills28k1 repos~587Automated safety check: PassMIT

Similar skills

  • Merge

    remotion-dev/remotion

    Official

    Wait for a Remotion pull request to become mergeable, handle merge conflicts, distinguish genuine CI failures from flakes, rerun flaky checks through the flake skill, and merge the PR.

    62k GitHub stars~508 tokensUpdated today
    DevelopmentAuto-check passed
  • Merge

    withastro/astro

    Official

    Handle main-to-next merge tasks including conflict resolution, changeset cleanup, and CI fix-ups.

    63k GitHub stars~153 tokensUpdated today
    Auto-check passed
  • Shipping and Launch Checklist

    addyosmani/agent-skills

    Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.

    103k GitHub starsUsed in 1 repo~2.8k tokens
    DevOps & CloudAuto-check passed
  • Merge Up

    symfony/symfony

    Cascade-merge maintained Symfony branches from oldest to newest (e.g.

    31k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Merge

    alirezarezvani/claude-skills

    Merge the winning agent's branch into base, archive losers, and clean up worktrees.

    28k GitHub starsUsed in 1 repo~587 tokens
    DevelopmentAuto-check passed
  • Ecc Recipes

    affaan-m/ECC

    Map a described workflow to the right ECC command group with run-order and stop condition, or browse all command-group recipe families read live from the commands directory.

    275k GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Opt Merge Multi Launch Kernels

What does Opt Merge Multi Launch Kernels do?

Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). Opt Merge Multi Launch Kernels is an agent skill from Xilinx/mlir-air.launch kernels into one multi-launch ELF (single XRT invocation).

How do I install Opt Merge Multi Launch Kernels in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a claude-code`. Or copy the skill folder (.claude/skills/opt-merge-multi-launch-kernels in Xilinx/mlir-air) into .claude/skills/opt-merge-multi-launch-kernels in your project. Claude Code loads it when a task matches its description.

How do I install Opt Merge Multi Launch Kernels in Codex?

Run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a codex`. Or copy the skill folder (.claude/skills/opt-merge-multi-launch-kernels in Xilinx/mlir-air) into .agents/skills/opt-merge-multi-launch-kernels in your project. Codex loads it when a task matches its description.

Can I use Opt Merge Multi Launch Kernels in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opt-merge-multi-launch-kernels, .gemini/skills/opt-merge-multi-launch-kernels, .github/skills/opt-merge-multi-launch-kernels and .opencode/skills/opt-merge-multi-launch-kernels in your project.

What does Opt Merge Multi Launch Kernels need to run?

SKILL.md names no scripts, command-line tools or credentials: Opt Merge Multi Launch Kernels is instructions for the agent only.

Does Opt Merge Multi Launch Kernels access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Opt Merge Multi Launch Kernels safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opt Merge Multi Launch Kernels use?

Opt Merge Multi Launch Kernels is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opt Merge Multi Launch Kernels use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opt Merge Multi Launch Kernels?

Skills that share tags, products or a category with Opt Merge Multi Launch Kernels: Merge (remotion-dev/remotion, 62k stars), Merge (withastro/astro, 63k stars), Shipping and Launch Checklist (addyosmani/agent-skills, 103k stars) and Merge Up (symfony/symfony, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opt Merge Multi Launch Kernels?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.