Merge
remotion-dev/remotion
Wait for a Remotion pull request to become mergeable, handle merge conflicts, distinguish genuine CI failures from flakes, rerun flaky checks through the flake skill, and merge the PR.
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .claude/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .claude/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernelsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .agents/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .agents/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .cursor/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .cursor/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/opt-merge-multi-launch-kernels--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .gemini/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .gemini/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernelsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .github/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .github/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air opt-merge-multi-launch-kernels --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/opt-merge-multi-launch-kernels .opencode/skills/opt-merge-multi-launch-kernels && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "opt-merge-multi-launch-kernels" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-merge-multi-launch-kernels into .opencode/skills/opt-merge-multi-launch-kernels/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-merge-multi-launch-kernels", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
opt-merge-multi-launch-kernelsProcedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
Opt Merge Multi Launch Kernels is an agent skill from Xilinx/mlir-air. Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). Invoked by phase-4-prefill-optimization and phase-5-decode-optimization to fuse kernel groups when building NEW model-specific fused ELFs (kernel-first path). Reduces XRT dispatch overhead (~50–200 µs per call on NPU2).
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Opt Merge Multi Launch Kernels loads about 1.8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 795 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 795 words, ~1,819 tokens.
.claude/skills/opt-merge-multi-launch-kernels/SKILL.md (or your agent's skills folder).When a deployment goes the kernel-first path in Phase 4/5 (model
has new ops or stitching boundaries differ from llama), build new
model-specific fused ELFs by stitching multiple air.launch kernels
into a single func.func. This skill is the procedural recipe.
The reference llama3 deployment went 10 launches/layer → 3 launches/layer using this recipe — the dominant prefill perf win.
A merge is "successful" when ALL hold:
rtol=1e-3 is too tight for
K=2048 BF16; use cosine + max_abs/max_rel informational logging
(see phase-1-kernel-validation PASS criteria for the convention).If (1) fails → invoke debug-multi-launch-merge.
If (2) fails → bisect by un-merging the last-added kernel.
If (3) fails (compile slower than unmerged) → not a correctness bug;
document and accept or revert.
programming_examples/llms/llama_kernel_builder/stitching.py —
the helpers this recipe uses (_rename_all, _fix_launch_func_args,
_wrap_ir_in_launch, _rename_all_with_externs, _rename_all_gemv)programming_examples/llms/llama32_1b/multi_launch_builder/rms_gemms_rope_multi.py
— reference 6-launch prefill merge (RMSNorm + Q/K/V GEMM + RoPE Q/K)programming_examples/llms/llama32_1b/multi_launch_builder/o_ffn_multi.py
— reference 8-launch prefill merge (O + add + RMSNorm + Gate/Up + SwiGLU + Down + add)programming_examples/llms/llama32_1b/multi_launch_builder/o_gemv_ffn_multi.py
— decode merge with 2-K extern rename (the pattern to extend to 3-K
when n_heads*head_dim != emb_dim)programming_examples/kernel_registry/supported_kernels.md
— per-kernel constraints; the FA section notes FlashAttention stays a
separate XRT call (does NOT merge into the fused block)List the per-layer kernel sequence. Mark each as one of:
full_block.md)Group consecutive mergeable launches between hard-stops. For a typical decoder-only LLM:
rms_gemms_rope_multi.py shape)o_ffn_multi.py shape)For decode, replace GEMM with GEMV throughout.
For each group, create <model>/multi_launch_builder/<group_name>_multi.py
that:
@module_builderfrom llama_kernel_builder.stitching import (_rename_all, _fix_launch_func_args, _wrap_ir_in_launch, ...) (resolved against the shared programming_examples/llms/llama_kernel_builder/ via sys.path)_extract_between_func_and_return(ir_text)_rename_all(body, prefix=...) to avoid collisions across the merged module_fix_launch_func_args(body, prefix, arg_map)func.func. If any
sub-kernel emits a bare air.herd (RMSNorm herd_x>1, Eltwise
Add at herd_x>1), wrap each via _wrap_ir_in_launch(...) BEFORE
stitching — otherwise the lowering's airrt-to-npu pass drops the
bare herd_rename_all_with_externs with per-launch extern allowlists to keep
different mv_*.o symbols distinct (extend the 2-K rename in
o_gemv_ffn_multi.py to 3-K when n_heads*head_dim != emb_dim)The canonical reference is rms_gemms_rope_multi.py — copy from there.
Use KernelCache.compile_and_cache(name=<group_name>, builder=<your_builder_function>, ...).
If compile fails → invoke debug-multi-launch-merge.
Run the merged ELF on a fixed input. Run the same kernels as separate XRT calls (the unmerged baseline). Compare the merged output to the unmerged output:
max_abs / max_rel informationalIf cosine fails: bisect by un-merging the last-added kernel from Step 3. If still fails after un-merging that one, the bug is in an EARLIER merge addition — bisect further.
Profile the merged version vs unmerged. Expected at NPU2 dispatch
overhead levels: ≥ 20% reduction per merged group at moderate scale
(llama3 saw 16-XRT-call/layer prefill → 3-XRT-call/layer = much
larger reduction). If the gain is much smaller than expected, the
per-call XRT overhead may not have been the bottleneck — the
opt-buffer-object-reuse skill (static weight BOs) is often the missing piece
in that case.
Record gain in <model>/docs/development_progress/phase{4,5}_*.md.
| Symptom | Likely cause | Where to look |
|---|---|---|
| Compile fails (BD exhaustion, channel routing, herd shape conflict, IR validation) | Multi-launch hardware-resource collision | Invoke debug-multi-launch-merge for the diagnostic decision tree |
| Output mismatch vs unmerged baseline | One of the per-kernel stitch operations corrupted SSA names, arg mapping, or layout boundary | Bisect by un-merging the last-added kernel; the first un-merge that restores correctness identifies the offender |
| BO corruption after merge (NaN on 2nd+ call, stale values) | static_input_indices / intermediate_indices not propagated to the merged ELF's cache.load_and_run() | Invoke debug-bo-corruption |
| Bare-herd kernel produces all-zero in merged ELF | RMSNorm or Eltwise at herd_x>1 emits bare air.herd; needs _wrap_ir_in_launch BEFORE stitching | Wrap in Step 3 step 6 |
| Merged compile slower than unmerged baseline | Compile-time scaling with ELF size; not a correctness bug | Document in TODO; accept or drop the most-recent merge addition |
Append to <model>/docs/development_progress/phase{4,5}_*.md:
## Multi-launch merge: <group_name>
- Sub-kernels merged: N
- Latency before: X ms
- Latency after: Y ms
- Gain: Z%
- Cosine vs unmerged: <value>
- max_abs / max_rel: <value> / <value>© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/opt-merge-multi-launch-kernels of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Opt Merge Multi Launch Kernels next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Opt Merge Multi Launch Kernels this skillXilinx/mlir-air | 150 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Mergeremotion-dev/remotion | 62k | — | ~508 | Automated safety check: Pass | Custom licence | |
| Mergewithastro/astro | 63k | — | ~153 | Automated safety check: Pass | Custom licence | |
| Shipping and Launch Checklistaddyosmani/agent-skills | 103k | 1 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Merge Upsymfony/symfony | 31k | — | ~4k | Automated safety check: Pass | MIT | |
| Mergealirezarezvani/claude-skills | 28k | 1 repos | ~587 | Automated safety check: Pass | MIT |
remotion-dev/remotion
Wait for a Remotion pull request to become mergeable, handle merge conflicts, distinguish genuine CI failures from flakes, rerun flaky checks through the flake skill, and merge the PR.
withastro/astro
Handle main-to-next merge tasks including conflict resolution, changeset cleanup, and CI fix-ups.
addyosmani/agent-skills
Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.
symfony/symfony
Cascade-merge maintained Symfony branches from oldest to newest (e.g.
alirezarezvani/claude-skills
Merge the winning agent's branch into base, archive losers, and clean up worktrees.
affaan-m/ECC
Map a described workflow to the right ECC command group with run-order and stop condition, or browse all command-group recipe families read live from the commands directory.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). Opt Merge Multi Launch Kernels is an agent skill from Xilinx/mlir-air.launch kernels into one multi-launch ELF (single XRT invocation).
Run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a claude-code`. Or copy the skill folder (.claude/skills/opt-merge-multi-launch-kernels in Xilinx/mlir-air) into .claude/skills/opt-merge-multi-launch-kernels in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a codex`. Or copy the skill folder (.claude/skills/opt-merge-multi-launch-kernels in Xilinx/mlir-air) into .agents/skills/opt-merge-multi-launch-kernels in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill opt-merge-multi-launch-kernels -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opt-merge-multi-launch-kernels, .gemini/skills/opt-merge-multi-launch-kernels, .github/skills/opt-merge-multi-launch-kernels and .opencode/skills/opt-merge-multi-launch-kernels in your project.
SKILL.md names no scripts, command-line tools or credentials: Opt Merge Multi Launch Kernels is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Opt Merge Multi Launch Kernels is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Opt Merge Multi Launch Kernels: Merge (remotion-dev/remotion, 62k stars), Merge (withastro/astro, 63k stars), Shipping and Launch Checklist (addyosmani/agent-skills, 103k stars) and Merge Up (symfony/symfony, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.