Kernel Organization
sgl-project/sglang
Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .claude/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .claude/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignmentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .agents/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .agents/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .cursor/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .cursor/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/opt-layout-alignment--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .gemini/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .gemini/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air opt-layout-alignmentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .github/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .github/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .opencode/skills/opt-layout-alignment && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "opt-layout-alignment" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/opt-layout-alignment into .opencode/skills/opt-layout-alignment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opt-layout-alignment", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
opt-layout-alignmentOptimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Opt Layout Alignment is an agent skill from Xilinx/mlir-air. Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. Canonical case: seq-first (seq, nheads·headdim) so RMSNorm → RoPE → FlashAttention → O-proj stay seq-first, eliminating 1–4 host transposes per layer. Invoked by phase-4-prefill-optimization (and phase-5 when decode introduces a transpose phase-4 didn't fix).
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Opt Layout Alignment loads about 1k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 439 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 439 words, ~1,013 tokens.
.claude/skills/opt-layout-alignment/SKILL.md (or your agent's skills folder).When two consecutive kernels disagree on activation layout, the host inserts a
transpose between them — a data round-trip that adds up across many per-layer
calls. This skill removes those transposes by choosing layouts that let
consecutive kernels hand off directly on-device. The canonical alignment is
seq-first activations (seq, n_heads·head_dim), which keeps RoPE,
FlashAttention, and the O projection on the same layout the GEMMs/RMSNorm
already produce — no host transpose between them.
Most inheritance deployments already run seq-first end-to-end (nothing to do — skip). This skill applies when the deployment still has a host transpose between two kernels.
Applying this skill is "successful" when ALL hold:
max_abs / max_rel informational.make verify still PASSES (end-to-end gate).If (1)/(2) regress → a kernel did not actually accept the new layout; revert.
programming_examples/flash_attention/kernel_fusion_based/attn_npu2_seqfirst.py
— the seq-first FlashAttention variant (head_dim ≤ 64).programming_examples/flash_attention/kernel_fusion_based/ — the
head-first kernel + wrapper used for head_dim ≥ 128 (see caveat below)..claude/skills/debug-fa-runtime-failure — owns the why of the
head_dim ≥ 128 routing.Profile / read the per-layer host code. Each np.transpose /
ascontiguousarray between two NPU kernel calls is a candidate. Note which
kernel boundary it bridges (typically RoPE→FA or FA→O-proj).
attn_npu2_seqfirst.py
for head_dim ≤ 64.The seq-first dk_chunks > 1 path has known runtime issues at head_dim ≥ 128.
Route head_dim ≥ 128 attention through the head-first wrapper (it does the
host transpose precisely so the rest of the pipeline stays seq-first). Do NOT
debug FA inline here — for the why and the discrimination of the failure
modes, invoke debug-fa-runtime-failure.
make verify → must still PASS.| Symptom | Likely cause | Where to look |
|---|---|---|
| Cosine drops after switching layout | a kernel didn't actually consume seq-first; the transpose was masking a real layout mismatch | Revert; confirm each kernel's accepted layout before deleting the transpose |
FA hang (ERT_CMD_STATE_TIMEOUT) or NaN at head_dim ≥ 128 | seq-first dk_chunks > 1 path bug | Route through the head-first wrapper; invoke debug-fa-runtime-failure |
| Transpose removed but no wall-time gain | the transpose wasn't on the hot path | Document; revert or keep for cleanliness |
For any failure not in the table, invoke superpowers:systematic-debugging.
Append to <model>/docs/development_progress/phase{4,5}_*.md:
## Layout alignment
- Transposes removed: <which boundaries>
- head_dim ≥ 128 routed head-first: yes/no/N-A
- Wall time before: X ms
- Wall time after: Y ms
- Cosine vs baseline: <value> | make verify: PASS/FAIL© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/opt-layout-alignment of Xilinx/mlir-air.
Open the folder on GitHubat commit bca27e5
Opt Layout Alignment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Opt Layout Alignment this skillXilinx/mlir-air | 150 | — | ~1k | Automated safety check: Pass | MIT | |
| Kernel Organizationsgl-project/sglang | 37k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence | |
| Ccs Alignthedotmack/claude-mem | 97k | — | ~6.1k | Automated safety check: Pass | Apache-2.0 | |
| Tabler Page Layoutstabler/tabler | 42k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Bio Read Alignment Bowtie2 AlignmentGPTomics/bioSkills | 1.2k | 1 repos | ~3.6k | Automated safety check: Pass | MIT |
sgl-project/sglang
Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
thedotmack/claude-mem
Run the CCS Align seat's hourly breathing cycle — prove the local claude-mem worker is healthy, pull needle observations through search → timeline → getobservations, land them in a seat-owned middle…
tabler/tabler
Picks and configures the right layout for a Tabler preview page and covers changing or adding layouts in shared/layouts, including DefaultLayout props and the page-header slot.
GPTomics/bioSkills
Aligns DNA short reads to a reference with Bowtie2, choosing end-to-end (whole read must align) vs local (soft-clip read ends) mode and a sensitivity preset; the de-facto aligner for ChIP-seq…
sickn33/agentic-awesome-skills
CRITICAL: Use for Makepad layout system. An agent skill from sickn33/agentic-awesome-skills.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. Opt Layout Alignment is an agent skill from Xilinx/mlir-air. Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a claude-code`. Or copy the skill folder (.claude/skills/opt-layout-alignment in Xilinx/mlir-air) into .claude/skills/opt-layout-alignment in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a codex`. Or copy the skill folder (.claude/skills/opt-layout-alignment in Xilinx/mlir-air) into .agents/skills/opt-layout-alignment in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opt-layout-alignment, .gemini/skills/opt-layout-alignment, .github/skills/opt-layout-alignment and .opencode/skills/opt-layout-alignment in your project.
Going by SKILL.md and its folder, Opt Layout Alignment needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Opt Layout Alignment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Opt Layout Alignment: Kernel Organization (sgl-project/sglang, 37k stars), Metal Kernel (pytorch/pytorch, 104k stars), Ccs Align (thedotmack/claude-mem, 97k stars) and Tabler Page Layouts (tabler/tabler, 42k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.