Agent skill

Opt Layout Alignment

by Xilinx in Xilinx/mlir-air

Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

MITAuto-check passed

Install Opt Layout Alignment

skills CLI
$ npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air opt-layout-alignment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/opt-layout-alignment .claude/skills/opt-layout-alignment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opt-layout-alignment
GitHub stars
150
Token cost
~1k tokens
SKILL.md length
439 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

  • Works in 4 steps: Find the host transposes → Make the kernels accept seq-first → head_dim ≥ 128 caveat → …
  • SKILL.md covers Purpose, Success criteria, Knowledge base references and Workflow, plus 2 more sections
  • Calls make

What it does

Opt Layout Alignment is an agent skill from Xilinx/mlir-air. Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. Canonical case: seq-first (seq, nheads·headdim) so RMSNorm → RoPE → FlashAttention → O-proj stay seq-first, eliminating 1–4 host transposes per layer. Invoked by phase-4-prefill-optimization (and phase-5 when decode introduces a transpose phase-4 didn't fix).

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

Example prompts

  • “/opt-layout-alignment”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Find the host transposes
  2. Make the kernels accept seq-first
  3. head_dim ≥ 128 caveat
  4. Validate + measure

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opt Layout Alignment loads about 1k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 439 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 439 words, ~1,013 tokens.

Download SKILL.mdSave it as .claude/skills/opt-layout-alignment/SKILL.md (or your agent's skills folder).
name
opt-layout-alignment
description
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. Canonical case: seq-first (seq, n_heads·head_dim) so RMSNorm → RoPE → FlashAttention → O-proj stay seq-first, eliminating 1–4 host transposes per layer. Invoked by phase-4-prefill-optimization (and phase-5 when decode introduces a transpose phase-4 didn't fix).

Purpose

When two consecutive kernels disagree on activation layout, the host inserts a transpose between them — a data round-trip that adds up across many per-layer calls. This skill removes those transposes by choosing layouts that let consecutive kernels hand off directly on-device. The canonical alignment is seq-first activations (seq, n_heads·head_dim), which keeps RoPE, FlashAttention, and the O projection on the same layout the GEMMs/RMSNorm already produce — no host transpose between them.

Most inheritance deployments already run seq-first end-to-end (nothing to do — skip). This skill applies when the deployment still has a host transpose between two kernels.

Success criteria

Applying this skill is "successful" when ALL hold:

  1. Output cosine ≥ 0.99 vs the pre-alignment baseline (changing layout must not change the math). Log max_abs / max_rel informational.
  2. make verify still PASSES (end-to-end gate).
  3. The targeted host transpose(s) are gone — fewer host ops, lower wall time.

If (1)/(2) regress → a kernel did not actually accept the new layout; revert.

Knowledge base references

  • programming_examples/flash_attention/kernel_fusion_based/attn_npu2_seqfirst.py — the seq-first FlashAttention variant (head_dim ≤ 64).
  • programming_examples/flash_attention/kernel_fusion_based/ — the head-first kernel + wrapper used for head_dim ≥ 128 (see caveat below).
  • .claude/skills/debug-fa-runtime-failure — owns the why of the head_dim ≥ 128 routing.

Workflow

Step 1: Find the host transposes

Profile / read the per-layer host code. Each np.transpose / ascontiguousarray between two NPU kernel calls is a candidate. Note which kernel boundary it bridges (typically RoPE→FA or FA→O-proj).

Step 2: Make the kernels accept seq-first
  • RoPE: accept seq-first input.
  • FlashAttention: accept seq-first Q, K, V — this is attn_npu2_seqfirst.py for head_dim ≤ 64.
  • Verify the producer kernel emits the layout the consumer expects, so the transpose can be deleted (not just moved).
Show full SKILL.md (174 more words)Show less
Step 3: head_dim ≥ 128 caveat

The seq-first dk_chunks > 1 path has known runtime issues at head_dim ≥ 128. Route head_dim ≥ 128 attention through the head-first wrapper (it does the host transpose precisely so the rest of the pipeline stays seq-first). Do NOT debug FA inline here — for the why and the discrimination of the failure modes, invoke debug-fa-runtime-failure.

Step 4: Validate + measure
  • Compare output to the pre-alignment baseline → cosine ≥ 0.99.
  • Re-run make verify → must still PASS.
  • Confirm the transpose is gone and wall time dropped.

Failure modes

SymptomLikely causeWhere to look
Cosine drops after switching layouta kernel didn't actually consume seq-first; the transpose was masking a real layout mismatchRevert; confirm each kernel's accepted layout before deleting the transpose
FA hang (ERT_CMD_STATE_TIMEOUT) or NaN at head_dim ≥ 128seq-first dk_chunks > 1 path bugRoute through the head-first wrapper; invoke debug-fa-runtime-failure
Transpose removed but no wall-time gainthe transpose wasn't on the hot pathDocument; revert or keep for cleanliness

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

Append to <model>/docs/development_progress/phase{4,5}_*.md:

## Layout alignment
- Transposes removed: <which boundaries>
- head_dim ≥ 128 routed head-first: yes/no/N-A
- Wall time before: X ms
- Wall time after:  Y ms
- Cosine vs baseline: <value>  | make verify: PASS/FAIL

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/opt-layout-alignment of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Opt Layout Alignment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opt Layout Alignment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opt Layout Alignment this skillXilinx/mlir-air150—~1kAutomated safety check: PassMIT
Kernel Organizationsgl-project/sglang37k—~1.3kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence
Ccs Alignthedotmack/claude-mem97k—~6.1kAutomated safety check: PassApache-2.0
Tabler Page Layoutstabler/tabler42k—~1.5kAutomated safety check: PassMIT
Bio Read Alignment Bowtie2 AlignmentGPTomics/bioSkills1.2k1 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Kernel Organization

    sgl-project/sglang

    Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.

    37k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ccs Align

    thedotmack/claude-mem

    Run the CCS Align seat's hourly breathing cycle — prove the local claude-mem worker is healthy, pull needle observations through search → timeline → getobservations, land them in a seat-owned middle…

    97k GitHub stars~6.1k tokensUpdated today
    Auto-check passed
  • Tabler Page Layouts

    tabler/tabler

    Picks and configures the right layout for a Tabler preview page and covers changing or adding layouts in shared/layouts, including DefaultLayout props and the page-header slot.

    42k GitHub stars~1.5k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Aligns DNA short reads to a reference with Bowtie2, choosing end-to-end (whole read must align) vs local (soft-clip read ends) mode and a sensitivity preset; the de-facto aligner for ChIP-seq…

    1.2k GitHub starsUsed in 1 repo~3.6k tokens
    Research & ScienceAuto-check passed
  • Makepad Layout

    sickn33/agentic-awesome-skills

    CRITICAL: Use for Makepad layout system. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Opt Layout Alignment

What does Opt Layout Alignment do?

Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. Opt Layout Alignment is an agent skill from Xilinx/mlir-air. Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

How do I install Opt Layout Alignment in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a claude-code`. Or copy the skill folder (.claude/skills/opt-layout-alignment in Xilinx/mlir-air) into .claude/skills/opt-layout-alignment in your project. Claude Code loads it when a task matches its description.

How do I install Opt Layout Alignment in Codex?

Run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a codex`. Or copy the skill folder (.claude/skills/opt-layout-alignment in Xilinx/mlir-air) into .agents/skills/opt-layout-alignment in your project. Codex loads it when a task matches its description.

Can I use Opt Layout Alignment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill opt-layout-alignment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opt-layout-alignment, .gemini/skills/opt-layout-alignment, .github/skills/opt-layout-alignment and .opencode/skills/opt-layout-alignment in your project.

What does Opt Layout Alignment need to run?

Going by SKILL.md and its folder, Opt Layout Alignment needs the command-line tools its instructions call (make).

Does Opt Layout Alignment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Opt Layout Alignment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opt Layout Alignment use?

Opt Layout Alignment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opt Layout Alignment use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opt Layout Alignment?

Skills that share tags, products or a category with Opt Layout Alignment: Kernel Organization (sgl-project/sglang, 37k stars), Metal Kernel (pytorch/pytorch, 104k stars), Ccs Align (thedotmack/claude-mem, 97k stars) and Tabler Page Layouts (tabler/tabler, 42k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opt Layout Alignment?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.