Agent skill

Phase 4 Prefill Optimization

by Xilinx in Xilinx/mlir-air

Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout).

MITAuto-check passedDevOps & Cloud

Install Phase 4 Prefill Optimization

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-4-prefill-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-4-prefill-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-4-prefill-optimization .claude/skills/phase-4-prefill-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-4-prefill-optimization
GitHub stars
150
Token cost
~2.3k tokens
SKILL.md length
1,046 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout).

  • Tasks that involve Deployment
  • SKILL.md covers Purpose, Phase 4 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls make

What it does

Phase 4 Prefill Optimization is an agent skill from Xilinx/mlir-air. Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout). Thin orchestrator that dispatches opt-merge-multi-launch-kernels, opt-buffer-object-reuse, and opt-layout-alignment. Each step preserves correctness by re-running the Phase 3 gate — make verify (token-set vs HF bf16) is the PASS/FAIL gate; make diagnosis per-layer cosine is the informational lens used to localize a regression…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. The licence is MIT.

When your agent uses it

  • Tasks that involve Deployment

Example prompts

  • “/phase-4-prefill-optimization”

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 4 Prefill Optimization loads about 2.3k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 1,046 words, ~2,295 tokens.

Download SKILL.mdSave it as .claude/skills/phase-4-prefill-optimization/SKILL.md (or your agent's skills folder).
name
phase-4-prefill-optimization
description
Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout). Thin orchestrator that dispatches `opt-merge-multi-launch-kernels`, `opt-buffer-object-reuse`, and `opt-layout-alignment`. Each step preserves correctness by re-running the Phase 3 gate — `make verify` (token-set vs HF bf16) is the PASS/FAIL gate; `make diagnosis` per-layer cosine is the informational lens used to localize a regression. Invoked after Phase 3 PASS.

Purpose

Phase 1-3 produced a functionally correct NPU pipeline (kernel-by-kernel in Phase 1, layer-by-layer in Phase 2-3, all numerically aligned with the HF bf16 reference). Phase 4 keeps that correctness while applying the shared optimization skillset to reduce prefill latency. This phase is a thin orchestrator: it dispatches three now-independent optimization skills — opt-merge-multi-launch-kernels (ELF-merging), opt-buffer-object-reuse (host↔NPU runtime-overhead reduction), and opt-layout-alignment (host-side layout) — each of which owns its own recipe and failure modes. These wins are things you compose from the kernels, not behaviors inherited from a reference; the reference's builders are worked examples. Every optimization is an experiment: apply, re-measure, re-run Phase 3 gate; revert if correctness regresses.

For scale, the reference deployment llama3.2-1B took prefill from 18.67 s → 1.30 s (14×) by composing these optimizations — an illustrative datapoint for what they buy on one model, not a target every deployment must hit.

Phase 4 PASS criteria (HARD GATES)

  1. Correctness preserved: after every applied optimization, make verify (the token-set gate vs HF bf16) still PASSES — this is the Phase 3 correctness gate, re-run between optimization skills. make diagnosis per-layer cosine is NOT a gate (the verify subsystem retired threshold-based diagnosis; compare_pair reports cosine with no pass/fail); run it only to localize a regression when verify breaks (which optimization / which layer the cosine cliff appears at). If make verify regresses to FAIL, revert the change and document why it doesn't apply to this model.
  2. Prefill kernel time strictly < Phase 3 baseline, measured at the same canonical prompt + seq_len with 5-warmup + 20-iter profile. Recording wall time too (kernel + host overhead) is good but the gate is on kernel time (host overhead optimization is a separate concern).
  3. Per-skill outcome documented in <model>/docs/development_progress/phase4_prefill.md: for each optimization skill it invoked, record applied / skipped / reverted, the latency delta, and a one-line reason.

The "≥ N optimization skills applied" check is NOT a gate — some models legitimately need only 1-2 (e.g., the model is already seq-first by construction → opt-layout-alignment N/A). The gate is the outcome (perf improved + correctness preserved), not the process count.

Knowledge base references

PRIMARY:

  • programming_examples/llms/llama32_1b/docs/profile.md — the reference deployment's profiling breakdown; the reference for what "good" looks like
  • programming_examples/llms/<model>/docs/development_progress/phase3_full.md — Phase 3 baseline timings + cosine numbers (the "before" state to measure against and preserve)
  • programming_examples/llms/llama_kernel_builder/cache.py — the KernelCache host-optimization knobs (static_input_indices, intermediate_indices); the opt-buffer-object-reuse skill owns how to wire them, this is just the source file it touches

REFERENCE EXEMPLARS (read/mirror to compose your own fused ELFs; import directly only on a bit-for-bit kernel-sequence match):

  • programming_examples/llms/llama32_1b/multi_launch_builder/ — the full set of fused-ELF builders, the worked example of how registry leaf kernels stitch into multi-launch ELFs. Mirror these for your model's kernel sequence. Two representative ones:
    • rms_gemms_rope_multi.py — fused 6-launch ELF for RMSNorm + Q/K/V GEMM + RoPE Q/K
    • o_ffn_multi.py — fused 8-launch ELF for O + add + RMSNorm + Gate/Up + SwiGLU + Down + add
  • programming_examples/llms/llama_kernel_builder/ — the shared toolkit (KernelCache, stitching, external_kernels) every fused-ELF build uses, inheritance or kernel-first alike.

Workflow

Step 1: Measure Phase 3 baseline

Before invoking any optimization skill, capture the baseline prefill time — this is the number every skill must beat:

bash
cd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make profile

Record: kernel time (ms) + wall time as the baseline. Also note the Phase 3 per-layer cosine table as a baseline to compare against if a later optimization breaks verify (it's the localization reference, not a gate).

Show full SKILL.md (509 more words)Show less
Step 2: Apply optimization skills

prefill draws on the shared optimization skillset. For each, decide if it applies, invoke the skill, then re-run the gate (Step 3). Skip with a logged reason if the trigger condition isn't met — "≥ N applied" is NOT the gate; the gate is the outcome (faster + make verify still PASSES).

Optimization skillWhen it applies to prefillWhat it does
opt-merge-multi-launch-kernelsalmost always (the dominant win)stitch each leaf kernel's air.launch into one fused ELF per kernel-group → one xrt.run() per group instead of per kernel (llama3: 16→3 calls/layer). Build the model's multi_launch_builder/ (kernel-first) or reuse llama's fused ELFs (bit-for-bit inheritance — the verdict made in phase-2-single-block-validation Step 1).
opt-buffer-object-reusealwayspre-load per-layer weight BOs once (static_input_indices) + reuse intermediate BOs (intermediate_indices); removes redundant host↔NPU uploads.
opt-layout-alignmentonly if a host transpose still sits between two kernelschoose seq-first layouts so RoPE/FA/O-proj hand off on-device; skip if the model already runs seq-first end-to-end (most inheritance deployments do).

Each skill owns its own recipe, success self-check, and failure modes — this phase does not restate them. Invoke the skill, read its result, then gate.

Step 3: Re-run Phase 3 gate after each optimization skill

After every applied (or attempted) optimization skill, re-run the gate:

bash
flock -x -w 1800 /tmp/mlir-air-npu.lock make verify      # GATE: token-set, exit 1 on FAIL

make verify PASS is the correctness gate. If it regresses to FAIL, revert the change and document why.

Only when verify FAILs, run diagnosis to localize the break:

bash
flock -x -w 1800 /tmp/mlir-air-npu.lock make diagnosis   # informational: per-layer cosine table

A cosine cliff at layer i points at the broken assumption there (e.g. a transpose opt-layout-alignment removed that this model actually needed). diagnosis does not PASS/FAIL — it is the microscope, verify is the gate.

Failure modes

SymptomLikely causeWhere to look
Multi-launch merge compile fails (BD exhaustion, channel routing, herd shape conflict, bare-herd, DMA stride, IR/compile blowup)non-1024-aligned dim (see the kernel's details/<Kernel>_bf16.md) OR wrong stitching boundaryInvoke debug-multi-launch-merge — it discriminates the 6 known compile blockers
Output corruption after BO pre-loading (correct first call, NaN/garbage on subsequent calls)Per-layer BO key collision OR static_input_indices set wrongInvoke debug-bo-corruption
FA hang (ERT_CMD_STATE_TIMEOUT) at head_dim ≥ 128Seq-first dk_chunks > 1 path bugInvoke debug-fa-runtime-failure; the head-first wrapper (routed by opt-layout-alignment) is the workaround
FA all-NaN at runtimeCompile-flag mismatch on attn_npu2.cc macros (LESSON 3 — -Dlqp must be per-tile, not per-launch)Invoke debug-fa-runtime-failure; compile_attn_npu2_split derives correct flags
Cosine drops after an optimization skillthe skill has a layout/type assumption your model violatesRevert the change; check whether the assumption (e.g., seq-first only, all weights pre-transposed) holds
Latency unchanged after opt-merge-multi-launch-kernelsMulti-launch ELF compiled but XRT call count didn't dropCheck xrt-smi top for actual call count; verify the new fused ELF is what _run_cached actually invokes (not falling back to per-kernel path)

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

On Phase 4 PASS:

  • <model>/docs/development_progress/phase4_prefill.md: per-skill table with applied / skipped / reverted, latency delta, reason
  • <model>/TODO.md: mark Phase 4, append final prefill kernel time + speedup vs Phase 3 baseline
  • If a new fused ELF was built (kernel-first path of opt-merge-multi-launch-kernels), surface to Phase 6 for potential promotion to a shared location if a second deployment validates the same pattern

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-4-prefill-optimization of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Phase 4 Prefill Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 4 Prefill Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 4 Prefill Optimization this skillXilinx/mlir-air150—~2.3kAutomated safety check: PassMIT
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb6.7k—~4kAutomated safety check: NotesApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Vercelremotion-dev/remotion62k—~1.2kAutomated safety check: PassCustom licence
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated 7 days ago
    DevOps & CloudAuto-check: notes
  • GreptimeDB Dev Docker Image

    GreptimeTeam/greptimedb

    Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.

    6.7k GitHub stars~4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check: notes
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Vercel

    remotion-dev/remotion

    Official

    Set up a Codex monitor for Vercel deployments and preview URLs.

    62k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    259 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Categories

Questions about Phase 4 Prefill Optimization

What does Phase 4 Prefill Optimization do?

Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout). Phase 4 Prefill Optimization is an agent skill from Xilinx/mlir-air. Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout).

When should I use Phase 4 Prefill Optimization?

Phase 4 Prefill Optimization fits situations like: tasks that involve Deployment.

How do I install Phase 4 Prefill Optimization in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-4-prefill-optimization -a claude-code`. Or copy the skill folder (.claude/skills/phase-4-prefill-optimization in Xilinx/mlir-air) into .claude/skills/phase-4-prefill-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Phase 4 Prefill Optimization in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-4-prefill-optimization -a codex`. Or copy the skill folder (.claude/skills/phase-4-prefill-optimization in Xilinx/mlir-air) into .agents/skills/phase-4-prefill-optimization in your project. Codex loads it when a task matches its description.

Can I use Phase 4 Prefill Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-4-prefill-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-4-prefill-optimization, .gemini/skills/phase-4-prefill-optimization, .github/skills/phase-4-prefill-optimization and .opencode/skills/phase-4-prefill-optimization in your project.

What does Phase 4 Prefill Optimization need to run?

Going by SKILL.md and its folder, Phase 4 Prefill Optimization needs the command-line tools its instructions call (make).

Does Phase 4 Prefill Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phase 4 Prefill Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 4 Prefill Optimization use?

Phase 4 Prefill Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 4 Prefill Optimization use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 4 Prefill Optimization?

Skills that share tags, products or a category with Phase 4 Prefill Optimization: Kubeshark Installer (kubeshark/kubeshark, 12k stars), GreptimeDB Dev Docker Image (GreptimeTeam/greptimedb, 6.7k stars), KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars) and Vercel (remotion-dev/remotion, 62k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 4 Prefill Optimization?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.