SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .claude/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .claude/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .agents/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .agents/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .cursor/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .cursor/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/phase-1-kernel-validation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .gemini/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .gemini/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .github/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .github/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .opencode/skills/phase-1-kernel-validation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phase-1-kernel-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-1-kernel-validation into .opencode/skills/phase-1-kernel-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-1-kernel-validation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phase-1-kernel-validationPhase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.
Phase 1 Kernel Validation is an agent skill from Xilinx/mlir-air. Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. Primary gate where a standalone harness exists: the harness's full-output element-wise np.isclose check at that kernel's rtol/atol vs an FP32 reference (the same PASS/FAIL make run prints). Fallback for a kernel with no harness: make diagnosis per-layer cosine vs the HF bf16 reference. Record each verified (kernel, shape) as a row in that kernel's…
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Deployment. It works with vLLM. The licence is MIT.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phase 1 Kernel Validation loads about 3.4k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 1,561 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,561 words, ~3,367 tokens.
.claude/skills/phase-1-kernel-validation/SKILL.md (or your agent's skills folder).Enumerate every (kernel, shape) the model needs and verify each is numerically correct on real NPU2 before integrating them into a block. Phase 1 isolates per-kernel correctness from integration bugs and extends the shared kernel registry with the shapes this model exercises, so Phase 2+ (and future deployments) can rely on a verified shape catalog.
The core loop is small:
kernel_registry/supported_kernels.md).mean_rel_L1kernel_registry/supported_kernels.md and its
details/<Kernel>_bf16.md (Used by = <model>); record the full
per-kernel results under <model>/docs/.Two sources of the per-kernel correctness verdict (use whichever exists):
make run that
compiles ONE kernel/block to an ELF and checks its full output
element-wise against an FP32 reference: np.isclose(|out−ref| ≤ atol + rtol·|ref|), every element must pass. The harness prints
[precision] mean_rel_L1=… rtol=… atol=… PASS!/failed. — that PASS/FAIL
IS the gate. This is the GPU/vLLM-aligned standard (rtol=1.6e-2 is
PyTorch/vLLM's canonical bf16 tolerance; atol is sized per kernel to the
measured datapath error). The registry's details/<Kernel>_bf16.md
§Tolerances documents each kernel's exact rtol/atol. Upstream ships a
harness for the FFN block (programming_examples/llms/llama_kernel_builder/ffn_swiglu/) plus
the top-level kernel examples matrix_multiplication/bf16_in_bf16_out
and matrix_multiplication/bf16_in_fp32_out (the BF16 GEMM, split by
output dtype — the legacy matrix_multiplication/bf16 is kept for NPU1),
matrix_vector_multiplication/bf16, flash_attention/kernel_fusion_based,
eltwise_add, weighted_rms_norm, rms_norm, rope_lut.make diagnosis: per-layer ffn_out cosine vs the HF
bf16 reference. The per-layer cosine reflects every kernel in that
layer. Cosine is the fallback lens only — not the gate for any kernel that
has a harness.Every (kernel, shape) the model needs must satisfy all four. Each catches a different bug class:
A correctness signal exists for the kernel × shape: either a
standalone harness make run (preferred — an isolated leaf check at the
registry's atol/rtol) OR the kernel is covered by make diagnosis
per-layer cosine. A kernel with neither is a silently-unverified
kernel — not allowed.
Numerical correctness (the real correctness gate — not theoretical compile-time rules):
np.isclose check PASSES at that kernel's rtol/atol
(details/<Kernel>_bf16.md §Tolerances) vs the FP32 reference. This is
the GPU/vLLM-aligned standard and the gate for every kernel that has a
harness.make diagnosis per-layer
cosine vs HF bf16 staying healthy (no cliff at the layer exercising
this kernel).
Catches: silent-corruption tile configs (e.g. GEMM
N % (tile_n × herd_n) != 0, which the builder does NOT assert). Trust
the measured atol/rtol verdict, not the rule.Record mean_rel_L1 (the headline metric) plus max_abs / max_rel
when the harness prints them. mean_rel_L1 is the registry's per-shape
accuracy column and a cheap regression baseline for future deployments;
the gate itself is the harness's pass/fail at rtol/atol.
Tile utilization documented: each (kernel, shape) records its herd config. Targets:
Registry row written: every verified (kernel, shape) that is new to
the registry is appended to that kernel's "tested shapes" table in
programming_examples/kernel_registry/supported_kernels.md and its
details/<Kernel>_bf16.md, with Used by = <model>, matching the
existing rows' column schema. The full per-kernel results
(mean_rel_L1 + harness PASS or diagnosis cosine, max_abs/max_rel,
perf, tile config) also go to <model>/docs/.
Catches: deployments that pass without leaving a reusable record.
Failure on ANY criterion blocks Phase 2.
PRIMARY (read before starting):
programming_examples/kernel_registry/supported_kernels.md
— index of every supported leaf kernel + its "tested shapes" table
(shape, tile config, perf, mean_rel_L1, Used by, status) across
deployments. This is the menu of known-good shapes to copy from and the
table you extend.programming_examples/kernel_registry/details/<Kernel>_bf16.md
— per-kernel detail: the numerical datapath, Tunable parameters
(knobs + hard constraints + tradeoffs), tolerances, per-shape data, and
the reproduce commands (which harness, how to run). Read the page for
each kernel your model needs. GEMM is the exception: it is split by
output dtype into details/GEMM_bf16_in_bf16_out.md and
details/GEMM_bf16_in_fp32_out.md (BF16-out has a --high-precision
tier — fused-cast / drain = FP32-accumulate + single cast, GPU-standard
~9.3e-3; F32-out always FP32-accumulates).programming_examples/kernel_registry/registry_lookup.py — the
machine-readable half of the registry. gemm_config(M,K,N, output_dtype, precision) returns the registry's best measured
{method, tile, gflops, mean_rel_L1} for a shape from the companion
details/*.json, and raises (no silent guess) for an unmeasured
shape — so Phase 1 recording a new GEMM shape is what unlocks the
programmatic lookup that Phase 4 builders consume. The .md tables
mirror these .json files for humans.Read the HF config.json (or the model's <model>_weights.py:Config
dataclass after Phase 0). Map each kernel call site to its shape using
standard transformer identities (Q proj output dim = n_heads * head_dim,
etc.). The llama-3.2-1B rows in kernel_registry/supported_kernels.md
(and programming_examples/llms/llama32_1b/'s prefill/decode call sites) are the worked
example of that call-site → shape mapping.
Write the working shape list — one row per (kernel, shape) with
mean_rel_L1 + verdict + perf + status columns empty (Step 2 fills
them) — under <model>/docs/
(e.g. <model>/docs/development_progress/phase1_kernels.md).
This is scratch tracking for the deployment, not a registry file; the
registry rows get written in Step 3.
Before running anything, read each needed kernel's details/<Kernel>_bf16.md
Tunable parameters / constraints section and check the model's dims
against it (alignment, max-K, placeability notes) — flag likely walls
BEFORE compile time.
(If the model has unusual ops — Q/K Norm, post-norm, ops between
currently-fused launches — flag in <model>/TODO.md as a Phase 2
prerequisite. The actual integration decision lives in
phase-2-single-block-validation Step 0a.)
For each (kernel, shape) in your Step 1 shape list:
a. Pick the correctness source + initial tile config.
Each kernel's details/<Kernel>_bf16.md has its reproduce commands (which
harness, how to run) and a Tunable parameters table (knobs + hard
constraints + tradeoffs). If your shape exists in that kernel's "tested
shapes" table in supported_kernels.md → reuse the tile config. Else →
mirror the nearest-shape entry; verify the hard constraints hold for your
shape; adjust if not (e.g., GEMM N % (tile_n × herd_n) != 0 → pick
smaller tile_n).
b. Run correctness (+ profile where a harness exists).
If a standalone harness or top-level example covers the shape:
cd programming_examples/<harness_or_example_dir>
flock -x -w 1800 /tmp/mlir-air-npu.lock make run # element-wise atol/rtol vs FP32 ref → PASS!/failed.
flock -x -w 1800 /tmp/mlir-air-npu.lock make profile # timingIf no standalone harness exists for the kernel, use the deployment's per-layer diagnosis:
cd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make diagnosis # per-layer ffn_out cosine vs HF bf16NPU is shared on this machine — every NPU command must be flock-wrapped
(see project memory). Compile-only steps don't need the lock.
If the harness reports failed. (or, for a no-harness kernel, the
diagnosis cosine drops) → see "Failure modes" below. Bound: 1 retry per
recipe per shape; if still failing, escalate to TODO.md "Active blockers".
c. Record results in <model>/docs/. Fill the scratch row:
mean_rel_L1 (+ harness PASS / diagnosis cosine), profile (ms / GFLOPS)
where measured, tile config used, tiles in flight, status, source (harness
vs diagnosis), and Notes (especially when tiles-in-flight is below
target — justify why).
For each (kernel, shape) verified in Step 2 that's new (not already
in that kernel's "tested shapes" table in supported_kernels.md), append
a row to both supported_kernels.md and the kernel's
details/<Kernel>_bf16.md per-shape table: shape + tile config +
tiles-in-flight, "Used by" listing your new model, mean_rel_L1 +
profile from Step 2, status. Match the existing rows' column schema for
that kernel exactly.
This grows the registry organically: each verified deployment extends the
menu of known-good shapes future deployments can copy from. (Adding a
new kernel — not just a new shape of an existing one — is the heavier
add-kernel workflow with its own parity checklist; Phase 1 only adds
shapes to kernels the registry already covers.)
When a test fails (harness failed. / diagnosis cosine drop, compile
error, hang), match the symptom to a likely cause. These are debug
starting points, not gates — trust the measured atol/rtol verdict, not the
rule. The relevant kernel's details/<Kernel>_bf16.md (constraints /
placeability section) is the authority for that kernel's hard limits.
| Symptom | Likely cause | Where to look |
|---|---|---|
'aiex.npu.push_queue' op Repeat count exceeds [0:255] | GEMV K too large; auto-split outer dim ≥ 256 | Set k_split so K = k_split × inner and k_split ≤ 255; see details/GEMV_bf16.md |
Allocator exhausted available buffer descriptor IDs | BD pool exhausted (non-aligned dim, or non-monotonic placeability of the iteration count) | Pad dim to 1024-aligned (GQA-aware reindexed padding) OR use kernel-first split-ELF path; see the placeability notes in the kernel's details/ page |
L2 capacity exceeded (matvec builder assert) | GEMV staged buffer > 512 KiB | Reduce tile_m (e.g., 8 → 2 for K=8192) or herd_m; see details/GEMV_bf16.md |
| Output all-zero / cosine = NaN | Bare-herd kernel without launch+segment wrapper | Wrap the bare herd in air.launch/air.segment (see ffn_swiglu harness for the multi-launch pattern) |
| Cosine = 0.02 or other small | GEMM N % (tile_n × herd_n) != 0 silent corruption (builder does NOT assert) | Pick tile_n so divisibility holds at this N |
| FA all-NaN at runtime | Compile-flag mismatch on attn_npu2.cc macros; OR seq-first dk_chunks>1 path at head_dim≥128 | see debug-fa-runtime-failure skill |
| Compile hangs > 10 min | Compiler scaling issue at large multi-launch | Cap and document; don't retry |
For any failure not in the table, invoke superpowers:systematic-debugging.
On Phase 1 PASS:
<model>/TODO.md, append "(N/N kernels PASSED)"<model>/docs/development_progress/progress.md
(per-kernel mean_rel_L1 + verdict, total time) — the per-model durable recordkernel_registry/supported_kernels.md + details/<Kernel>_bf16.md
carry a "tested shapes" row (Used by = <model>) for every new
(kernel, shape) this model exercises — the cross-model durable record© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/phase-1-kernel-validation of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Phase 1 Kernel Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phase 1 Kernel Validation this skillXilinx/mlir-air | 150 | — | ~3.4k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Aqua Troubleshootingoracle/accelerated-data-science | 125 | — | ~1.8k | Automated safety check: Pass | UPL-1.0 | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Aqua Deploymentoracle/accelerated-data-science | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | |
| Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
oracle/accelerated-data-science
Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
oracle/accelerated-data-science
Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Works with
Categories
Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. Phase 1 Kernel Validation is an agent skill from Xilinx/mlir-air. Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.
Phase 1 Kernel Validation fits situations like: tasks that involve LLM inference and serving; tasks that involve Deployment.
Run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-1-kernel-validation in Xilinx/mlir-air) into .claude/skills/phase-1-kernel-validation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a codex`. Or copy the skill folder (.claude/skills/phase-1-kernel-validation in Xilinx/mlir-air) into .agents/skills/phase-1-kernel-validation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-1-kernel-validation, .gemini/skills/phase-1-kernel-validation, .github/skills/phase-1-kernel-validation and .opencode/skills/phase-1-kernel-validation in your project.
Going by SKILL.md and its folder, Phase 1 Kernel Validation needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phase 1 Kernel Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phase 1 Kernel Validation: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Aqua Troubleshooting (oracle/accelerated-data-science, 125 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Aqua Deployment (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.