Kubeshark Installer
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .claude/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .claude/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .agents/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .agents/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .cursor/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .cursor/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/phase-2-single-block-validation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .gemini/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .gemini/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .github/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .github/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .opencode/skills/phase-2-single-block-validation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phase-2-single-block-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-2-single-block-validation into .opencode/skills/phase-2-single-block-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-2-single-block-validation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phase-2-single-block-validationPhase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…
Phase 2 Single Block Validation is an agent skill from Xilinx/mlir-air. Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/ diagnosis lens, promoted to a gate at layer 0). Catches integration bugs (layout mismatches, missing transposes, type drops between kernel boundaries) before scaling to N layers.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Deployment. The licence is MIT.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phase 2 Single Block Validation loads about 3.4k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 1,543 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,543 words, ~3,372 tokens.
.claude/skills/phase-2-single-block-validation/SKILL.md (or your agent's skills folder).Phase 1 verified each kernel × shape standalone. Phase 2 wires those kernels into a single transformer block on NPU and compares the block-level output to the HF bf16 reference at the same input. This catches integration bugs (wrong tensor layouts at kernel boundaries, dropped biases, missing padding) without the cost of running N layers.
The comparison reuses the verify/ subsystem's diagnosis lens:
verify/verify_runner.py:_run_diagnosis already computes per-layer
ffn_out cosine (NPU vs HF bf16) for every layer. Diagnosis itself is
informational (no thresholds); Phase 2's job is to promote layer 0's
result to a hard gate with a head_dim-scaled threshold. The reference
is HF transformers in bf16 (same dtype as the NPU — fair fight), obtained
via HfRunner in diagnosis mode (lite_mode=False), NOT a hand-written
CPU forward.
Run the diagnosis lens on the canonical prompt and gate on layer 0 (the single block under test). All four must hold:
Whole-tensor cosine ≥ 0.99 between NPU layer-0 ffn_out and HF
bf16 layer-0 ffn_out (hf.layer_intermediates[0]['ffn_out']).
Catches: coarse integration breaks (wrong layout, missing op).
Per-position cosine ≥ THRESHOLD(head_dim) at each real-token
position ([:real_len], NOT padded positions — those are
out-of-distribution and amplify BF16 noise unhelpfully). The
diagnosis comparator already returns per-position cosine
{min, p5, median, mean} via per_position_cosine + aggregate; gate
on min. Catches: per-row dropouts the whole-tensor cosine averages over.
| head_dim | per-position min |
|---|---|
| ≤ 64 | 0.99 |
| 128 | 0.98 |
| ≥ 256 | 0.97 |
Threshold scales with head_dim because BF16 accumulation noise
grows as √(head_dim · K): hd=64 deployments hit per-pos min ≈ 0.998;
hd=128 hits ≈ 0.980 with ~5× LOWER MAE — the larger cosine drop is
geometric, not a bug.
No NaN anywhere in NPU output.
Result documented in
<model>/docs/development_progress/phase2_block.md (cosine numbers
Also record max_abs and max_rel error alongside the cosine numbers
(informational, not gated — absolute thresholds depend on input
distribution). The diagnosis comparator's error_metrics reports these.
The current BF16-output GEMM production path is the registry's
high-precision tier (FP32-accumulate + a single epilogue cast, fused-cast
or drain) at ~9.3e-3 mean_rel_L1 — the GPU-standard accuracy, single-
sourced from details/GEMM_bf16_in_bf16_out.json; the low-precision
direct-codegen tier (1.3e-2–1.9e-2) is the fallback. Across 7 GEMMs +
softmax + RoPE + RMSNorm the block stays in that high-precision band.
Recording them gives
future deployments a regression baseline (NPU max_abs ≤ 1.5× reference
deployment's measured value at same shape signals no regression).
Why cosine here, and why it's only an interim gate. A 2026 survey of industry practice (vLLM, HF transformers, llama.cpp, MLPerf, TensorRT-LLM, the GPTQ/AWQ/SmoothQuant literature) found that per-layer activation cosine is NOT a standard correctness gate — everyone gates either end-to-end (vLLM token-set, llama.cpp logit-KL, MLPerf ≥99%-of-reference) or per-tensor on element-wise atol/rtol / SQNR vs an FP32 reference. The one stack that does gate on per-layer cosine (TPU-MLIR) uses 0.99, not a loose value. So treat this gate honestly:
- At Phase 2 there is no token output yet to score, so we still need some numeric tripwire on the block output — cosine vs HF bf16 is that tripwire. It catches gross integration bugs (layout, missing transpose, dtype drop), which is all Phase 2 needs.
- It is bf16-vs-bf16 (NPU bf16 vs HF bf16) — there is no FP32 ground truth at this layer, so a tight element-wise atol/rtol (HF's FP32-parity bar) would mis-fire on benign rounding-order differences; cosine's direction-only nature is why it tolerates that. That tolerance is also its weakness (it can pass on magnitude errors), which is exactly why it's interim, not the final word.
- The real correctness gate is end-to-end (Phase 3/6 token-set top-5, which mirrors vLLM
check_logprobs_close). Once the full model exists, that gate — not cosine — decides correctness. A future upgrade could add SQNR (≥40 dB, PyTorch Numeric Suite's "very good alignment" bar) as a more defensible per-block number; thresholds unchanged for now.
PRIMARY:
programming_examples/llms/llama_kernel_builder/ — the shared toolkit
(KernelCache, stitching, external_kernels) you compose the block FROM
(kernel-first default).programming_examples/llms/llama32_1b/multi_launch_builder/ +
llama32_1b_prefill.py:run_transformer_block — the reference exemplar:
read to see how the leaf kernels stitch into a block. On a bit-for-bit
kernel-sequence match you may call run_transformer_block directly
(inheritance shortcut).programming_examples/llms/verify/verify_runner.py:_run_diagnosis
— the per-layer NPU-vs-HF-bf16 cosine lens Phase 2 promotes to a gate.programming_examples/llms/verify/runners/hf_runner.py — how the
HF bf16 reference exposes per-layer ffn_out (lite_mode=False).<model>/docs/development_progress/ (Phase 1 output) +
programming_examples/kernel_registry/supported_kernels.md rows with
Used by = <model> — Phase 1's verified (kernel, shape) list; Phase 2
must wire ALL of them.programming_examples/kernel_registry/details/<Kernel>_bf16.md
— kernel-by-kernel reference (datapath, tile rules, constraints, layouts).WORKAROUNDS (apply when model config triggers them — re-derive from the HF reference impl; the patterns below describe the technique, not a shipped file):
Kernel-first (default). Derive the model's per-layer kernel sequence
from its config and build the block by composing the registry leaf kernels
(verified in Phase 1) into model-specific multi-launch ELFs under
<model>/multi_launch_builder/, using the shared llama_kernel_builder
toolkit (KernelCache, stitching, external_kernels). This is the general
path — it does not assume the model resembles llama, so it generalizes to
any decoder-only architecture in scope.
Read llama32_1b's assembly (llama32_1b_prefill.run_transformer_block
and llama32_1b/multi_launch_builder/*) as a worked exemplar of how
the leaf kernels stitch into a block — mirror its structure, adapting the
kernel sequence and shapes to your model.
Inheritance (shortcut). ONLY when the model's per-layer kernel
sequence matches llama's bit-for-bit —
RMSNorm → Q/K/V GEMM → RoPE → FA → O → add → RMSNorm → Gate/Up → SwiGLU → Down → add —
you may skip writing builders and call
llama32_1b_prefill.run_transformer_block directly with the new shape
parameters. This is an optimization for genuine llama variants, not the
starting assumption. Any of these breaks the bit-for-bit match and forces
the kernel-first path:
(a) NEW op type (e.g., Qwen3's Q/K Norm — per-head RMSNorm with
(head_dim,) weight)
(b) NEW op needs to land BETWEEN currently-fused launches (e.g.,
Q/K Norm sits between Q/K projection and RoPE, but
rms_gemv_rope fuses both)
(c) Op REORDER (post-norm vs pre-norm)
Either way, don't write new C kernels speculatively — almost always the
leaf kernel exists in the registry; the trick is the right way to STITCH.
For Q/K Norm specifically, weighted_rms_norm with the heads-as-M trick
(M=n_heads, N=head_dim, sharing the (head_dim,) weight across rows) IS
the op.
Two known triggers from model config (NOT from upstream phases):
Non-1024-aligned emb_dim or hidden_dim → BD pool exhaustion
risk at long seq (see the kernel's details/<Kernel>_bf16.md placeability
notes). Use GQA-aware
reindexed padding: pad up to a 1024-aligned multiple by inserting phantom
Q heads INSIDE each KV group (not at the end — naive padding breaks GQA
semantics by changing n_heads / n_kv_heads = group_size). CPU-only
sanity test the padded vs orig forward FIRST (cosine should be 0.999998+)
before touching NPU.
qkv_bias=True (Qwen2 / Qwen3 family) → host-side post-RoPE bias add,
exploiting RoPE's linearity: RoPE(q + bq) = RoPE(q) + RoPE(bq). The
rms_gemms_rope ELF stays bias-free; bias is added on host after the ELF
returns.
If both: padding determines the n_heads count the bias precompute uses.
In <model>/<model>_prefill.py, implement
run_single_block(layer_idx=0, hidden, weights, ...):
run_transformer_block_<model>(...) that runs the per-model multi-launch
ELFs in order via the shared KernelCache. Minimal skeleton:from llama_kernel_builder.cache import KernelCache
cache = KernelCache() # compile-once, run-many
def run_transformer_block_<model>(hidden, weights, cfg, cache):
# one _run_cached per fused ELF you built in Step 1, in order:
x = cache._run_cached("rms_qkv_rope", hidden, weights.qkv, ...) # RMSNorm+Q/K/V+RoPE
x = cache._run_cached("attn", x, ...) # FA
x = cache._run_cached("o_ffn", x, weights.o, weights.ffn, ...) # O+add+RMSNorm+SwiGLU+Down+add
return xllama32_1b_prefill.run_transformer_block(...) with this model's shape
parametersUse the canonical prompt from the deployment's verify prompt set
(verify/prompts/{base,instruct}.txt). Get the layer-0 reference from the
HF bf16 runner and compare:
from verify.runners.hf_runner import HfRunner
hf = HfRunner(hf_model_id, config, max_seq, lite_mode=False)
hf_pf = hf.prefill(prompt_tokens)
ref_block0 = hf_pf.layer_intermediates[0]["ffn_out"] # HF bf16, layer-0 output
npu_block0 = run_single_block(layer_idx=0, hidden=x, weights=weights, ...)Compute whole-tensor cosine and per-position cosines (real-token positions
only) with the diagnosis comparator
(verify/comparators.py:per_position_cosine + aggregate). Check against
the PASS criteria above. The simplest route is to run make diagnosis and
read layer 0's row from the report; the manual snippet above is for when
you need to gate inside a Phase-2 test script.
If cosine fails, the integration is broken at one specific kernel
boundary. Bisect by swapping NPU kernels back to a CPU equivalent one at a
time (use <model>_cpu_helpers.py for the ops that have a helper —
rms_norm, attention_reference — and a small inline numpy for the rest):
walk forward through the block, replacing npu_<kernel>(...) with the CPU
equivalent, recompute cosine. The first replacement that pushes cosine
above threshold identifies the offender — that's where the layout / type /
argument mismatch lives. Invoke superpowers:systematic-debugging on it.
Record the bisect table (per-step cosine) in phase2_block.md so future
deployments learn from this specific failure.
| Symptom | Likely cause | Where to look |
|---|---|---|
| Cosine drops at Q/K/V GEMM | weight loading / tensor layout (seq-first vs heads-first) | Compare NPU output shape to reference's; check np.ascontiguousarray() after weight load |
| Cosine drops at FlashAttention | causal masking missing / wrong dk_chunks compile flag | See debug-fa-runtime-failure |
| Cosine drops at Down GEMM | BF16 truncation; running the low-precision (direct-codegen) tier instead of high-precision | confirm the GEMM uses the registry's high-precision path (fused-cast / drain = FP32-accumulate + single cast), not --high-precision false; see details/GEMM_bf16_in_bf16_out.md |
| NaN in output | uninitialized BO / reused stale buffer | Invoke debug-bo-corruption |
| Cosine drops at residual add | bias forgotten on padded path / GQA reindex bug | If padding+bias model: re-run CPU sanity test on padded forward (Step 2) |
| Whole-tensor cosine OK but per-position min low | one bad position run; check whether last few positions diverge (causal mask edge case) | Print per-position cosine, look for contiguous bad runs |
For any failure not in the table, invoke superpowers:systematic-debugging.
On Phase 2 PASS:
<model>/docs/development_progress/phase2_block.md<model>/TODO.md© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/phase-2-single-block-validation of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Phase 2 Single Block Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phase 2 Single Block Validation this skillXilinx/mlir-air | 150 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Kubeshark Installerkubeshark/kubeshark | 12k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | |
| GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb | 6.7k | — | ~4k | Automated safety check: Notes | Apache-2.0 | |
| KubeSphere ServiceMesh Managerkubesphere/kubesphere | 17k | — | ~2.4k | Automated safety check: Pass | Custom licence | |
| Vercelremotion-dev/remotion | 62k | — | ~1.2k | Automated safety check: Pass | Custom licence | |
| AWS Cdk Developmentzxkane/aws-skills | 367 | 2 repos | ~2.5k | Automated safety check: Pass | MIT |
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
GreptimeTeam/greptimedb
Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.
kubesphere/kubesphere
Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.
remotion-dev/remotion
Set up a Codex monitor for Vercel deployments and preview URLs.
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Categories
Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…. Phase 2 Single Block Validation is an agent skill from Xilinx/mlir-air. Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/ diagnosis lens, promoted to a gate at layer 0).
Phase 2 Single Block Validation fits situations like: tasks that involve Deployment.
Run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-2-single-block-validation in Xilinx/mlir-air) into .claude/skills/phase-2-single-block-validation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a codex`. Or copy the skill folder (.claude/skills/phase-2-single-block-validation in Xilinx/mlir-air) into .agents/skills/phase-2-single-block-validation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-2-single-block-validation, .gemini/skills/phase-2-single-block-validation, .github/skills/phase-2-single-block-validation and .opencode/skills/phase-2-single-block-validation in your project.
Going by SKILL.md and its folder, Phase 2 Single Block Validation needs the command-line tools its instructions call (make). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phase 2 Single Block Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phase 2 Single Block Validation: Kubeshark Installer (kubeshark/kubeshark, 12k stars), GreptimeDB Dev Docker Image (GreptimeTeam/greptimedb, 6.7k stars), KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars) and Vercel (remotion-dev/remotion, 62k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.