Kubeshark Installer
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/…
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learn --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .claude/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .claude/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learnType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learn --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .agents/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .agents/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learn --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .cursor/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .cursor/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/phase-6-finalize-and-learn--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learn --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .gemini/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .gemini/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learnInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .github/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .github/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air phase-6-finalize-and-learn --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/phase-6-finalize-and-learn .opencode/skills/phase-6-finalize-and-learn && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phase-6-finalize-and-learn" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-6-finalize-and-learn into .opencode/skills/phase-6-finalize-and-learn/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-6-finalize-and-learn", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phase-6-finalize-and-learnPhase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/…
Phase 6 Finalize And Learn is an agent skill from Xilinx/mlir-air. Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/ subsystem + a Makefile (run / verify / verify-full / diagnosis / profile), and confirm make verify (top-k token-set gate vs HF bf16) PASSES. That gate is the production-readiness check. Capture lessons learned. Invoked after Phase 5 PASS.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Deployment. The licence is MIT.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phase 6 Finalize And Learn loads about 3.5k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 1,458 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,458 words, ~3,516 tokens.
.claude/skills/phase-6-finalize-and-learn/SKILL.md (or your agent's skills folder).Phase 4-5 produced an optimized prefill kernel and an optimized decode
kernel — but they may live in separate scripts with optimization-time
warmup hacks scattered through the main flow. Phase 6 integrates them
into a single clean <model>_inference.py and wires the deployment into
the shared programming_examples/llms/verify/ subsystem via a per-model verify_adapter.py
(mirroring llama32_1b/verify_adapter.py) so the deployment has:
setup() BEFORE the timed region).Makefile with the same targets the reference deployment and the
Phase 7 evaluator rely on:make run — runs inference, prints TTFT (prefill ms) + TPS (tokens/sec)make verify — top-k token-set gate vs HF bf16 (2 prompts × 32 tokens, k=5) — the production-readiness gatemake verify-full — same gate over the full prompt setmake diagnosis — per-layer ffn_out cosine vs HF bf16 (informational)make profile — per-phase + per-key-kernel breakdownLESSONS.md for future deployments.Why the token-set gate (not a hand-written CPU greedy match): the
reference is HF transformers in bf16, and the gate (compute_topk_set_check
in verify/comparators.py) checks that at the first divergence between
NPU and HF greedy sequences, each side's chosen token is in the other's
top-5. This catches decode-side KV-cache bugs that prefill-only checks
miss, while tolerating benign bf16 top-1 flips within the top-5 band. It
runs the production prefill/decode path, so it exercises the real
deployment, not a separate verify-only code path.
This top-k token-level inclusion check mirrors vLLM's correctness
methodology — it is the GPU/industry-standard end-to-end signal, which
is why it (not a per-tensor cosine) is the production-readiness gate. The
whole verify/ subsystem is shared across all models under
programming_examples/llms/verify/; a new deployment hooks in with a thin verify_adapter.py
rather than copying the runner, so every model is judged by the identical
gate. See programming_examples/llms/verify/README.md.
<model>_inference.py exists with the clean structure: setup()
(one-time preprocess, weight pre-load, BO allocation) called ONCE
before the profiled prefill() + decode_loop() region. No warmup
hacks, cache prime calls, or timing resets in the main flow.<model>/verify_adapter.py exists, mirroring
llama32_1b/verify_adapter.py: it provides the adapter interface the
shared programming_examples/llms/verify/verify_runner.py calls — resolve_model,
hf_reference, build_config, build_runner, and an NpuRunner
(implementing .prefill() / .decode_step()) that calls THIS model's
production run_npu_prefill / run_npu_decode_step. The verify runner,
comparators, report, and HF runner are NOT copied — they live once in
programming_examples/llms/verify/.make run works: invokes inference at default --n-tokens 100,
prints TTFT (prefill kernel ms) + TPS (tokens/sec).make verify PASSES the token-set gate (production-readiness):compute_topk_set_check, GATE_K=5, GATE_N_TOKENS=32).report.has_failure() is False → exit 0.make diagnosis runs without error (informational, not a gate —
the verify subsystem retired threshold-based diagnosis; compare_pair
reports per-layer cosine with no pass/fail). Eyeball the per-layer
cosine table against the Phase 3 baseline to confirm the finalized
integration didn't perturb the numerical alignment; if criterion 4
(make verify) FAILs, this table is the localization lens. The gate is
criterion 4, not this.make profile works: outputs per-phase total (setup / prefill
total / per-token decode avg / LM head) + key-kernel ms (FA, Down
GEMM/GEMV, LM Head — the bottleneck candidates).LESSONS.md updated with any new experiences (informational; not
gating the technical artifacts above).If make verify fails, the deployment is NOT production-ready regardless
of how good Phase 4/5 perf numbers look.
PRIMARY:
programming_examples/llms/llama32_1b/llama32_1b_inference.py — reference
clean inference structure (setup / prefill / decode_loop); copy fromprogramming_examples/llms/llama32_1b/verify_adapter.py — reference
adapter to mirror (the interface the shared verify runner calls)programming_examples/llms/verify/ — the shared verify subsystem
(runner + comparators + report + HF runner + prompts); hooked into, not copiedprogramming_examples/llms/llama32_1b/Makefile — reference target set
(run / verify / verify-full / diagnosis / profile / compile / clean)programming_examples/llms/<model>/docs/development_progress/{phase4_prefill,phase5_decode}.md
— Phase 4/5 outputs: which integration path was used + the optimized
prefill/decode runners to integrateSECONDARY:
programming_examples/llms/verify/README.md — verify methodologyprogramming_examples/kernel_registry/supported_kernels.md
— Phase 1's registry rows (Phase 6 confirms "Used by" reflects this model)<model>_inference.pyCopy programming_examples/llms/llama32_1b/llama32_1b_inference.py as the
starting point. The structure should be:
def setup(weights, config):
"""ONE-TIME preprocess: pre-load weight BOs, allocate caches,
install head-first FA wrapper if head_dim ≥ 128, etc. Everything
that should NOT be inside the profiled scope."""
...
def run_npu_prefill(input_ids, ...):
"""Phase 4 prefill — clean of warmup hacks."""
...
def run_npu_decode_step(token, pos, ...):
"""Phase 5 single decode step — clean. Called by both the decode
loop AND the NpuRunner in verify_adapter.py."""
...Audit Phase 4/5 scripts for warmup hacks that crept into the main
flow — pre-warm cache calls, dummy runs, timing resets — and move them
into setup() (or delete if no longer needed). The profiled scope must
be ONLY prefill + decode_loop.
Crucial: run_npu_prefill / run_npu_decode_step are the SAME
functions the model's verify_adapter.py NpuRunner calls. The verify
gate exercises the production path precisely because it imports these — do
not fork a verify-only copy.
verify_adapter.pyThe verify runner/comparators/report/HF runner are shared in
programming_examples/llms/verify/ — you do NOT copy them. You write one per-model file,
<model>/verify_adapter.py, mirroring llama32_1b/verify_adapter.py. It
provides the adapter interface the shared verify_runner.py loads via
--runner=<model>.verify_adapter:
resolve_model(choice) / hf_reference(name) → map --model
(base/instruct) to the HF checkpoint id.build_config() → return THIS model's Config.build_runner(...) → load weights, compile kernels, return the runner.class NpuRunner → .prefill() / .decode_step() calling THIS model's
production run_npu_prefill / run_npu_decode_step.The HF reference path, the comparators (compute_topk_set_check), and the
report all stay in programming_examples/llms/verify/ — unchanged, model-agnostic. If the
default chat template differs, point resolve_model / the prompt choice
at the right programming_examples/llms/verify/prompts/*.txt.
Mirror programming_examples/llms/llama32_1b/Makefile — note the verify
runner is the shared one at ../verify/, selected via --runner:
RUNNER_ADAPTER := <model>.verify_adapter
VERIFY_RUNNER := $(srcdir)/../verify/verify_runner.py
run:
flock -x -w 1800 /tmp/mlir-air-npu.lock \
bash -c 'cd $(BUILD_DIR) && python3 $(srcdir)/<model>_inference.py --n-tokens 100'
verify:
flock -x -w 1800 /tmp/mlir-air-npu.lock \
bash -c 'cd $(BUILD_DIR) && python3 $(VERIFY_RUNNER) --runner=$(RUNNER_ADAPTER) --prompts topk_token --model $(MODEL) --max-prompts 2'
verify-full:
flock -x -w 1800 /tmp/mlir-air-npu.lock \
bash -c 'cd $(BUILD_DIR) && python3 $(VERIFY_RUNNER) --runner=$(RUNNER_ADAPTER) --prompts topk_token --model $(MODEL)'
diagnosis:
flock -x -w 1800 /tmp/mlir-air-npu.lock \
bash -c 'cd $(BUILD_DIR) && python3 $(VERIFY_RUNNER) --runner=$(RUNNER_ADAPTER) --prompts single --model $(MODEL)'
profile:
flock -x -w 1800 /tmp/mlir-air-npu.lock \
bash -c 'cd $(BUILD_DIR) && python3 $(srcdir)/<model>_inference.py --profile --n-tokens 20'MODEL defaults to instruct (matches what production stacks deploy);
MODEL=base selects the base prompt set.
make verify — the production-readiness gatecd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make verifyThe runner: both NPU and HF bf16 greedy-decode each prompt × 32 tokens,
then compute_topk_set_check compares the two sequences. PASS = no
npu_vs_hf record is FAIL → exit 0; the report under verify/reports/
records the first divergence and the top-5 sets on each side.
If it FAILS at token i ≥ 1 (but make diagnosis per-layer cosine on
prefill is fine), the divergence is in the decode path / KV-cache — print
K/V cache values after token i-1 and compare to HF.
make run, capture TTFT + TPSRecord final numbers in <model>/docs/development_progress/phase6_finalize.md:
| Metric | Value | vs reference llama32_1b |
|---|---|---|
| TTFT (prefill kernel ms) | X | Y× / Y% |
| TPS (tokens/sec) | A | B× / B% |
| Decode ms/token | T | — |
make profile, capture breakdown--profile mode prints per-phase totals (setup / prefill / per-token
decode / LM head) + key-kernel ms (FA, Down GEMM/GEMV, LM Head — the
bottleneck candidates). Record in phase6_finalize.md.
Append to <model>/docs/development_progress/LESSONS.md for any new
experience (debug techniques used, surprising failures, configs that
mattered).
Then audit for promotion candidates (don't promote speculatively — only if 2+ uses):
<model>/multi_launch_builder/
(kernel-first path)? Cross-reference other deployments — if a 2nd
uses the same pattern, it's a candidate for a future shared location.kernel_registry/details/<Kernel>_bf16.md?.claude/skills/<phase>/SKILL.md directly (git history is
the change trail). Surface in <model>/TODO.md if not done inline.Sanity check that Phase 1's registry step completed:
kernel_registry/supported_kernels.md (and each details/<Kernel>_bf16.md)
has a "tested shapes" row with Used by = <model> for every new
(kernel, shape) this model exercises<model>/docs/development_progress/ has the full per-kernel results
(cosine, max_abs/max_rel, profile, status) backing those rowsIf gaps, fix here before Phase 7.
| Symptom | Likely cause | Where to look |
|---|---|---|
make verify FAILS at token i ≥ 1 but make diagnosis prefill cosine is fine | KV cache update bug at decode time (diagnosis only probes prefill) | Print K/V cache values after token i-1 vs HF; usually a layout / write-offset bug in the decode kernel |
make verify FAILS at token 0 | Prefill-side issue (LM Head precision, final norm) | Re-run make diagnosis; root cause is in prefill, not decode |
make verify errors importing the NPU runner | verify_adapter.py's NpuRunner not wired to THIS model's prefill/decode functions | Confirm the import points at <model>_inference.py, not the llama32_1b copy |
| TTFT regressed vs Phase 4 baseline | Integration introduced overhead (warmup hack creep, redundant setup in main flow) | Compare Phase 4 standalone profile to current make profile setup section |
| TPS regressed vs Phase 5 baseline | Same as above for decode | Compare Phase 5 standalone profile |
make profile shows huge "Setup" time inside profiled scope | setup() called inside the timed region instead of once before | Refactor — inference() must call setup() BEFORE t0 = time.time() |
For any failure not in the table, invoke superpowers:systematic-debugging.
On Phase 6 PASS:
<model>/docs/development_progress/phase6_finalize.md: TTFT + TPS +
profile breakdown + LESSONS summary<model>/TODO.md: mark Phase 6 PASSED<model>/ARCHITECTURE.md: write or update with final summary (model
config, key file map, perf headline). NOTE: use ARCHITECTURE.md, not
CLAUDE.md — the top-level .gitignore excludes CLAUDE.md, so it
would not ship in the PR.phase-7-independent-evaluator to re-derive every claim from scratch.
Phase 6 is "deployment is internally complete"; Phase 7 is "deployment
is independently audited".© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/phase-6-finalize-and-learn of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Phase 6 Finalize And Learn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phase 6 Finalize And Learn this skillXilinx/mlir-air | 150 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Kubeshark Installerkubeshark/kubeshark | 12k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | |
| GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb | 6.7k | — | ~4k | Automated safety check: Notes | Apache-2.0 | |
| KubeSphere ServiceMesh Managerkubesphere/kubesphere | 17k | — | ~2.4k | Automated safety check: Pass | Custom licence | |
| Vercelremotion-dev/remotion | 62k | — | ~1.2k | Automated safety check: Pass | Custom licence | |
| AWS Cdk Developmentzxkane/aws-skills | 367 | 2 repos | ~2.5k | Automated safety check: Pass | MIT |
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
GreptimeTeam/greptimedb
Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.
kubesphere/kubesphere
Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.
remotion-dev/remotion
Set up a Codex monitor for Vercel deployments and preview URLs.
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Categories
Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/…. Phase 6 Finalize And Learn is an agent skill from Xilinx/mlir-air.py hooking into the shared programmingexamples/llms/verify/ subsystem + a Makefile (run / verify / verify-full / diagnosis / profile), and confirm make verify (top-k token-set gate vs HF bf16) PASSES.
Phase 6 Finalize And Learn fits situations like: tasks that involve Deployment.
Run `npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a claude-code`. Or copy the skill folder (.claude/skills/phase-6-finalize-and-learn in Xilinx/mlir-air) into .claude/skills/phase-6-finalize-and-learn in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a codex`. Or copy the skill folder (.claude/skills/phase-6-finalize-and-learn in Xilinx/mlir-air) into .agents/skills/phase-6-finalize-and-learn in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-6-finalize-and-learn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-6-finalize-and-learn, .gemini/skills/phase-6-finalize-and-learn, .github/skills/phase-6-finalize-and-learn and .opencode/skills/phase-6-finalize-and-learn in your project.
Going by SKILL.md and its folder, Phase 6 Finalize And Learn needs the command-line tools its instructions call (make). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phase 6 Finalize And Learn is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phase 6 Finalize And Learn: Kubeshark Installer (kubeshark/kubeshark, 12k stars), GreptimeDB Dev Docker Image (GreptimeTeam/greptimedb, 6.7k stars), KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars) and Vercel (remotion-dev/remotion, 62k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.