Agent skill

Phase 3 Full Model Validation

by Xilinx in Xilinx/mlir-air

Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…

MITAuto-check passed

Install Phase 3 Full Model Validation

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .claude/skills/phase-3-full-model-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-3-full-model-validation
GitHub stars
150
Token cost
~2.2k tokens
SKILL.md length
1,121 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…

  • Tasks that involve End-to-end testing
  • SKILL.md covers Purpose, Phase 3 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls make

What it does

Phase 3 Full Model Validation is an agent skill from Xilinx/mlir-air. Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level top-5 set-inclusion via its token-set gate) at canonical prompts. Catches accumulated drift, KV cache bugs, layer-indexed weight loading errors. Invoked after Phase 2 gate.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

When your agent uses it

  • Tasks that involve End-to-end testing

Example prompts

  • “/phase-3-full-model-validation”

What it can do on your machine

Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 3 Full Model Validation loads about 2.2k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,121 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,121 words, ~2,248 tokens.

Download SKILL.mdSave it as .claude/skills/phase-3-full-model-validation/SKILL.md (or your agent's skills folder).
name
phase-3-full-model-validation
description
Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared `programming_examples/llms/verify/` diagnosis lens + token-level top-5 set-inclusion via its token-set gate) at canonical prompts. Catches accumulated drift, KV cache bugs, layer-indexed weight loading errors. Invoked after Phase 2 gate.

Purpose

Phase 2 verified one transformer block. Phase 3 scales to all N layers and confirms the full prefill stays numerically aligned with the HF bf16 reference end-to-end. Catches: accumulated BF16 drift across deep stacks, KV cache layout bugs, layer-indexed weight loading errors, LM head precision drops.

Phase 3 reuses the verify/ subsystem's two lenses:

  • make diagnosis — per-layer ffn_out cosine (NPU vs HF bf16) for ALL layers. Informational by default; Phase 3 promotes it to a gate.
  • make verify — token-level top-5 set-inclusion (NPU vs HF bf16 greedy sequences). Already a hard gate (exit 0/1).

There is no hand-written CPU full-model forward — the reference is HF transformers bf16 throughout.

Phase 3 PASS criteria (HARD GATES)

Run make diagnosis (per-layer cosine, all layers) and make verify (token-set gate) on the canonical prompts. All must hold:

Semantic correctness (token-level, vs HF bf16 — THE GATE)

  1. make verify PASSES: at the first divergence between NPU and HF greedy sequences, NPU's chosen token is in HF's top-5 AND HF's chosen token is in NPU's top-5 (the compute_topk_set_check gate, GATE_K=5, GATE_N_TOKENS=32), measured against HF bf16 on generated tokens. bf16 noise can flip top-1 even between mathematically equivalent implementations, but almost never displaces a token out of the top-5. This top-k token-set gate mirrors vLLM's check_logprobs_close — the GPU/industry-standard end-to-end signal, and the same gate Phase 6/7 re-run. This is the binding correctness verdict for Phase 3.

Numerical sanity (per-layer, vs HF bf16 — diagnostic, NOT the gate)

  1. Per-layer cosine table recorded, eyeballed for gross failure. This is bf16-vs-bf16 (no FP32 ground truth at the layer level), and per-layer activation cosine is not an industry-standard correctness gate (see the note below) — so it does not PASS/FAIL the phase. Use it to localize a problem when criterion 1 fails: a layer where cosine collapses (e.g. < 0.85, or a sudden cliff |cos[i+1] − cos[i]| > 0.05) points at a layer-indexed bug (wrong bo_key=f"kernel_L{i}", weight-load shifted by a layer). If make verify PASSES, a gently drifting cosine (a deep stack reaching ~0.88 at the last layer) is expected bf16 geometry, not a bug.
  2. Record max_abs / max_rel alongside the cosines (the diagnosis error_metrics reports them) as a regression baseline for future deployments at the same shape.

Hygiene

  1. No NaN anywhere in the stack.
  2. Per-layer cos table (from diagnosis) + make verify PASS/FAIL for every canonical prompt documented in <model>/docs/development_progress/phase3_full.md.

Why this split. A 2026 survey of industry practice (vLLM, HF transformers, llama.cpp, MLPerf, TensorRT-LLM, GPTQ/AWQ/SmoothQuant) found that correctness is gated end-to-end (vLLM token-set, llama.cpp logit-KL, MLPerf ≥99%-of-reference) — per-layer activation cosine is used for localization, not as a pass/fail gate, and where it is gated at all (TPU-MLIR) the bar is 0.99, not a loose value. Hence the token-set check is THE gate and cosine is the lens. A future upgrade could add a logit KL-divergence / perplexity-≥99%-of-HF check (llama.cpp / MLPerf style) as an even stronger end-to-end signal.

Knowledge base references

  • programming_examples/llms/llama32_1b/llama32_1b_inference.py:run_npu_prefill — reference full-stack pipeline (loops Phase 2's per-layer block)
  • programming_examples/llms/verify/verify_runner.py — both lenses: _run_diagnosis (per-layer cosine) and the topk_token gate (compute_topk_set_check)
  • programming_examples/llms/verify/README.md — the verify methodology (HF bf16 reference, top-k token-set gate, cosine as diagnosis)
  • <model>/docs/development_progress/ + the kernel_registry "tested shapes" rows with Used by = <model> — Phase 1 verified kernels (Phase 3 confirms they compose at scale)

Workflow

Step 1: Wire all N layers

Implement run_full_prefill(input_ids, weights, config) (or reuse the deployment's run_npu_prefill):

  • Kernel-first path (default): loop run_transformer_block_<model>(...) (the per-model block runner Phase 2 produced) for layer_idx in range(config.n_layers), then apply final RMSNorm + LM head.
  • Inheritance path (shortcut, bit-for-bit match only): loop llama32_1b_prefill.run_transformer_block(...) the same way.

Either way, this is just iterating the Phase 2 block runner N times plus head — no new kernels.

Step 2: Canonical prompts

Use the deployment's verify prompt set (verify/prompts/{base,instruct}.txt — base for base checkpoints, instruct for instruct/chat checkpoints). make verify runs 2 prompts × 32 tokens by default; make verify-full runs the full set. make diagnosis runs a single prompt for the per-layer cosine table.

The token-set gate handles the old "decisive vs competitive prompt" distinction automatically: a top-1 flip within the top-5 band is treated as benign drift (not a failure), while a token leaving the top-5 entirely is a failure. No manual per-prompt probability classification is needed.

Show full SKILL.md (434 more words)Show less
Step 3: Run both lenses + collect metrics
bash
cd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make diagnosis    # per-layer cosine, all layers
flock -x -w 1800 /tmp/mlir-air-npu.lock make verify       # token-set gate (exit 0/1)

From verify: PASS/FAIL (criterion 1, the gate); the report under verify/reports/ records the first divergence and the top-5 sets on each side. From diagnosis: read the per-layer cosine table as the localization lens (criterion 2) — not a pass/fail, but where you look when verify fails.

Step 4: Bisect on FAIL

If make verify fails, the diagnosis cosine table localizes where:

  • Sudden cliff at layer i (cos[i] >> cos[i+1]) → layer-indexed bug at i+1 (weight load shifted, wrong bo_key, wrong wq for that layer)
  • Gradual drift but a layer < 0.85 → BF16 accumulator saturating; check whether the production GEMM path uses F32 internal accumulate; cross-check make verify — if the token-set gate still PASSES, the drift is geometric not a bug; if it FAILS, real bug
  • Layer 0 already low → integration error in the single-block runner itself; revisit Phase 2

If make verify fails but per-layer cosine looks fine, the divergence is likely in the decode path / KV-cache (verify generates 32 tokens; the per-layer cosine only probes prefill). Validate KV cache values at end of prefill against HF. Within an offending layer, bisect kernel-by-kernel using Phase 2's CPU-fallback technique.

Failure modes

SymptomLikely causeWhere to look
Per-layer cos cliff at one layerLayer-indexed weight load bug or wrong bo_key=f"kernel_L{i}"Print weight shapes per layer; compare K/V cache layout at boundary
Per-layer cos drifts gradually < 0.85 by last layerReal BF16 saturation, OR Down GEMM missing F32 accumulatorRun make verify — if token-set gate PASSES, drift is geometric not a bug; otherwise check F32 accumulator pattern
Per-layer cos OK but make verify FAILS at predictionLM head precision drop (BF16 truncation in vocab projection)Apply F32 accumulator pattern to LM head GEMM
make verify FAILS: NPU top-1 NOT in HF top-5 at allReal correctness bugStep 4 bisect
NaN in outputUninitialized BO / reused stale bufferInvoke debug-bo-corruption
Per-layer cos OK, prefill prediction OK, but multi-token generation diverges quicklyKV cache update bug at decode time (per-layer cosine only probes prefill; verify's 32-token generation catches it)Validate KV cache values at end of prefill match HF

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

On Phase 3 PASS, this is the end-to-end correctness milestone — NPU is now numerically faithful to the HF bf16 reference at full N-layer scale. Update:

  • <model>/docs/development_progress/phase3_full.md: per-layer cos table (diagnosis) + make verify PASS per prompt
  • <model>/docs/development_progress/progress.md: Phase 3 summary
  • <model>/TODO.md: mark Phase 3, advance to perf phases

Phase 4 (prefill perf) and Phase 5 (decode perf) MUST preserve these gates (make diagnosis per-layer cosine + make verify token-set) after every optimization — perf cannot trade correctness.

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-3-full-model-validation of Xilinx/mlir-air.

Open the folder on GitHubat commit 6e81ce1

Compare with similar skills

Phase 3 Full Model Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 3 Full Model Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 3 Full Model Validation this skillXilinx/mlir-air150—~2.2kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Electron App Automationvercel-labs/agent-browser44k5 repos~1.7kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
playwright-cli Browser Automationgithub/gh-aw5.4k24 repos~2.8kAutomated safety check: PassMIT
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Electron App Automation

    vercel-labs/agent-browser

    Official

    Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port.

    44k GitHub starsUsed in 5 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.4k GitHub starsUsed in 24 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Phase 3 Full Model Validation

What does Phase 3 Full Model Validation do?

Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…. Phase 3 Full Model Validation is an agent skill from Xilinx/mlir-air. Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level top-5 set-inclusion via its token-set gate) at canonical prompts.

When should I use Phase 3 Full Model Validation?

Phase 3 Full Model Validation fits situations like: tasks that involve End-to-end testing.

How do I install Phase 3 Full Model Validation in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-3-full-model-validation in Xilinx/mlir-air) into .claude/skills/phase-3-full-model-validation in your project. Claude Code loads it when a task matches its description.

How do I install Phase 3 Full Model Validation in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a codex`. Or copy the skill folder (.claude/skills/phase-3-full-model-validation in Xilinx/mlir-air) into .agents/skills/phase-3-full-model-validation in your project. Codex loads it when a task matches its description.

Can I use Phase 3 Full Model Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-3-full-model-validation, .gemini/skills/phase-3-full-model-validation, .github/skills/phase-3-full-model-validation and .opencode/skills/phase-3-full-model-validation in your project.

What does Phase 3 Full Model Validation need to run?

Going by SKILL.md and its folder, Phase 3 Full Model Validation needs the command-line tools its instructions call (make).

Does Phase 3 Full Model Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phase 3 Full Model Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 3 Full Model Validation use?

Phase 3 Full Model Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 3 Full Model Validation use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 3 Full Model Validation?

Skills that share tags, products or a category with Phase 3 Full Model Validation: Web Application Testing (anthropics/skills, 180k stars), Electron App Automation (vercel-labs/agent-browser, 44k stars), OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars) and playwright-cli Browser Automation (github/gh-aw, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 3 Full Model Validation?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.