Web Application Testing
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .claude/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .claude/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .agents/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .agents/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .cursor/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .cursor/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/phase-3-full-model-validation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .gemini/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .gemini/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .github/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .github/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air phase-3-full-model-validation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/phase-3-full-model-validation .opencode/skills/phase-3-full-model-validation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phase-3-full-model-validation" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-3-full-model-validation into .opencode/skills/phase-3-full-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-3-full-model-validation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phase-3-full-model-validationPhase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…
Phase 3 Full Model Validation is an agent skill from Xilinx/mlir-air. Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level top-5 set-inclusion via its token-set gate) at canonical prompts. Catches accumulated drift, KV cache bugs, layer-indexed weight loading errors. Invoked after Phase 2 gate.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phase 3 Full Model Validation loads about 2.2k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,121 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,121 words, ~2,248 tokens.
.claude/skills/phase-3-full-model-validation/SKILL.md (or your agent's skills folder).Phase 2 verified one transformer block. Phase 3 scales to all N layers and confirms the full prefill stays numerically aligned with the HF bf16 reference end-to-end. Catches: accumulated BF16 drift across deep stacks, KV cache layout bugs, layer-indexed weight loading errors, LM head precision drops.
Phase 3 reuses the verify/ subsystem's two lenses:
make diagnosis — per-layer ffn_out cosine (NPU vs HF bf16) for
ALL layers. Informational by default; Phase 3 promotes it to a gate.make verify — token-level top-5 set-inclusion (NPU vs HF bf16
greedy sequences). Already a hard gate (exit 0/1).There is no hand-written CPU full-model forward — the reference is HF transformers bf16 throughout.
Run make diagnosis (per-layer cosine, all layers) and make verify
(token-set gate) on the canonical prompts. All must hold:
Semantic correctness (token-level, vs HF bf16 — THE GATE)
make verify PASSES: at the first divergence between NPU and HF
greedy sequences, NPU's chosen token is in HF's top-5 AND HF's chosen
token is in NPU's top-5 (the compute_topk_set_check gate, GATE_K=5,
GATE_N_TOKENS=32), measured against HF bf16 on generated tokens. bf16
noise can flip top-1 even between mathematically equivalent
implementations, but almost never displaces a token out of the top-5.
This top-k token-set gate mirrors vLLM's check_logprobs_close —
the GPU/industry-standard end-to-end signal, and the same gate Phase 6/7
re-run. This is the binding correctness verdict for Phase 3.Numerical sanity (per-layer, vs HF bf16 — diagnostic, NOT the gate)
|cos[i+1] − cos[i]| > 0.05)
points at a layer-indexed bug (wrong bo_key=f"kernel_L{i}",
weight-load shifted by a layer). If make verify PASSES, a gently
drifting cosine (a deep stack reaching ~0.88 at the last layer) is
expected bf16 geometry, not a bug.max_abs / max_rel alongside the cosines (the diagnosis
error_metrics reports them) as a regression baseline for future
deployments at the same shape.Hygiene
make verify PASS/FAIL for
every canonical prompt documented in
<model>/docs/development_progress/phase3_full.md.Why this split. A 2026 survey of industry practice (vLLM, HF transformers, llama.cpp, MLPerf, TensorRT-LLM, GPTQ/AWQ/SmoothQuant) found that correctness is gated end-to-end (vLLM token-set, llama.cpp logit-KL, MLPerf ≥99%-of-reference) — per-layer activation cosine is used for localization, not as a pass/fail gate, and where it is gated at all (TPU-MLIR) the bar is 0.99, not a loose value. Hence the token-set check is THE gate and cosine is the lens. A future upgrade could add a logit KL-divergence / perplexity-≥99%-of-HF check (llama.cpp / MLPerf style) as an even stronger end-to-end signal.
programming_examples/llms/llama32_1b/llama32_1b_inference.py:run_npu_prefill
— reference full-stack pipeline (loops Phase 2's per-layer block)programming_examples/llms/verify/verify_runner.py — both lenses:
_run_diagnosis (per-layer cosine) and the topk_token gate
(compute_topk_set_check)programming_examples/llms/verify/README.md — the verify
methodology (HF bf16 reference, top-k token-set gate, cosine as
diagnosis)<model>/docs/development_progress/ + the kernel_registry
"tested shapes" rows with Used by = <model>
— Phase 1 verified kernels (Phase 3 confirms they compose at scale)Implement run_full_prefill(input_ids, weights, config) (or reuse the
deployment's run_npu_prefill):
run_transformer_block_<model>(...)
(the per-model block runner Phase 2 produced) for
layer_idx in range(config.n_layers), then apply final RMSNorm + LM head.llama32_1b_prefill.run_transformer_block(...) the same way.Either way, this is just iterating the Phase 2 block runner N times plus head — no new kernels.
Use the deployment's verify prompt set (verify/prompts/{base,instruct}.txt
— base for base checkpoints, instruct for instruct/chat checkpoints).
make verify runs 2 prompts × 32 tokens by default; make verify-full
runs the full set. make diagnosis runs a single prompt for the per-layer
cosine table.
The token-set gate handles the old "decisive vs competitive prompt" distinction automatically: a top-1 flip within the top-5 band is treated as benign drift (not a failure), while a token leaving the top-5 entirely is a failure. No manual per-prompt probability classification is needed.
cd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make diagnosis # per-layer cosine, all layers
flock -x -w 1800 /tmp/mlir-air-npu.lock make verify # token-set gate (exit 0/1)From verify: PASS/FAIL (criterion 1, the gate); the report under
verify/reports/ records the first divergence and the top-5 sets on each
side. From diagnosis: read the per-layer cosine table as the localization
lens (criterion 2) — not a pass/fail, but where you look when verify fails.
If make verify fails, the diagnosis cosine table localizes where:
bo_key, wrong wq for that layer)make verify — if the token-set gate still PASSES, the
drift is geometric not a bug; if it FAILS, real bugIf make verify fails but per-layer cosine looks fine, the divergence is
likely in the decode path / KV-cache (verify generates 32 tokens; the
per-layer cosine only probes prefill). Validate KV cache values at end of
prefill against HF. Within an offending layer, bisect kernel-by-kernel
using Phase 2's CPU-fallback technique.
| Symptom | Likely cause | Where to look |
|---|---|---|
| Per-layer cos cliff at one layer | Layer-indexed weight load bug or wrong bo_key=f"kernel_L{i}" | Print weight shapes per layer; compare K/V cache layout at boundary |
| Per-layer cos drifts gradually < 0.85 by last layer | Real BF16 saturation, OR Down GEMM missing F32 accumulator | Run make verify — if token-set gate PASSES, drift is geometric not a bug; otherwise check F32 accumulator pattern |
Per-layer cos OK but make verify FAILS at prediction | LM head precision drop (BF16 truncation in vocab projection) | Apply F32 accumulator pattern to LM head GEMM |
make verify FAILS: NPU top-1 NOT in HF top-5 at all | Real correctness bug | Step 4 bisect |
| NaN in output | Uninitialized BO / reused stale buffer | Invoke debug-bo-corruption |
| Per-layer cos OK, prefill prediction OK, but multi-token generation diverges quickly | KV cache update bug at decode time (per-layer cosine only probes prefill; verify's 32-token generation catches it) | Validate KV cache values at end of prefill match HF |
For any failure not in the table, invoke superpowers:systematic-debugging.
On Phase 3 PASS, this is the end-to-end correctness milestone — NPU is now numerically faithful to the HF bf16 reference at full N-layer scale. Update:
<model>/docs/development_progress/phase3_full.md: per-layer cos table
(diagnosis) + make verify PASS per prompt<model>/docs/development_progress/progress.md: Phase 3 summary<model>/TODO.md: mark Phase 3, advance to perf phasesPhase 4 (prefill perf) and Phase 5 (decode perf) MUST preserve these gates
(make diagnosis per-layer cosine + make verify token-set) after every
optimization — perf cannot trade correctness.
© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/phase-3-full-model-validation of Xilinx/mlir-air.
Open the folder on GitHubat commit 6e81ce1
Phase 3 Full Model Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phase 3 Full Model Validation this skillXilinx/mlir-air | 150 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| Electron App Automationvercel-labs/agent-browser | 44k | 5 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| playwright-cli Browser Automationgithub/gh-aw | 5.4k | 24 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Write and Verify Playwright Testsappsmithorg/appsmith | 41k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 |
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
vercel-labs/agent-browser
Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port.
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
github/gh-aw
Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.
appsmithorg/appsmith
Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.
langgenius/dify
Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…. Phase 3 Full Model Validation is an agent skill from Xilinx/mlir-air. Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level top-5 set-inclusion via its token-set gate) at canonical prompts.
Phase 3 Full Model Validation fits situations like: tasks that involve End-to-end testing.
Run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-3-full-model-validation in Xilinx/mlir-air) into .claude/skills/phase-3-full-model-validation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a codex`. Or copy the skill folder (.claude/skills/phase-3-full-model-validation in Xilinx/mlir-air) into .agents/skills/phase-3-full-model-validation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-3-full-model-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-3-full-model-validation, .gemini/skills/phase-3-full-model-validation, .github/skills/phase-3-full-model-validation and .opencode/skills/phase-3-full-model-validation in your project.
Going by SKILL.md and its folder, Phase 3 Full Model Validation needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phase 3 Full Model Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phase 3 Full Model Validation: Web Application Testing (anthropics/skills, 180k stars), Electron App Automation (vercel-labs/agent-browser, 44k stars), OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars) and playwright-cli Browser Automation (github/gh-aw, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.