Agent skill

Phase 7 Independent Evaluator

by Xilinx in Xilinx/mlir-air

Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…

MITAuto-check passedAgent Workflows

Install Phase 7 Independent Evaluator

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .claude/skills/phase-7-independent-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-7-independent-evaluator
GitHub stars
150
Token cost
~2.4k tokens
SKILL.md length
1,005 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…

  • Tasks that involve Subagents
  • SKILL.md covers Purpose, Phase 7 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls make
  • Tasks that involve Deployment

What it does

Phase 7 Independent Evaluator is an agent skill from Xilinx/mlir-air. Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the production path vs HF bf16), then re-runs it as the primary gate. Produces a structured evaluationreport.md a human can read in 2 minutes to know the full deployment state. Invoke as /phase-7-independent-evaluator <modeldir or auto-spawn from deploy-new-llm after Phase 6 PASS.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents and Deployment. The licence is MIT.

When your agent uses it

  • Tasks that involve Subagents
  • Tasks that involve Deployment

Example prompts

  • “/phase-7-independent-evaluator”

What it can do on your machine

Read from SKILL.md and the folder at commit 416edae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 7 Independent Evaluator loads about 2.4k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,005 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 416edae, republished under its MIT licence (© Xilinx). 1,005 words, ~2,361 tokens.

Download SKILL.mdSave it as .claude/skills/phase-7-independent-evaluator/SKILL.md (or your agent's skills folder).
name
phase-7-independent-evaluator
description
Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the `make verify` implementation (anti-reward-hacking: confirms the token-set gate runs the production path vs HF bf16), then re-runs it as the primary gate. Produces a structured `evaluation_report.md` a human can read in 2 minutes to know the full deployment state. Invoke as `/phase-7-independent-evaluator <model_dir>` or auto-spawn from deploy-new-llm after Phase 6 PASS.

Purpose

The deploy-new-llm chain is autonomous. Phase 6 wires up a make verify gate that's supposed to compare NPU vs the HF bf16 reference via the verify/ subsystem's top-k token-set check. But the deployment agent wrote (or copied) both the gate AND the production code it exercises — nothing structurally prevents the gate from being mocked, pointed at a stale baseline, or wired to a verify-only code path that bypasses the real kernels. Phase 7 closes this by:

  1. Spawning a FRESH subagent with no context from the deployment
  2. Auditing the make verify code path (anti-reward-hacking) — confirming the gate actually drives the production kernels vs HF bf16 and gates on the token-set check, not a mocked / stale / verify-only path. This is what makes the independent re-run meaningful: without it, "the agent re-ran the gate it wrote itself" proves nothing.
  3. Re-running the end-to-end make verify independently — the shared verify/ top-k token-set check (vLLM-aligned) of NPU vs HF bf16 on the real production prefill/decode path. A clean independent PASS, on a gate the audit confirmed is real, is the verdict.
  4. Producing a short evaluation report humans can read in 2 minutes.

The deployment is NOT considered trustworthy until this report is written and its overall verdict is PASS or PASS-with-warnings.

Phase 7 PASS criteria (HARD GATES)

  1. make verify audited (anti-reward-hacking): subagent has read the Makefile target + the model's verify_adapter.py + the shared programming_examples/llms/verify/{verify_runner,comparators,runners/hf_runner}.py, and confirmed the gate (a) drives the model's PRODUCTION run_npu_prefill / run_npu_decode_step via the NpuRunner (not a mock), (b) compares against HF transformers in bf16 (torch_dtype=torch.bfloat16), (c) gates on compute_topk_set_check (top-5 set inclusion at first divergence), not a hardcoded return True / > 0 threshold.
  2. make verify PASSES under the fresh subagent run: token-set gate exits 0 (no npu_vs_hf FAIL), on the audited-real gate.
  3. Short evaluation report written with the audit verdict + the re-run result + the measured numbers behind them. Verdict PASS or PASS-with-warnings.

If the subagent declares PASS without showing measured numbers OR without auditing the verify code path, the report is REJECTED — re-spawn with stricter instructions.

Knowledge base references

  • programming_examples/llms/verify/README.md — the verify methodology the audit checks against (HF bf16 reference, top-k token-set gate, cosine-as-diagnosis)
  • programming_examples/llms/verify/comparators.py — the gate implementation (compute_topk_set_check) the audit confirms is real
  • <model>/docs/development_progress/ + the kernel_registry rows with Used by = <model> — Phase 1 catalog (subagent compares its measurements against this)
  • programming_examples/kernel_registry/supported_kernels.md — kernel-by-kernel ground truth

Workflow

Step 1: Spawn fresh subagent with independence + audit constraints

Use the general-purpose Agent type. Critical instructions in the spawn prompt:

  • Independence: do NOT read <model>/docs/development_progress/{LESSONS,progress,phaseN_*}.md BEFORE forming your own measurements. You may CITE them AFTER measuring (compare your numbers vs claimed).
  • Re-derivation: every PASS/FAIL verdict must be backed by a number YOU measured or a code path YOU read during this audit, not a number copied from a deployment doc.
  • Reward-hacking smell test: if the deployment-claimed numbers look surprisingly good, audit the gate implementation FIRST — does make verify really drive the production prefill/decode vs HF bf16 and gate on the token-set check, or did the agent shortcut to "if NPU output exists → PASS" / point the verify_adapter NpuRunner at a mock / compare against a stale cached baseline?
Step 2: Audit make verify — anti-reward-hacking

BEFORE running anything, READ:

  1. The Makefile: what command does make verify invoke?
    bash
    grep -A 5 "^verify:" <model_dir>/Makefile
    It should run the shared runner ../verify/verify_runner.py --runner=<model>.verify_adapter --prompts topk_token ... — not a model-local copy of the runner.
  2. <model_dir>/verify_adapter.py: this is the model's only verify code. Confirm its NpuRunner (the .prefill() / .decode_step() methods) imports and calls THIS model's production run_npu_prefill / run_npu_decode_step from <model>_inference.py — NOT a stub, NOT the llama32_1b copy. (The runner/comparators/report/HF runner live once in the shared programming_examples/llms/verify/; the adapter is the per-model hook.)
  3. programming_examples/llms/verify/runners/hf_runner.py (shared): confirm the reference loads with AutoModelForCausalLM.from_pretrained(..., torch_dtype=torch.bfloat16).
  4. programming_examples/llms/verify/comparators.py + programming_examples/llms/verify/report.py (shared): confirm the gate is compute_topk_set_check (top-5 set inclusion at first divergence) and report.has_failure() drives a real exit code — not a hardcoded return True / > 0 threshold.
  5. If any of (2)-(4) fails: report [FAIL] verify gate is reward-hacked and stop here. Tag the deployment as needs-remediation.
Show full SKILL.md (346 more words)Show less
Step 3: Run make verify — the PRIMARY gate
bash
cd <model_dir>
flock -x -w 1800 /tmp/mlir-air-npu.lock make verify

Expected (per Phase 6 / verify subsystem design):

  • Both NPU and HF bf16 greedy-decode the prompt set × 32 tokens.
  • At the first divergence, NPU's chosen token ∈ HF top-5 AND HF's chosen token ∈ NPU top-5 → PASS; exit 0; no npu_vs_hf FAIL in the report.
  • The report written under the model's build dir (reports/, by the shared programming_examples/llms/verify/ runner) records the first divergence + top-5 sets.

Also run make diagnosis to capture the per-layer cosine table as an informational sanity signal — eyeball it for a gross cliff or a NaN layer. It is NOT a gate: the verify subsystem retired threshold-based diagnosis (compare_pair reports cosine with no pass/fail), so do not fail the evaluation on a cosine number. make verify (the token-set gate) is the sole binding numeric verdict; the cosine table just helps localize if verify fails.

If the gate fails: record exact failure + cite the divergence position and which token left the top-5. Verdict = FAIL.

Step 4: Write the evaluation report

Output: <model_dir>/docs/evaluation_report.md. Keep it short — its job is to let a human know, in 2 minutes, whether the deployment is trustworthy and why. Use the structure below:

markdown
# Evaluation Report: <Model> on NPU2

## Verdict: <PASS / PASS-with-warnings / FAIL>  (<date>)

## 1. Gate audit (anti-reward-hacking)
- `make verify` drives production `run_npu_prefill` / `run_npu_decode_step`
  via `<model>/verify_adapter.py` NpuRunner (not a mock): <yes/no + evidence>
- Reference is HF transformers bf16 (`torch_dtype=torch.bfloat16`): <yes/no>
- Gate is `compute_topk_set_check` (top-5 at first divergence), not a
  hardcoded pass / `> 0` threshold: <yes/no>
- Conclusion: gate is REAL / reward-hacked

## 2. Independent `make verify` re-run
- Result: PASS / FAIL — first divergence at token <i>; NPU & HF chosen
  tokens both in the other's top-5
- `make diagnosis` per-layer cosine (informational only): <X→Y; cliff/NaN?>

## 3. Manual reproduce
    cd <model_dir>
    flock -x -w 1800 /tmp/mlir-air-npu.lock make verify

Failure modes

SymptomLikely causeWhat to do
Step 2 audit fails — make verify doesn't drive the production path vs HF bf16Reward-hacked gate (mock NpuRunner, stale baseline, hardcoded pass)Report [FAIL] verify gate is reward-hacked; deployment needs remediation; do NOT mark PASS
make verify returns PASS suspiciously fast (<10 s for a full N-layer model)Token count too low, or NpuRunner not actually invoking kernelsCheck the shared verify_runner.py GATE_N_TOKENS + the model's NpuRunner generates 32 tokens through the real prefill/decode
Subagent reads LESSONS/progress before measuringSkill prompt wasn't strict enoughReject the report; re-spawn with stricter instructions

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

This is the terminal verification phase. On Phase 7 PASS or PASS-with-warnings:

  • <model_dir>/docs/evaluation_report.md is the durable artifact
  • Append to <model>/TODO.md: "Independently evaluated YYYY-MM-DD: <verdict>"
  • Reference the report from <model>/docs/development_progress/progress.md

If FAIL: deployment cannot be tagged. Issues must be remediated and Phase 7 re-run.

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-7-independent-evaluator of Xilinx/mlir-air.

Open the folder on GitHubat commit 416edae

Compare with similar skills

Phase 7 Independent Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 7 Independent Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 7 Independent Evaluator this skillXilinx/mlir-air150—~2.4kAutomated safety check: PassMIT
Olore Claude Code Latestolorehq/olore104—~901Automated safety check: PassMIT
Team TopologyCotal-AI/Cotal322—~2.9kAutomated safety check: PassApache-2.0
Agent OrchestrationLeoYeAI/openclaw-master-skills2.2k—~4.4kAutomated safety check: PassMIT
Files Memory SystemLeoYeAI/openclaw-master-skills2.2k—~3.8kAutomated safety check: PassMIT
Hitl Approvalnwiizo/ccswarm153—~290Automated safety check: PassMIT

Similar skills

  • Local claude-code documentation reference (latest). An agent skill from olorehq/olore.

    104 GitHub stars~901 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Team Topology

    Cotal-AI/Cotal

    Define a multi-agent team for ANY task on ANY system as an explicit deployment topology - pick the shape from the task's dominant risk, specify the runtime/communication/trust layers, place model…

    322 GitHub stars~2.9k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agent Orchestration

    LeoYeAI/openclaw-master-skills

    Multi-agent orchestration patterns for production deployments.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Hitl Approval

    nwiizo/ccswarm

    Human-in-the-loop approval workflow for high-risk agent operations (file deletions, deployments, config changes).

    153 GitHub stars~290 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Convex Self Hosting

    waynesutton/markdown-site

    Integrate Convex static self hosting into existing apps using the latest upstream instructions from get-convex/self-hosting every time.

    627 GitHub stars~1.4k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed

Questions about Phase 7 Independent Evaluator

What does Phase 7 Independent Evaluator do?

Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…. Phase 7 Independent Evaluator is an agent skill from Xilinx/mlir-air. Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the production path vs HF bf16), then re-runs it as the primary gate.

When should I use Phase 7 Independent Evaluator?

Phase 7 Independent Evaluator fits situations like: tasks that involve Subagents; tasks that involve Deployment.

How do I install Phase 7 Independent Evaluator in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a claude-code`. Or copy the skill folder (.claude/skills/phase-7-independent-evaluator in Xilinx/mlir-air) into .claude/skills/phase-7-independent-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Phase 7 Independent Evaluator in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a codex`. Or copy the skill folder (.claude/skills/phase-7-independent-evaluator in Xilinx/mlir-air) into .agents/skills/phase-7-independent-evaluator in your project. Codex loads it when a task matches its description.

Can I use Phase 7 Independent Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-7-independent-evaluator, .gemini/skills/phase-7-independent-evaluator, .github/skills/phase-7-independent-evaluator and .opencode/skills/phase-7-independent-evaluator in your project.

What does Phase 7 Independent Evaluator need to run?

Going by SKILL.md and its folder, Phase 7 Independent Evaluator needs the command-line tools its instructions call (make).

Does Phase 7 Independent Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phase 7 Independent Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 7 Independent Evaluator use?

Phase 7 Independent Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 7 Independent Evaluator use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 7 Independent Evaluator?

Skills that share tags, products or a category with Phase 7 Independent Evaluator: Olore Claude Code Latest (olorehq/olore, 104 stars), Team Topology (Cotal-AI/Cotal, 322 stars), Agent Orchestration (LeoYeAI/openclaw-master-skills, 2.2k stars) and Files Memory System (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 7 Independent Evaluator?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 10, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.