Olore Claude Code Latest
olorehq/olore
Local claude-code documentation reference (latest). An agent skill from olorehq/olore.
Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .claude/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .claude/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .agents/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .agents/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .cursor/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .cursor/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/phase-7-independent-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .gemini/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .gemini/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .github/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .github/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air phase-7-independent-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/phase-7-independent-evaluator .opencode/skills/phase-7-independent-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "phase-7-independent-evaluator" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/phase-7-independent-evaluator into .opencode/skills/phase-7-independent-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "phase-7-independent-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
phase-7-independent-evaluatorPhase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…
Phase 7 Independent Evaluator is an agent skill from Xilinx/mlir-air. Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the production path vs HF bf16), then re-runs it as the primary gate. Produces a structured evaluationreport.md a human can read in 2 minutes to know the full deployment state. Invoke as /phase-7-independent-evaluator <modeldir or auto-spawn from deploy-new-llm after Phase 6 PASS.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Subagents and Deployment. The licence is MIT.
Read from SKILL.md and the folder at commit 416edae. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Phase 7 Independent Evaluator loads about 2.4k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,005 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit 416edae, republished under its MIT licence (© Xilinx). 1,005 words, ~2,361 tokens.
.claude/skills/phase-7-independent-evaluator/SKILL.md (or your agent's skills folder).The deploy-new-llm chain is autonomous. Phase 6 wires up a make verify
gate that's supposed to compare NPU vs the HF bf16 reference via the
verify/ subsystem's top-k token-set check. But the deployment agent
wrote (or copied) both the gate AND the production code it exercises —
nothing structurally prevents the gate from being mocked, pointed at a
stale baseline, or wired to a verify-only code path that bypasses the
real kernels. Phase 7 closes this by:
make verify code path (anti-reward-hacking) —
confirming the gate actually drives the production kernels vs HF bf16
and gates on the token-set check, not a mocked / stale / verify-only
path. This is what makes the independent re-run meaningful: without it,
"the agent re-ran the gate it wrote itself" proves nothing.make verify independently — the shared
verify/ top-k token-set check (vLLM-aligned) of NPU vs HF bf16 on the
real production prefill/decode path. A clean independent PASS, on a gate
the audit confirmed is real, is the verdict.The deployment is NOT considered trustworthy until this report is written and its overall verdict is PASS or PASS-with-warnings.
make verify audited (anti-reward-hacking): subagent has read the
Makefile target + the model's verify_adapter.py + the shared
programming_examples/llms/verify/{verify_runner,comparators,runners/hf_runner}.py, and
confirmed the gate (a) drives the model's PRODUCTION run_npu_prefill /
run_npu_decode_step via the NpuRunner (not a mock), (b) compares
against HF transformers in bf16 (torch_dtype=torch.bfloat16), (c)
gates on compute_topk_set_check (top-5 set inclusion at first
divergence), not a hardcoded return True / > 0 threshold.make verify PASSES under the fresh subagent run: token-set gate
exits 0 (no npu_vs_hf FAIL), on the audited-real gate.If the subagent declares PASS without showing measured numbers OR without auditing the verify code path, the report is REJECTED — re-spawn with stricter instructions.
programming_examples/llms/verify/README.md — the verify
methodology the audit checks against (HF bf16 reference, top-k
token-set gate, cosine-as-diagnosis)programming_examples/llms/verify/comparators.py — the gate
implementation (compute_topk_set_check) the audit confirms is real<model>/docs/development_progress/ + the kernel_registry rows with
Used by = <model>
— Phase 1 catalog (subagent compares its measurements against this)programming_examples/kernel_registry/supported_kernels.md
— kernel-by-kernel ground truthUse the general-purpose Agent type. Critical instructions in the
spawn prompt:
<model>/docs/development_progress/{LESSONS,progress,phaseN_*}.md
BEFORE forming your own measurements. You may CITE them AFTER
measuring (compare your numbers vs claimed).make verify really drive the production prefill/decode vs HF bf16
and gate on the token-set check, or did the agent shortcut to "if NPU
output exists → PASS" / point the verify_adapter NpuRunner at a mock /
compare against a stale cached baseline?make verify — anti-reward-hackingBEFORE running anything, READ:
make verify invoke?grep -A 5 "^verify:" <model_dir>/Makefile../verify/verify_runner.py --runner=<model>.verify_adapter --prompts topk_token ... — not a
model-local copy of the runner.<model_dir>/verify_adapter.py: this is the model's only verify
code. Confirm its NpuRunner (the .prefill() / .decode_step()
methods) imports and calls THIS model's production run_npu_prefill /
run_npu_decode_step from <model>_inference.py — NOT a stub, NOT the
llama32_1b copy. (The runner/comparators/report/HF runner live once in
the shared programming_examples/llms/verify/; the adapter is the per-model hook.)programming_examples/llms/verify/runners/hf_runner.py (shared): confirm the reference
loads with AutoModelForCausalLM.from_pretrained(..., torch_dtype=torch.bfloat16).programming_examples/llms/verify/comparators.py + programming_examples/llms/verify/report.py (shared):
confirm the gate is compute_topk_set_check (top-5 set inclusion at
first divergence) and report.has_failure() drives a real exit code —
not a hardcoded return True / > 0 threshold.[FAIL] verify gate is reward-hacked and stop here. Tag the deployment as needs-remediation.make verify — the PRIMARY gatecd <model_dir>
flock -x -w 1800 /tmp/mlir-air-npu.lock make verifyExpected (per Phase 6 / verify subsystem design):
npu_vs_hf FAIL in the report.reports/, by the
shared programming_examples/llms/verify/ runner) records the first divergence + top-5 sets.Also run make diagnosis to capture the per-layer cosine table as an
informational sanity signal — eyeball it for a gross cliff or a NaN
layer. It is NOT a gate: the verify subsystem retired threshold-based
diagnosis (compare_pair reports cosine with no pass/fail), so do not
fail the evaluation on a cosine number. make verify (the token-set gate)
is the sole binding numeric verdict; the cosine table just helps localize
if verify fails.
If the gate fails: record exact failure + cite the divergence position and which token left the top-5. Verdict = FAIL.
Output: <model_dir>/docs/evaluation_report.md. Keep it short — its job is
to let a human know, in 2 minutes, whether the deployment is trustworthy
and why. Use the structure below:
# Evaluation Report: <Model> on NPU2
## Verdict: <PASS / PASS-with-warnings / FAIL> (<date>)
## 1. Gate audit (anti-reward-hacking)
- `make verify` drives production `run_npu_prefill` / `run_npu_decode_step`
via `<model>/verify_adapter.py` NpuRunner (not a mock): <yes/no + evidence>
- Reference is HF transformers bf16 (`torch_dtype=torch.bfloat16`): <yes/no>
- Gate is `compute_topk_set_check` (top-5 at first divergence), not a
hardcoded pass / `> 0` threshold: <yes/no>
- Conclusion: gate is REAL / reward-hacked
## 2. Independent `make verify` re-run
- Result: PASS / FAIL — first divergence at token <i>; NPU & HF chosen
tokens both in the other's top-5
- `make diagnosis` per-layer cosine (informational only): <X→Y; cliff/NaN?>
## 3. Manual reproduce
cd <model_dir>
flock -x -w 1800 /tmp/mlir-air-npu.lock make verify| Symptom | Likely cause | What to do |
|---|---|---|
Step 2 audit fails — make verify doesn't drive the production path vs HF bf16 | Reward-hacked gate (mock NpuRunner, stale baseline, hardcoded pass) | Report [FAIL] verify gate is reward-hacked; deployment needs remediation; do NOT mark PASS |
make verify returns PASS suspiciously fast (<10 s for a full N-layer model) | Token count too low, or NpuRunner not actually invoking kernels | Check the shared verify_runner.py GATE_N_TOKENS + the model's NpuRunner generates 32 tokens through the real prefill/decode |
| Subagent reads LESSONS/progress before measuring | Skill prompt wasn't strict enough | Reject the report; re-spawn with stricter instructions |
For any failure not in the table, invoke superpowers:systematic-debugging.
This is the terminal verification phase. On Phase 7 PASS or PASS-with-warnings:
<model_dir>/docs/evaluation_report.md is the durable artifact<model>/TODO.md: "Independently evaluated YYYY-MM-DD: <verdict>"<model>/docs/development_progress/progress.mdIf FAIL: deployment cannot be tagged. Issues must be remediated and Phase 7 re-run.
© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/phase-7-independent-evaluator of Xilinx/mlir-air.
Open the folder on GitHubat commit 416edae
Phase 7 Independent Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Phase 7 Independent Evaluator this skillXilinx/mlir-air | 150 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Olore Claude Code Latestolorehq/olore | 104 | — | ~901 | Automated safety check: Pass | MIT | |
| Team TopologyCotal-AI/Cotal | 322 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Agent OrchestrationLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.4k | Automated safety check: Pass | MIT | |
| Files Memory SystemLeoYeAI/openclaw-master-skills | 2.2k | — | ~3.8k | Automated safety check: Pass | MIT | |
| Hitl Approvalnwiizo/ccswarm | 153 | — | ~290 | Automated safety check: Pass | MIT |
olorehq/olore
Local claude-code documentation reference (latest). An agent skill from olorehq/olore.
Cotal-AI/Cotal
Define a multi-agent team for ANY task on ANY system as an explicit deployment topology - pick the shape from the task's dominant risk, specify the runtime/communication/trust layers, place model…
LeoYeAI/openclaw-master-skills
Multi-agent orchestration patterns for production deployments.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
nwiizo/ccswarm
Human-in-the-loop approval workflow for high-risk agent operations (file deletions, deployments, config changes).
waynesutton/markdown-site
Integrate Convex static self hosting into existing apps using the latest upstream instructions from get-convex/self-hosting every time.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Entry point for deploying a new decoder-only LLM on AMD NPU2.
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Categories
Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…. Phase 7 Independent Evaluator is an agent skill from Xilinx/mlir-air. Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the production path vs HF bf16), then re-runs it as the primary gate.
Phase 7 Independent Evaluator fits situations like: tasks that involve Subagents; tasks that involve Deployment.
Run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a claude-code`. Or copy the skill folder (.claude/skills/phase-7-independent-evaluator in Xilinx/mlir-air) into .claude/skills/phase-7-independent-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a codex`. Or copy the skill folder (.claude/skills/phase-7-independent-evaluator in Xilinx/mlir-air) into .agents/skills/phase-7-independent-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-7-independent-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-7-independent-evaluator, .gemini/skills/phase-7-independent-evaluator, .github/skills/phase-7-independent-evaluator and .opencode/skills/phase-7-independent-evaluator in your project.
Going by SKILL.md and its folder, Phase 7 Independent Evaluator needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Phase 7 Independent Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Phase 7 Independent Evaluator: Olore Claude Code Latest (olorehq/olore, 104 stars), Team Topology (Cotal-AI/Cotal, 322 stars), Agent Orchestration (LeoYeAI/openclaw-master-skills, 2.2k stars) and Files Memory System (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 10, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.