Agent skill

Phase 0 Build Cpu Reference

by Xilinx in Xilinx/mlir-air

Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline…

MITAuto-check passed

Install Phase 0 Build Cpu Reference

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-0-build-cpu-reference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-0-build-cpu-reference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-0-build-cpu-reference .claude/skills/phase-0-build-cpu-reference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-0-build-cpu-reference
GitHub stars
150
Token cost
~2.8k tokens
SKILL.md length
1,235 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline…

  • SKILL.md covers Purpose, Phase 0 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls pip and huggingface-cli

What it does

Phase 0 Build Cpu Reference is an agent skill from Xilinx/mlir-air. Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline loads and runs via the shared programmingexamples/llms/verify/ subsystem's HfRunner. Downstream phases compare NPU against HF transformers in bf16 directly; there is no hand-written full-model FP32 oracle.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with NumPy. The licence is MIT.

Example prompts

  • “/phase-0-build-cpu-reference”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 0 Build Cpu Reference loads about 2.8k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 1,235 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 1,235 words, ~2,751 tokens.

Download SKILL.mdSave it as .claude/skills/phase-0-build-cpu-reference/SKILL.md (or your agent's skills folder).
name
phase-0-build-cpu-reference
description
Phase 0 of LLM deployment — produce `<model>_weights.py` (HF weight loader) and `<model>_cpu_helpers.py` (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline loads and runs via the shared `programming_examples/llms/verify/` subsystem's HfRunner. Downstream phases compare NPU against HF transformers in bf16 directly; there is no hand-written full-model FP32 oracle.

Purpose

Phase 0 establishes the inputs every downstream phase needs. It does NOT build a full-model CPU oracle — the reference is HuggingFace transformers in bf16, accessed through the verify/ subsystem's HfRunner. Phase 0 produces three things:

  1. <model>_weights.py — Config dataclass + HF weight loader. Maps HF safetensors names to the per-layer LayerWeights the NPU pipeline consumes.
  2. <model>_cpu_helpers.py — the small set of NumPy helpers that production prefill/decode import (default: rms_norm, attention_reference, softmax). These are NOT a per-kernel oracle catalog — each leaf kernel ships its own NumPy reference inside its llama_kernel_builder/<kernel>/run.py harness (Phase 1 uses those).
  3. HF bf16 baseline confirmation — verify the target model loads with torch_dtype=torch.bfloat16 and runs the canonical prompt through HfRunner, producing a sane top-1 and non-degenerate logits. This is the Phase 0 gate.
Why HF bf16 directly, and what still needs a NumPy helper

The reference is HF transformers in bf16 — same dtype as the NPU, so NPU-vs-reference is a fair fight (bf16 vs bf16), and there is no hand-written 480-line full-model forward to keep correct. HF's per-layer hidden states (used by Phase 2/3) are captured by HfRunner in diagnosis mode (lite_mode=False), which returns layer_intermediates[].ffn_out plus final_hidden_normed.

NumPy helpers are still required for two narrow cases that HF (a black box that only exposes end-to-end forward + per-layer hidden states) cannot serve:

  1. Per-kernel references (Phase 1): HF cannot expose a single kernel's output (RMSNorm alone, RoPE alone, etc.). Each leaf kernel's standalone harness (llama_kernel_builder/<kernel>/run.py) carries its own NumPy F32 reference for that kernel. <model>_cpu_helpers.py only holds the helpers production code itself imports at runtime.
  2. CPU fallbacks (production): e.g. prefill cpu_attn=True uses attention_reference when the NPU FlashAttention kernel is unavailable for the configured head_dim; the LM-head final norm uses rms_norm.

Phase 0 PASS criteria (HARD GATES)

Three checks, each catching a different bug class:

  1. Weights loadable (catches weight-name / shape mismatches): <model>_weights.py loads every expected tensor; every layer index 0..n_layers-1 has q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj, attn_norm, ffn_norm (plus q/k/v bias if qkv_bias=true).
  2. HF baseline runnable (catches HF auth / config / dtype issues): HfRunner(model_name, config, max_seq, lite_mode=True).prefill(tokens) on the canonical prompt runs without error, returns a sane top-1 token, and produces logits with no NaN and a non-degenerate distribution.
  3. Config consistency (catches config-extraction errors): every field extracted in Step 1 (n_layers, emb_dim, n_heads, n_kv_heads, head_dim, hidden_dim, vocab_size, rope_theta, tie_word_embeddings) matches HF's config.json exactly.

The per-layer and final-logits cosine checks that used to live in Phase 0 now live in Phase 2 (single-block) and Phase 3 (full-model), comparing NPU against HF bf16 intermediates via the verify/ diagnosis path. Phase 0 only confirms the HF baseline and helpers/weights are in place.

Knowledge base references

Read these BEFORE acting:

  • programming_examples/llms/llama32_1b/llama32_1b_weights.py — reference Config dataclass + HF weight loading pattern to copy.
  • programming_examples/llms/llama32_1b/llama32_1b_cpu_helpers.py — the canonical small NumPy helper file; mirror its scope (only production-imported helpers), not a per-kernel catalog.
  • programming_examples/llms/verify/runners/hf_runner.py — the bf16 HF reference runner; this is how Phase 0 confirms the baseline and how Phase 2/3 obtain per-layer intermediates.
  • programming_examples/llms/verify/README.md — the verify subsystem methodology (HF bf16 reference, top-k token-set gate, cosine as diagnosis).

Workflow

Step 1: Read the HF config

Fetch config.json for the target model from HuggingFace. Extract:

  • num_hidden_layers → n_layers
  • hidden_size → emb_dim
  • num_attention_heads → n_heads
  • num_key_value_heads → n_kv_heads (default to n_heads if absent → MHA)
  • intermediate_size → hidden_dim
  • vocab_size → vocab_size
  • rope_theta → rope_base (default 10000.0 if absent)
  • head_dim (compute as emb_dim // n_heads if absent)
  • tie_word_embeddings (affects whether to load lm_head.weight)
Step 2: Architecture compatibility check

Confirm the model is in-scope before scaffolding anything downstream:

  • Architecture must be in ["LlamaForCausalLM", "MistralForCausalLM" (only if no sliding window), "Qwen2ForCausalLM", "Qwen3ForCausalLM"] — i.e., a decoder-only with RMSNorm + SwiGLU + RoPE + GQA/MHA
  • Reject if: MoE layers, sliding-window attention, MLA, encoder-decoder
  • Reject explicitly with a clear message; don't scaffold a model the rest of the pipeline can't handle
Step 3: Generate <model>_weights.py

Copy programming_examples/llms/llama32_1b/llama32_1b_weights.py to <model>/<model>_weights.py. Modify:

  • LlamaConfig dataclass defaults → match Step 1 values
  • HF weight name remapping in load_weights() — most LLAMA-derived models share the same names (model.layers.<i>.self_attn.q_proj.weight etc.); confirm via inspecting the safetensors index. If different, write an explicit mapping.
  • generate_rope_lut() — verify rope_base is parameterized and uses the new value
  • LM head: load lm_head.weight only if tie_word_embeddings is False
  • If qkv_bias=true (Qwen2 / Qwen3 family): add bq, bk, bv fields to LayerWeights, parallel to wq/wk/wv. The bias is loaded just like the projection weights but with a different HF key (q_proj.bias etc.). The bias is applied on the HOST around the bias-free NPU kernels (exploiting RoPE linearity: RoPE(q + bq) = RoPE(q) + RoPE(bq)); the application detail belongs to Phase 2 (phase-2-single-block-validation Step 2) — surface it in TODO.md as a Phase 2 prerequisite. Re-derive the bias-on-host wrapper from the HF reference impl.
Show full SKILL.md (477 more words)Show less
Step 4: Generate <model>_cpu_helpers.py

Copy programming_examples/llms/llama32_1b/llama32_1b_cpu_helpers.py to <model>/<model>_cpu_helpers.py. Keep ONLY the helpers your model's production prefill/decode actually import:

  • rms_norm — almost always needed (LM-head final norm in inference).
  • attention_reference — needed if prefill exposes a cpu_attn=True fallback (GQA attention in F32 on host).
  • softmax — keep only if attention_reference is kept.

Rule for what belongs here: a helper goes in <model>_cpu_helpers.py ONLY if production code imports it at runtime. Per-kernel verification references do NOT go here — they live in each llama_kernel_builder/<kernel>/run.py. If a later phase (4/5) promotes a new CPU fallback op, add its helper then, not preemptively.

Modify:

  • Imports / config references → use the new model's names and config.
  • For unfamiliar architectures (different attention masking, post-norm vs pre-norm) — adapt attention_reference carefully and cross-check against HF's modeling_<arch>.py for the exact computation order.
Step 5: Confirm the HF bf16 baseline (HARD GATE)

Confirm the reference baseline using the verify/ subsystem's HfRunner, NOT a hand-written full-model comparison. Minimal confirmation:

python
# from programming_examples/llms/<model>/, with the shared programming_examples/llms/verify/ reachable via verify_adapter
from verify.runners.hf_runner import HfRunner
from <model>_weights import LlamaConfig  # or your Config class

config = LlamaConfig()                     # Step-1 values
runner = HfRunner(hf_model_id, config, max_seq=64, lite_mode=True)
rec = runner.prefill(tokenize(canonical_prompt))
# Assert: rec.top1_token is a sane token; rec.logits_at_pred has no NaN;
# config fields match HF config.json.

Canonical prompt: use one of the verify/prompts/{base,instruct}.txt prompts (keep Phase 0 consistent with the gate the later phases run). For a base model use base.txt; for an instruct model use instruct.txt.

PASS = all three §"Phase 0 PASS criteria" gates hold. If any fails, see "Failure modes".

(The per-layer cosine sanity that HF can provide via lite_mode=False is not run here — Phase 2/3 own that comparison against NPU output.)

Failure modes

SymptomLikely causeWhere to look
HF model won't loadMissing transformers/torch, or HF auth needed for gated modelpip install -r requirements.txt; for gated models huggingface-cli login
Weights load raises KeyError / shape mismatchHF weight names differ from the llama32_1b remap, or head_dim/n_kv_heads wrongInspect the safetensors index; print weight shapes after load; re-check Step 1 config
Config field mismatch vs HF config.jsonStep 1 extraction error (e.g. head_dim defaulted wrong, tie_word_embeddings missed)Re-read config.json; do not assume defaults
HfRunner.prefill returns NaN logitsdtype/precision issue in HF load, or a corrupt downloadRe-download the HF snapshot; confirm torch_dtype=bfloat16 path
top-1 token looks nonsensicaltokenizer mismatch (wrong chat template / BOS handling)Confirm tokenizer matches the model; check base vs instruct prompt set

For any failure not in the table, invoke superpowers:systematic-debugging.

(Per-layer cosine drops — RoPE convention, norm order, KV layout — are no longer a Phase 0 concern; they surface in Phase 2/3 when NPU output is compared against HF bf16 intermediates. The debug recipes live there.)

Update protocol

On Phase 0 PASS, append a brief Phase 0 entry to <model>/docs/development_progress/progress.md recording: HF model id, resolved config (n_layers / emb_dim / n_heads / n_kv_heads / head_dim / hidden_dim / vocab / rope_base / tie_word_embeddings), which cpu_helpers were kept, and the HF baseline confirmation result (top-1 token on the canonical prompt, logits OK).

Mark Phase 0 in <model>/TODO.md.

<model>_weights.py and <model>_cpu_helpers.py are now the stable inputs for Phases 1-6. The correctness oracle for downstream cosine/token checks is HF transformers bf16 via the verify/ subsystem — there is no hand-written reference file to keep in sync.

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-0-build-cpu-reference of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Phase 0 Build Cpu Reference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 0 Build Cpu Reference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 0 Build Cpu Reference this skillXilinx/mlir-air150—~2.8kAutomated safety check: PassMIT
Tushare Datazillionare/zillionare3182 repos~2.3kAutomated safety check: PassNone
FAISS Similarity SearchOrchestra-Research/AI-Research-SKILLs13k7 repos~1.3kAutomated safety check: PassMIT
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Python Performance Optimizationwshobson/agents40k12 repos~814Automated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2052 repos~3.4kAutomated safety check: NotesMIT

Similar skills

  • Tushare Data

    zillionare/zillionare

    面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。

    318 GitHub starsUsed in 2 repos~2.3k tokens
    Business, Finance & HRAuto-check passed
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 7 repos~1.3k tokens
    DatabasesAuto-check passed
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 12 repos~814 tokens
    DevelopmentAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    205 GitHub starsUsed in 2 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • UAV Trajectory Overlay from Video

    XXLiu-HNU/visualize_uav_trajectory

    Composites several moments from real drone footage into one still with ghost trails, then lays out paper figures and an editable PowerPoint file.

    242 GitHub stars~535 tokensUpdated 11 days ago
    Media & CreativeAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Questions about Phase 0 Build Cpu Reference

What does Phase 0 Build Cpu Reference do?

Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline…. Phase 0 Build Cpu Reference is an agent skill from Xilinx/mlir-air.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline loads and runs via the shared programmingexamples/llms/verify/ subsystem's HfRunner.

How do I install Phase 0 Build Cpu Reference in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-0-build-cpu-reference -a claude-code`. Or copy the skill folder (.claude/skills/phase-0-build-cpu-reference in Xilinx/mlir-air) into .claude/skills/phase-0-build-cpu-reference in your project. Claude Code loads it when a task matches its description.

How do I install Phase 0 Build Cpu Reference in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-0-build-cpu-reference -a codex`. Or copy the skill folder (.claude/skills/phase-0-build-cpu-reference in Xilinx/mlir-air) into .agents/skills/phase-0-build-cpu-reference in your project. Codex loads it when a task matches its description.

Can I use Phase 0 Build Cpu Reference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-0-build-cpu-reference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-0-build-cpu-reference, .gemini/skills/phase-0-build-cpu-reference, .github/skills/phase-0-build-cpu-reference and .opencode/skills/phase-0-build-cpu-reference in your project.

What does Phase 0 Build Cpu Reference need to run?

Going by SKILL.md and its folder, Phase 0 Build Cpu Reference needs the command-line tools its instructions call (pip and huggingface-cli). Our summary lists: Python 3.

Does Phase 0 Build Cpu Reference access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Phase 0 Build Cpu Reference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 0 Build Cpu Reference use?

Phase 0 Build Cpu Reference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 0 Build Cpu Reference use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 0 Build Cpu Reference?

Skills that share tags, products or a category with Phase 0 Build Cpu Reference: Tushare Data (zillionare/zillionare, 318 stars), FAISS Similarity Search (Orchestra-Research/AI-Research-SKILLs, 13k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Python Performance Optimization (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 0 Build Cpu Reference?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.