Agent skill

Phase 2 Single Block Validation

by Xilinx in Xilinx/mlir-air

Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…

MITAuto-check passedDevOps & Cloud

Install Phase 2 Single Block Validation

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-2-single-block-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-2-single-block-validation .claude/skills/phase-2-single-block-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-2-single-block-validation
GitHub stars
150
Token cost
~3.4k tokens
SKILL.md length
1,543 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…

  • Tasks that involve Deployment
  • SKILL.md covers Purpose, Phase 2 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls make

What it does

Phase 2 Single Block Validation is an agent skill from Xilinx/mlir-air. Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/ diagnosis lens, promoted to a gate at layer 0). Catches integration bugs (layout mismatches, missing transposes, type drops between kernel boundaries) before scaling to N layers.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. The licence is MIT.

When your agent uses it

  • Tasks that involve Deployment

Example prompts

  • “/phase-2-single-block-validation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 2 Single Block Validation loads about 3.4k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 1,543 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,543 words, ~3,372 tokens.

Download SKILL.mdSave it as .claude/skills/phase-2-single-block-validation/SKILL.md (or your agent's skills folder).
name
phase-2-single-block-validation
description
Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared `programming_examples/llms/verify/` diagnosis lens, promoted to a gate at layer 0). Catches integration bugs (layout mismatches, missing transposes, type drops between kernel boundaries) before scaling to N layers.

Purpose

Phase 1 verified each kernel × shape standalone. Phase 2 wires those kernels into a single transformer block on NPU and compares the block-level output to the HF bf16 reference at the same input. This catches integration bugs (wrong tensor layouts at kernel boundaries, dropped biases, missing padding) without the cost of running N layers.

The comparison reuses the verify/ subsystem's diagnosis lens: verify/verify_runner.py:_run_diagnosis already computes per-layer ffn_out cosine (NPU vs HF bf16) for every layer. Diagnosis itself is informational (no thresholds); Phase 2's job is to promote layer 0's result to a hard gate with a head_dim-scaled threshold. The reference is HF transformers in bf16 (same dtype as the NPU — fair fight), obtained via HfRunner in diagnosis mode (lite_mode=False), NOT a hand-written CPU forward.

Phase 2 PASS criteria (HARD GATES)

Run the diagnosis lens on the canonical prompt and gate on layer 0 (the single block under test). All four must hold:

  1. Whole-tensor cosine ≥ 0.99 between NPU layer-0 ffn_out and HF bf16 layer-0 ffn_out (hf.layer_intermediates[0]['ffn_out']). Catches: coarse integration breaks (wrong layout, missing op).

  2. Per-position cosine ≥ THRESHOLD(head_dim) at each real-token position ([:real_len], NOT padded positions — those are out-of-distribution and amplify BF16 noise unhelpfully). The diagnosis comparator already returns per-position cosine {min, p5, median, mean} via per_position_cosine + aggregate; gate on min. Catches: per-row dropouts the whole-tensor cosine averages over.

    head_dimper-position min
    ≤ 640.99
    1280.98
    ≥ 2560.97

    Threshold scales with head_dim because BF16 accumulation noise grows as √(head_dim · K): hd=64 deployments hit per-pos min ≈ 0.998; hd=128 hits ≈ 0.980 with ~5× LOWER MAE — the larger cosine drop is geometric, not a bug.

  3. No NaN anywhere in NPU output.

  4. Result documented in <model>/docs/development_progress/phase2_block.md (cosine numbers

    • tile/integration choice + any bisect findings if Step 4 fired).

Also record max_abs and max_rel error alongside the cosine numbers (informational, not gated — absolute thresholds depend on input distribution). The diagnosis comparator's error_metrics reports these. The current BF16-output GEMM production path is the registry's high-precision tier (FP32-accumulate + a single epilogue cast, fused-cast or drain) at ~9.3e-3 mean_rel_L1 — the GPU-standard accuracy, single- sourced from details/GEMM_bf16_in_bf16_out.json; the low-precision direct-codegen tier (1.3e-2–1.9e-2) is the fallback. Across 7 GEMMs + softmax + RoPE + RMSNorm the block stays in that high-precision band. Recording them gives future deployments a regression baseline (NPU max_abs ≤ 1.5× reference deployment's measured value at same shape signals no regression).

Why cosine here, and why it's only an interim gate. A 2026 survey of industry practice (vLLM, HF transformers, llama.cpp, MLPerf, TensorRT-LLM, the GPTQ/AWQ/SmoothQuant literature) found that per-layer activation cosine is NOT a standard correctness gate — everyone gates either end-to-end (vLLM token-set, llama.cpp logit-KL, MLPerf ≥99%-of-reference) or per-tensor on element-wise atol/rtol / SQNR vs an FP32 reference. The one stack that does gate on per-layer cosine (TPU-MLIR) uses 0.99, not a loose value. So treat this gate honestly:

  • At Phase 2 there is no token output yet to score, so we still need some numeric tripwire on the block output — cosine vs HF bf16 is that tripwire. It catches gross integration bugs (layout, missing transpose, dtype drop), which is all Phase 2 needs.
  • It is bf16-vs-bf16 (NPU bf16 vs HF bf16) — there is no FP32 ground truth at this layer, so a tight element-wise atol/rtol (HF's FP32-parity bar) would mis-fire on benign rounding-order differences; cosine's direction-only nature is why it tolerates that. That tolerance is also its weakness (it can pass on magnitude errors), which is exactly why it's interim, not the final word.
  • The real correctness gate is end-to-end (Phase 3/6 token-set top-5, which mirrors vLLM check_logprobs_close). Once the full model exists, that gate — not cosine — decides correctness. A future upgrade could add SQNR (≥40 dB, PyTorch Numeric Suite's "very good alignment" bar) as a more defensible per-block number; thresholds unchanged for now.

Knowledge base references

PRIMARY:

  • programming_examples/llms/llama_kernel_builder/ — the shared toolkit (KernelCache, stitching, external_kernels) you compose the block FROM (kernel-first default).
  • programming_examples/llms/llama32_1b/multi_launch_builder/ + llama32_1b_prefill.py:run_transformer_block — the reference exemplar: read to see how the leaf kernels stitch into a block. On a bit-for-bit kernel-sequence match you may call run_transformer_block directly (inheritance shortcut).
  • programming_examples/llms/verify/verify_runner.py:_run_diagnosis — the per-layer NPU-vs-HF-bf16 cosine lens Phase 2 promotes to a gate.
  • programming_examples/llms/verify/runners/hf_runner.py — how the HF bf16 reference exposes per-layer ffn_out (lite_mode=False).
  • <model>/docs/development_progress/ (Phase 1 output) + programming_examples/kernel_registry/supported_kernels.md rows with Used by = <model> — Phase 1's verified (kernel, shape) list; Phase 2 must wire ALL of them.
  • programming_examples/kernel_registry/details/<Kernel>_bf16.md — kernel-by-kernel reference (datapath, tile rules, constraints, layouts).

WORKAROUNDS (apply when model config triggers them — re-derive from the HF reference impl; the patterns below describe the technique, not a shipped file):

  • GQA-aware reindexed padding for non-1024-aligned dims (see Step 2)
  • Host-side post-RoPE bias add for QKV-bias models (see Step 2)

Workflow

Step 1: Choose integration path

Kernel-first (default). Derive the model's per-layer kernel sequence from its config and build the block by composing the registry leaf kernels (verified in Phase 1) into model-specific multi-launch ELFs under <model>/multi_launch_builder/, using the shared llama_kernel_builder toolkit (KernelCache, stitching, external_kernels). This is the general path — it does not assume the model resembles llama, so it generalizes to any decoder-only architecture in scope.

Read llama32_1b's assembly (llama32_1b_prefill.run_transformer_block and llama32_1b/multi_launch_builder/*) as a worked exemplar of how the leaf kernels stitch into a block — mirror its structure, adapting the kernel sequence and shapes to your model.

Inheritance (shortcut). ONLY when the model's per-layer kernel sequence matches llama's bit-for-bit — RMSNorm → Q/K/V GEMM → RoPE → FA → O → add → RMSNorm → Gate/Up → SwiGLU → Down → add — you may skip writing builders and call llama32_1b_prefill.run_transformer_block directly with the new shape parameters. This is an optimization for genuine llama variants, not the starting assumption. Any of these breaks the bit-for-bit match and forces the kernel-first path:

(a) NEW op type (e.g., Qwen3's Q/K Norm — per-head RMSNorm with (head_dim,) weight) (b) NEW op needs to land BETWEEN currently-fused launches (e.g., Q/K Norm sits between Q/K projection and RoPE, but rms_gemv_rope fuses both) (c) Op REORDER (post-norm vs pre-norm)

Either way, don't write new C kernels speculatively — almost always the leaf kernel exists in the registry; the trick is the right way to STITCH. For Q/K Norm specifically, weighted_rms_norm with the heads-as-M trick (M=n_heads, N=head_dim, sharing the (head_dim,) weight across rows) IS the op.

Show full SKILL.md (584 more words)Show less
Step 2: Apply config-specific prereqs (only if model needs them)

Two known triggers from model config (NOT from upstream phases):

Non-1024-aligned emb_dim or hidden_dim → BD pool exhaustion risk at long seq (see the kernel's details/<Kernel>_bf16.md placeability notes). Use GQA-aware reindexed padding: pad up to a 1024-aligned multiple by inserting phantom Q heads INSIDE each KV group (not at the end — naive padding breaks GQA semantics by changing n_heads / n_kv_heads = group_size). CPU-only sanity test the padded vs orig forward FIRST (cosine should be 0.999998+) before touching NPU.

qkv_bias=True (Qwen2 / Qwen3 family) → host-side post-RoPE bias add, exploiting RoPE's linearity: RoPE(q + bq) = RoPE(q) + RoPE(bq). The rms_gemms_rope ELF stays bias-free; bias is added on host after the ELF returns.

If both: padding determines the n_heads count the bias precompute uses.

Step 3: Wire one block + numerical check

In <model>/<model>_prefill.py, implement run_single_block(layer_idx=0, hidden, weights, ...):

  • Kernel-first path (default): call your new run_transformer_block_<model>(...) that runs the per-model multi-launch ELFs in order via the shared KernelCache. Minimal skeleton:
    python
    from llama_kernel_builder.cache import KernelCache
    cache = KernelCache()                      # compile-once, run-many
    def run_transformer_block_<model>(hidden, weights, cfg, cache):
        # one _run_cached per fused ELF you built in Step 1, in order:
        x = cache._run_cached("rms_qkv_rope", hidden, weights.qkv, ...)   # RMSNorm+Q/K/V+RoPE
        x = cache._run_cached("attn",         x, ...)                     # FA
        x = cache._run_cached("o_ffn",        x, weights.o, weights.ffn, ...)  # O+add+RMSNorm+SwiGLU+Down+add
        return x
  • Inheritance path (shortcut, bit-for-bit match only): call llama32_1b_prefill.run_transformer_block(...) with this model's shape parameters

Use the canonical prompt from the deployment's verify prompt set (verify/prompts/{base,instruct}.txt). Get the layer-0 reference from the HF bf16 runner and compare:

python
from verify.runners.hf_runner import HfRunner
hf = HfRunner(hf_model_id, config, max_seq, lite_mode=False)
hf_pf = hf.prefill(prompt_tokens)
ref_block0 = hf_pf.layer_intermediates[0]["ffn_out"]   # HF bf16, layer-0 output

npu_block0 = run_single_block(layer_idx=0, hidden=x, weights=weights, ...)

Compute whole-tensor cosine and per-position cosines (real-token positions only) with the diagnosis comparator (verify/comparators.py:per_position_cosine + aggregate). Check against the PASS criteria above. The simplest route is to run make diagnosis and read layer 0's row from the report; the manual snippet above is for when you need to gate inside a Phase-2 test script.

Step 4: Bisect on FAIL

If cosine fails, the integration is broken at one specific kernel boundary. Bisect by swapping NPU kernels back to a CPU equivalent one at a time (use <model>_cpu_helpers.py for the ops that have a helper — rms_norm, attention_reference — and a small inline numpy for the rest): walk forward through the block, replacing npu_<kernel>(...) with the CPU equivalent, recompute cosine. The first replacement that pushes cosine above threshold identifies the offender — that's where the layout / type / argument mismatch lives. Invoke superpowers:systematic-debugging on it.

Record the bisect table (per-step cosine) in phase2_block.md so future deployments learn from this specific failure.

Failure modes

SymptomLikely causeWhere to look
Cosine drops at Q/K/V GEMMweight loading / tensor layout (seq-first vs heads-first)Compare NPU output shape to reference's; check np.ascontiguousarray() after weight load
Cosine drops at FlashAttentioncausal masking missing / wrong dk_chunks compile flagSee debug-fa-runtime-failure
Cosine drops at Down GEMMBF16 truncation; running the low-precision (direct-codegen) tier instead of high-precisionconfirm the GEMM uses the registry's high-precision path (fused-cast / drain = FP32-accumulate + single cast), not --high-precision false; see details/GEMM_bf16_in_bf16_out.md
NaN in outputuninitialized BO / reused stale bufferInvoke debug-bo-corruption
Cosine drops at residual addbias forgotten on padded path / GQA reindex bugIf padding+bias model: re-run CPU sanity test on padded forward (Step 2)
Whole-tensor cosine OK but per-position min lowone bad position run; check whether last few positions diverge (causal mask edge case)Print per-position cosine, look for contiguous bad runs

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

On Phase 2 PASS:

  • Append cosine + per-position min + integration-path choice to <model>/docs/development_progress/phase2_block.md
  • Mark Phase 2 in <model>/TODO.md
  • If Step 2 padding or bias workarounds were used, surface as a Phase 4/5 prerequisite ("perf optimization must preserve the padded/bias wrappers")

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-2-single-block-validation of Xilinx/mlir-air.

Open the folder on GitHubat commit 6e81ce1

Compare with similar skills

Phase 2 Single Block Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 2 Single Block Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 2 Single Block Validation this skillXilinx/mlir-air150—~3.4kAutomated safety check: PassMIT
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb6.7k—~4kAutomated safety check: NotesApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Vercelremotion-dev/remotion62k—~1.2kAutomated safety check: PassCustom licence
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • GreptimeDB Dev Docker Image

    GreptimeTeam/greptimedb

    Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.

    6.7k GitHub stars~4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Vercel

    remotion-dev/remotion

    Official

    Set up a Codex monitor for Vercel deployments and preview URLs.

    62k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Categories

Questions about Phase 2 Single Block Validation

What does Phase 2 Single Block Validation do?

Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…. Phase 2 Single Block Validation is an agent skill from Xilinx/mlir-air. Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/ diagnosis lens, promoted to a gate at layer 0).

When should I use Phase 2 Single Block Validation?

Phase 2 Single Block Validation fits situations like: tasks that involve Deployment.

How do I install Phase 2 Single Block Validation in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-2-single-block-validation in Xilinx/mlir-air) into .claude/skills/phase-2-single-block-validation in your project. Claude Code loads it when a task matches its description.

How do I install Phase 2 Single Block Validation in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a codex`. Or copy the skill folder (.claude/skills/phase-2-single-block-validation in Xilinx/mlir-air) into .agents/skills/phase-2-single-block-validation in your project. Codex loads it when a task matches its description.

Can I use Phase 2 Single Block Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-2-single-block-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-2-single-block-validation, .gemini/skills/phase-2-single-block-validation, .github/skills/phase-2-single-block-validation and .opencode/skills/phase-2-single-block-validation in your project.

What does Phase 2 Single Block Validation need to run?

Going by SKILL.md and its folder, Phase 2 Single Block Validation needs the command-line tools its instructions call (make). Our summary lists: Python 3.

Does Phase 2 Single Block Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phase 2 Single Block Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 2 Single Block Validation use?

Phase 2 Single Block Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 2 Single Block Validation use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 2 Single Block Validation?

Skills that share tags, products or a category with Phase 2 Single Block Validation: Kubeshark Installer (kubeshark/kubeshark, 12k stars), GreptimeDB Dev Docker Image (GreptimeTeam/greptimedb, 6.7k stars), KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars) and Vercel (remotion-dev/remotion, 62k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 2 Single Block Validation?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.