Agent skill

Phase 1 Kernel Validation

by Xilinx in Xilinx/mlir-air

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

MITAuto-check passedAI & LLM Engineering

Install Phase 1 Kernel Validation

skills CLI
$ npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air phase-1-kernel-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phase-1-kernel-validation .claude/skills/phase-1-kernel-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phase-1-kernel-validation
GitHub stars
150
Token cost
~3.4k tokens
SKILL.md length
1,561 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers Purpose, Phase 1 PASS criteria (HARD…, Knowledge base references and Workflow, plus 2 more sections
  • Calls make
  • Tasks that involve Deployment

What it does

Phase 1 Kernel Validation is an agent skill from Xilinx/mlir-air. Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. Primary gate where a standalone harness exists: the harness's full-output element-wise np.isclose check at that kernel's rtol/atol vs an FP32 reference (the same PASS/FAIL make run prints). Fallback for a kernel with no harness: make diagnosis per-layer cosine vs the HF bf16 reference. Record each verified (kernel, shape) as a row in that kernel's…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Deployment. It works with vLLM. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Deployment

Example prompts

  • “s full-output element-wise np.isclose check at that kernel”
  • “tested shapes”
  • “/phase-1-kernel-validation”

What it can do on your machine

Read from SKILL.md and the folder at commit 6e81ce1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phase 1 Kernel Validation loads about 3.4k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 1,561 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~196
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit 6e81ce1, republished under its MIT licence (© Xilinx). 1,561 words, ~3,367 tokens.

Download SKILL.mdSave it as .claude/skills/phase-1-kernel-validation/SKILL.md (or your agent's skills folder).
name
phase-1-kernel-validation
description
Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. Primary gate where a standalone harness exists: the harness's full-output element-wise `np.isclose` check at that kernel's `rtol`/`atol` vs an FP32 reference (the same PASS/FAIL `make run` prints). Fallback for a kernel with no harness: `make diagnosis` per-layer cosine vs the HF bf16 reference. Record each verified (kernel, shape) as a row in that kernel's "tested shapes" table in `kernel_registry/supported_kernels.md` + `details/<Kernel>_bf16.md` (Used by = `<model>`); per-model progress stays in `<model>/docs/`. Hard gate before integration. Invoked by deploy-new-llm after Phase 0.

Purpose

Enumerate every (kernel, shape) the model needs and verify each is numerically correct on real NPU2 before integrating them into a block. Phase 1 isolates per-kernel correctness from integration bugs and extends the shared kernel registry with the shapes this model exercises, so Phase 2+ (and future deployments) can rely on a verified shape catalog.

The core loop is small:

  1. Derive the kernel × shape list from the model's config (worked example: the llama-3.2-1B rows in kernel_registry/supported_kernels.md).
  2. For each (kernel, shape) → obtain a correctness verdict (harness atol/rtol PASS, or diagnosis cosine) → record verdict + mean_rel_L1
    • perf.
  3. For each new (kernel, shape), append a row to that kernel's "tested shapes" table in kernel_registry/supported_kernels.md and its details/<Kernel>_bf16.md (Used by = <model>); record the full per-kernel results under <model>/docs/.

Two sources of the per-kernel correctness verdict (use whichever exists):

  • Standalone harness (primary) — a self-contained make run that compiles ONE kernel/block to an ELF and checks its full output element-wise against an FP32 reference: np.isclose(|out−ref| ≤ atol + rtol·|ref|), every element must pass. The harness prints [precision] mean_rel_L1=… rtol=… atol=… PASS!/failed. — that PASS/FAIL IS the gate. This is the GPU/vLLM-aligned standard (rtol=1.6e-2 is PyTorch/vLLM's canonical bf16 tolerance; atol is sized per kernel to the measured datapath error). The registry's details/<Kernel>_bf16.md §Tolerances documents each kernel's exact rtol/atol. Upstream ships a harness for the FFN block (programming_examples/llms/llama_kernel_builder/ffn_swiglu/) plus the top-level kernel examples matrix_multiplication/bf16_in_bf16_out and matrix_multiplication/bf16_in_fp32_out (the BF16 GEMM, split by output dtype — the legacy matrix_multiplication/bf16 is kept for NPU1), matrix_vector_multiplication/bf16, flash_attention/kernel_fusion_based, eltwise_add, weighted_rms_norm, rms_norm, rope_lut.
  • In-context diagnosis (fallback) — for a kernel with no standalone harness, there is no isolated output to run element-wise atol/rtol on, so fall back to make diagnosis: per-layer ffn_out cosine vs the HF bf16 reference. The per-layer cosine reflects every kernel in that layer. Cosine is the fallback lens only — not the gate for any kernel that has a harness.

Phase 1 PASS criteria (HARD GATES)

Every (kernel, shape) the model needs must satisfy all four. Each catches a different bug class:

  1. A correctness signal exists for the kernel × shape: either a standalone harness make run (preferred — an isolated leaf check at the registry's atol/rtol) OR the kernel is covered by make diagnosis per-layer cosine. A kernel with neither is a silently-unverified kernel — not allowed.

  2. Numerical correctness (the real correctness gate — not theoretical compile-time rules):

    • Harness-backed kernel → the harness's full-output element-wise np.isclose check PASSES at that kernel's rtol/atol (details/<Kernel>_bf16.md §Tolerances) vs the FP32 reference. This is the GPU/vLLM-aligned standard and the gate for every kernel that has a harness.
    • No-harness kernel → fall back to the make diagnosis per-layer cosine vs HF bf16 staying healthy (no cliff at the layer exercising this kernel). Catches: silent-corruption tile configs (e.g. GEMM N % (tile_n × herd_n) != 0, which the builder does NOT assert). Trust the measured atol/rtol verdict, not the rule.

    Record mean_rel_L1 (the headline metric) plus max_abs / max_rel when the harness prints them. mean_rel_L1 is the registry's per-shape accuracy column and a cheap regression baseline for future deployments; the gate itself is the harness's pass/fail at rtol/atol.

  3. Tile utilization documented: each (kernel, shape) records its herd config. Targets:

    • Compute-bound (GEMM, FA): full 8×4 = 32 tiles
    • Row-parallel (RMSNorm, RoPE, GEMV-decode, SiLU+Mul, Eltwise): 8×1 = 8 tiles When achieved < target, justify in the catalog's Notes column (e.g. "M=1 decode RMSNorm uses 1 tile because batch=1 has no row-parallelism"). Catches: silent under-utilization.
  4. Registry row written: every verified (kernel, shape) that is new to the registry is appended to that kernel's "tested shapes" table in programming_examples/kernel_registry/supported_kernels.md and its details/<Kernel>_bf16.md, with Used by = <model>, matching the existing rows' column schema. The full per-kernel results (mean_rel_L1 + harness PASS or diagnosis cosine, max_abs/max_rel, perf, tile config) also go to <model>/docs/. Catches: deployments that pass without leaving a reusable record.

Failure on ANY criterion blocks Phase 2.

Knowledge base references

PRIMARY (read before starting):

  • programming_examples/kernel_registry/supported_kernels.md — index of every supported leaf kernel + its "tested shapes" table (shape, tile config, perf, mean_rel_L1, Used by, status) across deployments. This is the menu of known-good shapes to copy from and the table you extend.
  • programming_examples/kernel_registry/details/<Kernel>_bf16.md — per-kernel detail: the numerical datapath, Tunable parameters (knobs + hard constraints + tradeoffs), tolerances, per-shape data, and the reproduce commands (which harness, how to run). Read the page for each kernel your model needs. GEMM is the exception: it is split by output dtype into details/GEMM_bf16_in_bf16_out.md and details/GEMM_bf16_in_fp32_out.md (BF16-out has a --high-precision tier — fused-cast / drain = FP32-accumulate + single cast, GPU-standard ~9.3e-3; F32-out always FP32-accumulates).
  • programming_examples/kernel_registry/registry_lookup.py — the machine-readable half of the registry. gemm_config(M,K,N, output_dtype, precision) returns the registry's best measured {method, tile, gflops, mean_rel_L1} for a shape from the companion details/*.json, and raises (no silent guess) for an unmeasured shape — so Phase 1 recording a new GEMM shape is what unlocks the programmatic lookup that Phase 4 builders consume. The .md tables mirror these .json files for humans.

Workflow

Step 1: Derive the model's shape list

Read the HF config.json (or the model's <model>_weights.py:Config dataclass after Phase 0). Map each kernel call site to its shape using standard transformer identities (Q proj output dim = n_heads * head_dim, etc.). The llama-3.2-1B rows in kernel_registry/supported_kernels.md (and programming_examples/llms/llama32_1b/'s prefill/decode call sites) are the worked example of that call-site → shape mapping.

Write the working shape list — one row per (kernel, shape) with mean_rel_L1 + verdict + perf + status columns empty (Step 2 fills them) — under <model>/docs/ (e.g. <model>/docs/development_progress/phase1_kernels.md). This is scratch tracking for the deployment, not a registry file; the registry rows get written in Step 3.

Before running anything, read each needed kernel's details/<Kernel>_bf16.md Tunable parameters / constraints section and check the model's dims against it (alignment, max-K, placeability notes) — flag likely walls BEFORE compile time.

(If the model has unusual ops — Q/K Norm, post-norm, ops between currently-fused launches — flag in <model>/TODO.md as a Phase 2 prerequisite. The actual integration decision lives in phase-2-single-block-validation Step 0a.)

Show full SKILL.md (607 more words)Show less
Step 2: Per-kernel verification

For each (kernel, shape) in your Step 1 shape list:

a. Pick the correctness source + initial tile config. Each kernel's details/<Kernel>_bf16.md has its reproduce commands (which harness, how to run) and a Tunable parameters table (knobs + hard constraints + tradeoffs). If your shape exists in that kernel's "tested shapes" table in supported_kernels.md → reuse the tile config. Else → mirror the nearest-shape entry; verify the hard constraints hold for your shape; adjust if not (e.g., GEMM N % (tile_n × herd_n) != 0 → pick smaller tile_n).

b. Run correctness (+ profile where a harness exists).

If a standalone harness or top-level example covers the shape:

bash
cd programming_examples/<harness_or_example_dir>
flock -x -w 1800 /tmp/mlir-air-npu.lock make run       # element-wise atol/rtol vs FP32 ref → PASS!/failed.
flock -x -w 1800 /tmp/mlir-air-npu.lock make profile   # timing

If no standalone harness exists for the kernel, use the deployment's per-layer diagnosis:

bash
cd programming_examples/llms/<model>
flock -x -w 1800 /tmp/mlir-air-npu.lock make diagnosis  # per-layer ffn_out cosine vs HF bf16

NPU is shared on this machine — every NPU command must be flock-wrapped (see project memory). Compile-only steps don't need the lock.

If the harness reports failed. (or, for a no-harness kernel, the diagnosis cosine drops) → see "Failure modes" below. Bound: 1 retry per recipe per shape; if still failing, escalate to TODO.md "Active blockers".

c. Record results in <model>/docs/. Fill the scratch row: mean_rel_L1 (+ harness PASS / diagnosis cosine), profile (ms / GFLOPS) where measured, tile config used, tiles in flight, status, source (harness vs diagnosis), and Notes (especially when tiles-in-flight is below target — justify why).

Step 3: Extend the kernel registry

For each (kernel, shape) verified in Step 2 that's new (not already in that kernel's "tested shapes" table in supported_kernels.md), append a row to both supported_kernels.md and the kernel's details/<Kernel>_bf16.md per-shape table: shape + tile config + tiles-in-flight, "Used by" listing your new model, mean_rel_L1 + profile from Step 2, status. Match the existing rows' column schema for that kernel exactly.

This grows the registry organically: each verified deployment extends the menu of known-good shapes future deployments can copy from. (Adding a new kernel — not just a new shape of an existing one — is the heavier add-kernel workflow with its own parity checklist; Phase 1 only adds shapes to kernels the registry already covers.)

Failure modes

When a test fails (harness failed. / diagnosis cosine drop, compile error, hang), match the symptom to a likely cause. These are debug starting points, not gates — trust the measured atol/rtol verdict, not the rule. The relevant kernel's details/<Kernel>_bf16.md (constraints / placeability section) is the authority for that kernel's hard limits.

SymptomLikely causeWhere to look
'aiex.npu.push_queue' op Repeat count exceeds [0:255]GEMV K too large; auto-split outer dim ≥ 256Set k_split so K = k_split × inner and k_split ≤ 255; see details/GEMV_bf16.md
Allocator exhausted available buffer descriptor IDsBD pool exhausted (non-aligned dim, or non-monotonic placeability of the iteration count)Pad dim to 1024-aligned (GQA-aware reindexed padding) OR use kernel-first split-ELF path; see the placeability notes in the kernel's details/ page
L2 capacity exceeded (matvec builder assert)GEMV staged buffer > 512 KiBReduce tile_m (e.g., 8 → 2 for K=8192) or herd_m; see details/GEMV_bf16.md
Output all-zero / cosine = NaNBare-herd kernel without launch+segment wrapperWrap the bare herd in air.launch/air.segment (see ffn_swiglu harness for the multi-launch pattern)
Cosine = 0.02 or other smallGEMM N % (tile_n × herd_n) != 0 silent corruption (builder does NOT assert)Pick tile_n so divisibility holds at this N
FA all-NaN at runtimeCompile-flag mismatch on attn_npu2.cc macros; OR seq-first dk_chunks>1 path at head_dim≥128see debug-fa-runtime-failure skill
Compile hangs > 10 minCompiler scaling issue at large multi-launchCap and document; don't retry

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

On Phase 1 PASS:

  • Mark Phase 1 in <model>/TODO.md, append "(N/N kernels PASSED)"
  • Append summary to <model>/docs/development_progress/progress.md (per-kernel mean_rel_L1 + verdict, total time) — the per-model durable record
  • kernel_registry/supported_kernels.md + details/<Kernel>_bf16.md carry a "tested shapes" row (Used by = <model>) for every new (kernel, shape) this model exercises — the cross-model durable record

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phase-1-kernel-validation of Xilinx/mlir-air.

Open the folder on GitHubat commit 6e81ce1

Compare with similar skills

Phase 1 Kernel Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phase 1 Kernel Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phase 1 Kernel Validation this skillXilinx/mlir-air150—~3.4kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Aqua Troubleshootingoracle/accelerated-data-science125—~1.8kAutomated safety check: PassUPL-1.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Aqua Deploymentoracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.0
Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Aqua Troubleshooting

    oracle/accelerated-data-science

    Official

    Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.

    125 GitHub stars~1.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Aqua Deployment

    oracle/accelerated-data-science

    Official

    Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

    125 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

    150 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Questions about Phase 1 Kernel Validation

What does Phase 1 Kernel Validation do?

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. Phase 1 Kernel Validation is an agent skill from Xilinx/mlir-air. Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

When should I use Phase 1 Kernel Validation?

Phase 1 Kernel Validation fits situations like: tasks that involve LLM inference and serving; tasks that involve Deployment.

How do I install Phase 1 Kernel Validation in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a claude-code`. Or copy the skill folder (.claude/skills/phase-1-kernel-validation in Xilinx/mlir-air) into .claude/skills/phase-1-kernel-validation in your project. Claude Code loads it when a task matches its description.

How do I install Phase 1 Kernel Validation in Codex?

Run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a codex`. Or copy the skill folder (.claude/skills/phase-1-kernel-validation in Xilinx/mlir-air) into .agents/skills/phase-1-kernel-validation in your project. Codex loads it when a task matches its description.

Can I use Phase 1 Kernel Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill phase-1-kernel-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phase-1-kernel-validation, .gemini/skills/phase-1-kernel-validation, .github/skills/phase-1-kernel-validation and .opencode/skills/phase-1-kernel-validation in your project.

What does Phase 1 Kernel Validation need to run?

Going by SKILL.md and its folder, Phase 1 Kernel Validation needs the command-line tools its instructions call (make).

Does Phase 1 Kernel Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phase 1 Kernel Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phase 1 Kernel Validation use?

Phase 1 Kernel Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phase 1 Kernel Validation use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phase 1 Kernel Validation?

Skills that share tags, products or a category with Phase 1 Kernel Validation: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Aqua Troubleshooting (oracle/accelerated-data-science, 125 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Aqua Deployment (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phase 1 Kernel Validation?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.