Deploy Setup
garrytan/gstack
Detects where an app deploys, its production URL and health checks, then saves the deploy configuration in CLAUDE.md for /land-and-deploy.
Entry point for deploying a new decoder-only LLM on AMD NPU2.
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Xilinx/mlir-air deploy-new-llm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/deploy-new-llm .claude/skills/deploy-new-llm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .claude/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Xilinx/mlir-air deploy-new-llm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/deploy-new-llm .agents/skills/deploy-new-llm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .agents/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Xilinx/mlir-air deploy-new-llm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/deploy-new-llm .cursor/skills/deploy-new-llm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .cursor/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Xilinx/mlir-air.git --path .claude/skills/deploy-new-llm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Xilinx/mlir-air deploy-new-llm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/deploy-new-llm .gemini/skills/deploy-new-llm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .gemini/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Xilinx/mlir-air deploy-new-llmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/deploy-new-llm .github/skills/deploy-new-llm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .github/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Xilinx/mlir-air --skill deploy-new-llm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Xilinx/mlir-air deploy-new-llm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/deploy-new-llm .opencode/skills/deploy-new-llm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "deploy-new-llm" agent skill from https://github.com/Xilinx/mlir-air/tree/main/.claude/skills/deploy-new-llm into .opencode/skills/deploy-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-new-llm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
deploy-new-llmEntry point for deploying a new decoder-only LLM on AMD NPU2.
Deploy New LLM is an agent skill from Xilinx/mlir-air. Entry point for deploying a new decoder-only LLM on AMD NPU2. Invoked by the user as /deploy-new-llm <hfmodelid [--name <dirname] [--target npu2|npu1] [--dtype bf16|fp16]. Bootstraps the per-model workspace, validates architecture is in scope, and dispatches the 7 per-phase skills with the gate of each phase enforced by that phase's skill.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a7e4d00. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makehuggingface-cligitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Deploy New LLM loads about 4.8k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 1,827 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Xilinx/mlir-air at commit a7e4d00, republished under its MIT licence (© Xilinx). 1,827 words, ~4,835 tokens.
.claude/skills/deploy-new-llm/SKILL.md (or your agent's skills folder).Single user-facing entry point that scaffolds a new model deployment and orchestrates the per-phase skills. The orchestrator does NOT do correctness work itself — every gate is enforced by the corresponding per-phase skill. This skill's job is workflow coordination + workspace bootstrap.
This skill is "successful" when:
phase-7-independent-evaluator) verdict = PASS or
PASS-with-warningsIf any phase's gate fails irrecoverably, the deployment is marked
needs-human-review in TODO.md and the orchestrator stops — the human
triages.
programming_examples/llms/llama32_1b/ — the reference Tier-A deployment
(everything in scope today inherits from this)programming_examples/llms/verify/ — the shared verify subsystem
(HF bf16 reference; make verify token-set check is the PASS/FAIL gate,
make diagnosis per-layer cosine is the informational lens); every model
hooks in via its own verify_adapter.py, not by copying this.programming_examples/llms/llama_kernel_builder/ — the shared kernel
builder (KernelCache, external kernels, stitching, the ffn_swiglu/
harness); per-model scripts import from it.programming_examples/kernel_registry/ — the model-agnostic kernel
registry. Human half: supported_kernels.md (index) + details/<Kernel>_bf16.md
per kernel (GEMM is split by output dtype: GEMM_bf16_in_bf16_out.md +
GEMM_bf16_in_fp32_out.md). Machine half: details/*.json +
registry_lookup.py (gemm_config(...) returns the best measured tile
config; raises for unmeasured shapes). Phase 1 appends this model's
verified (kernel, shape) rows (Used by = <model>); Phase 4 builders read
tile configs back via the lookup instead of hardcoding them.This skill assumes the user already has:
cd programming_examples/llms/llama32_1b && make help prints
the target list.meta-llama/Llama-3.2-3B, run huggingface-cli login and accept the
model card on huggingface.co before invoking this skill. ~6 GB disk
per BF16 3B model in ~/.cache/huggingface/hub/.If any are missing, halt and ask the user to address them. Do NOT try to install MLIR-AIR / set up XRT / log into HF on the user's behalf.
meta-llama/Llama-3.2-3B)--name <dirname> (default: derived from model ID,
lowercased, slashes → underscores)--target npu2|npu1 (default: npu2)--dtype bf16|fp16 (default: bf16)Fetch HF config.json. Reject if any of:
MixtralForCausalLM, gpt-oss class)sliding_window set in config AND
use_sliding_window=true)QKV bias is supported (e.g. Qwen2-family with qkv_bias=true) — added
on the host around the bias-free kernels; the technique lives in
phase-2-single-block-validation Step 2. Surface it in TODO.md as a Phase 2
prerequisite.
If rejected, print clear message and do NOT proceed.
test -d programming_examples/llms/llama_kernel_builder && \
test -d programming_examples/llms/verify && \
test -d programming_examples/llms/llama32_1b && echo OK || echo MISSINGThe first two are required: every deployment composes kernels via the
shared llama_kernel_builder toolkit and gates on the shared verify/
subsystem. The third, llama32_1b, is the reference exemplar — read
to mirror its assembly, and imported directly on a bit-for-bit match. If
any is missing, halt and instruct the human.
<model>/ directory — kernel-first, minimalThe model lives at programming_examples/llms/<dirname>/, a sibling of
programming_examples/llms/llama32_1b/ (the reference exemplar) and the shared programming_examples/llms/verify/
programming_examples/llms/llama_kernel_builder/ (the toolkit every deployment builds on).Default mindset: build this model up from registry kernels using the
shared llama_kernel_builder toolkit (KernelCache, stitching,
external_kernels). The per-phase skills write the model's own
<model>_prefill.py / <model>_decode.py / multi_launch_builder/ by
composing the Phase-1-verified leaf kernels, reading llama32_1b's
assembly as the worked exemplar. This generalizes to any in-scope
architecture — it does not assume the model resembles llama.
Do NOT cp -r llama32_1b <model>. Two reasons depending on path
(the kernel-first-vs-inheritance decision + the bit-for-bit match rule are
owned by phase-2-single-block-validation Step 1 — Phase 2 makes the call):
llama32_1b_*.py resolve via sys.path to ../llama32_1b/, so there's
nothing to copy; a local copy would silently use outdated logic and miss
upstream bug fixes.The minimal Tier-A scaffold is:
programming_examples/llms/<dirname>/
├── .gitignore # copy from llama32_1b/.gitignore + add *.o, *kernel_cache/
├── Makefile # template-render with model name (run / verify / verify-full / diagnosis / profile + compile / clean)
├── README.md # placeholder; final version written by phase-6-finalize-and-learn
├── ARCHITECTURE.md # model-specific guide (NOT CLAUDE.md — top-level .gitignore excludes it, so it would not ship)
├── TODO.md # phase status (template in Step 5)
├── verify_adapter.py # written by phase-6-finalize-and-learn; hooks this model into the shared programming_examples/llms/verify/
└── docs/development_progress/
├── progress.md # header-only; phases append as they pass
├── LESSONS.md # header-only; appended on novel failures
└── debug_log.md # header-only; appended on debug-recipe firingsPer-model scripts use this sys.path block to resolve the shared
programming_examples/llms/ packages (always) and the llama32_1b reference (as exemplar, or
to import directly on a bit-for-bit match):
from pathlib import Path
import sys
_THIS_DIR = Path(__file__).resolve().parent
_LLMS_DIR = _THIS_DIR.parent # programming_examples/llms/
for p in (_LLMS_DIR, _LLMS_DIR / "llama32_1b", _THIS_DIR):
if str(p) not in sys.path:
sys.path.insert(0, str(p))
# ALWAYS — the shared kernel toolkit you compose this model FROM:
from llama_kernel_builder.external_kernels import compile_all_external_kernels
from llama_kernel_builder.cache import KernelCache
# DEFAULT (kernel-first) — write <model>_prefill.py / multi_launch_builder/
# that assemble the registry leaf kernels for THIS model's sequence,
# mirroring llama32_1b's builders as the worked example.
# SHORTCUT (bit-for-bit llama variant ONLY) — skip writing builders and
# reuse the reference's assembly directly:
# from llama32_1b_prefill import run_transformer_block, ...
# from llama32_1b_decode import run_decode_block, compile_decode_kernelsFor models that fork one reference module (partial arch divergence), copy
ONLY the file being forked, rename it <model>_prefill.py, compose the
divergent part kernel-first, and import the unchanged rest from
../llama32_1b/. Don't bulk-copy.
Per-phase skills produce these model-specific files:
phase-0-build-cpu-reference): <model>_weights.py, <model>_cpu_helpers.py<model>_phaseN_test.py per phasephase-6-finalize-and-learn): <model>_inference.py (clean
end-to-end NPU runner: setup → prefill → decode_loop) + <model>/verify_adapter.py
(hooks this model into the shared programming_examples/llms/verify/; mirrors
llama32_1b/verify_adapter.py)Makefile template (mirror programming_examples/llms/llama32_1b/Makefile):
make compile — compile all kernelsmake run — <model>_inference.py --n-tokens 100make verify / make verify-full — ../verify/verify_runner.py --runner=<model>.verify_adapter token-set gatemake diagnosis — ../verify/verify_runner.py --runner=<model>.verify_adapter per-layer cosinemake profile — <model>_inference.py --profilePROMPT, N_TOKENS, MODEL (base/instruct) plumbed throughmake clean — remove *kernel_cache/, air_project/, build_*/, *.o, verify/reports/<model>/TODO.mdTemplate (filled with Step 2 config):
# Deployment: <model_name>
## Phase status
- [ ] 0: Build CPU Reference
- [ ] 1: Kernel Validation
- [ ] 2: Single-Block Validation
- [ ] 3: Full-Model Validation
- [ ] 4: Prefill Optimization
- [ ] 5: Decode Optimization
- [ ] 6: Finalize & Learn
- [ ] 7: Independent Evaluation
## Active blockers
(none yet)
## Resolved config (pulled from HF)
n_layers: <N>, emb_dim: <D>, n_heads: <H>, n_kv_heads: <K>,
head_dim: <hd>, hidden_dim: <F>, vocab_size: <V>, rope_theta: <R>Create <model>/docs/development_progress/:
progress.md (header only)LESSONS.md (header only)debug_log.md (header only)phase_timing.md (REQUIRED — per-phase wall-clock log; schema below)phase_timing.md schema (per-phase effort breakdown — useful for
understanding which architectural axes are genuinely hard). Capture the
deployment-session start timestamp (date +"%s") at scaffold time.
Update at every phase boundary:
# <Model> deployment — per-phase wall-clock log
## Baselines
- Deployment session start: <YYYY-MM-DD HH:MM:SS TZ> (epoch=<N>)
- Scaffold complete: <YYYY-MM-DD HH:MM:SS TZ> (epoch=<N>)
## Phase log
### Phase N — <Name> (PENDING / PASS / PASS-with-warnings / BLOCKED, YYYY-MM-DD)
- start_ts: <epoch s> (HH:MM:SS TZ)
- end_ts: <epoch s> (HH:MM:SS TZ)
- wall_min: **<N>**
- npu_compile_min: <N> (sum of NPU kernel compile times in this phase)
- npu_runtime_s: <N> (sum of XRTRunner / inference NPU run time)
- dev_min: **<N>** ≈ wall - compile - runtime (agent thinking/code/debug)
- notable_events: <bullets — record honestly even if "stuck on debug for K min">
## Summary table (filled at deployment end)
| Phase | wall_min | npu_compile_min | npu_runtime_s | dev_min | notes |
|---|---:|---:|---:|---:|---|
| Scaffold + Step 0-3 | | | | | |
| 0: CPU Oracle | | | | | |
| ... | | | | | |
| **Total** | | | | | |Why this matters: dev_min (vs npu_compile_min / npu_runtime_s)
is the real "agentic deployment cost". Even debug-stuck phases should be
honestly recorded — high dev_min on a phase reveals which
architectural axes are genuinely hard, which informs future deployments.
Phase → skill mapping (each gate enforced by the per-phase skill):
| Phase | Skill | Gate (in 1 line — see the skill itself for full criteria) |
|---|---|---|
| 0 | phase-0-build-cpu-reference | <model>_weights.py + <model>_cpu_helpers.py produced; HF bf16 baseline loads & runs canonical prompt via verify/ HfRunner (sane top-1, no NaN); config matches HF config.json |
| 1 | phase-1-kernel-validation | Every leaf kernel × shape: harness atol/rtol element-wise check vs FP32 ref PASSES (GPU/vLLM standard), or make diagnosis cosine vs HF bf16 for no-harness kernels; each new shape recorded as a kernel_registry row (Used by = <model>) + full results in <model>/docs/ |
| 2 | phase-2-single-block-validation | Single transformer block on NPU: per-layer cosine vs HF bf16 (diagnosis lens) ≥ 0.99 (whole-tensor) + per-position min ≥ head_dim-scaled threshold |
| 3 | phase-3-full-model-validation | Full N layers: make verify token-set gate (top-5 inclusion vs HF bf16) PASSES — the binding gate; make diagnosis per-layer cosine is the localization lens, not a pass/fail |
| 4 | phase-4-prefill-optimization | Apply optimization patterns; correctness preserved (make verify token-set still PASSES — diagnosis cosine is the localization lens, not the gate) AND prefill kernel time strictly < Phase 3 baseline |
| 5 | phase-5-decode-optimization | Same shape: correctness preserved (make verify token-set still PASSES) AND decode time/token strictly < Phase 4 baseline |
| 6 | phase-6-finalize-and-learn | Clean <model>_inference.py + <model>/verify_adapter.py (shared programming_examples/llms/verify/) + Makefile; make verify (top-5 token-set vs HF bf16) PASSES |
| 7 | phase-7-independent-evaluator | Fresh subagent: audit make verify (anti-reward-hacking) + re-run as primary gate; produce structured evaluation_report.md |
Report current state to the human:
"Workspace scaffolded at
programming_examples/llms/<dirname>/. Resolved config:<summary>. Ready to start Phase 0 (Build CPU Reference). Invokephase-0-build-cpu-referenceto begin, or say 'go' for me to invoke it now."
For each phase (0 → 6):
date +"%s : %Y-%m-%d %H:%M:%S %Z",
record in phase_timing.md under that phase's start_ts: linewall_min,
npu_compile_min (from skill output), npu_runtime_s (from skill
output), dev_min (residual). Record in phase_timing.md under
that phase's section. Even if the phase was stuck on debug for most
of the time, record honestly — it reveals which axes are hard.After Phase 6 PASSES but BEFORE the final hand-off, spawn the
phase-7-independent-evaluator skill as Phase 7. It re-derives every
correctness claim with a fresh subagent and produces
<model>/docs/evaluation_report.md.
Why: Phases 0-6 are autonomous and self-reporting. The deployment agent has no incentive to cheat, but also no incentive to catch its own silent regressions (preload errors, fallback gates, etc.). Phase 7 is the independence check.
"Spawning Phase 7 — independent evaluation. The evaluator subagent will audit
make verify(anti-reward-hacking), re-run it as the primary gate, and write a structured report. Expected runtime: 15-30 min."
Then call the phase-7-independent-evaluator skill with <model_dir> as input.
If the evaluator reports:
needs-human-review in TODO.md and STOP.
Do NOT hand off. Surface specific failures.If this deployment changed shared infra (
kernel_registry/,programming_examples/llms/llama_kernel_builder/,programming_examples/llms/verify/, or the referenceprogramming_examples/llms/llama32_1b/), re-runmake verifyon the OTHER deployments underprogramming_examples/llms/<model>/before tagging — a shared-infra change can silently break a sibling. NPU is a singleton, so run sequentially withflock.
Once Phase 6 PASS AND Phase 7 PASS (or PASS-with-warnings), report:
"Deployment complete. See:
programming_examples/llms/<dirname>/docs/development_progress/progress.md— phase summaryprogramming_examples/llms/<dirname>/docs/development_progress/phase_timing.md— per-phase wall+dev timeprogramming_examples/llms/<dirname>/docs/evaluation_report.md— independent auditprogramming_examples/kernel_registry/supported_kernels.md— the kernel × shape rows (Used by =<dirname>) this deployment added"
Optional: tag the deployment if the project workflow uses git tags
(git tag -a deployment-<dirname>-v1 -m "..."). Most deployments
don't tag — git log + commit messages are the durable record.
| Symptom | Likely cause | What to do |
|---|---|---|
| Architecture rejected at Step 2 | MoE / sliding-window / MLA / encoder-decoder model | Halt; tell the user this model is out of scope |
llama32_1b/ missing at Step 3 | Reference deployment not present | Halt; instruct human (per-model scripts resolve imports against it) |
| Per-phase gate fails | Per-phase skill should escalate via TODO.md "Active blockers" | Don't try to fix here; the per-phase skill's failure-mode table is the right place |
| Phase 7 = FAIL | Evaluator surfaced a real correctness or reward-hacking issue | Mark needs-human-review in TODO.md; STOP; do NOT hand off |
| Cross-deployment regression at Phase 7 | This deployment's shared-infra change broke another deployment | Revert or fix the shared-infra change before tagging |
| User skips a phase to "save time" | Skipped phases mean later phases verify against unverified upstream | Refuse to advance past the skipped gate; explain the dependency chain (Phase 1 → 2 → 3 → 4/5 → 6 each need the previous PASS) |
For any failure not in the table, escalate to the human (this skill is
orchestration; debugging belongs in the per-phase skills' failure-mode
tables or the cross-cutting debug-* skills).
This skill primarily reads TODO.md and dispatches; it doesn't write
to progress files itself (per-phase skills do that). On all-PASS:
<model>/TODO.md reflects all 7 phases checked<model>/docs/development_progress/progress.md has each phase's
summary entry (written by per-phase skills)<model>/docs/development_progress/phase_timing.md complete with
Summary table filled (this skill writes it via Step 7 phase-boundary
timestamp captures)<model>/docs/evaluation_report.md exists (written by Phase 7)© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/deploy-new-llm of Xilinx/mlir-air.
Open the folder on GitHubat commit a7e4d00
Deploy New LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Deploy New LLM this skillXilinx/mlir-air | 150 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Deploy Setupgarrytan/gstack | 136k | — | ~11k | Automated safety check: Notes | MIT | |
| Deployment Patternsaffaan-m/ECC | 276k | 1 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Smart Contract Entry Point Analyzertrailofbits/skills | 7.4k | 1 repos | ~2.4k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Vercel Deploybytedance/deer-flow | 84k | 10 repos | ~797 | Automated safety check: Pass | MIT | |
| Land and Deploygarrytan/gstack | 136k | — | ~18k | Automated safety check: Notes | MIT |
garrytan/gstack
Detects where an app deploys, its production URL and health checks, then saves the deploy configuration in CLAUDE.md for /land-and-deploy.
affaan-m/ECC
Deployment iş akışları, CI/CD pipeline kalıpları, Docker konteynerizasyonu, sağlık kontrolleri, rollback stratejileri ve web uygulamaları için üretim hazırlığı kontrol listeleri.
trailofbits/skills
Maps the state-changing entry points of a smart contract codebase and sorts them by access level, producing a structured audit report that leaves out read-only functions.
bytedance/deer-flow
Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.
garrytan/gstack
Merges a pull request, waits for CI and the deploy, then verifies production health with canary checks, picking up where /ship leaves off.
affaan-m/ECC
Covers rolling, blue-green and canary deployments, multi-stage Dockerfiles, a GitHub Actions pipeline, health checks and production readiness for web apps.
Xilinx/mlir-air
A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.
Xilinx/mlir-air
A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.
Xilinx/mlir-air
A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…
Xilinx/mlir-air
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.
Xilinx/mlir-air
Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.
Xilinx/mlir-air
Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).
Entry point for deploying a new decoder-only LLM on AMD NPU2. Deploy New LLM is an agent skill from Xilinx/mlir-air. Entry point for deploying a new decoder-only LLM on AMD NPU2.
Run `npx skills add Xilinx/mlir-air --skill deploy-new-llm -a claude-code`. Or copy the skill folder (.claude/skills/deploy-new-llm in Xilinx/mlir-air) into .claude/skills/deploy-new-llm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Xilinx/mlir-air --skill deploy-new-llm -a codex`. Or copy the skill folder (.claude/skills/deploy-new-llm in Xilinx/mlir-air) into .agents/skills/deploy-new-llm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill deploy-new-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deploy-new-llm, .gemini/skills/deploy-new-llm, .github/skills/deploy-new-llm and .opencode/skills/deploy-new-llm in your project.
Going by SKILL.md and its folder, Deploy New LLM needs the command-line tools its instructions call (make, huggingface-cli and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Deploy New LLM is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Deploy New LLM: Deploy Setup (garrytan/gstack, 136k stars), Deployment Patterns (affaan-m/ECC, 276k stars), Smart Contract Entry Point Analyzer (trailofbits/skills, 7.4k stars) and Vercel Deploy (bytedance/deer-flow, 84k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.