Repository
Xilinx/mlir-air agent skills
- skills
- 15
- GitHub stars
- 150
- Stars
- 150 (50 forks)
- Licence
- MIT
- Last push
- Oct 2026
- Created
- Sep 2021
Install all skills
npx skills add Xilinx/mlir-airAdd --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).
Skills in Xilinx/mlir-air, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. | Xilinx/ | 150 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 2 | A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128. | Xilinx/ | 150 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 3 | A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA… | Xilinx/ | 150 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 4 | Entry point for deploying a new decoder-only LLM on AMD NPU2. | Xilinx/ | 150 | — | ~4.8k | Automated safety check: Pass | MIT | today |
| 5 | Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them. | Xilinx/ | 150 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 6 | Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. | Xilinx/ | 150 | — | ~1k | Automated safety check: Pass | MIT | today |
| 7 | Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). | Xilinx/ | 150 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 8 | Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline… | Xilinx/ | 150 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 9 | Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 10 | Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/… | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 11 | Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level… | Xilinx/ | 150 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 12 | Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout). | Xilinx/ | 150 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 13 | Phase 5 of LLM deployment — apply the shared optimization skillset to a Phase-4-correct decode pipeline (multi-launch merge with N-way extern rename, static weight BOs, on-device layout). | Xilinx/ | 150 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 14 | Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/… | Xilinx/ | 150 | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 15 | Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the… | Xilinx/ | 150 | — | ~2.4k | Automated safety check: Pass | MIT | today |
Questions, answered from the data.
What is the best skill in Xilinx/mlir-air?
Debug Bo Corruption from Xilinx/mlir-air ranks first of the 15 skills in Xilinx/mlir-air listed here, with the highest score: its repository has 150 GitHub stars, its SKILL.md loads about 1.5k tokens and it passes the automated safety check with no findings. Next come Debug Fa Runtime Failure and Debug Multi Launch Merge.
Are the skills in Xilinx/mlir-air official?
None yet. All 15 skills in Xilinx/mlir-air listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.
How do I install all skills from Xilinx/mlir-air?
Run npx skills add Xilinx/mlir-air in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.