Search
By Xilinx
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline. | Xilinx/ | 150 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 2 | A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128. | Xilinx/ | 150 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 3 | A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA… | Xilinx/ | 150 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 4 | Entry point for deploying a new decoder-only LLM on AMD NPU2. | Xilinx/ | 150 | — | ~4.8k | Automated safety check: Pass | MIT | today |
| 5 | Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them. | Xilinx/ | 150 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 6 | Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose. | Xilinx/ | 150 | — | ~1k | Automated safety check: Pass | MIT | today |
| 7 | Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation). | Xilinx/ | 150 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 8 | Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline… | Xilinx/ | 150 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 9 | Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 10 | Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/… | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 11 | Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level… | Xilinx/ | 150 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 12 | Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout). | Xilinx/ | 150 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 13 | Phase 5 of LLM deployment — apply the shared optimization skillset to a Phase-4-correct decode pipeline (multi-launch merge with N-way extern rename, static weight BOs, on-device layout). | Xilinx/ | 150 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 14 | Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/… | Xilinx/ | 150 | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 15 | Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the… | Xilinx/ | 150 | — | ~2.4k | Automated safety check: Pass | MIT | today |