Search

By Xilinx

15 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

Xilinx/mlir-air150—~1.5kAutomated safety check: PassMITtoday
2

A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

Xilinx/mlir-air150—~1.9kAutomated safety check: PassMITtoday
3

A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

Xilinx/mlir-air150—~1.8kAutomated safety check: PassMITtoday
4

Entry point for deploying a new decoder-only LLM on AMD NPU2.

Xilinx/mlir-air150—~4.8kAutomated safety check: PassMITtoday
5

Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

Xilinx/mlir-air150—~1.3kAutomated safety check: PassMITtoday
6

Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

Xilinx/mlir-air150—~1kAutomated safety check: PassMITtoday
7

Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

Xilinx/mlir-air150—~1.8kAutomated safety check: PassMITtoday
8

Phase 0 of LLM deployment — produce <modelweights.py (HF weight loader) and <modelcpuhelpers.py (the few NumPy helpers production prefill/decode import), then confirm the HF bf16 reference baseline…

Xilinx/mlir-air150—~2.8kAutomated safety check: PassMITtoday
9

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITtoday
10

Phase 2 of LLM deployment — wire the verified Phase 1 kernels into one transformer block on NPU and verify per-layer cosine vs the HF bf16 reference (the shared programmingexamples/llms/verify/…

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITtoday
11

Phase 3 of LLM deployment — wire all N layers and verify NPU matches the HF bf16 reference end-to-end (per-layer cosine via the shared programmingexamples/llms/verify/ diagnosis lens + token-level…

Xilinx/mlir-air150—~2.2kAutomated safety check: PassMITtoday
12

Phase 4 of LLM deployment — apply the shared optimization skillset to a Phase-3-correct prefill pipeline (multi-launch merge, BO pre-loading + intermediate buffer reuse, seq-first layout).

Xilinx/mlir-air150—~2.3kAutomated safety check: PassMITtoday
13

Phase 5 of LLM deployment — apply the shared optimization skillset to a Phase-4-correct decode pipeline (multi-launch merge with N-way extern rename, static weight BOs, on-device layout).

Xilinx/mlir-air150—~2.3kAutomated safety check: PassMITtoday
14

Phase 6 of LLM deployment — integrate Phase 4 prefill + Phase 5 decode into a clean <modelinference.py, write the model's verifyadapter.py hooking into the shared programmingexamples/llms/verify/…

Xilinx/mlir-air150—~3.5kAutomated safety check: PassMITtoday
15

Phase 7 of LLM deployment — spawn a fresh subagent that treats the deployment as UNTRUSTED, audits the make verify implementation (anti-reward-hacking: confirms the token-set gate runs the…

Xilinx/mlir-air150—~2.4kAutomated safety check: PassMITtoday