Matlab Deploy AI Model
matlab/matlab-agentic-toolkit
Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder.
Extract Intel GPU ISA (assembly) from any XPU kernel. An agent skill from intel/torch-xpu-ops.
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .claude/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .claude/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .agents/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .agents/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .cursor/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .cursor/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/intel/torch-xpu-ops.git --path .claude/skills/extract-xpu-kernel-asm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .gemini/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .gemini/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .github/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .github/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install intel/torch-xpu-ops extract-xpu-kernel-asm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/extract-xpu-kernel-asm .opencode/skills/extract-xpu-kernel-asm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "extract-xpu-kernel-asm" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/extract-xpu-kernel-asm into .opencode/skills/extract-xpu-kernel-asm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract-xpu-kernel-asm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
extract-xpu-kernel-asmExtract Intel GPU ISA (assembly) from any XPU kernel. An agent skill from intel/torch-xpu-ops.
Extract Xpu Kernel Asm is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Extract Intel GPU ISA (assembly) from any XPU kernel. Classifies the codegen path (SYCL AOT, SYCL JIT, Triton, or oneDNN ngen) and delegates to the matching extraction skill. Use when asked to extract ASM, disassemble XPU kernels, get GPU ISA for an aten op, dump shader for PyTorch XPU, or disassemble a standalone DPC++/Triton binary.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Project scaffolding and Deep learning. It works with C++ and PyTorch. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a033aa5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash, python and cpp).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Extract Xpu Kernel Asm loads about 2.9k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 717 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from intel/torch-xpu-ops at commit a033aa5, republished under its Apache-2.0 licence (© intel). 717 words, ~2,947 tokens.
.claude/skills/extract-xpu-kernel-asm/SKILL.md (or your agent's skills folder).Classify which codegen path produced the kernel and delegate to the matching atomic skill to extract its Intel GPU ISA.
All XPU kernel compilation ultimately produces the same thing: GPU ISA bytes. The skills are separated not by compilation logic (which is largely identical), but by where the zebin lives:
┌─────────────────────────────────────────────────────────────────┐
│ Shared Compilation Stack │
│ SYCL/C++ → LLVM IR → SPIR-V → IGC → zebin (GPU ISA) │
│ │
│ • -g / -gline-tables-only acts at frontend (SYCL → LLVM IR) │
│ • Debug info flows: !dbg → OpLine → DebugLoc → .debug_line │
│ • JIT vs AOT: same IR pipeline, different WHEN it runs │
└─────────────────────────────────────────────────────────────────┘
┌───────────────┬──────────────────────────────────────────────────┐
│ Scenario │ Where is the zebin? How to get it? │
├───────────────┼──────────────────────────────────────────────────┤
│ sycl-aot │ Embedded in host binary at build time │
│ │ → DumpZEBin=1 NEOReadDebugKeys=1 → ocloc disasm │
├───────────────┼──────────────────────────────────────────────────┤
│ sycl-jit │ Generated at first launch, only in memory │
│ │ → IGC_ShaderDumpEnable=1 → dump dir → .asm/.elf │
├───────────────┼──────────────────────────────────────────────────┤
│ triton │ Triton compiler → SPIR-V → IGC JIT at launch │
│ │ → IGC_ShaderDumpEnable=1 (same as sycl-jit) │
├───────────────┼──────────────────────────────────────────────────┤
│ onednn (ngen) │ BYPASSES the entire SPIR-V/IGC stack │
│ │ Own JIT: ngen → raw ISA bytes (not zebin ELF) │
│ │ → ONEDNN_JIT_DUMP=1 → IGA ctypes disassembly │
└───────────────┴──────────────────────────────────────────────────┘Key implications for downstream:
sycl-aot, sycl-jit, triton all produce zebin ELF → same
ocloc disasm / readelf / .debug_line workflow applies.onednn produces raw ISA bytes (no ELF wrapper, no .debug_line)
→ requires IGA ctypes, pattern-recognition only for source mapping.@triton.jit, standalone DPC++.~/.triton/cache/.Key tools are NOT always on PATH. Probe before proceeding:
# Detect oneAPI root: check env vars first, then common install locations
ONEAPI=${ONEAPI_ROOT:-${CMPLR_ROOT:+${CMPLR_ROOT%/*}}}
if [ -z "$ONEAPI" ]; then
for d in /opt/intel/oneapi ~/intel/oneapi /usr/local/oneapi; do
[ -d "$d" ] && ONEAPI="$d" && break
done
fi
# ocloc (for disassembling zebin ELFs)
command -v ocloc >/dev/null || echo "ocloc not found; source oneapi-vars.sh"
# libiga64.so (for oneDNN ngen disassembly only)
IGA_LIB=$(test -n "$ONEAPI" && find "$ONEAPI" -name 'libiga64.so' 2>/dev/null | head -1)Identify which kernel was executed on the GPU. The kernel name determines which sub-skill to dispatch to (oneDNN vs Triton vs SYCL).
Preferred: unitrace (covers ALL code paths including native-handle):
unitrace -d <repro_cmd>Parse the == L0 Backend == table — each data row has the kernel name as the
first double-quoted field. Skip rows starting with ze* (those are API calls,
not GPU kernels). Extract and deduplicate kernel names.
Probe unitrace location in order: $UNITRACE env var → command -v unitrace
→ $UNITRACE_HOME/unitrace → <pti-gpu-build>/tools/unitrace/build/unitrace.
Fallback: SYCL_UR_TRACE (zero-dep, but has a blind spot):
SYCL_UR_TRACE=-1 <repro_cmd> 2>&1 | grep -oP 'pKernelName = 0x[0-9a-f]+ \(\K[^)]+'WARNING: SYCL_UR_TRACE is BLIND to kernels created via
urKernelCreateWithNativeHandle (oneDNN ngen, SYCL-TLA, Triton-xpu ≥ 3.7.0).
If the list is empty but the workload clearly ran GPU kernels, you MUST install
unitrace before proceeding. Do NOT extract ASM blindly.
| Signal | Scenario |
|---|---|
Kernel = gemm_kernel / gen_conv_kernel / routed via mkldnn::* | onednn |
Kernel = triton_* or standalone @triton.jit | triton |
Kernel = _ZTS… AND AOT path active for current device | sycl-aot |
Kernel = _ZTS… AND JIT path active (no AOT or target mismatch) | sycl-jit |
For _ZTS… kernels, determine AOT vs JIT:
A binary may contain __CLANG_OFFLOAD_BUNDLE but its AOT targets may not cover
the current device (e.g. built for PVC, running on BMG → runtime falls back to
JIT via SPIR-V). Test which path is actually active:
Method A (preferred): unitrace Kernel Properties
If unitrace was already used in Step 1, check the Compiled column:
AOT → scenario = sycl-aotJIT → scenario = sycl-jitMethod B: IGC dump probe
IGC is only invoked at runtime for JIT compilation. If IGC dump files appear, the kernel was JIT-compiled:
# Quick test: if IGC produces dump files, JIT path is active
rm -rf /tmp/igc_probe && mkdir -p /tmp/igc_probe
IGC_ShaderDumpEnable=1 IGC_ShaderDumpPidDisable=1 IGC_DumpToCustomDir=/tmp/igc_probe <repro_cmd> 2>/dev/null
if ls /tmp/igc_probe/*.asm 2>/dev/null | grep -q .; then
SCENARIO=sycl-jit
else
SCENARIO=sycl-aot
fi
rm -rf /tmp/igc_probeNOTE: Do NOT use DumpZEBin=1 for classification — it produces .elf files
for both AOT and JIT scenarios (the runtime always has a zebin to submit,
regardless of how it was produced).
If ambiguous → ask the user.
onednn → extract-asm-onednn
triton → extract-asm-triton
sycl-aot → extract-asm-syclkernel-aot
sycl-jit → extract-asm-syclkernel-jitAll fields REQUIRED:
| Field | Description |
|---|---|
input-kernel | User's kernel identifier (echoed verbatim) |
scenario | onednn / triton / sycl-aot / sycl-jit |
asm-dir | Absolute path to output directory |
asm-file | Absolute path to the chosen .asm file |
kernel-name | Kernel name as it appears in asm-file |
launch-evidence | ≥3 sentences explaining WHY this is the correct kernel |
Each example shows the dispatcher's classification — how to identify the scenario and what output to expect. The detailed extraction steps are in the respective sub-skill; these examples only demonstrate the classification signal and final result.
aten::matmulRepro:
import torch
a = torch.randn(4096, 4096, dtype=torch.bfloat16, device='xpu')
b = torch.randn(4096, 4096, dtype=torch.bfloat16, device='xpu')
c = a @ b; torch.xpu.synchronize()Classification signal: unitrace shows gemm_kernel → scenario = onednn
→ delegate to extract-asm-onednn.
Expected result:
scenario: onednn
asm-file: <workdir>/gemm.asm
kernel-name: gemm_kernel
validation: grep -c dpas gemm.asm → non-zero (GEMM uses dpas instructions)torch.compile(softmax)Repro:
import torch
@torch.compile
def fn(x): return torch.softmax(x, dim=-1)
x = torch.randn(1024, 1024, device='xpu')
fn(x); torch.xpu.synchronize()Classification signal: TORCH_LOGS=output_code shows a triton_per_fused_*softmax*
kernel name → scenario = triton → delegate to extract-asm-triton.
Expected result:
scenario: triton
asm-file: <igc_dump>/OCL_asm*_simd*_entry_*.asm
kernel-name: triton_per_fused_*softmax* (exact name varies by PyTorch version)
validation: grep 'libdevice.exp' <asm-file> confirms softmax exp computationRepro:
// vec_add.cpp
#include <sycl/sycl.hpp>
class VecAddKernel;
int main() {
sycl::queue q;
constexpr int N = 1 << 24;
float *a = sycl::malloc_device<float>(N, q);
float *c = sycl::malloc_device<float>(N, q);
q.parallel_for<VecAddKernel>(N, [=](int i) { c[i] = a[i] + 1.0f; }).wait();
sycl::free(a, q); sycl::free(c, q);
}icpx -fsycl -O2 -fsycl-targets=spir64_gen -Xs "-device <dev>" vec_add.cpp -o vec_addClassification signal: unitrace shows VecAddKernel with Compiled = AOT
→ scenario = sycl-aot → delegate to extract-asm-syclkernel-aot.
Expected result:
scenario: sycl-aot
asm-file: <workdir>/<name>_dump/.text._ZTS12VecAddKernel.asm
kernel-name: _ZTS12VecAddKernel
validation: c++filt _ZTS12VecAddKernel → VecAddKernellibtorch_xpu.soRepro:
import torch
x = torch.randn(1024, dtype=torch.float, device='xpu')
y = x + 1.0; torch.xpu.synchronize()Classification signal: unitrace shows the kernel with Compiled = AOT
→ scenario = sycl-aot → delegate to extract-asm-syclkernel-aot.
Expected result:
scenario: sycl-aot
asm-file: <workdir>/<name>_dump/.text._ZTSN2at6native3xpu...E.asm
kernel-name: matches the demangled kernel from c++filtRepro:
// shift_reduce.cpp
#include <sycl/sycl.hpp>
class ShiftReduceKernel;
int main() {
sycl::queue q;
auto *buf = sycl::malloc_device<int>(1024, q);
q.parallel_for<ShiftReduceKernel>(
sycl::nd_range<1>(1024, 32), [=](sycl::nd_item<1> it) {
int val = buf[it.get_global_id(0)];
val += sycl::shift_group_left(it.get_sub_group(), val, 1);
buf[it.get_global_id(0)] = val;
}).wait();
sycl::free(buf, q);
}icpx -fsycl -O2 -g -fsycl-targets=spir64 shift_reduce.cpp -o shift_reduceClassification signal: unitrace shows ShiftReduceKernel with Compiled = JIT
→ scenario = sycl-jit → delegate to extract-asm-syclkernel-jit.
Alternatively: IGC_ShaderDumpEnable=1 ./shift_reduce produces .asm files
in the IGC dump directory (confirming IGC was invoked at runtime = JIT).
Expected result:
scenario: sycl-jit
asm-file: <igc_dump>/OCL_asm*_simd*_entry_*.asm
kernel-name: _ZTS17ShiftReduceKernel
validation: compiled with -g → grep '// Line' shows source annotationsRepro:
// vec_add.cpp (same as Example 3)# Compile AOT for PVC, but run on BMG → runtime falls back to JIT via SPIR-V
icpx -fsycl -O2 -g -fsycl-targets=spir64_gen -Xs "-device pvc" vec_add.cpp -o vec_add_pvcClassification signal: unitrace may crash on cross-device AOT binaries.
Use IGC probe instead: IGC_ShaderDumpEnable=1 ./vec_add_pvc produces .asm
files → IGC was called at runtime → JIT fallback confirmed.
Expected result:
scenario: sycl-jit
asm-file: <igc_dump>/OCL_asm*_simd*_entry_*.asm
kernel-name: _ZTSN4sycl3_V16detail19__pf_kernel_wrapperI12VecAddKernelEE
validation: compiled with -g → grep '// Line' shows source annotations
.platform shows XE2 (BMG), not PVC© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/extract-xpu-kernel-asm of intel/torch-xpu-ops.
Open the folder on GitHubat commit a033aa5
Extract Xpu Kernel Asm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Extract Xpu Kernel Asm this skillintel/torch-xpu-ops | 115 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Matlab Deploy AI Modelmatlab/matlab-agentic-toolkit | 1.1k | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Embedded AI Deploymentmatlab/agent-skills-playground | 181 | 1 repos | ~3.4k | Automated safety check: Pass | Custom licence | |
| Qualcomm QNN Backend Developmentpytorch/executorch | 5.1k | — | ~1.8k | Automated safety check: Pass | Custom licence |
matlab/matlab-agentic-toolkit
Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
matlab/agent-skills-playground
Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).
pytorch/executorch
Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.
pytorch/pytorch
Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.
intel/torch-xpu-ops
Select the Intel GPU device to use when a system has multiple Intel GPU devices.
intel/torch-xpu-ops
Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
intel/torch-xpu-ops
Review pull requests for XPU operator or backend code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Guide users through creating Agent Skills for Claude Code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Read the evidence a nightly UT run produced, decide which failures share a root cause and which are machine breakage rather than product bugs, and write one issue draft per root cause to drafts.json.
Categories
Extract Intel GPU ISA (assembly) from any XPU kernel. An agent skill from intel/torch-xpu-ops. Extract Xpu Kernel Asm is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Extract Intel GPU ISA (assembly) from any XPU kernel.
Extract Xpu Kernel Asm fits situations like: asked to extract ASM; disassemble XPU kernels; get GPU ISA for an aten op; dump shader for PyTorch XPU.
Run `npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a claude-code`. Or copy the skill folder (.claude/skills/extract-xpu-kernel-asm in intel/torch-xpu-ops) into .claude/skills/extract-xpu-kernel-asm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a codex`. Or copy the skill folder (.claude/skills/extract-xpu-kernel-asm in intel/torch-xpu-ops) into .agents/skills/extract-xpu-kernel-asm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/torch-xpu-ops --skill extract-xpu-kernel-asm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract-xpu-kernel-asm, .gemini/skills/extract-xpu-kernel-asm, .github/skills/extract-xpu-kernel-asm and .opencode/skills/extract-xpu-kernel-asm in your project.
SKILL.md names no scripts, command-line tools or credentials: Extract Xpu Kernel Asm is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Extract Xpu Kernel Asm is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Extract Xpu Kernel Asm: Matlab Deploy AI Model (matlab/matlab-agentic-toolkit, 1.1k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars), Paddle Op Dev (PaddlePaddle/Paddle, 24k stars) and Embedded AI Deployment (matlab/agent-skills-playground, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
intel (a GitHub organization, an official publisher) maintains it in intel/torch-xpu-ops, which has 115 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.
Source: intel/torch-xpu-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.