Official agent skill

Extract Asm Onednn

by intel in intel/torch-xpu-ops

Extract GPU ISA from oneDNN ngen-JIT kernels. An agent skill from intel/torch-xpu-ops.

OfficialApache-2.0Auto-check passedDevelopment

Install Extract Asm Onednn

skills CLI
$ npx skills add intel/torch-xpu-ops --skill extract-asm-onednn -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/torch-xpu-ops extract-asm-onednn --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/extract-asm-onednn .claude/skills/extract-asm-onednn && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extract-asm-onednn
GitHub stars
115
Token cost
~1k tokens
SKILL.md length
236 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract GPU ISA from oneDNN ngen-JIT kernels. An agent skill from intel/torch-xpu-ops.

  • Works in 3 steps: Dump raw ISA via ONEDNN_JIT_DUMP. → Disassemble with IGA ctypes. → Pin the actually-invoked kernel.
  • Extracting ASM from oneDNN kernels (gemmkernel
  • SKILL.md covers How this differs from the…, When to use, When NOT to use and Steps
  • Calls python3

What it does

Extract Asm Onednn is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Extract GPU ISA from oneDNN ngen-JIT kernels. This is the ONLY codegen path that bypasses the standard SYCL/SPIR-V/IGC stack. oneDNN uses its own native code generator (ngen) that directly emits GPU ISA bytes — no SPIR-V, no IGC, no zebin ELF, no .debugline. Use when extracting ASM from oneDNN kernels (gemmkernel, genconvkernel), matmul, linear, conv, or SDPA-graph ops dispatched via mkldnn.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Project scaffolding. The licence is Apache-2.0.

When your agent uses it

  • Extracting ASM from oneDNN kernels (gemmkernel
  • SDPA-graph ops dispatched via mkldnn

Example prompts

  • “/extract-asm-onednn”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Dump raw ISA via ONEDNN_JIT_DUMP.
  2. Disassemble with IGA ctypes.
  3. Pin the actually-invoked kernel.

What it can do on your machine

Read from SKILL.md and the folder at commit a033aa5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extract Asm Onednn loads about 1k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 236 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/torch-xpu-ops at commit a033aa5, republished under its Apache-2.0 licence (© intel). 236 words, ~1,022 tokens.

Download SKILL.mdSave it as .claude/skills/extract-asm-onednn/SKILL.md (or your agent's skills folder).
name
extract-asm-onednn
description
Extract GPU ISA from oneDNN ngen-JIT kernels. This is the ONLY codegen path that bypasses the standard SYCL/SPIR-V/IGC stack. oneDNN uses its own native code generator (ngen) that directly emits GPU ISA bytes — no SPIR-V, no IGC, no zebin ELF, no .debug_line. Use when extracting ASM from oneDNN kernels (gemm_kernel, gen_conv_kernel), matmul, linear, conv, or SDPA-graph ops dispatched via mkldnn.

Extract ASM from oneDNN ngen-JIT Kernels

This is the ONLY path that does NOT use the standard compilation stack. All other scenarios (SYCL AOT/JIT, Triton) go through SPIR-V → IGC → zebin. oneDNN ngen bypasses all of that.

How this differs from the standard stack

Standard stack (SYCL/Triton):
  Source → LLVM IR → SPIR-V → IGC → zebin ELF (.text + .debug_line)
                                         ↓
                                    ocloc disasm → .asm

oneDNN ngen (THIS skill):
  oneDNN C++ templates → ngen JIT → RAW ISA BYTES (no ELF, no DWARF)
                                         ↓
                                    IGA ctypes → .asm

Key consequences:

  • No zebin — output is raw bytes, not ELF
  • No .debug_line — source mapping can only use pattern recognition
  • No IGC — IGC_ShaderDumpEnable is useless
  • No SPIR-V — ngen directly emits machine instructions
  • Dump mechanism: ONEDNN_JIT_DUMP=1 (writes .bin files)
  • Disassembly: IGA library via ctypes (not ocloc disasm)

When to use

  • Op dispatched through oneDNN (mkldnn::*): linear / matmul / mm / bmm / conv* / _scaled_dot_product_attention (oneDNN-graph)

When NOT to use

  • Kernel is triton_* → use extract-asm-triton
  • Kernel is _ZTS… (SYCL) → use extract-asm-syclkernel-{aot,jit}
  • You only need primitive-level timing → use benchdnn directly

Steps

Step 0: Locate tools
bash
# libiga64.so: shipped with oneAPI debugger component
# Detect oneAPI root: check env vars first, then common install locations
ONEAPI=${ONEAPI_ROOT:-${CMPLR_ROOT:+${CMPLR_ROOT%/*}}}
if [ -z "$ONEAPI" ]; then
  for d in /opt/intel/oneapi ~/intel/oneapi /usr/local/oneapi; do
    [ -d "$d" ] && ONEAPI="$d" && break
  done
fi
IGA_LIB=$(find ${ONEAPI:?"oneAPI not found; set ONEAPI_ROOT"} -name 'libiga64.so' 2>/dev/null | head -1)
test -n "$IGA_LIB" || { echo "libiga64.so not found under $ONEAPI; install oneAPI debugger"; exit 1; }
export IGA_LIB
  1. Dump raw ISA via ONEDNN_JIT_DUMP.

    bash
    OUT="<workdir>/onednn_$(date +%Y%m%d_%H%M%S)"
    mkdir -p "$OUT" && cd "$OUT"
    
    ONEDNN_JIT_DUMP=1 \
    ONEAPI_DEVICE_SELECTOR=level_zero:0 \
      <repro_cmd> 2>&1 | tee run.log
    
    ls dnnl_dump_gpu_*.bin

    These are raw ISA bytes (not zebin ELF). ocloc disasm cannot read them.

  2. Disassemble with IGA ctypes.

    oneDNN .bin files are raw ISA bytes (not ELF). Use libiga64.so via Python ctypes to disassemble. No separate script needed — run inline:

    bash
    for bin in dnnl_dump_gpu_*.bin; do
      name=$(basename "$bin" .bin)
      python3 -c "
    import ctypes, sys, pathlib
    iga = ctypes.CDLL('$IGA_LIB')
    raw = pathlib.Path('$bin').read_bytes()
    buf = ctypes.create_string_buffer(1 << 20)  # 1MB output buffer

Platform IDs: BMG/Xe2=0x2000000, PVC=0x30000, DG2=0x30004

iga.iga_disassemble(0x2000000, raw, len(raw), buf, len(buf)) sys.stdout.write(buf.value.decode()) " > "${name}.asm" done


**Platform ID selection:**
- BMG / Xe2 / LNL: `0x2000000`
- PVC (Ponte Vecchio): `0x30000`
- DG2 (Alchemist): `0x30004`

If `iga_disassemble` returns 0 bytes, the platform ID likely doesn't
match. Try the next one in the list above.

3. **Pin the actually-invoked kernel.**

```bash
ONEDNN_VERBOSE=1 <repro_cmd> 2>&1 | grep -E '^onednn_verbose.*exec' \
  | nl -ba | tee dnnl_verbose.log
# Nth (1-indexed) exec line → dnnl_dump_gpu_*_kernel.<N-1>.bin

The largest .bin by size is typically the GEMM kernel. Cross-check with grep -c dpas <asm> (GEMM has high dpas density).

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/extract-asm-onednn of intel/torch-xpu-ops.

Open the folder on GitHubat commit a033aa5

Compare with similar skills

Extract Asm Onednn next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extract Asm Onednn compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extract Asm Onednn this skillintel/torch-xpu-ops115—~1kAutomated safety check: PassApache-2.0
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence
PonytailDavidObando/gsharp5658 repos~1.7kAutomated safety check: PassMIT
Run Nx Generatornrwl/nx29k2 repos~592Automated safety check: NotesMIT
Conductor Setupgemini-cli-extensions/conductor3.8k—~4.2kAutomated safety check: PassApache-2.0
Mirage VFS Adapter Authoringstrukto-ai/mirage3.7k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    565 GitHub starsUsed in 8 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Run Nx generators with prioritization for workspace-plugin generators.

    29k GitHub starsUsed in 2 repos~592 tokens
    DevelopmentAuto-check: notes
  • Conductor Setup

    gemini-cli-extensions/conductor

    Scaffolds the project and sets up the Conductor environment.

    3.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Builds or extends a custom Mirage virtual filesystem adapter for an API, database, object store or app data, with a working mount configuration and filesystem tests.

    3.7k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Enforces this repository's TypeScript backend module architecture under server/: feature folders, barrel exports, and where shared types and utilities belong.

    14k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from intel/torch-xpu-ops

All 29 skills in this repo
  • Intel GPU Device Selection

    intel/torch-xpu-ops

    Official

    Select the Intel GPU device to use when a system has multiple Intel GPU devices.

    115 GitHub stars~508 tokensUpdated today
    Auto-check passed
  • Xpu CI Health Check

    intel/torch-xpu-ops

    Official

    Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…

    115 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • At Dispatch V2

    intel/torch-xpu-ops

    Official

    Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

    115 GitHub starsUsed in 3 repos~2.2k tokens
    Auto-check passed
  • PR Review

    intel/torch-xpu-ops

    Official

    Review pull requests for XPU operator or backend code. An agent skill from intel/torch-xpu-ops.

    115 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Skill Writer

    intel/torch-xpu-ops

    Official

    Guide users through creating Agent Skills for Claude Code. An agent skill from intel/torch-xpu-ops.

    115 GitHub starsUsed in 3 repos~2.4k tokens
    Auto-check passed
  • Ut Issue Authoring

    intel/torch-xpu-ops

    Official

    Read the evidence a nightly UT run produced, decide which failures share a root cause and which are machine breakage rather than product bugs, and write one issue draft per root cause to drafts.json.

    115 GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Extract Asm Onednn

What does Extract Asm Onednn do?

Extract GPU ISA from oneDNN ngen-JIT kernels. An agent skill from intel/torch-xpu-ops. Extract Asm Onednn is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Extract GPU ISA from oneDNN ngen-JIT kernels.

When should I use Extract Asm Onednn?

Extract Asm Onednn fits situations like: extracting ASM from oneDNN kernels (gemmkernel; SDPA-graph ops dispatched via mkldnn.

How do I install Extract Asm Onednn in Claude Code?

Run `npx skills add intel/torch-xpu-ops --skill extract-asm-onednn -a claude-code`. Or copy the skill folder (.claude/skills/extract-asm-onednn in intel/torch-xpu-ops) into .claude/skills/extract-asm-onednn in your project. Claude Code loads it when a task matches its description.

How do I install Extract Asm Onednn in Codex?

Run `npx skills add intel/torch-xpu-ops --skill extract-asm-onednn -a codex`. Or copy the skill folder (.claude/skills/extract-asm-onednn in intel/torch-xpu-ops) into .agents/skills/extract-asm-onednn in your project. Codex loads it when a task matches its description.

Can I use Extract Asm Onednn in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/torch-xpu-ops --skill extract-asm-onednn -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract-asm-onednn, .gemini/skills/extract-asm-onednn, .github/skills/extract-asm-onednn and .opencode/skills/extract-asm-onednn in your project.

What does Extract Asm Onednn need to run?

Going by SKILL.md and its folder, Extract Asm Onednn needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Extract Asm Onednn access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Extract Asm Onednn safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extract Asm Onednn use?

Extract Asm Onednn is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extract Asm Onednn use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extract Asm Onednn?

Skills that share tags, products or a category with Extract Asm Onednn: Nx Generate (nomcopter/react-mosaic, 4.8k stars), Ponytail (DavidObando/gsharp, 565 stars), Run Nx Generator (nrwl/nx, 29k stars) and Conductor Setup (gemini-cli-extensions/conductor, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extract Asm Onednn?

intel (a GitHub organization, an official publisher) maintains it in intel/torch-xpu-ops, which has 115 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: intel/torch-xpu-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.