Official agent skill

At Dispatch V2

by intel in intel/torch-xpu-ops

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install At Dispatch V2

skills CLI
$ npx skills add intel/torch-xpu-ops --skill at-dispatch-v2 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/torch-xpu-ops at-dispatch-v2 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/at-dispatch-v2 .claude/skills/at-dispatch-v2 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
at-dispatch-v2
GitHub stars
115
Used in
3 other repos
Token cost
~2.2k tokens
SKILL.md length
534 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

  • Works in 7 steps: Add the Dispatch_v2.h include → Identify the old dispatch pattern → Map the old macro to type groups → …
  • Porting ATDISPATCHALLTYPESAND
  • SKILL.md covers When to use this skill, Quick reference, Key transformations and Instructions, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

At Dispatch V2 is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. Use when porting ATDISPATCHALLTYPESAND, ATDISPATCHFLOATINGTYPES, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning and GPU and accelerator computing. It works with PyTorch, C++ and CUDA. The licence is Apache-2.0.

When your agent uses it

  • Porting ATDISPATCHALLTYPESAND
  • ATDISPATCHFLOATINGTYPES
  • Other dispatch macros to the new v2 API

Example prompts

  • “/at-dispatch-v2”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Add the Dispatch_v2.h include
  2. Identify the old dispatch pattern
  3. Map the old macro to type groups
  4. Extract the individual types
  5. Transform to AT_DISPATCH_V2
  6. Handle multi-line lambdas
  7. Verify the conversion

What it can do on your machine

Read from SKILL.md and the folder at commit 0187b3b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are cpp).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

At Dispatch V2 loads about 2.2k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 534 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/torch-xpu-ops at commit 0187b3b, republished under its Apache-2.0 licence (© intel). 534 words, ~2,193 tokens.

Download SKILL.mdSave it as .claude/skills/at-dispatch-v2/SKILL.md (or your agent's skills folder).
name
at-dispatch-v2
description
Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.

AT_DISPATCH to AT_DISPATCH_V2 Converter

This skill helps convert PyTorch's legacy AT_DISPATCH macros to the new AT_DISPATCH_V2 format, as defined in aten/src/ATen/Dispatch_v2.h.

When to use this skill

Use this skill when:

  • Converting AT_DISPATCH_* macros to AT_DISPATCH_V2
  • Porting ATen kernels to use the new dispatch API
  • Working with files in src/ that use dispatch macros
  • User mentions "AT_DISPATCH", "dispatch v2", "Dispatch_v2.h", or macro conversion

Quick reference

Old format:

cpp
AT_DISPATCH_ALL_TYPES_AND3(kBFloat16, kHalf, kBool, dtype, "kernel_name", [&]() {
  // lambda body
});

New format:

cpp
AT_DISPATCH_V2(dtype, "kernel_name", AT_WRAP([&]() {
  // lambda body
}), AT_EXPAND(AT_ALL_TYPES), kBFloat16, kHalf, kBool);

Key transformations

  1. Reorder arguments: scalar_type and name come first, then lambda, then types
  2. Wrap the lambda: Use AT_WRAP(lambda) to handle internal commas
  3. Expand type groups: Use AT_EXPAND(AT_ALL_TYPES) instead of implicit expansion
  4. List individual types: Add extra types (kHalf, kBFloat16, etc.) after expanded groups
  5. Add include: #include <ATen/Dispatch_v2.h> near other Dispatch includes

Instructions

Step 1: Add the Dispatch_v2.h include

Add the v2 header near the existing #include <ATen/Dispatch.h>:

cpp
#include <ATen/Dispatch.h>
#include <ATen/Dispatch_v2.h>

Keep the old Dispatch.h include for now (other code may still need it).

Step 2: Identify the old dispatch pattern

Common patterns to convert:

  • AT_DISPATCH_ALL_TYPES_AND{2,3,4}(type1, type2, ..., scalar_type, name, lambda)
  • AT_DISPATCH_FLOATING_TYPES_AND{2,3}(type1, type2, ..., scalar_type, name, lambda)
  • AT_DISPATCH_ALL_TYPES_AND_COMPLEX_AND{2,3}(type1, ..., scalar_type, name, lambda)
  • AT_DISPATCH_FLOATING_AND_COMPLEX_TYPES_AND{2,3}(type1, ..., scalar_type, name, lambda)
Step 3: Map the old macro to type groups

Identify which type group macro corresponds to the base types:

Old macro baseAT_DISPATCH_V2 type group
ALL_TYPESAT_EXPAND(AT_ALL_TYPES)
FLOATING_TYPESAT_EXPAND(AT_FLOATING_TYPES)
INTEGRAL_TYPESAT_EXPAND(AT_INTEGRAL_TYPES)
COMPLEX_TYPESAT_EXPAND(AT_COMPLEX_TYPES)
ALL_TYPES_AND_COMPLEXAT_EXPAND(AT_ALL_TYPES_AND_COMPLEX)

For combined patterns, use multiple AT_EXPAND() entries:

cpp
// Old: AT_DISPATCH_ALL_TYPES_AND_COMPLEX_AND2(...)
// New: AT_EXPAND(AT_ALL_TYPES), AT_EXPAND(AT_COMPLEX_TYPES), type1, type2
Step 4: Extract the individual types

From AT_DISPATCH_*_AND2(type1, type2, ...) or AT_DISPATCH_*_AND3(type1, type2, type3, ...), extract the individual types (type1, type2, etc.).

These become the trailing arguments after the type group:

cpp
AT_DISPATCH_V2(..., AT_EXPAND(AT_ALL_TYPES), kBFloat16, kHalf, kBool)
                                             ^^^^^^^^^^^^^^^^^^^^^^^^
                                             Individual types from AND3
Step 5: Transform to AT_DISPATCH_V2

Apply the transformation:

Pattern:

cpp
AT_DISPATCH_V2(
  scalar_type,           // 1st: The dtype expression
  "name",                // 2nd: The debug string
  AT_WRAP(lambda),       // 3rd: The lambda wrapped in AT_WRAP
  type_groups,           // 4th+: Type groups with AT_EXPAND()
  individual_types       // Last: Individual types
)

Example transformation:

cpp
// BEFORE
AT_DISPATCH_ALL_TYPES_AND3(
    kBFloat16, kHalf, kBool,
    iter.dtype(),
    "min_values_cuda",
    [&]() {
      min_values_kernel_cuda_impl<scalar_t>(iter);
    }
);

// AFTER
AT_DISPATCH_V2(
    iter.dtype(),
    "min_values_cuda",
    AT_WRAP([&]() {
      min_values_kernel_cuda_impl<scalar_t>(iter);
    }),
    AT_EXPAND(AT_ALL_TYPES),
    kBFloat16, kHalf, kBool
);
Step 6: Handle multi-line lambdas

For lambdas with internal commas or complex expressions, AT_WRAP is essential:

cpp
AT_DISPATCH_V2(
    dtype,
    "complex_kernel",
    AT_WRAP([&]() {
      gpu_reduce_kernel<scalar_t, scalar_t>(
        iter,
        MinOps<scalar_t>{},
        thrust::pair<scalar_t, int64_t>(upper_bound(), 0)  // Commas inside!
      );
    }),
    AT_EXPAND(AT_ALL_TYPES)
);
Step 7: Verify the conversion

Check that:

  • AT_WRAP() wraps the entire lambda
  • Type groups use AT_EXPAND()
  • Individual types don't have AT_EXPAND() (just kBFloat16, not AT_EXPAND(kBFloat16))
  • Argument order is: scalar_type, name, lambda, types
  • Include added: #include <ATen/Dispatch_v2.h>
Show full SKILL.md (218 more words)Show less

Type group reference

Available type group macros (use with AT_EXPAND()):

cpp
AT_INTEGRAL_TYPES      // kByte, kChar, kInt, kLong, kShort
AT_FLOATING_TYPES      // kDouble, kFloat
AT_COMPLEX_TYPES       // kComplexDouble, kComplexFloat
AT_QINT_TYPES         // kQInt8, kQUInt8, kQInt32
AT_ALL_TYPES          // INTEGRAL_TYPES + FLOATING_TYPES
AT_ALL_TYPES_AND_COMPLEX  // ALL_TYPES + COMPLEX_TYPES
AT_INTEGRAL_TYPES_V2  // INTEGRAL_TYPES + unsigned types
AT_BAREBONES_UNSIGNED_TYPES  // kUInt16, kUInt32, kUInt64
AT_FLOAT8_TYPES       // Float8 variants

Common patterns

Pattern: AT_DISPATCH_ALL_TYPES_AND2
cpp
// Before
AT_DISPATCH_ALL_TYPES_AND2(kHalf, kBFloat16, dtype, "op", [&]() {
  kernel<scalar_t>(data);
});

// After
AT_DISPATCH_V2(dtype, "op", AT_WRAP([&]() {
  kernel<scalar_t>(data);
}), AT_EXPAND(AT_ALL_TYPES), kHalf, kBFloat16);
Pattern: AT_DISPATCH_FLOATING_TYPES_AND3
cpp
// Before
AT_DISPATCH_FLOATING_TYPES_AND3(kHalf, kBFloat16, kFloat8_e4m3fn,
    tensor.scalar_type(), "float_op", [&] {
  process<scalar_t>(tensor);
});

// After
AT_DISPATCH_V2(tensor.scalar_type(), "float_op", AT_WRAP([&] {
  process<scalar_t>(tensor);
}), AT_EXPAND(AT_FLOATING_TYPES), kHalf, kBFloat16, kFloat8_e4m3fn);
Pattern: AT_DISPATCH_ALL_TYPES_AND_COMPLEX_AND2
cpp
// Before
AT_DISPATCH_ALL_TYPES_AND_COMPLEX_AND2(
    kComplexHalf, kHalf,
    self.scalar_type(),
    "complex_op",
    [&] {
      result = compute<scalar_t>(self);
    }
);

// After
AT_DISPATCH_V2(
    self.scalar_type(),
    "complex_op",
    AT_WRAP([&] {
      result = compute<scalar_t>(self);
    }),
    AT_EXPAND(AT_ALL_TYPES),
    AT_EXPAND(AT_COMPLEX_TYPES),
    kComplexHalf,
    kHalf
);

Edge cases

Case 1: No extra types (rare)
cpp
// Before
AT_DISPATCH_ALL_TYPES(dtype, "op", [&]() { kernel<scalar_t>(); });

// After
AT_DISPATCH_V2(dtype, "op", AT_WRAP([&]() {
  kernel<scalar_t>();
}), AT_EXPAND(AT_ALL_TYPES));
Case 2: Many individual types (AND4, AND5, etc.)
cpp
// Before
AT_DISPATCH_FLOATING_TYPES_AND4(kHalf, kBFloat16, kFloat8_e4m3fn, kFloat8_e5m2,
    dtype, "float8_op", [&]() { kernel<scalar_t>(); });

// After
AT_DISPATCH_V2(dtype, "float8_op", AT_WRAP([&]() {
  kernel<scalar_t>();
}), AT_EXPAND(AT_FLOATING_TYPES), kHalf, kBFloat16, kFloat8_e4m3fn, kFloat8_e5m2);
Case 3: Lambda with no captures
cpp
// Before
AT_DISPATCH_ALL_TYPES_AND2(kHalf, kBool, dtype, "op", []() {
  static_kernel<scalar_t>();
});

// After
AT_DISPATCH_V2(dtype, "op", AT_WRAP([]() {
  static_kernel<scalar_t>();
}), AT_EXPAND(AT_ALL_TYPES), kHalf, kBool);

Benefits of AT_DISPATCH_V2

  1. No arity in macro name: Don't need different macros for AND2, AND3, AND4
  2. Composable type sets: Mix and match type groups with AT_EXPAND()
  3. Extensible: Easy to add more types without hitting macro limits
  4. Clearer: Type groups are explicit, not implicit in macro name

Important notes

  • Keep #include <ATen/Dispatch.h> - other code may need it
  • The AT_WRAP() is mandatory - prevents comma parsing issues in the lambda
  • Type groups need AT_EXPAND(), individual types don't
  • The v2 API is in aten/src/ATen/Dispatch_v2.h - refer to it for full docs
  • See the header file for the Python script to regenerate the macro implementation

Workflow

When asked to convert AT_DISPATCH macros:

  1. Read the file to identify all AT_DISPATCH uses
  2. Add #include <ATen/Dispatch_v2.h> if not present
  3. For each dispatch macro:
    • Identify the pattern and extract components
    • Map the base type group
    • Extract individual types
    • Construct the AT_DISPATCH_V2 call
    • Apply with Edit tool
  4. Show the user the complete converted file
  5. Explain what was changed

Do NOT compile or test the code - focus on accurate conversion only.

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/at-dispatch-v2 of intel/torch-xpu-ops.

Open the folder on GitHubat commit 0187b3b

Used in 3 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in intel/torch-xpu-ops, which our catalogue first saw on October 7, 2026.

Compare with similar skills

At Dispatch V2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

At Dispatch V2 compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
At Dispatch V2 this skillintel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from intel/torch-xpu-ops

All 29 skills in this repo
  • Intel GPU Device Selection

    intel/torch-xpu-ops

    Official

    Select the Intel GPU device to use when a system has multiple Intel GPU devices.

    115 GitHub stars~508 tokensUpdated yesterday
    Auto-check passed
  • Xpu CI Health Check

    intel/torch-xpu-ops

    Official

    Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…

    115 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • PR Review

    intel/torch-xpu-ops

    Official

    Review pull requests for XPU operator or backend code. An agent skill from intel/torch-xpu-ops.

    115 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Skill Writer

    intel/torch-xpu-ops

    Official

    Guide users through creating Agent Skills for Claude Code. An agent skill from intel/torch-xpu-ops.

    115 GitHub starsUsed in 3 repos~2.4k tokens
    Auto-check passed
  • Ut Issue Authoring

    intel/torch-xpu-ops

    Official

    Read the evidence a nightly UT run produced, decide which failures share a root cause and which are machine breakage rather than product bugs, and write one issue draft per root cause to drafts.json.

    115 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Ut Refactor Review

    intel/torch-xpu-ops

    Official

    Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests.

    115 GitHub stars~917 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about At Dispatch V2

What does At Dispatch V2 do?

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. At Dispatch V2 is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

When should I use At Dispatch V2?

At Dispatch V2 fits situations like: porting ATDISPATCHALLTYPESAND; ATDISPATCHFLOATINGTYPES; other dispatch macros to the new v2 API.

How do I install At Dispatch V2 in Claude Code?

Run `npx skills add intel/torch-xpu-ops --skill at-dispatch-v2 -a claude-code`. Or copy the skill folder (.claude/skills/at-dispatch-v2 in intel/torch-xpu-ops) into .claude/skills/at-dispatch-v2 in your project. Claude Code loads it when a task matches its description.

How do I install At Dispatch V2 in Codex?

Run `npx skills add intel/torch-xpu-ops --skill at-dispatch-v2 -a codex`. Or copy the skill folder (.claude/skills/at-dispatch-v2 in intel/torch-xpu-ops) into .agents/skills/at-dispatch-v2 in your project. Codex loads it when a task matches its description.

Can I use At Dispatch V2 in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/torch-xpu-ops --skill at-dispatch-v2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/at-dispatch-v2, .gemini/skills/at-dispatch-v2, .github/skills/at-dispatch-v2 and .opencode/skills/at-dispatch-v2 in your project.

What does At Dispatch V2 need to run?

SKILL.md names no scripts, command-line tools or credentials: At Dispatch V2 is instructions for the agent only. Our summary lists: Python 3.

Does At Dispatch V2 access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is At Dispatch V2 safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does At Dispatch V2 use?

At Dispatch V2 is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does At Dispatch V2 use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to At Dispatch V2?

Skills that share tags, products or a category with At Dispatch V2: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Metal Kernel (pytorch/pytorch, 104k stars) and Ako4all (TongmingLAIC/AKO4ALL, 369 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains At Dispatch V2?

intel (a GitHub organization, an official publisher) maintains it in intel/torch-xpu-ops, which has 115 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.

Source: intel/torch-xpu-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.