Cuda Index Width
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-ops --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .claude/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .claude/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-opsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-ops --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .agents/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .agents/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-ops --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .cursor/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .cursor/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/meta-pytorch/attention-gym.git --path .agents/skills/validating-pytorch-custom-ops--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-ops --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .gemini/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .gemini/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-opsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .github/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .github/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install meta-pytorch/attention-gym validating-pytorch-custom-ops --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/validating-pytorch-custom-ops .opencode/skills/validating-pytorch-custom-ops && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "validating-pytorch-custom-ops" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/validating-pytorch-custom-ops into .opencode/skills/validating-pytorch-custom-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validating-pytorch-custom-ops", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
validating-pytorch-custom-opsEnsures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.
Validating Pytorch Custom Ops is an agent skill from meta-pytorch/attention-gym. Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. Use when adding or reviewing an implementation or custom kernel; covers customop, tritonop, fake tensors, autograd, opcheck, full-graph compilation, dynamic shapes, and CUDA Graphs.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch and CUDA. The repository describes itself as: Helpful tools and examples for working with flex-attention. The licence is BSD-3-Clause.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0beac51. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.pytorch.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Validating Pytorch Custom Ops loads about 6.4k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 2,530 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from meta-pytorch/attention-gym at commit 0beac51, republished under its BSD-3-Clause licence (© meta-pytorch). 2,530 words, ~6,440 tokens.
.claude/skills/validating-pytorch-custom-ops/SKILL.md (or your agent's skills folder).Use this workflow whenever adding a backend under attn_gym/linear/<variant>/impl/ or
attn_gym/sparse/<variant>/impl/.
define/impl
pair plus an autograd wrapper.opcheck with compact inputs and a dense permuted layout of each tensor input
(p.t().contiguous().t(), or token-innermost x.movedim(1, -1).contiguous().movedim(-1, 1)).
Padded or sliced inputs do not substitute: empty_like of a non-dense tensor is already
contiguous, so a stride-copying fake still passes.fullgraph=True compile on every supported architecture; Hopper and
Blackwell often dispatch backends that allocate differently.torch.compile.torch.compile. Prefer
torch.library.triton_op with torch.library.wrap_triton when a stable operator boundary is
useful and compiler subsystems should remain able to inspect the implementation.torch.library.define/impl, below) when the launcher cannot be traced correctly.The documented public function remains an ordinary Python function that owns semantic validation, mode selection, and backend dispatch. Compiling only a private launcher does not prove that the public operation is compile-friendly.
Every registered-operator boundary costs CPU time per call. Microbenchmark (torch 2.14 nightly, B200, identical trivial launcher, CPU wall-clock per call, medians over interleaved rounds):
| boundary | inference | fwd, requires_grad=True |
|---|---|---|
| raw Python | 1.2 us | 3.7 us (+autograd.Function) |
define/impl | 2.6 us | 5.3 us (+autograd.Function) |
custom_op | 5.5 us | 8.2 us (+autograd.Function) |
impl(op, "Autograd") kernel | 6.8 us | 7.4 us |
custom_op + register_autograd | — | 10.2 us |
Nesting one registered op inside another pays the boundary again (nested custom_op 7.1 us,
nested define/impl 3.7 us). Wrapping the real kda_l2norm_fwd launcher in custom_op added
~11 us per forward call over invoking the launcher directly.
Consequences for launch-bound paths (small kernels, high call rate):
autograd.Function over the launcher with no torch.library
registration (real l2norm op: ~17 us cheaper per gradient-tracking forward and ~58 us per
fwd+bwd than the custom_op boundary), at the cost of fullgraph=True support.torch.library.custom_op costs ~3 us
more per call for schema inference and mutation checking and is only preferable for
prototypes.autograd.Function over register_autograd for hot training ops; it was the cheapest
measured autograd boundary. Registering a Python kernel at the Autograd dispatch key
(torch.library.impl(op, "Autograd")) makes the raw op differentiable but runs Python on
every call, including inference; use it only when direct torch.ops calls must be
differentiable.fwd/bwd pair, pass the variant as a schema argument (str kind, bool, float), and
select the backend inside the CUDA implementation on real tensors. Do not create one op pair
per backend or thread a use_<backend> flag through autograd.Function.apply: selection that
runs under Dynamo tracing specializes into the graph, trips the functools.lru_cache warning
for cached device-property helpers, and multiplies fakes, handles, and Functions. Make every
backend return the same output contract (e.g. already-reduced parameter gradients) so one fake
describes all of them. attn_gym::_gate_transform_{fwd,bwd} is the reference shape.define/impl op has no autograd kernel: backprop through a direct call warns and produces
no gradient, so route all differentiable use through the autograd.Function wrapper.torch.library.opcheck accepts torch.ops.attn_gym.op.default, so the validation workflow
below is unchanged.Private backends should return tensors or a fixed tensor tuple. The public API may wrap that result
in a documented NamedTuple so users can use attributes or unpack it:
from typing import NamedTuple
import torch
class OperationOutput(NamedTuple):
output: torch.Tensor
final_state: torch.Tensor | None = None
def operation(...) -> OperationOutput:
output, final_state = backend_forward(...)
return OperationOutput(output, final_state)A registered operator's output structure and arity must agree with its schema. Do not switch between returning a tensor and a tuple based on an argument. Represent optional outputs explicitly and match them in the fake implementation.
The repo standard for opaque kernel boundaries. Write the schema string yourself; nothing is inferred from annotations:
import torch
from torch import Tensor
torch.library.define(
"attn_gym::example_fwd",
"(Tensor query, Tensor key, Tensor value, Tensor? initial_state) -> (Tensor, Tensor)",
)
torch.library.define("attn_gym::example_bwd", "(Tensor query, Tensor grad_output) -> Tensor")
def _example_fwd_cuda(
query: Tensor, key: Tensor, value: Tensor, initial_state: Tensor | None
) -> tuple[Tensor, Tensor]:
return launch_backend(query, key, value, initial_state)
# Register with an explicit call, not as a decorator: the decorator form returns None,
# and keeping the launcher callable lets benchmarks and tests bypass the boundary.
torch.library.impl("attn_gym::example_fwd", "CUDA", _example_fwd_cuda)
@torch.library.register_fake("attn_gym::example_fwd")
def _example_fwd_fake(
query: Tensor, key: Tensor, value: Tensor, initial_state: Tensor | None
) -> tuple[Tensor, Tensor]:
return value.new_empty(value.shape), query.new_empty(query.shape[:2])
_example_fwd = torch.ops.attn_gym.example_fwd.default
_example_bwd = torch.ops.attn_gym.example_bwd.default # impl/fake registered the same way
class _Example(torch.autograd.Function):
@staticmethod
def forward(ctx, query: Tensor, key: Tensor, value: Tensor) -> Tensor:
output, state = _example_fwd(query, key, value, None)
ctx.save_for_backward(query)
return output
@staticmethod
@torch.autograd.function.once_differentiable
def backward(ctx, grad_output: Tensor):
(query,) = ctx.saved_tensors
return _example_bwd(query, grad_output), None, NoneSchema rules:
Tensor?; scalars are int, float, bool; defaults may be embedded
(Tensor? cu_seqlens=None).(Tensor(a!) state, Tensor value) -> ().custom_op annotation inference accepts, adding
e.g. ScalarType, Layout, MemoryFormat, Generator, SymInt, and int[2]. Neither flow
accepts arbitrary Python objects, dataclasses, or callables; flatten configs to scalars or
specialize per-config outside the op.Tensor? and, as of torch 2.15 nightly (verified 2026-08-16:
plain fullgraph compile, autograd.Function wrapping + backward, and opcheck all pass for
both the tensor and None branches), no longer hard-fail the compile stack; on older
releases the None return handling was unreliable. Keep the fixed-arity convention anyway:
for "N or N+1 returns depending on a flag", define two fixed-arity schemas sharing one
launcher (see kda_chunk_fwd / kda_chunk_fwd_with_state) and let the autograd.Function
branch on the flag; its output arity is free to vary. The bool flag specializes into separate
graphs regardless, so a merged optional-output schema saves nothing under compile, while a
Tensor? return forces None-narrowing on every caller and extra schema boxing costs real
time on launch-bound eager paths (collapsing schema pairs measured 3.22% slower on the dense
forward probe, PR #314).kda_chunk_bwd schema takes
Tensor? cu_seqlens, Tensor? chunk_offsets for both dense and ragged and passes the strict
fullgraph matrix. The forward keeps separate dense/ragged ops as a measured eager-dispatch
decision, not a compile requirement.autograd.Function via ctx.save_for_backward, not part of the wrapper's
user-facing return. Unlike register_autograd, the Function is not limited to saving
user-visible outputs.torch.ops.attn_gym.<op>.default to a module-level name and call that; resolving the
torch.ops attribute chain per call adds overhead.The fake implementation describes output metadata without running the kernel. It must:
x.new_empty(shape) / torch.empty(...)), not with
torch.empty_like(x): it copies a dense permuted input's strides, so the fake can disagree
with the kernel (Inductor assert_size_stride failure) and a kernel indexing outputs with
contiguous offsets writes the wrong elements;.item();For genuinely data-dependent output dimensions, use torch.library.get_ctx() and
new_dynamic_size() rather than inspecting input data.
Registration must exist before graph capture. Keep registration deterministic and lightweight; do not rely on registration side effects occurring inside a compiled region. Continue importing optional kernel dependencies lazily.
Expose the weakest layout contract the kernel actually needs. Do not require full contiguity when
only one mode must be contiguous, and do not treat alignment as a runtime repair mechanism:
assumed_align is a promise to codegen that may enable wide loads.
Use the shared helpers in attn_gym._backends.cute:
tensor_supports_contiguous_dim(tensor, dim=-1, alignment_bytes=...) checks that one mode has
stride 1 and that every slice origin satisfies the promised byte alignment;make_fake_strided_tensor(dtype, shape, contiguous_dim=-1, ...) creates the matching TVM-FFI
fake signature with dynamic other strides and honest minimum alignment;tensor_supports_tma(tensor) is the 16-byte CUDA/TMA specialization of the same predicate;requires_int64_abi(...) still selects the independent wide-stride/address specialization.Prefer two measured signatures when vectorization needs stronger assumptions:
Normalize only unsupported layouts. A noncontiguous required mode may need a compact copy, but an
outer-strided or misaligned tensor belongs on the general signature if the generated loads are safe.
Note that .contiguous() returns an already-contiguous tensor unchanged and therefore does not fix
a misaligned storage offset.
For every new CuTe tensor ABI, test:
Express views through native slicing and local_tile/partition algebra. Reconstruct a tensor from
.iterator only for a deliberate storage reinterpretation such as a TMA mode permutation, and
explain that contract at the call site.
Declare every mutated argument in the schema string:
torch.library.define("attn_gym::update_state", "(Tensor(a!) state, Tensor value) -> ()")Do not declare an operator functional if its launcher writes into an input, including recurrent state, cache, workspace, or output buffers. Prefer functional operators when practical.
Training backends attach autograd with a torch.autograd.Function outside the registered ops, as
in the pattern above: forward calls the forward op (autograd is already disabled inside
Function.forward, so no redispatch guard is needed), backward calls the backward op, and the
public API routes through Function.apply. Notes:
@torch.autograd.function.once_differentiable.None for nondifferentiable inputs.ctx.mark_non_differentiable(...) for auxiliary outputs and
ctx.set_materialize_grads(False) when the backward handles None grads.torch.ops calls that backprop through it
warn and produce no gradient. Route all differentiable use through the wrapper.CustomOpDef.register_autograd (slowest measured path) or an Autograd dispatch-key
kernel (taxes every inference call) unless raw op calls must be differentiable.import torch
import triton
import triton.language as tl
from torch import Tensor
@triton.jit
def add_kernel(x_ptr, y_ptr, output_ptr, size, BLOCK: tl.constexpr):
block = tl.program_id(0)
offsets = block * BLOCK + tl.arange(0, BLOCK)
mask = offsets < size
x = tl.load(x_ptr + offsets, mask=mask)
y = tl.load(y_ptr + offsets, mask=mask)
tl.store(output_ptr + offsets, x + y, mask=mask)
@torch.library.triton_op("attn_gym::add", mutates_args={})
def add(x: Tensor, y: Tensor) -> Tensor:
output = torch.empty_like(x)
size = x.numel()
grid = lambda meta: (triton.cdiv(size, meta["BLOCK"]),)
torch.library.wrap_triton(add_kernel)[grid](
x,
y,
output,
size,
BLOCK=256,
)
return outputKeep cross-variant Triton infrastructure in attn_gym._backends.triton only after it has multiple
callers. Variant-specific kernels, schemas, fake behavior, and autograd remain in the variant
backend.
Optimized kernels must preserve an int32 default path and select a separate int64 specialization
when relative element offsets can exceed signed int32. Inventory every input, output, optional
tensor, manually reconstructed view, and manual load/store offset; checking only the primary input
or numel() is insufficient.
For Triton, use attn_gym._backends.triton.requires_int64_offsets in a
@triton.heuristics USE_INT64_OFFSETS constexpr. It checks reachable storage cosize for the
project's nonnegative-strided layouts. In the wide branch, cast program IDs, loaded token/chunk
origins, and other address indices to tl.int64 before the first potentially overflowing multiply
or addition. Casting a completed offset or pointer is too late. Bounded routing arrays may stay
int32, but widen their loaded values before using them in wide pointer arithmetic.
For CuTeDSL TVM-FFI, use attn_gym._backends.cute.requires_int64_abi. In addition to reachable
cosize, it checks every declared ABI stride because a size-one dimension can carry an unreachable
stride larger than INT32_MAX that the compiled signature must still represent. Carry
use_int64_offsets through the op, compile-cache key, and stable kernel name; use matching
cute.sym_int64 fake-signature fields in the wide variant. Widen values before multiplication when
constructing layouts or adding to iterators.
Required validation has three layers:
INT32_MAX, compared with an
equivalent compact layout. Singleton-stride tests prove ABI routing but not wide device pointer
arithmetic.Also assert ordinary layouts select int32 and measure both variants before claiming no regression;
int64 address arithmetic can increase instructions and registers. Use test/test_kda_int64_offsets.py
as the project reference.
opchecktorch.library.opcheck validates operator registration. Its default utilities are:
test_schema: runtime mutation and aliasing agree with the schema;test_autograd_registration: autograd is registered correctly;test_faketensor: fake execution matches real output metadata;test_aot_dispatch_dynamic: AOT dispatch works with dynamic shapes.Run it through pytest after importing the registration module:
import pytest
import torch
OPCHECK_UTILITIES = (
"test_schema",
"test_autograd_registration",
"test_faketensor",
"test_aot_dispatch_dynamic",
)
@pytest.mark.parametrize("dtype", [torch.float16, torch.bfloat16])
@pytest.mark.parametrize("requires_grad", [False, True])
def test_example_forward_registration(dtype, requires_grad):
query, key, value = make_inputs(dtype=dtype, requires_grad=requires_grad)
torch.library.opcheck(
example_forward,
(query, key, value),
test_utils=OPCHECK_UTILITIES,
)Pass keyword arguments as the third argument:
torch.library.opcheck(example_forward, args, {"causal": True})Create separate cases for materially different registration paths:
Keep raise_exception=True, the default, in committed tests. Use raise_exception=False only while
diagnosing multiple failures:
results = torch.library.opcheck(example_forward, args, raise_exception=False)
for utility, result in results.items():
print(utility, result)opcheck does not validate numerical correctness. A kernel can pass every utility and still compute
wrong values. Pair it with reference forward tests and gradient tests.
For unordered index outputs, AOT opcheck's exact tensor comparison can reject a valid permutation. Keep all utilities on a stable-output case (for example, selecting every candidate), then test the selective production path under fullgraph/dynamic compilation and replay with range, uniqueness, valid-count and reference-score checks. Do not widen integer tolerances or require deterministic ordering merely to satisfy opcheck when ordering and boundary tie choices are unspecified.
Every optimized backend needs a trusted eager/reference oracle. Compare:
Use torch.autograd.gradcheck when double precision is supported:
def test_example_forward_gradcheck():
inputs = make_inputs(dtype=torch.double, requires_grad=True)
assert torch.autograd.gradcheck(example_forward, inputs)For lower-precision-only GPU kernels, compare gradients against the reference with explicit atol/rtol values justified by dtype and reduction order.
Compile the documented function with strict graph capture:
def test_compiled_forward_and_backward():
eager_inputs = make_inputs(requires_grad=True)
compiled_inputs = clone_inputs(eager_inputs)
expected = operation(*eager_inputs, backend="triton")
compiled_operation = torch.compile(operation, fullgraph=True)
actual = compiled_operation(*compiled_inputs, backend="triton")
torch.testing.assert_close(actual.output, expected.output, atol=ATOL, rtol=RTOL)
torch.testing.assert_close(
actual.final_state,
expected.final_state,
atol=ATOL,
rtol=RTOL,
)
expected_gradients = torch.autograd.grad(loss(expected), eager_inputs)
actual_gradients = torch.autograd.grad(loss(actual), compiled_inputs)
for actual_gradient, expected_gradient in zip(actual_gradients, expected_gradients):
torch.testing.assert_close(
actual_gradient,
expected_gradient,
atol=GRAD_ATOL,
rtol=GRAD_RTOL,
)Include every mode that selects a materially different implementation, initial state present and
absent, and final state requested and omitted. A graph break under fullgraph=True fails the
advertised compilation contract.
If dynamic shapes are supported, compile once and reuse the callable across multiple sizes:
compiled_operation = torch.compile(operation, fullgraph=True, dynamic=True)
for sequence_length in (127, 193):
inputs = make_inputs(sequence_length)
expected = operation(*inputs)
actual = compiled_operation(*clone_inputs(inputs))
assert_outputs_close(actual, expected)Use TORCH_LOGS="recompiles" to verify dimensions advertised as dynamic are reused rather than
silently specialized.
When CUDA Graph support is claimed:
Do not infer CUDA Graph compatibility from a successful torch.compile call.
TORCH_LOGS="graph_breaks,recompiles"; locate whether validation,
dispatch, registration, or the backend launcher caused it.assert_size_stride failure only in a warm cache: Inductor's graph caches do not key on fake
implementations, so after changing a fake's strides a persistent TORCHINDUCTOR_CACHE_DIR
replays graphs built from the old fake. Rerun with a fresh cache dir before debugging; after
such a change lands, clear the Modal attention-gym-compile-cache Volume tarballs.Report these independently:
opcheck cases and utilities exercised;torch.compile(fullgraph=True) forward/backward results;Do not present lint, imports, registration-only checks, or opcheck alone as behavioral correctness.
torch.library APIs, including custom_op, register_fake, register_autograd, and opcheck:
https://docs.pytorch.org/docs/stable/library.htmltorch.compile:
https://docs.pytorch.org/tutorials/recipes/torch_compile_user_defined_triton_kernel_tutorial.html© meta-pytorch, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/validating-pytorch-custom-ops of meta-pytorch/attention-gym.
Open the folder on GitHubat commit 0beac51
Validating Pytorch Custom Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Validating Pytorch Custom Ops this skillmeta-pytorch/attention-gym | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
matlab/agent-skills-playground
Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).
meta-pytorch/attention-gym
Sets up an isolated per-worktree Python environment for attention-gym development using nightly PyTorch and the CI-mirroring uv flow.
meta-pytorch/attention-gym
Adds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking.
Categories
Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. Validating Pytorch Custom Ops is an agent skill from meta-pytorch/attention-gym.compile-friendly and correctly registered.
Validating Pytorch Custom Ops fits situations like: reviewing an implementation; covers customop; full-graph compilation.
Run `npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a claude-code`. Or copy the skill folder (.agents/skills/validating-pytorch-custom-ops in meta-pytorch/attention-gym) into .claude/skills/validating-pytorch-custom-ops in your project. Claude Code loads it when a task matches its description.
Run `npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a codex`. Or copy the skill folder (.agents/skills/validating-pytorch-custom-ops in meta-pytorch/attention-gym) into .agents/skills/validating-pytorch-custom-ops in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add meta-pytorch/attention-gym --skill validating-pytorch-custom-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validating-pytorch-custom-ops, .gemini/skills/validating-pytorch-custom-ops, .github/skills/validating-pytorch-custom-ops and .opencode/skills/validating-pytorch-custom-ops in your project.
SKILL.md names no scripts, command-line tools or credentials: Validating Pytorch Custom Ops is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: docs.pytorch.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Validating Pytorch Custom Ops is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Validating Pytorch Custom Ops: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Metal Kernel (pytorch/pytorch, 104k stars) and Ako4all (TongmingLAIC/AKO4ALL, 369 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
meta-pytorch (a GitHub organization) maintains it in meta-pytorch/attention-gym, which has 1,254 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 6, 2026.
Source: meta-pytorch/attention-gym on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.