Add Column Type
simstudioai/sim
Add a new table column type to Sim — registry entry, icon, storage shape, coercion, and the behavioral hooks the grid and API read.
Adds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking.
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-template --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .claude/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .claude/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-templateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-template --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .agents/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .agents/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-template --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .cursor/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .cursor/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/meta-pytorch/attention-gym.git --path .agents/skills/cutedsl-tunable-kernel-template--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-template --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .gemini/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .gemini/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-templateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .github/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .github/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install meta-pytorch/attention-gym cutedsl-tunable-kernel-template --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/meta-pytorch/attention-gym.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/cutedsl-tunable-kernel-template .opencode/skills/cutedsl-tunable-kernel-template && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cutedsl-tunable-kernel-template" agent skill from https://github.com/meta-pytorch/attention-gym/tree/main/.agents/skills/cutedsl-tunable-kernel-template into .opencode/skills/cutedsl-tunable-kernel-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutedsl-tunable-kernel-template", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cutedsl-tunable-kernel-templateAdds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking.
Cutedsl Tunable Kernel Template is an agent skill from meta-pytorch/attention-gym. Adds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking. Use after the core kernel exists and needs a public default, explicit-config, or autotune path.
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `copy_reads_example.py`).
The repository describes itself as: Helpful tools and examples for working with flex-attention. The licence is BSD-3-Clause.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 70fd810. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cutedsl Tunable Kernel Template loads about 4.1k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,793 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from meta-pytorch/attention-gym at commit 70fd810, republished under its BSD-3-Clause licence (© meta-pytorch). 1,793 words, ~4,099 tokens.
.claude/skills/cutedsl-tunable-kernel-template/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this after cutedsl-kernel-template establishes the core kernel. This skill owns the Attention Gym adapter and parent/compiler-process boundary. Load cutedsl-performance separately to design the search space, measure candidates, and accept a selector.
The canonical Attention Gym helpers are:
from attn_gym._backends.cute import benchmark_gpu, compile_tvm_ffi, jit_cache, run_tunable, tuneIf another repo has equivalent helpers, reuse them. Do not rebuild cache/process orchestration inside each op.
Keep each value in one domain:
| Value | Owner |
|---|---|
Real torch.Tensor inputs and outputs | Parent process and public wrapper |
| Cohesive runtime tensor/output bundle | One parent-only NamedTuple or dataclass when the flat argument list grows |
| Candidate generation from shapes/dtypes/device facts | configs(*runtime_args) in the parent |
| Benchmark-winner reuse identity | Host-static tuning_key(*runtime_args, target=...) result |
| Static, pickleable specialization values | Config record and compile_call(...) result |
| Fake CuTe tensors matching the runtime ABI | Cached compile(...) method |
| Fake environment stream and typed TVM-FFI option | compile_tvm_ffi(...) |
| Compiled callable launch and benchmarking | Parent process through launch(...) |
Never pass real tensors to compile(...) or compiler workers. compile_call(...) is the explicit projection from runtime inputs to static compile arguments.
Use a module-scope NamedTuple or frozen dataclass for codegen choices. Module scope makes the config importable and pickleable by fresh compiler processes.
When warps have named protocol responsibilities, represent them with a module-scope IntEnum
such as WarpRole.TMA_PRODUCER; do not compare warp indices with unexplained integer literals. Name
the responsibility precisely when one warp performs more than one role.
Make approximation choices such as fastmath explicit compile-time arguments. Default them to
False unless the public numerical contract deliberately chooses approximate math, encode them in
cache and profiler names, and correctness-test every exposed mode.
Put ordinary specialization defaults directly in the owning constructor or public function signature. Avoid module constants that only alias those defaults or a one-kernel policy limit; keep fixed compile-time expressions next to the device code or schedule that consumes them. Derive one-use schedule values in the owning op rather than adding free helper functions.
Do not repeat dtype assertions inside a CuTeDSL entrypoint when the cached compiler boundary already
constructs an exact fake-tensor ABI and the runtime operator validates that ABI. Keep small helpers
only when they are reused or mark a real protocol, cache, or runtime-ABI boundary; use established
integer utilities such as ceildiv directly instead of wrapping one arithmetic expression.
from typing import NamedTuple
from attn_gym._backends.cute.target import CompileTarget
class MyConfig(NamedTuple):
threads: int
tile_size: intA tunable adapter supplies seven distinct responsibilities. When its launch ABI has several values,
make that ABI an Args type owned by the adapter rather than a separate module-level private type:
class MyOp:
class Args(NamedTuple):
q: torch.Tensor
k: torch.Tensor
output: torch.Tensor
workspace: torch.Tensor
@staticmethod
def default_config(args: Args, *, target: CompileTarget) -> MyConfig:
"""Return a deterministic valid config without benchmarking."""
return MyConfig(threads=128, tile_size=256)
@staticmethod
def tuning_key(args: Args, *, target: CompileTarget) -> tuple[int]:
"""Identify workloads that may safely reuse one benchmark winner."""
return (args.q.shape[1],)
@staticmethod
def configs(args: Args) -> tuple[MyConfig, ...]:
"""Return valid candidates derived from actual runtime inputs."""
...
@staticmethod
@jit_cache
def compile(*static_args):
"""Construct the fake ABI and compile one specialization."""
...
@staticmethod
def compile_call(config: MyConfig, args: Args) -> tuple:
"""Project launch arguments to static arguments for compile(...)."""
...
@staticmethod
def launch(compiled, config: MyConfig, args: Args):
"""Launch in the parent with real tensors and return the public result."""
...Keep these responsibilities distinct, but they need not all live on the CuTeDSL op class. A constructor-heavy DSL op may stay focused on layouts and device code while a small adapter owns the seven static/class methods above.
default_config(*runtime_args, target=...) owns deterministic no-tune selection. It may inspect
tensor metadata, static operation arguments, and target facts, and must return a valid config for
every input accepted by the adapter without benchmarking or reading device-resident tensor values.
Keep this policy on the adapter rather than selecting a default in the parent wrapper. A constant
fallback simply ignores its arguments. Name canonical mode configs after the mode rather than
calling an architecture-specific config the default.tuning_key(*runtime_args, target=...) defines when an autotuned winner may be reused. Return
host-visible tensor metadata and target facts only; never read device values or synchronize, because
the key must remain valid during CUDA Graph replay. Use () when compile_call(...) already
distinguishes every relevant workload. Exact keys are the safe starting point; introduce buckets
only after measurements establish stable winner regions.configs(...) may inspect runtime shape, stride, dtype, alignment, or device facts.compile(...) owns fake tensors and the cached artifact boundary.compile_call(...) prevents runtime tensors from leaking into cache keys or subprocess payloads.launch(...) binds the compiled callable to real tensors for correctness checks and timing.The protocol accepts positional runtime arguments for small ABIs. Keep one or two naturally named
values positional; wrapping them in Args would add ceremony without clarifying ownership. Do not,
however, grow a positional list of inputs, outputs, and static semantics indefinitely. Once several
values form one launch ABI, pass one parent-only Args value. Nest Args in its adapter when no
other component owns that ABI; use a descriptive top-level type when multiple components genuinely
share it. Do not create a module-level _Runtime or _Args type merely to signal privacy. Reserve
“runtime” for an execution environment or lifecycle-bearing object, not a passive argument tuple.
compile_call(...) projects Args to static values, while launch(...) binds its real tensors to
the compiled callable.
Nesting communicates ownership, not public API status. If an adapter has a normal class name but is
an implementation detail, define the module's __all__ explicitly and list only the supported
wrapper and configuration types. Do not use leading underscores as a substitute for deciding and
documenting the module's actual public surface.
Users call one ordinary PyTorch function; they do not construct op objects or compiled callables.
def my_op(
q: torch.Tensor,
k: torch.Tensor,
*,
config: MyConfig | None = None,
tune: bool = False,
configs: Iterable[MyConfig] | None = None,
) -> torch.Tensor:
args = MyOp.Args(
q,
k,
torch.empty_like(q),
torch.empty_like(q, dtype=torch.float32),
)
return run_tunable(
MyOp,
args,
config=config,
autotune=tune,
configs=configs,
)[0]This gives three intentional modes:
my_op(q, k) # conservative default
my_op(q, k, config=MyConfig(64, 128)) # force one specialization
my_op(q, k, tune=True) # input-aware candidate method
my_op(q, k, tune=True, configs=(cfg_a, cfg_b)) # explicit candidate overrideReject config= with tuning and reject configs= without tuning rather than silently ignoring either argument.
If target metadata is supplied explicitly, install it before generating candidates so configs(...)
and compilation observe the same target. On heterogeneous multi-GPU hosts, derive that target from
the runtime tensor's device rather than the process's current device. The installed target is
process-global and sticky; callers that temporarily override it must restore the previous target.
Inside compile(...):
jit_cache can encode them for warm process-local lookups before
constructing a persistent hash or cache path. When the call tuple is not the right specialization
identity, pass cache_key= a pure function returning a complete hashable tuple or frozen config.
Return the structural key itself, never hash(key); jit_cache adds function, target, and source
identity and uses the same structural key for memory and disk caching.cute.sym_int()/cute.sym_int64()
when their dependent strides also remain dynamic. Bake a dimension into compile_call(...) only
when it changes layout/stride address arithmetic, tiling, vectorization, block shape, shared-memory
sizing, compile-time control flow, or another generated-code decision.compile_tvm_ffi(...) a stable lowercase name encoding every static compile argument.
Class entrypoints may expose get_name() instead of passing name= explicitly.@staticmethod
@jit_cache
def compile(config: MyConfig):
num_elements = cute.sym_int()
source = cute.runtime.make_fake_compact_tensor(
..., (num_elements,), stride_order=(0,), assumed_align=16
)
destination = cute.runtime.make_fake_compact_tensor(
..., (num_elements,), stride_order=(0,), assumed_align=16
)
op = _MyOp()
return compile_tvm_ffi(
op._jit_entrypoint,
source,
destination,
config.threads,
config.tile_size,
name=op.get_name(config),
)compile_tvm_ffi owns the typed TVM-FFI option and fake environment stream. Do not append another stream or pass string compiler options at call sites. compile_call(...) returns exactly one tuple of positional static arguments; use (config,) when compile(...) accepts only a config.
Keep ordinary layouts on the int32 address path. At the parent wrapper, call
attn_gym._backends.cute.requires_int64_abi on every ABI-visible input, output, and optional tensor,
then pass the result through compile_call(...) as a static use_int64_offsets argument. The CuTe
predicate is intentionally stricter than reachable cosize: TVM-FFI must represent every declared
stride, including a stride larger than INT32_MAX on a size-one mode that cannot reach it.
The width bool must participate in the jit_cache key and profiler/artifact name. In the wide
specialization, use cute.sym_int64 for each dynamic fake-tensor dimension or stride involved in
addressing; the fake signature and runtime tensor ABI must agree. Inside the op, widen each dynamic
index or origin before its first potentially overflowing multiply or addition. Casting the final
layout stride, iterator offset, or pointer is too late. Audit manually rebuilt layouts, iterator
addition, chunk/program indices, and stores as well as ordinary tensor loads. Bounded routing arrays
may remain int32 when their values and their own addressing are independently proven safe.
Do not infer width from numel() and do not make int64 unconditional: wider arithmetic can add
instructions and register pressure. Validate predicate-only oversized singleton strides, force the
i64 specialization on small inputs and compare it with i32, and, when memory permits, execute an
active offset beyond INT32_MAX against an equivalent compact layout. See
test/test_kda_int64_offsets.py for the project pattern.
Validate reachable inputs and static semantics once at the eager boundary used by each path: the
public tune path or the private custom-op implementation. Validate again in compile(...), because
cache/compiler-worker calls can bypass both wrappers; downstream launch helpers may then rely on the
allocated runtime ABI. Keep FakeTensor-incompatible checks such as data_ptr() alignment inside the
opaque/eager launcher rather than a trace-time validator.
For tune=True, run_tunable should:
configs= iterable, otherwise call kernel.configs(*runtime_args) once.compile_call(...).Before enabling tuning, force every generated candidate through config= in correctness tests and compare it with an independent reference; a benchmark cannot detect a wrong fast candidate. Direct benchmarking also assumes repeatable launches.
run_tunable(...) performs a final launch after benchmarking. For a destructive or accumulating op,
call tune(...) directly, restore all inputs and outputs after it returns, then execute the winner once
through the ordinary explicit-config path. A benchmark callback that restores only between samples is
insufficient because it cannot prepare state for the final launch.
A fully warm run must not start the compiler process. It still benchmarks requested candidates unless a separate baked selector chooses one without tuning.
torch.compile BoundaryCache lookup, target discovery, compilation, and tuning are eager host operations. When a public
function must support strict Dynamo capture, hide the ordinary no-tune launcher behind a private
functional torch.library.custom_op and register a fake implementation whose output shapes derive
symbolically from input shapes and static scalars.
Project a config to schema-supported scalars at that boundary; decode it inside the opaque op. Use an
optional scalar for target-resolved automatic selection rather than querying the target in the traced
wrapper. Keep tune=True eager-only and bypass the custom op. If output shape changes at a static
bucket boundary, Dynamo may legitimately compile another graph even with dynamic=True.
Read or run copy_reads_example.py for a complete toy kernel using an input-aware ReadConfig search space and a single public copy_reads(...) entrypoint.
© meta-pytorch, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .agents/skills/cutedsl-tunable-kernel-template of meta-pytorch/attention-gym.
Open the folder on GitHubat commit 70fd810
Cutedsl Tunable Kernel Template next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cutedsl Tunable Kernel Template this skillmeta-pytorch/attention-gym | 1.3k | — | ~4.1k | Automated safety check: Pass | BSD-3-Clause | |
| Add Column Typesimstudioai/sim | 30k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Add Sgl Kernelsgl-project/sglang | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| Add Jit Kernelsgl-project/sglang | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | |
| Add Pallas Kernelmarin-community/marin | 3.9k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Rust Path Typesopeninterpreter/openinterpreter | 69k | 2 repos | ~605 | Automated safety check: Pass | Apache-2.0 |
simstudioai/sim
Add a new table column type to Sim — registry entry, icon, storage shape, coercion, and the behavioral hooks the grid and API read.
sgl-project/sglang
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
sgl-project/sglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups
marin-community/marin
Add or change a named Pallas/Mosaic kernel, including its reference implementation, correctness tests, wrapper, or requested tuning.
openinterpreter/openinterpreter
Rules for choosing Rust types for filesystem paths in new Codex code, covering protocol types, internal use and model tool arguments.
facebook/pyrefly
Port a PyTorch model to use pyrefly's tensor shape type system (Tensor[[B, C, H, W]], Int[T]).
meta-pytorch/attention-gym
Sets up an isolated per-worktree Python environment for attention-gym development using nightly PyTorch and the CI-mirroring uv flow.
meta-pytorch/attention-gym
Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.
Adds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking. Cutedsl Tunable Kernel Template is an agent skill from meta-pytorch/attention-gym. Adds an Attention Gym tuning adapter to a CuTeDSL op using typed input-aware configs, cached fake-tensor TVM-FFI compilation, parallel candidate compilation, and sequential GPU benchmarking.
Run `npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a claude-code`. Or copy the skill folder (.agents/skills/cutedsl-tunable-kernel-template in meta-pytorch/attention-gym) into .claude/skills/cutedsl-tunable-kernel-template in your project. Claude Code loads it when a task matches its description.
Run `npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a codex`. Or copy the skill folder (.agents/skills/cutedsl-tunable-kernel-template in meta-pytorch/attention-gym) into .agents/skills/cutedsl-tunable-kernel-template in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add meta-pytorch/attention-gym --skill cutedsl-tunable-kernel-template -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cutedsl-tunable-kernel-template, .gemini/skills/cutedsl-tunable-kernel-template, .github/skills/cutedsl-tunable-kernel-template and .opencode/skills/cutedsl-tunable-kernel-template in your project.
Going by SKILL.md and its folder, Cutedsl Tunable Kernel Template needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cutedsl Tunable Kernel Template is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cutedsl Tunable Kernel Template: Add Column Type (simstudioai/sim, 30k stars), Add Sgl Kernel (sgl-project/sglang, 37k stars), Add Jit Kernel (sgl-project/sglang, 37k stars) and Add Pallas Kernel (marin-community/marin, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
meta-pytorch (a GitHub organization) maintains it in meta-pytorch/attention-gym, which has 1,255 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: meta-pytorch/attention-gym on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.