Aoti Debug
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code.
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .claude/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .claude/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .agents/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .agents/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .cursor/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .cursor/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tilegym-cutile-autotuning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .gemini/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .gemini/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .github/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .github/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .opencode/skills/tilegym-cutile-autotuning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tilegym-cutile-autotuning" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-cutile-autotuning into .opencode/skills/tilegym-cutile-autotuning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-cutile-autotuning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tilegym-cutile-autotuningA skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code.
Tilegym Cutile Autotuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code, autotune in filenames, or correctness/performance issues in autotuned CuTile kernels. Covers: tune-once/cache/launch pattern, per-architecture configs (sm80–sm120), parameter space design (tile sizes, occupancy, numctas), and 7 common pitfalls with solutions.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including reference files and assets (for example `BENCHMARK.md`, `assets/examples/01_rmsnorm_occupancy_only/autotuned_launch.py` and `assets/examples/01_rmsnorm_occupancy_only/fixed_launch.py`).
It sits in Development. It works with CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
bashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tilegym Cutile Autotuning loads about 4.6k tokens when it runs, and up to ~33k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,811 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,811 words, ~4,617 tokens.
.claude/skills/tilegym-cutile-autotuning/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Add autotuning to CuTile kernels using the exhaustive_search API with tune-once/cache/direct-launch pattern.
Follow the decision tree to classify the kernel, design a search space, implement the tune-once/cache/launch pattern, and validate performance.
references/kernel-type-templates.md; prune to ≤ 30 configs in the final code via arch filters (directed exploration probes may temporarily exceed this — see Design Philosophy)exhaustive_search + cache + ct.launch following the Step-by-Step Workflow; handle in-place writes with split-buffer if neededDISABLE_AUTOTUNE=1references/search-strategies.md| What are you trying to do? | Go to |
|---|---|
| Add autotune to a new kernel (most common) | Quick Reference below → Workflow: Adding Autotune → references/kernel-type-templates.md (pick by kernel type: T1=elementwise, T2=in-place, T3=matmul, T4=persistent, T5=FMHA, T6=FP8, T7=grouped GEMM, T8=varlen attention, T9=dual-GEMM fusion) |
| Debug: data corruption / wrong results after first run | Pitfall #1 (In-Place Kernel) |
| Debug: autotune taking 5+ minutes | Pitfall #2 (Compilation Timeout) |
| Debug: search space generator returning zero configs | Pitfall #5 first; also check arch filters, size guards, and num_ctas constraints |
| Optimize an existing autotune config | Workflow: Optimizing an Existing Config |
Most CuTile kernels (elementwise, reduction, LayerNorm) need only occupancy tuning. Copy this pattern:
from types import SimpleNamespace
from cuda.tile.tune import exhaustive_search
import cuda.tile as ct
import torch
def _my_autotune_configs():
for occ in [1, 2, 4, 8]:
yield SimpleNamespace(occupancy=occ)
# Module-level cache: tune once, launch fast forever after
_autotune_cache = {}
def my_op(x, output):
stream = torch.cuda.current_stream()
NUM_SM = torch.cuda.get_device_properties(x.device).multi_processor_count
# Cache key: anything that affects optimal config (use str() for device)
cache_key = (x.shape, x.dtype, str(x.device))
if cache_key not in _autotune_cache:
configs = list(_my_autotune_configs())
result = exhaustive_search(
configs,
stream,
grid_fn=lambda cfg: (min(NUM_SM * cfg.occupancy, M), 1, 1),
kernel=my_kernel,
args_fn=lambda cfg: (x, output, ...),
hints_fn=lambda cfg: {"occupancy": cfg.occupancy},
)
best_cfg = result.best.config
tuned_kernel = my_kernel.replace_hints(occupancy=best_cfg.occupancy)
_autotune_cache[cache_key] = (best_cfg, tuned_kernel) # cache BOTH
cfg, tuned_kernel = _autotune_cache[cache_key]
grid = (min(NUM_SM * cfg.occupancy, M), 1, 1)
ct.launch(stream, grid, tuned_kernel, (x, output, ...))Key rules:
exhaustive_search runs only on first call per shape; subsequent calls use cached config + ct.launch with zero overheadexhaustive_search requires a Sequence (list/tuple) — convert generators with list()When to use this pattern: Kernel has fixed block size (not tile-size tunable). Includes: elementwise (SwiGLU, GeGLU), reduction (RMSNorm, LayerNorm), RoPE, and persistent kernels with heuristic block sizes (grouped GEMM).
For complex kernels (matmul with tile sizes, FMHA, FP8 with num_ctas), read the full guide below + kernel-type-templates.md.
⚠️ Three pitfalls catch almost everyone — check before submitting:
replace_hintson hot path? → Cache BOTH config AND kernel object fromexhaustive_search. Callingreplace_hints()every invocation recompiles (100–500× slower) → Pitfall #7- In-place kernel (writes back to input tensor)? → MUST use split-buffer pattern during search → Pitfall #1
- Search space empty? → Check arch filters and
num_ctasconstraints → Pitfall #5
Minimum coverage: On sm100+, FMHA/matmul/varlen search spaces must include both
num_ctas=1andnum_ctas=2. For core dimensions (tile sizes, occupancy), keep at least 2 distinct values even if unsure which is better — letexhaustive_searchdecide.
When to stop tuning: A mean speedup in [0.98, 1.02] means your current search space isn't helping — but doesn't mean no config will help. Before stopping, check whether you've covered the key dimensions for this kernel type (consult
references/kernel-type-templates.md). If the search space already covers the template's recommended dimensions and the best result is still noise-floor, then stop — further micro-adjustments won't help. If key dimensions are missing (e.g., never triednum_ctas=2for a dual-GEMM kernel), expand the search space rather than giving up.Once correctness tests pass and the autotuned kernel shows speedup over the fixed-config baseline, stop — do not re-run to "confirm". GPU kernel timing fluctuates ±5–10 % between invocations due to clock scaling and OS scheduling; a subsequent timing dip does not mean your code is wrong.
To improve speedup, only modify the autotune search space (configs, tile sizes, occupancy, num_ctas). Do not modify other code (Python wrapper, stream management, etc.) to chase speedup — kernel performance is determined by the config selection, not by host-side code.
references/ docs. For in-place kernels, also read Pitfall #1.references/ docs.5-step summary: Classify kernel → Design search space (parameter-space-design.md) → Implement using template (kernel-type-templates.md) → Validate with A/B test → Check Pitfall Checklist.
Reading references: Read only the reference relevant to your kernel type — e.g., for FMHA, read the Template 5 section in references/kernel-type-templates.md; for hardware constraints, read only the target architecture's section. Avoid reading all references end-to-end when a targeted lookup suffices.
Build a small, precise search space bottom-up — not a large space trimmed down. CuTile compilation is much heavier than Triton (~0.5-1s per config), so the final code should contain ≤ 30 configs. The approach is: classify the kernel type first, then construct only the relevant configs for that type and architecture.
Directed exploration during development: If the initial template configs yield speedup < 1.0, you may run a temporary larger probe (30–100 configs) via bash + python3 -c to identify which dimensions matter — but this probe must be directional, not a blind cartesian product. Use the kernel type classification to decide which dimensions to vary (e.g. for dual-GEMM, probe num_ctas × occupancy while fixing tile sizes; for FMHA, probe TILE_M × num_ctas while fixing TILE_N). Once the probe identifies the winning region, lock the final code's search space to ≤ 8 top candidates. Do NOT write the large probe into the source file — it is a one-shot diagnostic tool.
All kernels should have autotuning added. The question is not whether to autotune, but what dimensions to search:
What type of kernel is this?
├── Compute-bound (matmul, GEMM, FMHA) → Does it have multiple tunable dimensions (tile sizes)?
│ ├── YES → Is it a fused multi-GEMM kernel (dual-GEMM, e.g. Linear+GLUAct)?
│ │ ├── YES → Template 9: low occupancy (1–2), conservative tiles (2× SHMEM/register pressure)
│ │ └── NO → Full search: TILE_M × TILE_N × (TILE_K) × occupancy × num_ctas
│ │ (see matmul/FMHA templates in kernel-type-templates.md)
│ └── NO → Occupancy-only search: [1, 2, 4, 8]
│ (see Quick Reference above)
├── Balanced (LayerNorm, reduction + compute) →
│ Occupancy-only search: [1, 2, 4, 8]
│ Expected benefit: 2-15%
└── Memory-bound (CE Loss, pure elementwise) →
Occupancy-only search: [1, 2, 4, 8]
Expected benefit: 0-15% (varies by kernel; zero-cost after tuning)Why memory-bound kernels only search occupancy (not num_ctas or tile sizes):
num_ctas has zero benefit: num_ctas > 1 enables TMA multicast, where multiple CTAs share tile data in shared memory (e.g., matmul A/B tiles reused across CTAs). Memory-bound kernels use per-element ct.gather/ct.scatter with no tile reuse — multi-CTA cooperation adds overhead with no data sharing benefit.Evidence — CE Loss experiment: A 12-config search (occupancy × num_ctas) on Cross-Entropy Loss yielded only 2.5% gain (0.79x → 0.81x vs Triton). The
num_ctasdimension contributed nothing; the result was reverted because compilation cost outweighed the marginal benefit. Occupancy-only (4 configs) achieves the same result at 3x less compilation time.
Note on memory-bound kernels: Adding occupancy-only autotune is always worthwhile because:
Occupancy controls how many CTAs run concurrently per SM. Use this as a starting point when designing the occupancy search space:
| Occupancy Range | Best For | Example Kernels |
|---|---|---|
| 1–4 | Compute-bound (heavy math) | Complex transforms, matmul |
| 4–8 | Balanced (GEMM, TMA) | Matrix multiply, FMHA |
| 8–16 | Memory-bound (reductions) | Softmax, LayerNorm |
| 16–32 | Very light (copies, casts) | Type conversions, elementwise |
Use these ranges to seed your initial search space. For occupancy-only kernels, [1, 2, 4, 8] covers most cases — see Quick Reference above.
See references/api-reference.md for the full
exhaustive_search API surface — current signature, TuningResult, the
tune-once/cache/launch pattern, replace_hints, kernel hints, search_space
design, and grid_fn patterns.
See references/workflow.md for the end-to-end
workflow — adding autotune to a new kernel, handling existing
multi-architecture configs, integration with torch.autograd.Function,
cross-backend config transfer (Triton → CuTile), and optimizing an existing
config.
See references/pitfalls.md for the full list of
common pitfalls — in-place data corruption, compilation timeout, cold-cache
performance skew, NCU profiling interference, search_space generator
exhaustion, FP8 precision loss, and replace_hints recompilation on hot
paths.
This skill covers only autotune configuration: search space design, exhaustive_search invocation, caching, and ct.launch with tuned hints. It does not modify kernel code.
In scope (autotune config):
exhaustive_search() calls and result handlingkernel.replace_hints() for applying tuned hintsct.launch() with tuned kernelDISABLE_AUTOTUNE fallback pathOut of scope (kernel code modifications — do NOT make these changes):
After adding autotuning, the following kernel-level optimizations may yield additional gains. These are outside the scope of this skill — mention them to the user as potential next steps, but do not implement them as part of autotuning:
flush_to_zero=True + rounding_mode=APPROX can provide 34-72% improvement for FMHA-class kernels (set via environment variables TILEIR_ENABLE_FTZ=1 TILEIR_ENABLE_APPROX=1 or in kernel code). Causal chain: larger tiles initially decrease performance by 18-43% due to subnormal handling overhead; enabling FTZ+APPROX rescues this and flips the result to +34-72%. Math flags are therefore a prerequisite for large-tile configs to be effective on FMHA-class kernels.slice_hint, buffer_depth, copy_config — requires modifying kernel IR codect.load) instead of ct.gather; removing unnecessary bounds checks (check_bounds=False when safe)padding_value parameter instead of manual ct.where masking; removing safe_offsKey differences: Triton uses @triton.autotune decorator with Config(...) objects; CuTile uses exhaustive_search() with SimpleNamespace configs + separate cache + ct.launch. CuTile has no num_warps/num_stages (compiler decides) — only tile sizes + occupancy + num_ctas. CuTile compilation is heavier (keep ≤30 configs in final code). CuTile cache is user-managed in-memory (no automatic persistence). CuTile separates args_fn (kernel args) from hints_fn (compiler hints).
| Category | Document | Content |
|---|---|---|
| API Reference | api-reference.md | exhaustive_search signature, TuningResult, tune-once/cache/launch pattern, replace_hints, kernel hints, search_space design, grid_fn patterns |
| Workflow | workflow.md | End-to-end workflow: adding autotune to a new kernel, multi-architecture configs, torch.autograd.Function integration, Triton→CuTile transfer, optimizing existing configs |
| Pitfalls | pitfalls.md | Common pitfalls: in-place corruption, compilation timeout, cold-cache skew, NCU interference, search_space exhaustion, FP8 precision, replace_hints recompilation |
| Parameter Design | parameter-space-design.md | Per-kernel-type parameter spaces, cross-arch patterns, grid_fn patterns, pruning rules |
| Search Strategies | search-strategies.md | Exhaustive search, A/B test methodology, DISABLE_AUTOTUNE pattern |
| Templates | kernel-type-templates.md | Copy-paste autotune templates for 8 kernel types |
| Hardware | hardware-constraints.md | Per-architecture constraints, tile size ranges, num_ctas rules, TMA requirements |
Key files: ops/cutile/matmul.py (matmul autotune), ops/cutile/attention.py (FMHA autotune), suites/unsloth/cutile/ct_ops.py (shared autotune_configs() occupancy=[1,2,4,8]), suites/unsloth/cutile/swiglu.py (elementwise example), suites/unsloth/cutile/rope_embedding.py (split-buffer pattern), suites/unsloth/cutile/grouped_gemm.py (persistent GEMM, occupancy-only).
Each example shows the before → after pattern: fixed_launch.py (hardcoded ct.launch) and autotuned_launch.py (refactored to tune-once/cache/launch).
| Directory | Kernel | Autotune Pattern | Complexity | Key Teaching Point |
|---|---|---|---|---|
assets/examples/01_rmsnorm_occupancy_only/ | RMSNorm (reduction) | Occupancy-only [1,2,4,8] | Low | Most common pattern — no tile tuning, just find best occupancy. Grid = NUM_SM * cfg.occupancy. Not in-place. |
assets/examples/02_matmul_full_search/ | GEMM C=A@B | Full: TILE_M/N/K + occupancy + num_ctas (sm90+) | High | Compute-bound kernel with multiple tunable dimensions. args_fn passes tile sizes as ct.Constant[int]. grid_fn depends on cfg. ≤30 configs. |
assets/examples/03_rope_inplace_splitbuffer/ | RoPE embedding (in-place) | Occupancy-only, with split-buffer | Medium | In-place kernel MUST use split-buffer during search to avoid corruption. Search writes to scratch; final ct.launch uses real in-place args. |
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (references, assets) in skills/tilegym-cutile-autotuning of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Tilegym Cutile Autotuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tilegym Cutile Autotuning this skillNVIDIA/skills | 3.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| Debug Distributed Hangsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Create Cuda Python Pull RequestNVIDIA/cuda-python | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| ONNX Runtime Source Buildmicrosoft/onnxruntime | 22k | — | ~1.4k | Automated safety check: Pass | MIT |
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
NVIDIA/cuda-python
Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code. Tilegym Cutile Autotuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when adding, modifying, optimizing, or debugging CuTile autotuning code.
Tilegym Cutile Autotuning fits situations like: debugging CuTile autotuning code; signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code; autotune in filenames; correctness/performance issues in autotuned CuTile kernels.
Run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a claude-code`. Or copy the skill folder (skills/tilegym-cutile-autotuning in NVIDIA/skills) into .claude/skills/tilegym-cutile-autotuning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a codex`. Or copy the skill folder (skills/tilegym-cutile-autotuning in NVIDIA/skills) into .agents/skills/tilegym-cutile-autotuning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilegym-cutile-autotuning, .gemini/skills/tilegym-cutile-autotuning, .github/skills/tilegym-cutile-autotuning and .opencode/skills/tilegym-cutile-autotuning in your project.
Going by SKILL.md and its folder, Tilegym Cutile Autotuning needs Python for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tilegym Cutile Autotuning is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tilegym Cutile Autotuning: Aoti Debug (pytorch/pytorch, 104k stars), Debug Distributed Hang (sgl-project/sglang, 37k stars), Create Cuda Python Pull Request (NVIDIA/cuda-python, 3.4k stars) and CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.