Official agent skill

Tilegym Cutile Autotuning

by NVIDIA in NVIDIA/skills

A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code.

OfficialApache-2.0Auto-check passedDevelopment

Install Tilegym Cutile Autotuning

skills CLI
$ npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tilegym-cutile-autotuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilegym-cutile-autotuning .claude/skills/tilegym-cutile-autotuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tilegym-cutile-autotuning
GitHub stars
3.5k
Token cost
~4.6k tokens
SKILL.md length
1,811 words
Files
18 (incl. references, assets)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code.

  • Works in 6 steps: Classify — use the Decision Tree to… → Design search space — select the… → Implement — add exhaustive_search +… → …
  • Debugging CuTile autotuning code
  • SKILL.md covers Instructions, Task Router — Jump to What You…, Quick Reference —… and Reading Guide, plus 12 more sections
  • Runs Python scripts from its folder; calls bash

What it does

Tilegym Cutile Autotuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code, autotune in filenames, or correctness/performance issues in autotuned CuTile kernels. Covers: tune-once/cache/launch pattern, per-architecture configs (sm80–sm120), parameter space design (tile sizes, occupancy, numctas), and 7 common pitfalls with solutions.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including reference files and assets (for example `BENCHMARK.md`, `assets/examples/01_rmsnorm_occupancy_only/autotuned_launch.py` and `assets/examples/01_rmsnorm_occupancy_only/fixed_launch.py`).

It sits in Development. It works with CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Debugging CuTile autotuning code
  • Signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code
  • Autotune in filenames
  • Correctness/performance issues in autotuned CuTile kernels

Example prompts

  • “/tilegym-cutile-autotuning”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Classify — use the Decision Tree to determine search dimensions (occupancy-only vs full tile search)
  2. Design search space — select the matching template from references/kernel-type-templates.md; prune to ≤ 30 configs in the final code via…
  3. Implement — add exhaustive_search + cache + ct.launch following the Step-by-Step Workflow; handle in-place writes with split-buffer if…
  4. Test — run correctness with autotune enabled and with DISABLE_AUTOTUNE=1
  5. Validate — A/B benchmark against fixed best-known config; see references/search-strategies.md
  6. Shrink — prune dead-weight configs that never win, targeting ≤ 8 configs per architecture to minimize compilation cost (Step 10)

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tilegym Cutile Autotuning loads about 4.6k tokens when it runs, and up to ~33k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,811 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~33k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,811 words, ~4,617 tokens.

Download SKILL.mdSave it as .claude/skills/tilegym-cutile-autotuning/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
tilegym-cutile-autotuning
description
Use when adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: `exhaustive_search` / `replace_hints` / `hints_fn` / `cuda.tile.tune` in code, `autotune` in filenames, or correctness/performance issues in autotuned CuTile kernels. Covers: tune-once/cache/launch pattern, per-architecture configs (sm80–sm120), parameter space design (tile sizes, occupancy, num_ctas), and 7 common pitfalls with solutions.
license
CC-BY-4.0 AND Apache-2.0

CuTile Autotuning

Add autotuning to CuTile kernels using the exhaustive_search API with tune-once/cache/direct-launch pattern.

Instructions

Follow the decision tree to classify the kernel, design a search space, implement the tune-once/cache/launch pattern, and validate performance.

  1. Classify — use the Decision Tree to determine search dimensions (occupancy-only vs full tile search)
  2. Design search space — select the matching template from references/kernel-type-templates.md; prune to ≤ 30 configs in the final code via arch filters (directed exploration probes may temporarily exceed this — see Design Philosophy)
  3. Implement — add exhaustive_search + cache + ct.launch following the Step-by-Step Workflow; handle in-place writes with split-buffer if needed
  4. Test — run correctness with autotune enabled and with DISABLE_AUTOTUNE=1
  5. Validate — A/B benchmark against fixed best-known config; see references/search-strategies.md
  6. Shrink — prune dead-weight configs that never win, targeting ≤ 8 configs per architecture to minimize compilation cost (Step 10)

Task Router — Jump to What You Need

What are you trying to do?Go to
Add autotune to a new kernel (most common)Quick Reference below → Workflow: Adding Autotune → references/kernel-type-templates.md (pick by kernel type: T1=elementwise, T2=in-place, T3=matmul, T4=persistent, T5=FMHA, T6=FP8, T7=grouped GEMM, T8=varlen attention, T9=dual-GEMM fusion)
Debug: data corruption / wrong results after first runPitfall #1 (In-Place Kernel)
Debug: autotune taking 5+ minutesPitfall #2 (Compilation Timeout)
Debug: search space generator returning zero configsPitfall #5 first; also check arch filters, size guards, and num_ctas constraints
Optimize an existing autotune configWorkflow: Optimizing an Existing Config

Quick Reference — Occupancy-Only Autotune (Tune-Once/Cache/Launch)

Most CuTile kernels (elementwise, reduction, LayerNorm) need only occupancy tuning. Copy this pattern:

python
from types import SimpleNamespace
from cuda.tile.tune import exhaustive_search
import cuda.tile as ct
import torch

def _my_autotune_configs():
    for occ in [1, 2, 4, 8]:
        yield SimpleNamespace(occupancy=occ)

# Module-level cache: tune once, launch fast forever after
_autotune_cache = {}

def my_op(x, output):
    stream = torch.cuda.current_stream()
    NUM_SM = torch.cuda.get_device_properties(x.device).multi_processor_count

    # Cache key: anything that affects optimal config (use str() for device)
    cache_key = (x.shape, x.dtype, str(x.device))

    if cache_key not in _autotune_cache:
        configs = list(_my_autotune_configs())
        result = exhaustive_search(
            configs,
            stream,
            grid_fn=lambda cfg: (min(NUM_SM * cfg.occupancy, M), 1, 1),
            kernel=my_kernel,
            args_fn=lambda cfg: (x, output, ...),
            hints_fn=lambda cfg: {"occupancy": cfg.occupancy},
        )
        best_cfg = result.best.config
        tuned_kernel = my_kernel.replace_hints(occupancy=best_cfg.occupancy)
        _autotune_cache[cache_key] = (best_cfg, tuned_kernel)  # cache BOTH

    cfg, tuned_kernel = _autotune_cache[cache_key]
    grid = (min(NUM_SM * cfg.occupancy, M), 1, 1)
    ct.launch(stream, grid, tuned_kernel, (x, output, ...))

Key rules:

  • Tune once, cache, launch directly — exhaustive_search runs only on first call per shape; subsequent calls use cached config + ct.launch with zero overhead
  • For in-place kernels use split-buffer during search (separate input/output tensors)
  • Keep ≤ 30 configs in final code (see Design Philosophy for temporary directed probes)
  • exhaustive_search requires a Sequence (list/tuple) — convert generators with list()
  • Search space must include the original fixed config — this guarantees autotuning never makes performance worse

When to use this pattern: Kernel has fixed block size (not tile-size tunable). Includes: elementwise (SwiGLU, GeGLU), reduction (RMSNorm, LayerNorm), RoPE, and persistent kernels with heuristic block sizes (grouped GEMM).

For complex kernels (matmul with tile sizes, FMHA, FP8 with num_ctas), read the full guide below + kernel-type-templates.md.

⚠️ Three pitfalls catch almost everyone — check before submitting:

  • replace_hints on hot path? → Cache BOTH config AND kernel object from exhaustive_search. Calling replace_hints() every invocation recompiles (100–500× slower) → Pitfall #7
  • In-place kernel (writes back to input tensor)? → MUST use split-buffer pattern during search → Pitfall #1
  • Search space empty? → Check arch filters and num_ctas constraints → Pitfall #5

Minimum coverage: On sm100+, FMHA/matmul/varlen search spaces must include both num_ctas=1 and num_ctas=2. For core dimensions (tile sizes, occupancy), keep at least 2 distinct values even if unsure which is better — let exhaustive_search decide.

When to stop tuning: A mean speedup in [0.98, 1.02] means your current search space isn't helping — but doesn't mean no config will help. Before stopping, check whether you've covered the key dimensions for this kernel type (consult references/kernel-type-templates.md). If the search space already covers the template's recommended dimensions and the best result is still noise-floor, then stop — further micro-adjustments won't help. If key dimensions are missing (e.g., never tried num_ctas=2 for a dual-GEMM kernel), expand the search space rather than giving up.

Once correctness tests pass and the autotuned kernel shows speedup over the fixed-config baseline, stop — do not re-run to "confirm". GPU kernel timing fluctuates ±5–10 % between invocations due to clock scaling and OS scheduling; a subsequent timing dip does not mean your code is wrong.

To improve speedup, only modify the autotune search space (configs, tile sizes, occupancy, num_ctas). Do not modify other code (Python wrapper, stream management, etc.) to chase speedup — kernel performance is determined by the config selection, not by host-side code.

Reading Guide

  • Occupancy-only kernels (elementwise, reduction, persistent with fixed block sizes): Quick Reference + Pitfall Checklist is sufficient — skip references/ docs. For in-place kernels, also read Pitfall #1.
  • Complex kernels (matmul with tunable tile sizes, FMHA, FP8 with num_ctas): Quick Reference → Decision Tree → API Reference → Step-by-Step Workflow → relevant references/ docs.

5-step summary: Classify kernel → Design search space (parameter-space-design.md) → Implement using template (kernel-type-templates.md) → Validate with A/B test → Check Pitfall Checklist.

Reading references: Read only the reference relevant to your kernel type — e.g., for FMHA, read the Template 5 section in references/kernel-type-templates.md; for hardware constraints, read only the target architecture's section. Avoid reading all references end-to-end when a targeted lookup suffices.

Design Philosophy

Build a small, precise search space bottom-up — not a large space trimmed down. CuTile compilation is much heavier than Triton (~0.5-1s per config), so the final code should contain ≤ 30 configs. The approach is: classify the kernel type first, then construct only the relevant configs for that type and architecture.

Directed exploration during development: If the initial template configs yield speedup < 1.0, you may run a temporary larger probe (30–100 configs) via bash + python3 -c to identify which dimensions matter — but this probe must be directional, not a blind cartesian product. Use the kernel type classification to decide which dimensions to vary (e.g. for dual-GEMM, probe num_ctas × occupancy while fixing tile sizes; for FMHA, probe TILE_M × num_ctas while fixing TILE_N). Once the probe identifies the winning region, lock the final code's search space to ≤ 8 top candidates. Do NOT write the large probe into the source file — it is a one-shot diagnostic tool.

Decision Tree: What Search Dimensions Does This Kernel Need?

All kernels should have autotuning added. The question is not whether to autotune, but what dimensions to search:

What type of kernel is this?
├── Compute-bound (matmul, GEMM, FMHA) → Does it have multiple tunable dimensions (tile sizes)?
│   ├── YES → Is it a fused multi-GEMM kernel (dual-GEMM, e.g. Linear+GLUAct)?
│   │   ├── YES → Template 9: low occupancy (1–2), conservative tiles (2× SHMEM/register pressure)
│   │   └── NO  → Full search: TILE_M × TILE_N × (TILE_K) × occupancy × num_ctas
│   │             (see matmul/FMHA templates in kernel-type-templates.md)
│   └── NO  → Occupancy-only search: [1, 2, 4, 8]
│             (see Quick Reference above)
├── Balanced (LayerNorm, reduction + compute) →
│   Occupancy-only search: [1, 2, 4, 8]
│   Expected benefit: 2-15%
└── Memory-bound (CE Loss, pure elementwise) →
    Occupancy-only search: [1, 2, 4, 8]
    Expected benefit: 0-15% (varies by kernel; zero-cost after tuning)

Why memory-bound kernels only search occupancy (not num_ctas or tile sizes):

  • num_ctas has zero benefit: num_ctas > 1 enables TMA multicast, where multiple CTAs share tile data in shared memory (e.g., matmul A/B tiles reused across CTAs). Memory-bound kernels use per-element ct.gather/ct.scatter with no tile reuse — multi-CTA cooperation adds overhead with no data sharing benefit.
  • Tile sizes are pre-determined: BLOCK_SIZE for memory-bound kernels is determined by offline sweep (e.g., 1024 is globally optimal on B200 across [256, 512, 1024, 2048, 4096, 8192]). This is a constant, not a runtime tunable.
  • Occupancy is the only effective knob: Higher occupancy lets the GPU hide memory latency by switching to another CTA while one is stalled on a memory request.

Evidence — CE Loss experiment: A 12-config search (occupancy × num_ctas) on Cross-Entropy Loss yielded only 2.5% gain (0.79x → 0.81x vs Triton). The num_ctas dimension contributed nothing; the result was reverted because compilation cost outweighed the marginal benefit. Occupancy-only (4 configs) achieves the same result at 3x less compilation time.

Note on memory-bound kernels: Adding occupancy-only autotune is always worthwhile because:

  • The tune-once/cache/launch pattern has zero runtime overhead after the first call
  • The search space is tiny (4 configs, ~2-4s compilation)
  • Even small improvements have value at scale
Show full SKILL.md (695 more words)Show less

Occupancy Selection Guide

Occupancy controls how many CTAs run concurrently per SM. Use this as a starting point when designing the occupancy search space:

Occupancy RangeBest ForExample Kernels
1–4Compute-bound (heavy math)Complex transforms, matmul
4–8Balanced (GEMM, TMA)Matrix multiply, FMHA
8–16Memory-bound (reductions)Softmax, LayerNorm
16–32Very light (copies, casts)Type conversions, elementwise

Use these ranges to seed your initial search space. For occupancy-only kernels, [1, 2, 4, 8] covers most cases — see Quick Reference above.

exhaustive_search API Reference

See references/api-reference.md for the full exhaustive_search API surface — current signature, TuningResult, the tune-once/cache/launch pattern, replace_hints, kernel hints, search_space design, and grid_fn patterns.

Step-by-Step Workflow

See references/workflow.md for the end-to-end workflow — adding autotune to a new kernel, handling existing multi-architecture configs, integration with torch.autograd.Function, cross-backend config transfer (Triton → CuTile), and optimizing an existing config.

Pitfall Checklist

See references/pitfalls.md for the full list of common pitfalls — in-place data corruption, compilation timeout, cold-cache performance skew, NCU profiling interference, search_space generator exhaustion, FP8 precision loss, and replace_hints recompilation on hot paths.

Scope and Boundaries

This skill covers only autotune configuration: search space design, exhaustive_search invocation, caching, and ct.launch with tuned hints. It does not modify kernel code.

In scope (autotune config):

  • Search space generator functions
  • exhaustive_search() calls and result handling
  • kernel.replace_hints() for applying tuned hints
  • Cache logic (key design, dict management)
  • ct.launch() with tuned kernel
  • DISABLE_AUTOTUNE fallback path

Out of scope (kernel code modifications — do NOT make these changes):

  • Math flags (flush_to_zero, rounding_mode)
  • Performance Hints (slice_hint, buffer_depth, copy_config)
  • Memory access patterns (2D→1D gather/scatter conversion)
  • Codegen optimizations (safe_offs → padding_value)
  • Algorithm changes (K-loop split, load balancing)

Further Optimization Suggestions

After adding autotuning, the following kernel-level optimizations may yield additional gains. These are outside the scope of this skill — mention them to the user as potential next steps, but do not implement them as part of autotuning:

  • Math flags: flush_to_zero=True + rounding_mode=APPROX can provide 34-72% improvement for FMHA-class kernels (set via environment variables TILEIR_ENABLE_FTZ=1 TILEIR_ENABLE_APPROX=1 or in kernel code). Causal chain: larger tiles initially decrease performance by 18-43% due to subnormal handling overhead; enabling FTZ+APPROX rescues this and flips the result to +34-72%. Math flags are therefore a prerequisite for large-tile configs to be effective on FMHA-class kernels.
  • Performance Hints: slice_hint, buffer_depth, copy_config — requires modifying kernel IR code
  • Memory access patterns: Using TMA loads (ct.load) instead of ct.gather; removing unnecessary bounds checks (check_bounds=False when safe)
  • Codegen quality: Using padding_value parameter instead of manual ct.where masking; removing safe_offs
  • Algorithm restructuring: K-loop split, load balancing, algebraic simplification

Differences from Triton Autotune

Key differences: Triton uses @triton.autotune decorator with Config(...) objects; CuTile uses exhaustive_search() with SimpleNamespace configs + separate cache + ct.launch. CuTile has no num_warps/num_stages (compiler decides) — only tile sizes + occupancy + num_ctas. CuTile compilation is heavier (keep ≤30 configs in final code). CuTile cache is user-managed in-memory (no automatic persistence). CuTile separates args_fn (kernel args) from hints_fn (compiler hints).

Reference Documents

CategoryDocumentContent
API Referenceapi-reference.mdexhaustive_search signature, TuningResult, tune-once/cache/launch pattern, replace_hints, kernel hints, search_space design, grid_fn patterns
Workflowworkflow.mdEnd-to-end workflow: adding autotune to a new kernel, multi-architecture configs, torch.autograd.Function integration, Triton→CuTile transfer, optimizing existing configs
Pitfallspitfalls.mdCommon pitfalls: in-place corruption, compilation timeout, cold-cache skew, NCU interference, search_space exhaustion, FP8 precision, replace_hints recompilation
Parameter Designparameter-space-design.mdPer-kernel-type parameter spaces, cross-arch patterns, grid_fn patterns, pruning rules
Search Strategiessearch-strategies.mdExhaustive search, A/B test methodology, DISABLE_AUTOTUNE pattern
Templateskernel-type-templates.mdCopy-paste autotune templates for 8 kernel types
Hardwarehardware-constraints.mdPer-architecture constraints, tile size ranges, num_ctas rules, TMA requirements

Source Code References

Key files: ops/cutile/matmul.py (matmul autotune), ops/cutile/attention.py (FMHA autotune), suites/unsloth/cutile/ct_ops.py (shared autotune_configs() occupancy=[1,2,4,8]), suites/unsloth/cutile/swiglu.py (elementwise example), suites/unsloth/cutile/rope_embedding.py (split-buffer pattern), suites/unsloth/cutile/grouped_gemm.py (persistent GEMM, occupancy-only).

Worked Examples

Each example shows the before → after pattern: fixed_launch.py (hardcoded ct.launch) and autotuned_launch.py (refactored to tune-once/cache/launch).

DirectoryKernelAutotune PatternComplexityKey Teaching Point
assets/examples/01_rmsnorm_occupancy_only/RMSNorm (reduction)Occupancy-only [1,2,4,8]LowMost common pattern — no tile tuning, just find best occupancy. Grid = NUM_SM * cfg.occupancy. Not in-place.
assets/examples/02_matmul_full_search/GEMM C=A@BFull: TILE_M/N/K + occupancy + num_ctas (sm90+)HighCompute-bound kernel with multiple tunable dimensions. args_fn passes tile sizes as ct.Constant[int]. grid_fn depends on cfg. ≤30 configs.
assets/examples/03_rope_inplace_splitbuffer/RoPE embedding (in-place)Occupancy-only, with split-bufferMediumIn-place kernel MUST use split-buffer during search to avoid corruption. Search writes to scratch; final ct.launch uses real in-place args.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (references, assets) in skills/tilegym-cutile-autotuning of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • assets/examples/01_rmsnorm_occupancy_only/autotuned_launch.py
  • assets/examples/01_rmsnorm_occupancy_only/fixed_launch.py
  • assets/examples/02_matmul_full_search/autotuned_launch.py
  • assets/examples/02_matmul_full_search/fixed_launch.py
  • assets/examples/03_rope_inplace_splitbuffer/autotuned_launch.py
  • assets/examples/03_rope_inplace_splitbuffer/fixed_launch.py
  • evals/evals.json
  • references/api-reference.md
  • references/hardware-constraints.md
  • references/kernel-type-templates.md
  • references/parameter-space-design.md
  • references/pitfalls.md
  • … and 4 more

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Tilegym Cutile Autotuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tilegym Cutile Autotuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tilegym Cutile Autotuning this skillNVIDIA/skills3.5k—~4.6kAutomated safety check: PassApache-2.0
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
Create Cuda Python Pull RequestNVIDIA/cuda-python3.4k—~1.1kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
ONNX Runtime Source Buildmicrosoft/onnxruntime22k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.

    3.4k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • ONNX Runtime Source Build

    microsoft/onnxruntime

    Official

    Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

    22k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Works with

Categories

Questions about Tilegym Cutile Autotuning

What does Tilegym Cutile Autotuning do?

A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code. Tilegym Cutile Autotuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when adding, modifying, optimizing, or debugging CuTile autotuning code.

When should I use Tilegym Cutile Autotuning?

Tilegym Cutile Autotuning fits situations like: debugging CuTile autotuning code; signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code; autotune in filenames; correctness/performance issues in autotuned CuTile kernels.

How do I install Tilegym Cutile Autotuning in Claude Code?

Run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a claude-code`. Or copy the skill folder (skills/tilegym-cutile-autotuning in NVIDIA/skills) into .claude/skills/tilegym-cutile-autotuning in your project. Claude Code loads it when a task matches its description.

How do I install Tilegym Cutile Autotuning in Codex?

Run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a codex`. Or copy the skill folder (skills/tilegym-cutile-autotuning in NVIDIA/skills) into .agents/skills/tilegym-cutile-autotuning in your project. Codex loads it when a task matches its description.

Can I use Tilegym Cutile Autotuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tilegym-cutile-autotuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilegym-cutile-autotuning, .gemini/skills/tilegym-cutile-autotuning, .github/skills/tilegym-cutile-autotuning and .opencode/skills/tilegym-cutile-autotuning in your project.

What does Tilegym Cutile Autotuning need to run?

Going by SKILL.md and its folder, Tilegym Cutile Autotuning needs Python for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: Python 3.

Does Tilegym Cutile Autotuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tilegym Cutile Autotuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tilegym Cutile Autotuning use?

Tilegym Cutile Autotuning is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tilegym Cutile Autotuning use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.

What are the alternatives to Tilegym Cutile Autotuning?

Skills that share tags, products or a category with Tilegym Cutile Autotuning: Aoti Debug (pytorch/pytorch, 104k stars), Debug Distributed Hang (sgl-project/sglang, 37k stars), Create Cuda Python Pull Request (NVIDIA/cuda-python, 3.4k stars) and CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tilegym Cutile Autotuning?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.