Official agent skill

Tilegym Improve Cutile Kernel Perf

by NVIDIA in NVIDIA/skills

Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning.

OfficialApache-2.0Auto-check passedDevelopment

Install Tilegym Improve Cutile Kernel Perf

skills CLI
$ npx skills add NVIDIA/skills --skill tilegym-improve-cutile-kernel-perf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tilegym-improve-cutile-kernel-perf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilegym-improve-cutile-kernel-perf .claude/skills/tilegym-improve-cutile-kernel-perf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tilegym-improve-cutile-kernel-perf
GitHub stars
3.5k
Token cost
~2k tokens
SKILL.md length
781 words
Files
11 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning.

  • Works in 7 steps: Create a fresh git branch: Propose a… → Locate the target kernel → Classify the kernel → …
  • Asked to optimize cutile kernel
  • SKILL.md covers Instructions, Setup, Experimentation and The experiment loop
  • Calls python and git

What it does

Tilegym Improve Cutile Kernel Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning. Covers tile sizes, occupancy, autotune configs, TMA, latency hints, persistent scheduling, numctas, flushtozero, and IR-level debugging. Use when asked to "optimize cutile kernel", "improve kernel perf", "tune cutile performance", "make kernel faster", or iteratively benchmark and refine a cuTile GPU kernel in the TileGym project.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/cutile-api-reference.md`).

It sits in Development, covering Performance optimization. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Asked to optimize cutile kernel
  • Improve kernel perf
  • Tune cutile performance
  • Make kernel faster

Example prompts

  • “optimize cutile kernel”
  • “improve kernel perf”
  • “tune cutile performance”
  • “/tilegym-improve-cutile-kernel-perf”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Create a fresh git branch: Propose a branch name, e.g., cutile-perf-- from current branch. Checkout git checkout -b
  2. Locate the target kernel
  3. Classify the kernel
  4. Check GPU environment
  5. Study related references
  6. Create @sandbox/perf_results.md to track progress. The first run will write a baseline
  7. Confirm and go: Once you get confirmation, kick off the experimentation

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tilegym Improve Cutile Kernel Perf loads about 2k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 781 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~24k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 781 words, ~1,952 tokens.

Download SKILL.mdSave it as .claude/skills/tilegym-improve-cutile-kernel-perf/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
tilegym-improve-cutile-kernel-perf
description
Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning. Covers tile sizes, occupancy, autotune configs, TMA, latency hints, persistent scheduling, num_ctas, flush_to_zero, and IR-level debugging. Use when asked to "optimize cutile kernel", "improve kernel perf", "tune cutile performance", "make kernel faster", or iteratively benchmark and refine a cuTile GPU kernel in the TileGym project.
license
CC-BY-4.0 AND Apache-2.0
metadata.author
TileGym Team <TileGym@nvidia.com>
metadata.version
2026.4.11
metadata.environment
IDE: Claude Code, Cursor (Agent mode); model: Opus 4.6
metadata.requires
GPU node Blackwell, Hopper and Ampere for benchmarking
metadata.tags
cutile, performance, optimization, kernel, profiling

Iterative cuTile Kernel Performance Optimization

Systematically profile, diagnose bottlenecks, and iteratively tune a cuTile kernel's performance in the TileGym repository.

Instructions

Follow the three phases in order: Setup the environment and baseline, run the Experimentation loop with a tracked log, then iterate The experiment loop until perf goals are met or further gains plateau.

Setup

Work with user to prepare optimization environment:

  1. Create a fresh git branch: Propose a branch name, e.g., cutile-perf-<kernel_name>-<date> from current branch. Checkout git checkout -b <branch name>

  2. Locate the target kernel:

    • cuTile kernels live under src/tilegym/suites/<suite>/cutile/ or src/tilegym/ops/cutile/
    • Read the kernel file and identify: the @ct.kernel decorated function(s), the launch wrapper (ct.launch() or ct_experimental.autotune_launch()), the @register_impl registration, and current autotune configs (if any)
  3. Classify the kernel:

    • Arithmetic Intensity < 10 -> Memory-bound
    • Arithmetic Intensity 10-50 -> Balanced
    • Arithmetic Intensity > 50 -> Compute-bound

    Note: classification is only used to pick the optimization priority order in the experiment loop. The core metric is always latency (ms).

  4. Check GPU environment:

    • Ensure a GPU node (Blackwell or Ampere GPU) is available
    • All subsequent benchmark commands should run on the GPU node
  5. Study related references:

    • references/optimization-playbook.md: Step-by-step recipes for each optimization (A through J) with before/after code examples
    • references/perf-knobs-catalog.md: Complete catalog of all tunable parameters (TMA, persistent scheduling, occupancy, tile sizes, latency hints, etc.)
    • references/cutile-api-reference.md: cuTile API reference and 18 critical rules
    • references/performance-model.md: Roofline/performance model, bottleneck diagnosis, autotuning
    • references/ir-dump-guide.md: IR dump, analysis, and error diagnosis
    • references/cutile-patterns-reference.md: Common cuTile patterns and conversion quick-reference
  6. Create @sandbox/perf_results.md to track progress. The first run will write a baseline

  7. Confirm and go: Once you get confirmation, kick off the experimentation

Experimentation

Every experiment iteration applies ONE optimization to the target kernel, verifies correctness, re-benchmarks, and records results. Each iteration should be enforced to finish within 10 minutes.

The goal
  • Improve the core metric: reduce latency (ms)
  • Subject to the core constraint: Correctness shall not regress — every optimization MUST preserve numerical correctness. latency (ms) shall not regress > 2% compared to baseline.
What you can change
  • The target kernel file under src/tilegym/suites/<suite>/cutile/ or src/tilegym/ops/cutile/: kernel body, tile sizes, occupancy, num_ctas, TMA usage, latency hints, flush_to_zero, autotune configs, persistent scheduling, and other cuTile-specific parameters
  • The kernel's launch wrapper: grid computation, autotune config space
  • @sandbox/: Feel free to add new files or modify files created by you, but don't check to git
What you can NOT change
  • Kernel functional semantics (inputs, outputs, and numerical behavior within tolerance)
  • Test infrastructure and benchmark harness
  • Anything not listed above
What to expect from experiment outputs
Correctness test:
bash
python -m pytest tests/suites/.../test_<kernel_name>.py -k "test_ and cutile and not test_perf" -v
Performance benchmark:

For each iteration:

  1. Run pytest benchmark: python -m pytest ... --print-record → extract latency (ms)
  2. Record latency in perf_results.md

Benchmark cmdlines:

bash
python -m pytest tests/suites/.../test_<kernel_name>.py -k "test_perf and cutile" --print-record -v

latency sample:

Cutile: {'forward': {'mean': 3.7903138461538455, 'std': 0.0016941310873207053, 'rel_std': 0.044696327430505396, 'median': 3.789880999999999, 'min': 3.7883389999999992, 'max': 3.7941230000000004, 'nrep': 13, 'peak_mem_mb': 913}} ms
Show full SKILL.md (338 more words)Show less
Track experiment progress

Use @sandbox/perf_results.md to record each iteration's results. It should only contain a Markdown table with 5 columns:

  • iteration: iteration number, starting from 0 (baseline)
  • optimization: what was applied (e.g., "baseline", "TMA replace gather", "persistent scheduling")
  • latency_ms: kernel latency in milliseconds, six decimal points
  • correctness: PASS or FAIL
  • status: Whether this iteration was keep, revert, or crash

Example content:

markdown
| iteration | optimization       | latency_ms | correctness | status |
|----------:|:-------------------|-----------:|:------------|-------:|
| 0         | baseline           |   0.820000 | PASS        | keep   |
| 1         | TMA replace gather |   0.390000 | PASS        | keep   |

Create the tabular header if the file was empty. Append one line for each iteration.

The baseline

The first iteration (iteration 0) will not change any code and simply run the correctness test and performance benchmark. Results will be listed at the first row as baseline.

The experiment loop

Core methodology is to apply ONE optimization per iteration from the playbook, verify correctness, benchmark, and decide whether to keep or revert. Try one optimization at a time, and have clean experiment records.

LOOP:

  1. Check git status: Current git branch/commit we're on

  2. Select and apply ONE optimization from references/optimization-playbook.md:

  3. Verify correctness — if fails, revert immediately. Common causes: flush_to_zero/rounding_mode=APPROX changed results, tile size OOB, allow_tma=False semantics, persistent loop bound error

  4. Re-benchmark and compare against current baseline

  5. Git commit

  6. Record results to @sandbox/perf_results.md

  7. Decision rules:

    OutcomeAction
    Improvement(latency (ms)) >= 5%Accept as new baseline, continue
    Improvement 2-5%Accept, lower priority for next iteration
    Improvement < 2%Accept but stop unless user wants more
    Regression on any configRevert immediately, try next optimization
    No improvement after 2 consecutive iterationsStop
    Root cause is scheduling or unknownEscalate to user
  8. If keeping, advance the baseline numbers and continue loop

  9. If reverting, git reset back to where you started and try the next optimization in priority order UNTIL: all attempts are finished, or more than 25 iterations have occurred, or the user interrupts

Be autonomous: Ask user clarifications at setup phase. Once stepped into the experiment loop, do not pause to ask user feedback: Use your best judgement for decision making, consult the optimization playbook and perf knobs catalog promptly, and think harder if stuck.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/tilegym-improve-cutile-kernel-perf of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/cutile-api-reference.md
  • references/cutile-patterns-reference.md
  • references/ir-dump-guide.md
  • references/optimization-playbook.md
  • references/perf-knobs-catalog.md
  • references/performance-model.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Tilegym Improve Cutile Kernel Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tilegym Improve Cutile Kernel Perf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tilegym Improve Cutile Kernel Perf this skillNVIDIA/skills3.5k—~2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Cmux Debugging Guidemanaflow-ai/cmux28k1 repos~1.1kAutomated safety check: PassCustom licence
Electron Heap Snapshot Analysiskeybase/client9.3k—~875Automated safety check: PassBSD-3-Clause

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.

    9.3k GitHub stars~875 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about Tilegym Improve Cutile Kernel Perf

What does Tilegym Improve Cutile Kernel Perf do?

Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning. Tilegym Improve Cutile Kernel Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning.

When should I use Tilegym Improve Cutile Kernel Perf?

Tilegym Improve Cutile Kernel Perf fits situations like: asked to optimize cutile kernel; improve kernel perf; tune cutile performance; make kernel faster.

How do I install Tilegym Improve Cutile Kernel Perf in Claude Code?

Run `npx skills add NVIDIA/skills --skill tilegym-improve-cutile-kernel-perf -a claude-code`. Or copy the skill folder (skills/tilegym-improve-cutile-kernel-perf in NVIDIA/skills) into .claude/skills/tilegym-improve-cutile-kernel-perf in your project. Claude Code loads it when a task matches its description.

How do I install Tilegym Improve Cutile Kernel Perf in Codex?

Run `npx skills add NVIDIA/skills --skill tilegym-improve-cutile-kernel-perf -a codex`. Or copy the skill folder (skills/tilegym-improve-cutile-kernel-perf in NVIDIA/skills) into .agents/skills/tilegym-improve-cutile-kernel-perf in your project. Codex loads it when a task matches its description.

Can I use Tilegym Improve Cutile Kernel Perf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tilegym-improve-cutile-kernel-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilegym-improve-cutile-kernel-perf, .gemini/skills/tilegym-improve-cutile-kernel-perf, .github/skills/tilegym-improve-cutile-kernel-perf and .opencode/skills/tilegym-improve-cutile-kernel-perf in your project.

What does Tilegym Improve Cutile Kernel Perf need to run?

Going by SKILL.md and its folder, Tilegym Improve Cutile Kernel Perf needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Tilegym Improve Cutile Kernel Perf access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tilegym Improve Cutile Kernel Perf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tilegym Improve Cutile Kernel Perf use?

Tilegym Improve Cutile Kernel Perf is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tilegym Improve Cutile Kernel Perf use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.

What are the alternatives to Tilegym Improve Cutile Kernel Perf?

Skills that share tags, products or a category with Tilegym Improve Cutile Kernel Perf: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Py (crazyguitar/pysheeet, 8.2k stars) and Cmux Debugging Guide (manaflow-ai/cmux, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tilegym Improve Cutile Kernel Perf?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.