Official agent skill

Tilegym Cutile Python

by NVIDIA in NVIDIA/skills

Expert cuTile programming assistant. An agent skill from NVIDIA/skills.

OfficialApache-2.0Auto-check passedAgent Workflows

Install Tilegym Cutile Python

skills CLI
$ npx skills add NVIDIA/skills --skill tilegym-cutile-python -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tilegym-cutile-python --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilegym-cutile-python .claude/skills/tilegym-cutile-python && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tilegym-cutile-python
GitHub stars
3.5k
Token cost
~5k tokens
SKILL.md length
2,196 words
Files
43
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Expert cuTile programming assistant. An agent skill from NVIDIA/skills.

  • Works in 7 steps: Search Examples and Consult References… → Understand the Problem → Design Kernel Architecture → …
  • Tasks that involve Multi-agent orchestration
  • SKILL.md covers Overview, When to Use This Skill, Reference Documentation and Examples, plus 7 more sections
  • Runs Python scripts from its folder; calls python

What it does

Tilegym Cutile Python is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Expert cuTile programming assistant. Write high-performance GPU kernels using cuTile's tile-based programming model with proper validation and optimization. Supports deep agent orchestration for complex multi-kernel tasks.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 48 other files (for example `BENCHMARK.md`, `evals/evals.json` and `examples/convolution/README.md`).

It sits in Agent Workflows, covering Multi-agent orchestration. It works with Python and NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Multi-agent orchestration

Example prompts

  • “/tilegym-cutile-python”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Search Examples and Consult References (MANDATORY)
  2. Understand the Problem
  3. Design Kernel Architecture
  4. Prepare Type System and Constants
  5. Implement the Kernel
  6. Prepare and Launch
  7. Validate and Test

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tilegym Cutile Python loads about 5k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 2,196 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 2,196 words, ~4,995 tokens.

Download SKILL.mdSave it as .claude/skills/tilegym-cutile-python/SKILL.md (or your agent's skills folder). This skill also uses 42 other files; get the full folder from GitHub.
name
tilegym-cutile-python
description
Expert cuTile programming assistant. Write high-performance GPU kernels using cuTile's tile-based programming model with proper validation and optimization. Supports deep agent orchestration for complex multi-kernel tasks.
version
1.3.0
license
CC-BY-4.0 AND Apache-2.0
metadata.author
TileGym Team <TileGym@nvidia.com>
metadata.tags
cutile, gpu-kernels, cuda

cuTile Python Programming Skill

You are an expert in cuTile programming, specializing in writing high-performance GPU kernels using cuTile's tile-based programming model. This skill provides comprehensive guidance for creating, debugging, and optimizing cuTile kernels.

Overview

cuTile is a parallel programming model for NVIDIA GPUs with a Python-based DSL that automatically leverages advanced hardware capabilities like tensor cores. This skill helps you write efficient, correct cuTile code.

When to Use This Skill

Invoke this skill when you need to:

  • Write cuTile GPU kernels from scratch
  • Convert tensor operations to cuTile implementations
  • Debug or fix cuTile kernel code
  • Optimize cuTile kernels for performance
  • Understand cuTile API and programming patterns
  • Validate cuTile implementations
  • Find and adapt examples from available reference sources

Optionally specify when invoking:

  • Target tensor shapes
  • Data types (default: float16)
  • Performance requirements
  • Any special constraints

Reference Documentation

cuTile Language Specification — https://docs.nvidia.com/cuda/cutile-python. Covers the execution model, data and memory models, debugging, compilation, and every public op (load/store, factories, reductions, scans, matmul, selection, math, bitwise, comparisons, atomics, metaprogramming, classes, enums, autotuning).

Implementation Guidelines (in the guidelines/ directory):

Examples

Before starting any cuTile programming task, always search for existing examples first. TileGym is the primary reference; the packaged examples/ directory complements it for ops TileGym does not yet cover (convolution, pooling, scan, GEMV, 4D matmul, split-k GEMM, group_norm).

The skill supports two installation contexts:

  • Inside a TileGym checkout (<repo>/skills/tilegym-cutile-python/, or <repo>/.agents/skills/tilegym-cutile-python/ / <repo>/.claude/skills/tilegym-cutile-python/ via the backward-compat symlinks) — TileGym ops are at <repo>/src/tilegym/ops/cutile/.
  • Installed elsewhere (e.g. ~/.agents/skills/tilegym-cutile-python/, ~/.claude/skills/tilegym-cutile-python/, or inside a different repo) — clone TileGym once to ${TILEGYM_SKILL_CACHE_DIR:-~/.cache/tilegym}/TileGym and use its src/tilegym/ops/cutile/.

See examples/tilegym_and_examples_guide.md for the full search order, directory layout, and cache-vs-repo decision procedure.

When to Clarify Before Implementation

For complex or ambiguous tasks, present approach options to the user before coding. This prevents wasted effort on the wrong implementation.

Clarify for These Task Types
Task TypeWhy ClarifyExample Questions
Optimization requests"Make this faster" has many pathsWhich bottleneck? Memory-bound vs compute-bound? Target speedup?
Architecture changesStructural decisions affect everythingData parallel vs model parallel? Persistent kernel vs standard?
Ambiguous operationsSame name, different implementationsFlash attention vs standard? Causal vs bidirectional? Grouped vs depthwise conv?
Performance vs correctness tradeoffsUser must chooseUse TF32 for speed? Approximate math functions? Reduced precision accumulation?
Missing constraintsCan't optimize without targetsTarget tensor shapes? Batch size range? Memory budget?
Act Directly for These Task Types
  • Clear, specific requests: "Write a ReLU kernel for shape (1024, 1024)"
  • Bug fixes with reproduction: "This kernel crashes on line 42"
  • API questions: "How do I use ct.gather?"
  • Example adaptations: "Adapt the TileGym softmax for my shapes"
How to Clarify

When clarification is needed:

  1. Briefly explain why multiple approaches exist
  2. Present 2-3 concrete options with tradeoffs
  3. Recommend one option if there's a clear best choice
  4. Ask the user to choose before proceeding

Example:

Your request "optimize this matmul" could go several directions:

1. **Persistent kernel** - Best for small matrices, faster, more complex code
2. **Tile size tuning** - Moderate gains, minimal code changes
3. **TMA prefetching** - Best for large matrices, requires Hopper+ GPU

I recommend option 2 for a first pass. Which approach would you like?

Complexity Assessment: Simple vs. Orchestrated Workflow

Before starting implementation, assess the complexity of the request to choose the right workflow.

Use the Simple Workflow (Steps 0-6 below) when:
  • Single kernel task (e.g., ReLU, softmax, one matmul)
  • Bug fix or optimization of an existing kernel
  • API question or example adaptation
  • Clear, single-operation request
Use the Deep Agent Orchestration Workflow when ANY of these apply:
  • 3+ distinct operations that need separate kernels (e.g., "implement a transformer block with attention, FFN, and layer norm")
  • Multiple user-defined functions in the input code (e.g., custom_activation(), custom_norm())
  • Inter-kernel data dependencies where output of one kernel feeds into another
  • PyTorch nn.Module with multiple layers in forward()
  • Explicit decomposition request (e.g., "break this into fused kernels")

When orchestration is needed, follow the Deep Agent Orchestration Workflow section. Otherwise, continue with the Instructions below.

Deep Agent Orchestration Workflow

For complex tasks requiring 3+ kernels, inter-kernel dependencies, or multi-layer nn.Module decomposition, use the orchestrated multi-agent pipeline. The main agent acts as an orchestrator (not a coder) — sub-agents handle reference reading and code generation.

Pipeline: Op Tracer (optional) → Analyzer → Kernel Agents (parallel) → Composer → Main Agent validates

For the complete step-by-step workflow (Steps O-0 through O-4), prompt templates, and error handling, see orchestration/workflow.md.

For the orchestration architecture, agent hierarchy, and kernel spec format, see orchestration/overview.md.


Instructions

Follow these steps when writing cuTile kernels (simple workflow for single-kernel tasks).

NOTE: Skip this entire section if using the Deep Agent Orchestration Workflow above. The orchestration workflow has its own steps (O-0 through O-4). Do NOT combine both workflows - that leads to the main agent reading all reference files AND spawning sub-agents, which wastes context.

Step 0: Search Examples and Consult References (MANDATORY)

Objective: Find existing examples and review relevant documentation

Example Search (Two-Step Strategy):

  1. Search TileGym (src/tilegym/ops/cutile/) first for similar cuTile kernel patterns.
  2. If TileGym has no match, search the packaged examples/ directory (part of this skill).
  3. Read relevant example files to understand implementation patterns.

Complex Algorithm Translation (flash attention, fused ops, etc.): When implementing complex algorithms, follow this systematic approach:

  1. Analyze the PyTorch implementation: Understand the mathematical operations, data flow, key computational patterns, memory access patterns, and any special optimizations or constraints.
  2. Study relevant cuTile examples: Review examples for similar operations — existing examples often provide the exact patterns you need. Copy and adapt working patterns rather than reinventing the wheel.
  3. Implement the cuTile version: Map PyTorch operations to cuTile primitives, apply kernel fusion where appropriate, ensure proper tile indexing and memory management, and validate against the PyTorch reference.

Reference Documentation:

Step 1: Understand the Problem

Objective: Clearly define what the kernel needs to compute

  • Identify input/output tensors and their shapes/dtypes
  • Understand the mathematical operations required
  • Determine data dependencies and computation flow
  • Analyze memory access patterns for optimization opportunities

Working with user-provided reference implementations:

  1. Preserve Reference Code: Keep the original PyTorch reference implementation intact. Only remove code that is clearly redundant or unnecessary.
  2. Conservative Approach: Do not modify or rewrite the reference implementation unless explicitly required. The reference serves as the ground truth for correctness validation.
  3. Seek Clarification: If you are uncertain about the correctness or intent of any part of the reference code, ask the user for clarification before proceeding.
  4. Maintain Functionality: Any changes to the reference code must preserve the original functionality and behavior.
Step 2: Design Kernel Architecture

Objective: Plan the kernel structure

  • Determine optimal block/tile sizes for parallelization (consider multiples of 32)
  • Calculate grid dimensions based on tensor sizes using ct.cdiv(size, block)
  • Design block indexing strategy using ct.bid()
  • Handle edge cases where tensor size is not divisible by block size
Step 3: Prepare Type System and Constants

Objective: Ensure proper type annotations

  • Identify all constant values that need type annotations
  • Add proper type annotations using ct.Constant[type] for all constants
  • Choose appropriate cuTile dtypes (ct.float32, ct.float16, ct.int32, etc.)
  • Ensure block sizes and other parameters are properly typed
Step 4: Implement the Kernel

Objective: Write the cuTile kernel function

  • Create @ct.kernel decorated kernel function with proper signature
  • Add required parameters (input tensors, output tensor, typed constants)
  • Implement block indexing with appropriate ct.bid() calls
  • Use ct.load() for input tensor access with proper indexing and tile shapes
  • Perform operations on loaded tiles using cuTile tile operations
  • Use ct.store() for output tensor writing with correct indexing
Step 5: Prepare and Launch

Objective: Set up tensor inputs and launch kernel

  • Ensure all input tensors are on CUDA device using .cuda() or .to("cuda")
  • Verify tensor dtypes are compatible with cuTile
  • Handle tensor contiguity requirements using .contiguous() if needed
  • Launch kernel with proper grid dimensions
Step 6: Validate and Test

Objective: Ensure correctness

  • Verify kernel compiles without errors
  • Test with various tensor sizes (aligned and unaligned to tile size)
  • Validate results against reference implementation if available
  • Check boundary conditions and edge cases

Validation Loop (MANDATORY)

IMPORTANT: After generating cuTile code, you MUST execute it to verify correctness. Do not just write the file - run it and fix any issues.

Validation Workflow
┌─────────────────────────────────────────────────────────────┐
│  1. Generate Code                                           │
│     - Write cuTile kernel with inline validation to file    │
│                                                             │
│  2. Execute Code                                            │
│     - Run: python <filename>.py                             │
│                                                             │
│  3. Check Results                                           │
│     ├─ Compilation error? → Fix syntax/type issues → Retry  │
│     ├─ Runtime error? → Fix kernel logic → Retry            │
│     ├─ Validation FAIL? → Fix numerical issues → Retry      │
│     └─ Validation PASS? → Done ✓                            │
└─────────────────────────────────────────────────────────────┘
Show full SKILL.md (897 more words)Show less
Execution Steps
  1. Write the generated code to a .py file
  2. Run the file using Bash: python <filename>.py
  3. Analyze the output:
    • If compilation error: Read error message, fix the code (check type annotations, syntax, API usage)
    • If runtime error: Check tensor shapes, grid dimensions, memory access patterns
    • If validation FAIL: Check numerical differences, tolerances, algorithm correctness
    • If validation PASS: Report success to user
  4. Iterate until PASS: Fix issues and re-run until validation passes (max 3 attempts)
Validation Output Best Practices
  • Don't print large tensors - Only print tensor contents when validation fails
  • Print summary stats - Show PASS/FAIL, max difference, tensor shape
  • Example validation pattern:
    python
    is_close = torch.allclose(cutile_output, reference_output, atol=1e-3, rtol=1e-3)
    if is_close:
        print("✓ Validation PASSED")
    else:
        max_diff = (cutile_output - reference_output).abs().max().item()
        print(f"✗ Validation FAILED - max diff: {max_diff}")
        print(f"  Expected: {reference_output}")
        print(f"  Got:      {cutile_output}")
Common Issues and Fixes
Error TypeTypical CauseFix
TypeError: missing Constant annotationMissing ct.Constant[int]Add type annotation to all constants
ValueError: tile dimension not power of 2Non-power-of-2 tile sizeUse 2**((size-1).bit_length())
IndexError / CUDA errorWrong grid dimensions or indicesCheck ct.cdiv usage, tile vs element indices
Validation FAIL: max diff = XNumerical mismatchCheck algorithm, increase tolerance, or fix logic
Default Tolerance Values

See guidelines/03_concepts.md → "Default Rules When User Does Not Specify" for tolerance values, default dtypes, and default tensor shapes.

Testing Checklist
  • ✓ Verify cuTile output matches reference implementation within tolerance
  • ✓ Test with various tensor sizes (aligned and unaligned to tile size)
  • ✓ Test boundary conditions and edge cases
  • ✓ Ensure all tensors are on CUDA device before kernel launch
  • ✓ Verify dtype consistency across inputs and outputs

Critical Requirements

Four essential requirements for all cuTile kernels:

  1. Pure cuTile forward path: Every compute op in forward()/composed_function() must go through @ct.kernel + ct.launch. Do not call nn.Conv2d()(x), F.conv2d(x, w), F.linear(x, w), or any other nn.*/F.* compute op as a runtime operation in the forward path.
    • Permitted in forward(): torch.empty, torch.zeros, torch.ones (allocation); tensor.reshape, tensor.view, tensor.permute, tensor.contiguous (rearrangement); torch.cat, torch.stack (concatenation); torch.sqrt, .sum(), .mean() (simple scalar ops between kernel launches).
    • Permitted in __init__(): Using nn.Conv2d, nn.Linear, etc. solely for weight initialization and storage is fine — as long as forward() extracts the weights (e.g., self.conv.weight.data) and passes them to ct.launch instead of calling self.conv(x).
    • See Rule 15 and Rule 17 in guidelines/02_code_generation_rules.md for common violations and detailed examples.
  2. Tile indices, not element indices: ct.load(A, index=(bid_m, k), shape=(BLOCK_M, K)) ✅ not (bid_m * BLOCK_M, k) ❌
  3. All tile dimensions must be powers of 2: Use 2**((size-1).bit_length()) to round up
  4. All constants need type annotations: BLOCK: ct.Constant[int] is required for compilation

For detailed guidelines on memory operations, tile sizing, common pitfalls, and optimization strategies, see the guidelines/ directory (01–03).

Performance Optimization

Key principle: Think in blocks of data rather than individual elements. Choose tile sizes that match hardware characteristics and maximize data reuse within tiles.

File Management Guidelines

IMPORTANT: Follow these rules for file creation:

  1. Single file by default: Generate a single .py file containing the kernel, validation, and test code unless the user explicitly requests multiple files
  2. No documentation files: Do NOT create README.md, documentation files, or separate example files unless explicitly requested
  3. Inline everything: Include the kernel implementation, validation logic, and test code in one cohesive file
  4. Minimal file creation: Only create what is absolutely necessary - prefer editing existing files over creating new ones
  5. No source citations: Do NOT include comments or docstrings mentioning TileGym files, reference files, or sources. The code should stand on its own without attribution
  6. Output to current working directory: All output .py files must be written to the current working directory where the user started the coding assistant. Run pwd at the start of the task. All generated .py files go directly in that directory (e.g. ./composed_foo.py), never in a subdirectory of the skill.
  7. Skill directory is read-only: <skill_dir> is passed to sub-agents solely so they can read references, examples, and orchestration instructions. No agent — main or sub — may ever write, create, or save any file under <skill_dir>. Use it only with read tools (Read, Glob, Grep, Bash cat/grep). Never pass it to Write, Edit, or any file-creating command.

Example structure for a single file:

python
import cuda.tile as ct
import torch

# Kernel implementation
@ct.kernel
def my_kernel(...):
    ...

# Validation function (if needed)
def validate(...):
    ...

# Test/demo code at bottom
if __name__ == "__main__":
    # Test the kernel
    ...

Success Criteria

Your implementation is successful when:

  1. ✅ Pure cuTile forward path: No nn.*/F.* compute calls in forward()/composed_function() — all compute routed through ct.launch (weight-init-only usage in __init__ is fine)
  2. ✅ Existing examples were searched before implementation
  3. ✅ Packaged examples/ were searched if TileGym had no match
  4. ✅ Only ONE .py file created (no READMEs, no separate examples unless requested)
  5. ✅ No source citations in code (no mentions of TileGym files or reference files in comments/docstrings)
  6. ✅ Generated cuTile code compiles without errors
  7. ✅ Numerical results match reference implementation within tolerance
  8. ✅ All constants have proper type annotations
  9. ✅ All tile dimensions are powers of 2
  10. ✅ Grid dimensions correctly cover all tensor elements
  11. ✅ Code includes inline validation and test code in the same file

Additional criteria when using orchestration (complex tasks):

  1. ✅ Complexity was assessed and orchestration was chosen for the right reasons
  2. ✅ Analyzer produced clear kernel specs with PyTorch references
  3. ✅ Independent kernels were generated in parallel (not sequentially)
  4. ✅ Each individual kernel was validated before composition
  5. ✅ Composed solution passes end-to-end validation against original PyTorch reference

Remember: Start by searching existing examples, follow the workflow systematically, and validate thoroughly. The reference files contain detailed rules and examples to guide you through every aspect of cuTile kernel development.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 42 other files in skills/tilegym-cutile-python of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • examples/convolution/README.md
  • examples/convolution/conv2d_with_bias_dilation_groups.py
  • examples/convolution/conv3d_with_bias_dilation_groups.py
  • examples/convolution/conv_transpose_2d.py
  • examples/convolution/conv_transpose_3d.py
  • examples/matmul/README.md
  • examples/matmul/matmul_4d_tensors.py
  • examples/matmul/matrix_vector_multiplication.py
  • examples/matmul/split_k_gemm.py
  • examples/normalization/README.md
  • examples/normalization/group_norm.py
  • examples/pooling/README.md
  • … and 28 more

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Tilegym Cutile Python next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tilegym Cutile Python compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tilegym Cutile Python this skillNVIDIA/skills3.5k—~5kAutomated safety check: PassApache-2.0
Adk Agent Buildergoogle/adk-python22k—~879Automated safety check: PassApache-2.0
Analyze Codebasedivar-ir/ai-doc-gen765—~899Automated safety check: PassMIT
Aicr Cross ReviewNVIDIA/aicr432—~15kAutomated safety check: PassApache-2.0
Team Swarmcatlog22/maestro-flow563—~2kAutomated safety check: NotesNone
Cao Pluginawslabs/cli-agent-orchestrator1.4k—~3.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Adk Agent Builder

    google/adk-python

    Official

    Builds ADK (Agent Development Kit) Python agents: LLM agents with tools, graph workflows of function and agent nodes, conditional routing, fan-out and join, schema-validated delegation between…

    22k GitHub stars~879 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Analyze Codebase

    divar-ir/ai-doc-gen

    Run a multi-agent deep analysis of a codebase, producing AI-readable analysis documents in .ai/docs/ covering structure, dependencies, data flow, request flow, and APIs.

    765 GitHub stars~899 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Official

    Multi-agent PR review using Claude Code, Codex, and CodeRabbit.

    432 GitHub stars~15k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Team Swarm

    catlog22/maestro-flow

    Swarm intelligence team skill — ACO-driven multi-agent exploration with hybrid LLM coordinator + Python optimization controller.

    563 GitHub stars~2k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • Cao Plugin

    awslabs/cli-agent-orchestrator

    Official

    Create a new CAO (CLI Agent Orchestrator) plugin. An agent skill from awslabs/cli-agent-orchestrator.

    1.4k GitHub stars~3.1k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • agystack Runtime Setup

    jtaroreh/agystack

    Configures agystack's model tiers per role and its execution runtime, choosing between local subagents and Cloud Run jobs for large parallel swarms.

    103 GitHub stars~1.7k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Categories

Questions about Tilegym Cutile Python

What does Tilegym Cutile Python do?

Expert cuTile programming assistant. An agent skill from NVIDIA/skills. Tilegym Cutile Python is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Expert cuTile programming assistant.

When should I use Tilegym Cutile Python?

Tilegym Cutile Python fits situations like: tasks that involve Multi-agent orchestration.

How do I install Tilegym Cutile Python in Claude Code?

Run `npx skills add NVIDIA/skills --skill tilegym-cutile-python -a claude-code`. Or copy the skill folder (skills/tilegym-cutile-python in NVIDIA/skills) into .claude/skills/tilegym-cutile-python in your project. Claude Code loads it when a task matches its description.

How do I install Tilegym Cutile Python in Codex?

Run `npx skills add NVIDIA/skills --skill tilegym-cutile-python -a codex`. Or copy the skill folder (skills/tilegym-cutile-python in NVIDIA/skills) into .agents/skills/tilegym-cutile-python in your project. Codex loads it when a task matches its description.

Can I use Tilegym Cutile Python in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tilegym-cutile-python -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilegym-cutile-python, .gemini/skills/tilegym-cutile-python, .github/skills/tilegym-cutile-python and .opencode/skills/tilegym-cutile-python in your project.

What does Tilegym Cutile Python need to run?

Going by SKILL.md and its folder, Tilegym Cutile Python needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Tilegym Cutile Python access the network?

SKILL.md names 1 domain. As links in the text: docs.nvidia.com. This is read from the text; nothing was executed.

Is Tilegym Cutile Python safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tilegym Cutile Python use?

Tilegym Cutile Python is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tilegym Cutile Python use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tilegym Cutile Python?

Skills that share tags, products or a category with Tilegym Cutile Python: Adk Agent Builder (google/adk-python, 22k stars), Analyze Codebase (divar-ir/ai-doc-gen, 765 stars), Aicr Cross Review (NVIDIA/aicr, 432 stars) and Team Swarm (catlog22/maestro-flow, 563 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tilegym Cutile Python?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.