Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Add a new cuTile GPU kernel operator to TileGym. An agent skill from NVIDIA/skills.
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .claude/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .claude/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .agents/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .agents/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .cursor/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .cursor/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tilegym-adding-cutile-kernel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .gemini/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .gemini/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .github/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .github/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tilegym-adding-cutile-kernel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tilegym-adding-cutile-kernel .opencode/skills/tilegym-adding-cutile-kernel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tilegym-adding-cutile-kernel" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel into .opencode/skills/tilegym-adding-cutile-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilegym-adding-cutile-kernel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tilegym-adding-cutile-kernelAdd a new cuTile GPU kernel operator to TileGym. An agent skill from NVIDIA/skills.
Tilegym Adding Cutile Kernel is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, init.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to register a new cuTile op.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).
It sits in AI & LLM Engineering. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pytestpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tilegym Adding Cutile Kernel loads about 2.1k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 329 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 329 words, ~2,055 tokens.
.claude/skills/tilegym-adding-cutile-kernel/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.End-to-end workflow for adding a new operator (e.g., my_op) with cuTile backend.
MUST follow these rules strictly:
completed after finishing, in_progress when startingcompleted with a note, do NOT silently skipMUST copy this checklist to TodoWrite at the start:
- [ ] Step 1: Register dispatch interface in ops.py
- [ ] Step 2: Implement cuTile backend
- [ ] Step 3: Register in __init__.py (cutile)
- [ ] Step 4: Add tests
- [ ] Step 5: Add benchmark to tests/benchmark
- [ ] Step 6: Verify (run pytest + lint)File: src/tilegym/ops/ops.py
Add a @dispatch function — this is the single entry point for all backends.
@dispatch(
"my_op",
)
def my_op(
input: torch.Tensor,
out: Optional[torch.Tensor] = None,
**kwargs: Any,
):
"""
Description of my_op.
Args:
input: Input tensor
out: Optional preallocated output tensor
**kwargs: Additional arguments for backend-specific configurations
Returns:
torch.Tensor
"""
raise NotImplementedError(f"my_op is not implemented for {get_current_backend()}")Key rules:
NotImplementedError**kwargs for backend-specific parametersReference: See existing ops in src/tilegym/ops/ops.py (e.g., silu_and_mul, softmax)
File: src/tilegym/ops/cutile/my_op.py
The file structure follows this template:
import torch
import cuda.tile as ct
from tilegym.backend import register_impl
@ct.kernel
def my_op_kernel_ct(x, output, n_elements: ct.Constant[int], BLOCK_SIZE: ct.Constant[int]):
bid = ct.bid(0)
indices = bid * BLOCK_SIZE + ct.arange(0, BLOCK_SIZE)
x_val = ct.gather(x, indices)
# ... compute ...
ct.scatter(output, indices, result)
@register_impl("my_op", backend="cutile")
def my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs) -> torch.Tensor:
n = input.numel()
if out is None:
out = torch.empty_like(input)
grid = ((n + 1023) // 1024,)
ct.launch(stream, grid, kernel, (some args, ...))
return outReference: src/tilegym/ops/cutile/silu_and_mul.py
__init__.py (CRITICAL)Missing this step means the cuTile backend implementation never gets loaded.
File: src/tilegym/ops/cutile/__init__.py
Add inside if is_backend_available("cutile"): block (alphabetically):
from . import my_opAnd in the function import section:
from .my_op import my_opAnd add "my_op" to __all__.
File: tests/ops/test_my_op.py
CRITICAL: Always import from tilegym.ops, NEVER from tilegym.ops.cutile.my_op.
import pytest
import torch
from tilegym.backend import is_backend_available, set_backend
from .. import common
_backends = ["cutile"]
class Test_MY_OP(common.PyTestCase):
@staticmethod
def reference(input):
"""Reference implementation using PyTorch."""
return torch.some_reference(input)
@pytest.mark.parametrize("shape, dtype", [
((1024,), torch.float16),
((1024, 512), torch.float32),
((64, 64, 64), torch.bfloat16),
])
@pytest.mark.parametrize("backend", _backends)
def test_op(self, shape, dtype, backend, arch):
if backend == "cutile" and not is_backend_available("cutile"):
pytest.skip("Cutile backend not available")
try:
set_backend(backend)
except Exception as e:
pytest.skip(f"Backend is not supported: {e}")
self.setUp()
from tilegym.ops import my_op
A = torch.randn(*shape, dtype=dtype, device="cuda")
self.assertCorrectness(
my_op, self.reference, {"input": A},
atol=1e-3, rtol=1e-3,
)Key patterns:
_backends = ["cutile"]test_op: use set_backend(backend) with try-except, call self.setUp()Reference: tests/ops/test_silu_and_mul.py
Below is the common errors.
1. Missing _backends list (inside class)
2. test_op / test_op_xxx — missing @pytest.mark.parametrize("backend", _backends), backend parameter, and tilegym.is_backend_available / tilegym.set_backend patternFile: tests/benchmark/bench_my_op.py
Key rules from benchmark_rules.md:
tilegym.ops.my_op(a, b, ..., backend=backend) — do not use set_backend.ALL_BACKENDS (include at least cutile and torch), filter with get_supported_backends().reference_my_op(...) and register it: register_impl("my_op", "torch")(reference_my_op).create_benchmark_config() to build triton.testing.Benchmark configs (e.g. by shape/dtype).@triton.testing.perf_report([...]) on bench_my_op(...); inside the bench function: correctness check with torch.testing.assert_close(fn(), ref(), ...), then ms = triton.testing.do_bench(fn) (or do_bench_cudagraph), compute GB/s or TFLOPS, and return the metric.if __name__ == "__main__": bench_my_op.run(print_data=True).Template structure:
import torch
import triton
import triton.testing
import tilegym
from tilegym.backend import is_backend_available, register_impl
ALL_BACKENDS = [
("cutile", "cuTile", ("orange", "-")) if is_backend_available("cutile") else None,
("torch", "PyTorch", ("green", "-")),
]
def get_supported_backends():
return [p for p in ALL_BACKENDS if p is not None]
def reference_my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs):
"""Reference implementation using PyTorch."""
...
register_impl("my_op", "torch")(reference_my_op)
def create_benchmark_config(datatype, ...):
available_backends = get_supported_backends()
if not available_backends:
return None
backends, names, styles = zip(*available_backends)
return triton.testing.Benchmark(
x_names=["M"], # or other dimension names
x_vals=[...],
line_arg="backend",
line_vals=list(backends),
line_names=list(names),
styles=list(styles),
ylabel="GB/s", # or TFLOPS
plot_name="my-op-...",
args={"datatype": datatype, ...},
)
@triton.testing.perf_report([
create_benchmark_config(datatype, ...)
for datatype in [torch.float16, torch.float32]
for ... in [...]
])
def bench_my_op(M, backend, datatype, ..., device="cuda"):
x = torch.randn(..., dtype=datatype, device=device)
fn = lambda: tilegym.ops.my_op(x, backend=backend)
ref = lambda: reference_my_op(x)
torch.testing.assert_close(fn(), ref(), rtol=1e-2, atol=1e-2)
ms = triton.testing.do_bench(fn) # or do_bench_cudagraph(fn)
# Compute metric (e.g. GB/s or TFLOPS) from ms and problem size
return metric
if __name__ == "__main__":
bench_my_op.run(print_data=True)Benchmark Plot Names: Must include -TFLOPS or -GBps suffix
plot_name=f"persistent-layer-norm-M{num_rows}-{dtype_name}-GBps"# Run tests
pytest tests/ops/test_my_op.py -v
# Run benchmark (optional)
python tests/benchmark/bench_my_op.py
# Lint
pre-commit run -a© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/tilegym-adding-cutile-kernel of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Tilegym Adding Cutile Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tilegym Adding Cutile Kernel this skillNVIDIA/skills | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 13 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Add a new cuTile GPU kernel operator to TileGym. An agent skill from NVIDIA/skills. Tilegym Adding Cutile Kernel is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Add a new cuTile GPU kernel operator to TileGym.
Tilegym Adding Cutile Kernel fits situations like: implementing a new cuTile operator/kernel in TileGym; asking how to register a new cuTile op.
Run `npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a claude-code`. Or copy the skill folder (skills/tilegym-adding-cutile-kernel in NVIDIA/skills) into .claude/skills/tilegym-adding-cutile-kernel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a codex`. Or copy the skill folder (skills/tilegym-adding-cutile-kernel in NVIDIA/skills) into .agents/skills/tilegym-adding-cutile-kernel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tilegym-adding-cutile-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilegym-adding-cutile-kernel, .gemini/skills/tilegym-adding-cutile-kernel, .github/skills/tilegym-adding-cutile-kernel and .opencode/skills/tilegym-adding-cutile-kernel in your project.
Going by SKILL.md and its folder, Tilegym Adding Cutile Kernel needs the command-line tools its instructions call (pytest and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tilegym Adding Cutile Kernel is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tilegym Adding Cutile Kernel: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.