Agent skill

Triton Skill

by slowlyC in slowlyC/agent-gpu-skills

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

MITAuto-check passedAI & LLM Engineering

Install Triton Skill

skills CLI
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/triton-skill .claude/skills/triton-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triton-skill
GitHub stars
169
Token cost
~1.3k tokens
SKILL.md length
413 words
Files
2
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

  • Works in 5 steps: Identify whether the task is Triton,… → Find the closest current tutorial or… → Verify every API used against its… → …
  • The task explicitly involves triton.jit
  • SKILL.md covers Locate the checkout, Choose the source surface, Query workflow and Implementation discipline, plus 1 more section
  • Calls rg, bash and python3

What it does

Triton Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl., Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, tritonkernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `quick-reference.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA, NVIDIA AI Platform and Python. The licence is MIT.

When your agent uses it

  • The task explicitly involves triton.jit
  • Triton.language
  • TensorDescriptor
  • Triton autotune

Example prompts

  • “/triton-skill”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify whether the task is Triton, Gluon, or compiler internals.
  2. Find the closest current tutorial or implementation for the operation and architecture.
  3. Verify every API used against its definition or another current call site.
  4. Preserve the target workload's shape, dtype, stride, layout, masking, and numerical contract.
  5. Establish correctness before changing launch parameters or optimization strategy.

What it can do on your machine

Read from SKILL.md and the folder at commit ae02d07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg
    • bash
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton Skill loads about 1.3k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 413 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from slowlyC/agent-gpu-skills at commit ae02d07, republished under its MIT licence (© slowlyC). 413 words, ~1,331 tokens.

Download SKILL.mdSave it as .claude/skills/triton-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
triton-skill
description
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl.*, Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, triton_kernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.

Triton and Gluon development

Use the local Triton checkout as the primary source for APIs and implementation patterns. Prefer current tutorials and source over remembered signatures because Triton and Gluon evolve quickly.

Locate the checkout

Resolve the directory containing this SKILL.md, then use its repos/triton/ child. The installer links that path to agent-gpu-skills/third_party/triton/ or to the checkout supplied through TRITON_REPO.

In commands below, replace TRITON_REPO with the resolved absolute path:

bash
TRITON_REPO=/absolute/path/to/triton-skill/repos/triton

If the checkout is missing, run this from the agent-gpu-skills repository and reinstall the Skill:

bash
bash update-repos.sh triton
bash install.sh --skill triton-skill

Choose the source surface

TaskStart here
Triton language syntax and introductory patternspython/tutorials/
Gluon layout and architecture-level patternspython/tutorials/gluon/
Complete example kernelspython/examples/
Production matmul, reduction, top-k and SwiGLUpython/triton_kernels/triton_kernels/
tl.* definitions and semanticspython/triton/language/
JIT, autotuning and runtime behaviorpython/triton/runtime/
Python compiler entry pointspython/triton/compiler/
Triton and GPU dialect definitionsinclude/triton/Dialect/
Compiler analyses, transforms and loweringlib/

Read quick-reference.md when choosing a tutorial, a complete example, or a production-kernel implementation.

Query workflow

  1. Identify whether the task is Triton, Gluon, or compiler internals.
  2. Find the closest current tutorial or implementation for the operation and architecture.
  3. Verify every API used against its definition or another current call site.
  4. Preserve the target workload's shape, dtype, stride, layout, masking, and numerical contract.
  5. Establish correctness before changing launch parameters or optimization strategy.

Discover current examples before relying on a remembered filename:

bash
find "$TRITON_REPO/python/tutorials" -maxdepth 2 -type f | sort
find "$TRITON_REPO/python/examples" -type f | sort

Query Triton language usage and definitions:

bash
rg -n 'tl\.dot|tl\.dot_scaled' "$TRITON_REPO/python/tutorials"
rg -n '@triton\.autotune' "$TRITON_REPO/python/tutorials"
rg -n '^def (load|store|dot|dot_scaled)' \
  "$TRITON_REPO/python/triton/language"

Query Gluon architecture patterns:

bash
rg -n '@gluon\.jit' "$TRITON_REPO/python/tutorials/gluon"
rg -n 'wgmma|tcgen05|mbarrier|tma' \
  "$TRITON_REPO/python/tutorials/gluon" \
  "$TRITON_REPO/python/examples"

Trace production kernels:

bash
rg -n 'persistent|TensorDescriptor' \
  "$TRITON_REPO/python/triton_kernels/triton_kernels/matmul_details"

rg -n 'mxfp|flexpoint' \
  "$TRITON_REPO/python/triton_kernels/triton_kernels/numerics_details"

Trace compiler definitions and lowering:

bash
rg -n 'def.*Op' "$TRITON_REPO/include/triton/Dialect/Triton/IR"
rg -n 'Encoding' "$TRITON_REPO/include/triton/Dialect/TritonGPU/IR"
rg -n 'wgmma|tma|tcgen05' \
  "$TRITON_REPO/include/triton/Dialect/TritonNvidiaGPU"
rg -n 'Pattern|Rewrite' "$TRITON_REPO/lib/Conversion/TritonGPUToLLVM"
Show full SKILL.md (167 more words)Show less

Implementation discipline

Keep these layers separate during diagnosis:

text
Python kernel and launch metadata
  → Triton/Gluon IR and compiler transforms
  → generated GPU code on the selected target

A source-level pattern does not prove that the compiled kernel uses the intended instruction or memory path. Inspect compiler output or profile data when that distinction matters. Add cuda-skill for PTX semantics, compute capability, Nsight, or Compute Sanitizer details.

For correctness work:

  • compare against an independent reference over representative and boundary shapes;
  • test masked tails, non-power-of-two dimensions, strides and dtype conversions;
  • separate compilation failures from runtime correctness and numerical tolerance;
  • reproduce the original dispatch and launch metadata before reducing the case.

For performance work:

  • freeze the benchmark shape set and measurement method;
  • confirm autotune keys cover every dimension that changes the best configuration;
  • change one tile, warp, stage, persistence, or specialization choice at a time;
  • verify the generated path before attributing a result to TMA, WGMMA, tcgen05, or warp specialization.

Updating the source

From the agent-gpu-skills repository:

bash
bash update-repos.sh triton
python3 scripts/validate_repo.py --require-sources

The checkout follows Triton main, while third_party/UPSTREAMS.toml records the commit last accepted by this Skill. Review source-map drift before updating that record.

© slowlyC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/triton-skill of slowlyC/agent-gpu-skills.

  • SKILL.md
  • quick-reference.md

Open the folder on GitHubat commit ae02d07

Compare with similar skills

Triton Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton Skill this skillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins916—~910Automated safety check: PassApache-2.0
Jetson Video SetupNVIDIA/skills3.6k1 repos~2.4kAutomated safety check: NotesApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0

Similar skills

  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Jetson Video Setup

    NVIDIA/skills

    Official

    A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

    3.6k GitHub starsUsed in 1 repo~2.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from slowlyC/agent-gpu-skills

  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Tilelang Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Cuda Skill

    slowlyC/agent-gpu-skills

    Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Triton Skill

What does Triton Skill do?

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Triton Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

When should I use Triton Skill?

Triton Skill fits situations like: the task explicitly involves triton.jit; triton.language; tensorDescriptor; triton autotune.

How do I install Triton Skill in Claude Code?

Run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a claude-code`. Or copy the skill folder (skills/triton-skill in slowlyC/agent-gpu-skills) into .claude/skills/triton-skill in your project. Claude Code loads it when a task matches its description.

How do I install Triton Skill in Codex?

Run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a codex`. Or copy the skill folder (skills/triton-skill in slowlyC/agent-gpu-skills) into .agents/skills/triton-skill in your project. Codex loads it when a task matches its description.

Can I use Triton Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-skill, .gemini/skills/triton-skill, .github/skills/triton-skill and .opencode/skills/triton-skill in your project.

What does Triton Skill need to run?

Going by SKILL.md and its folder, Triton Skill needs the command-line tools its instructions call (rg, bash and python3). Our summary lists: Python 3.

Does Triton Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton Skill use?

Triton Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton Skill use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton Skill?

Skills that share tags, products or a category with Triton Skill: Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Jetson Video Setup (NVIDIA/skills, 3.6k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton Skill?

slowlyC (a GitHub user) maintains it in slowlyC/agent-gpu-skills, which has 169 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 8, 2026.

Source: slowlyC/agent-gpu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.