Hyperpod Version Checker
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/triton-skill .claude/skills/triton-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .claude/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/triton-skill .agents/skills/triton-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .agents/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/triton-skill .cursor/skills/triton-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .cursor/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/slowlyC/agent-gpu-skills.git --path skills/triton-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/triton-skill .gemini/skills/triton-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .gemini/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install slowlyC/agent-gpu-skills triton-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/triton-skill .github/skills/triton-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .github/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install slowlyC/agent-gpu-skills triton-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/triton-skill .opencode/skills/triton-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "triton-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/triton-skill into .opencode/skills/triton-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
triton-skillWrite, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
Triton Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl., Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, tritonkernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `quick-reference.md`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA, NVIDIA AI Platform and Python. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit ae02d07. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
rgbashpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triton Skill loads about 1.3k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 413 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from slowlyC/agent-gpu-skills at commit ae02d07, republished under its MIT licence (© slowlyC). 413 words, ~1,331 tokens.
.claude/skills/triton-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use the local Triton checkout as the primary source for APIs and implementation patterns. Prefer current tutorials and source over remembered signatures because Triton and Gluon evolve quickly.
Resolve the directory containing this SKILL.md, then use its repos/triton/ child. The installer links that path to agent-gpu-skills/third_party/triton/ or to the checkout supplied through TRITON_REPO.
In commands below, replace TRITON_REPO with the resolved absolute path:
TRITON_REPO=/absolute/path/to/triton-skill/repos/tritonIf the checkout is missing, run this from the agent-gpu-skills repository and reinstall the Skill:
bash update-repos.sh triton
bash install.sh --skill triton-skill| Task | Start here |
|---|---|
| Triton language syntax and introductory patterns | python/tutorials/ |
| Gluon layout and architecture-level patterns | python/tutorials/gluon/ |
| Complete example kernels | python/examples/ |
| Production matmul, reduction, top-k and SwiGLU | python/triton_kernels/triton_kernels/ |
tl.* definitions and semantics | python/triton/language/ |
| JIT, autotuning and runtime behavior | python/triton/runtime/ |
| Python compiler entry points | python/triton/compiler/ |
| Triton and GPU dialect definitions | include/triton/Dialect/ |
| Compiler analyses, transforms and lowering | lib/ |
Read quick-reference.md when choosing a tutorial, a complete example, or a production-kernel implementation.
Discover current examples before relying on a remembered filename:
find "$TRITON_REPO/python/tutorials" -maxdepth 2 -type f | sort
find "$TRITON_REPO/python/examples" -type f | sortQuery Triton language usage and definitions:
rg -n 'tl\.dot|tl\.dot_scaled' "$TRITON_REPO/python/tutorials"
rg -n '@triton\.autotune' "$TRITON_REPO/python/tutorials"
rg -n '^def (load|store|dot|dot_scaled)' \
"$TRITON_REPO/python/triton/language"Query Gluon architecture patterns:
rg -n '@gluon\.jit' "$TRITON_REPO/python/tutorials/gluon"
rg -n 'wgmma|tcgen05|mbarrier|tma' \
"$TRITON_REPO/python/tutorials/gluon" \
"$TRITON_REPO/python/examples"Trace production kernels:
rg -n 'persistent|TensorDescriptor' \
"$TRITON_REPO/python/triton_kernels/triton_kernels/matmul_details"
rg -n 'mxfp|flexpoint' \
"$TRITON_REPO/python/triton_kernels/triton_kernels/numerics_details"Trace compiler definitions and lowering:
rg -n 'def.*Op' "$TRITON_REPO/include/triton/Dialect/Triton/IR"
rg -n 'Encoding' "$TRITON_REPO/include/triton/Dialect/TritonGPU/IR"
rg -n 'wgmma|tma|tcgen05' \
"$TRITON_REPO/include/triton/Dialect/TritonNvidiaGPU"
rg -n 'Pattern|Rewrite' "$TRITON_REPO/lib/Conversion/TritonGPUToLLVM"Keep these layers separate during diagnosis:
Python kernel and launch metadata
→ Triton/Gluon IR and compiler transforms
→ generated GPU code on the selected targetA source-level pattern does not prove that the compiled kernel uses the intended instruction or memory path. Inspect compiler output or profile data when that distinction matters. Add cuda-skill for PTX semantics, compute capability, Nsight, or Compute Sanitizer details.
For correctness work:
For performance work:
From the agent-gpu-skills repository:
bash update-repos.sh triton
python3 scripts/validate_repo.py --require-sourcesThe checkout follows Triton main, while third_party/UPSTREAMS.toml records the commit last accepted by this Skill. Review source-map drift before updating that record.
© slowlyC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/triton-skill of slowlyC/agent-gpu-skills.
Open the folder on GitHubat commit ae02d07
Triton Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triton Skill this skillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Hyperpod Version Checkerawslabs/agent-plugins | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Jetson Video SetupNVIDIA/skills | 3.6k | 1 repos | ~2.4k | Automated safety check: Notes | Apache-2.0 | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 |
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
NVIDIA/skills
A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
slowlyC/agent-gpu-skills
Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.
slowlyC/agent-gpu-skills
Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.
Works with
Categories
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Triton Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
Triton Skill fits situations like: the task explicitly involves triton.jit; triton.language; tensorDescriptor; triton autotune.
Run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a claude-code`. Or copy the skill folder (skills/triton-skill in slowlyC/agent-gpu-skills) into .claude/skills/triton-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a codex`. Or copy the skill folder (skills/triton-skill in slowlyC/agent-gpu-skills) into .agents/skills/triton-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slowlyC/agent-gpu-skills --skill triton-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-skill, .gemini/skills/triton-skill, .github/skills/triton-skill and .opencode/skills/triton-skill in your project.
Going by SKILL.md and its folder, Triton Skill needs the command-line tools its instructions call (rg, bash and python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Triton Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Triton Skill: Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Jetson Video Setup (NVIDIA/skills, 3.6k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
slowlyC (a GitHub user) maintains it in slowlyC/agent-gpu-skills, which has 169 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 8, 2026.
Source: slowlyC/agent-gpu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.