Agent skill

Triton Kernel

by vipshop in vipshop/cache-dit

Write optimized Triton GPU kernels for deep learning operations.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton Kernel

skills CLI
$ npx skills add vipshop/cache-dit --skill triton-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vipshop/cache-dit triton-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/triton-kernel .claude/skills/triton-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triton-kernel
GitHub stars
1.3k
Token cost
~1.1k tokens
SKILL.md length
372 words
Files
11
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Write optimized Triton GPU kernels for deep learning operations.

  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers Core Patterns (always apply), Quick Reference Examples, Performance Bottleneck… and Specialized Topics, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Deep learning

What it does

Triton Kernel is an agent skill from vipshop/cache-dit. Write optimized Triton GPU kernels for deep learning operations. Covers the full spectrum from basic vector ops to Flash Attention, persistent matmul, fused normalization, quantized GEMM, and memory-efficient patterns.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files (for example `triton-dynamic-launcher-tiling.md`, `triton-flash-attention-v2.md` and `triton-fused-epilogue-kernels.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing, Deep learning and Database schema design. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing
  • Tasks that involve Deep learning
  • Tasks that involve Database schema design

Example prompts

  • “/triton-kernel”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton Kernel loads about 1.1k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 372 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 372 words, ~1,131 tokens.

Download SKILL.mdSave it as .claude/skills/triton-kernel/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
triton-kernel
description
Write optimized Triton GPU kernels for deep learning operations. Covers the full spectrum from basic vector ops to Flash Attention, persistent matmul, fused normalization, quantized GEMM, and memory-efficient patterns.
user-invocable
true

Writing Optimized Triton GPU Kernels

Targets: Triton >= 2.1, any GPU with tl.dot support (SM70+/CDNA2+)

Core Patterns (always apply)

Kernel structure: Use @triton.jit decorator. Get block ID with tl.program_id(axis). Compute element offsets with tl.arange(0, BLOCK_SIZE). Build mask = offsets < n_elements for all loads/stores.

Block sizes: Strongly prefer powers of two (required for tl.arange; non-power-of-two may work but can reduce performance). Declare as tl.constexpr parameters. Use @triton.autotune to sweep BLOCK_SIZE_M/N/K configs per hardware.

Memory hierarchy: Keep intermediates in SRAM via block-level reductions (tl.sum, tl.max) before writing to global memory. Fuse multiple pointwise ops into one kernel to avoid DRAM round-trips.

Matmul: Use tl.dot(a, b) for tensor core operations. Always accumulate in tl.float32 when inputs are FP16. For L2 cache locality, use grouped tile ordering via group_id = pid // GROUP_SIZE.

Grid launching: Size grid dynamically: grid = lambda meta: (triton.cdiv(n, meta['BLOCK_SIZE']),).

Masking: ALWAYS mask boundary loads/stores: tl.load(ptr + offs, mask=offs < dim, other=0.0). Missing masks corrupt memory silently.

Benchmarking: Use triton.testing.Benchmark with x_names, x_vals, line_arg, line_vals to compare against PyTorch baselines.

Quick Reference Examples

Fused row-wise softmax — verified, based on official Triton tutorial:

python
@triton.jit
def fused_softmax(x_ptr, out_ptr, cols, BLOCK: tl.constexpr):
    row = tl.program_id(0)
    offs = tl.arange(0, BLOCK)
    mask = offs < cols
    x = tl.load(x_ptr + row * cols + offs, mask=mask, other=-1e9)
    x_max = tl.max(x, axis=0)
    ex = tl.exp(x - x_max)
    out = ex / tl.sum(ex, axis=0)
    tl.store(out_ptr + row * cols + offs, out, mask=mask)

Seed-based dropout — verified, based on official Triton tutorial:

python
@triton.jit
def dropout(x_ptr, out_ptr, seed, p, n, BLOCK: tl.constexpr):
    offs = tl.program_id(0) * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n
    x = tl.load(x_ptr + offs, mask=mask)
    r = tl.rand(seed, offs)  # Philox PRNG, deterministic
    keep = r > p
    tl.store(out_ptr + offs, x * keep / (1.0 - p), mask=mask)

Performance Bottleneck Quick-Reference

When optimizing an existing kernel, classify the bottleneck first (profile with ncu):

BottleneckDiagnosisFix
Memory-boundDRAM throughput > 60% of peak, compute < 30%PID swizzle, TMA, fuse ops to reduce loads
Compute-boundTensor core utilization > 60%, DRAM < 40%Persistent kernels, increase num_stages, warp specialization
UnderutilizedBoth < 60%, high stall metricsReduce register pressure, increase num_warps, autotune

See triton-gpu-kernel-optimization.md for specific NCU metric names and detailed strategies.

Show full SKILL.md (120 more words)Show less

Specialized Topics

Read these files for detailed guidance when the task involves these areas:

TaskFile to read
Flash Attention / fused self-attentiontriton-flash-attention-v2.md
Persistent kernels, warp specialization, TMAtriton-persistent-warp-matmul.md
LayerNorm, RMSNorm, GroupNorm (fwd + bwd)triton-fused-normalizations.md
FP4/FP8 quantized matmul, block scalingtriton-quantized-block-scaled-gemm.md
Kernel fusion, Philox dropout, recomputationtriton-memory-efficient-patterns.md
General tiled GEMM, autotune, benchmarkingtriton-gpu-kernel-optimization.md
Fusing normalization/gating/residual into attention or matmul epiloguetriton-fused-epilogue-kernels.md
Sequential stateful processing (LRU routing, mutable register state)triton-sequential-stateful-blocks.md
Launcher tile selection, num_stages/num_warps heuristicstriton-dynamic-launcher-tiling.md

When to read specialized files: Only read the relevant file when the user's task specifically involves that topic. The core patterns above are sufficient for basic kernels (vector ops, elementwise fusion, simple reductions).

Other references

  • triton-opt.md: For general optimization techniques while writing triton kernels.

© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files in .github/skills/triton-kernel of vipshop/cache-dit.

  • SKILL.md
  • triton-dynamic-launcher-tiling.md
  • triton-flash-attention-v2.md
  • triton-fused-epilogue-kernels.md
  • triton-fused-normalizations.md
  • triton-gpu-kernel-optimization.md
  • triton-memory-efficient-patterns.md
  • triton-opt.md
  • triton-persistent-warp-matmul.md
  • triton-quantized-block-scaled-gemm.md
  • triton-sequential-stateful-blocks.md

Open the folder on GitHubat commit a7898aa

Compare with similar skills

Triton Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton Kernel this skillvipshop/cache-dit1.3k—~1.1kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
DGX Spark Training Gotchaswshobson/agents40k—~2kAutomated safety check: PassMIT
Hugging Face AccelerateOrchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub stars~2k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Accelerate

    Orchestra-Research/AI-Research-SKILLs

    Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

    13k GitHub starsUsed in 5 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 2 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed

More from vipshop/cache-dit

  • High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…

    1.3k GitHub stars~11k tokensUpdated 10 days ago
    Auto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 10 days ago
    Auto-check passed
  • Cute Dsl Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

    1.3k GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check passed
  • Cutlass Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

    1.3k GitHub stars~2.2k tokensUpdated 10 days ago
    Auto-check passed
  • Operator Migration

    vipshop/cache-dit

    A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

    1.3k GitHub stars~3.8k tokensUpdated 10 days ago
    Auto-check passed
  • Ptq Workflow Integration

    vipshop/cache-dit

    A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

    1.3k GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check passed

Questions about Triton Kernel

What does Triton Kernel do?

Write optimized Triton GPU kernels for deep learning operations. Triton Kernel is an agent skill from vipshop/cache-dit. Write optimized Triton GPU kernels for deep learning operations.

When should I use Triton Kernel?

Triton Kernel fits situations like: tasks that involve GPU and accelerator computing; tasks that involve Deep learning; tasks that involve Database schema design.

How do I install Triton Kernel in Claude Code?

Run `npx skills add vipshop/cache-dit --skill triton-kernel -a claude-code`. Or copy the skill folder (.github/skills/triton-kernel in vipshop/cache-dit) into .claude/skills/triton-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Triton Kernel in Codex?

Run `npx skills add vipshop/cache-dit --skill triton-kernel -a codex`. Or copy the skill folder (.github/skills/triton-kernel in vipshop/cache-dit) into .agents/skills/triton-kernel in your project. Codex loads it when a task matches its description.

Can I use Triton Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill triton-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-kernel, .gemini/skills/triton-kernel, .github/skills/triton-kernel and .opencode/skills/triton-kernel in your project.

What does Triton Kernel need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton Kernel is instructions for the agent only. Our summary lists: Python 3.

Does Triton Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton Kernel use?

Triton Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton Kernel use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton Kernel?

Skills that share tags, products or a category with Triton Kernel: MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and DGX Spark Training Gotchas (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton Kernel?

vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.