Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

MITAuto-check passedAI & LLM Engineering

Install Triton Lang

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/triton-lang .claude/skills/triton-lang && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triton-lang
GitHub stars
252
Token cost
~1.8k tokens
SKILL.md length
308 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

  • Works in 8 steps: Minimal Triton kernel → Load/store and masking → Atomic operations → …
  • Writing Triton kernels with @triton.jit
  • SKILL.md covers Purpose, When to Use, Workflow and Common Problems, plus 1 more section
  • Calls python

What it does

Triton Lang is an agent skill from mohitmishra786/low-level-dev-skills. Triton language skill for Python GPU kernel authoring. Use when writing Triton kernels with @triton.jit, tl.load/store, masking, atomics, benchmarking with triton.testing, or integrating kernels into PyTorch. Activates on queries about Triton, tl.constexpr, block pointers, Triton benchmarking, or PyTorch custom ops.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning and GPU and accelerator computing. It works with PyTorch, Python and CUDA. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Writing Triton kernels with @triton.jit
  • Benchmarking with triton.testing
  • Integrating kernels into PyTorch

Example prompts

  • “/triton-lang”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Minimal Triton kernel
  2. Load/store and masking
  3. Atomic operations
  4. Shared memory via constexpr
  5. Autotuning
  6. Benchmarking
  7. PyTorch integration
  8. Debugging

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton Lang loads about 1.8k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 308 words, ~1,768 tokens.

Download SKILL.mdSave it as .claude/skills/triton-lang/SKILL.md (or your agent's skills folder).
name
triton-lang
description
Triton language skill for Python GPU kernel authoring. Use when writing Triton kernels with @triton.jit, tl.load/store, masking, atomics, benchmarking with triton.testing, or integrating kernels into PyTorch. Activates on queries about Triton, tl.constexpr, block pointers, Triton benchmarking, or PyTorch custom ops.

Triton

Purpose

Guide agents through writing GPU kernels in OpenAI Triton: the @triton.jit decorator, block-oriented tl.load/tl.store with masking, atomic operations, shared memory via tl.constexpr, benchmarking with triton.testing.Benchmark, PyTorch integration, and debugging with barriers.

When to Use

  • Writing custom PyTorch ops faster than pure PyTorch but without raw CUDA
  • Prototyping fused kernels (e.g., softmax + scale + bias)
  • Comparing block sizes and warp counts with Triton's autotuner
  • Porting NumPy-style elementwise ops to GPU
  • Learning GPU programming with higher-level Python syntax
  • Benchmarking kernel variants systematically

Workflow

1. Minimal Triton kernel
python
import torch
import triton
import triton.language as tl

@triton.jit
def add_kernel(x_ptr, y_ptr, out_ptr, n, BLOCK: tl.constexpr):
    pid = tl.program_id(0)
    offsets = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offsets < n
    x = tl.load(x_ptr + offsets, mask=mask)
    y = tl.load(y_ptr + offsets, mask=mask)
    tl.store(out_ptr + offsets, x + y, mask=mask)

def add(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
    n = x.numel()
    out = torch.empty_like(x)
    grid = lambda meta: (triton.cdiv(n, meta["BLOCK"]),)
    add_kernel[grid](x, y, out, n, BLOCK=1024)
    return out

Key concepts:

  • tl.program_id(0) — block index (like blockIdx.x)
  • tl.arange(0, BLOCK) — vector of thread indices within block
  • mask — predication for tail elements (no separate bounds kernel)
  • BLOCK: tl.constexpr — compile-time constant, enables unrolling
2. Load/store and masking
python
@triton.jit
def masked_load_example(ptr, n, BLOCK: tl.constexpr):
    pid = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n
    # masked load returns 0 for masked-off lanes
    vals = tl.load(ptr + offs, mask=mask, other=0.0)
    return vals

Block pointers (Triton 2.x+) for structured 2D access:

python
@triton.jit
def matvec_kernel(a_ptr, x_ptr, y_ptr, M, N, BLOCK_M: tl.constexpr, BLOCK_N: tl.constexpr):
    pid_m = tl.program_id(0)
    offs_m = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
    acc = tl.zeros((BLOCK_M,), dtype=tl.float32)
    for start_n in range(0, N, BLOCK_N):
        offs_n = start_n + tl.arange(0, BLOCK_N)
        a = tl.load(a_ptr + offs_m[:, None] * N + offs_n[None, :])
        x = tl.load(x_ptr + offs_n)
        acc += tl.sum(a * x[None, :], axis=1)
    tl.store(y_ptr + offs_m, acc, mask=offs_m < M)
3. Atomic operations
python
@triton.jit
def atomic_histogram(data_ptr, hist_ptr, n, BLOCK: tl.constexpr):
    pid = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n
    data = tl.load(data_ptr + offs, mask=mask)
    bucket = (data % 256).to(tl.int32)
    tl.atomic_add(hist_ptr + bucket, 1, mask=mask)

Use atomics sparingly — they serialize memory updates. Prefer block-level reduction then single atomic per block.

4. Shared memory via constexpr
python
@triton.jit
def reduce_kernel(x_ptr, out_ptr, n, BLOCK: tl.constexpr):
    pid = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n
    x = tl.load(x_ptr + offs, mask=mask, other=0.0)
    # Block reduction
    x = tl.sum(x, axis=0)
    tl.atomic_add(out_ptr, x)

BLOCK as tl.constexpr lets the compiler allocate shared memory and unroll loops at compile time.

5. Autotuning
python
@triton.autotune(
    configs=[
        triton.Config({"BLOCK": 128}, num_warps=4),
        triton.Config({"BLOCK": 256}, num_warps=4),
        triton.Config({"BLOCK": 512}, num_warps=8),
    ],
    key=["n"],
)
@triton.jit
def tuned_kernel(x_ptr, y_ptr, out_ptr, n, BLOCK: tl.constexpr):
    # ... kernel body ...
    pass

Autotuner benchmarks configs on first run and caches the best for each key shape.

6. Benchmarking
python
from triton.testing import Benchmark

def benchmark_add():
    n = 1024 * 1024
    x = torch.randn(n, device="cuda")
    y = torch.randn(n, device="cuda")

    def triton_add():
        return add(x, y)

    def torch_add():
        return x + y

    bench = Benchmark(
        x_names=["n"],
        x_vals=[2**i for i in range(10, 24)],
        line_arg="provider",
        line_vals=["triton", "torch"],
        line_names=["Triton", "PyTorch"],
        plot_name="add-bench",
        args={},
    )
    bench.run(lambda n, provider: {
        "triton": lambda: add(x[:n], y[:n]),
        "torch": lambda: x[:n] + y[:n],
    }[provider](), quantiles=[0.5, 0.9])
bash
# Quick timing in REPL
import triton.testing as tt
ms = tt.do_bench(lambda: add(x, y))
print(f"{ms:.3f} ms")
7. PyTorch integration
python
import torch
from torch.library import custom_op

@custom_op("mylib::triton_add", mutates_args=())
def triton_add_op(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
    return add(x, y)

@triton_add_op.register_fake
def _(x, y):
    return torch.empty_like(x)

# Use in model
class MyModule(torch.nn.Module):
    def forward(self, x, y):
        return triton_add_op(x, y)

For torch.compile compatibility, register fake/meta kernels and avoid Python side effects in the JIT function.

8. Debugging
python
@triton.jit
def debug_kernel(x_ptr, n, BLOCK: tl.constexpr):
    pid = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n
    x = tl.load(x_ptr + offs, mask=mask)
    # Synchronize threads within block for inspection
    tl.debug_barrier()
    tl.store(x_ptr + offs, x * 2.0, mask=mask)
bash
# Dump generated PTX/LLVM IR
TRITON_PRINT_AUTOTUNING=1 python script.py
# Set cache dir to inspect compiled kernels
export TRITON_CACHE_DIR=/tmp/triton_cache

Compare against PyTorch reference on small inputs before scaling up.

Common Problems

SymptomCauseFix
OutOfResourcesBLOCK too largeReduce BLOCK or num_warps
Wrong results on tailMissing maskAdd mask=offs < n to load/store
Slower than PyTorchSuboptimal BLOCKUse @triton.autotune
CompilationErrorType mismatchEnsure consistent dtypes; use .to(tl.float32)
NaN in outputUninitialized masked lanesPass other=0.0 to masked loads
torch.compile failsNo fake kernelRegister register_fake meta function
  • skills/gpu/cuda — underlying CUDA concepts when Triton limits are hit
  • skills/gpu/cuda-profiling — Nsight profiling of Triton-compiled kernels
  • skills/gpu/gpu-memory-model — coalescing and occupancy theory
  • skills/gpu/hip-rocm — AMD GPU path (Triton supports ROCm)
  • skills/hpc/openmp — CPU parallelism alongside GPU kernels

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu/triton-lang of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Triton Lang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton Lang compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton Lang this skillmohitmishra786/low-level-dev-skills252—~1.8kAutomated safety check: PassMIT
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins916—~910Automated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    252 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    252 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Triton Lang

What does Triton Lang do?

Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. Triton Lang is an agent skill from mohitmishra786/low-level-dev-skills. Triton language skill for Python GPU kernel authoring.

When should I use Triton Lang?

Triton Lang fits situations like: writing Triton kernels with @triton.jit; benchmarking with triton.testing; integrating kernels into PyTorch.

How do I install Triton Lang in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a claude-code`. Or copy the skill folder (skills/gpu/triton-lang in mohitmishra786/low-level-dev-skills) into .claude/skills/triton-lang in your project. Claude Code loads it when a task matches its description.

How do I install Triton Lang in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a codex`. Or copy the skill folder (skills/gpu/triton-lang in mohitmishra786/low-level-dev-skills) into .agents/skills/triton-lang in your project. Codex loads it when a task matches its description.

Can I use Triton Lang in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-lang, .gemini/skills/triton-lang, .github/skills/triton-lang and .opencode/skills/triton-lang in your project.

What does Triton Lang need to run?

Going by SKILL.md and its folder, Triton Lang needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Triton Lang access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton Lang safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton Lang use?

Triton Lang is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton Lang use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton Lang?

Skills that share tags, products or a category with Triton Lang: Magpie Kernel Evaluator (amd/skills, 408 stars), Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Cuda Index Width (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton Lang?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 252 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.