Magpie Kernel Evaluator
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/triton-lang .claude/skills/triton-lang && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .claude/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-langType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu/triton-lang .agents/skills/triton-lang && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .agents/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu/triton-lang .cursor/skills/triton-lang && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .cursor/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitmishra786/low-level-dev-skills.git --path skills/gpu/triton-lang--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu/triton-lang .gemini/skills/triton-lang && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .gemini/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitmishra786/low-level-dev-skills triton-langInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu/triton-lang .github/skills/triton-lang && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .github/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills triton-lang --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu/triton-lang .opencode/skills/triton-lang && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "triton-lang" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/triton-lang into .opencode/skills/triton-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-lang", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
triton-langTriton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.
Triton Lang is an agent skill from mohitmishra786/low-level-dev-skills. Triton language skill for Python GPU kernel authoring. Use when writing Triton kernels with @triton.jit, tl.load/store, masking, atomics, benchmarking with triton.testing, or integrating kernels into PyTorch. Activates on queries about Triton, tl.constexpr, block pointers, Triton benchmarking, or PyTorch custom ops.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning and GPU and accelerator computing. It works with PyTorch, Python and CUDA. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triton Lang loads about 1.8k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 308 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 308 words, ~1,768 tokens.
.claude/skills/triton-lang/SKILL.md (or your agent's skills folder).Guide agents through writing GPU kernels in OpenAI Triton: the @triton.jit decorator, block-oriented tl.load/tl.store with masking, atomic operations, shared memory via tl.constexpr, benchmarking with triton.testing.Benchmark, PyTorch integration, and debugging with barriers.
import torch
import triton
import triton.language as tl
@triton.jit
def add_kernel(x_ptr, y_ptr, out_ptr, n, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offsets = pid * BLOCK + tl.arange(0, BLOCK)
mask = offsets < n
x = tl.load(x_ptr + offsets, mask=mask)
y = tl.load(y_ptr + offsets, mask=mask)
tl.store(out_ptr + offsets, x + y, mask=mask)
def add(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
n = x.numel()
out = torch.empty_like(x)
grid = lambda meta: (triton.cdiv(n, meta["BLOCK"]),)
add_kernel[grid](x, y, out, n, BLOCK=1024)
return outKey concepts:
tl.program_id(0) — block index (like blockIdx.x)tl.arange(0, BLOCK) — vector of thread indices within blockmask — predication for tail elements (no separate bounds kernel)BLOCK: tl.constexpr — compile-time constant, enables unrolling@triton.jit
def masked_load_example(ptr, n, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < n
# masked load returns 0 for masked-off lanes
vals = tl.load(ptr + offs, mask=mask, other=0.0)
return valsBlock pointers (Triton 2.x+) for structured 2D access:
@triton.jit
def matvec_kernel(a_ptr, x_ptr, y_ptr, M, N, BLOCK_M: tl.constexpr, BLOCK_N: tl.constexpr):
pid_m = tl.program_id(0)
offs_m = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
acc = tl.zeros((BLOCK_M,), dtype=tl.float32)
for start_n in range(0, N, BLOCK_N):
offs_n = start_n + tl.arange(0, BLOCK_N)
a = tl.load(a_ptr + offs_m[:, None] * N + offs_n[None, :])
x = tl.load(x_ptr + offs_n)
acc += tl.sum(a * x[None, :], axis=1)
tl.store(y_ptr + offs_m, acc, mask=offs_m < M)@triton.jit
def atomic_histogram(data_ptr, hist_ptr, n, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < n
data = tl.load(data_ptr + offs, mask=mask)
bucket = (data % 256).to(tl.int32)
tl.atomic_add(hist_ptr + bucket, 1, mask=mask)Use atomics sparingly — they serialize memory updates. Prefer block-level reduction then single atomic per block.
@triton.jit
def reduce_kernel(x_ptr, out_ptr, n, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < n
x = tl.load(x_ptr + offs, mask=mask, other=0.0)
# Block reduction
x = tl.sum(x, axis=0)
tl.atomic_add(out_ptr, x)BLOCK as tl.constexpr lets the compiler allocate shared memory and unroll loops at compile time.
@triton.autotune(
configs=[
triton.Config({"BLOCK": 128}, num_warps=4),
triton.Config({"BLOCK": 256}, num_warps=4),
triton.Config({"BLOCK": 512}, num_warps=8),
],
key=["n"],
)
@triton.jit
def tuned_kernel(x_ptr, y_ptr, out_ptr, n, BLOCK: tl.constexpr):
# ... kernel body ...
passAutotuner benchmarks configs on first run and caches the best for each key shape.
from triton.testing import Benchmark
def benchmark_add():
n = 1024 * 1024
x = torch.randn(n, device="cuda")
y = torch.randn(n, device="cuda")
def triton_add():
return add(x, y)
def torch_add():
return x + y
bench = Benchmark(
x_names=["n"],
x_vals=[2**i for i in range(10, 24)],
line_arg="provider",
line_vals=["triton", "torch"],
line_names=["Triton", "PyTorch"],
plot_name="add-bench",
args={},
)
bench.run(lambda n, provider: {
"triton": lambda: add(x[:n], y[:n]),
"torch": lambda: x[:n] + y[:n],
}[provider](), quantiles=[0.5, 0.9])# Quick timing in REPL
import triton.testing as tt
ms = tt.do_bench(lambda: add(x, y))
print(f"{ms:.3f} ms")import torch
from torch.library import custom_op
@custom_op("mylib::triton_add", mutates_args=())
def triton_add_op(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
return add(x, y)
@triton_add_op.register_fake
def _(x, y):
return torch.empty_like(x)
# Use in model
class MyModule(torch.nn.Module):
def forward(self, x, y):
return triton_add_op(x, y)For torch.compile compatibility, register fake/meta kernels and avoid Python side effects in the JIT function.
@triton.jit
def debug_kernel(x_ptr, n, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < n
x = tl.load(x_ptr + offs, mask=mask)
# Synchronize threads within block for inspection
tl.debug_barrier()
tl.store(x_ptr + offs, x * 2.0, mask=mask)# Dump generated PTX/LLVM IR
TRITON_PRINT_AUTOTUNING=1 python script.py
# Set cache dir to inspect compiled kernels
export TRITON_CACHE_DIR=/tmp/triton_cacheCompare against PyTorch reference on small inputs before scaling up.
| Symptom | Cause | Fix |
|---|---|---|
OutOfResources | BLOCK too large | Reduce BLOCK or num_warps |
| Wrong results on tail | Missing mask | Add mask=offs < n to load/store |
| Slower than PyTorch | Suboptimal BLOCK | Use @triton.autotune |
CompilationError | Type mismatch | Ensure consistent dtypes; use .to(tl.float32) |
| NaN in output | Uninitialized masked lanes | Pass other=0.0 to masked loads |
| torch.compile fails | No fake kernel | Register register_fake meta function |
skills/gpu/cuda — underlying CUDA concepts when Triton limits are hitskills/gpu/cuda-profiling — Nsight profiling of Triton-compiled kernelsskills/gpu/gpu-memory-model — coalescing and occupancy theoryskills/gpu/hip-rocm — AMD GPU path (Triton supports ROCm)skills/hpc/openmp — CPU parallelism alongside GPU kernels© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu/triton-lang of mohitmishra786/low-level-dev-skills.
Open the folder on GitHubat commit bdc5847
Triton Lang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triton Lang this skillmohitmishra786/low-level-dev-skills | 252 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Magpie Kernel Evaluatoramd/skills | 408 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Hyperpod Version Checkerawslabs/agent-plugins | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 |
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
mohitmishra786/low-level-dev-skills
Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.
mohitmishra786/low-level-dev-skills
Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.
mohitmishra786/low-level-dev-skills
Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.
mohitmishra786/low-level-dev-skills
Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.
mohitmishra786/low-level-dev-skills
Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.
Categories
Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. Triton Lang is an agent skill from mohitmishra786/low-level-dev-skills. Triton language skill for Python GPU kernel authoring.
Triton Lang fits situations like: writing Triton kernels with @triton.jit; benchmarking with triton.testing; integrating kernels into PyTorch.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a claude-code`. Or copy the skill folder (skills/gpu/triton-lang in mohitmishra786/low-level-dev-skills) into .claude/skills/triton-lang in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a codex`. Or copy the skill folder (skills/gpu/triton-lang in mohitmishra786/low-level-dev-skills) into .agents/skills/triton-lang in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill triton-lang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-lang, .gemini/skills/triton-lang, .github/skills/triton-lang and .opencode/skills/triton-lang in your project.
Going by SKILL.md and its folder, Triton Lang needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Triton Lang is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Triton Lang: Magpie Kernel Evaluator (amd/skills, 408 stars), Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Cuda Index Width (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 252 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.
Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.