Agent skill

Triton Kernel Writing

by guqiong96 in guqiong96/Lvllm

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton Kernel Writing

skills CLI
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .claude/skills/triton-kernel-writing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triton-kernel-writing
GitHub stars
465
Used in
1 other repo
Token cost
~831 tokens
SKILL.md length
412 words
Files
2
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers Implementation, Launch and Indexing and Validation
  • Calls python
  • Tasks that involve LLM inference and serving

What it does

Triton Kernel Writing is an agent skill from guqiong96/Lvllm. Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

Its SKILL.md is about 830 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing and LLM inference and serving. It works with vLLM. The repository describes itself as: LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing
  • Tasks that involve LLM inference and serving

Example prompts

  • “/triton-kernel-writing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 43ffc42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • triton-lang.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton Kernel Writing loads about 831 tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 412 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~831

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guqiong96/Lvllm at commit 43ffc42, republished under its Apache-2.0 licence (© guqiong96). 412 words, ~831 tokens.

Download SKILL.mdSave it as .claude/skills/triton-kernel-writing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
triton-kernel-writing
description
Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

Triton Kernel Writing

Implementation

  • Follow the official Triton semantics. Check it when behavior may differ from Python or NumPy, especially type promotion, integer division and modulo, casts, broadcasting, and variable scoping.
  • Use the Triton kernel generated by torch.compile as a possible implementation to inspect. Print Inductor's generated code with TORCH_LOGS="output_code" .venv/bin/python <script> or enable torch._logging.set_logs(output_code=True) before the compiled function runs. Treat generated code as a reference, not as proof of correctness or optimality.
  • Find reasonable defaults for compile-time knobs such as BLOCK_SIZE, or use a small, legible heuristic when workloads need different choices. Use triton.autotune only when tuning is critical to performance, such as for a matrix multiplication. Otherwise prioritize simple code and fast startup.
  • Be careful to avoid unintended runtime JIT compilation. For example, put unimportant runtime integer scalars in do_not_specialize, especially those that may alternate between values such as 0 and 1, which can produce different specialization keys.
  • The Triton compiler does not guarantee safe ordering when a kernel writes to a pointer and subsequently reads from the same pointer. This pattern must have a tl.debug_barrier() between the write and read. The barrier synchronizes threads in the block; it does not synchronize separate program instances.
Show full SKILL.md (215 more words)Show less

Launch and Indexing

  • grid[1] and grid[2] must be at most 65,535. Choose or flatten the grid order so those dimensions cannot exceed the limit for supported shapes. For example, num_tokens is commonly 8K or 16K, but users may configure 32K or more. If num_tokens is a grid dimension, it is safe to put it in grid[0] (or tile it).
  • Use int64 for offset arithmetic when an index can exceed 32-bit range, especially for KV-cache addressing. Cast operands before multiplication or addition so an intermediate does not overflow in 32-bit arithmetic.
  • A [num_tokens, num_heads] grid can be a good low-latency mapping for decode, but it can be very slow for prefill. If the kernel serves prefill, consider tiling tokens or otherwise increasing the work and locality per program.

Validation

  • Check correctness at boundary shapes and at sizes that exercise masks and large offsets.
  • Choose accumulation and intermediate dtypes explicitly. Test numerically difficult inputs, not only random, well-scaled tensors.
  • Use $kernel-microbenchmark for benchmark construction, measurement, and interpretation.
  • Benchmark a sweep of num_tokens covering decode and representative prefill workloads. Include relevant head counts and dimensions when they affect the launch shape, and do not select an implementation or tuning heuristic from a single setup.
  • Include compilation or autotuning overhead when evaluating startup behavior; report steady-state kernel performance separately.

© guqiong96, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/triton-kernel-writing of guqiong96/Lvllm.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 43ffc42

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in guqiong96/Lvllm, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Triton Kernel Writing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton Kernel Writing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton Kernel Writing this skillguqiong96/Lvllm4651 repos~831Automated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
LLM Serving Capacity PlannerBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.5kAutomated safety check: PassNone
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    938 GitHub stars~2.5k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    938 GitHub stars~3.9k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

More from guqiong96/Lvllm

  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    Auto-check passed
  • Kernel Microbenchmark

    guqiong96/Lvllm

    Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

    465 GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed

Works with

Questions about Triton Kernel Writing

What does Triton Kernel Writing do?

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. Triton Kernel Writing is an agent skill from guqiong96/Lvllm. Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

When should I use Triton Kernel Writing?

Triton Kernel Writing fits situations like: tasks that involve GPU and accelerator computing; tasks that involve LLM inference and serving.

How do I install Triton Kernel Writing in Claude Code?

Run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a claude-code`. Or copy the skill folder (.agents/skills/triton-kernel-writing in guqiong96/Lvllm) into .claude/skills/triton-kernel-writing in your project. Claude Code loads it when a task matches its description.

How do I install Triton Kernel Writing in Codex?

Run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a codex`. Or copy the skill folder (.agents/skills/triton-kernel-writing in guqiong96/Lvllm) into .agents/skills/triton-kernel-writing in your project. Codex loads it when a task matches its description.

Can I use Triton Kernel Writing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-kernel-writing, .gemini/skills/triton-kernel-writing, .github/skills/triton-kernel-writing and .opencode/skills/triton-kernel-writing in your project.

What does Triton Kernel Writing need to run?

Going by SKILL.md and its folder, Triton Kernel Writing needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Triton Kernel Writing access the network?

SKILL.md names 1 domain. As links in the text: triton-lang.org. This is read from the text; nothing was executed.

Is Triton Kernel Writing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton Kernel Writing use?

Triton Kernel Writing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton Kernel Writing use?

About 831 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton Kernel Writing?

Skills that share tags, products or a category with Triton Kernel Writing: Hugging Face Local Model Evals (huggingface/skills, 11k stars), LLM Serving Capacity Planner (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton Kernel Writing?

guqiong96 (a GitHub user) maintains it in guqiong96/Lvllm, which has 465 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 22, 2026.

Source: guqiong96/Lvllm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.