Agent skill

Fastllm Triton Ops

by ztxz16 in ztxz16/fastllm

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Fastllm Triton Ops

skills CLI
$ npx skills add ztxz16/fastllm --skill fastllm-triton-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ztxz16/fastllm fastllm-triton-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ztxz16/fastllm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/fastllm-triton-ops .claude/skills/fastllm-triton-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-triton-ops
GitHub stars
5.1k
Token cost
~1.8k tokens
SKILL.md length
671 words
Files
2
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

  • Works in 6 steps: Inspect the existing CUDA op path in… → Add or extend the Python compile path in… → Add a CUDA launch wrapper in… → …
  • Modifying FastLLM CUDA op code to add
  • SKILL.md covers Overview, Workflow, Current Contract and C++ Pattern, plus 2 more sections
  • Calls bash

What it does

Fastllm Triton Ops is an agent skill from ztxz16/fastllm. Guide for adding Triton-backed CUDA operators to FastLLM. Use when modifying FastLLM CUDA op code to add, extend, debug, validate, or benchmark Triton-generated kernels through tools/fastllmtritonserver.py, src/devices/cuda/cudadevice.cpp, src/devices/cuda/fastllm-triton-cuda.cu, include/devices/cuda/fastllm-cuda.cuh, or related CMake wiring.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering. It works with C++ and CUDA. The repository describes itself as: fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tps,多并发可达60+。 The licence is Apache-2.0.

When your agent uses it

  • Modifying FastLLM CUDA op code to add
  • Benchmark Triton-generated kernels through tools/fastllmtritonserver.py
  • Src/devices/cuda/cudadevice.cpp
  • Src/devices/cuda/fastllm-triton-cuda.cu

Example prompts

  • “/fastllm-triton-ops”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Inspect the existing CUDA op path in src/devices/cuda/cudadevice.cpp.
  2. Add or extend the Python compile path in tools/fastllm_triton_server.py.
  3. Add a CUDA launch wrapper in src/devices/cuda/fastllm-triton-cuda.cu.
  4. Declare the wrapper in include/devices/cuda/fastllm-cuda.cuh.
  5. Wire the op in src/devices/cuda/cudadevice.cpp.
  6. Update build wiring only when needed.

What it can do on your machine

Read from SKILL.md and the folder at commit a2ff521. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Triton Ops loads about 1.8k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 671 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ztxz16/fastllm at commit a2ff521, republished under its Apache-2.0 licence (© ztxz16). 671 words, ~1,784 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-triton-ops/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
fastllm-triton-ops
description
Guide for adding Triton-backed CUDA operators to FastLLM. Use when modifying FastLLM CUDA op code to add, extend, debug, validate, or benchmark Triton-generated kernels through tools/fastllm_triton_server.py, src/devices/cuda/cudadevice.cpp, src/devices/cuda/fastllm-triton-cuda.cu, include/devices/cuda/fastllm-cuda.cuh, or related CMake wiring.

FastLLM Triton Ops

Overview

Use the existing Linear Triton prototype as the reference architecture: C++ decides whether an op is eligible, starts a local Python compiler server only on demand, asks it to emit a cached cubin plus metadata, then launches that cubin through the CUDA Driver API. Every Triton path must be environment-gated and must fall back to the original CUDA implementation on unsupported inputs or compile/launch failure.

Workflow

  1. Inspect the existing CUDA op path in src/devices/cuda/cudadevice.cpp.

    • Find the op's Reshape, CanRun, Run, and lower-level helper functions.
    • Identify the exact tensor layout, dtype combinations, shape variables, optional bias/scale tensors, and output aliasing behavior.
    • Keep the original implementation as the fallback path.
  2. Add or extend the Python compile path in tools/fastllm_triton_server.py.

    • Add a @triton.jit kernel for the new op.
    • Add <op>_cache_paths(payload) with a deterministic filename that includes op name, dtype/layout variants, SM arch, compile-time tile sizes, and feature flags.
    • Add compile_<op>(payload) that validates payload fields, compiles with ASTSource, writes .cubin, writes .json metadata, and returns the metadata.
    • Extend handle_compile(payload) by dispatching on payload["op"].
    • Keep /health and /compile stable; do not break existing "op": "linear" requests.
  3. Add a CUDA launch wrapper in src/devices/cuda/fastllm-triton-cuda.cu.

    • Reuse LoadTritonKernel for cubin/module/function caching.
    • Add one extern "C" wrapper per op, for example FastllmCudaTriton<Op>(...).
    • Use FastllmCudaPrepareInput, FastllmCudaPrepareOutput, FastllmCudaFinishInput, and FastllmCudaFinishOutput when passing Data buffers.
    • Match the Triton kernel argument order exactly, including Triton's hidden global_scratch and profile_scratch pointer arguments when needed by AOT metadata.
    • Return false on load or launch failure so the C++ caller can fall back.
  4. Declare the wrapper in include/devices/cuda/fastllm-cuda.cuh.

    • Keep the signature close to the CUDA fallback helper's shape arguments.
    • Pass metadata fields needed for launch, such as kernelName, shared, numWarps, and tile sizes.
  5. Wire the op in src/devices/cuda/cudadevice.cpp.

    • Add a small metadata struct for the op's .json fields.
    • Reuse common helpers: CudaTritonCacheDir, CudaTritonDataTypeName, CudaTritonHttpRequest, and CudaTritonEnsureServer.
    • Add CudaTriton<Op>BaseName, CudaTritonRead<Op>Meta, CudaTritonRequest<Op>Kernel, and TryCudaTriton<Op>.
    • Gate with global FASTLLM_CUDA_TRITON; add an op-specific override such as FASTLLM_CUDA_TRITON_<OP>=0.
    • Validate device, pointer presence, dtype, layout, shape, arch, and feature constraints before requesting a kernel.
    • In the original op helper, call TryCudaTriton<Op>(...) immediately before the original CUDA implementation.
  6. Update build wiring only when needed.

    • src/devices/cuda/fastllm-triton-cuda.cu is already in CMakeLists.txt.
    • If adding new files, update CMakeLists.txt and link dependencies without changing unrelated targets.
Show full SKILL.md (292 more words)Show less

Current Contract

The existing Triton infrastructure uses these environment variables:

  • FASTLLM_CUDA_TRITON=1: enable Triton-backed CUDA ops globally.
  • FASTLLM_CUDA_TRITON_<OP>=0: disable one op while keeping the global flag on, for example FASTLLM_CUDA_TRITON_LINEAR=0.
  • FASTLLM_CUDA_TRITON_CACHE_DIR: override cubin/json cache directory.
  • FASTLLM_CUDA_TRITON_SERVER_HOST, FASTLLM_CUDA_TRITON_SERVER_PORT: choose compiler server endpoint.
  • FASTLLM_CUDA_TRITON_PYTHON: choose Python interpreter.
  • FASTLLM_CUDA_TRITON_SERVER_SCRIPT: choose server script path.
  • FASTLLM_CUDA_TRITON_SERVER_LOG: choose compiler server log path.
  • FASTLLM_CUDA_TRITON_SERVER_WAIT_MS: choose startup wait timeout.
  • Per-op tile knobs should use FASTLLM_CUDA_TRITON_<OP>_<PARAM>, matching Linear's BLOCK_M, BLOCK_N, BLOCK_K, NUM_WARPS, and NUM_STAGES.

The metadata JSON returned by the compiler server should include at least:

json
{
  "ok": true,
  "op": "op_name",
  "cubin": "/path/to/kernel.cubin",
  "kernel": "compiled_kernel_name",
  "shared": 0,
  "num_warps": 4
}

Add op-specific launch fields, such as tile sizes, only when the C++ launcher needs them.

C++ Pattern

Keep the C++ control flow shaped like this:

cpp
if (!CudaEnvFlagEnabled("FASTLLM_CUDA_TRITON")) {
    return false;
}
const char *opEnv = std::getenv("FASTLLM_CUDA_TRITON_MYOP");
if (opEnv != nullptr && opEnv[0] != '\0' && !CudaEnvFlagEnabled("FASTLLM_CUDA_TRITON_MYOP")) {
    return false;
}
if (!inputs_are_supported) {
    return false;
}

Meta meta;
if (!ReadMeta(metaPath, meta)) {
    if (!RequestKernel(..., meta)) {
        return false;
    }
}
return FastllmCudaTritonMyOp(meta.cubinPath.c_str(), meta.kernelName.c_str(), ...);

Do not throw or call ErrorInFastLLM from the Triton trial path unless the original CUDA path would also fail. Unsupported Triton cases should return false.

Validation

Run validation in layers:

  1. Build:
bash
bash install.sh -DUSE_CUDA=ON
  1. Unit or op test, with Triton enabled and an isolated cache:
bash
FASTLLM_CUDA_TRITON=1 \
FASTLLM_CUDA_TRITON_CACHE_DIR=/tmp/fastllm-triton-optest \
../optest --op linear --device cuda:0 --param batch=4 --param in=8 --param out=6

Adjust the optest command for the op being added.

  1. End-to-end server smoke test:
bash
FASTLLM_CUDA_TRITON=1 \
FASTLLM_CUDA_TRITON_CACHE_DIR=/tmp/fastllm-triton-qwen \
ftllm server ~/hfmodels/Qwen3-8B/ --device cuda:0 --host 127.0.0.1 --port 18080 --tokens 8192 --hide_input

Send one non-streaming request to /v1/chat/completions and confirm it completes.

  1. Benchmark after warmup.
    • Always exclude first compile/server startup from performance numbers.
    • Compare against the same command without FASTLLM_CUDA_TRITON.
    • Report token throughput and wall time; state whether the measurement is kernel-only or end-to-end server throughput.

Guardrails

  • Keep Triton optional: default behavior must stay unchanged when FASTLLM_CUDA_TRITON is unset.
  • Prefer compile keys that are independent of runtime shapes when the kernel supports dynamic dimensions.
  • Include every compile-time specialization in the cache filename to prevent stale cubin reuse.
  • Serialize compilation in the Python server with the existing lock unless proving concurrent compilation is safe.
  • Keep C++ JSON parsing defensive; missing or invalid metadata should fall back.
  • Clean up test servers and compiler-server processes after benchmarks.

© ztxz16, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .codex/skills/fastllm-triton-ops of ztxz16/fastllm.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit a2ff521

Compare with similar skills

Fastllm Triton Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Triton Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Triton Ops this skillztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Leetcuda Cpp Kernelxlite-dev/LeetCUDA12k—~3.5kAutomated safety check: PassGPL-3.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Paddle Op DevPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Leetcuda Cpp Kernel

    xlite-dev/LeetCUDA

    LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、…

    12k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from ztxz16/fastllm

  • Git Commit Zh Split

    ztxz16/fastllm

    当用户要求提交代码、整理提交、准备 commit、拆分 commit、push,或指定提交与推送规范时使用。默认使用中文提交信息,将差异较大的改动拆分为多个提交;推送前先执行 fetch、stash、rebase、stash pop,再 push。

    5.1k GitHub stars~359 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Fastllm Triton Ops

What does Fastllm Triton Ops do?

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. Fastllm Triton Ops is an agent skill from ztxz16/fastllm. Guide for adding Triton-backed CUDA operators to FastLLM.

When should I use Fastllm Triton Ops?

Fastllm Triton Ops fits situations like: modifying FastLLM CUDA op code to add; benchmark Triton-generated kernels through tools/fastllmtritonserver.py; src/devices/cuda/cudadevice.cpp; src/devices/cuda/fastllm-triton-cuda.cu.

How do I install Fastllm Triton Ops in Claude Code?

Run `npx skills add ztxz16/fastllm --skill fastllm-triton-ops -a claude-code`. Or copy the skill folder (.codex/skills/fastllm-triton-ops in ztxz16/fastllm) into .claude/skills/fastllm-triton-ops in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Triton Ops in Codex?

Run `npx skills add ztxz16/fastllm --skill fastllm-triton-ops -a codex`. Or copy the skill folder (.codex/skills/fastllm-triton-ops in ztxz16/fastllm) into .agents/skills/fastllm-triton-ops in your project. Codex loads it when a task matches its description.

Can I use Fastllm Triton Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ztxz16/fastllm --skill fastllm-triton-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-triton-ops, .gemini/skills/fastllm-triton-ops, .github/skills/fastllm-triton-ops and .opencode/skills/fastllm-triton-ops in your project.

What does Fastllm Triton Ops need to run?

Going by SKILL.md and its folder, Fastllm Triton Ops needs the command-line tools its instructions call (bash). Our summary lists: Python 3.

Does Fastllm Triton Ops access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fastllm Triton Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Triton Ops use?

Fastllm Triton Ops is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Triton Ops use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Triton Ops?

Skills that share tags, products or a category with Fastllm Triton Ops: Paddle Build (PaddlePaddle/Paddle, 24k stars), Leetcuda Cpp Kernel (xlite-dev/LeetCUDA, 12k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars) and Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Triton Ops?

ztxz16 (a GitHub user) maintains it in ztxz16/fastllm, which has 5,101 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.

Source: ztxz16/fastllm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.