Official agent skill

Kernel Perf Testing

by facebookexperimental in facebookexperimental/triton

Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

OfficialMITAuto-check passedAI & LLM Engineering

Install Kernel Perf Testing

skills CLI
$ npx skills add facebookexperimental/triton --skill kernel-perf-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton kernel-perf-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kernel-perf-testing .claude/skills/kernel-perf-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kernel-perf-testing
GitHub stars
201
Token cost
~1.1k tokens
SKILL.md length
237 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

  • Works in 3 steps: Run nvidia-smi to check GPU occupancy. → Pick the GPU with the lowest memory usage. → Set CUDA_VISIBLE_DEVICES to that GPU.
  • User asks to benchmark
  • SKILL.md covers NVIDIA (Hopper / Blackwell), AMD (gfx950 / CDNA4, gfx1250), If tests hang and Interpreting results
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Kernel Perf Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs. Use when user asks to benchmark, profile, or measure performance of any TLX kernel (GEMM, Flash Attention, addmm+GLU, IKBO variants). Handles GPU selection, denoise wrapping (NVIDIA only), and version flags. Never run unless explicitly asked.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • User asks to benchmark
  • Measure performance of any TLX kernel (GEMM
  • Flash Attention

Example prompts

  • “/kernel-perf-testing”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run nvidia-smi to check GPU occupancy.
  2. Pick the GPU with the lowest memory usage.
  3. Set CUDA_VISIBLE_DEVICES to that GPU.

What it can do on your machine

Read from SKILL.md and the folder at commit 77b4a53. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kernel Perf Testing loads about 1.1k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 237 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 77b4a53, republished under its MIT licence (© facebookexperimental). 237 words, ~1,062 tokens.

Download SKILL.mdSave it as .claude/skills/kernel-perf-testing/SKILL.md (or your agent's skills folder).
name
kernel-perf-testing
description
Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs. Use when user asks to benchmark, profile, or measure performance of any TLX kernel (GEMM, Flash Attention, addmm+GLU, IKBO variants). Handles GPU selection, denoise wrapping (NVIDIA only), and version flags. Never run unless explicitly asked.
disable-model-invocation
true

Kernel Performance Testing

Never run performance tests unless the user explicitly asks.

Perf scripts live in third_party/tlx/tutorials/testing/test_<arch>_<op>_perf.py, one per (op, input-contract, arch). Each takes [--version ...] to select a kernel variant; with no --version it runs all variants for that op.

NVIDIA (Hopper / Blackwell)

GPU selection protocol
  1. Run nvidia-smi to check GPU occupancy.
  2. Pick the GPU with the lowest memory usage.
  3. Set CUDA_VISIBLE_DEVICES to that GPU.

Wrap benchmarks with denoise.sh (locks clocks/power) for stable results.

Hopper GPU
bash
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_hopper_gemm_perf.py [--version {ws|pipelined}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_hopper_fa_perf.py [--version ws_pipelined_pingpong]
Blackwell GPU
bash
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_gemm_perf.py [--version {ws|clc|pipelined|2cta}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_fa_perf.py [--version {ws|ws_persistent|ws_pipelined|ws_pipelined_persistent|clc}] [--mode fwd|bwd]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_blackwell_fa_mxfp8_perf.py

AMD (gfx950 / CDNA4, gfx1250)

denoise.sh works on AMD too: it runs the wrapped command with NUMA pinning. Its GPU clock/power lock is implemented with nvidia-smi, so on AMD that part is simply skipped (no clock lock) — expect a bit more run-to-run variance.

Pick a free GPU with rocm-smi (check GPU use (%) / VRAM). denoise.sh reads CUDA_VISIBLE_DEVICES, which torch on ROCm also honors for device selection.

gfx950 (CDNA4 / MI350-class)
bash
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_gemm_perf.py [--version {warp_pipeline|pipelined}] [--dtype fp16|bf16]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_fa_perf.py [--version {simple|prefetch|persistent}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_addmm_glu_perf.py [--version {tlx_baseline|tlx_simple_async|tlx_optimized_async|tlx_optimized|tlx_persistent}]
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_ikbo_fa_perf.py   # IKBO Flash Attention
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_ikbo_lce_perf.py  # IKBO LCE (distinct op, not attention)
gfx1250
bash
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/testing/test_amd_mxfp_gemm_perf.py [--version {tdm_pipelined}] [--transpose-b]

Each script self-gates (is_hip, is_hip_cdna4, or is_hip_gfx1250) and prints "Skipping benchmarks" on the wrong hardware.

Other kernels
bash
CUDA_VISIBLE_DEVICES=<gpu_id> third_party/tlx/denoise.sh python third_party/tlx/tutorials/<KERNEL.py>

If tests hang

Run third_party/tlx/killgpu.sh to kill GPU processes that have been running too long.

Interpreting results

  • Output reports TFLOPS for each problem size and configuration.
  • A reference baseline is printed alongside the Triton/TLX results: cuBLAS/rocBLAS for GEMM, SDPA for FA, the PyTorch pytorch_baseline for addmm+GLU.
  • Higher TFLOPS = better. Look for regressions relative to previous runs.
  • Check for consistency across runs — high variance suggests noisy measurements (ensure denoise.sh is being used).

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/kernel-perf-testing of facebookexperimental/triton.

Open the folder on GitHubat commit 77b4a53

Compare with similar skills

Kernel Perf Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kernel Perf Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kernel Perf Testing this skillfacebookexperimental/triton201—~1.1kAutomated safety check: PassMIT
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Yolo Detection 2026SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT
Sglang Diffusion Modelopt Quantsgl-project/sglang37k2 repos~5kAutomated safety check: PassApache-2.0

Similar skills

  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Yolo Detection 2026

    SharpAI/DeepCamera

    YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

    3.1k GitHub stars~1.5k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

    37k GitHub starsUsed in 2 repos~5k tokens
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    143 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Questions about Kernel Perf Testing

What does Kernel Perf Testing do?

Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs. Kernel Perf Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

When should I use Kernel Perf Testing?

Kernel Perf Testing fits situations like: user asks to benchmark; measure performance of any TLX kernel (GEMM; flash Attention.

How do I install Kernel Perf Testing in Claude Code?

Run `npx skills add facebookexperimental/triton --skill kernel-perf-testing -a claude-code`. Or copy the skill folder (.claude/skills/kernel-perf-testing in facebookexperimental/triton) into .claude/skills/kernel-perf-testing in your project. Claude Code loads it when a task matches its description.

How do I install Kernel Perf Testing in Codex?

Run `npx skills add facebookexperimental/triton --skill kernel-perf-testing -a codex`. Or copy the skill folder (.claude/skills/kernel-perf-testing in facebookexperimental/triton) into .agents/skills/kernel-perf-testing in your project. Codex loads it when a task matches its description.

Can I use Kernel Perf Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill kernel-perf-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kernel-perf-testing, .gemini/skills/kernel-perf-testing, .github/skills/kernel-perf-testing and .opencode/skills/kernel-perf-testing in your project.

What does Kernel Perf Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Kernel Perf Testing is instructions for the agent only. Our summary lists: Python 3.

Does Kernel Perf Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kernel Perf Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kernel Perf Testing use?

Kernel Perf Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kernel Perf Testing use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kernel Perf Testing?

Skills that share tags, products or a category with Kernel Perf Testing: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars) and Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kernel Perf Testing?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 7, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.