Official agent skill

Tlx Amd Testing

by facebookexperimental in facebookexperimental/triton

Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI.

OfficialMITAuto-check: notes

Install Tlx Amd Testing

skills CLI
$ npx skills add facebookexperimental/triton --skill tlx-amd-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton tlx-amd-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tlx-amd-testing .claude/skills/tlx-amd-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tlx-amd-testing
GitHub stars
201
Token cost
~1.3k tokens
SKILL.md length
419 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI.

  • Working on AMD TLX tutorial kernels — GEMM (warp-pipeline
  • SKILL.md covers Kernel inventory, Correctness, Perf and CI, plus 1 more section
  • Calls pytest and make
  • Flash Attention (simple

What it does

Tlx Amd Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI. Use when working on AMD TLX tutorial kernels — GEMM (warp-pipeline, LDS-pipelined, TDM, MXFP), Flash Attention (simple, prefetch, persistent), addmm+GLU, or IKBO (FA, LCE) — running their correctness or perf, checking arch gating (gfx950 vs gfx1250), or the MI350 CI workflow. Covers the standardized layout (one correctness file, one perf file per op×arch).

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • Working on AMD TLX tutorial kernels — GEMM (warp-pipeline
  • Flash Attention (simple
  • LCE) — running their correctness
  • Checking arch gating (gfx950 vs gfx1250)

Example prompts

  • “/tlx-amd-testing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 953bd20. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tlx Amd Testing loads about 1.3k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 419 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:67
    node, and resets both on exit. It needs sudo and is best-effort. It

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 953bd20, republished under its MIT licence (© facebookexperimental). 419 words, ~1,259 tokens.

Download SKILL.mdSave it as .claude/skills/tlx-amd-testing/SKILL.md (or your agent's skills folder).
name
tlx-amd-testing
description
Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI. Use when working on AMD TLX tutorial kernels — GEMM (warp-pipeline, LDS-pipelined, TDM, MXFP), Flash Attention (simple, prefetch, persistent), addmm+GLU, or IKBO (FA, LCE) — running their correctness or perf, checking arch gating (gfx950 vs gfx1250), or the MI350 CI workflow. Covers the standardized layout (one correctness file, one perf file per op×arch).

TLX-AMD Tutorial Kernel Testing

AMD tutorial kernels follow the same standardized layout as the NVIDIA (Hopper/Blackwell) reference: one shared correctness file (test_correctness.py, arch-gated) and one perf script per (op, arch). Each kernel is an importable module under third_party/tlx/tutorials/.

Kernel inventory

Kernel moduleOpCorrectness test(s)Perf scriptArch gate
amd_gemm_warp_pipeline.pyGEMMtest_amd_gemm_warp_pipelinetest_amd_gemm_perf.py (warp_pipeline)is_hip_cdna4
amd_gemm_pipelined.pyGEMM (LDS pipeline)test_amd_gemm_pipelinedtest_amd_gemm_perf.py (pipelined)is_hip
amd_gemm_gfx942.pyGEMM (MI300X, autotuned)test_amd_gemm_gfx942, test_amd_gemm_gfx942_odd_shapestest_amd_gemm_gfx942_perf.pyis_hip_cdna3
amd_addmm_gfx942.pyaddmm (MI300X, autotuned)test_amd_addmm_gfx942test_amd_addmm_gfx942_perf.pyis_hip_cdna3
amd_bmm_gfx942.pyBMM (MI300X, autotuned)test_amd_bmm_gfx942, test_amd_bmm_gfx942_distinct_atest_amd_bmm_gfx942_perf.pyis_hip_cdna3
amd_fa_pipelined.pyFlash Attentiontest_amd_fa_pipelinedtest_amd_fa_perf.py (simple, prefetch)is_hip_cdna4
amd_fa_persistent.pyFlash Attention (persistent)test_amd_fa_persistent, test_amd_fa_persistent_cross_attentiontest_amd_fa_perf.py (persistent)is_hip_cdna4
amd_addmm_glu.pyaddmm + GLU (gated linear unit, not GELU)test_amd_addmm_glutest_amd_addmm_glu_perf.pyis_hip_cdna4
ikbo/ikbo_fa_triton.pyIKBO Flash Attentiontest_ikbo_fatest_amd_ikbo_fa_perf.pynone (any HIP/CUDA)
ikbo/ikbo_lce_triton.pyIKBO LCE (logit cross-entropy — not attention)test_ikbo_lcetest_amd_ikbo_lce_perf.pynone (any HIP/CUDA)
amd_tdm_gemm_pipelined.pyGEMM (TDM)test_amd_tdm_gemm_pipelined—is_hip_gfx1250
amd_mxfp_gemm_tdm_pipelined.pyGEMM (MXFP, TDM)test_amd_mxfp_gemm_tdm_pipelinedtest_amd_mxfp_gemm_perf.pyis_hip_gfx1250

gfx950 = CDNA4 = MI350-class (is_hip_cdna4()). gfx942 = CDNA3 = MI300X-class (is_hip_cdna3()). gfx1250 is a separate, newer target (is_hip_gfx1250()). On gfx950, both the gfx1250-only and the gfx942-only GEMM tests auto-skip.

The MI300X kernel is the only gfx942 entry and has no CI runner — the MI350 workflow is gfx950, so test_amd_gemm_gfx942* always skips there. Run it by hand on an MI300X box.

Correctness

All AMD correctness lives in the single shared file; tests self-gate via @pytest.mark.skipif, so only the relevant cases run per GPU.

bash
# All AMD + IKBO (gfx1250-only cases auto-skip on gfx950):
pytest third_party/tlx/tutorials/testing/test_correctness.py -v -k "amd or ikbo"

# Whole file — Hopper/Blackwell cases auto-skip on AMD (what CI runs, no -k):
pytest third_party/tlx/tutorials/testing/test_correctness.py -v

-k "amd" alone does not select the IKBO tests (test_ikbo_* has no "amd" in its node id) — use -k "amd or ikbo".

Show full SKILL.md (171 more words)Show less

Perf

Never run perf unless explicitly asked. Use the kernel-perf-testing skill for run mechanics. denoise.sh does lock clocks on AMD: it identifies the part by PCI id (MI300X 0x74a0/0x74a1/gfx942, MI350X, MI355X), then applies rocm-smi --setperfdeterminism (default 2100 MHz, override DETERMINISM_CLK) and --setpoweroverdrive (750 W on MI300X, override DESIRED_POWER), NUMA-binds to the GPU's node, and resets both on exit. It needs sudo and is best-effort. It defaults HIP_VISIBLE_DEVICES to 4 — set it to a free GPU (rocm-smi) yourself.

CI

.github/workflows/mi350.yml runs on a gfx950 (MI350/CDNA4) runner and mirrors .github/workflows/h100.yml:

  • mi350-tlx-test — TLX unit tests (python/test/unit/language/test_tlx_*.py)
    • the tutorial correctness suite (test_correctness.py). AMD/IKBO run; Hopper/Blackwell and gfx1250 cases auto-skip.
  • mi350-meta-triton-test — TritonBench perf coverage (perf-regression lives here, not in the perf scripts above).

Nightly failures are filed as issues via report-nightly-failure.yml.

Local run note

After any C++ change (or a stale checkout), the in-tree libtriton.so can lag the Python source and every AMD kernel fails at compile with AttributeError: module '...amd.passes.ttgpuir' has no attribute '<pass>'. Fix: rebuild with make dev-install-llvm. If GPU tests hang, run third_party/tlx/killgpu.sh.

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tlx-amd-testing of facebookexperimental/triton.

Open the folder on GitHubat commit 953bd20

Compare with similar skills

Tlx Amd Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tlx Amd Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tlx Amd Testing this skillfacebookexperimental/triton201—~1.3kAutomated safety check: NotesMIT
Kernel Organizationsgl-project/sglang37k—~1.3kAutomated safety check: PassApache-2.0
Add Sgl Kernelsgl-project/sglang37k2 repos~3.4kAutomated safety check: PassApache-2.0
Add Jit Kernelsgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence
TutorialQ00/ouroboros6.2k—~1.6kAutomated safety check: PassMIT

Similar skills

  • Kernel Organization

    sgl-project/sglang

    Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.

    37k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Sgl Kernel

    sgl-project/sglang

    Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

    37k GitHub starsUsed in 2 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Jit Kernel

    sgl-project/sglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

    37k GitHub stars~13k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Tutorial

    Q00/ouroboros

    Interactive tutorial teaching Ouroboros hands-on. An agent skill from Q00/ouroboros.

    6.2k GitHub stars~1.6k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Tutorial Engineer

    sickn33/agentic-awesome-skills

    Creates step-by-step tutorials and educational content from code.

    47k GitHub starsUsed in 2 repos~4.1k tokens
    EducationAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Questions about Tlx Amd Testing

What does Tlx Amd Testing do?

Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI. Tlx Amd Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Test and run TLX-AMD tutorial kernels (gfx950/CDNA4 and gfx1250) and understand their CI.

When should I use Tlx Amd Testing?

Tlx Amd Testing fits situations like: working on AMD TLX tutorial kernels — GEMM (warp-pipeline; flash Attention (simple; LCE) — running their correctness; checking arch gating (gfx950 vs gfx1250).

How do I install Tlx Amd Testing in Claude Code?

Run `npx skills add facebookexperimental/triton --skill tlx-amd-testing -a claude-code`. Or copy the skill folder (.claude/skills/tlx-amd-testing in facebookexperimental/triton) into .claude/skills/tlx-amd-testing in your project. Claude Code loads it when a task matches its description.

How do I install Tlx Amd Testing in Codex?

Run `npx skills add facebookexperimental/triton --skill tlx-amd-testing -a codex`. Or copy the skill folder (.claude/skills/tlx-amd-testing in facebookexperimental/triton) into .agents/skills/tlx-amd-testing in your project. Codex loads it when a task matches its description.

Can I use Tlx Amd Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tlx-amd-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tlx-amd-testing, .gemini/skills/tlx-amd-testing, .github/skills/tlx-amd-testing and .opencode/skills/tlx-amd-testing in your project.

What does Tlx Amd Testing need to run?

Going by SKILL.md and its folder, Tlx Amd Testing needs the command-line tools its instructions call (pytest and make). Our summary lists: Python 3.

Does Tlx Amd Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tlx Amd Testing safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tlx Amd Testing use?

Tlx Amd Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tlx Amd Testing use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tlx Amd Testing?

Skills that share tags, products or a category with Tlx Amd Testing: Kernel Organization (sgl-project/sglang, 37k stars), Add Sgl Kernel (sgl-project/sglang, 37k stars), Add Jit Kernel (sgl-project/sglang, 37k stars) and Metal Kernel (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tlx Amd Testing?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 9, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.