Official agent skill

Sched2tlx Perf Testing

by facebookexperimental in facebookexperimental/triton

Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaledmm).

OfficialMITAuto-check passedProductivity & Automation

Install Sched2tlx Perf Testing

skills CLI
$ npx skills add facebookexperimental/triton --skill sched2tlx-perf-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton sched2tlx-perf-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/sched2tlx-perf-testing .claude/skills/sched2tlx-perf-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sched2tlx-perf-testing
GitHub stars
201
Token cost
~1.9k tokens
SKILL.md length
778 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaledmm).

  • Works in 2 steps: Buck first. If buck2 is available, use… → Repo venv fallback. Only when buck2 is…
  • The user asks to benchmark generated-vs-handwritten kernels
  • SKILL.md covers Performance testing…, Build and execution priority…, The one command: compare and Scheduler provenance, plus 2 more sections
  • Calls make, uv and python

What it does

Sched2tlx Perf Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaledmm). Use when the user asks to benchmark generated-vs-handwritten kernels, check corpus correctness, compare emitter revisions, or regenerate schedulegraph.json fixtures. Never run perf unless explicitly asked.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation. It works with C++. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • The user asks to benchmark generated-vs-handwritten kernels
  • Check corpus correctness
  • Compare emitter revisions
  • Regenerate schedulegraph.json fixtures

Example prompts

  • “/sched2tlx-perf-testing”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Buck first. If buck2 is available, use buck2 run to compile and
  2. Repo venv fallback. Only when buck2 is unavailable, build and run all

What it can do on your machine

Read from SKILL.md and the folder at commit 6f3dd70. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make
    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sched2tlx Perf Testing loads about 1.9k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 778 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 6f3dd70, republished under its MIT licence (© facebookexperimental). 778 words, ~1,917 tokens.

Download SKILL.mdSave it as .claude/skills/sched2tlx-perf-testing/SKILL.md (or your agent's skills folder).
name
sched2tlx-perf-testing
description
Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaled_mm). Use when the user asks to benchmark generated-vs-handwritten kernels, check corpus correctness, compare emitter revisions, or regenerate schedule_graph.json fixtures. Never run perf unless explicitly asked.
disable-model-invocation
true

sched2tlx Perf & Correctness Harness

Never run performance tests unless the user explicitly asks.

Harness: third_party/tlx/tools/sched2tlx/examples/testing/perf_regression/perf_harness.py Corpus: third_party/tlx/tools/sched2tlx/examples/case*/

Performance testing prerequisites

Before running any performance test for a C++ scheduler change, rebuild Triton, then regenerate both schedule_graph.json and generated.py for every case being tested. Do not benchmark stale fixtures.

Build and execution priority (critical)

Use one build path consistently for compilation, fixture regeneration, correctness, and timing:

  1. Buck first. If buck2 is available, use buck2 run to compile and run every performance test, including compare and per-case benchmarks. Load the running-with-buck skill for the required working directory, @mode/opt, beta-Triton modifier, GPU architecture, and CUDA flags. Do not silently fall back to the repo venv when Buck is available. If the requested benchmark has no runnable Buck target, report the missing target instead.
  2. Repo venv fallback. Only when buck2 is unavailable, build and run all performance tests with $REPO/.venv. The login shell's lmod modules break both build and runtime, so prefix every Python and triton-opt invocation with env -u LD_LIBRARY_PATH. Ignore the lua/posix noise every command prints. Use one of the repo-venv rebuild methods below.
Rebuilding Triton
  • Buck build. On the Buck path, buck2 build rebuilds the selected target and its changed C++ dependencies incrementally; buck2 run does the same before running it. Follow the running-with-buck skill and never guess a target name. A Buck rebuild does not update source-tree schedule_graph.json or generated.py unless the selected target explicitly regenerates them.
  • Incremental repo-venv build. After the development build has already been initialized, rebuild C++ changes from the Triton root with:
    bash
    env -u LD_LIBRARY_PATH \
      PATH="$REPO/.venv/bin:$HOME/.local/bin:/usr/local/bin:/usr/bin:/bin" \
      VIRTUAL_ENV="$REPO/.venv" PYTHON="$REPO/.venv/bin/python" make
  • Repo-venv editable rebuild. Use this to initialize or refresh the editable Triton installation used by the fallback workflow:
    bash
    env -u LD_LIBRARY_PATH \
      PATH="$REPO/.venv/bin:$HOME/.local/bin:/usr/local/bin:/usr/bin:/bin" \
      VIRTUAL_ENV="$REPO/.venv" CC=/usr/bin/gcc CXX=/usr/bin/g++ MAX_JOBS=14 \
      uv pip install -e . --no-build-isolation
    make dev-install-triton is the Makefile wrapper for the same editable installation flow when PYTHON points to $REPO/.venv/bin/python.

Regardless of the rebuild path, regenerate schedule_graph.json and generated.py for every selected case before performance testing a scheduler change.

The one command: compare

When Buck is available, run the applicable performance-runner target from fbsource/fbcode, following the running-with-buck skill, and pass compare, --rev, and --cases as program arguments after --.

Only when Buck is unavailable, run from examples/testing/perf_regression/:

env -u LD_LIBRARY_PATH $REPO/.venv/bin/python perf_harness.py compare \
    [--rev origin/main] [--cases case7_wgrad_bias,case9_scaled_mm/blockwise]

One row per case, four columns:

columnmeaning
casecase dir relative to examples/ (nested variants like case9_scaled_mm/blockwise included)
main (gen/hw)per-shape gen/handwritten throughput ratios for --rev's committed generated.py (default origin/main)
branch (gen/hw)the same for the working tree's generated.py
improvementper-shape % change of the branch's GENERATED-kernel throughput vs --rev's (positive = branch faster)

Semantics:

  • bench_spec.py files are discovered RECURSIVELY under examples/; top-level case*/ dirs without any spec are listed as (no bench_spec), never silently dropped. All of case1–case9 currently have specs.
  • Compare fixture identity using generated.py, not schedule_graph.json: JSON op ids are pointer-derived and unstable across regenerations, while byte-identical generated source means the kernels are identical. When the revision's and working tree's generated.py are byte-identical, benchmark the revision's gen/hw result once for the left column, skip a duplicate working-tree benchmark, show unchanged in the branch column, and show - for improvement.
  • Cases without a wired handwritten baseline (currently case8, which has no handwritten.py) show raw generated TFLOPS instead of a gen/hw ratio; the improvement column still works. case9_scaled_mm/blockwise wires hw_call to handwritten.blackwell_scaled_mm_ws, so it reports gen/hw ratios.
  • Correctness (vs torch reference, and vs handwritten output where present) is checked before timing; any failing shape appends FAIL to the cell. A kernel that raises shows an (error: ...) cell instead of crashing the table.
  • Both columns run under the CURRENT build — compare tests committed KERNEL fixtures, not toolchains. To evaluate a C++ scheduler change you must rebuild first and regenerate fixtures (below).
Show full SKILL.md (190 more words)Show less

Deep-dive per-case scripts (outside the harness): case4 perf_generated.py (gen vs no-WS vs handwritten WS) and run_generated.py (all three gradients); case8 bench_general.py (all three outputs + pool-vs-sum A/B); any case's run_*.py runner for correctness-only (case8's is run_triple_gemm_nows.py).

Scheduler provenance

The corpus fixtures (schedule_graph.json and the committed generated.py) are produced by the Modulo Scheduling pass.

Regenerating fixtures

When Buck is available, regenerate fixtures with the Buck-built beta triton-opt and the applicable sched2tlx Buck runner, following the same build-path rule above. The commands below are only for the no-Buck venv fallback:

TRITON_MODULO_DUMP_SCHEDULE=<case>/schedule_graph.json \
  build/cmake.*/bin/triton-opt -allow-unregistered-dialect \
  --nvgpu-modulo-schedule <case>/<kernel>_pre_modulo.ttgir -o /dev/null
env -u LD_LIBRARY_PATH PYTHONPATH=third_party/tlx/tools/sched2tlx \
  $REPO/.venv/bin/python -m sched2tlx <case>/schedule_graph.json -o <case>/generated.py

JSON op ids are pointer-derived and never byte-stable — regen always churns schedule_graph.json; the meaningful diff and benchmark-identity signal is generated.py. Known: case3 may need TRITON_MODULO_SELECT_VARIANT=2; case2 fixtures are ancient (fresh dumps differ, pre-existing); case8's committed generated.py predates the emitter's multiphase support landing (regen produces a single-phase kernel — don't "refresh" it casually).

Benchmark methodology

  • Use triton.testing.do_bench for every timing run. Configure a nonzero warmup before measurement; measured iterations must clear L2 before each invocation. Do not add ad-hoc CUDA-event timing loops.
  • Check nvidia-smi first; if a run hangs for minutes, run third_party/tlx/killgpu.sh.
  • One bench at a time — timing runs must not share the GPU.

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/sched2tlx-perf-testing of facebookexperimental/triton.

Open the folder on GitHubat commit 6f3dd70

Compare with similar skills

Sched2tlx Perf Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sched2tlx Perf Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sched2tlx Perf Testing this skillfacebookexperimental/triton201—~1.9kAutomated safety check: PassMIT
File Organization And Structurekitchen-engineer42/pdf2skills134—~500Automated safety check: PassNone
Update Milvus SDK Docsmilvus-io/web-content138—~13kAutomated safety check: PassApache-2.0
Dependency Watchtelegramdesktop/tdesktop33k1 repos~2.2kAutomated safety check: PassGPL-3.0
Process Inboxtelegramdesktop/tdesktop33k1 repos~5.4kAutomated safety check: PassGPL-3.0
Feishu Docopenclaw/openclaw392k—~516Automated safety check: PassMIT

Similar skills

  • File Organization And Structure

    kitchen-engineer42/pdf2skills

    Organize routines within files using blank line separation, consider alphabetical ordering when appropriate, and follow C++ standard file structure.

    134 GitHub stars~500 tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Update Milvus SDK Docs

    milvus-io/web-content

    Update the Milvus SDK API reference documentation under APIReference/ in the web-content repository so it reflects a new SDK release, using the SDK repository's git tags as ground truth.

    138 GitHub stars~13k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Dependency Watch

    telegramdesktop/tdesktop

    Audit Telegram Desktop dependencies on freshly fetched origin/dev for releases and security fixes, including upstream lag and backport candidates in patched forks.

    33k GitHub starsUsed in 1 repo~2.2k tokens
    Productivity & AutomationAuto-check passed
  • Process Inbox

    telegramdesktop/tdesktop

    Process the local ignored ai-tdesktop inbox into durable, independently testable Telegram Desktop task records while task execution worktrees remain active.

    33k GitHub starsUsed in 1 repo~5.4k tokens
    Productivity & AutomationAuto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Works with

Questions about Sched2tlx Perf Testing

What does Sched2tlx Perf Testing do?

Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaledmm). Sched2tlx Perf Testing is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run the sched2tlx perf/correctness harness over the modulo-scheduling example corpus (case1-9: GEMM, persistent GEMM, FA fwd/bwd, addmm+bias, LayerNorm, wgrad+bias, multiphase GEMM, scaledmm).

When should I use Sched2tlx Perf Testing?

Sched2tlx Perf Testing fits situations like: the user asks to benchmark generated-vs-handwritten kernels; check corpus correctness; compare emitter revisions; regenerate schedulegraph.json fixtures.

How do I install Sched2tlx Perf Testing in Claude Code?

Run `npx skills add facebookexperimental/triton --skill sched2tlx-perf-testing -a claude-code`. Or copy the skill folder (.claude/skills/sched2tlx-perf-testing in facebookexperimental/triton) into .claude/skills/sched2tlx-perf-testing in your project. Claude Code loads it when a task matches its description.

How do I install Sched2tlx Perf Testing in Codex?

Run `npx skills add facebookexperimental/triton --skill sched2tlx-perf-testing -a codex`. Or copy the skill folder (.claude/skills/sched2tlx-perf-testing in facebookexperimental/triton) into .agents/skills/sched2tlx-perf-testing in your project. Codex loads it when a task matches its description.

Can I use Sched2tlx Perf Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill sched2tlx-perf-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sched2tlx-perf-testing, .gemini/skills/sched2tlx-perf-testing, .github/skills/sched2tlx-perf-testing and .opencode/skills/sched2tlx-perf-testing in your project.

What does Sched2tlx Perf Testing need to run?

Going by SKILL.md and its folder, Sched2tlx Perf Testing needs the command-line tools its instructions call (make, uv and python). Our summary lists: Python 3.

Does Sched2tlx Perf Testing access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sched2tlx Perf Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sched2tlx Perf Testing use?

Sched2tlx Perf Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sched2tlx Perf Testing use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sched2tlx Perf Testing?

Skills that share tags, products or a category with Sched2tlx Perf Testing: File Organization And Structure (kitchen-engineer42/pdf2skills, 134 stars), Update Milvus SDK Docs (milvus-io/web-content, 138 stars), Dependency Watch (telegramdesktop/tdesktop, 33k stars) and Process Inbox (telegramdesktop/tdesktop, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sched2tlx Perf Testing?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 10, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.