Official agent skill

Whole Model Triton Kernel Local Numerics Drift Exploration

by facebookexperimental in facebookexperimental/triton

Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting.

OfficialMITAuto-check passedAI & LLM Engineering

Install Whole Model Triton Kernel Local Numerics Drift Exploration

skills CLI
$ npx skills add facebookexperimental/triton --skill whole-model-triton-kernel-local-numerics-drift-exploration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton whole-model-triton-kernel-local-numerics-drift-exploration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration .claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
whole-model-triton-kernel-local-numerics-drift-exploration
GitHub stars
201
Token cost
~792 tokens
SKILL.md length
431 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting.

  • Works in 6 steps: Request model context → Determine the kernel suite → Specify the A/B → …
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers 1. Request model context, 2. Determine the kernel suite, 3. Specify the A/B and 4. Run each kernel in isolation, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Whole Model Triton Kernel Local Numerics Drift Exploration is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting.

Its SKILL.md is about 790 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/whole-model-triton-kernel-local-numerics-drift-exploration”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Request model context
  2. Determine the kernel suite
  3. Specify the A/B
  4. Run each kernel in isolation
  5. Test every autotuning config
  6. Group results

What it can do on your machine

Read from SKILL.md and the folder at commit 6f3dd70. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Whole Model Triton Kernel Local Numerics Drift Exploration loads about 792 tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~792

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 6f3dd70, republished under its MIT licence (© facebookexperimental). 431 words, ~792 tokens.

Download SKILL.mdSave it as .claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration/SKILL.md (or your agent's skills folder).
name
whole-model-triton-kernel-local-numerics-drift-exploration
description
Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting.

Whole Model Triton Kernel Local Numerics Drift Exploration

Use this skill when asked to evaluate numerics risk of a Triton, TLX, or inductor change for a specific model.

1. Request model context

Ask for a doc or pointer that explains how to find information about the model, then collect:

  • Model name, entry point, and repro command.
  • Precision, dtypes, and representative shapes.
  • How the model is compiled (inductor settings, Triton/TLX path, custom kernels).
  • Where generated code or compile traces live.

Do not proceed on guesses. If the doc is missing fields above, ask follow-up questions.

2. Determine the kernel suite

From the model context, enumerate every Triton, TLX, or inductor-generated kernel the model produces:

  • Inspect inductor output code, compile logs, and autotune artifacts.
  • Record kernel name, source (Triton/TLX/inductor), signature, shapes, dtypes, and launch parameters.
  • This list is the test suite. Keep it explicit and reviewable.

3. Specify the A/B

Define baseline (A) and candidate (B) as Triton version plus commit hash:

  • Record version and commit hash for each side.
  • When swapping Triton versions, identify the commit hash from the version's build metadata, pinned dependency, or source checkout and record how it was determined.
  • State build or environment differences beyond the compiler change, if any.

4. Run each kernel in isolation

For each kernel in the suite:

  • Build an isolation harness using model-realistic shapes, dtypes, strides, and value ranges from step 1.
  • Reuse existing kernel-level accuracy tests where they exist.
  • Run A and B on identical inputs and compare outputs.
  • If a kernel has no accuracy test or the existing tests miss shapes, dtypes, or boundary conditions the model uses, add the missing test coverage before concluding.
Show full SKILL.md (156 more words)Show less

5. Test every autotuning config

Do not test only the autotuner winner:

  • Enumerate all autotuning configs for each kernel under both A and B.
  • Run each config individually in isolation.
  • Record per-config outputs and diffs so a config-specific regression is not hidden by winner selection.

6. Group results

Group every kernel into exactly one bucket:

  1. Bitwise equivalent: A and B outputs match exactly across all tested configs.
  2. Small numerics change, low probability of impacting model training: differences are within floating-point reassociation or rounding noise, small in magnitude and frequency, with no NaN/Inf divergence or distribution shift.
  3. Large numerics change, possible compiler bug or training risk: large max/mean error, systematic bias, NaN/Inf mismatch, shape- or config-dependent blowups, or any pattern inconsistent with benign reassociation.

Report per-kernel evidence: comparison metric, worst config, failing inputs, and bucket rationale. Flag bucket 3 kernels with repro commands and suspected scope (single kernel, autotune config, dtype, or shape class).

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration of facebookexperimental/triton.

Open the folder on GitHubat commit 6f3dd70

Compare with similar skills

Whole Model Triton Kernel Local Numerics Drift Exploration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Whole Model Triton Kernel Local Numerics Drift Exploration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Whole Model Triton Kernel Local Numerics Drift Exploration this skillfacebookexperimental/triton201—~792Automated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Liger Kernel Perflinkedin/Liger-Kernel6.7k—~1.5kAutomated safety check: PassBSD-2-Clause
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Perf

    linkedin/Liger-Kernel

    Optimizes the performance of existing Liger Kernel Triton kernels.

    6.7k GitHub stars~1.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Questions about Whole Model Triton Kernel Local Numerics Drift Exploration

What does Whole Model Triton Kernel Local Numerics Drift Exploration do?

Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting. Whole Model Triton Kernel Local Numerics Drift Exploration is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting.

When should I use Whole Model Triton Kernel Local Numerics Drift Exploration?

Whole Model Triton Kernel Local Numerics Drift Exploration fits situations like: tasks that involve GPU and accelerator computing.

How do I install Whole Model Triton Kernel Local Numerics Drift Exploration in Claude Code?

Run `npx skills add facebookexperimental/triton --skill whole-model-triton-kernel-local-numerics-drift-exploration -a claude-code`. Or copy the skill folder (.claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration in facebookexperimental/triton) into .claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration in your project. Claude Code loads it when a task matches its description.

How do I install Whole Model Triton Kernel Local Numerics Drift Exploration in Codex?

Run `npx skills add facebookexperimental/triton --skill whole-model-triton-kernel-local-numerics-drift-exploration -a codex`. Or copy the skill folder (.claude/skills/whole-model-triton-kernel-local-numerics-drift-exploration in facebookexperimental/triton) into .agents/skills/whole-model-triton-kernel-local-numerics-drift-exploration in your project. Codex loads it when a task matches its description.

Can I use Whole Model Triton Kernel Local Numerics Drift Exploration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill whole-model-triton-kernel-local-numerics-drift-exploration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whole-model-triton-kernel-local-numerics-drift-exploration, .gemini/skills/whole-model-triton-kernel-local-numerics-drift-exploration, .github/skills/whole-model-triton-kernel-local-numerics-drift-exploration and .opencode/skills/whole-model-triton-kernel-local-numerics-drift-exploration in your project.

What does Whole Model Triton Kernel Local Numerics Drift Exploration need to run?

SKILL.md names no scripts, command-line tools or credentials: Whole Model Triton Kernel Local Numerics Drift Exploration is instructions for the agent only.

Does Whole Model Triton Kernel Local Numerics Drift Exploration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Whole Model Triton Kernel Local Numerics Drift Exploration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Whole Model Triton Kernel Local Numerics Drift Exploration use?

Whole Model Triton Kernel Local Numerics Drift Exploration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Whole Model Triton Kernel Local Numerics Drift Exploration use?

About 792 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Whole Model Triton Kernel Local Numerics Drift Exploration?

Skills that share tags, products or a category with Whole Model Triton Kernel Local Numerics Drift Exploration: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Liger Kernel Perf (linkedin/Liger-Kernel, 6.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Whole Model Triton Kernel Local Numerics Drift Exploration?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 10, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.