Triton Kernel
vipshop/cache-dit
Write optimized Triton GPU kernels for deep learning operations.
Official agent skill
Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .claude/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .claude/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .agents/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .agents/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .agents/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .cursor/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .cursor/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/facebookexperimental/triton.git --path third_party/tlx/.claude/skills/tlx-kernel-optimization-agent--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .gemini/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .gemini/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .github/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .github/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .github/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .opencode/skills/tlx-kernel-optimization-agent && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tlx-kernel-optimization-agent" agent skill from https://github.com/facebookexperimental/triton/tree/main/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent into .opencode/skills/tlx-kernel-optimization-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-kernel-optimization-agent", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tlx-kernel-optimization-agentExecute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.
Tlx Kernel Optimization Agent is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Use this skill whenever the user says "use the TLX agent", "use the kernel optimization agent", "用 TLX agent 优化", or asks Carl/Claude to optimize a kernel with the repository agent. The required outcome is an actual agent CLI invocation and its measured result, not a walkthrough or reimplementation.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/input-contract.md`).
The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 612bd83. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tlx Kernel Optimization Agent loads about 3.5k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,546 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from facebookexperimental/triton at commit 612bd83, republished under its MIT licence (© facebookexperimental). 1,546 words, ~3,465 tokens.
.claude/skills/tlx-kernel-optimization-agent/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.The executable is third_party/tlx/tools/agents/kernel_optimization/cli.py.
Follow the layers below in order. Target-specific profiling rules supplement,
but never replace, the generic workflow.
--no-commit-winner only for an explicitly requested artifact-only run.
Submission remains a separate explicit action.--prior-run <output-dir>.
This imports prior evidence and source hashes for cross-run deduplication but
never adopts the old winner or replaces validation of the current kernel.--task authoring --op <op> --arch <arch> --suite <suite> to author
against the production kernel, harness, and cases. This does not tune the
heuristic policy.--tune-after-authoring only when source changes may alter the winning
configurations across the production suite and the user wants both steps.
The flag is opt-in and tuning runs only after authoring succeeds.--task tuning --op <op> --arch <arch> --suite <suite> when the kernel
implementation is already ready and only the heuristic policy needs work.Reading agent implementation is allowed only after the CLI reports an internal failure that requires diagnosis. The first attempt must use the public contract.
The CLI needs:
kernel.py complete source file
bundle/harness.py build, verify, benchmark, optional profile
bundle/cases.json workloads, weights, protected cases
bundle/target.json backend, architecture, device, environment
output/ fresh artifact directoryOptional inputs are reference_kernel.py and budget.json. The higher-level
coding agent owns target-bundle preparation; the TLX Agent consumes the bundle
as a frozen trust boundary and must never generate or modify it during the
optimization loop.
Before invoking the CLI, the higher-level agent must:
cases.json, hardware and environment in
target.json, and kernel-specific invariants, known failed experiments, and
evidence-to-knob guidance in target.json as optimization_guidance.Read references/input-contract.md for the complete construction and
validation checklist. Do not start the optimization loop until the bundle is
validated.
Candidate generation automatically receives trusted source-optimization skills
owned by third_party/tlx/tools/agents/kernel_optimization/skills/: every target
receives layout-conversion efficiency guidance, while CUDA/NVIDIA targets also
receive async TMA output publication and warp-barrier efficiency guidance. Known
Hopper and Blackwell targets additionally receive NVIDIA persistent pipeline
efficiency guidance, and Blackwell targets also receive persistent CLC scheduling
guidance. Unknown NVIDIA architectures receive the common NVIDIA guidance but
require an explicit allowlist update before receiving persistent pipeline guidance.
Canonical profiling workflow documentation lives under
third_party/tlx/tools/agents/kernel_optimization/docs/profiling/ and is not
injected as source guidance. Keep workload-specific
invariants and exclusions in target.json.optimization_guidance; they are applied
after the built-in target skills.
For every candidate, the standard loop is:
build -> verify -> benchmark -> profile -> decide -> repeatThe live log must include:
Do not print kernel source or private chain-of-thought. Continue through the configured round budget after an unpromoted round. Reject incorrect, duplicate, unstable, and materially regressing candidates.
The optimizer sends a profile request for the baseline (deep), every
correctness-passing candidate (summary, escalated to deep near the promotion
threshold), and the finalist (deep). Profiling is always on; the --profile
flag is vestigial. Every request carries tools=["proton_launch", "native_profiler"], where native_profiler resolves to ncu on CUDA/NVIDIA
targets automatically. Pass --diagnostic-proton-intra-kernel to additionally
collect warp-granularity Proton instrumentation traces for the baseline and
final winner only (diagnostic-only: never benchmark, promote, or commit
instrumented source or timing).
These requests produce data only if the bundle's harness.py::profile()
implements them. A stub that repackages endpoint timings yields
ncu=unavailable and zeroed proton.* fields on every line, and all
hypotheses degrade to endpoint latency plus source inspection. Before
launching, verify the harness profile path end to end:
third_party/tlx/tools/agents/kernel_optimization/docs/profiling/proton.md
and return profile()["proton"] with nonzero main_kernel_us for the
expected kernel launch.third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md
(CUDA/NVIDIA) and return profile()["ncu"] with non-null summary duration.
Report unsupported counters as JSON null with a diagnostic, never as zero.Harnesses that implement profile should accept a structured request with:
{
"level": "summary",
"tools": ["proton_launch", "native_profiler"],
"experiment_id": "stable candidate or baseline id",
"artifacts_dir": "/absolute/path/to/profile-artifacts",
"reason": "why this profile was requested",
"diagnostic_only": false
}The legacy two-argument profile(build_artifact, case) contract remains
supported and means summary profiling for the default tools. Return compact,
normalized JSON inline. Store raw .hatchet, .chrome_trace, .ncu-rep, CSV,
mapping, and command files as artifacts, and reference them with absolute paths.
Profiling has three distinct layers:
third_party/tlx/tools/agents/kernel_optimization/docs/profiling/proton.md and
collect for every correctness-passing candidate.native_profiler; each target harness
maps it to its platform tool. For CUDA/NVIDIA, this is NCU and the harness
should follow
third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md.Use Proton to attribute wrapper, benchmark phase, launch, and profiler overhead.
An ordinary Proton timeline with hook='triton' that shows one kernel launch
does not prove tlx.async_task overlap; it only shows launch/wrapper
attribution around the compiled kernel. Do not infer per-task overlap from that
timeline.
Diagnostic intra-kernel attribution must use Proton instrumentation mode with
backend='instrumentation', data='trace', granularity='warp', and explicit
Triton semantic enabled. Group warp lanes by known async-task warp ranges from
the source, generated metadata, or a saved mapping artifact. Do not request
warp_group granularity because the runtime rejects it today.
Instrumentation changes are diagnostic-only. Compiler transforms may move or merge scopes, so instrumented source and timing must never be benchmarked, promoted, committed, or used as speedup evidence.
Select target profiling guidance from target.json:
third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md.profile() contract without inventing NVIDIA requirements.Every correctness-passing candidate receives Proton attribution and the target guide's summary profile. Escalate to the target guide's deep profile for:
Do not spend deep-profile time on incorrect or clearly slow candidates.
Unsupported counters must be reported as unavailable with JSON null and a
diagnostic, never as zero.
Each candidate hypothesis must cite measured evidence and change one subsystem or tightly coupled invariant-preserving pair. Feed failed hypotheses and their metric regressions into subsequent prompts as exclusions.
Generic interpretation rules:
Use the target guide for vendor-specific counter interpretation.
Promote only when:
Revalidate the finalist with benchmark, Proton attribution, summary target
profile, and required deep target profile. Only then may the default commit
occur. Auto-commit must detect Git or Mercurial from the kernel path, preserve
unrelated dirty/staged work, include TLX agent authored in the commit body,
and log VCS, revision, repository, target, subject, and failure diagnostics. If
commit fails, keep all artifacts and return the distinct commit-failure status.
From the repository root:
PYTHONPATH=<repo-root> python -m third_party.tlx.tools.agents.kernel_optimization.cli \
--kernel <absolute-kernel.py> \
--harness <absolute-harness.py> \
--cases <absolute-cases.json> \
--target <absolute-target.json> \
--output-dir <absolute-output-dir> \
--prior-run <optional-previous-output-dir> \
--provider codex \
--max-rounds 5 \
--candidates-per-round 2 \
--max-candidate-seconds 600 \
--max-total-seconds 3600 \
--min-speedup 1.01 \
--max-cv 0.10 \
--benchmark-repetitions 10 \
--profileAdd --reference-kernel, --budget, --arch, --model, --vcs, or
--commit-message only when required. Do not invent a model name. Never run
optimizer.py directly.
The task is complete only after the actual CLI run reaches final revalidation
or a diagnosed blocking failure. Report the exact workload, GPU, baseline and
final latency, speedup, CV, correctness, stopping reason, commit result, and
absolute output directory. best_kernel.py by itself is not completion.
© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in third_party/tlx/.claude/skills/tlx-kernel-optimization-agent of facebookexperimental/triton.
Open the folder on GitHubat commit 612bd83
Tlx Kernel Optimization Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tlx Kernel Optimization Agent this skillfacebookexperimental/triton | 201 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Triton Kernelvipshop/cache-dit | 1.3k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Executealirezarezvani/claude-skills | 28k | — | ~831 | Automated safety check: Pass | MIT | |
| Debugging Executionsn8n-io/n8n | 207k | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Kernel Organizationsgl-project/sglang | 37k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence |
vipshop/cache-dit
Write optimized Triton GPU kernels for deep learning operations.
alirezarezvani/claude-skills
/cs:execute <decision — Generate a 90-day execution plan with weekly milestones, DRIs, and check-in cadence from an approved decision.
n8n-io/n8n
Debug failed or wrong-output workflow executions using executions tools.
sgl-project/sglang
Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
code-yeongyu/oh-my-openagent
Executes a written ulw-plan work plan with Boulder state, evidence ledger, worktree discipline, and parallel subagents.
facebookexperimental/triton
Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.
facebookexperimental/triton
Design and run Triton TTGIR debugging ablations using iroverride.
facebookexperimental/triton
Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.
facebookexperimental/triton
Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.
facebookexperimental/triton
Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).
facebookexperimental/triton
Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.
Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Tlx Kernel Optimization Agent is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.
Tlx Kernel Optimization Agent fits situations like: the user says use the TLX agent; use the kernel optimization agent; asks Carl/Claude to optimize a kernel with the repository agent.
Run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a claude-code`. Or copy the skill folder (third_party/tlx/.claude/skills/tlx-kernel-optimization-agent in facebookexperimental/triton) into .claude/skills/tlx-kernel-optimization-agent in your project. Claude Code loads it when a task matches its description.
Run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a codex`. Or copy the skill folder (third_party/tlx/.claude/skills/tlx-kernel-optimization-agent in facebookexperimental/triton) into .agents/skills/tlx-kernel-optimization-agent in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tlx-kernel-optimization-agent, .gemini/skills/tlx-kernel-optimization-agent, .github/skills/tlx-kernel-optimization-agent and .opencode/skills/tlx-kernel-optimization-agent in your project.
Going by SKILL.md and its folder, Tlx Kernel Optimization Agent needs the command-line tools its instructions call (python).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tlx Kernel Optimization Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tlx Kernel Optimization Agent: Triton Kernel (vipshop/cache-dit, 1.3k stars), Execute (alirezarezvani/claude-skills, 28k stars), Debugging Executions (n8n-io/n8n, 207k stars) and Kernel Organization (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 8, 2026.
Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.