Official agent skill

Tlx Kernel Optimization Agent

by facebookexperimental in facebookexperimental/triton

Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

OfficialMITAuto-check passed

Install Tlx Kernel Optimization Agent

skills CLI
$ npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton tlx-kernel-optimization-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/third_party/tlx/.claude/skills/tlx-kernel-optimization-agent .claude/skills/tlx-kernel-optimization-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tlx-kernel-optimization-agent
GitHub stars
201
Token cost
~3.5k tokens
SKILL.md length
1,546 words
Files
2 (incl. references)
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

  • Works in 8 steps: Invoke the public CLI. Do not manually… → Resolve the repository root, absolute… → Keep candidate generation isolated from… → …
  • The user says use the TLX agent
  • SKILL.md covers Layer 0: Invocation And Safety, Layer 1: Inputs And Standard…, Layer 2: Profiling Request And… and Layer 3: Proton Attribution, plus 4 more sections
  • Calls python

What it does

Tlx Kernel Optimization Agent is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Use this skill whenever the user says "use the TLX agent", "use the kernel optimization agent", "用 TLX agent 优化", or asks Carl/Claude to optimize a kernel with the repository agent. The required outcome is an actual agent CLI invocation and its measured result, not a walkthrough or reimplementation.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/input-contract.md`).

The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • The user says use the TLX agent
  • Use the kernel optimization agent
  • Asks Carl/Claude to optimize a kernel with the repository agent

Example prompts

  • “use the TLX agent”
  • “use the kernel optimization agent”
  • “用 TLX agent 优化”
  • “/tlx-kernel-optimization-agent”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Invoke the public CLI. Do not manually optimize the kernel or substitute a
  2. Resolve the repository root, absolute kernel path, and target bundle. Record
  3. Keep candidate generation isolated from the live checkout. Candidate source
  4. Run stdout and stderr separately. Tee stderr to a stable absolute live log;
  5. The CLI commits a revalidated winner by default. Use
  6. If the live kernel changes concurrently, do not overwrite it. Preserve the
  7. When continuing a completed optimization, pass --prior-run .
  8. Keep production kernel authoring and heuristic tuning independent

What it can do on your machine

Read from SKILL.md and the folder at commit 612bd83. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tlx Kernel Optimization Agent loads about 3.5k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 612bd83, republished under its MIT licence (© facebookexperimental). 1,546 words, ~3,465 tokens.

Download SKILL.mdSave it as .claude/skills/tlx-kernel-optimization-agent/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
tlx-kernel-optimization-agent
description
Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Use this skill whenever the user says "use the TLX agent", "use the kernel optimization agent", "用 TLX agent 优化", or asks Carl/Claude to optimize a kernel with the repository agent. The required outcome is an actual agent CLI invocation and its measured result, not a walkthrough or reimplementation.

Run The TLX Kernel Optimization Agent

The executable is third_party/tlx/tools/agents/kernel_optimization/cli.py. Follow the layers below in order. Target-specific profiling rules supplement, but never replace, the generic workflow.

Layer 0: Invocation And Safety

  1. Invoke the public CLI. Do not manually optimize the kernel or substitute a generic subagent before the first CLI attempt.
  2. Resolve the repository root, absolute kernel path, and target bundle. Record initial source-control status and treat existing bytes as the user baseline. Never clean or revert a dirty worktree.
  3. Keep candidate generation isolated from the live checkout. Candidate source may exist only in the provider's temporary workspace and output artifacts until final promotion.
  4. Run stdout and stderr separately. Tee stderr to a stable absolute live log; write stdout JSON to a separate artifact. When starting a run, respond only with the absolute log path unless the user asks for more.
  5. The CLI commits a revalidated winner by default. Use --no-commit-winner only for an explicitly requested artifact-only run. Submission remains a separate explicit action.
  6. If the live kernel changes concurrently, do not overwrite it. Preserve the winner artifact and report the conflict.
  7. When continuing a completed optimization, pass --prior-run <output-dir>. This imports prior evidence and source hashes for cross-run deduplication but never adopts the old winner or replaces validation of the current kernel.
  8. Keep production kernel authoring and heuristic tuning independent:
    • Use --task authoring --op <op> --arch <arch> --suite <suite> to author against the production kernel, harness, and cases. This does not tune the heuristic policy.
    • Add --tune-after-authoring only when source changes may alter the winning configurations across the production suite and the user wants both steps. The flag is opt-in and tuning runs only after authoring succeeds.
    • Use --task tuning --op <op> --arch <arch> --suite <suite> when the kernel implementation is already ready and only the heuristic policy needs work.

Reading agent implementation is allowed only after the CLI reports an internal failure that requires diagnosis. The first attempt must use the public contract.

Layer 1: Inputs And Standard Loop

The CLI needs:

text
kernel.py                  complete source file
bundle/harness.py          build, verify, benchmark, optional profile
bundle/cases.json          workloads, weights, protected cases
bundle/target.json         backend, architecture, device, environment
output/                    fresh artifact directory

Optional inputs are reference_kernel.py and budget.json. The higher-level coding agent owns target-bundle preparation; the TLX Agent consumes the bundle as a frozen trust boundary and must never generate or modify it during the optimization loop.

Before invoking the CLI, the higher-level agent must:

  1. Search for an existing bundle that exactly matches the kernel entry point, workload, mode, backend, and architecture. Do not silently reuse a nearby shape or provider.
  2. If no exact bundle exists, create a run-specific bundle from the nearest authoritative correctness test and production benchmark. Keep it outside the candidate workspace and do not modify the user's kernel permanently.
  3. Encode all requested workloads in cases.json, hardware and environment in target.json, and kernel-specific invariants, known failed experiments, and evidence-to-knob guidance in target.json as optimization_guidance.
  4. Run the bundle against the untouched kernel before launching the Agent. Confirm build succeeds, every protected case passes, repeated benchmark samples are stable, and summary profiling selects the intended kernel.
  5. Negative-test the trust boundary with an intentionally invalid or incorrect temporary candidate and confirm build or verification rejects it.
  6. Freeze the validated bundle for the duration of the run. Record its absolute paths and content hashes in the run log, and pass all paths explicitly to the public CLI.

Read references/input-contract.md for the complete construction and validation checklist. Do not start the optimization loop until the bundle is validated.

Candidate generation automatically receives trusted source-optimization skills owned by third_party/tlx/tools/agents/kernel_optimization/skills/: every target receives layout-conversion efficiency guidance, while CUDA/NVIDIA targets also receive async TMA output publication and warp-barrier efficiency guidance. Known Hopper and Blackwell targets additionally receive NVIDIA persistent pipeline efficiency guidance, and Blackwell targets also receive persistent CLC scheduling guidance. Unknown NVIDIA architectures receive the common NVIDIA guidance but require an explicit allowlist update before receiving persistent pipeline guidance. Canonical profiling workflow documentation lives under third_party/tlx/tools/agents/kernel_optimization/docs/profiling/ and is not injected as source guidance. Keep workload-specific invariants and exclusions in target.json.optimization_guidance; they are applied after the built-in target skills.

For every candidate, the standard loop is:

text
build -> verify -> benchmark -> profile -> decide -> repeat

The live log must include:

  • hypothesis and evidence;
  • one coherent source/configuration change;
  • expected effect and risk;
  • correctness, median, p95, CV, and speedup;
  • Proton attribution plus target summary profile deltas and a concise decision.

Do not print kernel source or private chain-of-thought. Continue through the configured round budget after an unpromoted round. Reject incorrect, duplicate, unstable, and materially regressing candidates.

Layer 2: Profiling Request And Artifacts

The optimizer sends a profile request for the baseline (deep), every correctness-passing candidate (summary, escalated to deep near the promotion threshold), and the finalist (deep). Profiling is always on; the --profile flag is vestigial. Every request carries tools=["proton_launch", "native_profiler"], where native_profiler resolves to ncu on CUDA/NVIDIA targets automatically. Pass --diagnostic-proton-intra-kernel to additionally collect warp-granularity Proton instrumentation traces for the baseline and final winner only (diagnostic-only: never benchmark, promote, or commit instrumented source or timing).

These requests produce data only if the bundle's harness.py::profile() implements them. A stub that repackages endpoint timings yields ncu=unavailable and zeroed proton.* fields on every line, and all hypotheses degrade to endpoint latency plus source inspection. Before launching, verify the harness profile path end to end:

  • Proton launch attribution: follow third_party/tlx/tools/agents/kernel_optimization/docs/profiling/proton.md and return profile()["proton"] with nonzero main_kernel_us for the expected kernel launch.
  • Target counters: follow third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md (CUDA/NVIDIA) and return profile()["ncu"] with non-null summary duration. Report unsupported counters as JSON null with a diagnostic, never as zero.

Harnesses that implement profile should accept a structured request with:

json
{
  "level": "summary",
  "tools": ["proton_launch", "native_profiler"],
  "experiment_id": "stable candidate or baseline id",
  "artifacts_dir": "/absolute/path/to/profile-artifacts",
  "reason": "why this profile was requested",
  "diagnostic_only": false
}

The legacy two-argument profile(build_artifact, case) contract remains supported and means summary profiling for the default tools. Return compact, normalized JSON inline. Store raw .hatchet, .chrome_trace, .ncu-rep, CSV, mapping, and command files as artifacts, and reference them with absolute paths.

Profiling has three distinct layers:

  • Proton wrapper/launch attribution: read third_party/tlx/tools/agents/kernel_optimization/docs/profiling/proton.md and collect for every correctness-passing candidate.
  • Target summary/deep profiling: request native_profiler; each target harness maps it to its platform tool. For CUDA/NVIDIA, this is NCU and the harness should follow third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md.
  • Diagnostic-only Proton intra-kernel instrumentation: use only for attribution questions that cannot be answered from wrapper timelines or target counters.
Show full SKILL.md (535 more words)Show less

Layer 3: Proton Attribution

Use Proton to attribute wrapper, benchmark phase, launch, and profiler overhead. An ordinary Proton timeline with hook='triton' that shows one kernel launch does not prove tlx.async_task overlap; it only shows launch/wrapper attribution around the compiled kernel. Do not infer per-task overlap from that timeline.

Diagnostic intra-kernel attribution must use Proton instrumentation mode with backend='instrumentation', data='trace', granularity='warp', and explicit Triton semantic enabled. Group warp lanes by known async-task warp ranges from the source, generated metadata, or a saved mapping artifact. Do not request warp_group granularity because the runtime rejects it today.

Instrumentation changes are diagnostic-only. Compiler transforms may move or merge scopes, so instrumented source and timing must never be benchmarked, promoted, committed, or used as speedup evidence.

Layer 4: Target Profiling

Select target profiling guidance from target.json:

  • CUDA/NVIDIA: read third_party/tlx/tools/agents/kernel_optimization/docs/profiling/nvidia-ncu.md.
  • Other backends: use a sibling target guide when present; otherwise use the vendor-neutral profile() contract without inventing NVIDIA requirements.

Every correctness-passing candidate receives Proton attribution and the target guide's summary profile. Escalate to the target guide's deep profile for:

  • the baseline before candidate generation;
  • a candidate within one percentage point of the promotion threshold;
  • disagreement between endpoint benchmark, Proton attribution, and target summary profile;
  • repeated rounds with no promoted candidate;
  • schedule or WS hypotheses that need stall, occupancy, spill, or memory-hierarchy evidence;
  • the final promoted candidate before commit.

Do not spend deep-profile time on incorrect or clearly slow candidates. Unsupported counters must be reported as unavailable with JSON null and a diagnostic, never as zero.

Layer 5: Evidence-To-Action Policy

Each candidate hypothesis must cite measured evidence and change one subsystem or tightly coupled invariant-preserving pair. Feed failed hypotheses and their metric regressions into subsequent prompts as exclusions.

Generic interpretation rules:

  • Higher kernel duration with lower compute and memory utilization indicates scheduling, serialization, synchronization, or insufficient parallelism.
  • A benchmark-only change inside the noise floor with unchanged kernel metrics is not optimization evidence.
  • Lower utilization without lower work, traffic, and duration is not a win.
  • Register-budget or buffering changes require spill/occupancy evidence and a producer-consumer/barrier proof.
  • Main-kernel profile time and public-wrapper benchmark time cover different scopes; do not subtract them to manufacture a bottleneck.

Use the target guide for vendor-specific counter interpretation.

Layer 6: Promotion, Revalidation, And Commit

Promote only when:

  • every protected case passes;
  • weighted speedup meets the configured threshold;
  • measurement variance is within budget;
  • target profiling shows no material main-kernel regression.

Revalidate the finalist with benchmark, Proton attribution, summary target profile, and required deep target profile. Only then may the default commit occur. Auto-commit must detect Git or Mercurial from the kernel path, preserve unrelated dirty/staged work, include TLX agent authored in the commit body, and log VCS, revision, repository, target, subject, and failure diagnostics. If commit fails, keep all artifacts and return the distinct commit-failure status.

Layer 7: Command And Completion

From the repository root:

bash
PYTHONPATH=<repo-root> python -m third_party.tlx.tools.agents.kernel_optimization.cli \
  --kernel <absolute-kernel.py> \
  --harness <absolute-harness.py> \
  --cases <absolute-cases.json> \
  --target <absolute-target.json> \
  --output-dir <absolute-output-dir> \
  --prior-run <optional-previous-output-dir> \
  --provider codex \
  --max-rounds 5 \
  --candidates-per-round 2 \
  --max-candidate-seconds 600 \
  --max-total-seconds 3600 \
  --min-speedup 1.01 \
  --max-cv 0.10 \
  --benchmark-repetitions 10 \
  --profile

Add --reference-kernel, --budget, --arch, --model, --vcs, or --commit-message only when required. Do not invent a model name. Never run optimizer.py directly.

The task is complete only after the actual CLI run reaches final revalidation or a diagnosed blocking failure. Report the exact workload, GPU, baseline and final latency, speedup, CV, correctness, stopping reason, commit result, and absolute output directory. best_kernel.py by itself is not completion.

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in third_party/tlx/.claude/skills/tlx-kernel-optimization-agent of facebookexperimental/triton.

  • SKILL.md
  • references/input-contract.md

Open the folder on GitHubat commit 612bd83

Compare with similar skills

Tlx Kernel Optimization Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tlx Kernel Optimization Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tlx Kernel Optimization Agent this skillfacebookexperimental/triton201—~3.5kAutomated safety check: PassMIT
Triton Kernelvipshop/cache-dit1.3k—~1.1kAutomated safety check: PassApache-2.0
Executealirezarezvani/claude-skills28k—~831Automated safety check: PassMIT
Debugging Executionsn8n-io/n8n207k—~2.6kAutomated safety check: PassCustom licence
Kernel Organizationsgl-project/sglang37k—~1.3kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence

Similar skills

  • Triton Kernel

    vipshop/cache-dit

    Write optimized Triton GPU kernels for deep learning operations.

    1.3k GitHub stars~1.1k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Execute

    alirezarezvani/claude-skills

    /cs:execute <decision — Generate a 90-day execution plan with weekly milestones, DRIs, and check-in cadence from an approved decision.

    28k GitHub stars~831 tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Official

    Debug failed or wrong-output workflow executions using executions tools.

    207k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Kernel Organization

    sgl-project/sglang

    Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations.

    37k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ulw Execute

    code-yeongyu/oh-my-openagent

    Executes a written ulw-plan work plan with Boulder state, evidence ledger, worktree discipline, and parallel subagents.

    70k GitHub stars~6.3k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed
  • Kernel Perf Testing

    facebookexperimental/triton

    Official

    Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

    201 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Tlx Kernel Optimization Agent

What does Tlx Kernel Optimization Agent do?

Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel. Tlx Kernel Optimization Agent is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

When should I use Tlx Kernel Optimization Agent?

Tlx Kernel Optimization Agent fits situations like: the user says use the TLX agent; use the kernel optimization agent; asks Carl/Claude to optimize a kernel with the repository agent.

How do I install Tlx Kernel Optimization Agent in Claude Code?

Run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a claude-code`. Or copy the skill folder (third_party/tlx/.claude/skills/tlx-kernel-optimization-agent in facebookexperimental/triton) into .claude/skills/tlx-kernel-optimization-agent in your project. Claude Code loads it when a task matches its description.

How do I install Tlx Kernel Optimization Agent in Codex?

Run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a codex`. Or copy the skill folder (third_party/tlx/.claude/skills/tlx-kernel-optimization-agent in facebookexperimental/triton) into .agents/skills/tlx-kernel-optimization-agent in your project. Codex loads it when a task matches its description.

Can I use Tlx Kernel Optimization Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tlx-kernel-optimization-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tlx-kernel-optimization-agent, .gemini/skills/tlx-kernel-optimization-agent, .github/skills/tlx-kernel-optimization-agent and .opencode/skills/tlx-kernel-optimization-agent in your project.

What does Tlx Kernel Optimization Agent need to run?

Going by SKILL.md and its folder, Tlx Kernel Optimization Agent needs the command-line tools its instructions call (python).

Does Tlx Kernel Optimization Agent access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tlx Kernel Optimization Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tlx Kernel Optimization Agent use?

Tlx Kernel Optimization Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tlx Kernel Optimization Agent use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Tlx Kernel Optimization Agent?

Skills that share tags, products or a category with Tlx Kernel Optimization Agent: Triton Kernel (vipshop/cache-dit, 1.3k stars), Execute (alirezarezvani/claude-skills, 28k stars), Debugging Executions (n8n-io/n8n, 207k stars) and Kernel Organization (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tlx Kernel Optimization Agent?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 8, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.