Official agent skill

Tlx API Reference

by facebookexperimental in facebookexperimental/triton

TLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton.

OfficialMITAuto-check passed

Install Tlx API Reference

skills CLI
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton tlx-api-reference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tlx-api-reference .claude/skills/tlx-api-reference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tlx-api-reference
GitHub stars
201
Token cost
~3.9k tokens
SKILL.md length
1,573 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

TLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton.

  • Works in 7 steps: Identify the protected buffer and… → Enumerate every task replica, warp,… → Verify num_warps * 32 * num_arrivals… → …
  • Modifying TLX kernel code that uses barriers (mbarrier
  • SKILL.md covers Warp Specialization, Memory Barriers, Memory Operations and Matrix Multiply (MMA), plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tlx API Reference is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. TLX DSL API reference for low-level GPU primitives. Use when writing or modifying TLX kernel code that uses barriers (mbarrier, named barriers), memory allocation (localalloc, SMEM, TMEM), TMA operations, warp specialization (asynctasks, asynctask), CLC (cluster launch control), or wgmma instructions. Covers Hopper and Blackwell hardware differences.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • Modifying TLX kernel code that uses barriers (mbarrier
  • Named barriers)
  • Memory allocation (localalloc
  • Warp specialization (asynctasks

Example prompts

  • “/tlx-api-reference”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Identify the protected buffer and confirm the barrier is signaled by explicit
  2. Enumerate every task replica, warp, lane, predicate, and arrival call that targets
  3. Verify num_warps * 32 * num_arrivals equals the exact number of unit arrivals.
  4. Prove every counted lane reaches the arrival after its final access to the buffer,
  5. Preserve buffer indices, phase calculations, wait sites, and arrival sites; change
  6. Keep full/TMA barriers ordinary even when the paired empty/reuse barriers are
  7. Run correctness and benchmark both forms. Warp barriers trade additional

What it can do on your machine

Read from SKILL.md and the folder at commit 612bd83. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tlx API Reference loads about 3.9k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,573 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 612bd83, republished under its MIT licence (© facebookexperimental). 1,573 words, ~3,895 tokens.

Download SKILL.mdSave it as .claude/skills/tlx-api-reference/SKILL.md (or your agent's skills folder).
name
tlx-api-reference
description
TLX DSL API reference for low-level GPU primitives. Use when writing or modifying TLX kernel code that uses barriers (mbarrier, named barriers), memory allocation (local_alloc, SMEM, TMEM), TMA operations, warp specialization (async_tasks, async_task), CLC (cluster launch control), or wgmma instructions. Covers Hopper and Blackwell hardware differences.

TLX API Quick Reference

Warp Specialization

FunctionDescriptionArch
tlx.async_tasks()Context manager wrapping all async task regionsBoth
tlx.async_task([task_ids])Assign code to specific task IDs (e.g., [0] = producer, [1,2] = consumers)Both
tlx.async_task(num_warps=N, num_regs=R)Explicit warp/register allocation for a taskBoth
tlx.async_task("default", num_regs=R)Default task for code outside explicit tasksBoth
tlx.async_task_replica_id()Returns replica ID inside an async regionBoth
Warp specialization skeleton
python
with tlx.async_tasks():
    with tlx.async_task([0]):       # Producer
        # TMA loads
    with tlx.async_task([1, 2]):    # Consumers
        # MMA compute

Memory Barriers

mbarrier (shared-memory allocated)
FunctionDescriptionArch
tlx.alloc_barriers(num_barriers, arrive_count=1)Allocate ordinary SMEM barriers. Software arrivals use leader-based lowering, which may synchronize participating threads before the leader arrives.Both
tlx.alloc_warp_barrier(num_barriers, num_warps=1, num_arrivals=1)Allocate SMEM barriers whose software arrivals are performed independently by every participating thread. The initialized count is num_warps * 32 * num_arrivals.NVIDIA
tlx.barrier_expect_bytes(bar, bytes, pred=None)Set expected transaction byte count on barrierBoth
tlx.barrier_wait(bar, phase, pred=None)Wait until barrier phase flips (LOCAL mbarrier only)Both
tlx.barrier_arrive(bar, arrive_count=1, remote_cta_rank=None)Signal arrival at barrier. remote_cta_rank signals a barrier in a remote CTA — only valid when ctas_per_cga > 1, causes "Unexpected buffer remote view in 1cta mode" otherwise. Guard with if USE_2CTA: when kernel supports both modes.Both
tlx.cluster_barrier()Full cluster-wide synchronization barrierBoth

Ordinary barrier count rules:

  • alloc_barriers(..., arrive_count=N) initializes the number of logical arrivals required to complete a phase. Transaction bytes registered with barrier_expect_bytes are an additional completion condition; do not infer the software arrival topology from the transaction byte count.
  • For TMA full/data-ready barriers where one task registers the transaction, use arrive_count=1 unless the surrounding protocol explicitly requires additional software arrivals.
  • For local software barriers shared by replicated tlx.async_task consumers, the ordinary count is typically the number of consumer replicas that arrive once per phase.
  • For cross-CTA software arrivals using remote_cta_rank, derive the count from the exact CTA protocol. Do not assume the local-task rule applies.
Warp-barrier semantics and safe conversion

alloc_warp_barrier changes the physical arrival protocol, not merely the spelling of the allocation. With an ordinary barrier, TLX selects a leader to perform the arrival and may synchronize the participating threads first. With a warp barrier, every participating thread performs its own arrival, avoiding that leader-path synchronization but issuing more mbarrier arrival operations.

The initialized count is:

text
expected arrivals per phase = num_warps * 32 * num_arrivals

Here num_warps is the number of warps in each participating task replica, and num_arrivals is the number of such all-thread arrival events targeting the same barrier in one phase. For example, two replicated four-warp consumers that each release one shared buffer slot use num_warps=4, num_arrivals=2, for 256 arrivals. A consumer-owned slot released by one four-warp replica uses num_warps=4, num_arrivals=1, for 128 arrivals. Keep the corresponding barrier_arrive calls at their original final-use points and normally use the default unit arrival count.

Before converting an ordinary barrier, classify its role and arrival source:

Barrier roleWarp-barrier candidate?Rule
Local software empty/reuse notificationYes, after proving the topologyEvery expected lane must arrive exactly once per declared arrival event, after its final read of the protected storage.
Local software data-ready notificationSometimesSafe only when readiness is produced by the same statically known all-thread topology; benchmark because per-thread arrivals are not universally faster.
TMA transaction/full barrierNoKeep an ordinary barrier so transaction completion remains tracked through barrier_expect_bytes and the TMA operation.
MMA/tensor-core completion barrierNo automatic conversionPreserve the completion mechanism required by the MMA API.
Named scheduling barrierNoPreserve the named-barrier ID, participant count, direction, and phase protocol.
Remote, multicast, or cross-CTA barrierNo automatic conversionKeep the established protocol unless backend support and the complete cluster-wide arrival topology are explicitly proven.
Divergently predicated arrivalNoA missing lane leaves the phase incomplete; use a warp barrier only when every counted lane is guaranteed to execute the required arrivals.

Safe-conversion checklist:

  1. Identify the protected buffer and confirm the barrier is signaled by explicit software barrier_arrive calls rather than TMA, MMA, multicast, or remote completion.
  2. Enumerate every task replica, warp, lane, predicate, and arrival call that targets the barrier during one phase.
  3. Verify num_warps * 32 * num_arrivals equals the exact number of unit arrivals.
  4. Prove every counted lane reaches the arrival after its final access to the buffer, including prologue, tail, and persistent-loop iterations.
  5. Preserve buffer indices, phase calculations, wait sites, and arrival sites; change only the allocator for the first A/B experiment.
  6. Keep full/TMA barriers ordinary even when the paired empty/reuse barriers are converted.
  7. Run correctness and benchmark both forms. Warp barriers trade additional per-thread arrivals for removal of leader-path synchronization, so conversion is an optimization candidate, not a universal rule.
Named barriers (hardware-allocated, indices 0–15)
FunctionDescriptionArch
tlx.named_barrier_wait(bar_id, num_threads)Wait until the total participant count reaches bar_idNVIDIA
tlx.named_barrier_arrive(bar_id, num_threads)Signal arrival at bar_id using the total participant countNVIDIA

num_threads is the total number of threads required to flip the barrier phase: num_waiting_threads + num_arriving_threads. Wait and Arrive calls for the same barrier phase must use the same value. The count must be a multiple of 32 (warp size); it is typically num_warp_groups * warps_per_group * 32.

Used for PingPong scheduling to prevent tensor core contention between consumer warp groups.

Memory Operations

SMEM / TMEM allocation
FunctionDescriptionArch
tlx.local_alloc(shape, dtype, num, storage=smem, reuse=None, layout=None)Allocate buffered tensor in SMEM or TMEMBoth (TMEM: Blackwell)
tlx.storage_alias_spec(storage=smem, buffer_size_bytes=None)Define shared buffer region for multiple local_alloc calls via reuseBoth
tlx.local_view(buf, index)Get view of a single buffer from a multi-buffered tensorBoth
tlx.local_slice(buf, start, end)Slice a sub-range of a buffered tensorBoth
tlx.subslice(tensor, dim, start, size)Subslice a tensor along a dimensionBoth
tlx.local_load(buf)Load from SMEM/TMEM buffer into registersBoth
tlx.local_store(val, buf)Store from registers into SMEM/TMEM bufferBoth
tlx.local_trans(buf)Transpose a shared memory bufferBoth
tlx.local_reinterpret(buf, dtype)Reinterpret buffer with a different dtypeBoth
tlx.remote_view(buf, remote_cta_rank)Get view of buffer in a remote CTA's SMEMBoth
tlx.remote_shmem_store(val, buf)Store to remote CTA's shared memoryBoth
tlx.async_remote_shmem_store(val, buf)Async store to remote CTA's shared memoryBoth
tlx.tmem_copy(src, dst)Copy between TMEM buffersBlackwell
tlx.fence_async_shared()Memory fence for async shared memory operationsBoth

Storage kinds: tlx.storage_kind.smem, tlx.storage_kind.tmem (Blackwell), tlx.storage_kind.smemCluster

Show full SKILL.md (589 more words)Show less
TMA (Tensor Memory Accelerator)
FunctionDescriptionArch
tlx.make_tensor_descriptor(ptr, shape, strides, block_shape)Create TMA descriptor from pointer (host-side)Hopper+
tlx.allocate_tensor_descriptor(ptr, shape, strides, block_shape, swizzle_mode)Allocate and fill TMA descriptor in SMEMHopper+
tlx.reinterpret_tensor_descriptor(desc, dtype)Reinterpret TMA descriptor with different dtypeHopper+
tlx.async_descriptor_load(desc, indices, barrier=None)Async TMA load from global → SMEM, tracked by barrierHopper+
tlx.async_descriptor_store(desc, val, indices)Async TMA store from registers → globalHopper+
tlx.async_descriptor_store_wait()Wait for all pending TMA stores to completeHopper+
tlx.async_load(ptr, buf, barrier)Async bulk copy global → SMEM (cp.async)Hopper+
tlx.async_load_commit_group()Commit async load groupHopper+
tlx.async_load_wait_group(n)Wait for async load groups (n pending allowed)Hopper+

Matrix Multiply (MMA)

FunctionDescriptionArch
tlx.async_dot(A, B, acc=None, use_acc=None, mBarriers=[], two_ctas=False)Warp-group MMA: D = A @ B + C. Maps to wgmma (Hopper) or tcgen05.mma (Blackwell)Both
tlx.async_dot_scaled(A, B, acc, A_scale, A_format, B_scale, B_format, ...)Scaled MMA with FP8 inputs: D = (Ascale_A) @ (Bscale_B) + DBlackwell
tlx.async_dot_wait(pendings, inp)Wait for N pending async dot operations to completeBoth
tlx.tcgen05_commit(mBarrier, two_ctas=False)Make mbarrier track completion of prior tcgen05 ops. Use a SEPARATE mbarrier from async_dotBlackwell

Minimum tile sizes for async_dot: M ≥ 64, K ≥ 16, N ≥ 32

Pair-CTA MMA (two_ctas=True): M must be 128 per CTA.

Multi-CTA (Cluster) Kernels

ctas_per_cga=(N,1,1) in triton.Config sets the cluster size. The grid specifies total CTAs; hardware divides by ctas_per_cga to get the number of clusters. E.g., grid=(2,1,1) with ctas_per_cga=(2,1,1) = 1 cluster of 2 CTAs.

2-CTA tile scheduling in attention backward

Each CTA in a cluster gets its own program_id and its own tile_id from CLC. Two CTAs in a cluster naturally get consecutive tiles (pid 0, pid 1). No special tile scheduling is needed for 2-CTA — start_n = pid works as-is. Grid size and n_tile_num do NOT change between 1-CTA and 2-CTA.

Think of 2-CTA as two independent 1-CTAs that handle their own K/V tiles and share Q/dO via multicast. For L2 efficiency, they process consecutive N-blocks.

2-CTA MMA semantics (two_ctas=True)
  • A operand (TMEM): per-CTA, each CTA has different data
  • B operand (SMEM): split across CTAs and combined by hardware via multicast
  • Output (TMEM): split across CTAs along the M dimension, written to both CTAs
  • Leader MMA writes to both leader TMEM and peer TMEM
2-CTA barrier patterns
  • TMA loads with two_ctas=True: only leader calls barrier_expect_bytes (guarded by if is_leader:). Use arrive_count=1.
  • Software arrives (barrier_arrive with remote_cta_rank=0): both CTAs arrive on leader's barrier. Use arrive_count=NUM_CTAS.
  • MMA mBarriers with two_ctas=True: hardware signals when input reads complete. The TMEM output write may still be in-flight.

input_precision options: tf32, tf32x3, ieee

CLC (Cluster Launch Control) — Blackwell only

FunctionDescription
tlx.clc_create_context(num_consumers, num_stages=1)Create CLC pipeline context (allocates barriers + response buffers)
tlx.clc_producer(context, p_producer, multi_ctas=False, k=0)Issue CLC try_cancel request from CTA 0
tlx.clc_consumer(context, p_consumer, multi_ctas=False, k=0, return_3d=False)Decode tile ID from CLC response, signal completion. Returns tile_id or -1. With return_3d=True, returns (ctaIdX, ctaIdY, ctaIdZ) tuple.

For 2-CTA mode: set multi_ctas=True (uses "arrive remote, wait local" pattern).

Utility

FunctionDescriptionArch
tlx.cluster_cta_rank()Unique CTA ID within a cluster (all dims)Both
tlx.thread_id(axis)Thread ID along axis 0, 1, or 2Both
tlx.dtype_of(tensor_or_desc)Get element type of tensor or tensor descriptorBoth
tlx.size_of(dtype)Size of dtype in bytesBoth
tlx.get_fp8_format_name(dtype)Get FP8 format string ("e5m2" or "e4m3") for scaled MMABoth
tlx.clock64()64-bit hardware clock value (for timing)Both
tlx.stoch_round(src, dst_ty, rand_bits)Hardware stochastic rounding FP32 → FP8/BF16/F16Blackwell

Common patterns

Producer-consumer with mbarrier (pipelined GEMM)
python
# Full/data-ready barriers track TMA transaction completion and remain ordinary.
bars_full = tlx.alloc_barriers(num_stages, arrive_count=1)

# Ordinary software-release form: one logical arrival per consumer replica.
bars_empty = tlx.alloc_barriers(
    num_stages,
    arrive_count=num_consumers,
)

# Optional optimized form when each consumer replica has four warps and every lane
# is statically guaranteed to execute one unit arrival after its final buffer read.
bars_empty_warp = tlx.alloc_warp_barrier(
    num_barriers=num_stages,
    num_warps=4,
    num_arrivals=num_consumers,
)

# Producer: wait for reuse, then start TMA load tracked by the ordinary full barrier.
tlx.barrier_wait(bar_empty, empty_phase)
tlx.barrier_expect_bytes(bar_full, nbytes)
tlx.async_descriptor_load(desc, indices, barrier=bar_full)

# Consumer: wait for TMA completion, consume the buffer, then release it.
tlx.barrier_wait(bar_full, full_phase)
acc = tlx.async_dot(A, B, acc)
acc = tlx.async_dot_wait(0, acc)  # Prove the final buffer read completed.
tlx.barrier_arrive(bar_empty)

When evaluating the optimized form, substitute bars_empty_warp for bars_empty at both the producer wait and consumer arrival sites. Do not convert bars_full.

PingPong with named barriers
python
# Consumer 0 waits for Consumer 1, then issues MMA
tlx.named_barrier_wait(9, 256)   # 256 = 2 warp groups * 4 warps * 32 threads
qk = tlx.async_dot(q, k)
tlx.named_barrier_arrive(10, 256)

# Consumer 1 waits for Consumer 0's MMA to finish
tlx.named_barrier_arrive(9, 256)
tlx.named_barrier_wait(10, 256)
qk = tlx.async_dot(q, k)

Deep-dive docs

  • API reference: third_party/tlx/README.md
  • Barriers: third_party/tlx/doc/tlx_barriers.md
  • Placeholder layouts: third_party/tlx/doc/PlaceholderLayouts.md
  • Storage alias design: third_party/tlx/doc/storage_alias_spec_design.md

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tlx-api-reference of facebookexperimental/triton.

Open the folder on GitHubat commit 612bd83

Compare with similar skills

Tlx API Reference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tlx API Reference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tlx API Reference this skillfacebookexperimental/triton201—~3.9kAutomated safety check: PassMIT
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
GPU Kubernetes Operationssickn33/agentic-awesome-skills47k2 repos~3.2kAutomated safety check: PassMIT
Makepad Dslsickn33/agentic-awesome-skills47k2 repos~1.2kAutomated safety check: PassMIT
Modal Serverless GPUOrchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT
ONNX Runtime GPU Transformers Testsmicrosoft/onnxruntime22k—~2.9kAutomated safety check: PassMIT

Similar skills

  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • GPU Kubernetes Operations

    sickn33/agentic-awesome-skills

    Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

    47k GitHub starsUsed in 2 repos~3.2k tokens
    DevOps & CloudAuto-check passed
  • Makepad Dsl

    sickn33/agentic-awesome-skills

    CRITICAL: Use for Makepad DSL syntax and inheritance. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Modal Serverless GPU

    Orchestra-Research/AI-Research-SKILLs

    Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs.

    13k GitHub starsUsed in 5 repos~2.1k tokens
    Backend & APIsAuto-check passed
  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • GPU Server Management

    sickn33/agentic-awesome-skills

    Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~2k tokens
    AI & LLM EngineeringAuto-check: notes

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Questions about Tlx API Reference

What does Tlx API Reference do?

TLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton. Tlx API Reference is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. TLX DSL API reference for low-level GPU primitives.

When should I use Tlx API Reference?

Tlx API Reference fits situations like: modifying TLX kernel code that uses barriers (mbarrier; named barriers); memory allocation (localalloc; warp specialization (asynctasks.

How do I install Tlx API Reference in Claude Code?

Run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a claude-code`. Or copy the skill folder (.claude/skills/tlx-api-reference in facebookexperimental/triton) into .claude/skills/tlx-api-reference in your project. Claude Code loads it when a task matches its description.

How do I install Tlx API Reference in Codex?

Run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a codex`. Or copy the skill folder (.claude/skills/tlx-api-reference in facebookexperimental/triton) into .agents/skills/tlx-api-reference in your project. Codex loads it when a task matches its description.

Can I use Tlx API Reference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tlx-api-reference, .gemini/skills/tlx-api-reference, .github/skills/tlx-api-reference and .opencode/skills/tlx-api-reference in your project.

What does Tlx API Reference need to run?

SKILL.md names no scripts, command-line tools or credentials: Tlx API Reference is instructions for the agent only. Our summary lists: Python 3.

Does Tlx API Reference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tlx API Reference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tlx API Reference use?

Tlx API Reference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tlx API Reference use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tlx API Reference?

Skills that share tags, products or a category with Tlx API Reference: Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars), GPU Kubernetes Operations (sickn33/agentic-awesome-skills, 47k stars), Makepad Dsl (sickn33/agentic-awesome-skills, 47k stars) and Modal Serverless GPU (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tlx API Reference?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 8, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.