Optimize For GPU
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
TLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton.
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install facebookexperimental/triton tlx-api-reference --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tlx-api-reference .claude/skills/tlx-api-reference && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .claude/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-referenceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install facebookexperimental/triton tlx-api-reference --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/tlx-api-reference .agents/skills/tlx-api-reference && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .agents/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install facebookexperimental/triton tlx-api-reference --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/tlx-api-reference .cursor/skills/tlx-api-reference && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .cursor/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/facebookexperimental/triton.git --path .claude/skills/tlx-api-reference--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install facebookexperimental/triton tlx-api-reference --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/tlx-api-reference .gemini/skills/tlx-api-reference && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .gemini/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install facebookexperimental/triton tlx-api-referenceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/tlx-api-reference .github/skills/tlx-api-reference && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .github/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tlx-api-reference -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install facebookexperimental/triton tlx-api-reference --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/tlx-api-reference .opencode/skills/tlx-api-reference && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tlx-api-reference" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tlx-api-reference into .opencode/skills/tlx-api-reference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tlx-api-reference", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tlx-api-referenceTLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton.
Tlx API Reference is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. TLX DSL API reference for low-level GPU primitives. Use when writing or modifying TLX kernel code that uses barriers (mbarrier, named barriers), memory allocation (localalloc, SMEM, TMEM), TMA operations, warp specialization (asynctasks, asynctask), CLC (cluster launch control), or wgmma instructions. Covers Hopper and Blackwell hardware differences.
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 612bd83. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tlx API Reference loads about 3.9k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,573 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from facebookexperimental/triton at commit 612bd83, republished under its MIT licence (© facebookexperimental). 1,573 words, ~3,895 tokens.
.claude/skills/tlx-api-reference/SKILL.md (or your agent's skills folder).| Function | Description | Arch |
|---|---|---|
tlx.async_tasks() | Context manager wrapping all async task regions | Both |
tlx.async_task([task_ids]) | Assign code to specific task IDs (e.g., [0] = producer, [1,2] = consumers) | Both |
tlx.async_task(num_warps=N, num_regs=R) | Explicit warp/register allocation for a task | Both |
tlx.async_task("default", num_regs=R) | Default task for code outside explicit tasks | Both |
tlx.async_task_replica_id() | Returns replica ID inside an async region | Both |
with tlx.async_tasks():
with tlx.async_task([0]): # Producer
# TMA loads
with tlx.async_task([1, 2]): # Consumers
# MMA compute| Function | Description | Arch |
|---|---|---|
tlx.alloc_barriers(num_barriers, arrive_count=1) | Allocate ordinary SMEM barriers. Software arrivals use leader-based lowering, which may synchronize participating threads before the leader arrives. | Both |
tlx.alloc_warp_barrier(num_barriers, num_warps=1, num_arrivals=1) | Allocate SMEM barriers whose software arrivals are performed independently by every participating thread. The initialized count is num_warps * 32 * num_arrivals. | NVIDIA |
tlx.barrier_expect_bytes(bar, bytes, pred=None) | Set expected transaction byte count on barrier | Both |
tlx.barrier_wait(bar, phase, pred=None) | Wait until barrier phase flips (LOCAL mbarrier only) | Both |
tlx.barrier_arrive(bar, arrive_count=1, remote_cta_rank=None) | Signal arrival at barrier. remote_cta_rank signals a barrier in a remote CTA — only valid when ctas_per_cga > 1, causes "Unexpected buffer remote view in 1cta mode" otherwise. Guard with if USE_2CTA: when kernel supports both modes. | Both |
tlx.cluster_barrier() | Full cluster-wide synchronization barrier | Both |
Ordinary barrier count rules:
alloc_barriers(..., arrive_count=N) initializes the number of logical arrivals
required to complete a phase. Transaction bytes registered with
barrier_expect_bytes are an additional completion condition; do not infer the
software arrival topology from the transaction byte count.arrive_count=1 unless the surrounding protocol explicitly requires additional
software arrivals.tlx.async_task consumers, the
ordinary count is typically the number of consumer replicas that arrive once per
phase.remote_cta_rank, derive the count from the
exact CTA protocol. Do not assume the local-task rule applies.alloc_warp_barrier changes the physical arrival protocol, not merely the spelling
of the allocation. With an ordinary barrier, TLX selects a leader to perform the
arrival and may synchronize the participating threads first. With a warp barrier,
every participating thread performs its own arrival, avoiding that leader-path
synchronization but issuing more mbarrier arrival operations.
The initialized count is:
expected arrivals per phase = num_warps * 32 * num_arrivalsHere num_warps is the number of warps in each participating task replica, and
num_arrivals is the number of such all-thread arrival events targeting the same
barrier in one phase. For example, two replicated four-warp consumers that each
release one shared buffer slot use num_warps=4, num_arrivals=2, for 256 arrivals.
A consumer-owned slot released by one four-warp replica uses
num_warps=4, num_arrivals=1, for 128 arrivals. Keep the corresponding
barrier_arrive calls at their original final-use points and normally use the
default unit arrival count.
Before converting an ordinary barrier, classify its role and arrival source:
| Barrier role | Warp-barrier candidate? | Rule |
|---|---|---|
| Local software empty/reuse notification | Yes, after proving the topology | Every expected lane must arrive exactly once per declared arrival event, after its final read of the protected storage. |
| Local software data-ready notification | Sometimes | Safe only when readiness is produced by the same statically known all-thread topology; benchmark because per-thread arrivals are not universally faster. |
| TMA transaction/full barrier | No | Keep an ordinary barrier so transaction completion remains tracked through barrier_expect_bytes and the TMA operation. |
| MMA/tensor-core completion barrier | No automatic conversion | Preserve the completion mechanism required by the MMA API. |
| Named scheduling barrier | No | Preserve the named-barrier ID, participant count, direction, and phase protocol. |
| Remote, multicast, or cross-CTA barrier | No automatic conversion | Keep the established protocol unless backend support and the complete cluster-wide arrival topology are explicitly proven. |
| Divergently predicated arrival | No | A missing lane leaves the phase incomplete; use a warp barrier only when every counted lane is guaranteed to execute the required arrivals. |
Safe-conversion checklist:
barrier_arrive calls rather than TMA, MMA, multicast, or remote
completion.num_warps * 32 * num_arrivals equals the exact number of unit arrivals.| Function | Description | Arch |
|---|---|---|
tlx.named_barrier_wait(bar_id, num_threads) | Wait until the total participant count reaches bar_id | NVIDIA |
tlx.named_barrier_arrive(bar_id, num_threads) | Signal arrival at bar_id using the total participant count | NVIDIA |
num_threads is the total number of threads required to flip the barrier phase:
num_waiting_threads + num_arriving_threads. Wait and Arrive calls for the same
barrier phase must use the same value. The count must be a multiple of 32 (warp
size); it is typically num_warp_groups * warps_per_group * 32.
Used for PingPong scheduling to prevent tensor core contention between consumer warp groups.
| Function | Description | Arch |
|---|---|---|
tlx.local_alloc(shape, dtype, num, storage=smem, reuse=None, layout=None) | Allocate buffered tensor in SMEM or TMEM | Both (TMEM: Blackwell) |
tlx.storage_alias_spec(storage=smem, buffer_size_bytes=None) | Define shared buffer region for multiple local_alloc calls via reuse | Both |
tlx.local_view(buf, index) | Get view of a single buffer from a multi-buffered tensor | Both |
tlx.local_slice(buf, start, end) | Slice a sub-range of a buffered tensor | Both |
tlx.subslice(tensor, dim, start, size) | Subslice a tensor along a dimension | Both |
tlx.local_load(buf) | Load from SMEM/TMEM buffer into registers | Both |
tlx.local_store(val, buf) | Store from registers into SMEM/TMEM buffer | Both |
tlx.local_trans(buf) | Transpose a shared memory buffer | Both |
tlx.local_reinterpret(buf, dtype) | Reinterpret buffer with a different dtype | Both |
tlx.remote_view(buf, remote_cta_rank) | Get view of buffer in a remote CTA's SMEM | Both |
tlx.remote_shmem_store(val, buf) | Store to remote CTA's shared memory | Both |
tlx.async_remote_shmem_store(val, buf) | Async store to remote CTA's shared memory | Both |
tlx.tmem_copy(src, dst) | Copy between TMEM buffers | Blackwell |
tlx.fence_async_shared() | Memory fence for async shared memory operations | Both |
Storage kinds: tlx.storage_kind.smem, tlx.storage_kind.tmem (Blackwell), tlx.storage_kind.smemCluster
| Function | Description | Arch |
|---|---|---|
tlx.make_tensor_descriptor(ptr, shape, strides, block_shape) | Create TMA descriptor from pointer (host-side) | Hopper+ |
tlx.allocate_tensor_descriptor(ptr, shape, strides, block_shape, swizzle_mode) | Allocate and fill TMA descriptor in SMEM | Hopper+ |
tlx.reinterpret_tensor_descriptor(desc, dtype) | Reinterpret TMA descriptor with different dtype | Hopper+ |
tlx.async_descriptor_load(desc, indices, barrier=None) | Async TMA load from global → SMEM, tracked by barrier | Hopper+ |
tlx.async_descriptor_store(desc, val, indices) | Async TMA store from registers → global | Hopper+ |
tlx.async_descriptor_store_wait() | Wait for all pending TMA stores to complete | Hopper+ |
tlx.async_load(ptr, buf, barrier) | Async bulk copy global → SMEM (cp.async) | Hopper+ |
tlx.async_load_commit_group() | Commit async load group | Hopper+ |
tlx.async_load_wait_group(n) | Wait for async load groups (n pending allowed) | Hopper+ |
| Function | Description | Arch |
|---|---|---|
tlx.async_dot(A, B, acc=None, use_acc=None, mBarriers=[], two_ctas=False) | Warp-group MMA: D = A @ B + C. Maps to wgmma (Hopper) or tcgen05.mma (Blackwell) | Both |
tlx.async_dot_scaled(A, B, acc, A_scale, A_format, B_scale, B_format, ...) | Scaled MMA with FP8 inputs: D = (Ascale_A) @ (Bscale_B) + D | Blackwell |
tlx.async_dot_wait(pendings, inp) | Wait for N pending async dot operations to complete | Both |
tlx.tcgen05_commit(mBarrier, two_ctas=False) | Make mbarrier track completion of prior tcgen05 ops. Use a SEPARATE mbarrier from async_dot | Blackwell |
Minimum tile sizes for async_dot: M ≥ 64, K ≥ 16, N ≥ 32
Pair-CTA MMA (two_ctas=True): M must be 128 per CTA.
ctas_per_cga=(N,1,1) in triton.Config sets the cluster size. The grid
specifies total CTAs; hardware divides by ctas_per_cga to get the number
of clusters. E.g., grid=(2,1,1) with ctas_per_cga=(2,1,1) = 1 cluster of
2 CTAs.
Each CTA in a cluster gets its own program_id and its own tile_id from
CLC. Two CTAs in a cluster naturally get consecutive tiles (pid 0, pid 1).
No special tile scheduling is needed for 2-CTA — start_n = pid works
as-is. Grid size and n_tile_num do NOT change between 1-CTA and 2-CTA.
Think of 2-CTA as two independent 1-CTAs that handle their own K/V tiles and share Q/dO via multicast. For L2 efficiency, they process consecutive N-blocks.
two_ctas=True)two_ctas=True: only leader calls barrier_expect_bytes
(guarded by if is_leader:). Use arrive_count=1.barrier_arrive with remote_cta_rank=0): both CTAs
arrive on leader's barrier. Use arrive_count=NUM_CTAS.mBarriers with two_ctas=True: hardware signals when input reads
complete. The TMEM output write may still be in-flight.input_precision options: tf32, tf32x3, ieee
| Function | Description |
|---|---|
tlx.clc_create_context(num_consumers, num_stages=1) | Create CLC pipeline context (allocates barriers + response buffers) |
tlx.clc_producer(context, p_producer, multi_ctas=False, k=0) | Issue CLC try_cancel request from CTA 0 |
tlx.clc_consumer(context, p_consumer, multi_ctas=False, k=0, return_3d=False) | Decode tile ID from CLC response, signal completion. Returns tile_id or -1. With return_3d=True, returns (ctaIdX, ctaIdY, ctaIdZ) tuple. |
For 2-CTA mode: set multi_ctas=True (uses "arrive remote, wait local" pattern).
| Function | Description | Arch |
|---|---|---|
tlx.cluster_cta_rank() | Unique CTA ID within a cluster (all dims) | Both |
tlx.thread_id(axis) | Thread ID along axis 0, 1, or 2 | Both |
tlx.dtype_of(tensor_or_desc) | Get element type of tensor or tensor descriptor | Both |
tlx.size_of(dtype) | Size of dtype in bytes | Both |
tlx.get_fp8_format_name(dtype) | Get FP8 format string ("e5m2" or "e4m3") for scaled MMA | Both |
tlx.clock64() | 64-bit hardware clock value (for timing) | Both |
tlx.stoch_round(src, dst_ty, rand_bits) | Hardware stochastic rounding FP32 → FP8/BF16/F16 | Blackwell |
# Full/data-ready barriers track TMA transaction completion and remain ordinary.
bars_full = tlx.alloc_barriers(num_stages, arrive_count=1)
# Ordinary software-release form: one logical arrival per consumer replica.
bars_empty = tlx.alloc_barriers(
num_stages,
arrive_count=num_consumers,
)
# Optional optimized form when each consumer replica has four warps and every lane
# is statically guaranteed to execute one unit arrival after its final buffer read.
bars_empty_warp = tlx.alloc_warp_barrier(
num_barriers=num_stages,
num_warps=4,
num_arrivals=num_consumers,
)
# Producer: wait for reuse, then start TMA load tracked by the ordinary full barrier.
tlx.barrier_wait(bar_empty, empty_phase)
tlx.barrier_expect_bytes(bar_full, nbytes)
tlx.async_descriptor_load(desc, indices, barrier=bar_full)
# Consumer: wait for TMA completion, consume the buffer, then release it.
tlx.barrier_wait(bar_full, full_phase)
acc = tlx.async_dot(A, B, acc)
acc = tlx.async_dot_wait(0, acc) # Prove the final buffer read completed.
tlx.barrier_arrive(bar_empty)When evaluating the optimized form, substitute bars_empty_warp for bars_empty at
both the producer wait and consumer arrival sites. Do not convert bars_full.
# Consumer 0 waits for Consumer 1, then issues MMA
tlx.named_barrier_wait(9, 256) # 256 = 2 warp groups * 4 warps * 32 threads
qk = tlx.async_dot(q, k)
tlx.named_barrier_arrive(10, 256)
# Consumer 1 waits for Consumer 0's MMA to finish
tlx.named_barrier_arrive(9, 256)
tlx.named_barrier_wait(10, 256)
qk = tlx.async_dot(q, k)third_party/tlx/README.mdthird_party/tlx/doc/tlx_barriers.mdthird_party/tlx/doc/PlaceholderLayouts.mdthird_party/tlx/doc/storage_alias_spec_design.md© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/tlx-api-reference of facebookexperimental/triton.
Open the folder on GitHubat commit 612bd83
Tlx API Reference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tlx API Reference this skillfacebookexperimental/triton | 201 | — | ~3.9k | Automated safety check: Pass | MIT | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| GPU Kubernetes Operationssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Makepad Dslsickn33/agentic-awesome-skills | 47k | 2 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Modal Serverless GPUOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | |
| ONNX Runtime GPU Transformers Testsmicrosoft/onnxruntime | 22k | — | ~2.9k | Automated safety check: Pass | MIT |
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
sickn33/agentic-awesome-skills
Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
sickn33/agentic-awesome-skills
CRITICAL: Use for Makepad DSL syntax and inheritance. An agent skill from sickn33/agentic-awesome-skills.
Orchestra-Research/AI-Research-SKILLs
Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs.
microsoft/onnxruntime
Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.
sickn33/agentic-awesome-skills
Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.
facebookexperimental/triton
Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.
facebookexperimental/triton
Design and run Triton TTGIR debugging ablations using iroverride.
facebookexperimental/triton
Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.
facebookexperimental/triton
Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.
facebookexperimental/triton
Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.
facebookexperimental/triton
Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).
TLX DSL API reference for low-level GPU primitives. An agent skill from facebookexperimental/triton. Tlx API Reference is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. TLX DSL API reference for low-level GPU primitives.
Tlx API Reference fits situations like: modifying TLX kernel code that uses barriers (mbarrier; named barriers); memory allocation (localalloc; warp specialization (asynctasks.
Run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a claude-code`. Or copy the skill folder (.claude/skills/tlx-api-reference in facebookexperimental/triton) into .claude/skills/tlx-api-reference in your project. Claude Code loads it when a task matches its description.
Run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a codex`. Or copy the skill folder (.claude/skills/tlx-api-reference in facebookexperimental/triton) into .agents/skills/tlx-api-reference in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tlx-api-reference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tlx-api-reference, .gemini/skills/tlx-api-reference, .github/skills/tlx-api-reference and .opencode/skills/tlx-api-reference in your project.
SKILL.md names no scripts, command-line tools or credentials: Tlx API Reference is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tlx API Reference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tlx API Reference: Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars), GPU Kubernetes Operations (sickn33/agentic-awesome-skills, 47k stars), Makepad Dsl (sickn33/agentic-awesome-skills, 47k stars) and Modal Serverless GPU (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 8, 2026.
Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.