Agent skill

Fla Triton To Gluon

by fla-org in fla-org/flash-linear-attention

Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

MITAuto-check passedDevelopment

Install Fla Triton To Gluon

skills CLI
$ npx skills add fla-org/flash-linear-attention --skill fla-triton-to-gluon -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fla-org/flash-linear-attention fla-triton-to-gluon --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fla-triton-to-gluon .claude/skills/fla-triton-to-gluon && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fla-triton-to-gluon
GitHub stars
5.8k
Token cost
~1.6k tokens
SKILL.md length
799 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

  • Works in 5 steps: Record the baseline commit and retain… → Translate loads, stores, and arithmetic… → Add async transfers only where load… → …
  • Tasks that involve Async programming
  • SKILL.md covers Check the target environment, Port incrementally, Layout and resource constraints and Async correctness, plus 1 more section
  • Calls git and python

What it does

Fla Triton To Gluon is an agent skill from fla-org/flash-linear-attention. Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Async programming and GPU and accelerator computing. It works with NVIDIA AI Platform. The repository describes itself as: 🚀 Efficient implementations for emerging model architectures. The licence is MIT.

When your agent uses it

  • Tasks that involve Async programming
  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/fla-triton-to-gluon”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Record the baseline commit and retain the validated tests, reference, and tolerances. Keep the Triton fallback available.
  2. Translate loads, stores, and arithmetic with explicit layouts first. Match each tensor's contiguous dimension and preserve masking. Pass…
  3. Add async transfers only where load latency or addressing pressure limits the kernel. Define buffer ownership, completion, and reuse for…
  4. Add architecture-specific MMA or scheduling when profiling supports it. Recheck register/shared-memory use and retune after each change.
  5. Run the full relevant correctness gate and same-hardware benchmarks, including supported variable-length and boundary cases. Measure…

What it can do on your machine

Read from SKILL.md and the folder at commit f3bec72. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • triton-lang.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fla Triton To Gluon loads about 1.6k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 799 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fla-org/flash-linear-attention at commit f3bec72, republished under its MIT licence (© fla-org). 799 words, ~1,604 tokens.

Download SKILL.mdSave it as .claude/skills/fla-triton-to-gluon/SKILL.md (or your agent's skills folder).
name
fla-triton-to-gluon
description
Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

Triton to Gluon

Use Gluon when profiling points to register pressure, layout conversion, or load/compute overlap that needs explicit control. A kernel already saturating its useful bandwidth or compute limit may not benefit. Keep the portable default path when adding hardware-specific implementations.

Follow fla-optimization-loop for the baseline and correctness gate, and fla-nvidia-performance for NVIDIA measurements.

Check the target environment

Gluon is experimental. Check the installed Triton API before copying upstream examples; their names and signatures can differ from the installed release.

Gate architecture-specific code appropriately: Ampere supports cp.async, Hopper adds TMA/WGMMA, and Blackwell adds TMEM/tcgen05. Code under Gluon's nvidia namespace requires NVIDIA hardware. Keep compile-time choices in explicit constexpr arguments so the kernel, launcher, and autotune pruning use the same values.

python
from triton.experimental import gluon
from triton.experimental.gluon import language as gl

Port incrementally

  1. Record the baseline commit and retain the validated tests, reference, and tolerances. Keep the Triton fallback available.
  2. Translate loads, stores, and arithmetic with explicit layouts first. Match each tensor's contiguous dimension and preserve masking. Pass forward, state, and supported backward comparisons before adding asynchrony.
  3. Add async transfers only where load latency or addressing pressure limits the kernel. Define buffer ownership, completion, and reuse for the prologue, steady state, and drain.
  4. Add architecture-specific MMA or scheduling when profiling supports it. Recheck register/shared-memory use and retune after each change.
  5. Run the full relevant correctness gate and same-hardware benchmarks, including supported variable-length and boundary cases. Measure production entry points as well as individual kernels.

A literal translation establishes parity; measure it before assuming a performance gain. Preserve the validated numerical algorithm and precision while changing execution layout or scheduling.

Layout and resource constraints

  • Pointer tensors and values need compatible layouts. A cross-warp gl.convert_layout may use shared memory; use assert_trivial=True when a conversion is intended to be free.
  • Broadcasts and oversized layout tiles can duplicate values and increase register use. Choose layouts from actual tensor strides rather than copying a tutorial's block sizes.
  • Account for all live shared-memory buffers and the target device's allocation limit. Prune infeasible autotune configurations; provide a streaming path when a resident design cannot fit supported shapes.
  • Large static unrolls multiplied by many autotune configurations can dominate compilation. Prune from the actual shape and live-memory budget, and reuse the worker and Triton cache during iteration.
  • Persistent schedules can trade launch overhead for worse L2 locality. Warp specialization also changes the total register budget; measure occupancy and stalls after changing either.

Async correctness

Transfers and buffer reuse

cp.async groups are ordered as a queue. wait_group(N) bounds all outstanding groups, so it cannot identify a particular buffer after unrelated prefetches are interleaved. For per-buffer scheduling, use separate mbarriers and match their arrival counts to the participating threads. In the installed API, check whether mbarrier_arrive increments the count; a preinitialized thread count requires the non-incrementing form.

Wait for a load before consuming its buffer, and protect reuse until all readers finish. Even same-lane shared-memory staging needs protection against overwriting data still being read. Track each mbarrier's phase as its buffer is reused; do not run more than one phase ahead or share a completion barrier between TMA and tcgen05 without reinitializing it.

TMA descriptors must satisfy the target's alignment and stride requirements. TMA stores may distinguish completion of the shared-memory read from completion of the global write. Check the installed wait API and require global completion before another operation reads the stored range.

Show full SKILL.md (241 more words)Show less
MMA and memory ordering

Hopper WGMMA requires its B operand in shared memory and uses register accumulators; consume the values returned by its wait operation so compiler dependencies remain explicit. Blackwell tcgen05 uses TMEM accumulators and mbarrier completion; respect the participating warpgroup and TMEM layout requirements. Initialize accumulators explicitly, including use_acc=False where supported.

Generic shared-memory accesses and async operations use different memory proxies. Apply the required proxy fence when handing a buffer to an async consumer; an mbarrier alone does not replace that fence. A completed TMA load establishes the ordering needed to read its destination. Check ordering across warp-specialized partitions as well as within a single producer/consumer loop.

Tails and exact-zero cases

Masked async copies can leave shared-memory elements uninitialized. Prevent invalid rows from entering reductions: zero-fill where supported, or use valid-row loads and explicitly remove every invalid contribution. Multiplying by zero does not neutralize a NaN. Keep output stores masked.

Do not rely on separately compiled reductions cancelling bitwise. For a mathematically exact-zero special case, such as a single-source softmax gradient, preserve the exact-zero result explicitly and test it against the reference.

Verification commands

Run from the repository root on the target GPU:

bash
FLA_BENCH_BASE=$(git rev-parse origin/main)
FLA_CI_ENV=0 python -m benchmarks.ops.verify --op chunk_kda --base "$FLA_BENCH_BASE"

Replace the operation with the one being ported. The command executes the current tests; it does not freeze them. Use --gate-k only for quick iteration, then run the full gate. Set backend environment flags before launching a fresh process so import-time choices cannot contaminate the comparison.

© fla-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/fla-triton-to-gluon of fla-org/flash-linear-attention.

Open the folder on GitHubat commit f3bec72

Compare with similar skills

Fla Triton To Gluon next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fla Triton To Gluon compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fla Triton To Gluon this skillfla-org/flash-linear-attention5.8k—~1.6kAutomated safety check: PassMIT
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Tilelang Developeryzlnew/infra-skills149—~2.4kAutomated safety check: PassNone
Cuda Debuggingmohitmishra786/low-level-dev-skills252—~1.5kAutomated safety check: PassMIT
Cuda Profilingmohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Tilelang Developer

    yzlnew/infra-skills

    Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

    149 GitHub stars~2.4k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Cuda Debugging

    mohitmishra786/low-level-dev-skills

    CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Cuda Profiling

    mohitmishra786/low-level-dev-skills

    CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    188 GitHub stars~547 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed

More from fla-org/flash-linear-attention

All 9 skills in this repo
  • Fla Ascend Performance

    fla-org/flash-linear-attention

    Profile and optimize FLA Triton-Ascend kernels using NPU traces, with guidance for UB capacity, memory movement, launch limits, and numerical correctness.

    5.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Fla Correctness Coverage

    fla-org/flash-linear-attention

    Select and run correctness coverage for FLA kernels and modules, including gradients, dispatch boundaries, and Triton addressing changes.

    5.8k GitHub stars~902 tokensUpdated today
    Auto-check passed
  • Fla Optimization Loop

    fla-org/flash-linear-attention

    Iterate on FLA kernel performance with a frozen correctness gate, a measured baseline, and reproducible candidate comparisons.

    5.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Fla Nvidia Performance

    fla-org/flash-linear-attention

    Profile and optimize FLA kernels on NVIDIA GPUs, with same-hardware benchmarks and targeted Nsight Compute analysis.

    5.8k GitHub stars~818 tokensUpdated today
    Auto-check passed
  • Fla PR Readiness

    fla-org/flash-linear-attention

    Prepare or update an FLA pull request with focused scope, verified evidence, and the current repository template.

    5.8k GitHub stars~856 tokensUpdated today
    Auto-check passed
  • Fla Design Coverage

    fla-org/flash-linear-attention

    Define supported inputs, numerical budgets, routing, and validation before designing an FLA kernel or numerical change.

    5.8k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Questions about Fla Triton To Gluon

What does Fla Triton To Gluon do?

Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck. Fla Triton To Gluon is an agent skill from fla-org/flash-linear-attention. Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

When should I use Fla Triton To Gluon?

Fla Triton To Gluon fits situations like: tasks that involve Async programming; tasks that involve GPU and accelerator computing.

How do I install Fla Triton To Gluon in Claude Code?

Run `npx skills add fla-org/flash-linear-attention --skill fla-triton-to-gluon -a claude-code`. Or copy the skill folder (.agents/skills/fla-triton-to-gluon in fla-org/flash-linear-attention) into .claude/skills/fla-triton-to-gluon in your project. Claude Code loads it when a task matches its description.

How do I install Fla Triton To Gluon in Codex?

Run `npx skills add fla-org/flash-linear-attention --skill fla-triton-to-gluon -a codex`. Or copy the skill folder (.agents/skills/fla-triton-to-gluon in fla-org/flash-linear-attention) into .agents/skills/fla-triton-to-gluon in your project. Codex loads it when a task matches its description.

Can I use Fla Triton To Gluon in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fla-org/flash-linear-attention --skill fla-triton-to-gluon -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fla-triton-to-gluon, .gemini/skills/fla-triton-to-gluon, .github/skills/fla-triton-to-gluon and .opencode/skills/fla-triton-to-gluon in your project.

What does Fla Triton To Gluon need to run?

Going by SKILL.md and its folder, Fla Triton To Gluon needs the command-line tools its instructions call (git and python). Our summary lists: Python 3.

Does Fla Triton To Gluon access the network?

SKILL.md names 1 domain. As links in the text: triton-lang.org. This is read from the text; nothing was executed.

Is Fla Triton To Gluon safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fla Triton To Gluon use?

Fla Triton To Gluon is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fla Triton To Gluon use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fla Triton To Gluon?

Skills that share tags, products or a category with Fla Triton To Gluon: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Tilelang Developer (yzlnew/infra-skills, 149 stars), Cuda Debugging (mohitmishra786/low-level-dev-skills, 252 stars) and Cuda Profiling (mohitmishra786/low-level-dev-skills, 252 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fla Triton To Gluon?

fla-org (a GitHub organization) maintains it in fla-org/flash-linear-attention, which has 5,842 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 10, 2026.

Source: fla-org/flash-linear-attention on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.