Agent skill

Fla Optimization Loop

by fla-org in fla-org/flash-linear-attention

Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

MITAuto-check passedTesting & QA

Install Fla Optimization Loop

skills CLI
$ npx skills add fla-org/flash-linear-attention --skill fla-optimization-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fla-org/flash-linear-attention fla-optimization-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fla-optimization-loop .claude/skills/fla-optimization-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fla-optimization-loop
GitHub stars
5.8k
Token cost
~2.6k tokens
SKILL.md length
1,276 words
Files
3 (incl. references)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

  • Works in 8 steps: The inviolable rule: the test file is a… → Write the task contract first (before… → Three phases → …
  • Iterating on fla/ops/ performance over multiple rounds
  • SKILL.md covers 0. The inviolable rule: the…, 1. Write the task contract…, 2. Three phases and 3. Iteration protocol, plus 4 more sections
  • Calls python and git

What it does

Fla Optimization Loop is an agent skill from fla-org/flash-linear-attention. Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness. Synthesizes the task-contract / three-phase / iteration-protocol / silent-bug-catalog discipline of agent kernel-optimization frameworks (KDA, the MLSys FlashInfer contest workflow, AKO4ALL/AKO4X), and anchors all of it on FLA's frozen pytest (forward AND backward, under NaN poisoning) as the immutable correctness gate. Use when iterating on fla/ops/ performance over multiple…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/TRAPS.md` and `references/opt-log-template.md`).

It sits in Testing & QA, covering Unit testing. It works with pytest. The repository describes itself as: 🚀 Efficient implementations for emerging model architectures. The licence is MIT.

When your agent uses it

  • Iterating on fla/ops/ performance over multiple rounds
  • Tasks that involve Unit testing

Example prompts

  • “/fla-optimization-loop”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. The inviolable rule: the test file is a frozen contract
  2. Write the task contract first (before any code)
  3. Three phases
  4. Iteration protocol
  5. Reproducibility
  6. Silent-bug & measurement traps
  7. Evidence records (scratch workspace)
  8. Promotion → MR

What it can do on your machine

Read from SKILL.md and the folder at commit b8ff848. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fla Optimization Loop loads about 2.6k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 1,276 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fla-org/flash-linear-attention at commit b8ff848, republished under its MIT licence (© fla-org). 1,276 words, ~2,568 tokens.

Download SKILL.mdSave it as .claude/skills/fla-optimization-loop/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
fla-optimization-loop
description
Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness. Synthesizes the task-contract / three-phase / iteration-protocol / silent-bug-catalog discipline of agent kernel-optimization frameworks (KDA, the MLSys FlashInfer contest workflow, AKO4ALL/AKO4X), and anchors all of it on FLA's frozen pytest (forward AND backward, under NaN poisoning) as the immutable correctness gate. Use when iterating on `fla/ops/**` performance over multiple rounds.

FLA Optimization Loop Skill

Use this when you are making an existing fla/ops/** kernel faster — or bringing a new kernel from correct to fast — across more than one iteration, in any FLA backend language (Triton, Gluon, TileLang, CuTe DSL). This skill is the search discipline that ties the other skills together; it does not replace them:

  • fla-nvidia-performance — how to profile (NCU), hardware baselines, MR-ready perf evidence.
  • fla-ascend-performance — how to profile (torch_npu), diagnose Cube/Vector/MTE/UB bottlenecks, and optimize Triton-Ascend kernels.
  • fla-correctness-coverage — how to design the test coverage matrix for an op.
  • fla-mr-readiness — how to package the promoted change into a PR.

This skill covers the loop around those: what to lock, how to iterate, what to record, and when to stop.

0. The inviolable rule: the test file is a frozen contract

The op's tests/ops/test_<op>.py and its fla/ops/<op>/naive.py reference are frozen for the entire optimization loop. The whole point of "faster" only means something if correctness — forward and backward, under the conftest NaN-memory poisoning — is held fixed. During a perf loop you may not:

  • edit the test file, or its naive.py reference;
  • loosen an assert_close tolerance, or widen rms_eps / dtype to make a diff pass;
  • drop, narrow, or skip parametrized shapes;
  • special-case the kernel on values it only sees in the test;
  • cache module-level tensors so trials reuse warm/identical data.

Run the gate with the unmodified test, every iteration:

bash
python -m benchmarks.ops.verify --op <op> [--gate-k <subset>]

verify.py runs the pytest file as a black box and refuses to report a speedup on a red gate. --gate-k only selects a shape subset for a fast signal — it never edits the test; promote only on a full (no -k) green gate.

Banned vs. allowed implementation (anti-reward-hacking):

BannedAllowed
Making the op a thin wrapper that delegates the whole compute to a vendor lib (a plain torch.matmul / F.scaled_dot_product_attention standing in as the operator) just to win latencyHand-written Triton / Gluon / TileLang / CuTe kernels; torch ops used as glue around a kernel you wrote
Returning uninitialized / partially-written outputs that happen to passFully initialized outputs (NaN poisoning will catch partial writes)
Stream tricks / monkey-patching the bench to dodge timingGenuine latency reduction measured by verify.py / run.py
One-sided numeric relaxation the baseline doesn't get — flipping allow_tf32 on, dropping the fp32 accumulator to bf16/tf32, a config that quietly changes the numeric pathSame accumulation precision and numeric flags on both sides; speed comes from the kernel, not from computing something less accurate

If you genuinely believe a test is wrong, that is a separate PR with its own justification — never bundled into a perf change. Stop and ask the user.

1. Write the task contract first (before any code)

Put a short docs/draft.md in your scratch workspace (see §6) stating:

  • Op + entry point — e.g. chunk_gla in fla.ops.gla.
  • Target — which shapes (from benchmarks/ops/registry.py SHAPE_CONFIGS), and a target speedup vs. main.
  • Allowed languages — Triton / Gluon / TileLang / CuTe (state any constraint).
  • Validation command — python -m benchmarks.ops.verify --op <op> (the frozen gate).
  • Benchmark command — python -m benchmarks.ops.verify --op <op> --base main.
  • Promotion criteria — full green gate, a measured repeatable win, and a profiler reading that explains it (§7).
  • Frozen scope — the test file, naive.py, and the public op signature.

Do not start editing kernels until the draft exists. (Borrowed from KDA: plan, then execute.)

2. Three phases

Run these in order; repeat 2 and 3 with progressively higher targets.

  • Phase 1 — correct baseline. Confirm the current kernel passes the full gate, and record baseline numbers: python -m benchmarks.ops.verify --op <op> --base main. For a brand-new kernel, get the gate green first; performance is secondary here.
  • Phase 2 — profile-guided optimization. Use fla-nvidia-performance for NCU evidence on NVIDIA backends, or fla-ascend-performance for Ascend NPU / Triton-Ascend backends. Enumerate candidate directions, rank them by expected benefit vs. implementation risk, and explore each for at most a few iterations. Keep, revise, or reject each with evidence — don't optimize blindly.
  • Phase 3 — shape specialization. Only when profiling shows different bottlenecks across shape regimes (short vs. long T, small vs. large D), add dispatch / specialized paths. Justify the added complexity with the measured win; validate on the full shape set, not the one you tuned on. Record each bucket (condition / entry point / per-bucket latency + speedup / reason) in dispatch.md (template in references/opt-log-template.md) — a specialized path without that evidence is unjustified complexity.
Show full SKILL.md (572 more words)Show less

3. Iteration protocol

Every iteration is exactly three steps, in order, with no telescoping into the next iteration between them:

  1. Make one change to the kernel.
  2. Run verify.py — gate must stay green; record the bench number.
  3. Append one row to OPT_LOG.md (see references/opt-log-template.md) and git commit.

A failed or no-change iteration is still an iteration: log it and commit before debugging the next direction. (Borrowed from AKO4ALL: bench → log → commit, the most-skipped step in practice.)

Stall handling. After 3 consecutive iterations with no improvement (≥ a few % over current best, above noise), stop and re-assess: re-profile, re-read OPT_LOG.md for which axes you've already tried, and search for known techniques for this op family before picking a new direction.

When to stop. A user-set iteration cap is reached; or re-assessment produces hard evidence of a floor (bandwidth-bound at HBM limit, launch-overhead dominated, timer-resolution limited — cite it in OPT_LOG.md); or you've documented ≥3 distinct directions tried with evidence. Don't stop silently because a tool (e.g. NCU) was unavailable — that's a re-assessment input.

No-go bar. If stopping means concluding there's no win to promote — a no-go — that verdict has its own bar: a first candidate losing doesn't clear it. A no-go needs a recorded baseline number, at least one reasoned candidate attempt (not a blind guess), the gate status, the bench evidence, and a named active bound or blocker (name which roofline/launch/timer limit, with the number). Without those five it's an unfinished loop.

4. Reproducibility

  • Seed — the tests already pin torch.manual_seed(42); don't undermine it.
  • Environment — verify.py prints GPU / CUDA / PyTorch / Triton / commit SHA; paste that line into OPT_LOG.md so a number is interpretable later.
  • Full-shape verdict before promotion — a --gate-k subset is a signal only.
  • Rank by the solution's own runtime for fast iteration signal; pay for the full --base main comparison only at the verdict. Clock noise on unlocked GPUs can swing absolute speedup — see references/TRAPS.md.

5. Silent-bug & measurement traps

Read references/TRAPS.md before trusting any number. It catalogs FLA-specific traps (NaN poisoning on partial writes, TF32 inflating fp32 reference diffs, assert_close relative tolerance, one-sided numeric-flag relaxation, autotune-cache staleness across edits, int64 address arithmetic, implausible speedups that signal a silently-skipped path). When you hit a new one, add it there with Fact / Why / How to apply so the next session doesn't re-learn it. (Borrowed from AKO4X's TRAPS.md.)

6. Evidence records (scratch workspace)

Keep your search artifacts in a git-ignored directory — profile/<op>-opt/ is already ignored:

text
profile/<op>-opt/
  docs/draft.md       # the task contract (§1)
  OPT_LOG.md          # one row per iteration (template in references/)
  dispatch.md         # one row per shape bucket — only if Phase 3 specializes (template in references/)
  TRAPS.md            # traps you hit this session (seed: references/TRAPS.md)
  trace/              # torch.profiler / NCU artifacts (kept out of git)

Each kept candidate's kernel gets a short header (Identity / Delta / Lessons / Dead-ends / Open-directions) per references/opt-log-template.md, so a later session can see what was tried and why.

7. Promotion → MR

A candidate is promotable only on a full green gate plus a measured, repeatable win on the target shapes, and a profiler/roofline reading that explains the win (or, for a no-go, the blocker) — a speedup you can't account for is a silent-skip suspect, not a result (see §5). Then:

  • Keep the diff minimal — change only what the win needs, plus light cleanups (CONTRIBUTING "Protect battle-tested paths; keep diffs minimal").
  • Collect perf evidence per backend skill — fla-nvidia-performance (before/after, NCU summary, dense + varlen coverage) or fla-ascend-performance (before/after, pipe/UB metrics, round summary template) — and a full final-claim stats block (median/mean/std/min/p10/p90 per shape, equal-weight geomean speedup, exact commands, baseline commit + candidate SHA, GPU id/model with idle-clock evidence — the last guards the clock-drift trap in §5).
  • Package the PR per fla-mr-readiness.

The scratch workspace under profile/<op>-opt/ stays local; it is not part of the PR.

© fla-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/fla-optimization-loop of fla-org/flash-linear-attention.

  • SKILL.md
  • references/TRAPS.md
  • references/opt-log-template.md

Open the folder on GitHubat commit b8ff848

Compare with similar skills

Fla Optimization Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fla Optimization Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fla Optimization Loop this skillfla-org/flash-linear-attention5.8k—~2.6kAutomated safety check: PassMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Pytest Runnersaleor/saleor23k—~251Automated safety check: PassBSD-3-Clause
Port Node Red Nodeoldrev/edgelinkd121—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Pytest Runner

    saleor/saleor

    Run pytest tests with automatic virtual environment activation. Use this skill whenever running tests, executing pytest, or when asked to "run tests", "test…

    23k GitHub stars~251 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Port Node Red Node

    oldrev/edgelinkd

    Port a Node-RED node into EdgeLinkd the way this repo does it: implement the node in Rust under crates/core/src/runtime/nodes, mirror Node-RED's mocha spec as pytest tests under tests/, register the…

    121 GitHub stars~3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed

More from fla-org/flash-linear-attention

All 9 skills in this repo
  • Fla Ascend Performance

    fla-org/flash-linear-attention

    Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.

    5.8k GitHub stars~5.6k tokensUpdated yesterday
    Auto-check passed
  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Fla Correctness Coverage

    fla-org/flash-linear-attention

    Guidelines for kernel correctness testing and coverage in fla/ops/ and related modules, including common Triton grid/addressing pitfalls.

    5.8k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Fla Design Coverage

    fla-org/flash-linear-attention

    Contract-first design and coverage discipline for FLA kernel and numerical changes.

    5.8k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Fla Dispatch Backends

    fla-org/flash-linear-attention

    Workflow for FLA backend dispatch decorators and backend implementations.

    5.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Fla Kda

    fla-org/flash-linear-attention

    FLA KDA kernel workflow and public technical notes. An agent skill from fla-org/flash-linear-attention.

    5.8k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Fla Optimization Loop

What does Fla Optimization Loop do?

Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness. Fla Optimization Loop is an agent skill from fla-org/flash-linear-attention. Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

When should I use Fla Optimization Loop?

Fla Optimization Loop fits situations like: iterating on fla/ops/ performance over multiple rounds; tasks that involve Unit testing.

How do I install Fla Optimization Loop in Claude Code?

Run `npx skills add fla-org/flash-linear-attention --skill fla-optimization-loop -a claude-code`. Or copy the skill folder (.agents/skills/fla-optimization-loop in fla-org/flash-linear-attention) into .claude/skills/fla-optimization-loop in your project. Claude Code loads it when a task matches its description.

How do I install Fla Optimization Loop in Codex?

Run `npx skills add fla-org/flash-linear-attention --skill fla-optimization-loop -a codex`. Or copy the skill folder (.agents/skills/fla-optimization-loop in fla-org/flash-linear-attention) into .agents/skills/fla-optimization-loop in your project. Codex loads it when a task matches its description.

Can I use Fla Optimization Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fla-org/flash-linear-attention --skill fla-optimization-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fla-optimization-loop, .gemini/skills/fla-optimization-loop, .github/skills/fla-optimization-loop and .opencode/skills/fla-optimization-loop in your project.

What does Fla Optimization Loop need to run?

Going by SKILL.md and its folder, Fla Optimization Loop needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Fla Optimization Loop access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fla Optimization Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fla Optimization Loop use?

Fla Optimization Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fla Optimization Loop use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Fla Optimization Loop?

Skills that share tags, products or a category with Fla Optimization Loop: Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars), Test Guard (amElnagdy/guard-skills, 1.3k stars) and Pytest Runner (saleor/saleor, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fla Optimization Loop?

fla-org (a GitHub organization) maintains it in fla-org/flash-linear-attention, which has 5,828 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 6, 2026.

Source: fla-org/flash-linear-attention on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.