Search

By fla-org

9 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.

fla-org/flash-linear-attention5.8k—~6.3kAutomated safety check: PassMITtoday
2

Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

fla-org/flash-linear-attention5.8k—~3.1kAutomated safety check: PassMITtoday
3

Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

fla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMITtoday
4

Guidelines for kernel correctness testing and coverage in fla/ops/ and related modules, including common Triton grid/addressing pitfalls.

fla-org/flash-linear-attention5.8k—~1.5kAutomated safety check: PassMITtoday
5

Contract-first design and coverage discipline for FLA kernel and numerical changes.

fla-org/flash-linear-attention5.8k—~3.6kAutomated safety check: PassMITtoday
6

FLA KDA kernel workflow and public technical notes. An agent skill from fla-org/flash-linear-attention.

fla-org/flash-linear-attention5.8k—~1.4kAutomated safety check: PassMITtoday
7

Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo.

fla-org/flash-linear-attention5.8k—~1.2kAutomated safety check: PassMITtoday
8

Prepare or update an FLA pull request with one concrete purpose, a concise description, verified tests and benchmarks, and justified size exceptions.

fla-org/flash-linear-attention5.8k—~1.2kAutomated safety check: PassMITtoday
9

Change FLA backend registration, dispatch, or verifiers while preserving routing and call contracts.

fla-org/flash-linear-attention5.8k—~649Automated safety check: PassMITtoday