Search

fla-org/flash-linear-attention

9 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Profile and optimize FLA Triton-Ascend kernels using NPU traces, with guidance for UB capacity, memory movement, launch limits, and numerical correctness.

fla-org/flash-linear-attention5.8k—~1.1kAutomated safety check: PassMITtoday
2

Select and run correctness coverage for FLA kernels and modules, including gradients, dispatch boundaries, and Triton addressing changes.

fla-org/flash-linear-attention5.8k—~902Automated safety check: PassMITtoday
3

Iterate on FLA kernel performance with a frozen correctness gate, a measured baseline, and reproducible candidate comparisons.

fla-org/flash-linear-attention5.8k—~1.2kAutomated safety check: PassMITtoday
4

Profile and optimize FLA kernels on NVIDIA GPUs, with same-hardware benchmarks and targeted Nsight Compute analysis.

fla-org/flash-linear-attention5.8k—~818Automated safety check: PassMITtoday
5

Prepare or update an FLA pull request with focused scope, verified evidence, and the current repository template.

fla-org/flash-linear-attention5.8k—~856Automated safety check: PassMITtoday
6

Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

fla-org/flash-linear-attention5.8k—~1.6kAutomated safety check: PassMITtoday
7

Define supported inputs, numerical budgets, routing, and validation before designing an FLA kernel or numerical change.

fla-org/flash-linear-attention5.8k—~1.3kAutomated safety check: PassMITtoday
8

Change FLA backend registration, dispatch, or verifiers while preserving routing and call contracts.

fla-org/flash-linear-attention5.8k—~998Automated safety check: PassMITtoday
9

Modify or review KDA gates, chunk and recurrent kernels, backend routes, and their correctness coverage.

fla-org/flash-linear-attention5.8k—~780Automated safety check: PassMITtoday