Agent skill

Fla Nvidia Performance

by fla-org in fla-org/flash-linear-attention

Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo.

MITAuto-check passedAI & LLM Engineering

Install Fla Nvidia Performance

skills CLI
$ npx skills add fla-org/flash-linear-attention --skill fla-nvidia-performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fla-org/flash-linear-attention fla-nvidia-performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fla-org/flash-linear-attention.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fla-nvidia-performance .claude/skills/fla-nvidia-performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fla-nvidia-performance
GitHub stars
5.8k
Token cost
~1.2k tokens
SKILL.md length
445 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo.

  • Works in 4 steps: Before / after benchmark → Profiling when needed → Workload coverage → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Hardware baseline, Detailed NCU workflow, Day-to-day development and Before opening a PR…, plus 2 more sections
  • Calls python

What it does

Fla Nvidia Performance is an agent skill from fla-org/flash-linear-attention. Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo. Covers profiling workflow, hardware baselines, and PR-ready performance evidence requirements. Uses an installed ncu-report-skill when a task needs detailed Nsight Compute collection and diagnosis.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: 🚀 Efficient implementations for emerging model architectures. The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “Use the fla-nvidia-performance skill to guideline for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo”
  • “/fla-nvidia-performance”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Before / after benchmark
  2. Profiling when needed
  3. Workload coverage
  4. Conclusion and risk

What it can do on your machine

Read from SKILL.md and the folder at commit 72ac946. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fla Nvidia Performance loads about 1.2k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fla-org/flash-linear-attention at commit 72ac946, republished under its MIT licence (© fla-org). 445 words, ~1,155 tokens.

Download SKILL.mdSave it as .claude/skills/fla-nvidia-performance/SKILL.md (or your agent's skills folder).
name
fla-nvidia-performance
description
Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo. Covers profiling workflow, hardware baselines, and PR-ready performance evidence requirements. Uses an installed ncu-report-skill when a task needs detailed Nsight Compute collection and diagnosis.

FLA NVIDIA Performance Skill

Use this skill when working on Triton, Gluon, TileLang, CUDA, or other NVIDIA GPU kernel optimizations, backend tuning, or any change that could affect throughput or latency in fla/ops/ or related modules.

Hardware baseline

  • Minimum effective baseline: datacenter NVIDIA GPUs with sm_90 or newer; H100 / H20 are accepted.
  • Preferred targets: datacenter NVIDIA GPUs with sm_100 or sm_103.
  • Reference only: A100 (sm_80), pre-sm_80 GPUs, and all consumer cards (including sm_86, sm_89, and sm_120). Do not use these numbers as the only PR performance conclusion.

Detailed NCU workflow

This repo intentionally does not vendor ncu-report-skill, add it as a submodule, or auto-clone it during agent work.

If a user-level ncu-report-skill is available, use it for:

  • Nsight Compute collection details (full, source, PM sampling, source counters);
  • sm_100 / sm_103 metric-name caveats;
  • report parsing helpers and diagnosis playbooks;
  • the final profiling report structure.

If it is not available, use the minimal NCU commands in this skill. Report any missing profiling evidence and its technical limitation; the availability of a helper skill does not belong in the PR description. Do not create untracked external clones inside this repo unless the user explicitly asks.

Day-to-day development

  • You are not required to run NCU for every incremental change.
  • Quick sanity checks with benchmark_training_throughput.py or benchmark_generation.py are enough to catch large regressions during development.
  • Prefer dense workloads for quick iteration; varlen workloads are checked before PR.
Show full SKILL.md (214 more words)Show less

Before opening a PR (performance evidence)

For NVIDIA performance-optimization PRs, collect the evidence below. Correctness-only kernel changes still need the tests and same-hardware benchmarks required by CONTRIBUTING.md; detailed profiling is needed when it explains a performance claim or unresolved regression.

  1. Before / after benchmark

    • Run the same benchmark script with the same workload on the same hardware.
    • Report throughput (tokens/s or iters/s) and, if relevant, peak memory.
  2. Profiling when needed

    • Use NCU when explaining a bottleneck, performance claim, or unresolved regression needs hardware metrics. It is not a mandatory check for every PR.
    • When collecting NCU evidence, run ncu with --set full and --set source for a representative changed kernel when Nsight Compute is available.
    • Capture the .ncu-rep locally; do not commit it to the repo.
    • In the PR description, paste a short summary of key metrics (e.g., memory throughput %, SOL, occupancy, top hot instructions).
  3. Workload coverage

    • Include at least one dense workload.
    • Include at least one variable-length workload if the op supports it.
  4. Conclusion and risk

    • State whether the change is an improvement, neutral, or a known trade-off.
    • Flag any backend or shape that regressed and explain why.

Profile artifact layout

Store local profile artifacts under:

text
profile/<run_name>/

For example:

text
profile/kda_chunk_bwd_20250603/
  ├── REPORT.md
  ├── reports/
  │   ├── full_<tag>.ncu-rep
  │   └── source_<tag>.ncu-rep
  └── analysis/

Keep .ncu-rep, .nsys-rep, and raw logs out of git.

Quick commands reference

bash
# Op microbenchmark
python -m benchmarks.ops.run --op chunk_kda --modes fwd

# Model training benchmark
python benchmarks/benchmark_training_throughput.py \
  --name kda --batch_size 2 --seq_len 8192

# Varlen training benchmark (if supported by the model/op path)
python benchmarks/benchmark_training_throughput.py \
  --name kda --batch_size 2 --seq_len 8192 --varlen

# NCU full profile
ncu --set full --section PmSampling --section PmSampling_WarpStates \
  -k "regex:<kernel_regex>" -c 1 \
  -o profile/<run_name>/reports/full_<tag> \
  python -m benchmarks.ops.run --op chunk_kda --modes fwd

# NCU source profile (for instruction-level analysis)
ncu --set source --section SourceCounters \
  -k "regex:<kernel_regex>" -c 1 \
  -o profile/<run_name>/reports/source_<tag> \
  python -m benchmarks.ops.run --op chunk_kda --modes fwd

© fla-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/fla-nvidia-performance of fla-org/flash-linear-attention.

Open the folder on GitHubat commit 72ac946

Compare with similar skills

Fla Nvidia Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fla Nvidia Performance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fla Nvidia Performance this skillfla-org/flash-linear-attention5.8k—~1.2kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent144—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    144 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    146 GitHub stars~547 tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed

More from fla-org/flash-linear-attention

All 9 skills in this repo
  • Fla Ascend Performance

    fla-org/flash-linear-attention

    Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.

    5.8k GitHub stars~6.3k tokensUpdated today
    Auto-check passed
  • Fla Optimization Loop

    fla-org/flash-linear-attention

    Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

    5.8k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Fla Correctness Coverage

    fla-org/flash-linear-attention

    Guidelines for kernel correctness testing and coverage in fla/ops/ and related modules, including common Triton grid/addressing pitfalls.

    5.8k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Fla Design Coverage

    fla-org/flash-linear-attention

    Contract-first design and coverage discipline for FLA kernel and numerical changes.

    5.8k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Fla Dispatch Backends

    fla-org/flash-linear-attention

    Workflow for FLA backend dispatch decorators and backend implementations.

    5.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Fla Nvidia Performance

What does Fla Nvidia Performance do?

Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo. Fla Nvidia Performance is an agent skill from fla-org/flash-linear-attention. Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo.

When should I use Fla Nvidia Performance?

Fla Nvidia Performance fits situations like: AI & LLM Engineering work in your project.

How do I install Fla Nvidia Performance in Claude Code?

Run `npx skills add fla-org/flash-linear-attention --skill fla-nvidia-performance -a claude-code`. Or copy the skill folder (.agents/skills/fla-nvidia-performance in fla-org/flash-linear-attention) into .claude/skills/fla-nvidia-performance in your project. Claude Code loads it when a task matches its description.

How do I install Fla Nvidia Performance in Codex?

Run `npx skills add fla-org/flash-linear-attention --skill fla-nvidia-performance -a codex`. Or copy the skill folder (.agents/skills/fla-nvidia-performance in fla-org/flash-linear-attention) into .agents/skills/fla-nvidia-performance in your project. Codex loads it when a task matches its description.

Can I use Fla Nvidia Performance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fla-org/flash-linear-attention --skill fla-nvidia-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fla-nvidia-performance, .gemini/skills/fla-nvidia-performance, .github/skills/fla-nvidia-performance and .opencode/skills/fla-nvidia-performance in your project.

What does Fla Nvidia Performance need to run?

Going by SKILL.md and its folder, Fla Nvidia Performance needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Fla Nvidia Performance access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fla Nvidia Performance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fla Nvidia Performance use?

Fla Nvidia Performance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fla Nvidia Performance use?

About 1.2k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fla Nvidia Performance?

Skills that share tags, products or a category with Fla Nvidia Performance: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fla Nvidia Performance?

fla-org (a GitHub organization) maintains it in fla-org/flash-linear-attention, which has 5,831 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: fla-org/flash-linear-attention on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.