Agent skill

Flash Attention

by AlexAI-MCP in AlexAI-MCP/hermes-CCC

Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.

MITAuto-check passedAI & LLM Engineering

Install Flash Attention

skills CLI
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/flash-attention .claude/skills/flash-attention && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flash-attention
GitHub stars
135
Token cost
~1.2k tokens
SKILL.md length
183 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.

  • Tasks that involve Deep learning
  • SKILL.md covers When to Use, Option 1: PyTorch Native SDPA…, Option 2: flash-attn Library… and Option 3: Enable in…, plus 5 more sections
  • Calls pip and python

What it does

Flash Attention is an agent skill from AlexAI-MCP/hermes-CCC. Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with CUDA and PyTorch. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/flash-attention”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flash Attention loads about 1.2k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 183 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 183 words, ~1,208 tokens.

Download SKILL.mdSave it as .claude/skills/flash-attention/SKILL.md (or your agent's skills folder).
name
flash-attention
description
Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.
version
1.0.0
author
hermes-CCC (ported from Hermes Agent by NousResearch)
license
MIT

Flash Attention — Fast Memory-Efficient Attention

Flash Attention provides 2-4x training speedup and 10-20x memory reduction by replacing the standard O(N²) attention with an IO-aware tiling algorithm.

When to Use

  • Sequences longer than 512 tokens → always use Flash Attention
  • GPU OOM during training → Flash Attention is usually the fix
  • Maximizing training throughput → 2-4x faster than standard attention
  • H100/A100 with FP8/BF16 → critical for efficiency

Option 1: PyTorch Native SDPA (Easiest — PyTorch 2.2+)

python
import torch
import torch.nn.functional as F

# PyTorch automatically uses Flash Attention if available
q = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)
k = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)
v = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)

# This automatically dispatches to Flash Attention on compatible hardware
output = F.scaled_dot_product_attention(q, k, v, dropout_p=0.0, is_causal=True)

# Check which kernel is being used
with torch.backends.cuda.sdp_kernel(
    enable_flash=True, enable_math=False, enable_mem_efficient=False
):
    output = F.scaled_dot_product_attention(q, k, v, is_causal=True)

Option 2: flash-attn Library (Maximum Performance)

bash
# Requires: CUDA toolkit, torch, ninja
pip install flash-attn --no-build-isolation

# If build fails:
pip install packaging ninja
pip install flash-attn --no-build-isolation --no-cache-dir
python
from flash_attn import flash_attn_func, flash_attn_varlen_func

# Basic usage: q,k,v shape: (batch, seqlen, nheads, headdim)
q = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)
k = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)
v = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)

output = flash_attn_func(q, k, v, dropout_p=0.0, causal=True)

Option 3: Enable in HuggingFace Transformers

python
from transformers import AutoModelForCausalLM

# Automatic Flash Attention 2 (recommended)
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    attn_implementation="flash_attention_2",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

# Or: eager (standard), sdpa (PyTorch SDPA)
model = AutoModelForCausalLM.from_pretrained(
    "...",
    attn_implementation="sdpa",  # no install needed
)

Hardware Requirements

GPUFP16BF16FP8
A100✅✅❌
H100✅✅✅
A10/A30✅✅❌
RTX 3090/4090✅✅❌
V100✅❌❌
T4✅❌❌

Requirements:

  • CUDA 11.6+
  • PyTorch 2.0+
  • Ampere GPU or newer (RTX 30 series, A100, H100) for best performance

Sliding Window Attention (Long Context)

python
from flash_attn.flash_attn_interface import flash_attn_varlen_func

# For sequences longer than GPU memory allows
# Use sliding window to limit attention span
output = flash_attn_func(
    q, k, v,
    causal=True,
    window_size=(512, 0)  # attend to last 512 tokens only
)

Memory Savings

Standard attention: O(N²) memory Flash Attention: O(N) memory (recomputes on backward pass)

Sequence length 4096:  ~6x memory savings
Sequence length 8192:  ~15x memory savings
Sequence length 32768: ~50x+ memory savings

Benchmarking

python
import time, torch

def bench(fn, warmup=3, reps=10):
    for _ in range(warmup):
        fn()
    torch.cuda.synchronize()
    t0 = time.time()
    for _ in range(reps):
        fn()
    torch.cuda.synchronize()
    return (time.time() - t0) / reps * 1000  # ms

q = torch.randn(4, 2048, 16, 64, device="cuda", dtype=torch.float16)
k, v = q.clone(), q.clone()

standard_ms = bench(lambda: F.scaled_dot_product_attention(q, k, v))
flash_ms = bench(lambda: flash_attn_func(q, k, v, causal=True))
print(f"Standard: {standard_ms:.1f}ms | Flash: {flash_ms:.1f}ms | Speedup: {standard_ms/flash_ms:.1f}x")

Common Issues

Build fails: Ensure CUDA toolkit matches PyTorch CUDA version (python -c "import torch; print(torch.version.cuda)")

Wrong dtype: Flash Attention requires FP16 or BF16, not FP32

OOM despite Flash Attention: Use gradient checkpointing additionally: model.gradient_checkpointing_enable()

Not faster on short sequences: Flash Attention shines at >512 tokens; below that overhead can dominate

© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/flash-attention of AlexAI-MCP/hermes-CCC.

Open the folder on GitHubat commit 8107e89

Compare with similar skills

Flash Attention next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flash Attention compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flash Attention this skillAlexAI-MCP/hermes-CCC135—~1.2kAutomated safety check: PassMIT
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Validating Pytorch Custom Opsmeta-pytorch/attention-gym1.3k—~6.4kAutomated safety check: PassBSD-3-Clause
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT

Similar skills

  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Validating Pytorch Custom Ops

    meta-pytorch/attention-gym

    Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.

    1.3k GitHub stars~6.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from AlexAI-MCP/hermes-CCC

All 44 skills in this repo
  • GitHub Code Review

    AlexAI-MCP/hermes-CCC

    Review GitHub pull requests with a findings-first engineering mindset.

    135 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • GitHub PR Workflow

    AlexAI-MCP/hermes-CCC

    Run a disciplined GitHub pull request workflow from branch creation through merge.

    135 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Memory

    AlexAI-MCP/hermes-CCC

    Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Route

    AlexAI-MCP/hermes-CCC

    Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Skill

    AlexAI-MCP/hermes-CCC

    Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Traj

    AlexAI-MCP/hermes-CCC

    Capture Claude Code interaction trajectories in training-friendly formats.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Flash Attention

What does Flash Attention do?

Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs. Flash Attention is an agent skill from AlexAI-MCP/hermes-CCC. Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.

When should I use Flash Attention?

Flash Attention fits situations like: tasks that involve Deep learning.

How do I install Flash Attention in Claude Code?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a claude-code`. Or copy the skill folder (skills/flash-attention in AlexAI-MCP/hermes-CCC) into .claude/skills/flash-attention in your project. Claude Code loads it when a task matches its description.

How do I install Flash Attention in Codex?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a codex`. Or copy the skill folder (skills/flash-attention in AlexAI-MCP/hermes-CCC) into .agents/skills/flash-attention in your project. Codex loads it when a task matches its description.

Can I use Flash Attention in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flash-attention, .gemini/skills/flash-attention, .github/skills/flash-attention and .opencode/skills/flash-attention in your project.

What does Flash Attention need to run?

Going by SKILL.md and its folder, Flash Attention needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3.

Does Flash Attention access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Flash Attention safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flash Attention use?

Flash Attention is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flash Attention use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flash Attention?

Skills that share tags, products or a category with Flash Attention: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Validating Pytorch Custom Ops (meta-pytorch/attention-gym, 1.3k stars) and Metal Kernel (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flash Attention?

AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.

Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.