Cuda Index Width
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/flash-attention .claude/skills/flash-attention && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .claude/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attentionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/flash-attention .agents/skills/flash-attention && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .agents/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/flash-attention .cursor/skills/flash-attention && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .cursor/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/AlexAI-MCP/hermes-CCC.git --path skills/flash-attention--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/flash-attention .gemini/skills/flash-attention && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .gemini/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attentionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/flash-attention .github/skills/flash-attention && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .github/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install AlexAI-MCP/hermes-CCC flash-attention --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/flash-attention .opencode/skills/flash-attention && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "flash-attention" agent skill from https://github.com/AlexAI-MCP/hermes-CCC/tree/master/skills/flash-attention into .opencode/skills/flash-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attention", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
flash-attentionOptimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.
Flash Attention is an agent skill from AlexAI-MCP/hermes-CCC. Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.
Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning. It works with CUDA and PyTorch. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.
Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flash Attention loads about 1.2k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 183 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 183 words, ~1,208 tokens.
.claude/skills/flash-attention/SKILL.md (or your agent's skills folder).Flash Attention provides 2-4x training speedup and 10-20x memory reduction by replacing the standard O(N²) attention with an IO-aware tiling algorithm.
import torch
import torch.nn.functional as F
# PyTorch automatically uses Flash Attention if available
q = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)
k = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)
v = torch.randn(2, 8, 512, 64, device="cuda", dtype=torch.float16)
# This automatically dispatches to Flash Attention on compatible hardware
output = F.scaled_dot_product_attention(q, k, v, dropout_p=0.0, is_causal=True)
# Check which kernel is being used
with torch.backends.cuda.sdp_kernel(
enable_flash=True, enable_math=False, enable_mem_efficient=False
):
output = F.scaled_dot_product_attention(q, k, v, is_causal=True)# Requires: CUDA toolkit, torch, ninja
pip install flash-attn --no-build-isolation
# If build fails:
pip install packaging ninja
pip install flash-attn --no-build-isolation --no-cache-dirfrom flash_attn import flash_attn_func, flash_attn_varlen_func
# Basic usage: q,k,v shape: (batch, seqlen, nheads, headdim)
q = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)
k = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)
v = torch.randn(2, 512, 8, 64, device="cuda", dtype=torch.float16)
output = flash_attn_func(q, k, v, dropout_p=0.0, causal=True)from transformers import AutoModelForCausalLM
# Automatic Flash Attention 2 (recommended)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct",
attn_implementation="flash_attention_2",
torch_dtype=torch.bfloat16,
device_map="auto",
)
# Or: eager (standard), sdpa (PyTorch SDPA)
model = AutoModelForCausalLM.from_pretrained(
"...",
attn_implementation="sdpa", # no install needed
)| GPU | FP16 | BF16 | FP8 |
|---|---|---|---|
| A100 | ✅ | ✅ | ❌ |
| H100 | ✅ | ✅ | ✅ |
| A10/A30 | ✅ | ✅ | ❌ |
| RTX 3090/4090 | ✅ | ✅ | ❌ |
| V100 | ✅ | ❌ | ❌ |
| T4 | ✅ | ❌ | ❌ |
Requirements:
from flash_attn.flash_attn_interface import flash_attn_varlen_func
# For sequences longer than GPU memory allows
# Use sliding window to limit attention span
output = flash_attn_func(
q, k, v,
causal=True,
window_size=(512, 0) # attend to last 512 tokens only
)Standard attention: O(N²) memory Flash Attention: O(N) memory (recomputes on backward pass)
Sequence length 4096: ~6x memory savings
Sequence length 8192: ~15x memory savings
Sequence length 32768: ~50x+ memory savingsimport time, torch
def bench(fn, warmup=3, reps=10):
for _ in range(warmup):
fn()
torch.cuda.synchronize()
t0 = time.time()
for _ in range(reps):
fn()
torch.cuda.synchronize()
return (time.time() - t0) / reps * 1000 # ms
q = torch.randn(4, 2048, 16, 64, device="cuda", dtype=torch.float16)
k, v = q.clone(), q.clone()
standard_ms = bench(lambda: F.scaled_dot_product_attention(q, k, v))
flash_ms = bench(lambda: flash_attn_func(q, k, v, causal=True))
print(f"Standard: {standard_ms:.1f}ms | Flash: {flash_ms:.1f}ms | Speedup: {standard_ms/flash_ms:.1f}x")Build fails: Ensure CUDA toolkit matches PyTorch CUDA version (python -c "import torch; print(torch.version.cuda)")
Wrong dtype: Flash Attention requires FP16 or BF16, not FP32
OOM despite Flash Attention: Use gradient checkpointing additionally: model.gradient_checkpointing_enable()
Not faster on short sequences: Flash Attention shines at >512 tokens; below that overhead can dominate
© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/flash-attention of AlexAI-MCP/hermes-CCC.
Open the folder on GitHubat commit 8107e89
Flash Attention next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flash Attention this skillAlexAI-MCP/hermes-CCC | 135 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Validating Pytorch Custom Opsmeta-pytorch/attention-gym | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT |
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
meta-pytorch/attention-gym
Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
AlexAI-MCP/hermes-CCC
Review GitHub pull requests with a findings-first engineering mindset.
AlexAI-MCP/hermes-CCC
Run a disciplined GitHub pull request workflow from branch creation through merge.
AlexAI-MCP/hermes-CCC
Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.
AlexAI-MCP/hermes-CCC
Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.
AlexAI-MCP/hermes-CCC
Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.
AlexAI-MCP/hermes-CCC
Capture Claude Code interaction trajectories in training-friendly formats.
Categories
Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs. Flash Attention is an agent skill from AlexAI-MCP/hermes-CCC. Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.
Flash Attention fits situations like: tasks that involve Deep learning.
Run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a claude-code`. Or copy the skill folder (skills/flash-attention in AlexAI-MCP/hermes-CCC) into .claude/skills/flash-attention in your project. Claude Code loads it when a task matches its description.
Run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a codex`. Or copy the skill folder (skills/flash-attention in AlexAI-MCP/hermes-CCC) into .agents/skills/flash-attention in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill flash-attention -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flash-attention, .gemini/skills/flash-attention, .github/skills/flash-attention and .opencode/skills/flash-attention in your project.
Going by SKILL.md and its folder, Flash Attention needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Flash Attention is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flash Attention: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Validating Pytorch Custom Ops (meta-pytorch/attention-gym, 1.3k stars) and Metal Kernel (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.
Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.