Agent skill

B200 Flash Attention4 Planner

by mirage-project in mirage-project/mirage

A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

Apache-2.0Auto-check passedDatabases

Install B200 Flash Attention4 Planner

skills CLI
$ npx skills add mirage-project/mirage --skill b200-flash-attention4-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-flash-attention4-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-flash-attention4-planner .claude/skills/b200-flash-attention4-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-flash-attention4-planner
GitHub stars
2.5k
Token cost
~1.9k tokens
SKILL.md length
901 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

  • Works in 3 steps: "Implement a causal FlashAttention… → "Help me draw the FA4 S/P/O TMEM regions… → "Add GQA or LSE output to this attention…
  • The user wants to design
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Flash Attention4 Planner is an agent skill from mirage-project/mirage. Use when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles, causal mask, GQA, tile scheduling, or final normalization. Outputs the algorithm state, tile graph, barrier graph, and validation plan. Not for cases that only use off-the-shelf framework operators, where the full backward is not yet defined, or for ordinary dense GEMM.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in Databases, covering Database schema design. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user wants to design
  • Extend a FlashAttention-style forward kernel on B200/Blackwell
  • Involving the two MMAs QKᵀ and PV
  • Tile scheduling

Example prompts

  • “/b200-flash-attention4-planner”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. "Implement a causal FlashAttention forward pass on B200."
  2. "Help me draw the FA4 S/P/O TMEM regions and the barrier graph."
  3. "Add GQA or LSE output to this attention kernel."

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Flash Attention4 Planner loads about 1.9k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 901 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 901 words, ~1,935 tokens.

Download SKILL.mdSave it as .claude/skills/b200-flash-attention4-planner/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-flash-attention4-planner
description
Use when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles, causal mask, GQA, tile scheduling, or final normalization. Outputs the algorithm state, tile graph, barrier graph, and validation plan. Not for cases that only use off-the-shelf framework operators, where the full backward is not yet defined, or for ordinary dense GEMM.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S14; S6–S9; S13
tags
b200, blackwell
related_skills
b200-tcgen05-mma-contract-builder, b200-tmem-lifecycle-planner, b200-mbarrier-protocol-auditor, b200-gemm-optimization-ladder, b200-tma-pipeline-designer…
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 FlashAttention-4 Planner

R — Source evidence (Reading, paraphrased)

  • [S14] Attention is not the same MMA repeated; it is a score MMA and a value MMA with online softmax, masking, and rescaling in between.
  • [S14] The streaming state is row_max, row_sum, and O; when a new maximum appears, the old denominator and O must both be rescaled to the same basis.
  • [S14] S, P, and O mainly reside in TMEM; softmax/correction reads or modifies them in registers, then writes back to TMEM.
  • [S14] Multiple warpgroups divide the work of driving TMA, MMA, softmax, and correction/epilogue; the barrier graph proves every tile can be safely consumed and reused.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

FA4 design should first draw the "algorithm state machine" and the "hardware tile graph" before writing code:

  • the algorithm state machine guarantees the online softmax math is correct;
  • the tile graph states where Q/K/V/S/P/O live in GMEM, SMEM, TMEM, and registers;
  • the role graph states which warp/warpgroup issues TMA/MMA or does softmax;
  • the barrier graph states the visibility and reuse along score→softmax→value MMA→epilogue.

Any implementation that omits the rescale, treats P as the final normalized matrix, or misses the O slot release may be fast but numerically wrong.


A1 — Applications in the source (Past Application)

Streaming online softmax

For each K/V block:

  • Compute the score S = QKᵀ.
  • Update m_new.
  • Scale row_sum and the old O by the old/new max difference.
  • Compute the current numerator P and accumulate P V.
  • Only at the very end do O / row_sum.
Role example
  • One warp handles the TMA loads.
  • One warp issues the score/value MMAs.
  • Two warpgroups handle the softmax for the two-stage Q pipeline.
  • One warpgroup does the O correction and epilogue.
  • One warp handles the final TMA store.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "Implement a causal FlashAttention forward pass on B200."
  2. "Help me draw the FA4 S/P/O TMEM regions and the barrier graph."
  3. "Add GQA or LSE output to this attention kernel."
Language signals
  • "Implement a causal FlashAttention forward pass on B200."
  • "Help me draw the FA4 S/P/O TMEM regions and the barrier graph."
  • "Add GQA or LSE output to this attention kernel."
Distinction from adjacent skills

Difference from b200-gemm-optimization-ladder: FA4 has two MMAs with softmax/rescaling in between and cannot be treated as a single GEMM loop. Combine with b200-tmem-lifecycle-planner for the region budget.


Show full SKILL.md (484 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must execute the following process:

  1. Fix the math semantics
    • Q/K/V shapes, layout, head_dim, causal, GQA ratio, scale, the output, and whether LSE is needed.
  2. Define the streaming state
    • The initial values and per-block update formulas of row_max, row_sum, and O for each row.
    • Be explicit about using natural exp or the equivalent exp2 scaling; the reference must match.
  3. Define the two MMAs
    • score MMA: Q×Kᵀ→S.
    • value MMA: P×V→O.
    • Write out both tile shapes, dtypes, SMEM operands, and TMEM outputs.
  4. Plan the S/P/O TMEM
    • After S is ready, read it into registers for mask/softmax; write P back to TMEM; accumulate O in TMEM and rescale when necessary.
    • Give the regions, strides, stages, and safe-reuse conditions.
  5. Assign warp roles
    • TMA load, MMA issue, softmax stage 0/1, correction/epilogue, TMA store.
    • Check that the collective scopes are complete.
  6. Build the barrier graph
    • Q/K/V ready→score/value MMA.
    • S ready→softmax.
    • P ready + O safe→value MMA.
    • final O ready→epilogue→store.
    • Draw the scalar mailbox's full/empty protocol separately.
  7. Implement mask/GQA
    • causal tiles: handle fully-skipped blocks, fully-valid blocks, and diagonal-boundary masking separately.
    • GQA: make the Q-head-to-KV-head mapping, reuse, and scheduler coordinates explicit.
  8. Implement rescale and writeback
    • The O correction is a full TMEM→register→TMEM tile operation; it cannot be removed.
    • After the loop ends: O / row_sum, cast, SMEM staging, TMA store drain.
  9. Validate
    • Compare against a high-precision PyTorch reference.
    • Cover causal/noncausal, different sequence lengths, head_dim, GQA ratios, tail tiles, and extreme logits.
    • If extending to the training forward, verify the LSE definition and scale are consistent; backward requires a separate design.
Required outputs
  1. Conclusion: the current choice/diagnosis; do not use a vague "it could be any of them".
  2. Evidence or assumptions: which items come from user data, and which are hypotheses awaiting verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: correctness tests, boundary tests, and one falsifiable experiment.
  5. Risks and fallback: alternative paths when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • Only a stable implementation already provided by PyTorch/FlashAttention needs to be called.
  • The user asks for the full backward, but the saved intermediates and gradient algorithm are not yet defined.
Failure modes
  • Not rescaling the old row_sum/O.
  • Normalizing P too early, breaking the streaming accumulation.
  • Reusing the S/P/O TMEM regions before their consumers finish.
  • Using the same slow path for the causal mask on both full blocks and boundary blocks.
  • Wrong GQA head mapping.
Limitations
  • This skill centers on the forward structure from the book; training backward, dropout, variable-length packed sequences, and distributed attention require additional design.

  • depends-on: b200-tcgen05-mma-contract-builder, b200-tmem-lifecycle-planner, b200-mbarrier-protocol-auditor
  • contrasts-with: b200-gemm-optimization-ladder
  • composes-with: b200-tma-pipeline-designer, b200-warp-specialized-debugger, b200-kernel-roofline-triage

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-flash-attention4-planner of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Flash Attention4 Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Flash Attention4 Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Flash Attention4 Planner this skillmirage-project/mirage2.5k—~1.9kAutomated safety check: PassApache-2.0
SQL Optimization Patternsynulihao/AgentSkillOS61811 repos~3.3kAutomated safety check: PassNone
Datamodellmnimbalyst/nimbalyst1.9k—~713Automated safety check: PassMIT
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~2.7kAutomated safety check: NotesMIT
Sqlite Schema Designfastrepl/anarlog9.5k—~1.9kAutomated safety check: PassMIT
Jazz Schema Designgarden-co/classic-jazz2.5k—~3.2kAutomated safety check: PassMIT

Similar skills

  • SQL Optimization Patterns

    ynulihao/AgentSkillOS

    Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.

    618 GitHub starsUsed in 11 repos~3.3k tokens
    DatabasesAuto-check passed
  • Datamodellm

    nimbalyst/nimbalyst

    Create visual data models for database schemas using Nimbalyst's DataModelLM editor.

    1.9k GitHub stars~713 tokensUpdated today
    DatabasesAuto-check passed
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    DatabasesAuto-check: notes
  • Sqlite Schema Design

    fastrepl/anarlog

    Design or review schemas for crates/cloudsync using SQLite Sync constraints, not generic SQLite advice.

    9.5k GitHub stars~1.9k tokensUpdated today
    DatabasesAuto-check passed
  • Jazz Schema Design

    garden-co/classic-jazz

    Design and implement collaborative data schemas using the Jazz framework.

    2.5k GitHub stars~3.2k tokensUpdated 1 mo ago
    DatabasesAuto-check passed
  • DB Migrations

    kurealnum/dotfiles

    A skill your agent uses when generating or regenerating Drizzle migration files, changing database schema tables or columns, resolving migration sequence conflicts after rebase, reviewing migration…

    290 GitHub stars~820 tokensUpdated 6 mo ago
    DatabasesAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • V2 Model Support

    mirage-project/mirage

    End-to-end pipeline for adding or porting a model to MPK Runtime-V2 — from a compute-graph spec (shapes + draw.io graph + HF checkpoint + TP/EP plan) to a working multi-GPU demo.

    2.5k GitHub stars~5.3k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about B200 Flash Attention4 Planner

What does B200 Flash Attention4 Planner do?

A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…. B200 Flash Attention4 Planner is an agent skill from mirage-project/mirage. Use when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles, causal mask, GQA, tile scheduling, or final normalization.

When should I use B200 Flash Attention4 Planner?

B200 Flash Attention4 Planner fits situations like: the user wants to design; extend a FlashAttention-style forward kernel on B200/Blackwell; involving the two MMAs QKᵀ and PV; tile scheduling.

How do I install B200 Flash Attention4 Planner in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-flash-attention4-planner -a claude-code`. Or copy the skill folder (.claude/skills/b200-flash-attention4-planner in mirage-project/mirage) into .claude/skills/b200-flash-attention4-planner in your project. Claude Code loads it when a task matches its description.

How do I install B200 Flash Attention4 Planner in Codex?

Run `npx skills add mirage-project/mirage --skill b200-flash-attention4-planner -a codex`. Or copy the skill folder (.claude/skills/b200-flash-attention4-planner in mirage-project/mirage) into .agents/skills/b200-flash-attention4-planner in your project. Codex loads it when a task matches its description.

Can I use B200 Flash Attention4 Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-flash-attention4-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-flash-attention4-planner, .gemini/skills/b200-flash-attention4-planner, .github/skills/b200-flash-attention4-planner and .opencode/skills/b200-flash-attention4-planner in your project.

What does B200 Flash Attention4 Planner need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Flash Attention4 Planner is instructions for the agent only.

Does B200 Flash Attention4 Planner access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Flash Attention4 Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Flash Attention4 Planner use?

B200 Flash Attention4 Planner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Flash Attention4 Planner use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Flash Attention4 Planner?

Skills that share tags, products or a category with B200 Flash Attention4 Planner: SQL Optimization Patterns (ynulihao/AgentSkillOS, 618 stars), Datamodellm (nimbalyst/nimbalyst, 1.9k stars), Experiment Audit (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars) and Sqlite Schema Design (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Flash Attention4 Planner?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.