Agent skill

B200 Gemm Optimization Ladder

by mirage-project in mirage-project/mirage

A skill your agent uses when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM.

Apache-2.0Auto-check passed

Install B200 Gemm Optimization Ladder

skills CLI
$ npx skills add mirage-project/mirage --skill b200-gemm-optimization-ladder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-gemm-optimization-ladder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-gemm-optimization-ladder .claude/skills/b200-gemm-optimization-ladder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-gemm-optimization-ladder
GitHub stars
2.5k
Token cost
~1.9k tokens
SKILL.md length
851 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM.

  • Works in 9 steps: A single 128×128 output tile. → K-loop accumulation. → Multiple CTAs covering the full M/N. → …
  • The user wants to implement from scratch
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Gemm Optimization Ladder is an agent skill from mirage-project/mirage. Use when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM. Advances level by level along "correct single tile→K loop→spatial tiling→TMA→multi-stage pipeline→persistent→warp specialization→2-CTA cluster→multi-consumer", with a correctness and performance gate at every level. Not for cases that only want to call a mature BLAS and need no custom fusion/layout.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user wants to implement from scratch
  • Systematically optimize a B200/Blackwell GEMM

Example prompts

  • “/b200-gemm-optimization-ladder”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. A single 128×128 output tile.
  2. K-loop accumulation.
  3. Multiple CTAs covering the full M/N.
  4. TMA async load/store.
  5. PIPE_DEPTH=2 software pipeline.
  6. Persistent kernel + tile scheduler.
  7. Warp specialization.
  8. 2-CTA cluster.
  9. Multi-consumer warp specialization.

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Gemm Optimization Ladder loads about 1.9k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 851 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 851 words, ~1,910 tokens.

Download SKILL.mdSave it as .claude/skills/b200-gemm-optimization-ladder/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-gemm-optimization-ladder
description
Use when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM. Advances level by level along "correct single tile→K loop→spatial tiling→TMA→multi-stage pipeline→persistent→warp specialization→2-CTA cluster→multi-consumer", with a correctness and performance gate at every level. Not for cases that only want to call a mature BLAS and need no custom fusion/layout.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S11; S12; S13; S3
tags
b200, blackwell
related_skills
b200-kernel-roofline-triage, b200-scope-layout-dispatch, b200-tma-pipeline-designer, b200-mbarrier-protocol-auditor, b200-tcgen05-mma-contract-builder…
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 GEMM Optimization Ladder

R — Source evidence (Reading, paraphrased)

  • [S11] Start from a minimal correct single tile, then add K accumulation and multi-CTA spatial tiling one step at a time, avoiding debugging all the complexity at once.
  • [S12] First let TMA take over the regular tile copies, then use multi-stage SMEM and a persistent scheduler to reduce waiting.
  • [S13] Warp specialization assigns load, MMA, and writeback to different roles; 2-CTA cluster and multi-consumer further remove serial bottlenecks.
  • [S3] Every level should be driven by roofline and measured evidence; a more complex structure is not guaranteed to be faster.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

GEMM optimization is not generating the "final kernel" in one shot, but a regression-testable, bisectable upgrade path. Every level must simultaneously satisfy:

  • correct against the reference;
  • an explainable layout/synchronization contract;
  • the performance change is measured;
  • if it regresses, you know where the added resource or serialization point is.

The agent should keep the version at every level; jumping straight to a complex warp-specialized cluster kernel and then blindly guessing at bugs is forbidden.


A1 — Applications in the source (Past Application)

The nine-level route in the book
  1. A single 128×128 output tile.
  2. K-loop accumulation.
  3. Multiple CTAs covering the full M/N.
  4. TMA async load/store.
  5. PIPE_DEPTH=2 software pipeline.
  6. Persistent kernel + tile scheduler.
  7. Warp specialization.
  8. 2-CTA cluster.
  9. Multi-consumer warp specialization.

Every level keeps the same basic data path: GMEM→SMEM→tcgen05→TMEM→register/SMEM→GMEM, changing only concurrency and scheduling.


A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "Start from a correct GEMM and optimize it step by step into a high-performance B200 version."
  2. "My GEMM is at Step 5 now; should the next step be persistent or warp specialization?"
  3. "Write correctness and performance acceptance criteria for each level."
Language signals
  • "Start from a correct GEMM and optimize it step by step into a high-performance B200 version."
  • "My GEMM is at Step 5 now; should the next step be persistent or warp specialization?"
  • "Write correctness and performance acceptance criteria for each level."
Distinction from adjacent skills

Difference from b200-kernel-roofline-triage: this skill is the GEMM-specific implementation route; the roofline skill decides whether compute/overlap optimization should continue to be pursued. Combine with the individual specialized skills to complete the concrete stages.


Show full SKILL.md (449 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must execute the following process:

  1. Establish the baseline contract
    • Fix the math definition, layout, dtype, reference, timing framework, and representative shape set.
  2. Level 1: single-tile correct path
    • Synchronous copy, one MMA, TMEM readback, store.
    • Acceptance: element-wise correct on small matrices, with every address space explainable.
  3. Level 2: K-loop
    • Correctly handle the initial accumulator, per-K-tile accumulation, and the K tail.
  4. Level 3: spatial tiling
    • grid→M/N tile mapping, boundary masks, full matrix coverage.
  5. Level 4: TMA
    • descriptor/swizzle, load barrier, store drain; compare against the synchronous version bitwise/within tolerance.
  6. Level 5: multi-stage pipeline
    • prologue/steady/epilogue, stage/phase ledger; measure the actual overlap.
  7. Level 6: persistent scheduler
    • A fixed set of resident CTAs processes multiple tiles; check the tail and tile order.
  8. Level 7: warp specialization
    • producer/MMA/writeback roles, the four classes of handoff, warpgroup-scoped sync.
  9. Level 8: 2-CTA cluster
    • cluster tile, DSMEM/multicast, cta_group::2, remote barrier.
  10. Level 9: multi-consumer
  • Multiple consumers partition the N/M/output regions, ensuring the same staged operand feeds more compute without write conflicts.
  1. Per-level gating
  • Correctness: random shapes, misalignment, K=1/multi-tile, NaN/Inf, error thresholds.
  • Performance: compare against the previous level, the library baseline, and the roofline.
  • Resource: register/SMEM/TMEM, active clusters, spill.
  1. Stopping rule
  • If you are already close to the practical roofline or added complexity no longer brings stable gains, stop upgrading and keep the simpler version.
Required outputs
  1. Conclusion: the current choice/diagnosis; do not use a vague "it could be any of them".
  2. Evidence or assumptions: which items come from user data, and which are hypotheses awaiting verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: correctness tests, boundary tests, and one falsifiable experiment.
  5. Risks and fallback: alternative paths when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • cuBLASLt/CUTLASS already meets the need and there is no fusion, special dtype/layout, or research purpose.
  • There is no correctness reference or reliable timing framework.
Failure modes
  • Jumping multiple levels at once, making the source of an error impossible to localize.
  • Testing only one neat large shape, ignoring small shapes and tail tiles.
  • Treating warp specialization as a guaranteed speedup, ignoring the occupancy of the added roles and the sync.
Limitations
  • The book's route is based mainly on TIRx and specific example shapes; when porting to CUDA/CUTLASS/Triton, keep the principles rather than copying the APIs verbatim.

  • depends-on: b200-kernel-roofline-triage, b200-scope-layout-dispatch
  • contrasts-with: none
  • composes-with: b200-tma-pipeline-designer, b200-mbarrier-protocol-auditor, b200-tcgen05-mma-contract-builder, b200-cluster-persistent-scheduler, b200-warp-specialized-debugger

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-gemm-optimization-ladder of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Gemm Optimization Ladder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Gemm Optimization Ladder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Gemm Optimization Ladder this skillmirage-project/mirage2.5k—~1.9kAutomated safety check: PassApache-2.0
Implementsickn33/agentic-awesome-skills47k5 repos~306Automated safety check: PassMIT
Implementcodewhale-hq/Codewhale41k—~190Automated safety check: PassMIT
Incremental Implementationaddyosmani/agent-skills103k1 repos~2.3kAutomated safety check: PassMIT
Implementbestofjs/bestofjs3.1k18 repos~109Automated safety check: PassMIT
ImplementAutomattic/simplenote-android1.9k—~1.1kAutomated safety check: PassGPL-2.0

Similar skills

  • Implement

    sickn33/agentic-awesome-skills

    Implement a piece of work based on a PRD or set of issues. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 5 repos~306 tokens
    Product & Project ManagementAuto-check passed
  • Implement

    codewhale-hq/Codewhale

    Carry an authorized, defined request or approved plan through scoped edits and proportionate verification.

    41k GitHub stars~190 tokensUpdated today
    Auto-check passed
  • Incremental Implementation

    addyosmani/agent-skills

    Delivers a change in thin vertical slices, each implemented, tested, verified and committed before the next, using vertical, contract-first or risk-first slicing.

    103k GitHub starsUsed in 1 repo~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Implement

    bestofjs/bestofjs

    Implement a piece of work based on a spec or set of tickets.

    3.1k GitHub starsUsed in 18 repos~109 tokens
    Auto-check passed
  • Implement

    Automattic/simplenote-android

    End-to-end implementation workflow: plan, implement, verify, commit, and open a draft PR.

    1.9k GitHub stars~1.1k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Sparc Implement

    ruvnet/ruflo

    Run the SPARC Pseudocode and Architecture phases (2 and 3) — write algorithm pseudocode, design module boundaries and API contracts, then implement

    74k GitHub stars~1.3k tokensUpdated today
    Backend & APIsAuto-check: notes

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about B200 Gemm Optimization Ladder

What does B200 Gemm Optimization Ladder do?

A skill your agent uses when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM. B200 Gemm Optimization Ladder is an agent skill from mirage-project/mirage. Use when the user wants to implement from scratch, port, or systematically optimize a B200/Blackwell GEMM.

When should I use B200 Gemm Optimization Ladder?

B200 Gemm Optimization Ladder fits situations like: the user wants to implement from scratch; systematically optimize a B200/Blackwell GEMM.

How do I install B200 Gemm Optimization Ladder in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-gemm-optimization-ladder -a claude-code`. Or copy the skill folder (.claude/skills/b200-gemm-optimization-ladder in mirage-project/mirage) into .claude/skills/b200-gemm-optimization-ladder in your project. Claude Code loads it when a task matches its description.

How do I install B200 Gemm Optimization Ladder in Codex?

Run `npx skills add mirage-project/mirage --skill b200-gemm-optimization-ladder -a codex`. Or copy the skill folder (.claude/skills/b200-gemm-optimization-ladder in mirage-project/mirage) into .agents/skills/b200-gemm-optimization-ladder in your project. Codex loads it when a task matches its description.

Can I use B200 Gemm Optimization Ladder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-gemm-optimization-ladder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-gemm-optimization-ladder, .gemini/skills/b200-gemm-optimization-ladder, .github/skills/b200-gemm-optimization-ladder and .opencode/skills/b200-gemm-optimization-ladder in your project.

What does B200 Gemm Optimization Ladder need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Gemm Optimization Ladder is instructions for the agent only.

Does B200 Gemm Optimization Ladder access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Gemm Optimization Ladder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Gemm Optimization Ladder use?

B200 Gemm Optimization Ladder is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Gemm Optimization Ladder use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Gemm Optimization Ladder?

Skills that share tags, products or a category with B200 Gemm Optimization Ladder: Implement (sickn33/agentic-awesome-skills, 47k stars), Implement (codewhale-hq/Codewhale, 41k stars), Incremental Implementation (addyosmani/agent-skills, 103k stars) and Implement (bestofjs/bestofjs, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Gemm Optimization Ladder?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,543 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.