Agent skill

B200 Scope Layout Dispatch

by mirage-project in mirage-project/mirage

A skill your agent uses when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and…

Apache-2.0Auto-check passed

Install B200 Scope Layout Dispatch

skills CLI
$ npx skills add mirage-project/mirage --skill b200-scope-layout-dispatch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-scope-layout-dispatch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-scope-layout-dispatch .claude/skills/b200-scope-layout-dispatch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-scope-layout-dispatch
GitHub stars
2.5k
Token cost
~1.9k tokens
SKILL.md length
876 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and…

  • Works in 3 steps: "Map this fused op onto a B200 kernel —… → "Why does this collective deadlock when… → "Help me draw the tile lifetime across…
  • The user wants to map an ML operator onto a B200/Blackwell kernel
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Scope Layout Dispatch is an agent skill from mirage-project/mirage. Use when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and whether to invoke thread code or TMA/tcgen05. Outputs a complete contract for scope, layout, dispatch, and handoff. Not for high-level model architecture design alone or pure hardware spec queries.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user wants to map an ML operator onto a B200/Blackwell kernel
  • Review which threads should execute a given tile primitive
  • Where the data should live
  • Whether to invoke thread code

Example prompts

  • “/b200-scope-layout-dispatch”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. "Map this fused op onto a B200 kernel — give me the scope/layout/dispatch first."
  2. "Why does this collective deadlock when placed inside a warp branch?"
  3. "Help me draw the tile lifetime across GMEM→SMEM→TMEM→register."

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Scope Layout Dispatch loads about 1.9k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 876 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 876 words, ~1,890 tokens.

Download SKILL.mdSave it as .claude/skills/b200-scope-layout-dispatch/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-scope-layout-dispatch
description
Use when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and whether to invoke thread code or TMA/tcgen05. Outputs a complete contract for scope, layout, dispatch, and handoff. Not for high-level model architecture design alone or pure hardware spec queries.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S2; S4; S6; S7; S8
tags
b200, blackwell
related_skills
b200-kernel-roofline-triage, b200-layout-contract-auditor, b200-tma-pipeline-designer, b200-tcgen05-mma-contract-builder, b200-mbarrier-protocol-auditor
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 Scope–Layout–Dispatch Designer

R — Source evidence (Reading, paraphrased)

  • [S2] Blackwell's execution hierarchy runs from thread, warp, warpgroup, CTA, and cluster up to grid; different hardware operations have different natural scopes.
  • [S2] TMA is usually issued by a single thread, TMEM readback is done cooperatively by a warpgroup, and a tcgen05 MMA is issued by one elected thread on behalf of the participating group.
  • [S4] The layout determines the mapping from logical index to physical location, thread/register ownership, and bank/coalescing behavior.
  • [S6–S9] Asynchronous hardware operations must have an explicit handoff; program order alone does not prove the data is ready.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

Treat kernel design as four mutually constraining tables, not one blob of mixed code:

  • Scope: who participates in or issues this operation? A lane, a warp, a warpgroup, a CTA, or a two-CTA cluster?
  • Layout: where does each element of the logical tile physically live? GMEM/SMEM/TMEM/register, and which lane/warp/CTA owns it?
  • Dispatch: which hardware path performs the same logical action — plain thread instructions, TMA, tcgen05, shuffle, or a collective?
  • Handoff: when does the producer prove the result is visible, what does the consumer wait on, and when is the storage allowed to be reused?

Whenever one table changes, the other three must be re-checked.


A1 — Applications in the source (Past Application)

Case 1: TMA-loading an operand tile
  • Scope: one designated thread issues it; consumers within the CTA use it.
  • Layout: a logical GMEM rectangle mapped to a swizzled SMEM tile.
  • Dispatch: TMA descriptor + async copy.
  • Handoff: a byte-count mbarrier.
Case 2: Blackwell MMA and epilogue
  • Scope: one elected thread issues it; a warpgroup/CTA or 2-CTA group participates.
  • Layout: A/B in SMEM, the accumulator in TMEM, the epilogue reading back into registers.
  • Dispatch: tcgen05.mma + tcgen05.ld.
  • Handoff: MMA commit to a barrier; the epilogue waits for the result to complete.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "Map this fused op onto a B200 kernel — give me the scope/layout/dispatch first."
  2. "Why does this collective deadlock when placed inside a warp branch?"
  3. "Help me draw the tile lifetime across GMEM→SMEM→TMEM→register."
Language signals
  • "Map this fused op onto a B200 kernel — give me the scope/layout/dispatch first."
  • "Why does this collective deadlock when placed inside a warp branch?"
  • "Help me draw the tile lifetime across GMEM→SMEM→TMEM→register."
Distinction from adjacent skills

Versus b200-layout-contract-auditor: this skill does the overall design; the layout auditor specifically verifies addresses, swizzle, banks, and Tensor Core contracts. Versus b200-mbarrier-protocol-auditor: this skill defines the handoffs; the latter verifies the protocol barrier by barrier.


Show full SKILL.md (430 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must follow this procedure:

  1. Decompose the operator graph
    • Break the logical operator into tile load, transform, MMA/reduction, epilogue, and store.
    • Done criterion: each node has exactly one primary output tile or scalar.
  2. Choose a scope for each node
    • Record the "participation scope" and the "issue scope"; the two may differ.
    • Check that collectives execute over their full participation scope.
  3. Draw data residency and ownership
    • For each tile write down: shape, dtype, storage space, layout, owner, live range.
    • Mark every ownership change; an ownership change usually implies a real data movement or a collective.
  4. Choose the dispatch
    • Regular rectangular GMEM↔SMEM: evaluate TMA first.
    • Tensor Core tiles: evaluate tcgen05 and its supported shapes/dtypes.
    • Small/irregular operations: plain threads, vectorized load/store, shuffle, or CUDA cores.
  5. Define the handoffs
    • For each asynchronous edge write down the producer, consumer, signal, phase, fence/drain, and the point at which reuse is allowed.
  6. Produce the contract table
    • Columns: operation / logical tile / scope / storage-layout / dispatch / readiness / release.
  7. Run a consistency audit
    • The TMA descriptor, SMEM layout, and MMA expectations agree.
    • The TMEM mapping agrees with the tcgen05.ld readback.
    • No CTA-wide collective is placed inside a partial-thread branch.
  8. Give the minimal implementation order
    • Build a synchronous correct version first, then replace pieces one by one with asynchronous/specialized paths.
Required outputs
  1. Conclusion: the current choice/diagnosis, never a vague "we may need to look at everything".
  2. Evidence or assumptions: which items come from user data and which are assumptions pending verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: a correctness test, a boundary test, and one falsifiable experiment.
  5. Risks and fallback: the alternative path when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • The user only needs to call cuBLAS/cuDNN and has no custom-kernel requirement.
  • The algorithm itself is still undecided and even the tile-level data dependencies cannot be described.
Failure modes
  • Mistaking "who issues the instruction" for "who participates in the result".
  • Drawing only logical tensor shapes and not physical ownership.
  • Changing the dispatch while keeping the old synchronization protocol.
Limitations
  • The framework guarantees the design is auditable, not that the optimal tile falls out automatically; you still need to compile, run correctness tests, and measure performance.

  • depends-on: b200-kernel-roofline-triage
  • contrasts-with: none
  • composes-with: b200-layout-contract-auditor, b200-tma-pipeline-designer, b200-tcgen05-mma-contract-builder, b200-mbarrier-protocol-auditor

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-scope-layout-dispatch of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Scope Layout Dispatch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Scope Layout Dispatch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Scope Layout Dispatch this skillmirage-project/mirage2.5k—~1.9kAutomated safety check: PassApache-2.0
Token Mapnexu-io/open-design100k—~1.4kAutomated safety check: PassApache-2.0
Operational Home Layoutkunchenguid/firstmate7.7k—~6kAutomated safety check: NotesMIT
Maps Geographyasgeirtj/system_prompts_leaks69k—~717Automated safety check: PassCC0-1.0
It Operationsdavila7/claude-code-templates32k1 repos~3.7kAutomated safety check: PassMIT
Tabler Page Layoutstabler/tabler42k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Token Map

    nexu-io/open-design

    Map an extracted Figma / source-code token bag onto the active OD design system, producing a deterministic mapping the generate stage can consume.

    100k GitHub stars~1.4k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Operational Home Layout

    kunchenguid/firstmate

    Load when locating, interpreting, or changing Firstmate home, config, data, state, project, or generated runtime paths.

    7.7k GitHub stars~6k tokensUpdated today
    Auto-check: notes
  • Maps Geography

    asgeirtj/system_prompts_leaks

    Accurate maps from real geo data — use for any map, or whenever geography would make a good graphic for a deliverable

    69k GitHub stars~717 tokensUpdated yesterday
    Auto-check passed
  • It Operations

    davila7/claude-code-templates

    Manages IT infrastructure, monitoring, incident response, and service reliability.

    32k GitHub starsUsed in 1 repo~3.7k tokens
    DevOps & CloudAuto-check passed
  • Tabler Page Layouts

    tabler/tabler

    Picks and configures the right layout for a Tabler preview page and covers changing or adding layouts in shared/layouts, including DefaultLayout props and the page-header slot.

    42k GitHub stars~1.5k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Operator approval contract with internal filing notices for agent-drafted outbound messages, hashed drafts, epoch-keyed decisions, durable delivery claims and receipts, and a pre-draft baseline gate.

    275k GitHub stars~3.3k tokensUpdated 3 days ago
    Auto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed

Questions about B200 Scope Layout Dispatch

What does B200 Scope Layout Dispatch do?

A skill your agent uses when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and…. B200 Scope Layout Dispatch is an agent skill from mirage-project/mirage. Use when the user wants to map an ML operator onto a B200/Blackwell kernel, or to review which threads should execute a given tile primitive, where the data should live, and whether to invoke thread code or TMA/tcgen05.

When should I use B200 Scope Layout Dispatch?

B200 Scope Layout Dispatch fits situations like: the user wants to map an ML operator onto a B200/Blackwell kernel; review which threads should execute a given tile primitive; where the data should live; whether to invoke thread code.

How do I install B200 Scope Layout Dispatch in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-scope-layout-dispatch -a claude-code`. Or copy the skill folder (.claude/skills/b200-scope-layout-dispatch in mirage-project/mirage) into .claude/skills/b200-scope-layout-dispatch in your project. Claude Code loads it when a task matches its description.

How do I install B200 Scope Layout Dispatch in Codex?

Run `npx skills add mirage-project/mirage --skill b200-scope-layout-dispatch -a codex`. Or copy the skill folder (.claude/skills/b200-scope-layout-dispatch in mirage-project/mirage) into .agents/skills/b200-scope-layout-dispatch in your project. Codex loads it when a task matches its description.

Can I use B200 Scope Layout Dispatch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-scope-layout-dispatch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-scope-layout-dispatch, .gemini/skills/b200-scope-layout-dispatch, .github/skills/b200-scope-layout-dispatch and .opencode/skills/b200-scope-layout-dispatch in your project.

What does B200 Scope Layout Dispatch need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Scope Layout Dispatch is instructions for the agent only.

Does B200 Scope Layout Dispatch access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Scope Layout Dispatch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Scope Layout Dispatch use?

B200 Scope Layout Dispatch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Scope Layout Dispatch use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Scope Layout Dispatch?

Skills that share tags, products or a category with B200 Scope Layout Dispatch: Token Map (nexu-io/open-design, 100k stars), Operational Home Layout (kunchenguid/firstmate, 7.7k stars), Maps Geography (asgeirtj/system_prompts_leaks, 69k stars) and It Operations (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Scope Layout Dispatch?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,541 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.