Agent skill

B200 Layout Contract Auditor

by mirage-project in mirage-project/mirage

A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused…

Apache-2.0Auto-check passedSecurity

Install B200 Layout Contract Auditor

skills CLI
$ npx skills add mirage-project/mirage --skill b200-layout-contract-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-layout-contract-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-layout-contract-auditor .claude/skills/b200-layout-contract-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-layout-contract-auditor
GitHub stars
2.5k
Token cost
~1.8k tokens
SKILL.md length
833 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused…

  • Works in 3 steps: "This TMA + tcgen05 kernel is… → "Help me find shared memory bank… → "How does this TMEM accumulator map back…
  • A B200/Blackwell kernel shows wrong results
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Layout Contract Auditor is an agent skill from mirage-project/mirage. Use when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused TMEM/register ownership. Audits shape–stride, thread distribution, swizzle, and the hardware operand contract layer by layer. Not for pure synchronization deadlocks or compile errors unrelated to memory layout.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in Security, covering Threat modeling and GPU and accelerator computing. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • A B200/Blackwell kernel shows wrong results
  • Uncoalesced global memory access
  • SMEM bank conflicts
  • A TMA swizzle that mismatches the Tensor Core read

Example prompts

  • “/b200-layout-contract-auditor”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. "This TMA + tcgen05 kernel is numerically wrong, but no address goes out of bounds."
  2. "Help me find shared memory bank conflicts and swizzle problems."
  3. "How does this TMEM accumulator map back to each lane's register fragment?"

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Layout Contract Auditor loads about 1.8k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 833 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 833 words, ~1,799 tokens.

Download SKILL.mdSave it as .claude/skills/b200-layout-contract-auditor/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-layout-contract-auditor
description
Use when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused TMEM/register ownership. Audits shape–stride, thread distribution, swizzle, and the hardware operand contract layer by layer. Not for pure synchronization deadlocks or compile errors unrelated to memory layout.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S4; S5; S6; S7; S8
tags
b200, blackwell
related_skills
b200-scope-layout-dispatch, b200-mbarrier-protocol-auditor, b200-tma-pipeline-designer, b200-tmem-lifecycle-planner, b200-tcgen05-mma-contract-builder…
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 Layout Contract Auditor

R — Source evidence (Reading, paraphrased)

  • [S4] A layout is the mapping from logical index to physical location; it directly determines coalescing, bank conflicts, and whether the hardware can read the tile at all.
  • [S5] The two constraints invariant across generations are global coalescing and shared-memory bank conflicts; Tensor Cores add specific operand layout contracts on top.
  • [S6] The TMA descriptor, the target SMEM swizzle, and the downstream MMA's interpretation of the layout must agree exactly.
  • [S8] TMEM is a two-dimensional TLane × TCol address space; it must not be understood as an ordinary SMEM byte array.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

A layout audit must proceed hop by hop along the data path:

logical tensor → GMEM access → SMEM tile/swizzle → Tensor Core operand → TMEM accumulator/scale → register fragment → output store

Each hop answers four things: the index formula, the physical stride, the owner, and the consumer's expectation. If any one of them disagrees, it can show up as degraded performance or a silent error. In particular, distinguish a "view/stride rewrite" from an "ownership change": the latter usually requires a real data movement.


A1 — Applications in the source (Past Application)

Case 1: TMA writes a 128B swizzle, but the MMA reads a different layout
  • The logical tile's values are unchanged, but the physical bank arrangement disagrees.
  • The hardware does not report a "layout mismatch"; it interprets the addresses it receives and computes on the wrong elements.
Case 2: a transpose mistakenly assumed to be free
  • A view over a single linear storage can change only the strides.
  • If the transpose also changes lane/register ownership or the SMEM swizzle, then a load/store/shuffle/specialized instruction must occur.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "This TMA + tcgen05 kernel is numerically wrong, but no address goes out of bounds."
  2. "Help me find shared memory bank conflicts and swizzle problems."
  3. "How does this TMEM accumulator map back to each lane's register fragment?"
Language signals
  • "This TMA + tcgen05 kernel is numerically wrong, but no address goes out of bounds."
  • "Help me find shared memory bank conflicts and swizzle problems."
  • "How does this TMEM accumulator map back to each lane's register fragment?"
Distinction from adjacent skills

Versus b200-mbarrier-protocol-auditor: this skill checks "where the bytes are and who owns them"; the barrier auditor checks "when they may be read and when they may be reused". Wrong results often require invoking both in combination.


Show full SKILL.md (406 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must follow this procedure:

  1. Freeze the logical semantics
    • Write down each tensor/tile's logical shape, axis meanings, transpose relationships, and the expected element formula.
  2. Audit GMEM coalescing
    • List lane→address for one warp; check contiguity, alignment, transaction count, and boundary tiles.
  3. Audit the SMEM bank mapping
    • Compute banks for the consuming instruction's access pattern; identify same-bank different-address conflicts.
    • When choosing a swizzle, prefer the largest atom that the tile's contiguous dimension can fill; drop to a smaller atom when it cannot.
  4. Check the three-way contract
    • The TMA tensor-map descriptor's shape/stride/tile/swizzle.
    • The SMEM layout declared by the DSL/code.
    • The MMA/specialized load's interpretation of the operand.
    • All three must agree item by item.
  5. Audit TMEM
    • Pin down the TLane, TCol mapping and the column base.
    • Distinguish the accumulator layout from the block-scale layout; both live in TMEM but usually differ.
  6. Audit the register fragment
    • Write down which elements each lane holds; confirm the tcgen05.ld shape/repeat matches the epilogue's expectation.
  7. Identify the real data-movement points
    • Label views that change only shape/stride separately from rearrangements that change owner/swizzle.
  8. Output fix recommendations and validation vectors
    • Use small, non-symmetric matrices (so transpose errors are not masked by symmetric data).
    • Design an observable sentinel pattern for each hop.
Required outputs
  1. Conclusion: the current choice/diagnosis, never a vague "we may need to look at everything".
  2. Evidence or assumptions: which items come from user data and which are assumptions pending verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: a correctness test, a boundary test, and one falsifiable experiment.
  5. Risks and fallback: the alternative path when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • A pure barrier deadlock with no sign of a data-layout issue.
  • Merely a Python API rename or a link failure.
Failure modes
  • Validating with all-0, all-1, or symmetric inputs, which masks misplacement.
  • Looking only at logical shapes without tracking lane/warp/CTA ownership.
  • Assuming every transpose/view is zero-copy.
Limitations
  • Some layout contracts depend on the exact PTX instruction forms and the compiler lowering; they must be verified against the generated CUDA/PTX.

  • depends-on: b200-scope-layout-dispatch
  • contrasts-with: b200-mbarrier-protocol-auditor
  • composes-with: b200-tma-pipeline-designer, b200-tmem-lifecycle-planner, b200-tcgen05-mma-contract-builder, b200-warp-specialized-debugger

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-layout-contract-auditor of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Layout Contract Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Layout Contract Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Layout Contract Auditor this skillmirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.0
Forensifyalexgreensh/repo-forensics190—~2.5kAutomated safety check: NotesCustom licence
Create Rulecartography-cncf/cartography4.1k—~3kAutomated safety check: PassApache-2.0
Commit Security Scancodexstar69/bug-hunter520—~629Automated safety check: PassMIT
Auditing Code For Vulnerabilitiestrilwu/secskills157—~3.2kAutomated safety check: PassMIT
Threat Mitigation Mappingwshobson/agents40k8 repos~742Automated safety check: PassMIT

Similar skills

  • Forensify

    alexgreensh/repo-forensics

    Cross-agent self-inspection of your AI-agent stack. An agent skill from alexgreensh/repo-forensics.

    190 GitHub stars~2.5k tokensUpdated 13 days ago
    SecurityAuto-check: notes
  • Create Rule

    cartography-cncf/cartography

    Author a Cartography security rule (one or more Cypher Facts plus a Pydantic Finding output model) under cartography/rules/data/rules/.

    4.1k GitHub stars~3k tokensUpdated today
    SecurityAuto-check passed
  • Commit Security Scan

    codexstar69/bug-hunter

    Scan code changes for security vulnerabilities using Bug Hunter-native artifacts and STRIDE context.

    520 GitHub stars~629 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis.

    157 GitHub stars~3.2k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Match identified threats to preventive, detective and corrective controls across network, application, data, endpoint and process layers to plan remediation.

    40k GitHub starsUsed in 8 repos~742 tokens
    SecurityAuto-check passed
  • Audit Browser Security Boundaries

    nordstjernen-web/northstar-browser

    Audit browser-engine changes that process untrusted content or cross native-memory, origin, network, storage, extension, decoder, sandbox, or operating-system boundaries.

    127 GitHub stars~920 tokensUpdated today
    SecurityAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about B200 Layout Contract Auditor

What does B200 Layout Contract Auditor do?

A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused…. B200 Layout Contract Auditor is an agent skill from mirage-project/mirage. Use when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused TMEM/register ownership.

When should I use B200 Layout Contract Auditor?

B200 Layout Contract Auditor fits situations like: A B200/Blackwell kernel shows wrong results; uncoalesced global memory access; SMEM bank conflicts; A TMA swizzle that mismatches the Tensor Core read.

How do I install B200 Layout Contract Auditor in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-layout-contract-auditor -a claude-code`. Or copy the skill folder (.claude/skills/b200-layout-contract-auditor in mirage-project/mirage) into .claude/skills/b200-layout-contract-auditor in your project. Claude Code loads it when a task matches its description.

How do I install B200 Layout Contract Auditor in Codex?

Run `npx skills add mirage-project/mirage --skill b200-layout-contract-auditor -a codex`. Or copy the skill folder (.claude/skills/b200-layout-contract-auditor in mirage-project/mirage) into .agents/skills/b200-layout-contract-auditor in your project. Codex loads it when a task matches its description.

Can I use B200 Layout Contract Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-layout-contract-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-layout-contract-auditor, .gemini/skills/b200-layout-contract-auditor, .github/skills/b200-layout-contract-auditor and .opencode/skills/b200-layout-contract-auditor in your project.

What does B200 Layout Contract Auditor need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Layout Contract Auditor is instructions for the agent only. Our summary lists: Python 3.

Does B200 Layout Contract Auditor access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Layout Contract Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Layout Contract Auditor use?

B200 Layout Contract Auditor is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Layout Contract Auditor use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Layout Contract Auditor?

Skills that share tags, products or a category with B200 Layout Contract Auditor: Forensify (alexgreensh/repo-forensics, 190 stars), Create Rule (cartography-cncf/cartography, 4.1k stars), Commit Security Scan (codexstar69/bug-hunter, 520 stars) and Auditing Code For Vulnerabilities (trilwu/secskills, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Layout Contract Auditor?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.