Agent skill

B200 Tcgen05 Mma Contract Builder

by mirage-project in mirage-project/mirage

A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an…

Apache-2.0Auto-check passedAI & LLM Engineering

Install B200 Tcgen05 Mma Contract Builder

skills CLI
$ npx skills add mirage-project/mirage --skill b200-tcgen05-mma-contract-builder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-tcgen05-mma-contract-builder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-tcgen05-mma-contract-builder .claude/skills/b200-tcgen05-mma-contract-builder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-tcgen05-mma-contract-builder
GitHub stars
2.5k
Token cost
~1.8k tokens
SKILL.md length
850 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an…

  • Works in 3 steps: "Choose the tcgen05 tile and cta_group… → "Where should the scale factors of an… → "How is the accumulator split across the…
  • The user needs to choose the tile
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Tcgen05 Mma Contract Builder is an agent skill from mirage-project/mirage. Use when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an mxfp8/nvfp4 block-scaled GEMM. Produces an auditable MMA contract and completion protocol. Not for ordinary CUDA-core matmul or non-Blackwell targets.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in AI & LLM Engineering. It works with CUDA. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user needs to choose the tile
  • SMEM operand layout
  • TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell
  • Implement an mxfp8/nvfp4 block-scaled GEMM

Example prompts

  • “/b200-tcgen05-mma-contract-builder”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. "Choose the tcgen05 tile and cta_group for this B200 GEMM."
  2. "Where should the scale factors of an nvfp4 block-scaled MMA go?"
  3. "How is the accumulator split across the two CTAs' TMEM under cta_group::2?"

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Tcgen05 Mma Contract Builder loads about 1.8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 850 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 850 words, ~1,837 tokens.

Download SKILL.mdSave it as .claude/skills/b200-tcgen05-mma-contract-builder/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-tcgen05-mma-contract-builder
description
Use when the user needs to choose the tile, dtype, `cta_group::1/2`, SMEM operand layout, or TMEM accumulator mapping for a `tcgen05` MMA on B200/Blackwell, or to implement an mxfp8/nvfp4 block-scaled GEMM. Produces an auditable MMA contract and completion protocol. Not for ordinary CUDA-core matmul or non-Blackwell targets.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S7; S8; S5; S16
tags
b200, blackwell
related_skills
b200-scope-layout-dispatch, b200-layout-contract-auditor, b200-tmem-lifecycle-planner, b200-mbarrier-protocol-auditor, b200-cluster-persistent-scheduler…
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 tcgen05 MMA Contract Builder

R — Source evidence (Reading, paraphrased)

  • [S7] tcgen05 is the Blackwell Tensor Core instruction family, issued by a single elected thread on behalf of the participating group; the operation itself is asynchronous.
  • [S7] A/B usually reside in SMEM and the accumulator in TMEM; cta_group::2 makes the two CTAs of the same cluster cooperate, each keeping its own accumulator fragment.
  • [S7] The block-scaled mode adds SFA/SFB: the data stays in SMEM, while the scale factors are supplied through TMEM.
  • [S16/S17] B200 is compute capability 10.0; when using architecture-conditional features, make the portability boundary explicit.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

A complete MMA contract contains at least:

  • the math shape M×N×K, the input/accumulate dtypes, and whether it is block-scaled;
  • the participating unit, cta_group::1 or cta_group::2;
  • which address space and layout each of A/B/scale/C lives in;
  • who issues, who participates, and when it completes;
  • the accumulator's mapping in TMEM and how the epilogue reads it back;
  • boundary tiles, alignment, and the toolchain target.

If only a PTX mnemonic is given, without this contract, the agent should not generate code that merely "looks like it compiles".


A1 — Applications in the source (Past Application)

Case 1: cta_group::1, M=128
  • A single CTA supplies the SMEM tiles of A/B.
  • The 128 M rows map directly onto TMEM's 128 Lane rows; N maps onto Col.
Case 2: cta_group::2, M=256
  • Two CTAs cooperate; each CTA owns 128 M rows and keeps the corresponding accumulator in its own TMEM.
  • The even CTA is responsible for issuing the operation and for pair completion.
Case 3: block-scaled nvfp4/mxfp8
  • The quantized A/B data is read from SMEM.
  • SFA follows A's M partitioning; SFB, because both CTAs share B, must be visible/multicast to the pair.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "Choose the tcgen05 tile and cta_group for this B200 GEMM."
  2. "Where should the scale factors of an nvfp4 block-scaled MMA go?"
  3. "How is the accumulator split across the two CTAs' TMEM under cta_group::2?"
Language signals
  • "Choose the tcgen05 tile and cta_group for this B200 GEMM."
  • "Where should the scale factors of an nvfp4 block-scaled MMA go?"
  • "How is the accumulator split across the two CTAs' TMEM under cta_group::2?"
Distinction from adjacent skills

Versus b200-tmem-lifecycle-planner: this skill first defines the MMA's math and hardware contract; the TMEM planner goes deeper into column budgeting and reuse. Versus b200-cluster-persistent-scheduler: this skill focuses on a single cooperative MMA; the latter focuses on cluster-level scheduling and tails.


Show full SKILL.md (414 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must execute the following procedure:

  1. Confirm the target and toolchain
    • Target GPU, compute capability, whether sm_100a architecture-conditional features are allowed.
    • dtype and numerical-error requirements.
  2. Define the math tile
    • M/N/K, transposes, accumulate semantics, boundary/remainder handling.
  3. Choose the CTA group
    • group::1: simpler single-CTA resources and synchronization.
    • group::2: a larger cooperative tile and cross-CTA sharing, but added cluster/DSMEM/remote-barrier complexity.
  4. Define operand placement
    • A/B SMEM layout, swizzle, and the slice each CTA holds.
    • For block-scaled, list the SFA/SFB shapes, the K block size, and the SMEM→TMEM copy.
  5. Define the accumulator mapping
    • Write out each CTA's TLane/TCol formulas.
    • For modes such as M=64/128/256, make the lane packing explicit; avoid assuming a contiguous mapping.
  6. Define issue and completion
    • The elected thread issues; the commit group is bound to a completion barrier.
    • Every TMEM consumer must run only after the barrier completes.
  7. Define the epilogue
    • tcgen05.ld fragment shape, register cast/fusion, store path.
  8. Run the three-contract check
    • The SMEM operand layout, the TMEM layout, and the async completion must all match.
  9. Output the implementation skeleton and validation
    • Start with small shapes and random asymmetric data.
    • For dense and block-scaled separately, build a high-precision reference, error thresholds, and boundary K-block tests.
Required outputs
  1. Conclusion: the current choice/diagnosis, without vague "could be any of them" hedging.
  2. Evidence or assumptions: which come from user data, and which are hypotheses awaiting verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: correctness tests, boundary tests, and one falsifiable experiment.
  5. Risks and fallback: alternative paths when hardware, version, or resource conditions are not met.

B — Boundaries (Boundary) ★

Do not use when
  • The target is Ampere/Hopper, or only CUDA cores are used.
  • The shape is too small and Tensor Core tile utilization extremely low; a custom MMA has not yet been shown to be worthwhile.
Failure modes
  • Assuming every thread should issue the MMA.
  • Treating the TMEM accumulator as a register fragment.
  • Splitting only the compute under cta_group::2 without clarifying each CTA's operand/accumulator ownership.
  • Wrong block-scale SFA/SFB layout or visibility.
Limitations
  • The exact shapes/dtypes the instruction supports and the compiler APIs evolve; consult the current toolchain reference before generating code.

  • depends-on: b200-scope-layout-dispatch, b200-layout-contract-auditor
  • contrasts-with: none
  • composes-with: b200-tmem-lifecycle-planner, b200-mbarrier-protocol-auditor, b200-cluster-persistent-scheduler, b200-gemm-optimization-ladder

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-tcgen05-mma-contract-builder of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Tcgen05 Mma Contract Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Tcgen05 Mma Contract Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Tcgen05 Mma Contract Builder this skillmirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Benchmark TuneMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Areno Develop Kernel

    inclusionAI/AReno

    Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

    323 GitHub stars~498 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about B200 Tcgen05 Mma Contract Builder

What does B200 Tcgen05 Mma Contract Builder do?

A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an…. B200 Tcgen05 Mma Contract Builder is an agent skill from mirage-project/mirage. Use when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an mxfp8/nvfp4 block-scaled GEMM.

When should I use B200 Tcgen05 Mma Contract Builder?

B200 Tcgen05 Mma Contract Builder fits situations like: the user needs to choose the tile; SMEM operand layout; TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell; implement an mxfp8/nvfp4 block-scaled GEMM.

How do I install B200 Tcgen05 Mma Contract Builder in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-tcgen05-mma-contract-builder -a claude-code`. Or copy the skill folder (.claude/skills/b200-tcgen05-mma-contract-builder in mirage-project/mirage) into .claude/skills/b200-tcgen05-mma-contract-builder in your project. Claude Code loads it when a task matches its description.

How do I install B200 Tcgen05 Mma Contract Builder in Codex?

Run `npx skills add mirage-project/mirage --skill b200-tcgen05-mma-contract-builder -a codex`. Or copy the skill folder (.claude/skills/b200-tcgen05-mma-contract-builder in mirage-project/mirage) into .agents/skills/b200-tcgen05-mma-contract-builder in your project. Codex loads it when a task matches its description.

Can I use B200 Tcgen05 Mma Contract Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-tcgen05-mma-contract-builder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-tcgen05-mma-contract-builder, .gemini/skills/b200-tcgen05-mma-contract-builder, .github/skills/b200-tcgen05-mma-contract-builder and .opencode/skills/b200-tcgen05-mma-contract-builder in your project.

What does B200 Tcgen05 Mma Contract Builder need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Tcgen05 Mma Contract Builder is instructions for the agent only.

Does B200 Tcgen05 Mma Contract Builder access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Tcgen05 Mma Contract Builder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Tcgen05 Mma Contract Builder use?

B200 Tcgen05 Mma Contract Builder is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Tcgen05 Mma Contract Builder use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Tcgen05 Mma Contract Builder?

Skills that share tags, products or a category with B200 Tcgen05 Mma Contract Builder: Esmfold2 (JimLiu/science-skills, 228 stars), MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), Benchmark Tune (Mesh-LLM/mesh-llm, 3.5k stars) and Cuda Kernel Optimizer (KernelFlow-ops/cuda-optimized-skill, 214 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Tcgen05 Mma Contract Builder?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.