Agent skill

B200 Kernel Roofline Triage

by mirage-project in mirage-project/mirage

A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA".

Apache-2.0Auto-check passedAI & LLM Engineering

Install B200 Kernel Roofline Triage

skills CLI
$ npx skills add mirage-project/mirage --skill b200-kernel-roofline-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-kernel-roofline-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-kernel-roofline-triage .claude/skills/b200-kernel-roofline-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-kernel-roofline-triage
GitHub stars
2.5k
Token cost
~2.2k tokens
SKILL.md length
1,020 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA".

  • Works in 3 steps: How much useful compute is done per… → How many bytes were moved, from which… → Does the current implementation actually…
  • The user asks why is this B200/Blackwell kernel slow
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Kernel Roofline Triage is an agent skill from mirage-project/mirage. Use when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA". Classifies the problem via arithmetic intensity, dataflow, and resource occupancy as memory-bandwidth-, compute-throughput-, latency/concurrency-, or scheduling-bound, and gives a minimal falsifiable experiment. Not for queries that only ask about hardware specs with no kernel/operator context.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user asks why is this B200/Blackwell kernel slow
  • What should I optimize first
  • Move to Tensor Cores/TMA

Example prompts

  • “/b200-kernel-roofline-triage”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. How much useful compute is done per output element? Estimate the FLOPs.
  2. How many bytes were moved, from which level of storage, to do that compute? Give at least the HBM accounting; add L2/SMEM accountings when…
  3. Does the current implementation actually convert the theoretical roof into hardware busyness? Even with high algorithmic arithmetic…

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Kernel Roofline Triage loads about 2.2k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 1,020 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,020 words, ~2,182 tokens.

Download SKILL.mdSave it as .claude/skills/b200-kernel-roofline-triage/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-kernel-roofline-triage
description
Use when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA". Classifies the problem via arithmetic intensity, dataflow, and resource occupancy as memory-bandwidth-, compute-throughput-, latency/concurrency-, or scheduling-bound, and gives a minimal falsifiable experiment. Not for queries that only ask about hardware specs with no kernel/operator context.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S3; S16
tags
b200, roofline, performance, triage
related_skills
b200-gemm-optimization-ladder, b200-layout-contract-auditor, b200-tma-pipeline-designer, b200-warp-specialized-debugger
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 Kernel Roofline Triage

R — Source evidence (Reading, paraphrased)

  • [S3] Split the kernel's ceiling into a "compute roof" and a "bandwidth roof", and use arithmetic intensity to decide which side is more likely the current constraint.
  • [S3] For low-arithmetic-intensity operators, prioritize reducing bytes, fusion, reuse, or a narrower dtype; for high-arithmetic-intensity GEMMs, the focus is keeping the Tensor Cores continuously busy.
  • [S16] Blackwell still inherits the general CUDA best practices: parallelize, reduce Host↔Device transfers, coalesce accesses, and reduce redundant accesses and warp divergence.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

Do not start optimizing from "this trick is new"; answer three questions first:

  1. How much useful compute is done per output element? Estimate the FLOPs.
  2. How many bytes were moved, from which level of storage, to do that compute? Give at least the HBM accounting; add L2/SMEM accountings when necessary.
  3. Does the current implementation actually convert the theoretical roof into hardware busyness? Even with high algorithmic arithmetic intensity, a wrong layout, serialized load/compute/store, resource pressure, or the launch shape can still leave the Tensor Cores idle.

The final diagnosis must not just say "memory-bound/compute-bound"; it must also give the evidence, the alternative explanations not yet ruled out, and the next minimal experiment.


A1 — Applications in the source (Past Application)

Case 1: a large GEMM that should be compute-bound yet measures low
  • Problem: the matrices are large enough that by arithmetic intensity it should sit to the right of the ridge under the compute roof, yet Tensor Core utilization is low.
  • How the methodology was used: after ruling out HBM bandwidth, check whether execution is still serialized as "load→compute→store" and whether TMA, software pipelining, or warp specialization is missing.
  • Conclusion: the bottleneck is not "move a little less HBM traffic" but idle compute engines and insufficient pipeline overlap.
Case 2: elementwise operators such as RMSNorm/GELU
  • Problem: adding more math optimizations barely changes performance.
  • How the methodology was used: their FLOPs/byte is very low, so first check coalescing, the number of reads/writes, fusion opportunities, and the dtype.
  • Conclusion: the goal is to approach the bandwidth roof, not to chase Tensor Core peak.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "This kernel only hits xx TFLOPS on B200 — where do I look first?"
  2. "Should this operator be fused, or switched to TMA/Tensor Cores?"
  3. "Do a roofline diagnosis for me and give an ordered sequence of optimization experiments."
Language signals
  • "This kernel only hits xx TFLOPS on B200 — where do I look first?"
  • "Should this operator be fused, or switched to TMA/Tensor Cores?"
  • "Do a roofline diagnosis for me and give an ordered sequence of optimization experiments."
Distinction from adjacent skills

Versus b200-gemm-optimization-ladder: this skill first determines the bottleneck and the optimization direction; the latter gives the step-by-step implementation route specifically for GEMM. Versus b200-layout-contract-auditor: this skill does global performance attribution; the latter audits addresses and hardware layout contracts in depth.


Show full SKILL.md (510 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must follow this procedure:

  1. Collect the minimal fact set
    • The operator formula, shapes, dtypes, batch, and whether inputs/outputs are reused.
    • The current implementation path (CUDA/Triton/TIRx/CUTLASS/framework op), the timing methodology, warmup, and whether communication is included.
    • The B200 model/power limit/clocks, the compile target, and a profiler summary.
    • Done criterion: you can write a one-line estimate of "useful FLOPs" and a one-line estimate of "HBM bytes".
  2. Compute at least one roofline accounting
    • AI_HBM = useful_FLOPs / HBM_bytes.
    • Prefer the measured bandwidth on the user's machine and the same-dtype peak; without them, only order-of-magnitude inference is possible — label the assumptions explicitly.
    • Done criterion: an initial memory-bound / compute-bound / near-ridge call.
  3. Check whether the implementation contradicts the initial call
    • memory-bound: check repeated reads/writes, intermediate tensors spilled to HBM, uncoalesced accesses, an overly wide dtype, and insufficient request concurrency.
    • compute-bound: check Tensor Core instructions, tile utilization, TMA/compute/store overlap, warp roles, tail tiles, and small shapes.
    • Neither fits: check launch latency, synchronization, CPU submission, communication, power/frequency, and occupancy resource pressure.
  4. Build the evidence matrix
    • For each candidate bottleneck write "supporting evidence / counter-evidence / measurement needed".
    • Done criterion: at least 3 candidates listed, and they must not all be the same class of micro-optimization.
  5. Design the minimal falsifiable experiment
    • Change only one factor at a time, e.g. disable fusion, switch to a contiguous layout, increase the pipeline depth, use a fixed shape, lock the clocks.
    • Stopping rule: if the experiment result contradicts the hypothesis, go back to step 3; do not keep stacking the same class of optimization.
  6. Output the optimization order
    • P0: correctness and timing credibility; P1: the roof-determined primary bottleneck; P2: secondary scheduling/resource issues; P3: fine-tuning.
Required outputs
  1. Conclusion: the current choice/diagnosis, never a vague "we may need to look at everything".
  2. Evidence or assumptions: which items come from user data and which are assumptions pending verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: a correctness test, a boundary test, and one falsifiable experiment.
  5. Risks and fallback: the alternative path when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • The user only asks about B200 memory capacity, price, or rack specs, with no operator or performance question.
  • With no shape, dtype, timing, or dataflow information at all, do not assert a bottleneck outright.
Failure modes
  • Treating the theoretical peak as a directly achievable promise.
  • Looking only at occupancy while ignoring that explicit pipelining can already hide the latency; or conversely, looking only at pipelining while ignoring that resources prevent residency in the first place.
  • Substituting a single profiler percentage for end-to-end evidence.
Limitations
  • Roofline has limited explanatory power for irregular accesses, short kernels, dependency chains, and cross-GPU communication; latency and communication models must be added when necessary.

  • depends-on: none
  • contrasts-with: b200-gemm-optimization-ladder
  • composes-with: b200-layout-contract-auditor, b200-tma-pipeline-designer, b200-warp-specialized-debugger

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-kernel-roofline-triage of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Kernel Roofline Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Kernel Roofline Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Kernel Roofline Triage this skillmirage-project/mirage2.5k—~2.2kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Liger Kernel Perflinkedin/Liger-Kernel6.7k—~1.5kAutomated safety check: PassBSD-2-Clause
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Perf

    linkedin/Liger-Kernel

    Optimizes the performance of existing Liger Kernel Triton kernels.

    6.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about B200 Kernel Roofline Triage

What does B200 Kernel Roofline Triage do?

A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA". B200 Kernel Roofline Triage is an agent skill from mirage-project/mirage. Use when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA".

When should I use B200 Kernel Roofline Triage?

B200 Kernel Roofline Triage fits situations like: the user asks why is this B200/Blackwell kernel slow; what should I optimize first; move to Tensor Cores/TMA.

How do I install B200 Kernel Roofline Triage in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-kernel-roofline-triage -a claude-code`. Or copy the skill folder (.claude/skills/b200-kernel-roofline-triage in mirage-project/mirage) into .claude/skills/b200-kernel-roofline-triage in your project. Claude Code loads it when a task matches its description.

How do I install B200 Kernel Roofline Triage in Codex?

Run `npx skills add mirage-project/mirage --skill b200-kernel-roofline-triage -a codex`. Or copy the skill folder (.claude/skills/b200-kernel-roofline-triage in mirage-project/mirage) into .agents/skills/b200-kernel-roofline-triage in your project. Codex loads it when a task matches its description.

Can I use B200 Kernel Roofline Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-kernel-roofline-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-kernel-roofline-triage, .gemini/skills/b200-kernel-roofline-triage, .github/skills/b200-kernel-roofline-triage and .opencode/skills/b200-kernel-roofline-triage in your project.

What does B200 Kernel Roofline Triage need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Kernel Roofline Triage is instructions for the agent only.

Does B200 Kernel Roofline Triage access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Kernel Roofline Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Kernel Roofline Triage use?

B200 Kernel Roofline Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Kernel Roofline Triage use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Kernel Roofline Triage?

Skills that share tags, products or a category with B200 Kernel Roofline Triage: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Liger Kernel Perf (linkedin/Liger-Kernel, 6.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Kernel Roofline Triage?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.