Agent skill

B200 Warp Specialized Debugger

by mirage-project in mirage-project/mirage

A skill your agent uses when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow".

Apache-2.0Auto-check passedEducation

Install B200 Warp Specialized Debugger

skills CLI
$ npx skills add mirage-project/mirage --skill b200-warp-specialized-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage b200-warp-specialized-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/b200-warp-specialized-debugger .claude/skills/b200-warp-specialized-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
b200-warp-specialized-debugger
GitHub stars
2.5k
Token cost
~1.9k tokens
SKILL.md length
840 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow".

  • Works in 3 steps: "This warp-specialized kernel deadlocks… → "The generated TIRx code is correct but… → "After an illegal access all subsequent…
  • A B200/TIRx/CUDA warp-specialized kernel fails to compile
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

B200 Warp Specialized Debugger is an agent skill from mirage-project/mirage. Use when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow". First verifies the environment and a minimal reproduction, then builds a roles/storage/handoff/lifetime worksheet from the generated CUDA/PTX, fixing one handoff at a time. Not for cases with no reproducible code yet or purely high-level model accuracy issues.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in Education, covering Educational content. It works with CUDA. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • A B200/TIRx/CUDA warp-specialized kernel fails to compile
  • Hits an illegal memory access
  • Produces wrong results
  • Is correct but slow

Example prompts

  • “correct but slow”
  • “/b200-warp-specialized-debugger”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. "This warp-specialized kernel deadlocks — how do I debug it systematically?"
  2. "The generated TIRx code is correct but very slow — help me look at the lowering."
  3. "After an illegal access all subsequent Python runs are broken — how do I get a minimal reproduction?"

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

B200 Warp Specialized Debugger loads about 1.9k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 840 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 840 words, ~1,873 tokens.

Download SKILL.mdSave it as .claude/skills/b200-warp-specialized-debugger/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
b200-warp-specialized-debugger
description
Use when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow". First verifies the environment and a minimal reproduction, then builds a roles/storage/handoff/lifetime worksheet from the generated CUDA/PTX, fixing one handoff at a time. Not for cases with no reproducible code yet or purely high-level model accuracy issues.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S15; S13; S14
tags
b200, blackwell
related_skills
b200-scope-layout-dispatch, b200-mbarrier-protocol-auditor, b200-layout-contract-auditor, blackwell-build-compatibility-auditor
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

B200 Warp-Specialized Kernel Debugger

R — Source evidence (Reading, paraphrased)

  • [S15] Before debugging, confirm the actually imported TVM/TIRx, the GPU capability, and the target; these examples target Blackwell sm_100a.
  • [S15] Runtime problems usually reduce to a broken handoff: an uninitialized barrier, a wrong arrival count/phase, a collective hidden inside a partial-role branch, a missing visibility fence, or storage reused too early.
  • [S15] Check role guards, mbarrier init, tcgen05, TMA, and CTA/warpgroup synchronization in the generated CUDA rather than rewriting the Python DSL first.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

The object of debugging in an asynchronous specialized kernel is not "lines of code" but the data handoffs across roles. The unified worksheet:

  • Roles: who issues each asynchronous operation?
  • Storage: where does the tile live in GMEM/SMEM/TMEM/register at each step?
  • Handoff: producer, consumer, signal, count, phase, fence/drain.
  • Lifetime: the earliest point it may be read, overwritten, or freed.

First classify the symptom as compile / deadlock / crash / wrong result / correct-but-slow, then follow a different evidence path for each.


A1 — Applications in the source (Past Application)

Case 1: deadlock

Common root causes: barrier init inside a role branch, a wrong arrival count, producer/consumer starting on the same initial phase, CTA-wide sync placed inside a warpgroup branch.

Case 2: wrong result

Common root causes: mismatched TMA/MMA/TMEM layouts, waiting on the wrong stage/phase, stores not drained, O/S/P regions reused too early.

Case 3: correct but slow

Common root causes: roles still running serially, pipeline bubbles, the producer becoming the submission bottleneck, resources leaving too few active clusters, the compiler not generating the expected specialized instructions.


A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "This warp-specialized kernel deadlocks — how do I debug it systematically?"
  2. "The generated TIRx code is correct but very slow — help me look at the lowering."
  3. "After an illegal access all subsequent Python runs are broken — how do I get a minimal reproduction?"
Language signals
  • "This warp-specialized kernel deadlocks — how do I debug it systematically?"
  • "The generated TIRx code is correct but very slow — help me look at the lowering."
  • "After an illegal access all subsequent Python runs are broken — how do I get a minimal reproduction?"
Distinction from adjacent skills

Versus b200-mbarrier-protocol-auditor: the debugger first builds global evidence organized by symptom; once a barrier is implicated, invoke the dedicated auditor. Versus the build compatibility skill: the former handles kernel runtime/lowering, the latter handles binary/arch compatibility.


Show full SKILL.md (418 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must follow this procedure:

  1. Verify the runtime context
    • Print the actual TVM/TIRx/torch paths and versions, the GPU name/capability, and the compile target.
    • After an illegal access, restart the process/context before reproducing.
  2. Narrow the reproduction
    • Smallest still-failing shape, fixed seed, unrelated fusion and concurrency turned off.
  3. Run minimal correctness first
    • Prove the reference before any performance work; on a compile fail, do not descend into runtime synchronization guesses.
  4. Save the generated code
    • Search for role guards, mbarrier_init, tcgen05, cp.async.bulk.tensor, cluster/CTA sync, and TMEM alloc/free.
  5. Fill in the four-column worksheet
    • roles / storage / handoff / lifetime.
  6. Branch by symptom
    • Compile: API, target, dispatch, buffer scope, unsupported shape.
    • Deadlock: init/count/phase/collective scope/commit arrival.
    • Crash: addresses, descriptors, TMEM/SMEM out-of-bounds, context poisoning.
    • Wrong: layout, stale phase, missing fence, premature reuse.
    • Slow: whether the specialized instructions were generated, role timelines, pipeline bubbles, resources/occupancy.
  7. Change only one handoff at a time
    • Record the minimal test and the generated-code diff before and after the change.
  8. Re-validation order
    • correctness → sanitizer/boundary → profiler → large-shape performance.
  9. Produce a reproducible report
    • Environment, commands, shapes, locations of the generated-code snippets, expected/actual, minimal diff.
Required outputs
  1. Conclusion: the current choice/diagnosis, never a vague "we may need to look at everything".
  2. Evidence or assumptions: which items come from user data and which are assumptions pending verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: a correctness test, a boundary test, and one falsifiable experiment.
  5. Risks and fallback: the alternative path when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • There is no code, log, shape, or reproduction steps — just "it feels slow".
  • A purely model-level training-loss anomaly that has not been shown to come from a custom kernel.
Failure modes
  • Rewriting the whole kernel as the first move.
  • Not restarting the context after an illegal access and treating the subsequent phantom errors as new evidence.
  • Reading only the DSL source and not the generated CUDA/PTX.
  • Modifying multiple barriers/layouts/roles at the same time.
Limitations
  • Some hardware-level hangs or compiler bugs require a minimal reproduction filed upstream; this skill can produce a high-quality issue, but it cannot guarantee a local workaround for every toolchain defect.

  • depends-on: b200-scope-layout-dispatch
  • contrasts-with: none
  • composes-with: b200-mbarrier-protocol-auditor, b200-layout-contract-auditor, blackwell-build-compatibility-auditor

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/b200-warp-specialized-debugger of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

B200 Warp Specialized Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

B200 Warp Specialized Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
B200 Warp Specialized Debugger this skillmirage-project/mirage2.5k—~1.9kAutomated safety check: PassApache-2.0
OpenMAIC Setup and ExtensionTHU-MAIC/OpenMAIC40k—~1.7kAutomated safety check: NotesMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Project Tutorrohitg00/ai-engineering-from-scratch66k—~1.6kAutomated safety check: PassMIT
Hung-Yi Lee Teaching Stylevoidful/hung-yi-lee-skill1.3k—~13kAutomated safety check: PassNone
Claude Code Self-Assessment Advisorlhfer/claude-howto-zh-cn2.3k—~1.7kAutomated safety check: PassMIT

Similar skills

  • Guides setup, classroom generation and secondary development for OpenMAIC, the multi-agent interactive classroom, one confirmed phase at a time.

    40k GitHub stars~1.7k tokensUpdated yesterday
    EducationAuto-check: notes
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Project Tutor

    rohitg00/ai-engineering-from-scratch

    Tutors a learner through one stage of a hands-on AI engineering project per session: lesson, prediction, code, grader run and reflection, with hints but never full solutions.

    66k GitHub stars~1.6k tokensUpdated yesterday
    EducationAuto-check passed
  • Hung-Yi Lee Teaching Style

    voidful/hung-yi-lee-skill

    Explains machine learning, LLMs, AI agents and speech modeling in a Hung-Yi Lee-inspired teaching style, drawing on a knowledge base built from his lectures and research references.

    1.3k GitHub stars~13k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Claude Code Self-Assessment Advisor

    lhfer/claude-howto-zh-cn

    Quizzes you on Claude Code in a quick or deep mode, scores your level across 10 topics and recommends what to learn next, in Chinese.

    2.3k GitHub stars~1.7k tokensUpdated 2 mo ago
    EducationAuto-check passed
  • AI Engineering Course Guide

    rohitg00/ai-engineering-from-scratch

    Routes a topic, question or bug to the exact lessons in the AI Engineering from Scratch curriculum and suggests the next command to run.

    66k GitHub stars~1.7k tokensUpdated yesterday
    EducationAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed

Works with

Categories

Questions about B200 Warp Specialized Debugger

What does B200 Warp Specialized Debugger do?

A skill your agent uses when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow". B200 Warp Specialized Debugger is an agent skill from mirage-project/mirage. Use when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow".

When should I use B200 Warp Specialized Debugger?

B200 Warp Specialized Debugger fits situations like: A B200/TIRx/CUDA warp-specialized kernel fails to compile; hits an illegal memory access; produces wrong results; is correct but slow.

How do I install B200 Warp Specialized Debugger in Claude Code?

Run `npx skills add mirage-project/mirage --skill b200-warp-specialized-debugger -a claude-code`. Or copy the skill folder (.claude/skills/b200-warp-specialized-debugger in mirage-project/mirage) into .claude/skills/b200-warp-specialized-debugger in your project. Claude Code loads it when a task matches its description.

How do I install B200 Warp Specialized Debugger in Codex?

Run `npx skills add mirage-project/mirage --skill b200-warp-specialized-debugger -a codex`. Or copy the skill folder (.claude/skills/b200-warp-specialized-debugger in mirage-project/mirage) into .agents/skills/b200-warp-specialized-debugger in your project. Codex loads it when a task matches its description.

Can I use B200 Warp Specialized Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill b200-warp-specialized-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/b200-warp-specialized-debugger, .gemini/skills/b200-warp-specialized-debugger, .github/skills/b200-warp-specialized-debugger and .opencode/skills/b200-warp-specialized-debugger in your project.

What does B200 Warp Specialized Debugger need to run?

SKILL.md names no scripts, command-line tools or credentials: B200 Warp Specialized Debugger is instructions for the agent only. Our summary lists: Python 3.

Does B200 Warp Specialized Debugger access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is B200 Warp Specialized Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does B200 Warp Specialized Debugger use?

B200 Warp Specialized Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does B200 Warp Specialized Debugger use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to B200 Warp Specialized Debugger?

Skills that share tags, products or a category with B200 Warp Specialized Debugger: OpenMAIC Setup and Extension (THU-MAIC/OpenMAIC, 40k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars), AI Engineering Project Tutor (rohitg00/ai-engineering-from-scratch, 66k stars) and Hung-Yi Lee Teaching Style (voidful/hung-yi-lee-skill, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains B200 Warp Specialized Debugger?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,543 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.