Agent skill

Blackwell Build Compatibility Auditor

by mirage-project in mirage-project/mirage

A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin…

Apache-2.0Auto-check passedTesting & QA

Install Blackwell Build Compatibility Auditor

skills CLI
$ npx skills add mirage-project/mirage --skill blackwell-build-compatibility-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage blackwell-build-compatibility-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/blackwell-build-compatibility-auditor .claude/skills/blackwell-build-compatibility-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
blackwell-build-compatibility-auditor
GitHub stars
2.5k
Token cost
~1.7k tokens
SKILL.md length
795 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin…

  • Works in 3 steps: whether the device/driver/toolchain… → whether the binary contains a native… → whether the code semantics depend on…
  • The user wants to confirm whether an existing CUDA extension/binary can run on B200
  • SKILL.md covers R — Source evidence (Reading,…, I — Methodology skeleton…, A1 — Applications in the… and A2 — Trigger scenarios (Future…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Blackwell Build Compatibility Auditor is an agent skill from mirage-project/mirage. Use when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin, CUDA Toolkit versions, JIT, and fatbin. Outputs compatibility evidence, build flags, and a smoke test. Not for kernel performance tuning once the kernel already runs correctly.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `test-prompts.json`).

It sits in Testing & QA, covering QA and bug reports and Deep learning. It works with CUDA. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • The user wants to confirm whether an existing CUDA extension/binary can run on B200
  • Configure compute100/sm100
  • The architecture-specific sm100a
  • Check PTX/cubin

Example prompts

  • “/blackwell-build-compatibility-auditor”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. whether the device/driver/toolchain recognizes the target;
  2. whether the binary contains a native cubin usable on B200 or JIT-able PTX;
  3. whether the code semantics depend on architecture-conditional features and legacy warp-synchronous assumptions.

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlc.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Blackwell Build Compatibility Auditor loads about 1.7k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 795 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 795 words, ~1,747 tokens.

Download SKILL.mdSave it as .claude/skills/blackwell-build-compatibility-auditor/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
blackwell-build-compatibility-auditor
description
Use when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure `compute_100/sm_100` or the architecture-specific `sm_100a`, or check PTX/cubin, CUDA Toolkit versions, JIT, and fatbin. Outputs compatibility evidence, build flags, and a smoke test. Not for kernel performance tuning once the kernel already runs correctly.
source_book
Modern GPU Programming For MLSys (MLC Community) + NVIDIA Blackwell Tuning/Compatibility Guides
source_chapter
S17; S16; S15
tags
blackwell, blackwell
related_skills
b200-kernel-roofline-triage, b200-warp-specialized-debugger, b200-tcgen05-mma-contract-builder
version
0.1.0
<!-- Distilled from "Modern GPU Programming for MLSys" — https://mlc.ai/modern-gpu-programming-for-mlsys/ -->

Blackwell Build Compatibility Auditor

R — Source evidence (Reading, paraphrased)

  • [S17] A cubin only runs within its compatible compute-capability range; PTX can be JIT-compiled on higher capabilities, so binaries are advised to retain PTX.
  • [S17] CUDA_FORCE_PTX_JIT=1 can be used to verify whether the application contains usable PTX; the variable must be unset after the test.
  • [S17] CUDA 12.8 can generate native cubin for Blackwell compute capability 10.0 while also retaining compute_100 PTX.
  • [S17] Architecture-conditional features using sm_100a/compute_100a have no general forward/backward compatibility.

Source: distilled from "Modern GPU Programming for MLSys" (https://mlc.ai/modern-gpu-programming-for-mlsys/) and the NVIDIA Blackwell tuning/compatibility guides. Short paraphrases only; no long passages are reproduced.


I — Methodology skeleton (Interpretation)

The compatibility audit has three layers:

  1. whether the device/driver/toolchain recognizes the target;
  2. whether the binary contains a native cubin usable on B200 or JIT-able PTX;
  3. whether the code semantics depend on architecture-conditional features and legacy warp-synchronous assumptions.

"It compiles on some machine" is not compatibility evidence; the artifacts and the actual load path must be checked.


A1 — Applications in the source (Past Application)

Case 1: artifacts from an old CUDA build
  • If the binary retains reasonably new PTX, B200 can JIT it at runtime.
  • If there are only old-architecture cubins and no compatible PTX, the kernel launch will fail and a rebuild is needed.
Case 2: a CUDA 12.8 build
  • Generating both the sm_100 native cubin and compute_100 PTX reduces first-run JIT while preserving future compatibility.
Case 3: sm_100a
  • Used when depending on specific Blackwell architecture-conditional features.
  • It should not be treated as a general Blackwell/PTX fallback; there must be a capability check and a plain path.

A2 — Trigger scenarios (Future Trigger) ★

In what situations will the user need this skill?
  1. "Can this PyTorch CUDA extension run directly on B200?"
  2. "How should nvcc's sm_100, compute_100, and sm_100a be configured?"
  3. "How do I prove the wheel/fatbin contains PTX/cubin usable on Blackwell?"
Language signals
  • "Can this PyTorch CUDA extension run directly on B200?"
  • "How should nvcc's sm_100, compute_100, and sm_100a be configured?"
  • "How do I prove the wheel/fatbin contains PTX/cubin usable on Blackwell?"
Distinction from adjacent skills

Difference from b200-warp-specialized-debugger: this skill first settles whether it can load/run correctly at all; the latter troubleshoots handoffs and performance once the kernel has already entered compilation or execution.


Show full SKILL.md (417 more words)Show less

E — Executable steps (Execution)

Once the skill is activated, the agent must execute the following process:

  1. Record the environment matrix
    • GPU/capability, driver, CUDA runtime/toolkit, compiler, framework, and extension versions.
  2. Inventory the binary forms
    • Check the fatbin/cubin/PTX targets; record whether sm_100, compute_100, or only old cubins are present.
  3. Run the PTX JIT verification
    • Temporarily set CUDA_FORCE_PTX_JIT=1 and run a minimal kernel smoke test.
    • Success: there is at least JIT-able PTX; failure: a rebuild or dependency fix is needed.
    • Explicitly unset it after the test.
  4. Generate the build flags
    • CUDA 12.8+: at minimum consider -gencode=arch=compute_100,code=sm_100 and code=compute_100.
    • Multi-architecture wheels: keep the cubin for each target and retain at least one PTX backend target.
  5. Audit the architecture-conditional paths
    • If sm_100a/compute_100a or tcgen05-specific features are used, establish device-capability detection and a non-a fallback.
  6. Check warp-synchronous assumptions
    • Do a migration audit of intra-warp communication without explicit synchronization and partial-mask collectives.
  7. Run the smoke matrix
    • Load the extension, a minimal kernel, representative dtypes/shapes, error-message capture, and first-JIT vs cached runs.
  8. Output the conclusion grade
    • A: native + PTX; B: runnable via PTX JIT; C: restricted architecture path only; F: no compatible code.
    • Attach the rebuild commands and risks.
Required outputs
  1. Conclusion: the current choice/diagnosis; do not use a vague "it could be any of them".
  2. Evidence or assumptions: which items come from user data, and which are hypotheses awaiting verification.
  3. Contract/table/timeline: the auditable intermediate artifacts corresponding to this skill.
  4. Minimal validation: correctness tests, boundary tests, and one falsifiable experiment.
  5. Risks and fallback: alternative paths when hardware, version, or resource requirements are not met.

B — Boundaries (Boundary) ★

Do not use when
  • The kernel already runs stably and the user only asks why it is not fast enough.
  • Non-CUDA package compatibility problems.
Failure modes
  • Looking only at the CUDA runtime version without checking the actual binary.
  • Forgetting to unset after forcing PTX JIT, misjudging later performance/startup time.
  • Treating sm_100a as ordinary PTX that can be JIT-compiled on any future GPU.
  • Wheel builds that keep only the dev machine's architecture cubin.
Limitations
  • Frameworks differ in the environment-variable names for architecture lists and in packaging logic, so the corresponding official build docs must be consulted; this skill provides the audit method, not a replacement for framework docs.

  • depends-on: none
  • contrasts-with: b200-kernel-roofline-triage
  • composes-with: b200-warp-specialized-debugger, b200-tcgen05-mma-contract-builder

Audit info

  • Validation passed: V1 ✓ / V2 ✓ / V3 ✓
  • Test definitions: 6 (3 should_trigger / 2 should_not_trigger / 1 edge_case)
  • Hardware validation: not performed; must be verified on a target B200
  • Distilled: 2026-06-25

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/blackwell-build-compatibility-auditor of mirage-project/mirage.

  • SKILL.md
  • test-prompts.json

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

Blackwell Build Compatibility Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Blackwell Build Compatibility Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Blackwell Build Compatibility Auditor this skillmirage-project/mirage2.5k—~1.7kAutomated safety check: PassApache-2.0
Scrub Issuepytorch/pytorch104k—~4.6kAutomated safety check: PassCustom licence
Nv Reason CxrNVIDIA/skills3.6k—~3.9kAutomated safety check: NotesApache-2.0
Minimal Run And Auditlllllllama/RigorPilot-Skills4971 repos~691Automated safety check: PassMIT
Ut Refactor Reviewintel/torch-xpu-ops115—~917Automated safety check: PassApache-2.0
Jetson Video SetupNVIDIA/skills3.6k1 repos~2.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Scrub Issue

    pytorch/pytorch

    Fetch, analyze, reproduce, and minimize GitHub issue reproductions.

    104k GitHub stars~4.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Nv Reason Cxr

    NVIDIA/skills

    Official

    Used for command-shape or live NV-Reason-CXR chest X-ray reasoning smoke tests.

    3.6k GitHub stars~3.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Minimal Run And Audit

    lllllllama/RigorPilot-Skills

    Rigor Run skill for README-first deep learning repo reproduction.

    497 GitHub starsUsed in 1 repo~691 tokens
    Testing & QAAuto-check passed
  • Ut Refactor Review

    intel/torch-xpu-ops

    Official

    Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests.

    115 GitHub stars~917 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Jetson Video Setup

    NVIDIA/skills

    Official

    A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

    3.6k GitHub starsUsed in 1 repo~2.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Blackwell Build Compatibility Auditor

What does Blackwell Build Compatibility Auditor do?

A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin…. Blackwell Build Compatibility Auditor is an agent skill from mirage-project/mirage. Use when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin, CUDA Toolkit versions, JIT, and fatbin.

When should I use Blackwell Build Compatibility Auditor?

Blackwell Build Compatibility Auditor fits situations like: the user wants to confirm whether an existing CUDA extension/binary can run on B200; configure compute100/sm100; the architecture-specific sm100a; check PTX/cubin.

How do I install Blackwell Build Compatibility Auditor in Claude Code?

Run `npx skills add mirage-project/mirage --skill blackwell-build-compatibility-auditor -a claude-code`. Or copy the skill folder (.claude/skills/blackwell-build-compatibility-auditor in mirage-project/mirage) into .claude/skills/blackwell-build-compatibility-auditor in your project. Claude Code loads it when a task matches its description.

How do I install Blackwell Build Compatibility Auditor in Codex?

Run `npx skills add mirage-project/mirage --skill blackwell-build-compatibility-auditor -a codex`. Or copy the skill folder (.claude/skills/blackwell-build-compatibility-auditor in mirage-project/mirage) into .agents/skills/blackwell-build-compatibility-auditor in your project. Codex loads it when a task matches its description.

Can I use Blackwell Build Compatibility Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill blackwell-build-compatibility-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/blackwell-build-compatibility-auditor, .gemini/skills/blackwell-build-compatibility-auditor, .github/skills/blackwell-build-compatibility-auditor and .opencode/skills/blackwell-build-compatibility-auditor in your project.

What does Blackwell Build Compatibility Auditor need to run?

SKILL.md names no scripts, command-line tools or credentials: Blackwell Build Compatibility Auditor is instructions for the agent only.

Does Blackwell Build Compatibility Auditor access the network?

SKILL.md names 1 domain. As links in the text: mlc.ai. This is read from the text; nothing was executed.

Is Blackwell Build Compatibility Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Blackwell Build Compatibility Auditor use?

Blackwell Build Compatibility Auditor is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Blackwell Build Compatibility Auditor use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Blackwell Build Compatibility Auditor?

Skills that share tags, products or a category with Blackwell Build Compatibility Auditor: Scrub Issue (pytorch/pytorch, 104k stars), Nv Reason Cxr (NVIDIA/skills, 3.6k stars), Minimal Run And Audit (lllllllama/RigorPilot-Skills, 497 stars) and Ut Refactor Review (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Blackwell Build Compatibility Auditor?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.