Agent skill

V2 Kernel Writing

by mirage-project in mirage-project/mirage

Runtime-V2 kernel-writing workflow. An agent skill from mirage-project/mirage.

Apache-2.0Auto-check passedAgent Workflows

Install V2 Kernel Writing

skills CLI
$ npx skills add mirage-project/mirage --skill v2-kernel-writing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage v2-kernel-writing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/v2-kernel-writing .claude/skills/v2-kernel-writing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
v2-kernel-writing
GitHub stars
2.5k
Token cost
~3.7k tokens
SKILL.md length
1,600 words
Files
10 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runtime-V2 kernel-writing workflow. An agent skill from mirage-project/mirage.

  • Works in 7 steps: LOAD (orchestrator, no code) → SPEC (designer subagent) → IMPLEMENT (implementer subagent) → …
  • Rewriting ANY Runtime-V2 task kernel (tasks/blackwellv2/.cuh + registration) — a new op
  • SKILL.md covers Environment prerequisites…, Reference docs (this skill's…, Dispatch pattern (subagent… and Stage 0 — LOAD (orchestrator,…, plus 7 more sections
  • Calls git

What it does

V2 Kernel Writing is an agent skill from mirage-project/mirage. Runtime-V2 kernel-writing workflow. Use when writing, porting, or rewriting ANY Runtime-V2 task kernel (tasks/blackwellv2/.cuh + registration) — a new op, a v1→v2 port, or a rewrite toward the reference linearsm100v2 warp-role pipeline idiom. Drives the staged loop SPEC→IMPLEMENT→WIRE→VALIDATE→PERF→REVIEW with per-stage subagents and the b200- sub-skills, and enforces the M=1 anti-loop evidence + the v2 protocol invariants (§1.1 dep-prefix, stale-arrival re-init, skipafterstep0, taskoffset wiring).

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `applications/attn-ffn-reference-rewrite-plan.md`, `applications/ferret_dispatch_w13w2.md` and `applications/ffn_item1_spec.md`).

It sits in Agent Workflows, covering Subagents. It works with Git. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • Rewriting ANY Runtime-V2 task kernel (tasks/blackwellv2/.cuh + registration) — a new op
  • A rewrite toward the reference linearsm100v2 warp-role pipeline idiom

Example prompts

  • “/v2-kernel-writing”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. LOAD (orchestrator, no code)
  2. SPEC (designer subagent)
  3. IMPLEMENT (implementer subagent)
  4. WIRE (orchestrator)
  5. VALIDATE (validator subagent + references/validation-debug.md)
  6. PERF (orchestrator or validator)
  7. REVIEW (orchestrator)

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

V2 Kernel Writing loads about 3.7k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 1,600 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,600 words, ~3,742 tokens.

Download SKILL.mdSave it as .claude/skills/v2-kernel-writing/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
v2-kernel-writing
description
Runtime-V2 kernel-writing workflow. Use when writing, porting, or rewriting ANY Runtime-V2 task kernel (tasks/blackwell_v2/*.cuh + registration) — a new op, a v1→v2 port, or a rewrite toward the reference linear_sm100_v2 warp-role pipeline idiom. Drives the staged loop SPEC→IMPLEMENT→WIRE→VALIDATE→PERF→REVIEW with per-stage subagents and the b200-* sub-skills, and enforces the M=1 anti-loop evidence + the v2 protocol invariants (§1.1 dep-prefix, stale-arrival re-init, skip_after_step0, task_offset wiring).

You are writing a Runtime-V2 task kernel for MPK (dsv3-decode-clean). The house style is the reference linear_sm100_v2.cuh warp-role pipeline (loader W4 TMA → launcher W5 tcgen05/TMEM → consumers W0-3 epilogue → storer W6 page release), with consumer-only as the sanctioned idiom for non-GEMM-shaped ops. Quality bar and every protocol invariant live in this skill's references — the loop below tells you when to read what and what to dispatch.

For MODEL-level bring-up (whole compute graph → demo) use v2-model-support; this skill is the per-KERNEL inner loop that pipeline dispatches into.

Environment prerequisites (what must exist on the machine)

Everything the loop REQUIRES travels with the repo: this skill's references/ + applications/ docs, the reference kernels (include/mirage/persistent_kernel/tasks/blackwell_v2/), the harness (tests/runtime_python/blackwell_v2/), and the wiring surface. The rest is machine-local and DEGRADES GRACEFULLY:

  • b200-* sub-skills (the Skill-tool names in the index below) are USER-global (~/.claude/skills/b200-*), NOT in this repo — a fresh clone still has them only under the same user account. If missing: proceed anyway; references/house-style.md + references/upstream-kernel-catalog.md carry the distilled protocol/layout contracts.
  • Upstream catalog reads (git show mirage-project/runtime_refactor:<path>) need the remote: git remote add mirage-project https://github.com/mirage-project/mirage.git && git fetch mirage-project runtime_refactor. Optional — the catalog doc is self-contained.
  • ~/ferret/ (TEMPLATE_v2.yaml, docs/v2_runtime_notes.md, workspace1..8) — required ONLY for the ferret-v2 engine; ~/kda-workspaces/ only for kda; ~/kernel_tools/ (ncu_profile.sh) only for the NCU bound-check. Absent ⇒ route Stage 2/5 work to v2-kernel-engineer (the default anyway) and use b200-kernel-roofline-triage/manual NCU.
  • Personal memory (~/.claude/projects/-home-muhengl-mirage/memory/) — optional context only (same-user machines). references/m1-decode-evidence.md is the self-contained distillation of the evidence rows; treat the memory dir as its (optional) source citations.
  • Hardware: Stages 0-3 need no GPU. Stage 4/5 need a B200 (sm_100a); TP8-geometry ops need the 8-GPU box (see v2-model-support/references/box-orchestration.md) — without it, deliver with local gates + an explicit "pending TIER-1".
  • Codex MCP (mcp__codex__codex) for Stage-6 double-checks — if unconfigured, the ablation-logic-reviewer pass still runs; note the missing second engine in the verdict.

Reference docs (this skill's folder)

DocWhat it isRead at
references/house-style.mdReference methodology spec (roles, SEM tables, SMEM regions, TMA/tcgen05 patterns, quality bar)Stage 0, and by every subagent
references/upstream-kernel-catalog.mdPer-kernel/family pattern catalog of upstream runtime_refactor@0eadb3fd (which family to copy, gotchas, sync-with-upstream list)Stage 0 (find your op's family) + Stage 1
references/m1-decode-evidence.mdANTI-LOOP map: DEAD / WIN / UNTESTED at M=1 decodeStage 0 + Stage 1
references/wiring-recipe.mdThe v2 8-file registration checklist + footgunsStage 3
references/validation-debug.mdGates, hang/crash tooling, TIER measurement, profilerStage 4 + 5
references/ferret-v2-dispatch.mdStage-2/5 engine routing (engineer|ferret-v2|kda) + the ferret-v2 flow/contractStage 2 + 5
applications/attn-ffn-reference-rewrite-plan.mdThe staged attn/FFN rewrite (user directive)when working that campaign
applications/ffn_item1_spec.mdWorked Stage-1 SPEC exemplar (W13/W2 per-tile pipeline) — the deliverable shape Stage 1 must produceStage 1 (as template)
applications/ferret_dispatch_w13w2.mdWorked ferret-v2 dispatch brief exemplar (targets/gate/protocol_frozen/budget)Stage 2/5 ferret dispatches (as template)

Dispatch pattern (subagent nesting)

The MAIN THREAD (or one lead subagent for a multi-kernel campaign) is the orchestrator: it runs Stages 0/3/6 itself and dispatches ONE subagent per heavy stage — a designer (Stage 1), an implementer (Stage 2 — use the v2-kernel-engineer agent if defined in .claude/agents/, else general-purpose with that discipline pasted in), a validator (Stage 4). Subagents do not dispatch subagents. Each dispatch prompt MUST name: the stage's reference docs (absolute paths), the sub-skills to load via the Skill tool, the op contract, and the exact deliverable. Stages are sequential; iterate 2↔4 on failures, 5→1 on a perf verdict that changes the design.

Hard rules for every stage: default build byte-identical (new task types additive; levers env-gated default-OFF); no repo-wide refactors; GPU safety (test-mode first, never crash-loop the megakernel); every non-trivial conclusion → Stage 6 review before acting on it.

Stage 0 — LOAD (orchestrator, no code)

Read references/house-style.md + references/m1-decode-evidence.md IN FULL before any design. Then classify the op: shape (M, N, K / attention / elementwise), dtype, per-token work, TP/EP sharding, where it sits in the layer DAG (producer/consumer events), and which evidence rows (D*/W*/U*) touch it. Match the op to an upstream family (A pipeline / B consumer+regions / C consumer+monolith / D consumer+no-SMEM / E sub-op helper) in references/upstream-kernel-catalog.md — the family names the file to crib from. If the op matches a DEAD row and no new mechanism is on offer — STOP and say so; that is a successful outcome of this stage.

Stage 1 — SPEC (designer subagent)

Deliverable: a spec.h-style design doc (markdown). Location convention: campaign items that should travel with the repo go to .claude/skills/v2-kernel-writing/applications/<item>_spec.md (exemplar: applications/ffn_item1_spec.md); throwaway/exploratory specs go to scratch/ (git-ignored, machine-local). The spec contains:

  1. Engine choice via the DECISION TREE:
    • GEMM-shaped & M ≥ tile (or a real K-pipeline: ≥ 4 K-iterations of streamed weight tiles; page count per tile is dtype-dependent — an fp8 128×128 tile is exactly 1 page and still qualifies) → reference pipeline w/ TMA + tcgen05. Load b200-tma-pipeline-designer (stage ring, swizzle, load-vs-store completion), b200-tcgen05-mma-contract-builder (tile/dtype/ cta_group, SMEM operand layout, I-desc), b200-tmem-lifecycle-planner (TMEM columns, alloc/dealloc, ld/wait). Anchor every choice to house-style §2/§5.
    • M=1 GEMV / memory-bound streaming → consumer-GEMV per m1-decode-evidence (D2/D3/D9); nwarps per the W1 wave-quant tie-test ceil(items/(ntasks*nwarps)).
    • Attention-shaped → consumer-only per upstream attention_sm100.cuh (house-style §0); for FA-style rewrites additionally load b200-flash-attention4-planner.
    • Consumer-only ops also pick a SMEM family (catalog §2-4): planner-region typed buffer struct (rmsnorm/argmax pattern — the DEFAULT for anything staging through SMEM); ONE monolithic honest region when porting a kernel with hand-rolled internal offsets (attention pattern — never declare regions the device won't address); NUM_REGIONS=0 spec for pure-GMEM elementwise (silu/embedding pattern). A tiny fused sub-op (norm/rope-style) may be a family-E SMEM-view helper inside a host kernel instead of a new task.
    • Unsure how to map the op at all → load b200-scope-layout-dispatch first.
  2. SMEM region plan + budget: named regions, sizes, can_pack, page math (16KB pages, ≤14 pages, total ≤ 224256 B), alignment 1024 — house-style §4 format.
  3. SEM ordinal table (house-style §3 format: ordinal, count, producer→consumer, meaning; ≤31 op-private) — or the tag-flag alternative (W4) with its flag layout, if multi-role handshakes on the consumer-only idiom.
  4. Stage count + role responsibilities per role, incl. who re-inits which async mbars and who releases which pages on which path (bounds-fail included).
  5. Task granularity: DEFAULT per-tile (tile_idx = task_offset). Grid-wide fused is the EXCEPTION — requires written justification + the monotonic-barrier + skip_after_step0 + num_tasks==num_workers contract (house-style §6, wiring-recipe §7/§8).
  6. Evidence check: per design choice, the D*/W*/U* row it rests on; for U* probes, the pre-registered predicted Δ + kill threshold.
Show full SKILL.md (592 more words)Show less

Stage 2 — IMPLEMENT (implementer subagent)

Engine choice first (routing rule + flow in references/ferret-v2-dispatch.md): v2-kernel-engineer = protocol-heavy house-style port / first bring-up (default); ferret-v2 (ferret-kernel-agent V2 MODE) = beat-a-numeric-TIER-2-target optimization loop once a faithful FROZEN gate exists (op wired + harness case + live anchor) — also the "ferret writes from spec" path over a wired stub; kda = verdict-grade honest transfer when the number decides a campaign verdict and over-claim is costly. Ferret-v2 rearranges the stages (S3-stub wiring precedes the run — the gate substrate is the in-tree harness); its friction escape falls back to engineer-shell + ferret-math-only-body.

Write <op>_v2.cuh + <op>_v2_spec.h to the spec. House-style code conventions:

  • namespace kernel { namespace <op>_v2 {; spec.h constants + static_asserts pinning every mirrored constant; role-split __device__ __noinline__ functions (one per role).
  • mbarrier protocol per house-style §2/§3: start-of-task re-init of async-arrived mbars by their arriving role; fence.mbarrier_init.release.cluster after inits; tcgen05.fence::after_thread_sync at MMA↔wait boundaries. Before finalizing, invoke b200-mbarrier-protocol-auditor on the barrier ledger (every mbar: init count, arrivers, tx-bytes, waiters, phase evolution, re-init ownership).
  • extern __shared__ __align__(1024); SMEM only via task_desc->smem_region_offset(REGION_*).
  • No __syncthreads() in role loops — named barriers (bar.sync <free-id>, 128) or tag-flags only; elect_sync() for single-thread issue; no blockIdx for identity.
  • Layout doubts (swizzle vs tcgen05 operand, coalescing, bank conflicts) → b200-layout-contract-auditor. Build-flag doubts (sm_100a, -rdc=true) → blackwell-build-compatibility-auditor.

Stage 3 — WIRE (orchestrator)

Follow references/wiring-recipe.md top to bottom — enum, task_header include, register fn (§1.1 dep-prefix is the first line of the consumer body — MANDATORY), graph.cc tuple, runtime.cc task_type_to_name + task_offset=bid.x list, py wrapper (num_tasks==num_workers gate if grid-wide), builder use_v2 branch, skip_after_step0 on any monotonic-barrier scratch. Tick the §10 ship checklist explicitly.

Stage 4 — VALIDATE (validator subagent + references/validation-debug.md)

In order, no skipping: (1) test-mode numeric vs torch in tests/runtime_python/blackwell_v2/ (cos ≥ 0.999, rel_max ≤ 3e-2, no NaN; v1-counterpart compare); (2) §1.1/protocol static audit; (3) in-MPK --layers 0-3 probe; (4) MULTI-STEP run, iter ≥ 3 — iter-0-fine/iter-1-hang = persistent-state re-init (skip_after_step0), NOT a missing event; (5) on any hang: watchdog (-DMPK_V2_BREADCRUMB + MPK_V2_HANG_WATCHDOG_S); on any crash: compute-sanitizer memcheck = ground truth (breadcrumb counts are base-rate-biased). Deadlock/wrong-result debugging → b200-warp-specialized-debugger (roles/storage/handoff/lifetime worksheet, one handoff at a time). Math-changing on TP8 → poison-fill gate, not token-identity.

Stage 5 — PERF (orchestrator or validator)

TIER hierarchy is the law: TIER 1 in-MPK %globaltimer slowCTA @ production grid = the only verdict-grade number; harness slowCTA corroborates; cudaEvent-wall / standalone-warm are diagnostic-only. Compare against the reference/v1 body anchor from the spec. Bottleneck classification → b200-kernel-roofline-triage (achievable-floor rules from m1-decode-evidence §4 apply — same-grid xor-consumer floor, never theoretical peak). For a pipeline kernel that is correct-but-slow, climb b200-gemm-optimization-ladder one rung at a time. Profiler: buffer = 120000*128 entries; export via scripts/v2_perfetto_export.py. For a sustained beat-a-numeric-target optimization loop on one kernel, dispatch ferret-v2 (references/ferret-v2-dispatch.md): it iterates the pair against the frozen harness gate (TIER-2 body_span) autonomously; TIER-1 in-MPK slowCTA stays the final verdict here. For a whole perf-optimization CAMPAIGN around this kernel (measure→plan→implement→re-measure →land, agent roster + history contract) use the sibling skill v2-perf-iteration.

Stage 6 — REVIEW (orchestrator)

  • EVERY non-trivial conclusion (root-cause, DEAD/ALIVE verdict, perf claim, "matches reference") → ablation-logic-reviewer subagent + Codex MCP double-check (default params) BEFORE acting on or reporting it.
  • Landing: mpk-correctness-gate for anything math-adjacent, then mpk-commit-reviewer before git commit (staged-path + byte-identity + message gates). Verdicts → mpk-memory-keeper (experiment_history INDEX + memory; update m1-decode-evidence sources).

Sub-skill index (load via Skill tool, exact names)

Sub-skillUse atFor
b200-scope-layout-dispatchS1op→kernel mapping: scope/layout/dispatch/handoff contract
b200-tma-pipeline-designerS1/S2TMA descriptors, stage ring, swizzle, completion protocol
b200-tcgen05-mma-contract-builderS1/S2MMA tile/dtype/descriptor contract
b200-tmem-lifecycle-plannerS1/S2TMEM columns, alloc/ld/wait/dealloc lifecycle
b200-flash-attention4-plannerS1 (attn rewrites)QKᵀ/PV + online-softmax tile & barrier graph
b200-mbarrier-protocol-auditorS2 gateper-barrier ledger audit before finalizing
b200-layout-contract-auditorS2/S4shape-stride/swizzle/operand-contract bugs
blackwell-build-compatibility-auditorS2/S3sm_100a flags, PTX/cubin, JIT
b200-warp-specialized-debuggerS4deadlock / IMA / wrong-result / correct-but-slow
b200-kernel-roofline-triageS5bound classification + minimal falsifying experiment
b200-gemm-optimization-ladderS5staged GEMM perf climb with gates

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in .claude/skills/v2-kernel-writing of mirage-project/mirage.

  • SKILL.md
  • applications/attn-ffn-reference-rewrite-plan.md
  • applications/ferret_dispatch_w13w2.md
  • applications/ffn_item1_spec.md
  • references/ferret-v2-dispatch.md
  • references/house-style.md
  • references/m1-decode-evidence.md
  • references/upstream-kernel-catalog.md
  • references/validation-debug.md
  • references/wiring-recipe.md

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

V2 Kernel Writing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

V2 Kernel Writing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
V2 Kernel Writing this skillmirage-project/mirage2.5k—~3.7kAutomated safety check: PassApache-2.0
O2 Review Loopopenobserve/openobserve22k—~3.7kAutomated safety check: PassAGPL-3.0
ClawTeam Multi-Agent Swarmwin4r/ClawTeam-OpenClaw1.5k1 repos~2.9kAutomated safety check: PassMIT
Agent Deckasheshgoplani/agent-deck1k—~1.7kAutomated safety check: PassMIT
Clawteamwin4r/ClawTeam-OpenClaw1.5k—~3.1kAutomated safety check: PassMIT
Puppetmaster Agent Orchestrationprofessorpalmer/Puppetmaster467—~3.2kAutomated safety check: PassMIT

Similar skills

  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • ClawTeam Multi-Agent Swarm

    win4r/ClawTeam-OpenClaw

    Launches a swarm of specialist Hermes agents in git-worktree-isolated tmux windows with a kanban board and file-based inboxes, using built-in templates like hedge-fund and code-review.

    1.5k GitHub starsUsed in 1 repo~2.9k tokens
    Agent WorkflowsAuto-check passed
  • Agent Deck

    asheshgoplani/agent-deck

    agent-deck, the terminal session manager for AI coding agents.

    1k GitHub stars~1.7k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Clawteam

    win4r/ClawTeam-OpenClaw

    Multi-agent swarm orchestration. An agent skill from win4r/ClawTeam-OpenClaw.

    1.5k GitHub stars~3.1k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Puppetmaster Agent Orchestration

    professorpalmer/Puppetmaster

    Operates and supervises Puppetmaster, a multi-agent orchestrator, through its MCP tools or CLI, picking the right verb for edits, reviews, audits and long-running jobs.

    467 GitHub stars~3.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Claude Statusbar

    leeguooooo/claude-code-usage-bar

    Manage cs (claude-statusbar) — switch theme/style/density, override severity colors, preview combinations, run doctor, reset config, install, upgrade (cs upgrade — the only supported upgrade path)…

    377 GitHub stars~2.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about V2 Kernel Writing

What does V2 Kernel Writing do?

Runtime-V2 kernel-writing workflow. An agent skill from mirage-project/mirage. V2 Kernel Writing is an agent skill from mirage-project/mirage. Runtime-V2 kernel-writing workflow.

When should I use V2 Kernel Writing?

V2 Kernel Writing fits situations like: rewriting ANY Runtime-V2 task kernel (tasks/blackwellv2/.cuh + registration) — a new op; A rewrite toward the reference linearsm100v2 warp-role pipeline idiom.

How do I install V2 Kernel Writing in Claude Code?

Run `npx skills add mirage-project/mirage --skill v2-kernel-writing -a claude-code`. Or copy the skill folder (.claude/skills/v2-kernel-writing in mirage-project/mirage) into .claude/skills/v2-kernel-writing in your project. Claude Code loads it when a task matches its description.

How do I install V2 Kernel Writing in Codex?

Run `npx skills add mirage-project/mirage --skill v2-kernel-writing -a codex`. Or copy the skill folder (.claude/skills/v2-kernel-writing in mirage-project/mirage) into .agents/skills/v2-kernel-writing in your project. Codex loads it when a task matches its description.

Can I use V2 Kernel Writing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill v2-kernel-writing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/v2-kernel-writing, .gemini/skills/v2-kernel-writing, .github/skills/v2-kernel-writing and .opencode/skills/v2-kernel-writing in your project.

What does V2 Kernel Writing need to run?

Going by SKILL.md and its folder, V2 Kernel Writing needs the command-line tools its instructions call (git).

Does V2 Kernel Writing access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is V2 Kernel Writing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does V2 Kernel Writing use?

V2 Kernel Writing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does V2 Kernel Writing use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to V2 Kernel Writing?

Skills that share tags, products or a category with V2 Kernel Writing: O2 Review Loop (openobserve/openobserve, 22k stars), ClawTeam Multi-Agent Swarm (win4r/ClawTeam-OpenClaw, 1.5k stars), Agent Deck (asheshgoplani/agent-deck, 1k stars) and Clawteam (win4r/ClawTeam-OpenClaw, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains V2 Kernel Writing?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,541 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.