Agent skill

Mpk Faithful Gate

by mirage-project in mirage-project/mirage

Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

Apache-2.0Auto-check passedTesting & QA

Install Mpk Faithful Gate

skills CLI
$ npx skills add mirage-project/mirage --skill mpk-faithful-gate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage mpk-faithful-gate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/mpk-faithful-gate .claude/skills/mpk-faithful-gate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mpk-faithful-gate
GitHub stars
2.5k
Token cost
~2.6k tokens
SKILL.md length
1,135 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

  • Works in 5 steps: Find the task_type_id. Get the numeric… → Build the input at PRODUCTION decode… → Report slowCTA + wall + sumCTA + cos,… → …
  • Tasks that involve Meeting notes and agendas
  • SKILL.md covers Why this exists (read first —…, The existing gates — your…, Recipe — stand up + run a… and Decode geometry (the critical…, plus 3 more sections
  • Calls python3

What it does

Mpk Faithful Gate is an agent skill from mirage-project/mirage. Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape. Use this WHENEVER you need to measure, gate, or head-to-head-optimize an MPK kernel's per-task latency (dense FP8 GEMM, routed group-GEMM W13/W2, MLA decode, router/topk, AllReduce, etc.), stand up a faithful gate for a NEW kernel/shape, or dispatch a KDA/Ferret kernel agent against a faithful measure — and ESPECIALLY before trusting any per-task µs number in a DSv3 decode perf…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Meeting notes and agendas and End-to-end testing. It works with DeepSeek. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Meeting notes and agendas
  • Tasks that involve End-to-end testing

Example prompts

  • “/mpk-faithful-gate”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Find the task_type_id. Get the numeric TASK_*_SM100 enum from runtime_header.h
  2. Build the input at PRODUCTION decode geometry (see "Decode geometry" below) — this
  3. Report slowCTA + wall + sumCTA + cos, and emit a single
  4. Per-worker-count sweep. The optimal kernel FLIPS with the grid (GEMV wins at high
  5. Run it via the broker (never pick a fixed GPU; faithful needs an EXCLUSIVE card)

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mpk Faithful Gate loads about 2.6k tokens when it runs. Until then it costs about 232 tokens; SKILL.md has 1,135 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~232
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,135 words, ~2,629 tokens.

Download SKILL.mdSave it as .claude/skills/mpk-faithful-gate/SKILL.md (or your agent's skills folder).
name
mpk-faithful-gate
description
Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape. Use this WHENEVER you need to measure, gate, or head-to-head-optimize an MPK kernel's per-task latency (dense FP8 GEMM, routed group-GEMM W13/W2, MLA decode, router/topk, AllReduce, etc.), stand up a faithful gate for a NEW kernel/shape, or dispatch a KDA/Ferret kernel agent against a faithful measure — and ESPECIALLY before trusting any per-task µs number in a DSv3 decode perf campaign. The faithful in-MPK slowCTA is the trusted measure; a standalone green-ctx bench MIS-RANKS and a whole-megakernel e2e number hides per-task cost — do not use either as the gate. Covers the slowCTA definition, the _faithful_helper reuse, the decode (M=1) input geometry, the GPU-broker + exclusive remote box, the gate watchdog, and the candidate-overlay bridge to faithful_eval.py.

MPK Faithful Per-Task Gate

Why this exists (read first — it's the whole point)

Two tempting measures LIE for the DSv3 decode campaign:

  • A standalone green-ctx kernel bench (ferret/cpp_examples) runs the kernel alone on a fresh CUDA context. It MIS-RANKS — a kernel that wins standalone can be NULL or a REGRESS once it's compiled into the shared megakernel at the production grid with co-resident tasks (3 confirmed "standalone-doesn't-transfer" cases: kv-up CUDA-core, o_proj split-K, dense fine-N TP8).
  • The whole-megakernel e2e per-token latency hides which task moved.

The trusted measure is the FAITHFUL in-MPK per-task number: compile the kernel INTO the real megakernel, launch the PRODUCTION persistent grid (grid.x = num_workers, 136 on a B200), turn the persistent profiler ON, and extract THIS task's per-instance span:

  • slowCTA = max-over-CTAs of (end − begin) = the slowest single-CTA BODY (per-instance compute). THIS is the verdict metric. Not P50, not per-kernel-aggregate, not wall.
  • wall = max(end) − min(begin) = the isolated makespan (includes dispatch stagger).
  • sumCTA = total busy CTA-time = the ANTI-SLOUGHING guard. A candidate that lowers slowCTA by spreading the SAME work onto more CTAs (no real speedup) shows a RISING sumCTA — reject it. Judge on slowCTA AND wall AND sumCTA together.
  • cos ≥ 0.99 vs a PyTorch reference (correctness; never trade it for speed).

The existing gates — your templates (copy structure, don't reinvent)

FamilyGate fileNotes
Dense FP8 GEMM (qkv_a/q_b/q_b_pe/kv_b/o_proj/shared_*/router)tests/runtime_python/blackwell/sm100_fp8_gemm_dense/test_fp8_gemm_dense_*_pk_testmode.py + _build_helper.py (SHAPE_REGISTRY) + _faithful_helper.pyM=1 decode; shapes are named in SHAPE_REGISTRY (add a row for a new dense shape)
Routed group-GEMM W13/W2tests/runtime_python/blackwell/sm100_fp8_group_gemm_decode/test_fp8_group_gemm_faithful_pertask.pydecode geometry via MPK_GG_DECODE_M1; active experts via meta row-1 mask
(template core, kernel-agnostic)sm100_fp8_gemm_dense/_faithful_helper.py → profiled_per_task_latency_by_id, resolve_num_workers, make_profiler_tensorimport these; do not rewrite
Candidate-overlay bridge~/ferret/scripts/faithful_eval.py (KIND map + _run_test + _run_group_gemm)drives baseline vs shadow-overlay candidate in two subprocesses
GPU broker~/gpu_broker/ (gpu_gate.sh, gpu_pool.conf, acquire/release; machine-local)shares local + remote box; flock-serialized; exclusivity-checked

When the task fits one of these families, EXTEND it (add a shape row / a mode), don't write a new harness. Only author a fresh test_*_faithful_pertask.py for a genuinely new family (e.g. attention) — and even then mirror the group-GEMM file's structure exactly.

Recipe — stand up + run a faithful gate for a new (kernel, shape)

  1. Find the task_type_id. Get the numeric TASK_*_SM100 enum from runtime_header.h / the kernel's register_* in src/kernel/task_register.cc. Match the profiler row by this NUMERIC id, NOT the name — many task names are missing from profiler_persistent.py's event_name_list so the CSV writes UNKNOWN_<id>. profiled_per_task_latency_by_id(pk, TASK_ID, iters, label=...) handles this.

  2. Build the input at PRODUCTION decode geometry (see "Decode geometry" below) — this is the #1 correctness trap. Register the kernel through PersistentKernel(test_mode=True, num_workers=resolve_num_workers(...)), attach inputs, pk.compile(), pk() once for correctness, then the profiled timing run.

  3. Report slowCTA + wall + sumCTA + cos, and emit a single FAITHFUL_RESULT {...} JSON line (mirror the group-GEMM gate's _lat_table / FAITHFUL_RESULT format) so faithful_eval.py can parse baseline vs candidate.

  4. Per-worker-count sweep. The optimal kernel FLIPS with the grid (GEMV wins at high occupancy nw≥128; tcgen05 MMA wins at low nw≤68). Sweep nw ∈ {8, 64, 68, 128, 136} (128 first), via MPK_TEST_NUM_WORKERS / the --num-workers plumb. The 64/68 points feed the group-GEMM ‖ shared-expert worker-partition design.

  5. Run it via the broker (never pick a fixed GPU; faithful needs an EXCLUSIVE card):

    ~/gpu_broker/gpu_gate.sh --holder <tag> -- \
      --kind <finen|gemv_m1|largem_compact> --shape <S> --num-workers <W> \
      --configs <S> [--candidate-role smallm] [--kernel <cand.cuh>] [--baseline] [--decode-m1]

    It auto-acquires a free local-or-remote slot, runs the gate, streams KERNEL_RESULT/FAITHFUL_RESULT back, and releases (trap-on-EXIT). --baseline = the in-tree kernel (ratio 1.0); --kernel <cand.cuh> overlays a candidate via a shadow MIRAGE_ROOT (production tree never written).

Decode geometry (the critical correctness point)

Production decode is bs=1 → M=1 per active unit. A gate that feeds an all-128-rows-real activation measures a PREFILL-like geometry and CANNOT reward an M=1-aware kernel.

  • Dense projections: M=1 (the dense gate is already M=1).
  • Routed group-GEMM: exactly 1 real token per ACTIVE expert (top-k routing). Build a_bf16 with only the live row (expert*128 + 0) of each active expert non-zero (the other 127 rows are pad/zero), and validate cos on the LIVE rows ONLY. This is the MPK_GG_DECODE_M1=1 mode. The active experts come from meta row-1's active_expert_mask (~4–8 active at bs=1 TP8 EP2).
  • MLA / attention: per-rank decode (1 query token, the real KV-seq length).
Show full SKILL.md (476 more words)Show less

Invariants & gotchas (these have each cost real hours — bake them in)

  1. Exclusive GPU only. A faithful slowCTA on a contended card is INVALID. The gate torch-probes + exclusivity-checks (refuses on a foreign compute proc). The broker free-checks before acquiring.
  2. Gate watchdog (MPK_GATE_TIMEOUT_S, default 420s). A buggy candidate can DEADLOCK the megakernel and wedge the (exclusive) box indefinitely. faithful_eval.py's _run_test wraps the test in subprocess.run(timeout=...) → on timeout it SIGKILLs the child = the megakernel-HOST process → frees the card → returns rc=124 VERDICT=TIMEOUT_HANG. The timeout MUST live in the .py (which runs the test as its subprocess); a wrapper-level timeout orphans the test grandchild (it keeps holding the GPU). Keep this.
  3. Candidate kernels: NO block barrier inside the persistent-sweep tile loop. In MPK each worker strides tiles by num_workers; workers with fewer tiles EARLY-EXIT the loop, so a __syncthreads()/bar.sync inside it deadlocks (the exited workers never reach it). Each lane/warp does its full K-sweep independently; only a final intra-warp __shfl reduction is allowed. (This deadlock wedged the exclusive box for 24 min.)
  4. Remote box runs its OWN repo copy. The exclusive remote box (<BOX_USER>@<BOX_IP> — site-specific, resolve per session per v2-model-support/references/box-orchestration.md §1-2; its mirage repo + ferret-scripts paths are box-local) does NOT see your local edits. After editing a gate test / _build_helper.py (new shape) / faithful_eval.py (watchdog, new flag), scp them there (keep a .bak). Verify with a remote python3 -c "import ast; ...".
  5. The remote box's faithful_eval.sh has an arg-allowlist (and is OFF-LIMITS to edit — along with cc-run.sh). It REJECTS unknown flags (--decode-m1 → "unknown arg"). Forward a new knob as an ENV var instead, via the machine-local box shim (~/nebius_gate/faithful_eval_<box>.sh, which translates e.g. --decode-m1 → MPK_GG_DECODE_M1=1 prefix on the remote command). The test reads the env var directly and _run_test does NOT pop it. Local-only: the ferret faithful_eval.sh similarly hardcodes a stale "GPU 5 = KDA job" refusal — route around it via the broker gpu_pool.conf (disable the blocked local slot), don't edit the .sh.
  6. Add --num-workers / new flags to faithful_eval.py (argparse → run() → _run_test/_run_group_gemm → the test env), NOT just an env var, so they cross the remote shim as args. (The shim then env-translates the ones the remote .sh rejects.)
  7. Per-(shape) one number. The gate measures ONE shape per invocation; mapping it onto multiple differently-shaped --configs mis-gates. Pass --configs matching --shape.

Dispatching a kernel agent against the gate

When handing a kernel to KDA/Ferret (the kda-kernel-agent / ferret-kernel-agent), the faithful gate IS the promotion authority. Give the agent: the exact gpu_gate.sh command (with --kind/--shape/--num-workers/--decode-m1), the measured in-tree baseline slowCTA (measure it FIRST), the vLLM ref + the ≥20% target (≤ vLLM_ref ÷ 1.2), the per-worker-count sweep set, and the barrier-free-tile-loop + alignment crash-safety constraints (gotcha #3). The agent's candidate is overlaid via --kernel; the shadow MIRAGE_ROOT keeps the prod tree clean.

Output: the FAITHFUL_RESULT contract

Each gate run prints one machine-parseable line the bridge/agents consume, e.g.:

FAITHFUL_RESULT {"kind":"...","shape":"...","num_active":8,"candidate":{"slowCTA_us":..,"wall_us":..,"sumCTA_us":..,"cos":1.0,"status":"PASS"},"baseline_in_tree":{...},"slowCTA_ratio_base_over_cand":..,"verdict":"PASS_GATE|BELOW_BAR|TIMEOUT_HANG|INVALID"}

Keep this shape stable — the ferret/KDA loop greps it.

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/mpk-faithful-gate of mirage-project/mirage.

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

Mpk Faithful Gate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mpk Faithful Gate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mpk Faithful Gate this skillmirage-project/mirage2.5k—~2.6kAutomated safety check: PassApache-2.0
Blockless Extension E2EFreakStudioCN/mpy-hardware-extension154—~1.2kAutomated safety check: NotesCustom licence
Mmsp DevPrism-Shadow/model-message-stream-protocol113—~7.2kAutomated safety check: PassApache-2.0
Local Platform E2Ecomputesdk/benchmarks126—~3kAutomated safety check: NotesMIT
playwright-cli Browser Automationgithub/gh-aw5.4k25 repos~2.8kAutomated safety check: PassMIT
Kane CLI Browser TestingLambdaTest/kane-cli249—~8.4kAutomated safety check: PassApache-2.0

Similar skills

  • Blockless Extension E2E

    FreakStudioCN/mpy-hardware-extension

    Run and debug the Blockless VS Code extension release gate: CI-equivalent API and extension tests, V0 protocol smoke, live DeepSeek full-stack e2e, VSIX packaging, local reinstall, direct…

    154 GitHub stars~1.2k tokensUpdated 13 days ago
    Testing & QAAuto-check: notes
  • Mmsp Dev

    Prism-Shadow/model-message-stream-protocol

    Fixed workflow for developing MMSP itself — adding or updating model support, and changing its pages.

    113 GitHub stars~7.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Local Platform E2E

    computesdk/benchmarks

    Stand up benchmarks-platform locally (Postgres + MinIO + ClickHouse in docker) and run a real @benchsdk/runner benchmark against it, with no cloud or provider credentials.

    126 GitHub stars~3k tokensUpdated yesterday
    DatabasesAuto-check: notes
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.4k GitHub starsUsed in 25 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    249 GitHub stars~8.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Tabbit

    Tabbit-Browser/dsh-tabbit

    Control Tabbit Browser in a task-isolated Playwright workspace; never switch browser backends.

    104 GitHub stars~780 tokensUpdated 3 days ago
    Testing & QAAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • V2 Model Support

    mirage-project/mirage

    End-to-end pipeline for adding or porting a model to MPK Runtime-V2 — from a compute-graph spec (shapes + draw.io graph + HF checkpoint + TP/EP plan) to a working multi-GPU demo.

    2.5k GitHub stars~5.3k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Mpk Faithful Gate

What does Mpk Faithful Gate do?

Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape. Mpk Faithful Gate is an agent skill from mirage-project/mirage. Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

When should I use Mpk Faithful Gate?

Mpk Faithful Gate fits situations like: tasks that involve Meeting notes and agendas; tasks that involve End-to-end testing.

How do I install Mpk Faithful Gate in Claude Code?

Run `npx skills add mirage-project/mirage --skill mpk-faithful-gate -a claude-code`. Or copy the skill folder (.claude/skills/mpk-faithful-gate in mirage-project/mirage) into .claude/skills/mpk-faithful-gate in your project. Claude Code loads it when a task matches its description.

How do I install Mpk Faithful Gate in Codex?

Run `npx skills add mirage-project/mirage --skill mpk-faithful-gate -a codex`. Or copy the skill folder (.claude/skills/mpk-faithful-gate in mirage-project/mirage) into .agents/skills/mpk-faithful-gate in your project. Codex loads it when a task matches its description.

Can I use Mpk Faithful Gate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill mpk-faithful-gate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mpk-faithful-gate, .gemini/skills/mpk-faithful-gate, .github/skills/mpk-faithful-gate and .opencode/skills/mpk-faithful-gate in your project.

What does Mpk Faithful Gate need to run?

Going by SKILL.md and its folder, Mpk Faithful Gate needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Mpk Faithful Gate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Mpk Faithful Gate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Mpk Faithful Gate use?

Mpk Faithful Gate is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mpk Faithful Gate use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mpk Faithful Gate?

Skills that share tags, products or a category with Mpk Faithful Gate: Blockless Extension E2E (FreakStudioCN/mpy-hardware-extension, 154 stars), Mmsp Dev (Prism-Shadow/model-message-stream-protocol, 113 stars), Local Platform E2E (computesdk/benchmarks, 126 stars) and playwright-cli Browser Automation (github/gh-aw, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mpk Faithful Gate?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.