Running Tests
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.
$ npx skills add mirage-project/mirage --skill test-mode -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mirage-project/mirage test-mode --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-mode .claude/skills/test-mode && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .claude/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-modeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mirage-project/mirage --skill test-mode -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mirage-project/mirage test-mode --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/test-mode .agents/skills/test-mode && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .agents/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill test-mode -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mirage-project/mirage test-mode --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/test-mode .cursor/skills/test-mode && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .cursor/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mirage-project/mirage.git --path .claude/skills/test-mode--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mirage-project/mirage --skill test-mode -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mirage-project/mirage test-mode --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/test-mode .gemini/skills/test-mode && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .gemini/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mirage-project/mirage test-modeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mirage-project/mirage --skill test-mode -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/test-mode .github/skills/test-mode && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .github/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill test-mode -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mirage-project/mirage test-mode --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/test-mode .opencode/skills/test-mode && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-mode" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/test-mode into .opencode/skills/test-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-mode", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-modeGuide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.
Test Mode is an agent skill from mirage-project/mirage. Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline. Use when writing layer tests, debugging kernel output, or validating a new task end-to-end.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Unit testing and Deep learning. It works with PyTorch. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Mode loads about 4.6k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 1,461 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,461 words, ~4,575 tokens.
.claude/skills/test-mode/SKILL.md (or your agent's skills folder).Test mode compiles and runs an MPK task graph exactly once and exits. It exercises the full pipeline — Python layer API, task registration, C++ code generation, nvcc compilation, runtime dispatch, and the persistent runtime's metadata setup (init_kernel + prepare_next_batch) — making it the primary tool for validating that a new layer or task works end-to-end.
Test mode is selected by setting params["test_mode"] = True at construction time. Internally this defines -DMPK_TEST_MODE for the launcher build, which:
qo_indptr_buffer, paged_kv_*, input_tokens, etc.).init_request_resources() and prepare_next_batch run normally — the same code paths production uses.prepare_next_batch's always-finalize shortcut on iter 1, which returns false and terminates the scheduler after exactly one task-graph pass.Every test mode file must include a PyTorch reference implementation that computes the same operation, and must compare the MPK output against it numerically. A test that only runs the kernel without checking correctness is not a valid test — it only proves the kernel doesn't crash.
The reference should:
torch ops (or torch.nn.functional) to implement the same math the layer performs.float32 if needed for a trustworthy reference).atol=1e-2, rtol=1e-2; fp16 similar; fp32 much tighter.Use torch.testing.assert_close(out, ref, atol=..., rtol=...) and/or print (out - ref).abs().max() so failures surface immediately rather than silently producing wrong numbers.
pytorch_reference.pyPer-layer test_mode files must import their PyTorch reference from pytorch_reference.py in the same folder, not redefine it inline. The folder layout is tests/runtime_python/<arch>/sm100_<layer>/, with one pytorch_reference.py per folder containing one function per in-scope layer. Both the new test_mode test (test_<layer>_testmode.py) and the existing kernel-wrapper test (test_<layer>.py) import from the same file, so they stay aligned on a single canonical reference.
If pytorch_reference.py does not yet exist for the layer, create it. If a kernel-wrapper test already exists with an inline reference, extract that reference into pytorch_reference.py and refactor the kernel-wrapper test to import from it.
import torch
import mirage
from mirage.mpk.persistent_kernel import PersistentKernel
# 1. Configure
num_workers, num_schedulers = mirage.get_configurations_from_gpu(0)
params = PersistentKernel.get_default_init_parameters()
params["test_mode"] = True
params["num_workers"] = num_workers
params["num_local_schedulers"] = num_schedulers
pk = PersistentKernel(**params)
# 2. Create tensors and attach (both inputs AND outputs)
x = torch.randn(16, 4096, dtype=torch.bfloat16, device="cuda")
w = torch.randn(4096, dtype=torch.bfloat16, device="cuda")
out = torch.zeros(16, 4096, dtype=torch.bfloat16, device="cuda")
x_dt = pk.attach_input(x, name="x")
w_dt = pk.attach_input(w, name="w")
out_dt = pk.attach_input(out, name="out")
# 3. Build layer(s)
block_dim = (256, 1, 1) if pk.target_cc >= 90 else (128, 1, 1)
pk.rmsnorm_layer(input=x_dt, weight=w_dt, output=out_dt,
grid_dim=(16, 1, 1), block_dim=block_dim)
# 4. Compile and run — same call as production; MPK_TEST_MODE is baked into the .so
pk.compile(output_dir="./test_output") # saves .cu and .json for debugging
pk()
torch.cuda.synchronize()
# 5. Compare against a PyTorch reference
def torch_rmsnorm(x, w, eps=1e-6):
x_f32 = x.to(torch.float32)
rms = x_f32.pow(2).mean(dim=-1, keepdim=True).add(eps).rsqrt()
return (x_f32 * rms * w.to(torch.float32)).to(x.dtype)
ref = torch_rmsnorm(x, w)
print("Max diff:", (out - ref).abs().max().item())
torch.testing.assert_close(out, ref, atol=1e-2, rtol=1e-2)
# 6. Cleanup
pk.finalize()The launcher's first scheduler event is always EVENT_END_OF_TASK_GRAPH, so prepare_next_batch fires before iter 0. Concretely:
init_kernel: zero step / request_ids / qo_indptr / paged_kv_indptr; seed page_queue
1st END_OF_TASK_GRAPH (iter_num=0): prepare_next_batch fills meta tensors for iter 0 → returns true
iter 0: the test — layer-under-test runs with valid meta tensors
2nd END_OF_TASK_GRAPH (iter_num=1): prepare_next_batch finalizes (MPK_TEST_MODE always-finalize)
→ no new prefills (next_request_id == total_num_requests)
→ returns false → terminateSo tests of meta-tensor-dependent layers (paged attention, MoE routing, embedding, sampling) work — they read the values that prepare_next_batch wrote.
PersistentKernel.get_default_init_parameters() (classmethod)Returns a dict with safe defaults for test mode. You must set params["test_mode"] = True — it is not in the defaults.
Commonly overridden keys:
| Key | Default | When to override |
|---|---|---|
test_mode | (not present) | Always set to True |
num_workers | 1 | Set from mirage.get_configurations_from_gpu(0) |
num_local_schedulers | 4 | Set from mirage.get_configurations_from_gpu(0) |
max_num_batched_tokens | 1 | Set to your test's batch size if the task kernel uses this compile-time constant |
max_num_batched_requests | 1 | Same as above |
max_num_pages / page_size / max_seq_length | 1 | Bump these so prepare_next_batch can fit your prefill (max_num_pages * page_size >= prompt_length) |
world_size / mpi_rank | 1 / 0 | For multi-GPU tests; set from mpi4py.MPI.COMM_WORLD |
use_cutlass_kernel | False | Set True if your layer uses CUTLASS-based kernels |
meta_tensors | {} | Auto-defaulted; override only the entries that drive your test scenario (typically prompt_lengths and/or tokens) — see "Meta-Tensor Defaults" below |
mirage.get_configurations_from_gpu(rank)Returns (num_workers, num_schedulers) tuned for the GPU at the given rank. Always use this rather than hardcoding — the values depend on SM count and architecture.
pk.attach_input(tensor, name)Registers a PyTorch CUDA tensor with the computation graph. Returns a DTensor for use in layer calls.
pk.compile(output_dir=None)Generates CUDA code, compiles with nvcc, loads the resulting .so module.
output_dir to save test_rank0.cu and task_graph_rank0.json — essential for debugging compilation errors or incorrect results.pk() — Launch the kernelSame call as production. In test mode the launcher was compiled with -DMPK_TEST_MODE so it terminates after one task-graph pass. The previous pk.run_test_mode() method has been removed; use pk() directly.
compile().torch.cuda.synchronize() before reading output tensors.default_stream=stream kwarg if you don't want the current stream.params["profiler_tensor"] and optional params["trace_name"] before compile(). After pk() returns, both <trace_name>.perfetto-trace and <trace_name>.csv are written. See "Profiling" below.pk.finalize()Frees GPU resources (queues, events, task/event storage). Call when done.
Test mode auto-allocates any of the 10 meta tensors that you don't pass:
| Key | Default shape | Default dtype | Default content |
|---|---|---|---|
tokens | (1, max_seq_length) | int64 | zeros |
step | (total_num_requests,) | int32 | zeros |
prompt_lengths | (total_num_requests,) | int32 | filled with max_num_batched_tokens |
input_tokens | (max_num_batched_tokens,) | int64 | zeros (filled by prepare_next_batch) |
output_tokens | (max_num_batched_tokens,) | int64 | zeros |
num_new_tokens | (1,) | int32 | zeros |
qo_indptr_buffer | (max_num_batched_requests + 1,) | int32 | zeros (filled by prepare_next_batch) |
paged_kv_indptr_buffer | (max_num_batched_requests + 1,) | int32 | zeros (filled by prepare_next_batch) |
paged_kv_indices_buffer | (max_num_pages,) | int32 | zeros (filled by prepare_next_batch) |
paged_kv_last_page_len_buffer | (max_num_batched_requests,) | int32 | zeros (filled by prepare_next_batch) |
total_num_requests is derived from tokens.shape[0] (defaults to 1).
Override only what your test scenario requires. Typical patterns:
# Single prefill of length N (controls qo_indptr_buffer / paged_kv_* via prepare_next_batch)
params["meta_tensors"] = {
"prompt_lengths": torch.tensor([N], dtype=torch.int32, device="cuda"),
}
# Specific prompt content (e.g. for embedding-layer tests that read input_tokens)
params["meta_tensors"] = {
"prompt_lengths": torch.tensor([N], dtype=torch.int32, device="cuda"),
"tokens": torch.tensor([[101, 7592, 2088, ...]], dtype=torch.int64, device="cuda"),
}
# Multi-request batch — total_num_requests inferred from tokens.shape[0]
params["meta_tensors"] = {
"tokens": torch.zeros((4, max_seq_length), dtype=torch.int64, device="cuda"),
"prompt_lengths": torch.tensor([16, 8, 32, 4], dtype=torch.int32, device="cuda"),
}The shape/dtype assertions that production runs through (e.g. tokens.shape[1] == max_seq_length, prompt_lengths.dtype == int32) all run in test mode too — defaults satisfy them by construction; user overrides will fail loudly if they don't match.
Multiple layers can be chained with intermediate tensors. From the Qwen3 dense MLP pattern:
# Gate+Up linear → SiLU-Mul → Down+Residual
# Attach weights separately, then shuffle for interleaved gate/up layout
w_gate_dt = pk.attach_input(w_gate, name="w_gate")
w_up_dt = pk.attach_input(w_up, name="w_up")
w_gatedup_dt = pk.shuffle_tensors(
inputs=[w_gate_dt, w_up_dt],
shuffled_dim=0,
num_groups=num_tasks // 2,
name="w_gatedup",
)
# Layer 1: Gate+Up fused linear
pk.linear_layer(input=input_dt, weight=w_gatedup_dt, output=mlp_mid_dt,
grid_dim=(num_tasks, 1, 1), block_dim=block_dim)
# Layer 2: SiLU activation * element-wise multiply
pk.silu_mul_layer(input=mlp_mid_dt, output=silu_out_dt,
grid_dim=(num_tasks // 2, 1, 1), block_dim=block_dim)
# Layer 3: Down projection + residual add
pk.linear_with_residual_layer(input=silu_out_dt, weight=w_down_dt,
residual=residual_dt, output=mlp_out_dt,
grid_dim=(hidden_size // 64, 1, 1), block_dim=block_dim)Key pattern: intermediate tensors (mlp_mid, silu_out) are pre-allocated and attached via attach_input so they can be inspected after execution if needed. For a runnable multi-task test see tests/runtime_python/test_mode/test_diamond_fork_join_testmode.py.
Test mode supports world_size > 1. Each rank is independent — auto-defaults are deterministic functions of kernel params, so they produce identical values on every rank.
from mpi4py import MPI
comm = MPI.COMM_WORLD
world_size = comm.Get_size()
rank = comm.Get_rank()
torch.cuda.set_device(rank)
params = PersistentKernel.get_default_init_parameters()
params["test_mode"] = True
params["world_size"] = world_size
params["mpi_rank"] = rank
# ... rest of setup is the same as single-GPU
pk = PersistentKernel(**params)
# ... attach + register layers (incl. NVSHMEM ops like pk.allreduce_layer)
pk.compile(output_dir=...)
pk()
torch.cuda.synchronize()Launch convention:
LD_PRELOAD=$NVSHMEM_HOME/lib/libnvshmem_host.so \
mpirun --np 2 -x LD_PRELOAD -x LD_LIBRARY_PATH -x NVSHMEM_HOME \
python tests/runtime_python/test_mode/test_multigpu_rmsnorm_testmode.pyThe LD_PRELOAD is required so dlopen()-loaded launcher modules resolve nvshmem_selected_device_transport and other NVSHMEM-versioned symbols. This is an existing NVSHMEM 3.x quirk, unrelated to test mode.
Test mode supports profiling because pk() runs the standard __call__ path. Profiling is opt-in: pass a profiler_tensor and the post-run hook writes both a Perfetto trace (for human inspection) and a CSV (for programmatic queries).
params = PersistentKernel.get_default_init_parameters()
params["test_mode"] = True
params["trace_name"] = "rmsnorm_smoke" # optional; defaults to f"mirage_{mpi_rank}"
params["profiler_tensor"] = torch.zeros(
3000 * 128, dtype=torch.uint64, device="cuda"
)
pk = PersistentKernel(**params)
# ... attach tensors, register layers, pk.compile(...) as usual ...
pk()
torch.cuda.synchronize()
# Two files now exist alongside the script:
# rmsnorm_smoke.perfetto-trace ← drag into ui.perfetto.dev
# rmsnorm_smoke.csv ← query with scripts/parse_profile.pyThe profiler buffer must be uint64 on CUDA. 3000 * 128 entries is the conventional size used by the demos and is plenty for short test-mode runs. Each task event consumes 2 entries (BEGIN + END); buffer overflow would surface later as a RuntimeError("dangling BEGIN ...") from the CSV writer.
One row per fully-paired task event (and one row per kInstant event with duration_ns=0):
| Column | Meaning |
|---|---|
task_type_id | TaskType enum value (e.g. 253) |
task_type_name | Symbolic name (e.g. TASK_LINEAR_SM100) |
block_idx, group_idx | Worker that executed the event |
event_no | Per-worker execution counter |
begin_ts, end_ts | Raw 32-bit %globaltimer_lo values (ns, wraps every ~4.3 s) |
duration_ns | (end_ts - begin_ts) mod 2^32 |
scripts/parse_profile.pyAll output is JSON; errors print {"error": "..."} and exit 2.
# What task types ran, with event counts
python scripts/parse_profile.py rmsnorm_smoke.csv --list
# Average runtime of one task type
python scripts/parse_profile.py rmsnorm_smoke.csv TASK_RMS_NORM_HOPPER --stat avg
# Min / max / avg / median in one shot
python scripts/parse_profile.py rmsnorm_smoke.csv TASK_LINEAR_SM100 --stat all
# Numeric task-type id is also accepted
python scripts/parse_profile.py rmsnorm_smoke.csv 253 --stat minSample output:
{"task_type": "TASK_LINEAR_SM100", "count": 7040, "min_ns": 11776, "max_ns": 193856, "avg_ns": 26731.58, "median_ns": 31184.0}For finer-grained analysis (per-worker breakdown, percentiles, outliers), pandas.read_csv(...) the file directly — the schema is stable and column names speak for themselves.
prepare_next_batch returns false on its second call. No multi-iteration scheduling.max_num_batched_tokens, max_num_batched_requests, max_num_pages, max_seq_length, etc.), so bump those if your test needs larger buffers.MPK_TEST_MODE is a compile-time flag — switching test_mode between True and False requires re-running pk.compile(); the same launcher .so isn't reusable across modes.Compilation fails:
<output_dir>/test_rank0.cu for the generated code. Search for your task name in the _execute_task() function.<output_dir>/task_graph_rank0.json for the task graph. This file might be extremely long; don't read it raw — use scripts/parse_task_graph.py.Incorrect dimension splitting:
input_map for each associated tensor to specify how dimensions are split across the grid. If grid/block dimensions don't divide tensor dimensions correctly, the kernel may read/write out of bounds, producing NaNs or wrong results.Incomplete task attributes:
runtime.cc. Missing/incorrect attributes cause undefined behavior or compilation errors.Kernel hangs / never terminates:
total_num_requests is set to match the number of in-flight test requests (typically 1, derived from tokens.shape[0]). If next_request_id never reaches total_num_requests, prepare_next_batch will keep returning true and iterations will not stop.mode is "offline" (the default). MPK_TEST_MODE is designed to layer on top of MODE_OFFLINE's prepare_next_batch; other modes are not supported.Verifying that prepare_next_batch actually ran:
pk() returns, read back pk.meta_tensors["step"][0]. It should equal prompt_lengths[0] — prepare_next_batch's Step 1.1 advances step by num_tokens on the second call. See test_prepare_next_batch_testmode.py for the canonical assertion.| File | What it tests |
|---|---|
tests/runtime_python/test_mode/test_rmsnorm_testmode.py | Single layer (RMSNorm), default meta tensors |
tests/runtime_python/test_mode/test_diamond_fork_join_testmode.py | Multi-task graph (synthetic fork+join) |
© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/test-mode of mirage-project/mirage.
Open the folder on GitHubat commit f9eb70c
Test Mode next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Mode this skillmirage-project/mirage | 2.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Running Testsbrendanhasz/probflow | 175 | — | ~657 | Automated safety check: Pass | MIT | |
| Ut Refactor Reviewintel/torch-xpu-ops | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | |
| Liger Kernel Devlinkedin/Liger-Kernel | 6.6k | — | ~799 | Automated safety check: Pass | BSD-2-Clause | |
| Scrub Issuepytorch/pytorch | 104k | — | ~4.6k | Automated safety check: Pass | Custom licence | |
| Fix Reproduceintel/torch-xpu-ops | 115 | — | ~7.7k | Automated safety check: Pass | Apache-2.0 |
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
intel/torch-xpu-ops
Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests.
linkedin/Liger-Kernel
Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.
pytorch/pytorch
Fetch, analyze, reproduce, and minimize GitHub issue reproductions.
intel/torch-xpu-ops
A skill your agent uses when asked to reproduce a bug, verify a nightly CI failure, or confirm a failure still exists on latest source.
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
mirage-project/mirage
Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.
mirage-project/mirage
Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).
mirage-project/mirage
A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…
mirage-project/mirage
Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.
mirage-project/mirage
A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…
mirage-project/mirage
End-to-end pipeline for adding or porting a model to MPK Runtime-V2 — from a compute-graph spec (shapes + draw.io graph + HF checkpoint + TP/EP plan) to a working multi-GPU demo.
Works with
Categories
Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline. Test Mode is an agent skill from mirage-project/mirage. Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.
Test Mode fits situations like: writing layer tests; debugging kernel output; validating a new task end-to-end.
Run `npx skills add mirage-project/mirage --skill test-mode -a claude-code`. Or copy the skill folder (.claude/skills/test-mode in mirage-project/mirage) into .claude/skills/test-mode in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mirage-project/mirage --skill test-mode -a codex`. Or copy the skill folder (.claude/skills/test-mode in mirage-project/mirage) into .agents/skills/test-mode in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill test-mode -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-mode, .gemini/skills/test-mode, .github/skills/test-mode and .opencode/skills/test-mode in your project.
Going by SKILL.md and its folder, Test Mode needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Mode is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Test Mode: Running Tests (brendanhasz/probflow, 175 stars), Ut Refactor Review (intel/torch-xpu-ops, 115 stars), Liger Kernel Dev (linkedin/Liger-Kernel, 6.6k stars) and Scrub Issue (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,541 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.
Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.