Perf Loop
ozontech/seq-db
Autonomous performance-optimization loop for seq-db driven by seqbazooka.
Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mirage-project/mirage v2-perf-iteration --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .claude/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .claude/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iterationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mirage-project/mirage v2-perf-iteration --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .agents/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .agents/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mirage-project/mirage v2-perf-iteration --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .cursor/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .cursor/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mirage-project/mirage.git --path .claude/skills/v2-perf-iteration--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mirage-project/mirage v2-perf-iteration --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .gemini/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .gemini/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mirage-project/mirage v2-perf-iterationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .github/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .github/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill v2-perf-iteration -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mirage-project/mirage v2-perf-iteration --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/v2-perf-iteration .opencode/skills/v2-perf-iteration && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "v2-perf-iteration" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/v2-perf-iteration into .opencode/skills/v2-perf-iteration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "v2-perf-iteration", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
v2-perf-iterationRuntime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.
V2 Perf Iteration is an agent skill from mirage-project/mirage. Runtime-V2 performance-iteration workflow. Use when running a perf-optimization campaign or iteration on the v2 runtime (--use-v2) — measuring a baseline, ranking bottlenecks, planning levers, implementing, re-measuring, and landing/recording the verdict. Drives the loop MEASURE→ANALYZE→PLAN→NEXT-MOVE→REVIEW→IMPLEMENT→VALIDATE+RE-MEASURE→LOOP-OR-LAND→RECORD with the mpk- subagent roster, the v2 profiler/perfetto toolchain, and the TIER-1 TP8 verdict discipline.
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/loop-agents.md`, `tools/perfetto_analyze.py` and `tools/perfetto_depgraph.py`).
It sits in Agent Workflows, covering Subagents and Performance optimization. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
V2 Perf Iteration loads about 4k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 121 tokens; SKILL.md has 1,591 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,591 words, ~3,959 tokens.
.claude/skills/v2-perf-iteration/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.This is the perf-optimization loop of the v1 multi-agent campaign that ran for months
(the old repo-root WORKFLOW.md is now a superseded stub pointing here), upgraded for
Runtime-V2's measurement reality. Siblings: v2-model-support
(bring-up; its Phase (d) is this skill), v2-kernel-writing (the per-KERNEL inner loop
this skill dispatches into when a lever is kernel-body work). The main thread (or one lead
orchestrator) runs the loop and does all edits/commits/box-ops; subagents measure-parse /
analyze / plan / review / record. No subagent dispatches another subagent, and box
operations never go inside a subagent (v2-model-support/references/box-orchestration.md).
Goal + verdict metric — pin ONE per campaign, then do not drift. The loop is
metric-agnostic; what is non-negotiable is that a single PRODUCTION verdict config is
declared up front and every lever's verdict-grade Δ is measured there. Worked example —
the concluded 2026-06/07 DSv3 campaign's instance: e2e decode tpot, bs=1, TP8 EP2, MTP off, toward
8 ms/token (SGLang 7.99 on the same box proves it reachable); v2 clean baseline
2026-07-07: 12.069 ms/tok vs v1 ~9.787; TP<8/local runs are triage only. A new campaign
(e.g. Qwen3-8B single-GPU bs=1024 throughput) writes its own goal line in this exact shape
— metric, config, target, why-reachable — and restates it in EVERY dispatch prompt.
Per-task numbers follow the TIER hierarchy
(v2-kernel-writing/references/validation-debug.md §8): TIER-1 = in-MPK
%globaltimer / per-position slowest-CTA at the production grid — the ONLY verdict tier;
faithful harness corroborates (TIER-2); cudaEvent-wall / standalone-warm are diagnostic only.
⚠️ The .claude/agents/mpk-* defs still carry v1-era DSv3 framing (150µs/MoE-layer, TP4,
per_position_grid.py, scratch/ helper scripts that are git-ignored/machine-local).
Every dispatch prompt MUST restate the CURRENT campaign goal, the v2 toolchain commands
from the quickstart below, and the artifact paths — otherwise the agent drifts to the
stale v1 pipeline (each def now carries a "V2 / new-campaign note" saying exactly this).
In-repo (travels with every clone): the parser python -m mirage.mpk.prof
(python/mirage/mpk/prof.py, tracked), the demo --profiling plumbing, this skill's
references/ + tools/. Machine-local / degradations:
scripts/v2_perfetto_export.py / perfetto_analyze.py / perfetto_depgraph.py are
UNTRACKED on this branch — a fresh clone will not have them under scripts/. Archived
copies travel in .claude/skills/v2-perf-iteration/tools/ — run them from there (or copy
back to scripts/, which stays git-ignored for local files). mirage.mpk.prof summary/check/pagewait is the tracked no-dependency fallback for text-table analysis.v2-model-support/references/ box-orchestration.md; §1-2 there are site-specific). Single-GPU campaigns run the whole
loop locally.experiment_history/ — git-ignored ⇒ empty on a fresh clone; create INDEX.md + the
journal at step 9 of the first iteration. The kernel-lever anti-loop that must survive
clones is v2-kernel-writing/references/m1-decode-evidence.md (in-repo).~/.claude/agents/: mpk-perf-analyzer, ablation-logic-reviewer,
codex-task-dispatcher) — same-account only; the rest of the roster is in-repo at
.claude/agents/. Codex MCP — machine-configured (.mcp.json is git-ignored); absent
⇒ reviews degrade to subagent-only (state it). Personal memory — optional context.~/ref_vllm_sglang.md (analyzer/planner per-kernel reference table) — machine-local;
absent ⇒ rank gaps against the SGLang/vLLM numbers recorded in the campaign goal line and
in-repo docs, and say the external table was unavailable.(1) MEASURE ─▶ (2) ANALYZE ─▶ (3) PLAN ─▶ (4) NEXT-MOVE ─▶ (5) REVIEW-BEFORE-ACT
▲ │
│ (new bottleneck / re-plan) ▼
(9) RECORD ◀─ (8) LOOP-OR-LAND ◀─ (7) VALIDATE + RE-MEASURE ◀─ (6) IMPLEMENT| # | Step | Who | In → Out |
|---|---|---|---|
| 1 | MEASURE | main thread (box) + mpk-profiler discipline | profiled --use-v2 run → per-task-type consumer-body table + per-position slowCTA + tpot |
| 2 | ANALYZE | mpk-perf-analyzer (Opus) | trace/tables → ranked gaps vs refs, kernel-level vs system-level split |
| 3 | PLAN | mpk-optimization-planner (Opus) | report + history → µs-derived ranked batch plan, [MAIN|ENGINEER|FERRET|CODEX] tags, 3-round Codex convergence |
| 4 | NEXT-MOVE | mpk-iterator | plan + report → reflection + the single next move w/ falsifiable predicted Δ |
| 5 | REVIEW | ablation-logic-reviewer + Codex MCP | the move/conclusion → first-principles audit (MANDATORY before acting) |
| 6 | IMPLEMENT | main thread (route by tag) | env-gated default-OFF change |
| 7 | VALIDATE | mpk-correctness-gate, then re-run (1) | PASS/FAIL + the TIER-1 TP8 Δ |
| 8 | LOOP-OR-LAND | main thread + mpk-commit-reviewer | commit (WIN) / revert+INDEX (NULL/REGRESS) / re-plan |
| 9 | RECORD | mpk-memory-keeper | journal + INDEX row + personal-memory lesson |
Full roster card (what each agent consumes/returns + key discipline): references/loop-agents.md.
1. MEASURE. The mpk-profiler pattern updated for v2: GPU-safety pre-flight → the
canonical config → profiled run → parse → cleanup+zombie-guard → report. At TP8 the box
session belongs to the MAIN THREAD (setup/poll split, retries, verify-STOPPED — follow
v2-model-support/references/box-orchestration.md; do NOT re-derive box mechanics here):
rsync → --use-v2 --profiling run (quickstart below) → retrieve the per-rank
*_v2prof.npy → v2_perfetto_export.py + python -m mirage.mpk.prof summary/check.
Dispatch mpk-profiler itself only for local-GPU triage runs or offline parsing of an
already-retrieved buffer — never for box ops. Report = tpot (n-of-N) + the per-task-type
consumer-body table (µs/instance × count × layers = ms and % of tpot) + per-position slowCTA
2. ANALYZE (optional on small iterations, mandatory on a fresh baseline). Ranked TODO
split kernel-level (body ≫ SOTA ref at M=1 shape) vs system-level (dep-wait, page-wait,
role-coordination overhead, AR/skew — the v2 runtime-overhead axis that made v2 12.07 vs v1
9.79). It reads experiment_history/INDEX.md first.
3. PLAN. µs-derived ranked batch plan; every lever: target position + arithmetic +
on-critical-path reasoning + correctness risk + dispatch tag. Anti-loop is MANDATORY: check
experiment_history/INDEX.md AND v2-kernel-writing/references/m1-decode-evidence.md
(the DEAD/WIN/UNTESTED map) — a dead lever is only re-proposable by naming what's different.
4. NEXT-MOVE. One concrete single-iteration move with a falsifiable predicted Δ ("tpot 12.07 → ~11.5 because attn consumer body 108 → ~99µs and attn is on the CP").
5. REVIEW-BEFORE-ACT (MANDATORY, user-locked). Every non-trivial conclusion — root-cause,
ablation verdict, dead/alive, ceiling, perf claim — goes through ablation-logic-reviewer
(first-principles re-derivation) AND a Codex MCP cross-check (mcp__codex__codex, DEFAULT
params) BEFORE you act on it or report it settled. When stuck: detailed multi-turn Codex
discussion BEFORE escalating to the user (escalate only when both agree there's no room).
6. IMPLEMENT (main thread routes by tag):
v2-kernel-writing skill
(SPEC→IMPLEMENT→WIRE→VALIDATE→PERF→REVIEW; its Stage 2 dispatches v2-kernel-engineer,
and ferret-kernel-agent/kda-kernel-agent are its beat-a-target engines).ferret-kernel-agent
(frozen-gate autonomous loop) or kda-kernel-agent (verdict-grade honest transfer);
routing one-liners in references/loop-agents.md.codex-task-dispatcher.7. VALIDATE + RE-MEASURE. Math-changing → mpk-correctness-gate (test-mode + non-null
MoE + the TP8-nondeterminism-aware gates: deterministic canary, poison-fill,
coherence-in-envelope — see validation-debug.md §7). Math-neutral → token-identity on a
deterministic config. Then re-run step (1); the verdict is the TIER-1 TP8 number, and the
predicted Δ is confirmed or refuted — say which.
8. LOOP-OR-LAND decision rules:
mpk-commit-reviewer gate (staged-path,
default-OFF byte-identity, message mechanism+Δ+sign-off) then commit. BLOCK → fix, re-gate.9. RECORD. mpk-memory-keeper appends the journal entry + INDEX one-liner (esp.
NULL/REGRESS) + folds structural lessons into personal memory. This closes the anti-loop:
steps (2)-(4) read what (9) wrote.
Profiled run (add to the canonical demo invocation, per-rank under mpirun):
demo/deepseek_v3/demo.py ... --use-v2 --profiling --trace-name <tag> \
[--profile-start-step N] # profile steady-state, not warmup--profiling compiles with -DMPK_ENABLE_PROFILING (persistent_kernel.py:504).V2_PROF_BUF_ENTRIES = 120000*128 (15.36M) uint64 entries
— demo.py auto-sizes this (V2_PROFILER_BUFFER_ENTRIES, demo.py:30) and HARD-RAISES if
MPK_PROFILER_BUFFER_ENTRIES is set smaller (demo.py:64-75): a smaller buffer = silent
device OOB (the v2 profiler writes tail accumulators at absolute end-of-buffer indices).
Don't override it; don't "fix" a >256-worker abort by shrinking the buffer.V2_PROF_WINDOW_ITERS = 25 decode steps are recorded; 8 tracks/SM
(consumer/loader/launcher/storer/controller + 3 phase tracks: dep-wait/page-wait/>2µs).<tag>_rank<r>_v2prof.npy (raw buffer) + <tag>_rank<r>.perfetto-trace
(v1 exporter output — garbage for v2 buffers, ignore it) + an auto text summary at
run end (prof.print_run_summary).Parse (usually rank0's npy, retrieved from the box):
python -m mirage.mpk.prof check <npy> # structural gate: needs "ALL CHECKS PASS",
# dropped events MUST be 0 (else trace truncated)
python -m mirage.mpk.prof summary <npy> # per-task-type consumer table: n/SM/it, dep-wait,
# suffix, body+disp, win-mean/p50 + busy ms/SM/step
python -m mirage.mpk.prof pagewait <npy> # page-protocol serialization (dead prefetch)
python scripts/v2_perfetto_export.py <npy> <out.json> --last-steps 2 [--sm N]
# Chrome-JSON for ui.perfetto.dev; NEVER --full
# (full window OOMs the UI); deeper analysis:
# scripts/perfetto_analyze.py / perfetto_depgraph.py
# (fresh clone: these are untracked — run the
# archived copies in this skill's tools/ dir)What number to quote: per-invocation task latency = ONE consumer slice; per-task-type
body = the consumer-group windows (summary's table / the consumer track in perfetto).
Never average the role tracks. Headline = e2e tpot + the per-task-type decomposition.
Worked example (2026-07-09 profile): attn 108µs × 61 layers ≈ 58% of tpot; ffn
(52+14)µs × 58 MoE layers ≈ 30%; AR ≈ 6% → the attn consumer body is the dominant axis,
AR is not — that ranking IS the plan input.
Hang/crash during a profiled run: the historical profiled-only wedges were the v2 runtime
races, ALL FIXED 2026-07-16 (689dadc5/7d271a01/025029a1/7b6ae2bb; former wedge windows
pass post-fix — see v2-kernel-writing/references/validation-debug.md §5.1), so profiled
measurement is first-class again; a hang on a ≥7b6ae2bb tree is a NEW bug. Watchdog
-DMPK_V2_BREADCRUMB + MPK_V2_HANG_WATCHDOG_S=<s> names the hung task; crash →
compute-sanitizer memcheck is ground truth (breadcrumb in-flight counts are base-rate
artifacts). Full triage table: validation-debug.md §5. Remember: instrumentation changes
tpot (breadcrumb cost ~5.3ms on full-61L) — never quote an instrumented run as the baseline.
| Doc | Content |
|---|---|
references/loop-agents.md | Roster card: every loop agent + the kernel-perf engines + routing |
tools/ | Archived copies of the untracked v2 perfetto toolchain (v2_perfetto_export.py, perfetto_analyze.py, perfetto_depgraph.py) — the clone-safe way to run them |
../v2-kernel-writing/references/validation-debug.md | TIER hierarchy §8, profiler contract §9, hang/crash triage §5 |
../v2-kernel-writing/references/m1-decode-evidence.md | The DEAD/WIN/UNTESTED anti-loop map for kernel levers |
../v2-model-support/references/box-orchestration.md | Box session playbook (TP8 runs live here) |
experiment_history/README.md + INDEX.md | The durable log contract + the anti-loop source |
© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in .claude/skills/v2-perf-iteration of mirage-project/mirage.
Open the folder on GitHubat commit f9eb70c
V2 Perf Iteration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| V2 Perf Iteration this skillmirage-project/mirage | 2.5k | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Perf Loopozontech/seq-db | 133 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Session Profilertamdogood/builder-essential-skills | 220 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Tracelens Analysis Orchestratoramd/skills | 406 | — | ~760 | Automated safety check: Pass | MIT | |
| O2 Review Loopopenobserve/openobserve | 22k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | |
| Clone App Pat Proper-simmons/clone-app-pat-pro-public | 259 | — | ~1.9k | Automated safety check: Notes | None |
ozontech/seq-db
Autonomous performance-optimization loop for seq-db driven by seqbazooka.
tamdogood/builder-essential-skills
Profile and debug Hermes sessions from their JSONL transcripts.
amd/skills
Orchestrates modular PyTorch profiler trace analysis with TraceLens: generates perf reports, prepares category data, runs system-level and compute-kernel subagents in parallel, validates outputs…
openobserve/openobserve
Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.
per-simmons/clone-app-pat-pro-public
Clones any web app pixel-for-pixel from a URL. An agent skill from per-simmons/clone-app-pat-pro-public.
CherryHQ/cherry-studio
Delegates one bounded repository task to Kimi Code in non-interactive prompt mode and reads back the final result from its JSON event stream.
mirage-project/mirage
Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).
mirage-project/mirage
A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…
mirage-project/mirage
Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.
mirage-project/mirage
A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…
mirage-project/mirage
Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.
mirage-project/mirage
End-to-end pipeline for adding or porting a model to MPK Runtime-V2 — from a compute-graph spec (shapes + draw.io graph + HF checkpoint + TP/EP plan) to a working multi-GPU demo.
Categories
Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage. V2 Perf Iteration is an agent skill from mirage-project/mirage. Runtime-V2 performance-iteration workflow.
V2 Perf Iteration fits situations like: running a perf-optimization campaign; iteration on the v2 runtime (--use-v; — measuring a baseline; ranking bottlenecks.
Run `npx skills add mirage-project/mirage --skill v2-perf-iteration -a claude-code`. Or copy the skill folder (.claude/skills/v2-perf-iteration in mirage-project/mirage) into .claude/skills/v2-perf-iteration in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mirage-project/mirage --skill v2-perf-iteration -a codex`. Or copy the skill folder (.claude/skills/v2-perf-iteration in mirage-project/mirage) into .agents/skills/v2-perf-iteration in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill v2-perf-iteration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/v2-perf-iteration, .gemini/skills/v2-perf-iteration, .github/skills/v2-perf-iteration and .opencode/skills/v2-perf-iteration in your project.
Going by SKILL.md and its folder, V2 Perf Iteration needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
V2 Perf Iteration is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with V2 Perf Iteration: Perf Loop (ozontech/seq-db, 133 stars), Session Profiler (tamdogood/builder-essential-skills, 220 stars), Tracelens Analysis Orchestrator (amd/skills, 406 stars) and O2 Review Loop (openobserve/openobserve, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,543 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.
Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.