Agent skill

Ferret Kernel System

by mirage-project in mirage-project/mirage

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…

Apache-2.0Auto-check passedAgent Workflows

Install Ferret Kernel System

skills CLI
$ npx skills add mirage-project/mirage --skill ferret-kernel-system -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mirage-project/mirage ferret-kernel-system --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ferret-kernel-system .claude/skills/ferret-kernel-system && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ferret-kernel-system
GitHub stars
2.5k
Token cost
~1.3k tokens
SKILL.md length
566 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…

  • Works in 6 steps: TARGET — the exact MPK op + what it… → REAL-MATH CONTRACT — enumerate EVERY… → SHAPES — exact (M,K,N) at the real TP/EP… → …
  • — phrasings like dispatch ferret to optimize a kernel
  • SKILL.md covers The problem it solves (read…, The 3 agents (in…, How to invoke (the main-thread… and The invariants that make it…, plus 2 more sections
  • Calls claude

What it does

Ferret Kernel System is an agent skill from mirage-project/mirage. Use when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer cheating its own correctness test. This is the frozen-gate Ferret system — a 3-agent orchestration (dispatcher → independent test-writer → optimizer) that makes "the optimizer ships a simplified-math kernel and self-reports a passing test" structurally impossible. Triggers — phrasings like "dispatch ferret to optimize a kernel", "have an…

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `DESIGN.md`).

It sits in Agent Workflows, covering GPU and accelerator computing and Multi-agent orchestration. It works with CUDA. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.

When your agent uses it

  • — phrasings like dispatch ferret to optimize a kernel
  • Have an agent write a kernel but guarantee it doesnt simplify the math
  • Generate a kernel the frozen-gate way: the user/main-thread wants a kernel optimized
  • Generated by an agent

Example prompts

  • “structurally impossible. Triggers — phrasings like”
  • “have an agent write a kernel but guarantee it doesn”
  • “generate a kernel the frozen-gate way”
  • “/ferret-kernel-system”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. TARGET — the exact MPK op + what it replaces; the baseline = the kernel being
  2. REAL-MATH CONTRACT — enumerate EVERY step the kernel must compute, NO
  3. SHAPES — exact (M,K,N) at the real TP/EP regime (derive from the builder/weights).
  4. PRODUCTION COMPILE FLAGS — -rdc=true / MPK_FORCE_RDC_TRUE=1, arch sm_100a,
  5. ABI — the device task_impl signature + NS/NE.
  6. CANONICAL REFERENCE SOURCE — the already-trusted oracle to compare against

What it can do on your machine

Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ferret Kernel System loads about 1.3k tokens when it runs. Until then it costs about 202 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~202
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 566 words, ~1,266 tokens.

Download SKILL.mdSave it as .claude/skills/ferret-kernel-system/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ferret-kernel-system
description
Use when you need a NEW or optimized MPK CUDA kernel (a per-task `.cuh` under include/mirage/persistent_kernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer cheating its own correctness test. This is the frozen-gate Ferret system — a 3-agent orchestration (dispatcher → independent test-writer → optimizer) that makes "the optimizer ships a simplified-math kernel and self-reports a passing test" structurally impossible. Triggers — phrasings like "dispatch ferret to optimize a kernel", "have an agent write a kernel but guarantee it doesn't simplify the math", "generate a kernel the frozen-gate way": the user/main-thread wants a kernel optimized or generated by an agent, especially when correctness/no-simplification matters more than fire-and-forget speed.

Ferret frozen-gate kernel system — how to invoke + why

The problem it solves (read first)

The OLD ferret wrote its OWN correctness test. It once shipped a DeepSeek-V3 attention kernel with SIMPLIFIED math (theta-10000 rope not YaRN, 1/sqrt(576) not the YaRN mscale scale, head-sum o_proj skipping W_UV, no kv_a_layernorm) and self-reported cosine 1.0 — against its own simplified reference. Marking its own homework. The refactor takes judging + constraints OUT of the optimizer.

The 3 agents (in .claude/agents/)

AgentRoleYou invoke?
ferret-kernel-agent (L1 dispatcher)The ENTRY POINT. Pins ALL constraints, freezes the gate via the test-writer, runs the optimizer in-session, drives Codex review each round, does in-MPK faithful acceptance.YES — this is the only one you invoke directly.
ferret-test-writer (L2a)Writes the FROZEN, hash-locked gate vs a CANONICAL reference (never re-derived), checking INTERMEDIATE tensors. Spawned by L1, before the optimizer.No (L1 spawns it)
ferret-optimizer (L2b)In-session optimizer (replaces claude -p), judged ONLY by the frozen gate, can't simplify. Spawned by L1 each round.No (L1 spawns it)

How to invoke (the main-thread → subagent contract)

Invoke the ferret-kernel-agent subagent (via the Agent tool) with a COMPLETE constraint contract — this is where nothing gets missed and no simplification slips in. Give it, explicitly:

  1. TARGET — the exact MPK op + what it replaces; the baseline = the kernel being replaced, benched the way MPK calls it (not an external SOTA unless it's the consumer).
  2. REAL-MATH CONTRACT — enumerate EVERY step the kernel must compute, NO simplification (the test-writer turns this into intermediate checks).
  3. SHAPES — exact (M,K,N) at the real TP/EP regime (derive from the builder/weights).
  4. PRODUCTION COMPILE FLAGS — -rdc=true / MPK_FORCE_RDC_TRUE=1, arch sm_100a, single-stream / no-CUDA-graph / no-cta_group::2. State that FINAL acceptance is the in-MPK faithful build; standalone -rdc=true is diagnostic only.
  5. ABI — the __device__ task_impl signature + NS/NE.
  6. CANONICAL REFERENCE SOURCE — the already-trusted oracle to compare against (the in-MPK task-chain output; the official HF model; an in-tree faithful test). NEVER "let the agent derive it." If the gate already exists (hash-locked in the workspace), tell L1 to reuse it (hash-verify) and go straight to the optimizer loop.
Show full SKILL.md (223 more words)Show less

The invariants that make it trustworthy (Codex-hardened — gate fidelity is load-bearing)

  • The gate is built by the independent test-writer, not the optimizer.
  • The reference is CANONICAL (validated against a trusted source), never re-derived.
  • The gate checks INTERMEDIATES (golden vectors per stage), not just a final cosine — a deep simplification (a dropped layernorm, a wrong rope base) is caught at the first diverging stage (first_failing_stage), not washed out.
  • Multiple metrics + edge cases (long-context, boundary positions).
  • The gate is sha256 hash-locked; L1 re-verifies the hash before EVERY round (tamper = abort).
  • The optimizer is judged ONLY by gate/check.py. Codex reviews each round on two axes — Integrity (did it simplify to pass?) + Plan (is the lever sound?).
  • FINAL acceptance = in-MPK faithful build (MPK_FORCE_RDC_TRUE=1, compiled into the real megakernel). A standalone number never ships.
  • Early stop = round incomplete, never success.

Proof it works

Re-run on the simplified attention: the optimizer (judged by the frozen gate) built a CORRECT real-DSv3 fused attention in round 1 — GATE_RESULT {pass:true} on all 5 cases / all intermediates, gate untouched, dump provenance verified (W_UV BMM present, real mscale²/sqrt(192) scale), all 4 simplifications fixed, Codex Integrity+Plan PASS. The old simplified-kernel failure is now structurally impossible.

Reference

Full design + the /cd-mechanics + the hardening rationale: DESIGN.md (bundled alongside this skill). The ferret runtime itself lives at ~/ferret/ (its CLAUDE.md is the optimizer's methodology). Pre-authored task specs: ~/ferret/tasks/.

© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/ferret-kernel-system of mirage-project/mirage.

  • SKILL.md
  • DESIGN.md

Open the folder on GitHubat commit f9eb70c

Compare with similar skills

Ferret Kernel System next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ferret Kernel System compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ferret Kernel System this skillmirage-project/mirage2.5k—~1.3kAutomated safety check: PassApache-2.0
Kernel VerificationZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNone
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Kernel Verification

    ZJLi2013/awesome-kernel-skills

    5-stage kernel correctness verification protocol for Triton and CUDA kernels.

    102 GitHub stars~702 tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from mirage-project/mirage

All 24 skills in this repo
  • V2 Perf Iteration

    mirage-project/mirage

    Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.

    2.5k GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Mpk Faithful Gate

    mirage-project/mirage

    Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.

    2.5k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Mpk Lever Cleanup

    mirage-project/mirage

    A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…

    2.5k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Test Mode

    mirage-project/mirage

    Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.

    2.5k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Ferret Kernel System

What does Ferret Kernel System do?

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…. Ferret Kernel System is an agent skill from mirage-project/mirage.cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer cheating its own correctness test.

When should I use Ferret Kernel System?

Ferret Kernel System fits situations like: — phrasings like dispatch ferret to optimize a kernel; have an agent write a kernel but guarantee it doesnt simplify the math; generate a kernel the frozen-gate way: the user/main-thread wants a kernel optimized; generated by an agent.

How do I install Ferret Kernel System in Claude Code?

Run `npx skills add mirage-project/mirage --skill ferret-kernel-system -a claude-code`. Or copy the skill folder (.claude/skills/ferret-kernel-system in mirage-project/mirage) into .claude/skills/ferret-kernel-system in your project. Claude Code loads it when a task matches its description.

How do I install Ferret Kernel System in Codex?

Run `npx skills add mirage-project/mirage --skill ferret-kernel-system -a codex`. Or copy the skill folder (.claude/skills/ferret-kernel-system in mirage-project/mirage) into .agents/skills/ferret-kernel-system in your project. Codex loads it when a task matches its description.

Can I use Ferret Kernel System in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill ferret-kernel-system -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ferret-kernel-system, .gemini/skills/ferret-kernel-system, .github/skills/ferret-kernel-system and .opencode/skills/ferret-kernel-system in your project.

What does Ferret Kernel System need to run?

Going by SKILL.md and its folder, Ferret Kernel System needs the command-line tools its instructions call (claude).

Does Ferret Kernel System access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ferret Kernel System safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ferret Kernel System use?

Ferret Kernel System is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ferret Kernel System use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ferret Kernel System?

Skills that share tags, products or a category with Ferret Kernel System: Kernel Verification (ZJLi2013/awesome-kernel-skills, 102 stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Debug Distributed Hang (sgl-project/sglang, 37k stars) and CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ferret Kernel System?

mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,541 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.

Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.