Agent skill

Megakernel Optimization

by RightNow-AI in RightNow-AI/AutoMegaKernel

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

MITAuto-check passedAI & LLM Engineering

Install Megakernel Optimization

skills CLI
$ npx skills add RightNow-AI/AutoMegaKernel --skill megakernel-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RightNow-AI/AutoMegaKernel megakernel-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RightNow-AI/AutoMegaKernel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/megakernel-optimization .claude/skills/megakernel-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
megakernel-optimization
GitHub stars
148
Token cost
~1.8k tokens
SKILL.md length
642 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

  • Works in 7 steps: Read the surface with amk_propose (or… → Baseline. amk_eval the incumbent.… → Propose ONE knob change (one knob per… → …
  • Generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK)
  • SKILL.md covers HARD HONESTY RULES (state and…, The edit surface (read it…, The eval verdict and The loop (drive this exactly), plus 2 more sections
  • Calls python and uv

What it does

Megakernel Optimization is an agent skill from RightNow-AI/AutoMegaKernel. Use when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop (or hands off to the unattended autoresearch driver).

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets, Autonomous loops and Deep learning. It works with CUDA and Hugging Face. The repository describes itself as: An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper…. The licence is MIT.

When your agent uses it

  • Generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK)
  • Drives the correctness-gated propose - eval - keep/revert loop (or hands off to the unattended autoresearch driver)

Example prompts

  • “/megakernel-optimization”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Read the surface with amk_propose (or amk propose). Note the incumbent
  2. Baseline. amk_eval the incumbent. Require valid AND correct. This is the bar to beat;
  3. Propose ONE knob change (one knob per trial, schedule knob OR one kernel_knobs field),
  4. Eval the candidate with amk_eval.
  5. Keep/revert. Keep ONLY if valid AND correct AND
  6. Record the outcome to the orchestrator: `amk_orchestrate_record(status, latency_us=...,
  7. Repeat from step 3 around the current best. Ask amk_orchestrate_next() (CLI

What it can do on your machine

Read from SKILL.md and the folder at commit 884534c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Megakernel Optimization loads about 1.8k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 642 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RightNow-AI/AutoMegaKernel at commit 884534c, republished under its MIT licence (© RightNow-AI). 642 words, ~1,756 tokens.

Download SKILL.mdSave it as .claude/skills/megakernel-optimization/SKILL.md (or your agent's skills folder).
name
megakernel-optimization
description
Use when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose -> eval -> keep/revert loop (or hands off to the unattended autoresearch driver).

AutoMegaKernel (AMK), megakernel schedule optimization

AMK compiles a HuggingFace Llama-family model into ONE persistent CUDA megakernel and tunes it with an AutoKernel-style loop: read the edit surface -> propose ONE knob change -> eval -> keep/revert -> record -> repeat. This skill drives Loop 2 (schedule + kernel_knobs search). You never write kernel code; you only edit a structured ScheduleConfig (plus its reserved kernel_knobs sub-object). The frozen VM lowers your config deterministically and the CPU ReferenceVM judges correctness vs eager PyTorch.

HARD HONESTY RULES (state and obey these every time)

  • Correctness FIRST. A latency is NEVER reported without a correctness PASS vs the CPU ReferenceVM. Keep a candidate only if it is correct AND >= 1% faster than the incumbent.
  • validate-before-launch. An unsafe ScheduleConfig is a clean REJECTED (a deadlock/race-free proof rejects it before launch), never a hung GPU.
  • The edit surface is ScheduleConfig + kernel_knobs ONLY, never raw kernel code, never vm/, never the frozen ABI.
  • Measured-gpu latency is drift-robust; physically-impossible sub-roofline latencies are withheld as artifacts.
  • All speedups are vs AMK's OWN baseline (default schedule), NOT a claim of beating cuBLAS/vLLM. AMK is currently within ~13% of cuBLAS at batch-1, behind it.

The edit surface (read it before proposing)

Read the surface programmatically, never guess knob names. Prefer the canonical MCP tool; fall back to the CLI if MCP is unavailable.

  • MCP: amk_propose(model, gpu="rtx5090") -> { schedule_config, schedule_id, search_space, ... }. search_space includes the kernel_knobs.* sub-surface.
  • CLI: amk propose <model> --gpu <arch> (or uv run python amk_cli.py propose <model> --gpu <arch>) prints the same surface as JSON on stdout.

The ScheduleConfig knobs (edit ONE per trial): tiling.gemv.N_tile, tiling.attention.kv_block, fusion_grouping, sm_assignment, pipelining_depth, page_allocation, threads_per_block, smem_bytes_per_block. The reserved kernel_knobs object holds GEMV build knobs: cols_per_warp, cpasync, cpa_stages, cpa_cols (these move MEASURED latency under device=cuda; the predicted/CPU path does not model them). A config WITHOUT kernel_knobs is byte-identical to the production incumbent.

<model> is toy / toy-2L (fully supported) or a HuggingFace id (best-effort). <arch> is a registered GpuTarget: rtx5090, b200, h100, a100.

The eval verdict

  • MCP: amk_eval(model, gpu, config, device="auto") where config is a JSON ScheduleConfig object (optionally carrying a kernel_knobs object).
  • CLI: write cfg.json, then amk eval <model> --gpu <arch> --config cfg.json (JSON-only on stdout; exit code 0 = valid+correct, 1 = rejected or incorrect).

The verdict carries valid, rejected_reason, correct, latency_us, latency_kind (measured-gpu | predicted), pct_of_roofline, bound_us, schedule_id. latency_us and latency_kind are null unless correct is true and the config was valid. eval never crashes, malformed knobs come back as a clean valid=false with a rejected_reason.

Show full SKILL.md (240 more words)Show less

The loop (drive this exactly)

  1. Read the surface with amk_propose (or amk propose). Note the incumbent schedule_config and the editable search_space.
  2. Baseline. amk_eval the incumbent. Require valid AND correct. This is the bar to beat; its latency_us is the incumbent latency.
  3. Propose ONE knob change (one knob per trial, schedule knob OR one kernel_knobs field), building the candidate config from the incumbent.
  4. Eval the candidate with amk_eval.
  5. Keep/revert. Keep ONLY if valid AND correct AND latency_us < incumbent_latency_us * 0.99 (a strict >= 1% win). Otherwise revert (keep the old incumbent). Tie-break: a measured-gpu number outranks a predicted one, then simpler config wins.
  6. Record the outcome to the orchestrator: amk_orchestrate_record(status, latency_us=..., pct_roofline=..., kind=..., config=..., description=...) with status one of kept/revert/failed/crash/timeout/rejected. (CLI: python amk_orchestrate.py record kept --latency-us ... --pct-roofline ... --kind ... --config cfg.json --description "...".)
  7. Repeat from step 3 around the current best. Ask amk_orchestrate_next() (CLI: python amk_orchestrate.py next) whether to continue or STOP (plateau / near-roofline / budget / >=3x speedup), and amk_orchestrate_status() for baseline/best/speedup/plateau.

To run the whole keep/revert loop in one call, use amk_loop(model, gpu, budget=8) (CLI: amk loop <model> --gpu <arch> --budget N). To run unattended for hours, hand off to amk_autoresearch(model, gpu, minutes=..., overnight=...) (CLI: amk autoresearch ...); see the /amk-autoresearch command.

Worked example (toy on rtx5090)

1. amk_propose("toy", "rtx5090")
   -> incumbent schedule_config (pipelining_depth=0, N_tile default, no kernel_knobs),
      search_space lists N_tile in {64,128,256,512}, pipelining_depth 0-4, kernel_knobs.cpasync {0,1}, ...

2. amk_eval("toy", "rtx5090", <incumbent cfg>, device="cuda")
   -> { valid:true, correct:true, latency_us: 1228.0, latency_kind:"measured-gpu", ... }
      incumbent_latency = 1228.0 us

3. ONE knob change: set pipelining_depth = 3 (hides the inter-op HBM bubble).
   cfg = { ...incumbent, "pipelining_depth": 3 }

4. amk_eval("toy", "rtx5090", cfg, device="cuda")
   -> { valid:true, correct:true, latency_us: 1010.0, latency_kind:"measured-gpu", ... }

5. 1010.0 < 1228.0 * 0.99  -> KEEP. New incumbent latency = 1010.0 us.
   amk_orchestrate_record("kept", latency_us=1010.0, pct_roofline=..., kind="measured-gpu",
                          config=cfg, description="pipelining_depth 0->3")

6. Next ONE knob change off the new best, e.g. kernel_knobs.cpasync = 1 / N_tile = 128. Eval.
   If a candidate is correct but only 0.4% faster -> record "revert" (correct but not kept).
   If a candidate is invalid (e.g. over-cap smem) -> record "rejected" and revert.

7. amk_orchestrate_next() until it says STOP. Speedup is vs AMK's own default schedule, NOT a
   cuBLAS/vLLM claim.

Canonical tool / CLI names (use these EXACT names)

  • MCP: amk_doctor, amk_propose, amk_eval, amk_loop, amk_autoresearch, amk_orchestrate_status, amk_orchestrate_next, amk_orchestrate_report, amk_orchestrate_record.
  • CLI: amk propose|eval|loop|autoresearch|compile|generate|doctor and python amk_orchestrate.py status|next|record|report.

Full contract: read HARNESS.md (terminology in README.md).

© RightNow-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/megakernel-optimization of RightNow-AI/AutoMegaKernel.

Open the folder on GitHubat commit 884534c

Compare with similar skills

Megakernel Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megakernel Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megakernel Optimization this skillRightNow-AI/AutoMegaKernel148—~1.8kAutomated safety check: PassMIT
Cosmos3 Post TrainingNVIDIA/cosmos-framework556—~2.7kAutomated safety check: PassCustom licence
Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Perforatedai Libraries TransformersPerforatedAI/PerforatedAI237—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Cosmos3 Post Training

    NVIDIA/cosmos-framework

    Official

    Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

    556 GitHub stars~2.7k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Mamba State-Space Models

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

    13k GitHub starsUsed in 3 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Perforatedai Libraries Transformers

    PerforatedAI/PerforatedAI

    HuggingFace Transformers integration for PerforatedAI. An agent skill from PerforatedAI/PerforatedAI.

    237 GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Moe Training

    Orchestra-Research/AI-Research-SKILLs

    Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace.

    13k GitHub starsUsed in 3 repos~3.7k tokens
    AI & LLM EngineeringAuto-check passed

Questions about Megakernel Optimization

What does Megakernel Optimization do?

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…. Megakernel Optimization is an agent skill from RightNow-AI/AutoMegaKernel. Use when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop (or hands off to the unattended autoresearch driver).

When should I use Megakernel Optimization?

Megakernel Optimization fits situations like: generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK); drives the correctness-gated propose - eval - keep/revert loop (or hands off to the unattended autoresearch driver).

How do I install Megakernel Optimization in Claude Code?

Run `npx skills add RightNow-AI/AutoMegaKernel --skill megakernel-optimization -a claude-code`. Or copy the skill folder (.claude/skills/megakernel-optimization in RightNow-AI/AutoMegaKernel) into .claude/skills/megakernel-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Megakernel Optimization in Codex?

Run `npx skills add RightNow-AI/AutoMegaKernel --skill megakernel-optimization -a codex`. Or copy the skill folder (.claude/skills/megakernel-optimization in RightNow-AI/AutoMegaKernel) into .agents/skills/megakernel-optimization in your project. Codex loads it when a task matches its description.

Can I use Megakernel Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RightNow-AI/AutoMegaKernel --skill megakernel-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/megakernel-optimization, .gemini/skills/megakernel-optimization, .github/skills/megakernel-optimization and .opencode/skills/megakernel-optimization in your project.

What does Megakernel Optimization need to run?

Going by SKILL.md and its folder, Megakernel Optimization needs the command-line tools its instructions call (python and uv). Our summary lists: Python 3.

Does Megakernel Optimization access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Megakernel Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megakernel Optimization use?

Megakernel Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megakernel Optimization use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megakernel Optimization?

Skills that share tags, products or a category with Megakernel Optimization: Cosmos3 Post Training (NVIDIA/cosmos-framework, 556 stars), Mamba State-Space Models (Orchestra-Research/AI-Research-SKILLs, 13k stars), Esmfold2 (JimLiu/science-skills, 227 stars) and Hugging Face Local Models (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megakernel Optimization?

RightNow-AI (a GitHub organization) maintains it in RightNow-AI/AutoMegaKernel, which has 148 GitHub stars. The repository was last updated on September 18, 2026.

Source: RightNow-AI/AutoMegaKernel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.