Agent skill

Karpathy

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

MITAuto-check passedAgent Workflows

Install Karpathy

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill karpathy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills karpathy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/karpathy .claude/skills/karpathy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
karpathy
GitHub stars
174
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
1,368 words
Files
2
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

  • Works in 6 steps: Choose — branches (one git commit per… → Open the run — branches: agree on a run… → Read the in-scope files — the repo is… → …
  • The user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code
  • SKILL.md covers When to use, Setup, The experiment loop and results.tsv (logging results), plus 2 more sections
  • Calls git and uv

What it does

Karpathy is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g. valbpb). One agent proposes one change at a time, runs training in the user's env, keeps it only if the metric improves (advancing a git branch) else reverts, and loops forever until the human interrupts. A faithful adaptation of Karpathy's autoresearch. Not for the analysis-first variant that profiles before editing (that is…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `examples/run.example.yaml`). Compatibility notes: Requires Python 3.9+

It sits in Agent Workflows, covering Autonomous loops and Git workflow. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • The user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code
  • Keeps changes that lower a single scalar metric (e.g

Example prompts

  • “/karpathy”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose — branches (one git commit per run; the original) or snapshots
  2. Open the run — branches: agree on a run tag from today's date (e.g. mar5) and create the
  3. Read the in-scope files — the repo is small; read them for full context: the README, the
  4. Verify the env/data exists — confirm can run (data shards, tokenizer, deps present).
  5. Initialize results.tsv — create it with just the header row; the baseline is recorded after the
  6. Confirm and go — confirm the setup looks right, then kick off the experimentation.

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+

    From compatibility in the SKILL.md frontmatter.

Context cost

Karpathy loads about 2.6k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 1,368 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,368 words, ~2,554 tokens.

Download SKILL.mdSave it as .claude/skills/karpathy/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
karpathy
description
Use when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g. val_bpb). One agent proposes one change at a time, runs training in the user's env, keeps it only if the metric improves (advancing a git branch) else reverts, and loops forever until the human interrupts. A faithful adaptation of Karpathy's autoresearch. Not for the analysis-first variant that profiles before editing (that is ml-autoresearch), and not for a budgeted, plateau-stopping refactor.
compatibility
Requires Python 3.9+
metadata.version
0.1.0

Karpathy Autoresearch

This is an experiment to have the LLM do its own research. You are a completely autonomous researcher: you hack the training code with an idea, run it, keep the change if the metric improves and revert it if it doesn't, advancing a branch as you go — and you repeat forever, until the human interrupts you. The artifact is the <editable_files>; the feedback signal is one scalar <metric> (lower is better, e.g. val_bpb) read from the run. Training runs in the user's own environment via <run_cmd> — this skill installs nothing and imports nothing; it edits code, shells out, and reads the metric from the log.

When to use

Use this to leave an agent running on a single training script, optimizing one scalar metric hands-off, where any improvement is kept and the loop never stops on its own. Default to broad freedom inside <editable_files>; the only hard limit is that the run finishes within the budget without crashing. Not for the analysis-first variant that reasons about the data before each edit (that is ml-autoresearch).

Setup

Resolve bindings interactively (load loop.run.yaml and skip if it already exists; else, on Claude Code infer + recommend each via AskUserQuestion, otherwise ask as quoted prompts; write loop.run.yaml). Then work with the user to set up a fresh run:

  1. Choose <iter_strategy> — branches (one git commit per run; the original) or snapshots (one folder per run under <sandbox_root>/). Snapshots are safer on a dirty or gitignored tree; branches mirror Karpathy. Either is fully supported throughout the loop.
  2. Open the run — branches: agree on a run tag from today's date (e.g. mar5) and create the branch git checkout -b autoresearch/<tag> (it must not already exist; this is a fresh run). snapshots: no branch — each iteration gets its own <sandbox_root>/iter<N>/.
  3. Read the in-scope files — the repo is small; read them for full context: the README, the read-only harness that defines the metric (the <metric> ground truth — do not modify), and the <editable_files> you will hack (model/optimizer/training loop).
  4. Verify the env/data exists — confirm <run_cmd> can run (data shards, tokenizer, deps present). If not, tell the human the one command to prepare it (e.g. uv run prepare.py).
  5. Initialize results.tsv — create it with just the header row; the baseline is recorded after the first run. Leave it untracked (never commit it).
  6. Confirm and go — confirm the setup looks right, then kick off the experimentation.
bindingmeaningdefaulthow to infer
<metric>scalar to minimize; must be printed by the run (e.g. val_bpb)—grep the in-scope files / README for val_bpb, val_loss, error…
<run_cmd> / <entrypoint>the command that launches one training run—e.g. uv run train.py; from pyproject.toml/.venv/README
<editable_files>the file(s) you may hack — everything else is read-only—the training script(s); exclude data, the eval harness, configs you must not touch
<sandbox_root>where results.tsv (+ snapshots) live./sandbox—
<iter_strategy>branches (a git commit per run) or snapshots (a folder per run)snapshotssnapshots is safer off a dirty/gitignored tree; branches mirrors Karpathy
<gate> / <budget>time (wall-clock) or epochs, and its value—the run's existing time/epoch setting

The experiment loop

Each experiment is one training run on a fixed budget (<gate>/<budget> — wall-clock time or a fixed epoch count, excluding startup/compile). You launch it simply: <run_cmd> (e.g. uv run train.py). Because the budget is fixed you don't need to worry about training time — every run gets the same budget.

What you CAN do: Modify <editable_files> — this is the only file you edit. Everything is fair game: model architecture, optimizer, hyperparameters, training loop, batch size, model size, etc.

What you CANNOT do: modify the read-only harness or the evaluation (the <metric> is the ground truth); install new packages or add dependencies (use only what's already available).

The goal is simple: get the lowest <metric>. Since the budget is fixed, you don't need to worry about training time — it's always the budget. Everything is fair game: change the architecture, the optimizer, the hyperparameters, the batch size, the model size. The only constraint is that the code runs without crashing and finishes within the budget.

VRAM is a soft constraint. Some increase is acceptable for meaningful <metric> gains, but it should not blow up dramatically.

Simplicity criterion: All else being equal, simpler is better. A small improvement that adds ugly complexity is not worth it. Conversely, removing something and getting equal or better results is a great outcome — that's a simplification win. When evaluating whether to keep a change, weigh the complexity cost against the improvement magnitude. A 0.001 <metric> improvement that adds 20 lines of hacky code? Probably not worth it. A 0.001 <metric> improvement from deleting code? Definitely keep. An improvement of ~0 but much simpler code? Keep.

The first run: Your very first run should always be to establish the baseline, so you will run the training script as is.

Copy this checklist; LOOP FOREVER:

  • 1. Look at the git state — the branch/commit you're on (snapshots: the next iter<N>/).
  • 2. Tune <editable_files> with one experimental idea by directly hacking the code.
  • 3. Commit it (git commit -am "<idea>"; snapshots: copy <editable_files> into iter<N>/code_snapshot/ first).
  • 4. Run the experiment: <run_cmd> > run.log 2>&1 (redirect everything — do NOT use tee or let output flood your context).
  • 5. Read the result: grep "^<metric>:" run.log (also grab peak memory if printed).
  • 6. If the grep is empty the run crashed — tail -n 50 run.log, read the trace, fix if it's something dumb (typo/missing import), else give up after a couple of tries.
  • 7. Record the result in results.tsv (do NOT commit it — leave it untracked).
  • 8. If <metric> improved (lower), advance — keep the commit.
  • 9. If it's equal or worse, git reset back to where you started (snapshots: restore from code_snapshot/). Go to 1.
Show full SKILL.md (418 more words)Show less

The idea is that you are a completely autonomous researcher trying things out. If they work, keep. If they don't, discard. And you're advancing the branch so that you can iterate. If you feel like you're getting stuck in some way, you can rewind but you should probably do this very very sparingly (if ever).

Output format. When the run finishes it prints a summary; the exact lines depend on what the user's script prints (<metric> is whatever you bound — val_bpb, val_loss, a perplexity, an error rate, …), e.g.:

<metric>:         0.997900
peak_vram_mb:     45060.2
num_params_M:     50.3

The numbers vary by machine since each run stops at the budget. Extract the metric with grep "^<metric>:" run.log.

Timeout. A run should take ~its budget plus a little eval overhead. If a time-gated run exceeds 2× <budget> minutes, kill it and treat it as a failure (discard and revert).

Crashes. Use judgement: something dumb and easy (a typo, a missing import) — fix it and re-run; an idea that's fundamentally broken — skip it, log crash as the status, and move on.

results.tsv (logging results)

<sandbox_root>/results.tsv, tab-separated (NOT comma-separated — commas break in descriptions). Header

  • 5 columns: the git commit (short, 7 chars; or iter in snapshots mode), <metric> (e.g. 1.234567, or 0.000000 for a crash), peak memory in GB (.1f, peak_vram_mb/1024; 0.0 for a crash), status ∈ {keep, discard, crash}, and a text description of what the experiment tried.
commit	<metric>	memory_gb	status	description
a1b2c3d	0.997900	44.0	keep	baseline
b2c3d4e	0.993200	44.2	keep	increase LR to 0.04
c3d4e5f	1.005000	44.0	discard	switch to GeLU activation
d4e5f6g	0.000000	0.0	crash	double model width (OOM)

Report the best run when interrupted, not necessarily the last.

Constraints

  • Only edit <editable_files>. The read-only harness that produces <metric> is the ground truth — editing it (or the eval) would corrupt the signal the loop is scored against.
  • One change per iteration, so each <metric> delta is attributable to a single idea.
  • Run in the user's env via <run_cmd>. Install nothing, add no dependencies — shell out and read the log. Always redirect output to run.log; never tee, never flood your context.
  • Don't commit results.tsv — leave it untracked. <sandbox_root>/ is self-contained (no ../).

Stops — NEVER STOP

Once the experiment loop has begun (after the initial setup), do NOT pause to ask the human if you should continue. Do NOT ask "should I keep going?" or "is this a good stopping point?". The human might be asleep, or gone from a computer and expects you to continue working indefinitely until you are manually stopped. You are autonomous. If you run out of ideas, think harder — read papers referenced in the code, re-read the in-scope files for new angles, try combining previous near-misses, try more radical architectural changes. The loop runs until the human interrupts you, period.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in loops/karpathy of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/run.example.yaml

Open the folder on GitHubat commit f1169e6

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in gaasher/Agent-Loop-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Karpathy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Karpathy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Karpathy this skillgaasher/Agent-Loop-Skills1741 repos~2.6kAutomated safety check: PassMIT
Rebuild Branchplatformplatform/PlatformPlatform441—~2.1kAutomated safety check: NotesMIT
Darwin SkillHHU3637kr/skills1451 repos~2.2kAutomated safety check: PassNone
AutoresearchFactory-AI/factory-plugins110—~4.2kAutomated safety check: PassNone
AutoResearch LoopLearnPrompt/andrej-karpathy-skills109—~1.4kAutomated safety check: PassMIT
Show Me Your Work Decision Logcursor/plugins10k9 repos~1.6kAutomated safety check: PassNone

Similar skills

  • Rebuild Branch

    platformplatform/PlatformPlatform

    Rebuild a stale branch by cherry-picking each commit onto a fresh branch off main, using a ralph-loop to validate each commit (build, test, format, lint, optional e2e) before moving on.

    441 GitHub stars~2.1k tokensUpdated 14 days ago
    Agent WorkflowsAuto-check: notes
  • Darwin Skill

    HHU3637kr/skills

    Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.

    145 GitHub starsUsed in 1 repo~2.2k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch

    Factory-AI/factory-plugins

    Autonomous experiment loop for optimization research. An agent skill from Factory-AI/factory-plugins.

    110 GitHub stars~4.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • AutoResearch Loop

    LearnPrompt/andrej-karpathy-skills

    Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

    109 GitHub stars~1.4k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 9 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed

More from gaasher/Agent-Loop-Skills

All 21 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Tournament Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an autonomous ML research loop that pressure-tests competing ideas before spending compute — several research subagents each propose one architecture…

    174 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Dueling Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

    174 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: warnings
  • Anomaly Investigation

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

    174 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Data Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted.

    174 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Karpathy

What does Karpathy do?

A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g. Karpathy is an agent skill from gaasher/Agent-Loop-Skills.g.

When should I use Karpathy?

Karpathy fits situations like: the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code; keeps changes that lower a single scalar metric (e.g.

How do I install Karpathy in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill karpathy -a claude-code`. Or copy the skill folder (loops/karpathy in gaasher/Agent-Loop-Skills) into .claude/skills/karpathy in your project. Claude Code loads it when a task matches its description.

How do I install Karpathy in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill karpathy -a codex`. Or copy the skill folder (loops/karpathy in gaasher/Agent-Loop-Skills) into .agents/skills/karpathy in your project. Codex loads it when a task matches its description.

Can I use Karpathy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill karpathy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/karpathy, .gemini/skills/karpathy, .github/skills/karpathy and .opencode/skills/karpathy in your project.

What does Karpathy need to run?

Going by SKILL.md and its folder, Karpathy needs the command-line tools its instructions call (git and uv). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+.

Does Karpathy access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Karpathy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Karpathy use?

Karpathy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Karpathy use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Karpathy?

Skills that share tags, products or a category with Karpathy: Rebuild Branch (platformplatform/PlatformPlatform, 441 stars), Darwin Skill (HHU3637kr/skills, 145 stars), Autoresearch (Factory-AI/factory-plugins, 110 stars) and AutoResearch Loop (LearnPrompt/andrej-karpathy-skills, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Karpathy?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.