Agent skill

AutoResearch Loop

by LearnPrompt in LearnPrompt/andrej-karpathy-skills

Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

MITAuto-check passedAgent Workflows

Install AutoResearch Loop

skills CLI
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/karpathy-autoresearch .claude/skills/karpathy-autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
karpathy-autoresearch
GitHub stars
110
Token cost
~1.4k tokens
SKILL.md length
153 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

  • Automating a series of ML experiments or hyperparameter runs
  • SKILL.md covers Core Principle, The AutoResearch Architecture, Setup Prompt (One-Time) and Per-Iteration Loop Prompt, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Setting up a research agent that proposes the next experiment itself

What it does

The core idea is that you change the prompt and the agent changes everything else: it reads the current state, runs an experiment, analyzes the results and proposes the next iteration, then waits for you to approve the prompt change before repeating. You stay in the hypothesis space and the agent stays in implementation and execution. The skill supplies a one-time setup prompt, a per-iteration prompt, a JSONL experiment log (id, branch, hypothesis, metric, runtime, notes) and a summary prompt to run after several iterations.

Each experiment lives on its own git branch, so earlier runs are never lost. The pattern is described as usable beyond machine learning, for prompt tuning, code speed, content formats and UX flows, each with its own metric. Safety rules tell the agent never to delete earlier data or commits and never to touch files outside the current experiment's scope. It is based on a post by Andrej Karpathy, and the excerpt is cut off at the prompt contract.

When your agent uses it

  • Automating a series of ML experiments or hyperparameter runs
  • Setting up a research agent that proposes the next experiment itself
  • Tuning a prompt or an implementation's speed through measured iterations
  • Summarizing an experiment log after several iterations

Example prompts

  • “Set up an autoresearch loop to tune the learning rate of our training script.”
  • “Run a hyperparameter search on the classifier, one change per iteration, and ask me before each new hypothesis.”
  • “Log each experiment on its own branch and summarize the results after ten iterations.”

Requirements

  • A git repository for the experiment branches

What it can do on your machine

Read from SKILL.md and the folder at commit 9e46dec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are jsonl).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • x.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AutoResearch Loop loads about 1.4k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 153 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LearnPrompt/andrej-karpathy-skills at commit 9e46dec, republished under its MIT licence (© LearnPrompt). 153 words, ~1,401 tokens.

Download SKILL.mdSave it as .claude/skills/karpathy-autoresearch/SKILL.md (or your agent's skills folder).
name
karpathy-autoresearch
description
Set up an autonomous research loop where an agent iterates on experiments, hypotheses, and code on git branches. Use this skill when the user wants to automate ML experiments, set up a research agent that self-improves, needs to run many variations of an experiment without manual intervention, or says "automate experiments", "agent research loop", "self-improving agent", "run hyperparameter search", "autoresearch". Based on Karpathy 28k-like AutoResearch post.
disable-model-invocation
false
user-invocable
true
related_skills
karpathy-llm-wiki, karpathy-output-evolution, karpathy-education-first, karpathy-system-prompt-learning

Skill 6: AutoResearch(自主研究循环)

Source: https://x.com/karpathy/status/2030371219518931079 "Autoresearch project — agent self-iterates training code" — 28k likes

Core Principle

You change the prompt. The agent changes everything else.

The loop: agent reads current state → runs experiment → analyzes results → proposes next iteration → you approve the prompt change → repeat.

You stay in the hypothesis space. Agent stays in the implementation + execution space.

The AutoResearch Architecture

Human Layer (you):
  - Define research question
  - Approve hypothesis changes
  - Set evaluation metric
  - Stop when satisfied

Agent Layer:
  - Implement current hypothesis
  - Run experiment (on git branch)
  - Analyze logs/results
  - Propose next hypothesis
  - Never change the research question

Setup Prompt (One-Time)

You are AutoResearcher for this project.

Research question: [WHAT YOU'RE TRYING TO LEARN/OPTIMIZE]
Evaluation metric: [HOW WE MEASURE SUCCESS — must be a single number]
Current best result: [BASELINE or "none yet"]

Constraints:
- Work on git branches: each experiment gets branch "exp/[short-description]"
- Never modify main branch
- Log all results to experiments.jsonl in format: {"id": N, "hypothesis": "...", "metric": value, "notes": "..."}
- Each iteration: implement → run → log → propose next

Starting hypothesis: [YOUR FIRST HYPOTHESIS TO TEST]

Begin iteration 1. Implement the hypothesis, run it, report the metric, then propose iteration 2.

Per-Iteration Loop Prompt

AutoResearcher: Iteration [N]

Current state:
- Branch: exp/[current]
- Metric so far: [RESULTS LOG]
- Best result: [BEST SO FAR]

Completed last iteration: [WHAT HAPPENED]
Metric result: [VALUE]

Your tasks this iteration:
1. Analyze: what does the result tell us about the hypothesis?
2. Propose: what's the next hypothesis? (one variable change at a time)
3. Implement: write the code change for this hypothesis
4. Run: execute and report the metric
5. Commit: git commit to branch exp/[new-name] with message: "exp: [hypothesis description]"

Constraint: change ONE variable at a time. If last experiment failed, back off to simpler hypothesis.

Experiment Log Format

jsonl
{"id": 1, "branch": "exp/baseline", "hypothesis": "default hyperparams", "metric": 0.423, "runtime_s": 45, "notes": "starting point"}
{"id": 2, "branch": "exp/lr-0.001", "hypothesis": "lower learning rate 0.01→0.001", "metric": 0.451, "runtime_s": 47, "notes": "small improvement, keep lowering?"}
{"id": 3, "branch": "exp/lr-0.0001", "hypothesis": "lower learning rate 0.001→0.0001", "metric": 0.438, "runtime_s": 46, "notes": "worse than id=2, sweet spot was 0.001"}

Research Summary Prompt (run after N iterations)

Summarize this research session.

Experiment log:
[PASTE experiments.jsonl]

Produce:
1. Best configuration found (all hyperparameters/settings)
2. What we learned (2-3 bullet points about the search space)
3. What we should try next (top 3 unexplored directions)
4. A clear hypothesis about WHY the best configuration works
5. Confidence level: how sure are we this is near-optimal?

Applying AutoResearch Beyond ML

This pattern works for any iterative optimization:

DomainResearch QuestionMetric
ML trainingBest architecture/hyperparamsValidation loss
Prompt engineeringBest system promptTask success rate
Code optimizationFastest implementationRuntime (ms)
Content creationMost engaging formatClick/read rate
Product designBest UX flowTask completion rate

Safety Rails

AutoResearcher rules (always include):
- NEVER delete data or commits from previous experiments
- NEVER modify any file outside the current experiment's scope
- NEVER run experiments that cost > [COST_LIMIT] without checking
- ALWAYS commit before starting the next iteration
- ALWAYS report the metric before proposing the next hypothesis
- If metric degrades 3 iterations in a row, STOP and ask for guidance

Workflow

属于工作流:研究到发布(内容创作者/研究者)

位置上游下游
入口用户有研究问题时直接触发karpathy-llm-wiki(把实验结果沉淀为知识页)

完整链路:autoresearch → llm-wiki → output-evolution → education-first

Prompt Contract

text
You are AutoResearcher. Research question: <QUESTION>. Evaluation metric: <SINGLE_NUMBER_METRIC>. Current baseline: <VALUE or "none">. Starting hypothesis: <FIRST_HYPOTHESIS>. Work on git branches (exp/<name>). Each iteration: implement → run → log result to experiments.jsonl → analyze → propose next hypothesis (one variable change). Never modify main. Stop and ask if metric degrades 3 iterations in a row.

Verification Checklist

  • 研究问题有明确的单一评估指标
  • 每轮实验只改一个变量
  • 所有结果都记录到 experiments.jsonl
  • 每轮在独立 git branch 上工作
  • 人类在方向性变更前被征求意见
  • 有明确的停止条件(指标连续恶化/达到目标)

© LearnPrompt, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in karpathy-autoresearch of LearnPrompt/andrej-karpathy-skills.

Open the folder on GitHubat commit 9e46dec

Compare with similar skills

AutoResearch Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AutoResearch Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AutoResearch Loop this skillLearnPrompt/andrej-karpathy-skills110—~1.4kAutomated safety check: PassMIT
Darwin SkillHHU3637kr/skills1451 repos~2.2kAutomated safety check: PassNone
Huashu Agent Swarmalchaincyf/huashu-skills1.7k—~576Automated safety check: PassMIT
Foremergenaw103/foremerge538—~2.4kAutomated safety check: PassApache-2.0
Session Catchup and Handoffposhan0126/dotclaude870—~811Automated safety check: PassMIT
Light Project StructureLight0305/Light-skills640—~3kAutomated safety check: NotesMIT

Similar skills

  • Darwin Skill

    HHU3637kr/skills

    Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.

    145 GitHub starsUsed in 1 repo~2.2k tokens
    Agent WorkflowsAuto-check passed
  • Huashu Agent Swarm

    alchaincyf/huashu-skills

    多Agent蜂群并行协作,纯git自组织,适合大型项目开发。当用户提到"蜂群模式"、"多agent"、"并行开发"、"agent swarm"时使用。

    1.7k GitHub stars~576 tokensUpdated 19 days ago
    Agent WorkflowsAuto-check passed
  • Foremerge

    naw103/foremerge

    Coordinate parallel coding agents with Foremerge's local Git-compatible CLI and MCP server.

    538 GitHub stars~2.4k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Session Catchup and Handoff

    poshan0126/dotclaude

    Rebuilds working context after /clear by reading a handoff note and the branch's git state, or writes that handoff note before a session ends.

    870 GitHub stars~811 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Light Project Structure

    Light0305/Light-skills

    Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.

    640 GitHub stars~3k tokensUpdated 3 mo ago
    DevelopmentAuto-check: notes
  • Rebuild Branch

    platformplatform/PlatformPlatform

    Rebuild a stale branch by cherry-picking each commit onto a fresh branch off main, using a ralph-loop to validate each commit (build, test, format, lint, optional e2e) before moving on.

    441 GitHub stars~2.1k tokensUpdated 17 days ago
    Agent WorkflowsAuto-check: notes

More from LearnPrompt/andrej-karpathy-skills

All 15 skills in this repo
  • Karpathy Methodology Index

    LearnPrompt/andrej-karpathy-skills

    Apply Andrej Karpathy AI methodology and principles from his 2023-2026 insights. Use this skill when the user wants to apply Karpathy-style thinking, needs…

    110 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Agentic Engineering

    LearnPrompt/andrej-karpathy-skills

    Apply Karpathy-style agentic engineering to any coding or building task.

    110 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Education First

    LearnPrompt/andrej-karpathy-skills

    Apply the education-first mindset — make everything you build teachable, create nano-project explanations, write for beginners.

    110 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Idea Files

    LearnPrompt/andrej-karpathy-skills

    Create and share ideas as abstract Gist-style specs instead of code — letting agents or others implement.

    110 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy LLM Simulator

    LearnPrompt/andrej-karpathy-skills

    Use LLM as a simulator of expert debates and opposing viewpoints instead of getting a single sycophantic answer.

    110 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy LLM Wiki

    LearnPrompt/andrej-karpathy-skills

    Build and maintain an LLM-powered personal knowledge base or wiki.

    110 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about AutoResearch Loop

What does AutoResearch Loop do?

Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change. The core idea is that you change the prompt and the agent changes everything else: it reads the current state, runs an experiment, analyzes the results and proposes the next iteration, then waits for you to approve the prompt change before repeating. You stay in the hypothesis space and the agent stays in implementation and execution.

When should I use AutoResearch Loop?

AutoResearch Loop fits situations like: automating a series of ML experiments or hyperparameter runs; setting up a research agent that proposes the next experiment itself; tuning a prompt or an implementation's speed through measured iterations; summarizing an experiment log after several iterations.

How do I install AutoResearch Loop in Claude Code?

Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a claude-code`. Or copy the skill folder (karpathy-autoresearch in LearnPrompt/andrej-karpathy-skills) into .claude/skills/karpathy-autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install AutoResearch Loop in Codex?

Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a codex`. Or copy the skill folder (karpathy-autoresearch in LearnPrompt/andrej-karpathy-skills) into .agents/skills/karpathy-autoresearch in your project. Codex loads it when a task matches its description.

Can I use AutoResearch Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/karpathy-autoresearch, .gemini/skills/karpathy-autoresearch, .github/skills/karpathy-autoresearch and .opencode/skills/karpathy-autoresearch in your project.

What does AutoResearch Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: AutoResearch Loop is instructions for the agent only. Our summary lists: A git repository for the experiment branches.

Does AutoResearch Loop access the network?

SKILL.md names 1 domain. As links in the text: x.com. This is read from the text; nothing was executed.

Is AutoResearch Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AutoResearch Loop use?

AutoResearch Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AutoResearch Loop use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AutoResearch Loop?

Skills that share tags, products or a category with AutoResearch Loop: Darwin Skill (HHU3637kr/skills, 145 stars), Huashu Agent Swarm (alchaincyf/huashu-skills, 1.7k stars), Foremerge (naw103/foremerge, 538 stars) and Session Catchup and Handoff (poshan0126/dotclaude, 870 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AutoResearch Loop?

LearnPrompt (a GitHub user) maintains it in LearnPrompt/andrej-karpathy-skills, which has 110 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on July 10, 2026.

Source: LearnPrompt/andrej-karpathy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.