Darwin Skill
HHU3637kr/skills
Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.
Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/karpathy-autoresearch .claude/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .claude/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/karpathy-autoresearch .agents/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .agents/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/karpathy-autoresearch .cursor/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .cursor/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LearnPrompt/andrej-karpathy-skills.git --path karpathy-autoresearch--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/karpathy-autoresearch .gemini/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .gemini/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/karpathy-autoresearch .github/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .github/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-autoresearch --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/karpathy-autoresearch .opencode/skills/karpathy-autoresearch && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "karpathy-autoresearch" agent skill from https://github.com/LearnPrompt/andrej-karpathy-skills/tree/main/karpathy-autoresearch into .opencode/skills/karpathy-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "karpathy-autoresearch", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
karpathy-autoresearchSets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.
The core idea is that you change the prompt and the agent changes everything else: it reads the current state, runs an experiment, analyzes the results and proposes the next iteration, then waits for you to approve the prompt change before repeating. You stay in the hypothesis space and the agent stays in implementation and execution. The skill supplies a one-time setup prompt, a per-iteration prompt, a JSONL experiment log (id, branch, hypothesis, metric, runtime, notes) and a summary prompt to run after several iterations.
Each experiment lives on its own git branch, so earlier runs are never lost. The pattern is described as usable beyond machine learning, for prompt tuning, code speed, content formats and UX flows, each with its own metric. Safety rules tell the agent never to delete earlier data or commits and never to touch files outside the current experiment's scope. It is based on a post by Andrej Karpathy, and the excerpt is cut off at the prompt contract.
Read from SKILL.md and the folder at commit 9e46dec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are jsonl).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
x.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AutoResearch Loop loads about 1.4k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 153 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LearnPrompt/andrej-karpathy-skills at commit 9e46dec, republished under its MIT licence (© LearnPrompt). 153 words, ~1,401 tokens.
.claude/skills/karpathy-autoresearch/SKILL.md (or your agent's skills folder).Source: https://x.com/karpathy/status/2030371219518931079 "Autoresearch project — agent self-iterates training code" — 28k likes
You change the prompt. The agent changes everything else.
The loop: agent reads current state → runs experiment → analyzes results → proposes next iteration → you approve the prompt change → repeat.
You stay in the hypothesis space. Agent stays in the implementation + execution space.
Human Layer (you):
- Define research question
- Approve hypothesis changes
- Set evaluation metric
- Stop when satisfied
Agent Layer:
- Implement current hypothesis
- Run experiment (on git branch)
- Analyze logs/results
- Propose next hypothesis
- Never change the research questionYou are AutoResearcher for this project.
Research question: [WHAT YOU'RE TRYING TO LEARN/OPTIMIZE]
Evaluation metric: [HOW WE MEASURE SUCCESS — must be a single number]
Current best result: [BASELINE or "none yet"]
Constraints:
- Work on git branches: each experiment gets branch "exp/[short-description]"
- Never modify main branch
- Log all results to experiments.jsonl in format: {"id": N, "hypothesis": "...", "metric": value, "notes": "..."}
- Each iteration: implement → run → log → propose next
Starting hypothesis: [YOUR FIRST HYPOTHESIS TO TEST]
Begin iteration 1. Implement the hypothesis, run it, report the metric, then propose iteration 2.AutoResearcher: Iteration [N]
Current state:
- Branch: exp/[current]
- Metric so far: [RESULTS LOG]
- Best result: [BEST SO FAR]
Completed last iteration: [WHAT HAPPENED]
Metric result: [VALUE]
Your tasks this iteration:
1. Analyze: what does the result tell us about the hypothesis?
2. Propose: what's the next hypothesis? (one variable change at a time)
3. Implement: write the code change for this hypothesis
4. Run: execute and report the metric
5. Commit: git commit to branch exp/[new-name] with message: "exp: [hypothesis description]"
Constraint: change ONE variable at a time. If last experiment failed, back off to simpler hypothesis.{"id": 1, "branch": "exp/baseline", "hypothesis": "default hyperparams", "metric": 0.423, "runtime_s": 45, "notes": "starting point"}
{"id": 2, "branch": "exp/lr-0.001", "hypothesis": "lower learning rate 0.01→0.001", "metric": 0.451, "runtime_s": 47, "notes": "small improvement, keep lowering?"}
{"id": 3, "branch": "exp/lr-0.0001", "hypothesis": "lower learning rate 0.001→0.0001", "metric": 0.438, "runtime_s": 46, "notes": "worse than id=2, sweet spot was 0.001"}Summarize this research session.
Experiment log:
[PASTE experiments.jsonl]
Produce:
1. Best configuration found (all hyperparameters/settings)
2. What we learned (2-3 bullet points about the search space)
3. What we should try next (top 3 unexplored directions)
4. A clear hypothesis about WHY the best configuration works
5. Confidence level: how sure are we this is near-optimal?This pattern works for any iterative optimization:
| Domain | Research Question | Metric |
|---|---|---|
| ML training | Best architecture/hyperparams | Validation loss |
| Prompt engineering | Best system prompt | Task success rate |
| Code optimization | Fastest implementation | Runtime (ms) |
| Content creation | Most engaging format | Click/read rate |
| Product design | Best UX flow | Task completion rate |
AutoResearcher rules (always include):
- NEVER delete data or commits from previous experiments
- NEVER modify any file outside the current experiment's scope
- NEVER run experiments that cost > [COST_LIMIT] without checking
- ALWAYS commit before starting the next iteration
- ALWAYS report the metric before proposing the next hypothesis
- If metric degrades 3 iterations in a row, STOP and ask for guidance属于工作流:研究到发布(内容创作者/研究者)
| 位置 | 上游 | 下游 |
|---|---|---|
| 入口 | 用户有研究问题时直接触发 | karpathy-llm-wiki(把实验结果沉淀为知识页) |
完整链路:autoresearch → llm-wiki → output-evolution → education-first
You are AutoResearcher. Research question: <QUESTION>. Evaluation metric: <SINGLE_NUMBER_METRIC>. Current baseline: <VALUE or "none">. Starting hypothesis: <FIRST_HYPOTHESIS>. Work on git branches (exp/<name>). Each iteration: implement → run → log result to experiments.jsonl → analyze → propose next hypothesis (one variable change). Never modify main. Stop and ask if metric degrades 3 iterations in a row.© LearnPrompt, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in karpathy-autoresearch of LearnPrompt/andrej-karpathy-skills.
Open the folder on GitHubat commit 9e46dec
AutoResearch Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AutoResearch Loop this skillLearnPrompt/andrej-karpathy-skills | 110 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Darwin SkillHHU3637kr/skills | 145 | 1 repos | ~2.2k | Automated safety check: Pass | None | |
| Huashu Agent Swarmalchaincyf/huashu-skills | 1.7k | — | ~576 | Automated safety check: Pass | MIT | |
| Foremergenaw103/foremerge | 538 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Session Catchup and Handoffposhan0126/dotclaude | 870 | — | ~811 | Automated safety check: Pass | MIT | |
| Light Project StructureLight0305/Light-skills | 640 | — | ~3k | Automated safety check: Notes | MIT |
HHU3637kr/skills
Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.
alchaincyf/huashu-skills
多Agent蜂群并行协作,纯git自组织,适合大型项目开发。当用户提到"蜂群模式"、"多agent"、"并行开发"、"agent swarm"时使用。
naw103/foremerge
Coordinate parallel coding agents with Foremerge's local Git-compatible CLI and MCP server.
poshan0126/dotclaude
Rebuilds working context after /clear by reading a handoff note and the branch's git state, or writes that handoff note before a session ends.
Light0305/Light-skills
Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.
platformplatform/PlatformPlatform
Rebuild a stale branch by cherry-picking each commit onto a fresh branch off main, using a ralph-loop to validate each commit (build, test, format, lint, optional e2e) before moving on.
LearnPrompt/andrej-karpathy-skills
Apply Andrej Karpathy AI methodology and principles from his 2023-2026 insights. Use this skill when the user wants to apply Karpathy-style thinking, needs…
LearnPrompt/andrej-karpathy-skills
Apply Karpathy-style agentic engineering to any coding or building task.
LearnPrompt/andrej-karpathy-skills
Apply the education-first mindset — make everything you build teachable, create nano-project explanations, write for beginners.
LearnPrompt/andrej-karpathy-skills
Create and share ideas as abstract Gist-style specs instead of code — letting agents or others implement.
LearnPrompt/andrej-karpathy-skills
Use LLM as a simulator of expert debates and opposing viewpoints instead of getting a single sycophantic answer.
LearnPrompt/andrej-karpathy-skills
Build and maintain an LLM-powered personal knowledge base or wiki.
Works with
Categories
Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change. The core idea is that you change the prompt and the agent changes everything else: it reads the current state, runs an experiment, analyzes the results and proposes the next iteration, then waits for you to approve the prompt change before repeating. You stay in the hypothesis space and the agent stays in implementation and execution.
AutoResearch Loop fits situations like: automating a series of ML experiments or hyperparameter runs; setting up a research agent that proposes the next experiment itself; tuning a prompt or an implementation's speed through measured iterations; summarizing an experiment log after several iterations.
Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a claude-code`. Or copy the skill folder (karpathy-autoresearch in LearnPrompt/andrej-karpathy-skills) into .claude/skills/karpathy-autoresearch in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a codex`. Or copy the skill folder (karpathy-autoresearch in LearnPrompt/andrej-karpathy-skills) into .agents/skills/karpathy-autoresearch in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/karpathy-autoresearch, .gemini/skills/karpathy-autoresearch, .github/skills/karpathy-autoresearch and .opencode/skills/karpathy-autoresearch in your project.
SKILL.md names no scripts, command-line tools or credentials: AutoResearch Loop is instructions for the agent only. Our summary lists: A git repository for the experiment branches.
SKILL.md names 1 domain. As links in the text: x.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AutoResearch Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with AutoResearch Loop: Darwin Skill (HHU3637kr/skills, 145 stars), Huashu Agent Swarm (alchaincyf/huashu-skills, 1.7k stars), Foremerge (naw103/foremerge, 538 stars) and Session Catchup and Handoff (poshan0126/dotclaude, 870 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LearnPrompt (a GitHub user) maintains it in LearnPrompt/andrej-karpathy-skills, which has 110 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on July 10, 2026.
Source: LearnPrompt/andrej-karpathy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.