Agent skill

Autoresearch Agent

by alirezarezvani in alirezarezvani/claude-skills

Autonomous experiment loop that optimizes any file by a measurable metric.

MITAuto-check passedAgent Workflows

Install Autoresearch Agent

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill autoresearch-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills autoresearch-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/autoresearch-agent/skills/autoresearch-agent .claude/skills/autoresearch-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autoresearch-agent
GitHub stars
28k
Used in
1 other repo
Token cost
~3k tokens
SKILL.md length
1,093 words
Files
6 (incl. scripts, references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Autonomous experiment loop that optimizes any file by a measurable metric.

  • Works in 4 steps: Read… → Read program.md for strategy,… → Read results.tsv for experiment history… → …
  • : user wants to optimize code speed
  • SKILL.md covers Slash Commands, When This Skill Activates, Setup and Agent Protocol, plus 5 more sections
  • Runs Python scripts from its folder; calls python, git and gemini

What it does

Autoresearch Agent is an agent skill from alirezarezvani/claude-skills. Autonomous experiment loop that optimizes any file by a measurable metric. Inspired by Karpathy's autoresearch. The agent edits a target file, runs a fixed evaluation, keeps improvements (git commit), discards failures (git reset), and loops indefinitely. Use when: user wants to optimize code speed, reduce bundle/image size, improve test pass rate, optimize prompts, improve content quality (headlines, copy, CTR), or run any measurable improvement loop. Requires: a target file, an evaluation command that outputs a…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/experiment-domains.md`, `references/program-template.md` and `scripts/log_results.py`).

It sits in Agent Workflows, covering Autonomous loops. It works with Git. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • : user wants to optimize code speed
  • Reduce bundle/image size
  • Improve test pass rate
  • Optimize prompts

Example prompts

  • “/autoresearch-agent”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read .autoresearch/{domain}/{name}/config.cfg to get
  2. Read program.md for strategy, constraints, and what you can/cannot change
  3. Read results.tsv for experiment history (columns: commit, metric, status, description)
  4. Checkout the experiment branch: git checkout autoresearch/{domain}/{name}

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git
    • gemini
    • cursor

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autoresearch Agent loads about 3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 1,093 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 1,093 words, ~3,047 tokens.

Download SKILL.mdSave it as .claude/skills/autoresearch-agent/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
autoresearch-agent
description
Autonomous experiment loop that optimizes any file by a measurable metric. Inspired by Karpathy's autoresearch. The agent edits a target file, runs a fixed evaluation, keeps improvements (git commit), discards failures (git reset), and loops indefinitely. Use when: user wants to optimize code speed, reduce bundle/image size, improve test pass rate, optimize prompts, improve content quality (headlines, copy, CTR), or run any measurable improvement loop. Requires: a target file, an evaluation command that outputs a metric, and a git repo.
license
MIT
metadata.version
2.0.0
metadata.author
Alireza Rezvani
metadata.category
engineering
metadata.updated
2026-03-13

Autoresearch Agent

You sleep. The agent experiments. You wake up to results.

Autonomous experiment loop inspired by Karpathy's autoresearch. The agent edits one file, runs a fixed evaluation, keeps improvements, discards failures, and loops indefinitely.

Not one guess — fifty measured attempts, compounding.


Slash Commands

CommandWhat it does
/ar:setupSet up a new experiment interactively
/ar:runRun a single experiment iteration
/ar:loopStart autonomous loop with configurable interval (10m, 1h, daily, weekly, monthly)
/ar:ar-statusShow dashboard and results
/ar:ar-resumeResume a paused experiment

When This Skill Activates

Recognize these patterns from the user:

  • "Make this faster / smaller / better"
  • "Optimize [file] for [metric]"
  • "Improve my [headlines / copy / prompts]"
  • "Run experiments overnight"
  • "I want to get [metric] from X to Y"
  • Any request involving: optimize, benchmark, improve, experiment loop, autoresearch

If the user describes a target file + a way to measure success → this skill applies.


Setup

First Time — Create the Experiment

Run the setup script. The user decides where experiments live:

Project-level (inside repo, git-tracked, shareable with team):

bash
python scripts/setup_experiment.py \
  --domain engineering \
  --name api-speed \
  --target src/api/search.py \
  --eval "pytest bench.py --tb=no -q" \
  --metric p50_ms \
  --direction lower \
  --scope project

User-level (personal, in ~/.autoresearch/):

bash
python scripts/setup_experiment.py \
  --domain marketing \
  --name medium-ctr \
  --target content/titles.md \
  --eval "python evaluate.py" \
  --metric ctr_score \
  --direction higher \
  --evaluator llm_judge_content \
  --scope user

The --scope flag determines where .autoresearch/ lives:

  • project (default) → .autoresearch/ in the repo root. Experiment definitions are git-tracked. Results are gitignored.
  • user → ~/.autoresearch/ in the home directory. Everything is personal.
What Setup Creates
.autoresearch/
├── config.yaml                        ← Global settings
├── .gitignore                         ← Ignores results.tsv, *.log
└── {domain}/{experiment-name}/
    ├── program.md                     ← Objectives, constraints, strategy
    ├── config.cfg                     ← Target, eval cmd, metric, direction
    ├── results.tsv                    ← Experiment log (gitignored)
    └── evaluate.py                    ← Evaluation script (if --evaluator used)

results.tsv columns: commit | metric | status | description

  • commit — short git hash
  • metric — float value or "N/A" for crashes
  • status — keep | discard | crash
  • description — what changed or why it crashed
Domains
DomainUse Cases
engineeringCode speed, memory, bundle size, test pass rate, build time
marketingHeadlines, social copy, email subjects, ad copy, engagement
contentArticle structure, SEO descriptions, readability, CTR
promptsSystem prompts, chatbot tone, agent instructions
customAnything else with a measurable metric
If program.md Already Exists

The user may have written their own program.md. If found in the experiment directory, read it. It overrides the template. Only ask for what's missing.


Agent Protocol

You are the loop. The scripts handle setup and evaluation — you handle the creative work.

Before Starting
  1. Read .autoresearch/{domain}/{name}/config.cfg to get:
    • target — the file you edit
    • evaluate_cmd — the command that measures your changes
    • metric — the metric name to look for in eval output
    • metric_direction — "lower" or "higher" is better
    • time_budget_minutes — max time per evaluation
  2. Read program.md for strategy, constraints, and what you can/cannot change
  3. Read results.tsv for experiment history (columns: commit, metric, status, description)
  4. Checkout the experiment branch: git checkout autoresearch/{domain}/{name}
Each Iteration
  1. Review results.tsv — what worked? What failed? What hasn't been tried?
  2. Decide ONE change to the target file. One variable per experiment.
  3. Edit the target file
  4. Commit: git add {target} && git commit -m "experiment: {description}"
  5. Evaluate: python scripts/run_experiment.py --experiment {domain}/{name} --single
  6. Read the output — it prints KEEP, DISCARD, or CRASH with the metric value
  7. Go to step 1
What the Script Handles (you don't)
  • Running the eval command with timeout
  • Parsing the metric from eval output
  • Comparing to previous best
  • Reverting the commit on failure (git reset --hard HEAD~1)
  • Logging the result to results.tsv
Starting an Experiment
bash
# Single iteration (the agent calls this repeatedly)
python scripts/run_experiment.py --experiment engineering/api-speed --single

# Dry run (test setup before starting)
python scripts/run_experiment.py --experiment engineering/api-speed --dry-run
Strategy Escalation
  • Runs 1-5: Low-hanging fruit (obvious improvements, simple optimizations)
  • Runs 6-15: Systematic exploration (vary one parameter at a time)
  • Runs 16-30: Structural changes (algorithm swaps, architecture shifts)
  • Runs 30+: Radical experiments (completely different approaches)
  • If no improvement in 20+ runs: update program.md Strategy section
Self-Improvement

After every 10 experiments, review results.tsv for patterns. Update the Strategy section of program.md with what you learned (e.g., "caching changes consistently improve by 5-10%", "refactoring attempts never improve the metric"). Future iterations benefit from this accumulated knowledge.

Stopping
  • Run until interrupted by the user, context limit reached, or goal in program.md is met
  • Before stopping: ensure results.tsv is up to date
  • On context limit: the next session can resume — results.tsv and git log persist
Show full SKILL.md (475 more words)Show less
Rules
  • One change per experiment. Don't change 5 things at once. You won't know what worked.
  • Simplicity criterion. A small improvement that adds ugly complexity is not worth it. Equal performance with simpler code is a win. Removing code that gets same results is the best outcome.
  • Never modify the evaluator. evaluate.py is the ground truth. Modifying it invalidates all comparisons. Hard stop if you catch yourself doing this.
  • Timeout. If a run exceeds 2.5× the time budget, kill it and treat as crash.
  • Crash handling. If it's a typo or missing import, fix and re-run. If the idea is fundamentally broken, revert, log "crash", move on. 5 consecutive crashes → pause and alert.
  • No new dependencies. Only use what's already available in the project.

Evaluators

Ready-to-use evaluation scripts. Copied into the experiment directory during setup with --evaluator.

Free Evaluators (no API cost)
EvaluatorMetricUse Case
benchmark_speedp50_ms (lower)Function/API execution time
benchmark_sizesize_bytes (lower)File, bundle, Docker image size
test_pass_ratepass_rate (higher)Test suite pass percentage
build_speedbuild_seconds (lower)Build/compile/Docker build time
memory_usagepeak_mb (lower)Peak memory during execution
LLM Judge Evaluators (uses your subscription)
EvaluatorMetricUse Case
llm_judge_contentctr_score 0-10 (higher)Headlines, titles, descriptions
llm_judge_promptquality_score 0-100 (higher)System prompts, agent instructions
llm_judge_copyengagement_score 0-10 (higher)Social posts, ad copy, emails

LLM judges call the CLI tool the user is already running (Claude, Codex, Gemini). The evaluation prompt is locked inside evaluate.py — the agent cannot modify it. This prevents the agent from gaming its own evaluator.

The user's existing subscription covers the cost:

  • Claude Code Max → unlimited Claude calls for evaluation
  • Codex CLI (ChatGPT Pro) → unlimited Codex calls
  • Gemini CLI (free tier) → free evaluation calls
Custom Evaluators

If no built-in evaluator fits, the user writes their own evaluate.py. Only requirement: it must print metric_name: value to stdout.

python
#!/usr/bin/env python3
# My custom evaluator — DO NOT MODIFY after experiment starts
import subprocess
result = subprocess.run(["my-benchmark", "--json"], capture_output=True, text=True)
# Parse and output
print(f"my_metric: {parse_score(result.stdout)}")

Viewing Results

bash
# Single experiment
python scripts/log_results.py --experiment engineering/api-speed

# All experiments in a domain
python scripts/log_results.py --domain engineering

# Cross-experiment dashboard
python scripts/log_results.py --dashboard

# Export formats
python scripts/log_results.py --experiment engineering/api-speed --format csv --output results.csv
python scripts/log_results.py --experiment engineering/api-speed --format markdown --output results.md
python scripts/log_results.py --dashboard --format markdown --output dashboard.md
Dashboard Output
DOMAIN          EXPERIMENT          RUNS  KEPT  BEST         Δ FROM START  STATUS
engineering     api-speed            47    14   185ms        -76.9%        active
engineering     bundle-size          23     8   412KB        -58.3%        paused
marketing       medium-ctr           31    11   8.4/10       +68.0%        active
prompts         support-tone         15     6   82/100       +46.4%        done
Export Formats
  • TSV — default, tab-separated (compatible with spreadsheets)
  • CSV — comma-separated, with proper quoting
  • Markdown — formatted table, readable in GitHub/docs

Proactive Triggers

Flag these without being asked:

  • No evaluation command works → Test it before starting the loop. Run once, verify output.
  • Target file not in git → git init && git add . && git commit -m 'initial' first.
  • Metric direction unclear → Ask: is lower or higher better? Must know before starting.
  • Time budget too short → If eval takes longer than budget, every run crashes.
  • Agent modifying evaluate.py → Hard stop. This invalidates all comparisons.
  • 5 consecutive crashes → Pause the loop. Alert the user. Don't keep burning cycles.
  • No improvement in 20+ runs → Suggest changing strategy in program.md or trying a different approach.

Installation

One-liner (any tool)
bash
git clone https://github.com/alirezarezvani/claude-skills.git
cp -r claude-skills/engineering/autoresearch-agent ~/.claude/skills/
Multi-tool install
bash
./scripts/convert.sh --skill autoresearch-agent --tool codex|gemini|cursor|windsurf|openclaw
OpenClaw
bash
clawhub install cs-autoresearch-agent

  • self-improving-agent — improves an agent's own memory/rules over time. NOT for structured experiment loops.
  • senior-ml-engineer — ML architecture decisions. Complementary — use for initial design, then autoresearch for optimization.
  • tdd-guide — test-driven development. Complementary — tests can be the evaluation function.
  • skill-security-auditor — audit skills before publishing. NOT for optimization loops.

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in engineering/autoresearch-agent/skills/autoresearch-agent of alirezarezvani/claude-skills.

  • SKILL.md
  • references/experiment-domains.md
  • references/program-template.md
  • scripts/log_results.py
  • scripts/run_experiment.py
  • scripts/setup_experiment.py

Open the folder on GitHubat commit 19392f7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in alirezarezvani/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Autoresearch Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autoresearch Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autoresearch Agent this skillalirezarezvani/claude-skills28k1 repos~3kAutomated safety check: PassMIT
Darwin SkillHHU3637kr/skills1451 repos~2.2kAutomated safety check: PassNone
AutoResearch LoopLearnPrompt/andrej-karpathy-skills110—~1.4kAutomated safety check: PassMIT
AutoresearchApocrathia/home-assistant-config179—~1.2kAutomated safety check: PassNone
Self-Improvement Tournament Loopzereight/gitlab-mcp2k1 repos~1.8kAutomated safety check: WarnMIT
Effective Harnessesliangdabiao/exa-research-mcp-skill110—~1.5kAutomated safety check: PassNone

Similar skills

  • Darwin Skill

    HHU3637kr/skills

    Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.

    145 GitHub starsUsed in 1 repo~2.2k tokens
    Agent WorkflowsAuto-check passed
  • AutoResearch Loop

    LearnPrompt/andrej-karpathy-skills

    Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

    110 GitHub stars~1.4k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Autoresearch

    Apocrathia/home-assistant-config

    Optional, operator-gated metric experiments under .scratch/ for Home Assistant Homelab.

    179 GitHub stars~1.2k tokensUpdated 18 days ago
    Agent WorkflowsAuto-check passed
  • Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.

    2k GitHub starsUsed in 1 repo~1.8k tokens
    DevelopmentAuto-check: warnings
  • Effective Harnesses

    liangdabiao/exa-research-mcp-skill

    Long-running agent project harness for Codex, OpenClaw, Claude Code, and other coding agents.

    110 GitHub stars~1.5k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Works with

Categories

Questions about Autoresearch Agent

What does Autoresearch Agent do?

Autonomous experiment loop that optimizes any file by a measurable metric. Autoresearch Agent is an agent skill from alirezarezvani/claude-skills. Autonomous experiment loop that optimizes any file by a measurable metric.

When should I use Autoresearch Agent?

Autoresearch Agent fits situations like: : user wants to optimize code speed; reduce bundle/image size; improve test pass rate; optimize prompts.

How do I install Autoresearch Agent in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill autoresearch-agent -a claude-code`. Or copy the skill folder (engineering/autoresearch-agent/skills/autoresearch-agent in alirezarezvani/claude-skills) into .claude/skills/autoresearch-agent in your project. Claude Code loads it when a task matches its description.

How do I install Autoresearch Agent in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill autoresearch-agent -a codex`. Or copy the skill folder (engineering/autoresearch-agent/skills/autoresearch-agent in alirezarezvani/claude-skills) into .agents/skills/autoresearch-agent in your project. Codex loads it when a task matches its description.

Can I use Autoresearch Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill autoresearch-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch-agent, .gemini/skills/autoresearch-agent, .github/skills/autoresearch-agent and .opencode/skills/autoresearch-agent in your project.

What does Autoresearch Agent need to run?

Going by SKILL.md and its folder, Autoresearch Agent needs Python for the scripts in its folder and the command-line tools its instructions call (python, git, gemini and cursor). Our summary lists: Python 3; Docker.

Does Autoresearch Agent access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Autoresearch Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Autoresearch Agent use?

Autoresearch Agent is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autoresearch Agent use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Autoresearch Agent?

Skills that share tags, products or a category with Autoresearch Agent: Darwin Skill (HHU3637kr/skills, 145 stars), AutoResearch Loop (LearnPrompt/andrej-karpathy-skills, 110 stars), Autoresearch (Apocrathia/home-assistant-config, 179 stars) and Self-Improvement Tournament Loop (zereight/gitlab-mcp, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autoresearch Agent?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,891 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.