Agent skill

Autoresearch

by lamm-mit in lamm-mit/scienceclaw

Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.

Apache-2.0Auto-check: notesAgent Workflows

Install Autoresearch

skills CLI
$ npx skills add lamm-mit/scienceclaw --skill autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lamm-mit/scienceclaw autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lamm-mit/scienceclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autoresearch .claude/skills/autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autoresearch
GitHub stars
244
Token cost
~1.2k tokens
SKILL.md length
502 words
Files
3 (incl. scripts)
Skills in repo
86
Repo updated
First seen
Licence
Apache-2.0

At a glance

Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.

  • Works in 6 steps: Read program.md for instructions → Modify train.py (hyperparameters,… → Run training for exactly 5 minutes → …
  • Tasks that involve Autonomous loops
  • SKILL.md covers autoresearch, Prerequisites, Installation and How to run, plus 1 more section
  • Runs Python scripts from its folder; calls uv, git and curl; reaches astral.sh

What it does

Autoresearch is an agent skill from lamm-mit/scienceclaw. Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/USAGE.md` and `scripts/autoresearch_client.py`).

It sits in Agent Workflows, covering Autonomous loops. It works with arXiv. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Autonomous loops

Example prompts

  • “/autoresearch”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read program.md for instructions
  2. Modify train.py (hyperparameters, architecture, optimizer, batch size, etc.)
  3. Run training for exactly 5 minutes
  4. Evaluate using val_bpb (validation bits per byte)
  5. Keep or discard changes based on improvement
  6. Repeat autonomously

What it can do on your machine

Read from SKILL.md and the folder at commit ab9aba1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • git
    • curl
    • sh
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • astral.sh

    Also links to:

    • github.com
    • docs.astral.sh
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autoresearch loads about 1.2k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 502 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePipes a well-known installer script into a shellSKILL.md:46
    curl -LsSf https://astral.sh/uv/install.sh | sh

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from lamm-mit/scienceclaw at commit ab9aba1, republished under its Apache-2.0 licence (© lamm-mit). 502 words, ~1,218 tokens.

Download SKILL.mdSave it as .claude/skills/autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
autoresearch
description
Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.
source_type
github
auth_required
false
repository_url
https://github.com/karpathy/autoresearch
reference_url
https://github.com/karpathy/nanochat

autoresearch

Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.

Code repository

https://github.com/karpathy/autoresearch

Use this as the implementation source: clone the repo and follow its README for install, dependencies, and how to run code or experiments. The generated client prints JSON with a suggested git clone command.

Primary resource (landing page)

https://github.com/karpathy/nanochat

This is the paper or artifact home from DOI/registry metadata — not a JSON API. If this URL is arXiv, the generated client can still fetch live Atom metadata (title, abstract, authors) without a BASE_URL. For other hosts, the client uses stub mode until you set a real BASE_URL for a REST service.

What “running” this client does

The *_client.py script prints JSON that combines a GitHub repository (clone URL + suggested git clone) with optional paper context from arXiv (live Atom metadata when reference_url is arXiv). Run the real code by cloning the repo and following its README — the skill is your agent-facing entrypoint, not a substitute for the repo’s install steps.

To call a REST API instead, set BASE_URL in scripts/autoresearch_client.py or wrap the upstream CLI with subprocess after clone.

How to run the method (from the source)

Extracted for operators and agents. Confirm against the upstream repository or paper before relying on it in production.

Prerequisites

  • Single NVIDIA GPU (tested on H100)
  • Python 3.10+
  • uv project manager

Installation

bash
# 1. Install uv project manager (if you don't already have it)
curl -LsSf https://astral.sh/uv/install.sh | sh

# 2. Install dependencies
uv sync

# 3. Download data and train tokenizer (one-time, ~2 min)
uv run prepare.py

How to run

Manual single training experiment (~5 min):

bash
uv run train.py

Autonomous agent mode:

Point your AI agent (Claude, Codex, etc.) to the program.md file and prompt:

Hi have a look at program.md and let's kick off a new experiment! let's do the setup first.

The agent will autonomously:

  1. Read program.md for instructions
  2. Modify train.py (hyperparameters, architecture, optimizer, batch size, etc.)
  3. Run training for exactly 5 minutes
  4. Evaluate using val_bpb (validation bits per byte)
  5. Keep or discard changes based on improvement
  6. Repeat autonomously
Show full SKILL.md (196 more words)Show less

Configuration

Key files to understand:

  • prepare.py — Fixed constants, one-time data prep (downloads training data, trains BPE tokenizer), runtime utilities. Do not modify.
  • train.py — Single file edited by the agent. Contains GPT model, optimizer (Muon + AdamW), training loop. Fair game: architecture, hyperparameters, batch size, optimizer settings.
  • program.md — Baseline instructions for agents. Edit this to customize agent behavior and research setup.

Training constraints:

  • Fixed 5-minute time budget (wall clock, excluding startup/compilation) regardless of compute platform
  • Metric: val_bpb (validation bits per byte, lower is better, vocab-size-independent)
  • Expected frequency: ~12 experiments/hour, ~100 experiments overnight

For smaller compute platforms (MacBook, etc.), tune in prepare.py and train.py:

  • Use smaller dataset (e.g., TinyStories)
  • Decrease vocab_size (from 8192 to 4096, 2048, or 256 bytes)
  • Lower MAX_SEQ_LEN (down to 256)
  • Reduce EVAL_TOKENS for faster validation
  • Lower DEPTH (default 8, try 4)
  • Use WINDOW_PATTERN: "L" instead of "SSSL"
  • Reduce TOTAL_BATCH_SIZE to 2**14 (~16K) or lower

Refer to notable forks for CPU/MacOS/Windows/AMD variants.

The same text lives in scripts/USAGE.md for tools that prefer reading files under scripts/.

Parameters

--time-budget (int) [optional, default=5] Fixed wall-clock training duration in minutes (default: 5) --metric (str) [optional, default=val_bpb] Optimization metric: val_bpb (validation bits per byte, lower is better)

Usage
bash
python3 scripts/autoresearch_client.py uv run train.py
Example Output
json
{"val_bpb": 1.234, "epoch": 1, "loss": 2.567}

© lamm-mit, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/autoresearch of lamm-mit/scienceclaw.

  • SKILL.md
  • scripts/USAGE.md
  • scripts/autoresearch_client.py

Open the folder on GitHubat commit ab9aba1

Compare with similar skills

Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autoresearch this skilllamm-mit/scienceclaw244—~1.2kAutomated safety check: NotesApache-2.0
Show Me Your Work Decision Logcursor/plugins10k8 repos~1.6kAutomated safety check: PassNone
Autoresearch Iteration Loopuditgoenka/autoresearch6.5k1 repos~2kAutomated safety check: PassMIT
Install Loop Engineeringcobusgreyling/loop-engineering11k1 repos~648Automated safety check: PassMIT
LoopyForward-Future/loopy3.2k—~3.9kAutomated safety check: PassMIT
AI Performance Improvement Plantanweai/pua20k2 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Install Loop Engineering

    cobusgreyling/loop-engineering

    Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.

    11k GitHub starsUsed in 1 repo~648 tokens
    Agent WorkflowsAuto-check passed
  • Loopy

    Forward-Future/loopy

    Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.

    3.2k GitHub stars~3.9k tokensUpdated 28 days ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.

    20k GitHub starsUsed in 2 repos~6.9k tokens
    Agent WorkflowsAuto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from lamm-mit/scienceclaw

All 86 skills in this repo
  • Fred Economic Data

    lamm-mit/scienceclaw

    Query FRED (Federal Reserve Economic Data) API for 800,000+ economic time series from 100+ sources.

    244 GitHub starsUsed in 4 repos~3k tokens
    Auto-check passed
  • Drug Research

    lamm-mit/scienceclaw

    Generates comprehensive drug research reports with compound disambiguation, evidence grading, and mandatory completeness sections.

    244 GitHub starsUsed in 3 repos~1.7k tokens
    Auto-check passed
  • Imaging Data Commons

    lamm-mit/scienceclaw

    Query and download public cancer imaging data from NCI Imaging Data Commons using idc-index.

    244 GitHub starsUsed in 5 repos~11k tokens
    Auto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~3.1k tokens
    Auto-check: warnings
  • Infographics

    lamm-mit/scienceclaw

    Create professional infographics using Nano Banana Pro AI with smart iterative refinement.

    244 GitHub starsUsed in 6 repos~4.4k tokens
    Auto-check: notes
  • Disease Research

    lamm-mit/scienceclaw

    Generate comprehensive disease research reports using 100+ ToolUniverse tools.

    244 GitHub stars~946 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Autoresearch

What does Autoresearch do?

Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte. Autoresearch is an agent skill from lamm-mit/scienceclaw. Autonomous AI agent that modifies and iteratively improves a GPT language model training setup, running experiments within a 5-minute time budget to optimize validation bits-per-byte.

When should I use Autoresearch?

Autoresearch fits situations like: tasks that involve Autonomous loops.

How do I install Autoresearch in Claude Code?

Run `npx skills add lamm-mit/scienceclaw --skill autoresearch -a claude-code`. Or copy the skill folder (skills/autoresearch in lamm-mit/scienceclaw) into .claude/skills/autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Autoresearch in Codex?

Run `npx skills add lamm-mit/scienceclaw --skill autoresearch -a codex`. Or copy the skill folder (skills/autoresearch in lamm-mit/scienceclaw) into .agents/skills/autoresearch in your project. Codex loads it when a task matches its description.

Can I use Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lamm-mit/scienceclaw --skill autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch, .gemini/skills/autoresearch, .github/skills/autoresearch and .opencode/skills/autoresearch in your project.

What does Autoresearch need to run?

Going by SKILL.md and its folder, Autoresearch needs Python for the scripts in its folder and the command-line tools its instructions call (uv, git, curl, sh and python3). Our summary lists: Python 3.

Does Autoresearch access the network?

SKILL.md names 4 domains. In commands or code: astral.sh; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, docs.astral.sh and huggingface.co. This is read from the text; nothing was executed.

Is Autoresearch safe to install?

Our automated static check of SKILL.md found notes only (pipes a well-known installer script into a shell), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Autoresearch use?

Autoresearch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autoresearch use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autoresearch?

Skills that share tags, products or a category with Autoresearch: Show Me Your Work Decision Log (cursor/plugins, 10k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autoresearch?

lamm-mit (a GitHub user) maintains it in lamm-mit/scienceclaw, which has 244 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on August 21, 2026.

Source: lamm-mit/scienceclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.