Agent skill

Autoresearch

by Factory-AI in Factory-AI/factory-plugins

Autonomous experiment loop for optimization research. An agent skill from Factory-AI/factory-plugins.

No licenceAuto-check passedAgent Workflows

Install Autoresearch

skills CLI
$ npx skills add Factory-AI/factory-plugins --skill autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Factory-AI/factory-plugins autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Factory-AI/factory-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/autoresearch/skills/autoresearch .claude/skills/autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autoresearch
GitHub stars
111
Token cost
~4.2k tokens
SKILL.md length
1,511 words
Files
2
Skills in repo
13
Repo updated
First seen
Licence
None found

At a glance

Autonomous experiment loop for optimization research. An agent skill from Factory-AI/factory-plugins.

  • Works in 9 steps: Gather Information → Create Branch and State Files → Initialize JSONL and Commit State Files → …
  • The user wants to: - Optimize a metric through systematic experimentation (ML training loss
  • SKILL.md covers Overview, Setup, The Experiment Loop and State Files Reference, plus 5 more sections
  • Runs Python scripts from its folder; calls git, python3 and pnpm

What it does

Autoresearch is an agent skill from Factory-AI/factory-plugins. Autonomous experiment loop for optimization research. Use when the user wants to: - Optimize a metric through systematic experimentation (ML training loss, test speed, bundle size, build time, etc.) - Run an automated research loop: try an idea, measure it, keep improvements, revert regressions, repeat - Set up autoresearch for any codebase with a measurable optimization target Implements the autoresearch pattern with MAD-based confidence scoring, git branch isolation, and structured experiment logging.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `autoresearch_helper.py`).

It sits in Agent Workflows, covering Autonomous loops, Git workflow and Web performance. The repository describes itself as: Official Factory plugins marketplace.

When your agent uses it

  • The user wants to: - Optimize a metric through systematic experimentation (ML training loss
  • Etc.) - Run an automated research loop: try an idea
  • Keep improvements
  • Revert regressions

Example prompts

  • “/autoresearch”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Gather Information
  2. Create Branch and State Files
  3. Initialize JSONL and Commit State Files
  4. Run Baseline
  5. Summarize Results
  6. Group Changes
  7. Resolve File Conflicts
  8. Create Clean Branches
  9. Verify and Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7166a07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • python3
    • pnpm
    • uv
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, pnpm and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autoresearch loads about 4.2k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 1,511 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,511 words (~4,188 tokens).

“Autonomous experiment loop: try ideas, keep what works, discard what doesn't, never stop.”

— opening of SKILL.md by Factory-AI
name
autoresearch
version
1.0.0

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in plugins/autoresearch/skills/autoresearch of Factory-AI/factory-plugins.

  • SKILL.md
  • autoresearch_helper.py

Open the folder on GitHubat commit 7166a07

Compare with similar skills

Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autoresearch this skillFactory-AI/factory-plugins111—~4.2kAutomated safety check: PassNone
Rebuild Branchplatformplatform/PlatformPlatform441—~2.1kAutomated safety check: NotesMIT
Darwin SkillHHU3637kr/skills1451 repos~2.2kAutomated safety check: PassNone
AutoResearch LoopLearnPrompt/andrej-karpathy-skills110—~1.4kAutomated safety check: PassMIT
Karpathygaasher/Agent-Loop-Skills174—~2.6kAutomated safety check: PassMIT
Slate Ar Recipeudecode/plate17k—~658Automated safety check: PassCustom licence

Similar skills

  • Rebuild Branch

    platformplatform/PlatformPlatform

    Rebuild a stale branch by cherry-picking each commit onto a fresh branch off main, using a ralph-loop to validate each commit (build, test, format, lint, optional e2e) before moving on.

    441 GitHub stars~2.1k tokensUpdated 15 days ago
    Agent WorkflowsAuto-check: notes
  • Darwin Skill

    HHU3637kr/skills

    Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.

    145 GitHub starsUsed in 1 repo~2.2k tokens
    Agent WorkflowsAuto-check passed
  • AutoResearch Loop

    LearnPrompt/andrej-karpathy-skills

    Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

    110 GitHub stars~1.4k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Slate Ar Recipe

    udecode/plate

    Slate v2 Autoresearch recipe picker. An agent skill from udecode/plate.

    17k GitHub stars~658 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed

More from Factory-AI/factory-plugins

All 13 skills in this repo
  • Ban Type Assertions

    Factory-AI/factory-plugins

    Ban as type assertions in a package via the @typescript-eslint/consistent-type-assertions lint rule, replacing them with compiler-verified type-safe alternatives.

    111 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed
  • Browser Navigation

    Factory-AI/factory-plugins

    Automate browser interactions for web testing, form filling, screenshots, and data extraction.

    111 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Commit Security Scan

    Factory-AI/factory-plugins

    Analyze code changes for security vulnerabilities using LLM reasoning and threat model patterns.

    111 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Create PR

    Factory-AI/factory-plugins

    Create a pull request with Conventional Commits formatting, a templated body, and local verification.

    111 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Droid Control

    Factory-AI/factory-plugins

    Control terminal TUIs, browsers, and native desktop apps for testing, demos, QA, and computer-use tasks.

    111 GitHub stars~3.6k tokensUpdated 2 days ago
    Auto-check: notes
  • Fix Knip Unused Exports

    Factory-AI/factory-plugins

    Fix knip "Unused exports" violations. An agent skill from Factory-AI/factory-plugins.

    111 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Autoresearch

What does Autoresearch do?

Autonomous experiment loop for optimization research. An agent skill from Factory-AI/factory-plugins. Autoresearch is an agent skill from Factory-AI/factory-plugins. Autonomous experiment loop for optimization research.

When should I use Autoresearch?

Autoresearch fits situations like: the user wants to: - Optimize a metric through systematic experimentation (ML training loss; etc.) - Run an automated research loop: try an idea; keep improvements; revert regressions.

How do I install Autoresearch in Claude Code?

Run `npx skills add Factory-AI/factory-plugins --skill autoresearch -a claude-code`. Or copy the skill folder (plugins/autoresearch/skills/autoresearch in Factory-AI/factory-plugins) into .claude/skills/autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Autoresearch in Codex?

Run `npx skills add Factory-AI/factory-plugins --skill autoresearch -a codex`. Or copy the skill folder (plugins/autoresearch/skills/autoresearch in Factory-AI/factory-plugins) into .agents/skills/autoresearch in your project. Codex loads it when a task matches its description.

Can I use Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Factory-AI/factory-plugins --skill autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch, .gemini/skills/autoresearch, .github/skills/autoresearch and .opencode/skills/autoresearch in your project.

What does Autoresearch need to run?

Going by SKILL.md and its folder, Autoresearch needs Python for the scripts in its folder and the command-line tools its instructions call (git, python3, pnpm, uv and bash). Our summary lists: Python 3.

Does Autoresearch access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Autoresearch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autoresearch use?

No licence was found for Autoresearch or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Autoresearch use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autoresearch?

Skills that share tags, products or a category with Autoresearch: Rebuild Branch (platformplatform/PlatformPlatform, 441 stars), Darwin Skill (HHU3637kr/skills, 145 stars), AutoResearch Loop (LearnPrompt/andrej-karpathy-skills, 110 stars) and Karpathy (gaasher/Agent-Loop-Skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autoresearch?

Factory-AI (a GitHub organization) maintains it in Factory-AI/factory-plugins, which has 111 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 7, 2026.

Source: Factory-AI/factory-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.