Agent skill

Autoresearch

by OpenLAIR in OpenLAIR/dr-claw

Autonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw.

MITAuto-check passedAgent Workflows

Install Autoresearch

skills CLI
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenLAIR/dr-claw autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autoresearch .claude/skills/autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autoresearch
GitHub stars
1.2k
Token cost
~7.5k tokens
SKILL.md length
2,853 words
Files
12 (incl. references)
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Autonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw.

  • Works in 3 steps: Check if the user provided ALL required… → If ANY required context is missing → you… → Each subcommand's reference file has an…
  • Tasks that involve Autonomous loops
  • SKILL.md covers MANDATORY: Interactive Setup…, Subcommands, When to Activate and Bounded Iterations, plus 5 more sections
  • Calls npm, git and gh

What it does

Autoresearch is an agent skill from OpenLAIR/dr-claw. Autonomous Goal-directed Iteration. Apply Karpathy's autoresearch principles to ANY task. Loops autonomously — modify, verify, keep/discard, repeat. 9 subcommands: plan, debug, fix, security, ship, scenario, predict, learn.

Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/autonomous-loop-protocol.md`, `references/core-principles.md` and `references/debug-workflow.md`).

It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power. The licence is MIT.

When your agent uses it

  • Tasks that involve Autonomous loops

Example prompts

  • “/autoresearch”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Check if the user provided ALL required context inline (Goal, Scope, Metric, flags, etc.)
  2. If ANY required context is missing → you MUST use AskUserQuestion to collect it BEFORE proceeding to any execution phase. DO NOT skip this…
  3. Each subcommand's reference file has an "Interactive Setup" section — follow it exactly when context is missing.

What it can do on your machine

Read from SKILL.md and the folder at commit d51b64e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • git
    • gh
    • kubectl
    • npx
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autoresearch loads about 7.5k tokens when it runs, and up to ~64k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 2,853 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~64k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenLAIR/dr-claw at commit d51b64e, republished under its MIT licence (© OpenLAIR). 2,853 words, ~7,509 tokens.

Download SKILL.mdSave it as .claude/skills/autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
autoresearch
description
Autonomous Goal-directed Iteration. Apply Karpathy's autoresearch principles to ANY task. Loops autonomously — modify, verify, keep/discard, repeat. 9 subcommands: plan, debug, fix, security, ship, scenario, predict, learn.
version
1.8.2
license
MIT
metadata.author
uditgoenka/autoresearch
metadata.version
1.8.2

Claude Autoresearch — Autonomous Goal-directed Iteration

Inspired by Karpathy's autoresearch. Applies constraint-driven autonomous iteration to ANY work — not just ML research.

Core idea: You are an autonomous agent. Modify → Verify → Keep/Discard → Repeat.

MANDATORY: Interactive Setup Gate

CRITICAL — READ THIS FIRST BEFORE ANY ACTION:

For ALL commands (/autoresearch, /autoresearch:plan, /autoresearch:debug, /autoresearch:fix, /autoresearch:security, /autoresearch:ship, /autoresearch:scenario, /autoresearch:predict, /autoresearch:learn):

  1. Check if the user provided ALL required context inline (Goal, Scope, Metric, flags, etc.)
  2. If ANY required context is missing → you MUST use AskUserQuestion to collect it BEFORE proceeding to any execution phase. DO NOT skip this step. DO NOT proceed without user input.
  3. Each subcommand's reference file has an "Interactive Setup" section — follow it exactly when context is missing.
CommandRequired ContextIf Missing → Ask
/autoresearchGoal, Scope, Metric, Direction, VerifyBatch 1 (4 questions) + Batch 2 (3 questions) from Setup Phase below
/autoresearch:planGoalAsk via AskUserQuestion per references/plan-workflow.md
/autoresearch:debugIssue/Symptom, Scope4 batched questions per references/debug-workflow.md
/autoresearch:fixTarget, Scope4 batched questions per references/fix-workflow.md
/autoresearch:securityScope, Depth3 batched questions per references/security-workflow.md
/autoresearch:shipWhat/Type, Mode3 batched questions per references/ship-workflow.md
/autoresearch:scenarioScenario, Domain4-8 adaptive questions per references/scenario-workflow.md
/autoresearch:predictScope, Goal3-4 batched questions per references/predict-workflow.md
/autoresearch:learnMode, Scope4 batched questions per references/learn-workflow.md

YOU MUST NOT start any loop, phase, or execution without completing interactive setup when context is missing. This is a BLOCKING prerequisite.

Subcommands

SubcommandPurpose
/autoresearchRun the autonomous loop (default)
/autoresearch:planInteractive wizard to build Scope, Metric, Direction & Verify from a Goal
/autoresearch:securityAutonomous security audit: STRIDE threat model + OWASP Top 10 + red-team (4 adversarial personas)
/autoresearch:shipUniversal shipping workflow: ship code, content, marketing, sales, research, or anything
/autoresearch:debugAutonomous bug-hunting loop: scientific method + iterative investigation until codebase is clean
/autoresearch:fixAutonomous fix loop: iteratively repair errors (tests, types, lint, build) until zero remain
/autoresearch:scenarioScenario-driven use case generator: explore situations, edge cases, and derivative scenarios
/autoresearch:predictMulti-persona swarm prediction: pre-analyze code from multiple expert perspectives before acting
/autoresearch:learnAutonomous codebase documentation engine: scout, learn, generate/update docs with validation-fix loop
/autoresearch:security — Autonomous Security Audit

Runs a comprehensive security audit using the autoresearch loop pattern. Generates a full STRIDE threat model, maps attack surfaces, then iteratively tests each vulnerability vector — logging findings with severity, OWASP category, and code evidence.

Load: references/security-workflow.md for full protocol.

What it does:

  1. Codebase Reconnaissance — scans tech stack, dependencies, configs, API routes
  2. Asset Identification — catalogs data stores, auth systems, external services, user inputs
  3. Trust Boundary Mapping — browser↔server, public↔authenticated, user↔admin, CI/CD↔prod
  4. STRIDE Threat Model — Spoofing, Tampering, Repudiation, Info Disclosure, DoS, Elevation of Privilege
  5. Attack Surface Map — entry points, data flows, abuse paths
  6. Autonomous Loop — iteratively tests each vector, validates with code evidence, logs findings
  7. Final Report — severity-ranked findings with mitigations, coverage matrix, iteration log

Key behaviors:

  • Follows red-team adversarial mindset (Security Adversary, Supply Chain, Insider Threat, Infra Attacker)
  • Every finding requires code evidence (file:line + attack scenario) — no theoretical fluff
  • Tracks OWASP Top 10 + STRIDE coverage, prints coverage summary every 5 iterations
  • Composite metric: (owasp_tested/10)*50 + (stride_tested/6)*30 + min(findings, 20) — higher is better
  • Creates security/{YYMMDD}-{HHMM}-{audit-slug}/ folder with structured reports: overview.md, threat-model.md, attack-surface-map.md, findings.md, owasp-coverage.md, dependency-audit.md, recommendations.md, security-audit-results.tsv

Flags:

FlagPurpose
--diffDelta mode — only audit files changed since last audit
--fixAfter audit, auto-fix confirmed Critical/High findings using autoresearch loop
--fail-on {severity}Exit non-zero if findings meet threshold (for CI/CD gating)

Usage:

# Unlimited — keep finding vulnerabilities until interrupted
/autoresearch:security

# Bounded — exactly 10 security sweep iterations
/autoresearch:security
Iterations: 10

# With focused scope
/autoresearch:security
Scope: src/api/**/*.ts, src/middleware/**/*.ts
Focus: authentication and authorization flows

# Delta mode — only audit changed files since last audit
/autoresearch:security --diff

# Auto-fix confirmed Critical/High findings after audit
/autoresearch:security --fix
Iterations: 15

# CI/CD gate — fail pipeline if any Critical findings
/autoresearch:security --fail-on critical
Iterations: 10

# Combined — delta audit + fix + gate
/autoresearch:security --diff --fix --fail-on critical
Iterations: 15

Inspired by:

  • Strix — AI-powered security testing with proof-of-concept validation
  • /plan red-team — adversarial review with hostile reviewer personas
  • OWASP Top 10 (2021) — industry-standard vulnerability taxonomy
  • STRIDE — Microsoft's threat modeling framework
/autoresearch:ship — Universal Shipping Workflow

Ship anything — code, content, marketing, sales, research, or design — through a structured 8-phase workflow that applies autoresearch loop principles to the last mile.

Load: references/ship-workflow.md for full protocol.

What it does:

  1. Identify — auto-detect what you're shipping (code PR, deployment, blog post, email campaign, sales deck, research paper, design assets)
  2. Inventory — assess current state and readiness gaps
  3. Checklist — generate domain-specific pre-ship gates (all mechanically verifiable)
  4. Prepare — autoresearch loop to fix failing checklist items until 100% pass
  5. Dry-run — simulate the ship action without side effects
  6. Ship — execute the actual delivery (merge, deploy, publish, send)
  7. Verify — post-ship health check confirms it landed
  8. Log — record shipment to ship-log.tsv for traceability

Supported shipment types:

TypeExample Ship Actions
code-prgh pr create with full description
code-releaseGit tag + GitHub release
deploymentCI/CD trigger, kubectl apply, push to deploy branch
contentPublish via CMS, commit to content branch
marketing-emailSend via ESP (SendGrid, Mailchimp)
marketing-campaignActivate ads, launch landing page
salesSend proposal, share deck
researchUpload to repository, submit paper
designExport assets, share with stakeholders

Flags:

FlagPurpose
--dry-runValidate everything but don't actually ship (stop at Phase 5)
--autoAuto-approve dry-run gate if no errors
--forceSkip non-critical checklist items (blockers still enforced)
--rollbackUndo the last ship action (if reversible)
--monitor NPost-ship monitoring for N minutes
--type <type>Override auto-detection with explicit shipment type
--checklist-onlyOnly generate and evaluate checklist (stop at Phase 3)

Usage:

# Auto-detect and ship (interactive)
/autoresearch:ship

# Ship code PR with auto-approve
/autoresearch:ship --auto

# Dry-run a deployment before going live
/autoresearch:ship --type deployment --dry-run

# Ship with post-deployment monitoring
/autoresearch:ship --monitor 10

# Prepare iteratively then ship
/autoresearch:ship
Iterations: 5

# Just check if something is ready to ship
/autoresearch:ship --checklist-only

# Ship a blog post
/autoresearch:ship
Target: content/blog/my-new-post.md
Type: content

# Ship a sales deck
/autoresearch:ship --type sales
Target: decks/q1-proposal.pdf

# Rollback a bad deployment
/autoresearch:ship --rollback

Composite metric (for bounded loops):

ship_score = (checklist_passing / checklist_total) * 80
           + (dry_run_passed ? 15 : 0)
           + (no_blockers ? 5 : 0)

Score of 100 = fully ready. Below 80 = not shippable.

Output directory: Creates ship/{YYMMDD}-{HHMM}-{ship-slug}/ with checklist.md, ship-log.tsv, summary.md.

/autoresearch:scenario — Scenario-Driven Use Case Generator

Autonomous scenario exploration engine that generates, expands, and stress-tests use cases from a seed scenario. Discovers edge cases, failure modes, and derivative scenarios that manual analysis misses.

Load: references/scenario-workflow.md for full protocol.

What it does:

  1. Seed Analysis — parse scenario, identify actors, goals, preconditions, components
  2. Decomposition — break into 12 exploration dimensions (happy path, error, edge case, abuse, scale, concurrent, temporal, data variation, permission, integration, recovery, state transition)
  3. Situation Generation — create one concrete situation per iteration from unexplored dimensions
  4. Classification — deduplicate (new/variant/duplicate/out-of-scope/low-value)
  5. Expansion — derive edge cases, what-ifs, failure modes from each kept situation
  6. Logging — record to scenario-results.tsv with dimension, severity, classification
  7. Repeat — pick next unexplored dimension/combination, iterate

Key behaviors:

  • Adaptive interactive setup: 4-8 questions based on how much context the user provides
  • 12 exploration dimensions ensure comprehensive coverage
  • Domain-specific templates (software, product, business, security, marketing)
  • Every situation requires concrete trigger, flow, and expected outcome — no vague "something goes wrong"
  • Composite metric: scenarios_generated*10 + edge_cases_found*15 + (dimensions_covered/12)*30 + unique_actors*5
  • Creates scenario/{YYMMDD}-{HHMM}-{slug}/ with: scenarios.md, use-cases.md, edge-cases.md, scenario-results.tsv, summary.md

Flags:

FlagPurpose
--domain <type>Set domain (software, product, business, security, marketing)
--depth <level>Exploration depth: shallow (10), standard (25), deep (50+)
--scope <glob>Limit to specific files/features
--format <type>Output: use-cases, user-stories, test-scenarios, threat-scenarios, mixed
--focus <area>Prioritize dimension: edge-cases, failures, security, scale

Usage:

# Unlimited — keep exploring until interrupted
/autoresearch:scenario

# Bounded with context
/autoresearch:scenario
Scenario: User attempts checkout with multiple payment methods
Domain: software
Depth: standard
Iterations: 25

# Quick edge case scan
/autoresearch:scenario --depth shallow --focus edge-cases
Scenario: File upload feature for profile pictures

# Security-focused
/autoresearch:scenario --domain security
Scenario: OAuth2 login flow with third-party providers
Iterations: 30

# Generate test scenarios
/autoresearch:scenario --format test-scenarios --domain software
Scenario: REST API pagination with filtering and sorting
/autoresearch:predict — Multi-Persona Swarm Prediction

Multi-perspective code analysis using swarm intelligence principles. Simulates 3-5 expert personas (Architect, Security Analyst, Performance Engineer, Reliability Engineer, Devil's Advocate) that independently analyze code, debate findings, and reach consensus — all within Claude's native context. Zero external dependencies.

Load: references/predict-workflow.md for full protocol.

What it does:

  1. Codebase Reconnaissance — scan files, extract entities, map dependencies into knowledge .md files
  2. Persona Generation — create 3-5 expert personas from codebase context
  3. Independent Analysis — each persona analyzes code from their unique perspective
  4. Structured Debate — 1-2 rounds of cross-examination with mandatory Devil's Advocate dissent
  5. Consensus — synthesizer aggregates findings with confidence scores + anti-herd check
  6. Knowledge Output — write predict/ folder with codebase-analysis.md, dependency-map.md, component-clusters.md
  7. Report — generate findings.md, hypothesis-queue.md, overview.md
  8. Handoff — write handoff.json for optional --chain to debug/security/fix/ship/scenario

Key behaviors:

  • File-based knowledge representation: .md files ARE the knowledge graph, zero external deps
  • Git-hash stamping: every output embeds commit SHA for staleness detection
  • Incremental updates: only re-analyzes files changed since last run
  • Anti-herd mechanism: Devil's Advocate mandatory, groupthink detection via flip rate + entropy
  • Empirical evidence always trumps swarm prediction when chained with autoresearch loop
  • Composite metric: findings_confirmed*15 + findings_probable*8 + minority_preserved*3 + (personas/total)*20 + (rounds/planned)*10 + anti_herd_passed*5
  • Creates predict/{YYMMDD}-{HHMM}-{slug}/ folder with: overview.md, codebase-analysis.md, dependency-map.md, component-clusters.md, persona-debates.md, hypothesis-queue.md, findings.md, predict-results.tsv, handoff.json

Flags:

FlagPurpose
--chain <targets>Chain to tools. Single: --chain debug. Multi: --chain scenario,debug,fix (sequential)
--personas NNumber of personas (default: 5, range: 3-8)
--rounds NDebate rounds (default: 2, range: 1-3)
--depth <level>Depth preset: shallow (3 personas, 1 round), standard (5, 2), deep (8, 3)
--adversarialUse adversarial persona set (Red Team, Blue Team, Insider, Supply Chain, Judge)
--budget <N>Max total findings across all personas (default: 40)
--fail-on <severity>Exit non-zero if findings at or above severity (for CI/CD)
--scope <glob>Limit analysis to specific files

Usage:

# Standard analysis
/autoresearch:predict
Scope: src/**/*.ts
Goal: Find reliability issues

# Quick security scan
/autoresearch:predict --depth shallow --chain security
Scope: src/api/**

# Deep analysis with adversarial debate
/autoresearch:predict --depth deep --adversarial
Goal: Pre-deployment quality audit

# CI/CD gate
/autoresearch:predict --fail-on critical --budget 20
Scope: src/**
Iterations: 1

# Chain to debug for hypothesis-driven investigation
/autoresearch:predict --chain debug
Scope: src/auth/**
Goal: Investigate intermittent 500 errors

# Multi-chain: predict → scenario → debug → fix (sequential pipeline)
/autoresearch:predict --chain scenario,debug,fix
Scope: src/**
Goal: Full quality pipeline for new feature
/autoresearch:learn — Autonomous Codebase Documentation Engine

Scouts codebase structure, learns patterns and architecture, generates/updates comprehensive documentation — then validates and iteratively improves until docs match codebase reality.

Load: references/learn-workflow.md for full protocol.

What it does:

  1. Scout — parallel codebase reconnaissance with scale awareness and monorepo detection
  2. Analyze — project type classification, tech stack detection, staleness measurement
  3. Map — dynamic doc discovery (docs/*.md), gap analysis, conditional doc selection
  4. Generate — spawn docs-manager with structured prompt template and full context
  5. Validate — mechanical verification (code refs, links, completeness, size compliance)
  6. Fix — validation-fix loop: re-generate failed docs with feedback (max 3 retries)
  7. Finalize — inventory check, git diff summary, size compliance
  8. Log — record results to learn-results.tsv

4 Modes:

ModePurposeAutoresearch Loop?
initLearn codebase from scratch, generate all docsYes — validate-fix cycle
updateLearn what changed, refresh existing docsYes — validate-fix cycle
checkRead-only health/staleness assessmentNo — diagnostic only
summarizeQuick codebase summary with file inventoryMinimal — size check only

Key behaviors:

  • Fully dynamic doc discovery — scans docs/*.md, no hardcoded file lists
  • State-aware mode detection — auto-selects init/update based on docs/ state
  • Project-type-adaptive — creates deployment-guide.md only if deployment config exists
  • Validation-fix loop capped at 3 retries — escalates to user if unresolved
  • Scale-aware scouting — adjusts parallelism for 5k+ file codebases
  • Composite metric: learn_score = validation%×0.5 + coverage%×0.3 + size_compliance%×0.2
  • Creates learn/{YYMMDD}-{HHMM}-{slug}/ with: learn-results.tsv, summary.md, validation-report.md, scout-context.md

Flags:

FlagPurpose
--mode <mode>Operation: init, update, check, summarize (default: auto-detect)
--scope <glob>Limit codebase learning to specific dirs
--depth <level>Doc comprehensiveness: quick, standard, deep
--scanForce fresh scout in summarize mode
--topics <list>Focus summarize on specific topics
--file <name>Selective update — target single doc
--no-fixSkip validation-fix loop
--format <fmt>Output format: markdown (default). Planned: confluence, rst, html

Usage:

# Auto-detect mode and learn
/autoresearch:learn

# Initialize docs for new project
/autoresearch:learn --mode init --depth deep

# Update docs after changes
/autoresearch:learn --mode update
Iterations: 3

# Read-only health check
/autoresearch:learn --mode check

# Quick summary
/autoresearch:learn --mode summarize --scan

# Selective update of one doc
/autoresearch:learn --mode update --file system-architecture.md

# Scoped learning
/autoresearch:learn --scope src/api/**
Iterations: 5
Show full SKILL.md (1,207 more words)Show less
/autoresearch:plan — Goal → Configuration Wizard

Converts a plain-language goal into a validated, ready-to-execute autoresearch configuration.

Load: references/plan-workflow.md for full protocol.

Quick summary:

  1. Capture Goal — ask what the user wants to improve (or accept inline text)
  2. Analyze Context — scan codebase for tooling, test runners, build scripts
  3. Define Scope — suggest file globs, validate they resolve to real files
  4. Define Metric — suggest mechanical metrics, validate they output a number
  5. Define Direction — higher or lower is better
  6. Define Verify — construct the shell command, dry-run it, confirm it works
  7. Confirm & Launch — present the complete config, offer to launch immediately

Critical gates:

  • Metric MUST be mechanical (outputs a parseable number, not subjective)
  • Verify command MUST pass a dry run on the current codebase before accepting
  • Scope MUST resolve to ≥1 file

Usage:

/autoresearch:plan
Goal: Make the API respond faster

/autoresearch:plan Increase test coverage to 95%

/autoresearch:plan Reduce bundle size below 200KB

After the wizard completes, the user gets a ready-to-paste /autoresearch invocation — or can launch it directly.

When to Activate

  • User invokes /autoresearch → run the loop
  • User invokes /autoresearch:plan → run the planning wizard
  • User invokes /autoresearch:security → run the security audit
  • User says "help me set up autoresearch", "plan an autoresearch run" → run the planning wizard
  • User says "security audit", "threat model", "OWASP", "STRIDE", "find vulnerabilities", "red-team" → run the security audit
  • User invokes /autoresearch:ship → run the ship workflow
  • User says "ship it", "deploy this", "publish this", "launch this", "get this out the door" → run the ship workflow
  • User invokes /autoresearch:debug → run the debug loop
  • User says "find all bugs", "hunt bugs", "debug this", "why is this failing", "investigate" → run the debug loop
  • User invokes /autoresearch:fix → run the fix loop
  • User says "fix all errors", "make tests pass", "fix the build", "clean up errors" → run the fix loop
  • User invokes /autoresearch:scenario → run the scenario loop
  • User says "explore scenarios", "generate use cases", "what could go wrong", "stress test this feature", "edge cases for" → run the scenario loop
  • User invokes /autoresearch:learn → run the learn workflow
  • User says "learn this codebase", "generate docs", "document this project", "create documentation", "update docs", "check docs", "docs health" → run the learn workflow
  • User invokes /autoresearch:predict → run the predict workflow
  • User says "predict", "multi-perspective", "swarm analysis", "what do multiple experts think", "analyze from different angles" → run the predict workflow
  • User says "work autonomously", "iterate until done", "keep improving", "run overnight" → run the loop
  • Any task requiring repeated iteration cycles with measurable outcomes → run the loop

Bounded Iterations

By default, autoresearch loops forever until manually interrupted. To run exactly N iterations, add Iterations: N to your inline config.

Unlimited (default):

/autoresearch
Goal: Increase test coverage to 90%

Bounded (N iterations):

/autoresearch
Goal: Increase test coverage to 90%
Iterations: 25

After N iterations Claude stops and prints a final summary with baseline → current best, keeps/discards/crashes. If the goal is achieved before N iterations, Claude prints early completion and stops.

When to Use Bounded Iterations
ScenarioRecommendation
Run overnight, review in morningUnlimited (default)
Quick 30-min improvement sessionIterations: 10
Targeted fix with known scopeIterations: 5
Exploratory — see if approach worksIterations: 15
CI/CD pipeline integration--iterations N flag (set N based on time budget)

Setup Phase (Do Once)

If the user provides Goal, Scope, Metric, and Verify inline → extract them and proceed to step 5.

CRITICAL: If ANY critical field is missing (Goal, Scope, Metric, Direction, or Verify), you MUST use AskUserQuestion to collect them interactively. DO NOT proceed to The Loop or any execution phase without completing this setup. This is a BLOCKING prerequisite.

Interactive Setup (when invoked without full config)

Scan the codebase first for smart defaults, then ask ALL questions in batched AskUserQuestion calls (max 4 per call). This gives users full clarity upfront.

Batch 1 — Core config (4 questions in one call):

Use a SINGLE AskUserQuestion call with these 4 questions:

#HeaderQuestionOptions (smart defaults from codebase scan)
1Goal"What do you want to improve?""Test coverage (higher)", "Bundle size (lower)", "Performance (faster)", "Code quality (fewer errors)"
2Scope"Which files can autoresearch modify?"Suggested globs from project structure (e.g. "src//*.ts", "content//*.md")
3Metric"What number tells you if it got better? (must be a command output, not subjective)"Detected options: "coverage % (higher)", "bundle size KB (lower)", "error count (lower)", "test pass count (higher)"
4Direction"Higher or lower is better?""Higher is better", "Lower is better"

Batch 2 — Verify + Guard + Launch (3 questions in one call):

#HeaderQuestionOptions
5Verify"What command produces the metric? (I'll dry-run it to confirm)"Suggested commands from detected tooling
6Guard"Any command that must ALWAYS pass? (prevents regressions)""npm test", "tsc --noEmit", "npm run build", "Skip — no guard"
7Launch"Ready to go?""Launch (unlimited)", "Launch with iteration limit", "Edit config", "Cancel"

After Batch 2: Dry-run the verify command. If it fails, ask user to fix or choose a different command. If it passes, proceed with launch choice.

IMPORTANT: You MUST call AskUserQuestion with batched questions — never ask one at a time, and never skip this step. Users should see all config choices together for full context. DO NOT proceed to Setup Steps or The Loop without completing interactive setup.

Setup Steps (after config is complete)
  1. Read all in-scope files for full context before any modification
  2. Define the goal — extracted from user input or inline config
  3. Define scope constraints — validated file globs
  4. Define guard (optional) — regression prevention command
  5. Create a results log — Track every iteration (see references/results-logging.md)
  6. Establish baseline — Run verification on current state AND guard (if set). Record as iteration #0
  7. Confirm and go — Show user the setup, get confirmation, then BEGIN THE LOOP

The Loop

Read references/autonomous-loop-protocol.md for full protocol details.

LOOP (FOREVER or N times):
  1. Review: Read current state + git history + results log
  2. Ideate: Pick next change based on goal, past results, what hasn't been tried
  3. Modify: Make ONE focused change to in-scope files
  4. Commit: Git commit the change (before verification)
  5. Verify: Run the mechanical metric (tests, build, benchmark, etc.)
  6. Guard: If guard is set, run the guard command
  7. Decide:
     - IMPROVED + guard passed (or no guard) → Keep commit, log "keep", advance
     - IMPROVED + guard FAILED → Revert, then try to rework the optimization
       (max 2 attempts) so it improves the metric WITHOUT breaking the guard.
       Never modify guard/test files — adapt the implementation instead.
       If still failing → log "discard (guard failed)" and move on
     - SAME/WORSE → Git revert, log "discard"
     - CRASHED → Try to fix (max 3 attempts), else log "crash" and move on
  8. Log: Record result in results log
  9. Repeat: Go to step 1.
     - If unbounded: NEVER STOP. NEVER ASK "should I continue?"
     - If bounded (N): Stop after N iterations, print final summary

Critical Rules

  1. Loop until done — Unbounded: loop until interrupted. Bounded: loop N times then summarize.
  2. Read before write — Always understand full context before modifying
  3. One change per iteration — Atomic changes. If it breaks, you know exactly why
  4. Mechanical verification only — No subjective "looks good". Use metrics
  5. Automatic rollback — Failed changes revert instantly. No debates
  6. Simplicity wins — Equal results + less code = KEEP. Tiny improvement + ugly complexity = DISCARD
  7. Git is memory — Every experiment committed with experiment: prefix. Use git revert (not git reset --hard) for rollbacks so failed experiments remain visible in history. Agent MUST read git log and git diff of kept commits to learn patterns before each iteration
  8. When stuck, think harder — Re-read files, re-read goal, combine near-misses, try radical changes. Don't ask for help unless truly blocked by missing access/permissions

Principles Reference

See references/core-principles.md for the 7 generalizable principles from autoresearch.

Adapting to Different Domains

DomainMetricScopeVerify CommandGuard
Backend codeTests pass + coverage %src/**/*.tsnpm test—
Frontend UILighthouse scoresrc/components/**npx lighthousenpm test
ML trainingval_bpb / losstrain.pyuv run train.py—
Blog/contentWord count + readabilitycontent/*.mdCustom script—
PerformanceBenchmark time (ms)Target filesnpm run benchnpm test
RefactoringTests pass + LOC reducedTarget modulenpm test && wc -lnpm run typecheck
SecurityOWASP + STRIDE coverage + findingsAPI/auth/middleware/autoresearch:security—
ShippingChecklist pass rate (%)Any artifact/autoresearch:shipDomain-specific
DebuggingBugs found + coverageTarget files/autoresearch:debug—
FixingError count (lower)Target files/autoresearch:fixnpm test
Scenario analysisScenario coverage score (higher)Feature/domain files/autoresearch:scenario—
ScenariosUse cases + edge cases + dimension coverageTarget feature/files/autoresearch:scenario—
PredictionFindings + hypotheses (higher)Target files/autoresearch:predict—
DocumentationValidation pass rate (higher)docs/*.md/autoresearch:learnnpm test

Adapt the loop to your domain. The PRINCIPLES are universal; the METRICS are domain-specific.

© OpenLAIR, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in skills/autoresearch of OpenLAIR/dr-claw.

  • SKILL.md
  • references/autonomous-loop-protocol.md
  • references/core-principles.md
  • references/debug-workflow.md
  • references/fix-workflow.md
  • references/learn-workflow.md
  • references/plan-workflow.md
  • references/predict-workflow.md
  • references/results-logging.md
  • references/scenario-workflow.md
  • references/security-workflow.md
  • references/ship-workflow.md

Open the folder on GitHubat commit d51b64e

Compare with similar skills

Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autoresearch this skillOpenLAIR/dr-claw1.2k—~7.5kAutomated safety check: PassMIT
Show Me Your Work Decision Logcursor/plugins11k8 repos~1.6kAutomated safety check: PassNone
Autoresearch Iteration Loopuditgoenka/autoresearch6.5k1 repos~2kAutomated safety check: PassMIT
Install Loop Engineeringcobusgreyling/loop-engineering11k1 repos~648Automated safety check: PassMIT
LoopyForward-Future/loopy3.2k—~3.9kAutomated safety check: PassMIT
AI Performance Improvement Plantanweai/pua20k2 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    11k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Install Loop Engineering

    cobusgreyling/loop-engineering

    Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.

    11k GitHub starsUsed in 1 repo~648 tokens
    Agent WorkflowsAuto-check passed
  • Loopy

    Forward-Future/loopy

    Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.

    3.2k GitHub stars~3.9k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.

    20k GitHub starsUsed in 2 repos~6.9k tokens
    Agent WorkflowsAuto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from OpenLAIR/dr-claw

All 36 skills in this repo
  • Analyzes reviewer comments and drafts venue-specific rebuttals for AI and computer science conferences, with an issue board, task list and paper edit plan.

    1.2k GitHub stars~4.9k tokensUpdated 23 days ago
    Auto-check: notes
  • Turns a research paper into a slide deck and, optionally, a narrated demo video, through script, slide generation, text-to-speech and video assembly stages you control.

    1.2k GitHub stars~1.5k tokensUpdated 23 days ago
    Auto-check passed
  • Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.

    1.2k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check: notes
  • ML Dataset Discovery

    OpenLAIR/dr-claw

    Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.

    1.2k GitHub stars~741 tokensUpdated 23 days ago
    Auto-check passed
  • Gemini Deep Research

    OpenLAIR/dr-claw

    Runs multi-source web research through Google's Gemini Deep Research Agent with a bundled Python script and saves a structured, cited report as files.

    1.2k GitHub stars~996 tokensUpdated 23 days ago
    Auto-check passed
  • Six-phase workflow for writing, revising and adapting grant proposals for NSF, NIH, DOE, DARPA, NASA and China's NSFC, from profiling through simulated peer review.

    1.2k GitHub stars~9.2k tokensUpdated 23 days ago
    Auto-check passed

Categories

Questions about Autoresearch

What does Autoresearch do?

Autonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw. Autoresearch is an agent skill from OpenLAIR/dr-claw. Autonomous Goal-directed Iteration.

When should I use Autoresearch?

Autoresearch fits situations like: tasks that involve Autonomous loops.

How do I install Autoresearch in Claude Code?

Run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a claude-code`. Or copy the skill folder (skills/autoresearch in OpenLAIR/dr-claw) into .claude/skills/autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Autoresearch in Codex?

Run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a codex`. Or copy the skill folder (skills/autoresearch in OpenLAIR/dr-claw) into .agents/skills/autoresearch in your project. Codex loads it when a task matches its description.

Can I use Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch, .gemini/skills/autoresearch, .github/skills/autoresearch and .opencode/skills/autoresearch in your project.

What does Autoresearch need to run?

Going by SKILL.md and its folder, Autoresearch needs the command-line tools its instructions call (npm, git, gh, kubectl, npx and uv).

Does Autoresearch access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Autoresearch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autoresearch use?

Autoresearch is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autoresearch use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 56k tokens, read only when the agent opens those files.

What are the alternatives to Autoresearch?

Skills that share tags, products or a category with Autoresearch: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autoresearch?

OpenLAIR (a GitHub organization) maintains it in OpenLAIR/dr-claw, which has 1,155 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 17, 2026.

Source: OpenLAIR/dr-claw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.