Show Me Your Work Decision Log
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
Autonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw.
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OpenLAIR/dr-claw autoresearch --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autoresearch .claude/skills/autoresearch && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .claude/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OpenLAIR/dr-claw autoresearch --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/autoresearch .agents/skills/autoresearch && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .agents/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OpenLAIR/dr-claw autoresearch --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/autoresearch .cursor/skills/autoresearch && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .cursor/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OpenLAIR/dr-claw.git --path skills/autoresearch--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OpenLAIR/dr-claw autoresearch --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/autoresearch .gemini/skills/autoresearch && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .gemini/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OpenLAIR/dr-claw autoresearchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/autoresearch .github/skills/autoresearch && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .github/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenLAIR/dr-claw --skill autoresearch -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OpenLAIR/dr-claw autoresearch --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/autoresearch .opencode/skills/autoresearch && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "autoresearch" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/autoresearch into .opencode/skills/autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
autoresearchAutonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw.
Autoresearch is an agent skill from OpenLAIR/dr-claw. Autonomous Goal-directed Iteration. Apply Karpathy's autoresearch principles to ANY task. Loops autonomously — modify, verify, keep/discard, repeat. 9 subcommands: plan, debug, fix, security, ship, scenario, predict, learn.
Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/autonomous-loop-protocol.md`, `references/core-principles.md` and `references/debug-workflow.md`).
It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d51b64e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmgitghkubectlnpxuvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Autoresearch loads about 7.5k tokens when it runs, and up to ~64k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 2,853 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from OpenLAIR/dr-claw at commit d51b64e, republished under its MIT licence (© OpenLAIR). 2,853 words, ~7,509 tokens.
.claude/skills/autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Inspired by Karpathy's autoresearch. Applies constraint-driven autonomous iteration to ANY work — not just ML research.
Core idea: You are an autonomous agent. Modify → Verify → Keep/Discard → Repeat.
CRITICAL — READ THIS FIRST BEFORE ANY ACTION:
For ALL commands (/autoresearch, /autoresearch:plan, /autoresearch:debug, /autoresearch:fix, /autoresearch:security, /autoresearch:ship, /autoresearch:scenario, /autoresearch:predict, /autoresearch:learn):
AskUserQuestion to collect it BEFORE proceeding to any execution phase. DO NOT skip this step. DO NOT proceed without user input.| Command | Required Context | If Missing → Ask |
|---|---|---|
/autoresearch | Goal, Scope, Metric, Direction, Verify | Batch 1 (4 questions) + Batch 2 (3 questions) from Setup Phase below |
/autoresearch:plan | Goal | Ask via AskUserQuestion per references/plan-workflow.md |
/autoresearch:debug | Issue/Symptom, Scope | 4 batched questions per references/debug-workflow.md |
/autoresearch:fix | Target, Scope | 4 batched questions per references/fix-workflow.md |
/autoresearch:security | Scope, Depth | 3 batched questions per references/security-workflow.md |
/autoresearch:ship | What/Type, Mode | 3 batched questions per references/ship-workflow.md |
/autoresearch:scenario | Scenario, Domain | 4-8 adaptive questions per references/scenario-workflow.md |
/autoresearch:predict | Scope, Goal | 3-4 batched questions per references/predict-workflow.md |
/autoresearch:learn | Mode, Scope | 4 batched questions per references/learn-workflow.md |
YOU MUST NOT start any loop, phase, or execution without completing interactive setup when context is missing. This is a BLOCKING prerequisite.
| Subcommand | Purpose |
|---|---|
/autoresearch | Run the autonomous loop (default) |
/autoresearch:plan | Interactive wizard to build Scope, Metric, Direction & Verify from a Goal |
/autoresearch:security | Autonomous security audit: STRIDE threat model + OWASP Top 10 + red-team (4 adversarial personas) |
/autoresearch:ship | Universal shipping workflow: ship code, content, marketing, sales, research, or anything |
/autoresearch:debug | Autonomous bug-hunting loop: scientific method + iterative investigation until codebase is clean |
/autoresearch:fix | Autonomous fix loop: iteratively repair errors (tests, types, lint, build) until zero remain |
/autoresearch:scenario | Scenario-driven use case generator: explore situations, edge cases, and derivative scenarios |
/autoresearch:predict | Multi-persona swarm prediction: pre-analyze code from multiple expert perspectives before acting |
/autoresearch:learn | Autonomous codebase documentation engine: scout, learn, generate/update docs with validation-fix loop |
Runs a comprehensive security audit using the autoresearch loop pattern. Generates a full STRIDE threat model, maps attack surfaces, then iteratively tests each vulnerability vector — logging findings with severity, OWASP category, and code evidence.
Load: references/security-workflow.md for full protocol.
What it does:
Key behaviors:
(owasp_tested/10)*50 + (stride_tested/6)*30 + min(findings, 20) — higher is bettersecurity/{YYMMDD}-{HHMM}-{audit-slug}/ folder with structured reports:
overview.md, threat-model.md, attack-surface-map.md, findings.md, owasp-coverage.md, dependency-audit.md, recommendations.md, security-audit-results.tsvFlags:
| Flag | Purpose |
|---|---|
--diff | Delta mode — only audit files changed since last audit |
--fix | After audit, auto-fix confirmed Critical/High findings using autoresearch loop |
--fail-on {severity} | Exit non-zero if findings meet threshold (for CI/CD gating) |
Usage:
# Unlimited — keep finding vulnerabilities until interrupted
/autoresearch:security
# Bounded — exactly 10 security sweep iterations
/autoresearch:security
Iterations: 10
# With focused scope
/autoresearch:security
Scope: src/api/**/*.ts, src/middleware/**/*.ts
Focus: authentication and authorization flows
# Delta mode — only audit changed files since last audit
/autoresearch:security --diff
# Auto-fix confirmed Critical/High findings after audit
/autoresearch:security --fix
Iterations: 15
# CI/CD gate — fail pipeline if any Critical findings
/autoresearch:security --fail-on critical
Iterations: 10
# Combined — delta audit + fix + gate
/autoresearch:security --diff --fix --fail-on critical
Iterations: 15Inspired by:
/plan red-team — adversarial review with hostile reviewer personasShip anything — code, content, marketing, sales, research, or design — through a structured 8-phase workflow that applies autoresearch loop principles to the last mile.
Load: references/ship-workflow.md for full protocol.
What it does:
ship-log.tsv for traceabilitySupported shipment types:
| Type | Example Ship Actions |
|---|---|
code-pr | gh pr create with full description |
code-release | Git tag + GitHub release |
deployment | CI/CD trigger, kubectl apply, push to deploy branch |
content | Publish via CMS, commit to content branch |
marketing-email | Send via ESP (SendGrid, Mailchimp) |
marketing-campaign | Activate ads, launch landing page |
sales | Send proposal, share deck |
research | Upload to repository, submit paper |
design | Export assets, share with stakeholders |
Flags:
| Flag | Purpose |
|---|---|
--dry-run | Validate everything but don't actually ship (stop at Phase 5) |
--auto | Auto-approve dry-run gate if no errors |
--force | Skip non-critical checklist items (blockers still enforced) |
--rollback | Undo the last ship action (if reversible) |
--monitor N | Post-ship monitoring for N minutes |
--type <type> | Override auto-detection with explicit shipment type |
--checklist-only | Only generate and evaluate checklist (stop at Phase 3) |
Usage:
# Auto-detect and ship (interactive)
/autoresearch:ship
# Ship code PR with auto-approve
/autoresearch:ship --auto
# Dry-run a deployment before going live
/autoresearch:ship --type deployment --dry-run
# Ship with post-deployment monitoring
/autoresearch:ship --monitor 10
# Prepare iteratively then ship
/autoresearch:ship
Iterations: 5
# Just check if something is ready to ship
/autoresearch:ship --checklist-only
# Ship a blog post
/autoresearch:ship
Target: content/blog/my-new-post.md
Type: content
# Ship a sales deck
/autoresearch:ship --type sales
Target: decks/q1-proposal.pdf
# Rollback a bad deployment
/autoresearch:ship --rollbackComposite metric (for bounded loops):
ship_score = (checklist_passing / checklist_total) * 80
+ (dry_run_passed ? 15 : 0)
+ (no_blockers ? 5 : 0)Score of 100 = fully ready. Below 80 = not shippable.
Output directory: Creates ship/{YYMMDD}-{HHMM}-{ship-slug}/ with checklist.md, ship-log.tsv, summary.md.
Autonomous scenario exploration engine that generates, expands, and stress-tests use cases from a seed scenario. Discovers edge cases, failure modes, and derivative scenarios that manual analysis misses.
Load: references/scenario-workflow.md for full protocol.
What it does:
Key behaviors:
scenarios_generated*10 + edge_cases_found*15 + (dimensions_covered/12)*30 + unique_actors*5scenario/{YYMMDD}-{HHMM}-{slug}/ with: scenarios.md, use-cases.md, edge-cases.md, scenario-results.tsv, summary.mdFlags:
| Flag | Purpose |
|---|---|
--domain <type> | Set domain (software, product, business, security, marketing) |
--depth <level> | Exploration depth: shallow (10), standard (25), deep (50+) |
--scope <glob> | Limit to specific files/features |
--format <type> | Output: use-cases, user-stories, test-scenarios, threat-scenarios, mixed |
--focus <area> | Prioritize dimension: edge-cases, failures, security, scale |
Usage:
# Unlimited — keep exploring until interrupted
/autoresearch:scenario
# Bounded with context
/autoresearch:scenario
Scenario: User attempts checkout with multiple payment methods
Domain: software
Depth: standard
Iterations: 25
# Quick edge case scan
/autoresearch:scenario --depth shallow --focus edge-cases
Scenario: File upload feature for profile pictures
# Security-focused
/autoresearch:scenario --domain security
Scenario: OAuth2 login flow with third-party providers
Iterations: 30
# Generate test scenarios
/autoresearch:scenario --format test-scenarios --domain software
Scenario: REST API pagination with filtering and sortingMulti-perspective code analysis using swarm intelligence principles. Simulates 3-5 expert personas (Architect, Security Analyst, Performance Engineer, Reliability Engineer, Devil's Advocate) that independently analyze code, debate findings, and reach consensus — all within Claude's native context. Zero external dependencies.
Load: references/predict-workflow.md for full protocol.
What it does:
Key behaviors:
findings_confirmed*15 + findings_probable*8 + minority_preserved*3 + (personas/total)*20 + (rounds/planned)*10 + anti_herd_passed*5predict/{YYMMDD}-{HHMM}-{slug}/ folder with: overview.md, codebase-analysis.md, dependency-map.md, component-clusters.md, persona-debates.md, hypothesis-queue.md, findings.md, predict-results.tsv, handoff.jsonFlags:
| Flag | Purpose |
|---|---|
--chain <targets> | Chain to tools. Single: --chain debug. Multi: --chain scenario,debug,fix (sequential) |
--personas N | Number of personas (default: 5, range: 3-8) |
--rounds N | Debate rounds (default: 2, range: 1-3) |
--depth <level> | Depth preset: shallow (3 personas, 1 round), standard (5, 2), deep (8, 3) |
--adversarial | Use adversarial persona set (Red Team, Blue Team, Insider, Supply Chain, Judge) |
--budget <N> | Max total findings across all personas (default: 40) |
--fail-on <severity> | Exit non-zero if findings at or above severity (for CI/CD) |
--scope <glob> | Limit analysis to specific files |
Usage:
# Standard analysis
/autoresearch:predict
Scope: src/**/*.ts
Goal: Find reliability issues
# Quick security scan
/autoresearch:predict --depth shallow --chain security
Scope: src/api/**
# Deep analysis with adversarial debate
/autoresearch:predict --depth deep --adversarial
Goal: Pre-deployment quality audit
# CI/CD gate
/autoresearch:predict --fail-on critical --budget 20
Scope: src/**
Iterations: 1
# Chain to debug for hypothesis-driven investigation
/autoresearch:predict --chain debug
Scope: src/auth/**
Goal: Investigate intermittent 500 errors
# Multi-chain: predict → scenario → debug → fix (sequential pipeline)
/autoresearch:predict --chain scenario,debug,fix
Scope: src/**
Goal: Full quality pipeline for new featureScouts codebase structure, learns patterns and architecture, generates/updates comprehensive documentation — then validates and iteratively improves until docs match codebase reality.
Load: references/learn-workflow.md for full protocol.
What it does:
docs/*.md), gap analysis, conditional doc selection4 Modes:
| Mode | Purpose | Autoresearch Loop? |
|---|---|---|
init | Learn codebase from scratch, generate all docs | Yes — validate-fix cycle |
update | Learn what changed, refresh existing docs | Yes — validate-fix cycle |
check | Read-only health/staleness assessment | No — diagnostic only |
summarize | Quick codebase summary with file inventory | Minimal — size check only |
Key behaviors:
docs/*.md, no hardcoded file listslearn_score = validation%×0.5 + coverage%×0.3 + size_compliance%×0.2learn/{YYMMDD}-{HHMM}-{slug}/ with: learn-results.tsv, summary.md, validation-report.md, scout-context.mdFlags:
| Flag | Purpose |
|---|---|
--mode <mode> | Operation: init, update, check, summarize (default: auto-detect) |
--scope <glob> | Limit codebase learning to specific dirs |
--depth <level> | Doc comprehensiveness: quick, standard, deep |
--scan | Force fresh scout in summarize mode |
--topics <list> | Focus summarize on specific topics |
--file <name> | Selective update — target single doc |
--no-fix | Skip validation-fix loop |
--format <fmt> | Output format: markdown (default). Planned: confluence, rst, html |
Usage:
# Auto-detect mode and learn
/autoresearch:learn
# Initialize docs for new project
/autoresearch:learn --mode init --depth deep
# Update docs after changes
/autoresearch:learn --mode update
Iterations: 3
# Read-only health check
/autoresearch:learn --mode check
# Quick summary
/autoresearch:learn --mode summarize --scan
# Selective update of one doc
/autoresearch:learn --mode update --file system-architecture.md
# Scoped learning
/autoresearch:learn --scope src/api/**
Iterations: 5Converts a plain-language goal into a validated, ready-to-execute autoresearch configuration.
Load: references/plan-workflow.md for full protocol.
Quick summary:
Critical gates:
Usage:
/autoresearch:plan
Goal: Make the API respond faster
/autoresearch:plan Increase test coverage to 95%
/autoresearch:plan Reduce bundle size below 200KBAfter the wizard completes, the user gets a ready-to-paste /autoresearch invocation — or can launch it directly.
/autoresearch → run the loop/autoresearch:plan → run the planning wizard/autoresearch:security → run the security audit/autoresearch:ship → run the ship workflow/autoresearch:debug → run the debug loop/autoresearch:fix → run the fix loop/autoresearch:scenario → run the scenario loop/autoresearch:learn → run the learn workflow/autoresearch:predict → run the predict workflowBy default, autoresearch loops forever until manually interrupted. To run exactly N iterations, add Iterations: N to your inline config.
Unlimited (default):
/autoresearch
Goal: Increase test coverage to 90%Bounded (N iterations):
/autoresearch
Goal: Increase test coverage to 90%
Iterations: 25After N iterations Claude stops and prints a final summary with baseline → current best, keeps/discards/crashes. If the goal is achieved before N iterations, Claude prints early completion and stops.
| Scenario | Recommendation |
|---|---|
| Run overnight, review in morning | Unlimited (default) |
| Quick 30-min improvement session | Iterations: 10 |
| Targeted fix with known scope | Iterations: 5 |
| Exploratory — see if approach works | Iterations: 15 |
| CI/CD pipeline integration | --iterations N flag (set N based on time budget) |
If the user provides Goal, Scope, Metric, and Verify inline → extract them and proceed to step 5.
CRITICAL: If ANY critical field is missing (Goal, Scope, Metric, Direction, or Verify), you MUST use AskUserQuestion to collect them interactively. DO NOT proceed to The Loop or any execution phase without completing this setup. This is a BLOCKING prerequisite.
Scan the codebase first for smart defaults, then ask ALL questions in batched AskUserQuestion calls (max 4 per call). This gives users full clarity upfront.
Batch 1 — Core config (4 questions in one call):
Use a SINGLE AskUserQuestion call with these 4 questions:
| # | Header | Question | Options (smart defaults from codebase scan) |
|---|---|---|---|
| 1 | Goal | "What do you want to improve?" | "Test coverage (higher)", "Bundle size (lower)", "Performance (faster)", "Code quality (fewer errors)" |
| 2 | Scope | "Which files can autoresearch modify?" | Suggested globs from project structure (e.g. "src//*.ts", "content//*.md") |
| 3 | Metric | "What number tells you if it got better? (must be a command output, not subjective)" | Detected options: "coverage % (higher)", "bundle size KB (lower)", "error count (lower)", "test pass count (higher)" |
| 4 | Direction | "Higher or lower is better?" | "Higher is better", "Lower is better" |
Batch 2 — Verify + Guard + Launch (3 questions in one call):
| # | Header | Question | Options |
|---|---|---|---|
| 5 | Verify | "What command produces the metric? (I'll dry-run it to confirm)" | Suggested commands from detected tooling |
| 6 | Guard | "Any command that must ALWAYS pass? (prevents regressions)" | "npm test", "tsc --noEmit", "npm run build", "Skip — no guard" |
| 7 | Launch | "Ready to go?" | "Launch (unlimited)", "Launch with iteration limit", "Edit config", "Cancel" |
After Batch 2: Dry-run the verify command. If it fails, ask user to fix or choose a different command. If it passes, proceed with launch choice.
IMPORTANT: You MUST call AskUserQuestion with batched questions — never ask one at a time, and never skip this step. Users should see all config choices together for full context. DO NOT proceed to Setup Steps or The Loop without completing interactive setup.
references/results-logging.md)Read references/autonomous-loop-protocol.md for full protocol details.
LOOP (FOREVER or N times):
1. Review: Read current state + git history + results log
2. Ideate: Pick next change based on goal, past results, what hasn't been tried
3. Modify: Make ONE focused change to in-scope files
4. Commit: Git commit the change (before verification)
5. Verify: Run the mechanical metric (tests, build, benchmark, etc.)
6. Guard: If guard is set, run the guard command
7. Decide:
- IMPROVED + guard passed (or no guard) → Keep commit, log "keep", advance
- IMPROVED + guard FAILED → Revert, then try to rework the optimization
(max 2 attempts) so it improves the metric WITHOUT breaking the guard.
Never modify guard/test files — adapt the implementation instead.
If still failing → log "discard (guard failed)" and move on
- SAME/WORSE → Git revert, log "discard"
- CRASHED → Try to fix (max 3 attempts), else log "crash" and move on
8. Log: Record result in results log
9. Repeat: Go to step 1.
- If unbounded: NEVER STOP. NEVER ASK "should I continue?"
- If bounded (N): Stop after N iterations, print final summaryexperiment: prefix. Use git revert (not git reset --hard) for rollbacks so failed experiments remain visible in history. Agent MUST read git log and git diff of kept commits to learn patterns before each iterationSee references/core-principles.md for the 7 generalizable principles from autoresearch.
| Domain | Metric | Scope | Verify Command | Guard |
|---|---|---|---|---|
| Backend code | Tests pass + coverage % | src/**/*.ts | npm test | — |
| Frontend UI | Lighthouse score | src/components/** | npx lighthouse | npm test |
| ML training | val_bpb / loss | train.py | uv run train.py | — |
| Blog/content | Word count + readability | content/*.md | Custom script | — |
| Performance | Benchmark time (ms) | Target files | npm run bench | npm test |
| Refactoring | Tests pass + LOC reduced | Target module | npm test && wc -l | npm run typecheck |
| Security | OWASP + STRIDE coverage + findings | API/auth/middleware | /autoresearch:security | — |
| Shipping | Checklist pass rate (%) | Any artifact | /autoresearch:ship | Domain-specific |
| Debugging | Bugs found + coverage | Target files | /autoresearch:debug | — |
| Fixing | Error count (lower) | Target files | /autoresearch:fix | npm test |
| Scenario analysis | Scenario coverage score (higher) | Feature/domain files | /autoresearch:scenario | — |
| Scenarios | Use cases + edge cases + dimension coverage | Target feature/files | /autoresearch:scenario | — |
| Prediction | Findings + hypotheses (higher) | Target files | /autoresearch:predict | — |
| Documentation | Validation pass rate (higher) | docs/*.md | /autoresearch:learn | npm test |
Adapt the loop to your domain. The PRINCIPLES are universal; the METRICS are domain-specific.
© OpenLAIR, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in skills/autoresearch of OpenLAIR/dr-claw.
Open the folder on GitHubat commit d51b64e
Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Autoresearch this skillOpenLAIR/dr-claw | 1.2k | — | ~7.5k | Automated safety check: Pass | MIT | |
| Show Me Your Work Decision Logcursor/plugins | 11k | 8 repos | ~1.6k | Automated safety check: Pass | None | |
| Autoresearch Iteration Loopuditgoenka/autoresearch | 6.5k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Install Loop Engineeringcobusgreyling/loop-engineering | 11k | 1 repos | ~648 | Automated safety check: Pass | MIT | |
| LoopyForward-Future/loopy | 3.2k | — | ~3.9k | Automated safety check: Pass | MIT | |
| AI Performance Improvement Plantanweai/pua | 20k | 2 repos | ~6.9k | Automated safety check: Pass | MIT |
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
uditgoenka/autoresearch
Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.
cobusgreyling/loop-engineering
Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.
Forward-Future/loopy
Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.
tanweai/pua
Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.
loopx-project/loopx
Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.
OpenLAIR/dr-claw
Analyzes reviewer comments and drafts venue-specific rebuttals for AI and computer science conferences, with an issue board, task list and paper edit plan.
OpenLAIR/dr-claw
Turns a research paper into a slide deck and, optionally, a narrated demo video, through script, slide generation, text-to-speech and video assembly stages you control.
OpenLAIR/dr-claw
Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.
OpenLAIR/dr-claw
Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.
OpenLAIR/dr-claw
Runs multi-source web research through Google's Gemini Deep Research Agent with a bundled Python script and saves a structured, cited report as files.
OpenLAIR/dr-claw
Six-phase workflow for writing, revising and adapting grant proposals for NSF, NIH, DOE, DARPA, NASA and China's NSFC, from profiling through simulated peer review.
Categories
Autonomous Goal-directed Iteration. An agent skill from OpenLAIR/dr-claw. Autoresearch is an agent skill from OpenLAIR/dr-claw. Autonomous Goal-directed Iteration.
Autoresearch fits situations like: tasks that involve Autonomous loops.
Run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a claude-code`. Or copy the skill folder (skills/autoresearch in OpenLAIR/dr-claw) into .claude/skills/autoresearch in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a codex`. Or copy the skill folder (skills/autoresearch in OpenLAIR/dr-claw) into .agents/skills/autoresearch in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenLAIR/dr-claw --skill autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch, .gemini/skills/autoresearch, .github/skills/autoresearch and .opencode/skills/autoresearch in your project.
Going by SKILL.md and its folder, Autoresearch needs the command-line tools its instructions call (npm, git, gh, kubectl, npx and uv).
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Autoresearch is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 56k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Autoresearch: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OpenLAIR (a GitHub organization) maintains it in OpenLAIR/dr-claw, which has 1,155 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 17, 2026.
Source: OpenLAIR/dr-claw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.