Loop Factory
JuliusBrussee/skills
Run a spec-driven agent loop where coding tasks live as markdown specs that move through inbox → active → archive, get implemented by Claude Code or Codex, and pass a review gate before they count…
Domain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch.
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jdrhyne/agent-skills autoresearch-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autoresearch-loop .claude/skills/autoresearch-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .claude/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jdrhyne/agent-skills autoresearch-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/autoresearch-loop .agents/skills/autoresearch-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .agents/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jdrhyne/agent-skills autoresearch-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/autoresearch-loop .cursor/skills/autoresearch-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .cursor/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jdrhyne/agent-skills.git --path skills/autoresearch-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jdrhyne/agent-skills autoresearch-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/autoresearch-loop .gemini/skills/autoresearch-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .gemini/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jdrhyne/agent-skills autoresearch-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/autoresearch-loop .github/skills/autoresearch-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .github/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jdrhyne/agent-skills autoresearch-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/autoresearch-loop .opencode/skills/autoresearch-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "autoresearch-loop" agent skill from https://github.com/jdrhyne/agent-skills/tree/main/skills/autoresearch-loop into .opencode/skills/autoresearch-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "autoresearch-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
autoresearch-loopDomain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch.
Autoresearch Loop is an agent skill from jdrhyne/agent-skills. Domain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch. Use when you want an agent to discover what to measure for a project/goal, then run a keep-or-revert experiment loop that proposes changes, measures them against an objective, keeps wins, discards regressions, and records implemented improvements. Adapts to code perf-auditing, codegen, bug-finding, ad optimization, or any artifact + measurable objective + trial. Trigger: 'autoresearch this', 'find and implement improvements to…
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 28 other files, including scripts and reference files (for example `DESIGN.md`, `README.md` and `adapters/bug-finding.md`).
It sits in Agent Workflows, covering Autonomous loops and Project scaffolding. The repository describes itself as: A collection of AI agent skills for Clawdbot, Claude Code, Codex. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 439cd3a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (JavaScript and Shell, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Autoresearch Loop loads about 2k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 991 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jdrhyne/agent-skills at commit 439cd3a, republished under its MIT licence (© jdrhyne). 991 words, ~1,966 tokens.
.claude/skills/autoresearch-loop/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.Generalize Karpathy's autoresearch into a domain-adaptive improvement loop. The agent discovers what to measure, then runs a disciplined propose → trial → keep-or-revert loop, maintaining an explicit ledger of what was tried, kept, discarded, and implemented.
Read DESIGN.md once at the start of a run for the full architecture and the domain-specific tensions (metric latency/noise/cost, Goodhart gaming, cost-per-trial, reversibility). The phases below are the operating procedure.
Runtime: the loop's mechanics (run a trial, parse the metric, score confidence, keep/commit or discard/revert) are handled by the arl CLI over a .auto/ session folder — a Claude-native port of pi-autoresearch's tools. Read references/runtime-contract.md for the .auto/ layout, the METRIC name=value contract, MAD confidence scoring, and the arl init|run|log|status commands. Invoke it as node scripts/arl.mjs <cmd> (or arl if on PATH).
Given {project, goal, context} — follow the procedure in references/metric-discovery.md (restate the goal as an outcome → enumerate candidates on the proxy→outcome spectrum → score on six axes → choose primary + guardrails + strategy → red-team for gaming). In brief:
adapters/*.md). If none fits, draft an inline adapter following references/domain-adapter-contract.md.deterministic-delta (fast, low-noise), significance-test, or bandit (noisy/delayed/expensive — e.g. ads).runs/<id>/CHARTER.md. Confirm it with the user before spending real budget if trials cost money or touch production.Write .auto/measure.sh (and .auto/checks.sh if guardrails require it). arl init with the primary metric + direction, run the baseline (arl run), and record it (arl log --status keep --metric <baseline> --desc baseline). If the baseline can't be measured cleanly and repeatably, stop (see "When NOT to run").
Before proposing any change, profile where the cost actually is, and confirm the benchmark stresses the IN-SCOPE artifact — not a dependency, a native/FFI call, an external engine, the network, or unrelated code. (Validated the hard way on two live runs: once the assumed hot path was wrong twice and 97% of time was in an out-of-scope library; once the in-scope managed code was only 0.4–3.8% of wall-time because a Rust NIF dominated — the correct loop output there was a true negative, "re-scope," not a sub-noise edit. See adapters/code-perf-audit.md → Pitfalls.) Spend one profiling run on the managed-vs-native/dependency split; a loop that optimizes code which isn't the bottleneck produces confident, useless churn — and proving "no in-scope headroom" cheaply is itself a successful outcome.
Each iteration:
.auto/prompt.md, the .auto/log.jsonl tail, and .auto/ideas.md — never re-propose an exhausted line..auto/ideas.md; append newly-imagined ideas there..auto/ is preserved).arl run — runs the trial harness, parses METRIC lines, runs guardrail checks.sh.fast-low-noise domains, watch the MAD confidence score (<1.0× = within noise, re-run before trusting). For delayed-expensive domains (ads), do NOT trust a single trial — reach the adapter's minimum sample and use significance/bandit logic. A primary win that regresses any guardrail is a discard.arl log --status keep|discard|... --metric <value> --desc "..." --asi <learning>. Keep auto-commits and advances the baseline; discard auto-reverts the code. Record cost with --cost.--asi with what was learned (survives a discard's revert); prune exhausted ideas.Respect concurrency reality: deterministic domains can run many fast sequential trials; noisy/delayed domains (ads) run few long concurrent trials and must reach a minimum sample before any verdict.
adapters/web-onboarding.md.At any stop, report (arl status summarizes most of it): metric baseline → current with the delta and confidence, the kept improvements (each with its delta and cost), what was tried and discarded with the --asi reasons from .auto/log.jsonl, the remaining promising ideas, and total cost. The durable record is .auto/ plus the git history of kept commits.
DESIGN.md — architecture, per-domain design tensions, and the pi-autoresearch prior-art decision (read once per run).references/metric-discovery.md — the Phase 0 procedure: deriving + scoring + red-teaming the metric.references/runtime-contract.md — the .auto/ layout, METRIC contract, MAD confidence, and arl commands.references/domain-adapter-contract.md — how to define a new domain adapter.references/journal-schema.md — the ledger record formats.adapters/*.md — concrete domain adapters (code-perf-audit, bug-finding, code-generation, google-ads, web-onboarding).© jdrhyne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 24 other files (scripts, references) in skills/autoresearch-loop of jdrhyne/agent-skills.
Open the folder on GitHubat commit 439cd3a
Autoresearch Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Autoresearch Loop this skilljdrhyne/agent-skills | 240 | — | ~2k | Automated safety check: Pass | MIT | |
| Loop FactoryJuliusBrussee/skills | 162 | — | ~2k | Automated safety check: Pass | MIT | |
| Harness Engineeringguanyang/open-agent-hub | 977 | 1 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Self Improvement Loopsguanyang/open-agent-hub | 977 | 1 repos | ~5.6k | Automated safety check: Pass | MIT | |
| Toolifycoreyhaines31/makerskills | 850 | — | ~2.9k | Automated safety check: Notes | MIT | |
| Spec Optimizeleo-kuang-ai/spec-first | 107 | — | ~13k | Automated safety check: Pass | MIT |
JuliusBrussee/skills
Run a spec-driven agent loop where coding tasks live as markdown specs that move through inbox → active → archive, get implemented by Claude Code or Codex, and pass a review gate before they count…
guanyang/open-agent-hub
This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR…
guanyang/open-agent-hub
This skill should be used when the harness, scaffold, workflow, or optimizer itself is the optimization target: recursive self-improvement (RSI) loops, meta-harnesses, self-improving harnesses that…
coreyhaines31/makerskills
When you want to integrate an external tool, API, MCP server, or service into a project — the wizard walks you through auth, config, env vars, client wrapper code, example usage, and an optional…
leo-kuang-ai/spec-first
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
hoodini/ai-agents-skills
Build a new AI agent on AWS and deploy it easily, OR wrap and deploy an agent you already have, using the Amazon Bedrock AgentCore harness.
jdrhyne/agent-skills
Gong API for searching calls, transcripts, and conversation intelligence.
jdrhyne/agent-skills
Generate beautifully designed PDF reports with a Nordic/Scandinavian aesthetic.
jdrhyne/agent-skills
Read Google Analytics 4 reporting data through the GA4 Data API.
jdrhyne/agent-skills
Bulk download images from login-protected gallery websites using an attached browser session.
jdrhyne/agent-skills
Read Google Search Console properties, Search Analytics, URL Inspection, and sitemaps.
jdrhyne/agent-skills
Upload, edit, and export documents via Nudocs.ai. An agent skill from jdrhyne/agent-skills.
Categories
Domain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch. Autoresearch Loop is an agent skill from jdrhyne/agent-skills. Domain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch.
Autoresearch Loop fits situations like: you want an agent to discover what to measure for a project/goal; then run a keep-or-revert experiment loop that proposes changes; measures them against an objective; discards regressions.
Run `npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a claude-code`. Or copy the skill folder (skills/autoresearch-loop in jdrhyne/agent-skills) into .claude/skills/autoresearch-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a codex`. Or copy the skill folder (skills/autoresearch-loop in jdrhyne/agent-skills) into .agents/skills/autoresearch-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jdrhyne/agent-skills --skill autoresearch-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch-loop, .gemini/skills/autoresearch-loop, .github/skills/autoresearch-loop and .opencode/skills/autoresearch-loop in your project.
Going by SKILL.md and its folder, Autoresearch Loop needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Node.js; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Autoresearch Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Autoresearch Loop: Loop Factory (JuliusBrussee/skills, 162 stars), Harness Engineering (guanyang/open-agent-hub, 977 stars), Self Improvement Loops (guanyang/open-agent-hub, 977 stars) and Toolify (coreyhaines31/makerskills, 850 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jdrhyne (a GitHub user) maintains it in jdrhyne/agent-skills, which has 240 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on August 30, 2026.
Source: jdrhyne/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.