Agent skill

Dueling Autoresearch

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

MITAuto-check: warningsAgent Workflows

Install Dueling Autoresearch

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/dueling-autoresearch .claude/skills/dueling-autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dueling-autoresearch
GitHub stars
174
Token cost
~2.6k tokens
SKILL.md length
1,182 words
Files
3
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

  • The user wants two approaches raced head-to-head on a single shared metric — e.g
  • SKILL.md covers When to use, Setup, The loop (duel) and Ledger, plus 1 more section
  • Calls python
  • Tasks that involve Autonomous loops

What it does

Dueling Autoresearch is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants two approaches raced head-to-head on a single shared metric — e.g. a classical/algorithmic lane vs an ML/learned lane, or any two strategies for the same task. Each lane runs its own analysis-first research loop confined to its lane, the lanes share a scoreboard and may borrow ideas across the boundary without abandoning their identity, and a shared eval keeps the head-to-head honest; loops until interrupted, reporting the current leader. Not for improving a single approach in isolation…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/run.example.yaml` and `roles/TrackAgent.md`). Compatibility notes: Requires Python 3.9+

It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • The user wants two approaches raced head-to-head on a single shared metric — e.g
  • Tasks that involve Autonomous loops

Example prompts

  • “/dueling-autoresearch”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+

    From compatibility in the SKILL.md frontmatter.

Context cost

Dueling Autoresearch loads about 2.6k tokens when it runs. Until then it costs about 167 tokens; SKILL.md has 1,182 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:28
    honest. Do not pause for permission once the loop is running.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,182 words, ~2,557 tokens.

Download SKILL.mdSave it as .claude/skills/dueling-autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dueling-autoresearch
description
Use when the user wants two approaches raced head-to-head on a single shared metric — e.g. a classical/algorithmic lane vs an ML/learned lane, or any two strategies for the same task. Each lane runs its own analysis-first research loop confined to its lane, the lanes share a scoreboard and may borrow ideas across the boundary without abandoning their identity, and a shared eval keeps the head-to-head honest; loops until interrupted, reporting the current leader. Not for improving a single approach in isolation (use a single-track research loop), and not for picking between two finished artifacts in one shot (that is a one-time comparison).
compatibility
Requires Python 3.9+
metadata.version
0.1.0

Dueling Autoresearch Loop

Two lanes work the same objective in parallel and race the same metric — by default a classical/algorithmic lane against an ML/learned lane (the lanes are user-named). Each lane runs its own analysis-first iteration via roles/TrackAgent.md, confined to its lane. Every round both lanes post to a shared duel_log.md scoreboard and may borrow ideas across the lane boundary — but each stays in its lane. The feedback signal is the shared <metric> on a shared eval: if the classical lane wins, that is a real result. Lanes support mixed code locations — a codebase lane edits existing repo files, a sandbox lane authors its own code — and an eval-parity gate keeps the scores comparable.

You are the orchestrator: each round you advance both lanes, update the scoreboard, and keep both honest. Do not pause for permission once the loop is running.

When to use

Use this to race two genuinely different approaches on one metric and keep them honest against the same eval — classical vs learned, two model families, two query strategies. Default to spawning both lanes in parallel and letting the scoreboard drive cross-lane idea borrowing; if a lane runs dry, push it to a more radical in-lane change or to borrow a fresh idea from the log. Not for tuning a single approach (use a single-track loop), and not for a one-shot comparison of two finished things.

The cast (all in this folder):

  • roles/TrackAgent.md — the per-lane researcher, instantiated once per lane.

Setup

Resolve bindings interactively. If loop.run.yaml exists in the working dir, load it, confirm the values in one line, and skip to the loop. Otherwise: on Claude Code (the AskUserQuestion tool is available) infer a likely value for each binding and present it as the recommended option; on other hosts ask each as a quoted plain-text prompt. Then write loop.run.yaml (format: examples/run.example.yaml) and confirm the values before creating any other files.

The host also decides spawn-or-degrade: on Claude Code spawn a real Agent per lane so the two run in parallel; otherwise adopt roles/TrackAgent.md inline and run the lanes sequentially.

Shared bindings (identical for both lanes — the honesty anchor):

bindingmeaningdefaulthow to infer
<metric>the single metric both lanes race; ground truth of the duel—ask; scan run logs for a printed score
<metric_direction>minimize or maximize—from the metric's nature (loss vs accuracy)
<gate>run budget unit: time or epochsepochsthe artifact's runner
<budget>epochs per run (or minutes if gate: time)5—
<sandbox_root>where snapshots, ledgers, and the duel log live./sandbox—
<iter_strategy>snapshots or branches (snapshots recommended — two lanes on one branch is simplest)snapshots—

Per-lane bindings (two lanes, default names classical and learned). Each lane has a code_location that decides which other fields it needs — never add code to the codebase:

fieldmeaningwhen
namelane namealways
code_locationcodebase or sandboxalways
run_cmdexisting entrypoint to run, e.g. python train.pyif codebase
editable_filesexisting repo files this lane may editif codebase
entrycommand run from inside <sandbox_root>/<lane>/iter<N>/, e.g. python run.pyif sandbox
  • codebase — the lane maps to existing code: edits its editable_files and runs run_cmd. Two codebase lanes must have non-overlapping editable_files.
  • sandbox — no implementation exists and none is added to the repo: the lane authors and runs its code inside <sandbox_root>/<lane>/iter<N>/ via entry.

Typical duel on a repo with one existing model: the learned lane is codebase (edits model.py/config.yaml, runs train.py); the classical lane is sandbox (authors its own code under <sandbox_root>/classical/iter<N>/). Nothing is added to the codebase, yet the classical lane is still built and iterated.

Eval-parity gate (the honesty anchor). Before starting, confirm both lanes report <metric> on the same held-out set, computed the same way, so the scores are comparable — state how each lane emits it (e.g. both print <metric>: to their run log). If they don't match, fix it first; the duel is meaningless otherwise. If gate: time, write a run_with_timeout.sh wrapper per lane (timeout $(( <budget> * 60 )) <entry-or-run_cmd> "$@").

Initialise the sandbox (after confirmation):

<sandbox_root>/
├── duel_log.md          ← shared channel + scoreboard (## Scoreboard, ## Round log; headers only)
├── <laneA>/results.tsv  ← lane A ledger, header only
└── <laneB>/results.tsv  ← lane B ledger, header only

Each lane's per-iteration work lives in <sandbox_root>/<lane>/iter<N>/ (analysis/, results/, the run log). A codebase lane's iter dir also holds code_snapshot/ (the pre-change copy for revert); a sandbox lane's iter dir holds the lane's actual code for that iteration (a kept iteration carries forward as the next one's starting point).

Show full SKILL.md (473 more words)Show less

The loop (duel)

Each round advances both lanes by one iteration. On Claude Code, spawn the two TrackAgents in parallel (one turn, two Agent calls); otherwise run lane A then lane B inline. A track is one analysis-first iteration confined to its lane — the 8 steps in roles/TrackAgent.md. Round 1 is each lane's baseline (a codebase lane runs unmodified; a sandbox lane authors its initial implementation in iter1/). One change per lane per round, so each metric delta is attributable.

Copy this checklist and tick items off each round:

  • State — note round N; read duel_log.md (both lanes' latest posts + the scoreboard).
  • Advance each lane — run a TrackAgent (roles/TrackAgent.md) per lane, given its lane bindings, the shared <metric>/<metric_direction>/<gate>/<budget>, and duel_log.md.
  • Track posts — each lane appends its round entry to duel_log.md (best <metric>, one key finding, any dead end, one idea the other lane could borrow).
  • Scoreboard — update ## Scoreboard: best <metric> per lane and the current leader (per <metric_direction>); optionally flag one cross-pollination suggestion for next round.
  • Continue — go to the next round; never pause to ask whether to continue.

Spawn-or-degrade per lane. Where the host supports it, spawn a real isolated TrackAgent per lane (Claude Code: an Agent per lane, both launched in one turn for parallelism). Otherwise adopt roles/TrackAgent.md inline and run the lanes sequentially. Each TrackAgent is confined to its lane and returns its iteration summary to the orchestrator.

Ledger

Two ledgers: a per-lane results.tsv for each lane's experiments, and the shared duel_log.md scoreboard + round posts.

Per-lane <sandbox_root>/<lane>/results.tsv (tab-separated, never commas in free text):

iter	<metric>	status	analysis_summary	description
1	0.6320	keep	baseline; classical features, logistic head	baseline
2	0.6610	keep	added HOG features; per-class gains on textured classes	add HOG feature extractor

status ∈ {keep, discard, crash} (0.000000 for <metric> on crash).

Shared <sandbox_root>/duel_log.md — scoreboard + per-round posts:

## Scoreboard
round	classical_best	learned_best	leader
1	0.6320	0.6480	learned
2	0.6610	0.7050	learned

## Round log
### Round 2
- **classical** — best 0.6610 (this iter 0.6610). Finding: HOG helps textured classes
  (results/per_class.txt). Dead end: raw-pixel kNN plateaus. Borrow: learned's augmentation
  could expand classical's training set.
- **learned** — best 0.7050. Finding: BN fixed conv2 saturation. Dead end: dropout hurt at this
  budget. Borrow: classical's HOG features as an aux input channel.

Report the current leader (per <metric_direction>), never a final winner — a lane that is behind can still come back. Leave results.tsv, duel_log.md, and iter*/ untracked (do not commit them).

Constraints

  • Never add code to the codebase. A codebase lane edits only its own <editable_files>; a sandbox lane lives entirely in <sandbox_root>/<lane>/. Lanes never touch each other's files — every other file is the evaluation ground truth.
  • Stay in lane. A lane borrows ideas, never converts into the other approach — a classical lane stays classical even if it borrows a loss/target idea from the learned lane.
  • Same metric, same eval. Both lanes optimize <metric> on the same held-out set computed the same way; never compare otherwise. Do not modify the evaluation/metric — it is the shared ground truth that makes the duel honest.
  • One change per lane per round, so each metric delta is attributable.
  • Do not install new packages or add dependencies the project lacks; helper code stays stdlib-only.
  • Redirect each run's output to its lane's run log; never use tee. The sandbox is self-contained — no ../ escapes.
  • Do not pause the loop to ask for direction; once running, keep both lanes iterating until manually interrupted, and report the current leader rather than declaring a final winner.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in loops/dueling-autoresearch of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/run.example.yaml
  • roles/TrackAgent.md

Open the folder on GitHubat commit f1169e6

Compare with similar skills

Dueling Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dueling Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dueling Autoresearch this skillgaasher/Agent-Loop-Skills174—~2.6kAutomated safety check: WarnMIT
Show Me Your Work Decision Logcursor/plugins11k8 repos~1.6kAutomated safety check: PassNone
Autoresearch Iteration Loopuditgoenka/autoresearch6.5k1 repos~2kAutomated safety check: PassMIT
Install Loop Engineeringcobusgreyling/loop-engineering11k1 repos~648Automated safety check: PassMIT
LoopyForward-Future/loopy3.2k—~3.9kAutomated safety check: PassMIT
AI Performance Improvement Plantanweai/pua20k2 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    11k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Install Loop Engineering

    cobusgreyling/loop-engineering

    Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.

    11k GitHub starsUsed in 1 repo~648 tokens
    Agent WorkflowsAuto-check passed
  • Loopy

    Forward-Future/loopy

    Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.

    3.2k GitHub stars~3.9k tokensUpdated 29 days ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.

    20k GitHub starsUsed in 2 repos~6.9k tokens
    Agent WorkflowsAuto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from gaasher/Agent-Loop-Skills

All 20 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub stars~3.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Anomaly Investigation

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

    174 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Data Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted.

    174 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Hypothesis Gen

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Dueling Autoresearch

What does Dueling Autoresearch do?

A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g. Dueling Autoresearch is an agent skill from gaasher/Agent-Loop-Skills.g.

When should I use Dueling Autoresearch?

Dueling Autoresearch fits situations like: the user wants two approaches raced head-to-head on a single shared metric — e.g; tasks that involve Autonomous loops.

How do I install Dueling Autoresearch in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a claude-code`. Or copy the skill folder (loops/dueling-autoresearch in gaasher/Agent-Loop-Skills) into .claude/skills/dueling-autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Dueling Autoresearch in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a codex`. Or copy the skill folder (loops/dueling-autoresearch in gaasher/Agent-Loop-Skills) into .agents/skills/dueling-autoresearch in your project. Codex loads it when a task matches its description.

Can I use Dueling Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dueling-autoresearch, .gemini/skills/dueling-autoresearch, .github/skills/dueling-autoresearch and .opencode/skills/dueling-autoresearch in your project.

What does Dueling Autoresearch need to run?

Going by SKILL.md and its folder, Dueling Autoresearch needs the command-line tools its instructions call (python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+.

Does Dueling Autoresearch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dueling Autoresearch safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Dueling Autoresearch use?

Dueling Autoresearch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dueling Autoresearch use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dueling Autoresearch?

Skills that share tags, products or a category with Dueling Autoresearch: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dueling Autoresearch?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.