Show Me Your Work Decision Log
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/dueling-autoresearch .claude/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .claude/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/loops/dueling-autoresearch .agents/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .agents/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/loops/dueling-autoresearch .cursor/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .cursor/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/gaasher/Agent-Loop-Skills.git --path loops/dueling-autoresearch--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/loops/dueling-autoresearch .gemini/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .gemini/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/loops/dueling-autoresearch .github/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .github/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install gaasher/Agent-Loop-Skills dueling-autoresearch --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/loops/dueling-autoresearch .opencode/skills/dueling-autoresearch && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dueling-autoresearch" agent skill from https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/dueling-autoresearch into .opencode/skills/dueling-autoresearch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dueling-autoresearch", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dueling-autoresearchA skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.
Dueling Autoresearch is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants two approaches raced head-to-head on a single shared metric — e.g. a classical/algorithmic lane vs an ML/learned lane, or any two strategies for the same task. Each lane runs its own analysis-first research loop confined to its lane, the lanes share a scoreboard and may borrow ideas across the boundary without abandoning their identity, and a shared eval keeps the head-to-head honest; loops until interrupted, reporting the current leader. Not for improving a single approach in isolation…
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/run.example.yaml` and `roles/TrackAgent.md`). Compatibility notes: Requires Python 3.9+
It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.
Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Python 3.9+
From compatibility in the SKILL.md frontmatter.
Dueling Autoresearch loads about 2.6k tokens when it runs. Until then it costs about 167 tokens; SKILL.md has 1,182 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
honest. Do not pause for permission once the loop is running.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,182 words, ~2,557 tokens.
.claude/skills/dueling-autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Two lanes work the same objective in parallel and race the same metric — by default a
classical/algorithmic lane against an ML/learned lane (the lanes are user-named). Each lane
runs its own analysis-first iteration via roles/TrackAgent.md, confined to its lane. Every
round both lanes post to a shared duel_log.md scoreboard and may borrow ideas across the lane
boundary — but each stays in its lane. The feedback signal is the shared <metric> on a shared
eval: if the classical lane wins, that is a real result. Lanes support mixed code locations — a
codebase lane edits existing repo files, a sandbox lane authors its own code — and an
eval-parity gate keeps the scores comparable.
You are the orchestrator: each round you advance both lanes, update the scoreboard, and keep both honest. Do not pause for permission once the loop is running.
Use this to race two genuinely different approaches on one metric and keep them honest against the same eval — classical vs learned, two model families, two query strategies. Default to spawning both lanes in parallel and letting the scoreboard drive cross-lane idea borrowing; if a lane runs dry, push it to a more radical in-lane change or to borrow a fresh idea from the log. Not for tuning a single approach (use a single-track loop), and not for a one-shot comparison of two finished things.
The cast (all in this folder):
roles/TrackAgent.md — the per-lane researcher, instantiated once per lane.Resolve bindings interactively. If loop.run.yaml exists in the working dir, load it, confirm the
values in one line, and skip to the loop. Otherwise: on Claude Code (the AskUserQuestion tool is
available) infer a likely value for each binding and present it as the recommended option; on other
hosts ask each as a quoted plain-text prompt. Then write loop.run.yaml (format:
examples/run.example.yaml) and confirm the values before creating any other files.
The host also decides spawn-or-degrade: on Claude Code spawn a real Agent per lane so the two
run in parallel; otherwise adopt roles/TrackAgent.md inline and run the lanes sequentially.
Shared bindings (identical for both lanes — the honesty anchor):
| binding | meaning | default | how to infer |
|---|---|---|---|
<metric> | the single metric both lanes race; ground truth of the duel | — | ask; scan run logs for a printed score |
<metric_direction> | minimize or maximize | — | from the metric's nature (loss vs accuracy) |
<gate> | run budget unit: time or epochs | epochs | the artifact's runner |
<budget> | epochs per run (or minutes if gate: time) | 5 | — |
<sandbox_root> | where snapshots, ledgers, and the duel log live | ./sandbox | — |
<iter_strategy> | snapshots or branches (snapshots recommended — two lanes on one branch is simplest) | snapshots | — |
Per-lane bindings (two lanes, default names classical and learned). Each lane has a
code_location that decides which other fields it needs — never add code to the codebase:
| field | meaning | when |
|---|---|---|
name | lane name | always |
code_location | codebase or sandbox | always |
run_cmd | existing entrypoint to run, e.g. python train.py | if codebase |
editable_files | existing repo files this lane may edit | if codebase |
entry | command run from inside <sandbox_root>/<lane>/iter<N>/, e.g. python run.py | if sandbox |
codebase — the lane maps to existing code: edits its editable_files and runs run_cmd.
Two codebase lanes must have non-overlapping editable_files.sandbox — no implementation exists and none is added to the repo: the lane authors and runs
its code inside <sandbox_root>/<lane>/iter<N>/ via entry.Typical duel on a repo with one existing model: the learned lane is
codebase(editsmodel.py/config.yaml, runstrain.py); the classical lane issandbox(authors its own code under<sandbox_root>/classical/iter<N>/). Nothing is added to the codebase, yet the classical lane is still built and iterated.
Eval-parity gate (the honesty anchor). Before starting, confirm both lanes report <metric>
on the same held-out set, computed the same way, so the scores are comparable — state how each lane
emits it (e.g. both print <metric>: to their run log). If they don't match, fix it first; the duel
is meaningless otherwise. If gate: time, write a run_with_timeout.sh wrapper per lane
(timeout $(( <budget> * 60 )) <entry-or-run_cmd> "$@").
Initialise the sandbox (after confirmation):
<sandbox_root>/
├── duel_log.md ← shared channel + scoreboard (## Scoreboard, ## Round log; headers only)
├── <laneA>/results.tsv ← lane A ledger, header only
└── <laneB>/results.tsv ← lane B ledger, header onlyEach lane's per-iteration work lives in <sandbox_root>/<lane>/iter<N>/ (analysis/, results/,
the run log). A codebase lane's iter dir also holds code_snapshot/ (the pre-change copy for
revert); a sandbox lane's iter dir holds the lane's actual code for that iteration (a kept
iteration carries forward as the next one's starting point).
Each round advances both lanes by one iteration. On Claude Code, spawn the two TrackAgents in
parallel (one turn, two Agent calls); otherwise run lane A then lane B inline. A track is one
analysis-first iteration confined to its lane — the 8 steps in roles/TrackAgent.md. Round 1 is
each lane's baseline (a codebase lane runs unmodified; a sandbox lane authors its initial
implementation in iter1/). One change per lane per round, so each metric delta is attributable.
Copy this checklist and tick items off each round:
duel_log.md (both lanes' latest posts + the scoreboard).roles/TrackAgent.md) per lane, given its lane
bindings, the shared <metric>/<metric_direction>/<gate>/<budget>, and duel_log.md.duel_log.md (best <metric>, one key
finding, any dead end, one idea the other lane could borrow).## Scoreboard: best <metric> per lane and the current leader (per
<metric_direction>); optionally flag one cross-pollination suggestion for next round.Spawn-or-degrade per lane. Where the host supports it, spawn a real isolated TrackAgent per lane
(Claude Code: an Agent per lane, both launched in one turn for parallelism). Otherwise adopt
roles/TrackAgent.md inline and run the lanes sequentially. Each TrackAgent is confined to its lane
and returns its iteration summary to the orchestrator.
Two ledgers: a per-lane results.tsv for each lane's experiments, and the shared
duel_log.md scoreboard + round posts.
Per-lane <sandbox_root>/<lane>/results.tsv (tab-separated, never commas in free text):
iter <metric> status analysis_summary description
1 0.6320 keep baseline; classical features, logistic head baseline
2 0.6610 keep added HOG features; per-class gains on textured classes add HOG feature extractorstatus ∈ {keep, discard, crash} (0.000000 for <metric> on crash).
Shared <sandbox_root>/duel_log.md — scoreboard + per-round posts:
## Scoreboard
round classical_best learned_best leader
1 0.6320 0.6480 learned
2 0.6610 0.7050 learned
## Round log
### Round 2
- **classical** — best 0.6610 (this iter 0.6610). Finding: HOG helps textured classes
(results/per_class.txt). Dead end: raw-pixel kNN plateaus. Borrow: learned's augmentation
could expand classical's training set.
- **learned** — best 0.7050. Finding: BN fixed conv2 saturation. Dead end: dropout hurt at this
budget. Borrow: classical's HOG features as an aux input channel.Report the current leader (per <metric_direction>), never a final winner — a lane that is
behind can still come back. Leave results.tsv, duel_log.md, and iter*/ untracked (do not
commit them).
codebase lane edits only its own <editable_files>; a
sandbox lane lives entirely in <sandbox_root>/<lane>/. Lanes never touch each other's files —
every other file is the evaluation ground truth.<metric> on the same held-out set computed the
same way; never compare otherwise. Do not modify the evaluation/metric — it is the shared ground
truth that makes the duel honest.tee. The sandbox is self-contained —
no ../ escapes.© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in loops/dueling-autoresearch of gaasher/Agent-Loop-Skills.
Open the folder on GitHubat commit f1169e6
Dueling Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dueling Autoresearch this skillgaasher/Agent-Loop-Skills | 174 | — | ~2.6k | Automated safety check: Warn | MIT | |
| Show Me Your Work Decision Logcursor/plugins | 11k | 8 repos | ~1.6k | Automated safety check: Pass | None | |
| Autoresearch Iteration Loopuditgoenka/autoresearch | 6.5k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Install Loop Engineeringcobusgreyling/loop-engineering | 11k | 1 repos | ~648 | Automated safety check: Pass | MIT | |
| LoopyForward-Future/loopy | 3.2k | — | ~3.9k | Automated safety check: Pass | MIT | |
| AI Performance Improvement Plantanweai/pua | 20k | 2 repos | ~6.9k | Automated safety check: Pass | MIT |
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
uditgoenka/autoresearch
Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.
cobusgreyling/loop-engineering
Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.
Forward-Future/loopy
Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.
tanweai/pua
Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.
loopx-project/loopx
Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…
gaasher/Agent-Loop-Skills
A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…
gaasher/Agent-Loop-Skills
A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.
Categories
A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g. Dueling Autoresearch is an agent skill from gaasher/Agent-Loop-Skills.g.
Dueling Autoresearch fits situations like: the user wants two approaches raced head-to-head on a single shared metric — e.g; tasks that involve Autonomous loops.
Run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a claude-code`. Or copy the skill folder (loops/dueling-autoresearch in gaasher/Agent-Loop-Skills) into .claude/skills/dueling-autoresearch in your project. Claude Code loads it when a task matches its description.
Run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a codex`. Or copy the skill folder (loops/dueling-autoresearch in gaasher/Agent-Loop-Skills) into .agents/skills/dueling-autoresearch in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill dueling-autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dueling-autoresearch, .gemini/skills/dueling-autoresearch, .github/skills/dueling-autoresearch and .opencode/skills/dueling-autoresearch in your project.
Going by SKILL.md and its folder, Dueling Autoresearch needs the command-line tools its instructions call (python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Dueling Autoresearch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dueling Autoresearch: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on June 30, 2026.
Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.