Agent skill

Compare Runs

by adrianco in adrianco/retort

Compare evaluated runs in a retort experiment along factor dimensions.

Apache-2.0Auto-check passedData & Analytics

Install Compare Runs

skills CLI
$ npx skills add adrianco/retort --skill compare-runs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adrianco/retort compare-runs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/compare-runs .claude/skills/compare-runs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compare-runs
GitHub stars
207
Token cost
~2.2k tokens
SKILL.md length
762 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compare evaluated runs in a retort experiment along factor dimensions.

  • Works in 7 steps: Discover evaluated runs → Load findings and metrics → Aggregate per cell → …
  • Tasks that involve Statistics
  • SKILL.md covers Overview, Parameters, Steps and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Compare Runs is an agent skill from adrianco/retort. Compare evaluated runs in a retort experiment along factor dimensions. Surfaces effects of each factor, aggregates across replicates, and highlights cells that diverge qualitatively — complementing (not replacing) retort's ANOVA analysis.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Statistics. The repository describes itself as: Platform Evolution Engine. Distill the best from the combinatorial mess. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Statistics

Example prompts

  • “/compare-runs”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Discover evaluated runs
  2. Load findings and metrics
  3. Aggregate per cell
  4. Surface factor effects
  5. Identify qualitative divergence
  6. Surface shared issues
  7. Write the report

What it can do on your machine

Read from SKILL.md and the folder at commit 1f75769. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compare Runs loads about 2.2k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 762 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adrianco/retort at commit 1f75769, republished under its Apache-2.0 licence (© adrianco). 762 words, ~2,161 tokens.

Download SKILL.mdSave it as .claude/skills/compare-runs/SKILL.md (or your agent's skills folder).
name
compare-runs
description
Compare evaluated runs in a retort experiment along factor dimensions. Surfaces effects of each factor, aggregates across replicates, and highlights cells that diverge qualitatively — complementing (not replacing) retort's ANOVA analysis.
type
anthropic-skill
version
1.0

Compare Runs

Overview

Retort is a Design of Experiments engine — every run sits at a point in the factor space, and replicates exist to measure noise. Unlike pourpoise's compare-attempts (which maintains a ranked leaderboard of ad-hoc submissions), comparing retort runs is about factor effects and qualitative divergence, not "who won".

This skill is the qualitative counterpart to retort analyze / retort report effects. The ANOVA tells you whether a factor is significant. This skill tells you what changed in the generated code as you varied it.

Parameters

  • experiment_dir (required): e.g. experiment-1/. Must contain retort.db and a runs/ directory of per-run archives, each already evaluated by evaluate-run (i.e. each rep<N>/ has an evaluation.md and findings.jsonl).
  • output_file (optional, default: {experiment_dir}/reports/comparison.md): Where to write the comparison report.
  • group_by (optional, default: all factors): Comma-separated factors whose effect you want to isolate. The other factors are aggregated over.
  • include_failed (optional, default: true): Whether to include failed runs in the comparison (they're still informative — failure itself is a finding).

Steps

1. Discover evaluated runs
bash
find {experiment_dir}/runs -name evaluation.md -not -path "*/salvaged-*/*" | sort

Parse each path to extract the cell (from the parent directory name, e.g. language=rust_model=opus_tooling=beads) and replicate (from the rep<N> or rep<N>-failed segment).

Constraints:

  • You MUST skip runs under salvaged-*/ — those are manually-archived artifacts from before auto-archival existed, and their layout is not regular.
  • You MUST handle both rep<N>/ (success) and rep<N>-failed/ (failure) directory names.
  • You SHOULD warn but not fail if some cells have no evaluation (evaluator hadn't run yet).
2. Load findings and metrics

For each evaluated run, load:

  • Factors (from the cell directory name or the run's stack.json)
  • Summary metrics (from evaluation.md header: requirements pass count, test counts, build status)
  • Findings (from findings.jsonl) — categorize by kind and severity
  • Architecture signal (from summary/index.md: the "Shape" one-liner)

Constraints:

  • You MUST NOT re-evaluate. If evaluation.md is stale relative to the source, add a stale tag but use what's there.
  • You MUST preserve replicates separately — do not average them into the cell yet.
3. Aggregate per cell

Group runs by cell (same factor combination across replicates):

CellReplicatesReq pass (mean ± sd)Test pass (mean ± sd)Build failuresDistinct "shapes"
lang=python,model=opus,tooling=none3/38.7±0.6 / 1012.3±0.5 / 1201 (all Flask+SQLite)
lang=python,model=opus,tooling=beads2/36.0±1.4 / 1010.0±1.0 / 1212 (Flask vs FastAPI)
..................

Constraints:

  • You MUST show mean ± sample standard deviation, not just mean. Variance across replicates is a first-class signal.
  • "Distinct shapes" counts unique architecture summaries across replicates — a cell where the agent chose a different framework each time is a signal worth surfacing.
  • Cells with only one replicate: drop the ±sd and mark as n=1.
4. Surface factor effects

For each factor in group_by, hold the others roughly fixed and report the effect:

markdown
## Effect of `tooling` (none vs beads)

Aggregating over language and model:

| Tooling | Mean req pass | Mean token cost | Mean build time | Build fail rate |
|---------|---------------|-----------------|-----------------|-----------------|
| none    | 8.1 / 10      | 289K tokens     | 124s            | 1/18 (6%)       |
| beads   | 7.2 / 10      | 441K tokens     | 156s            | 2/18 (11%)      |

Direction: beads costs ~52% more tokens and scored 11% lower on requirements.
This matches the p < 0.10 effect reported by `retort analyze`.

Qualitative: with `tooling: beads`, agents spent ~30% of their turns
on bd bookkeeping (visible in run transcripts). This accounts for the
token overhead.

Constraints:

  • You MUST NOT claim statistical significance yourself — cite retort analyze or note "directional only".
  • You MUST cross-reference the structured effect size reported by retort report effects when available at {experiment_dir}/reports/analysis/*.md.
  • Qualitative claims MUST cite specific run IDs or findings (see rep3 of lang=python,model=opus,tooling=beads).
Show full SKILL.md (280 more words)Show less
5. Identify qualitative divergence

Find cells where runs within a replicate group diverged in a way pure metrics miss:

markdown
## Qualitative divergence

### lang=typescript, model=sonnet, tooling=beads — 3 replicates, 3 different shapes

| Rep | Framework | Storage | Notes |
|-----|-----------|---------|-------|
| 1 | Express + better-sqlite3 | SQLite | Built, tests failed on migration |
| 2 | Fastify + Prisma | Postgres (!) | Didn't build — assumed a Postgres container |
| 3 | Raw http + in-memory | Map object | Simplest, passed all requirements |

Within-cell variance is high: sonnet+TS+beads doesn't converge on a single
architecture. Consider whether this cell needs more replicates or whether
the task spec is under-constrained.

Constraints:

  • You MUST highlight cells where replicates disagree on framework/library/storage, not just scores.
  • You SHOULD identify at most 5 "most divergent" cells — past that it becomes noise.
  • You MUST NOT label divergence as a failure — sometimes it's the signal the experiment was built to find.
6. Surface shared issues

Aggregate findings across all runs. A finding kind that shows up in >50% of runs is a property of the task or the model, not of the individual run.

markdown
## Shared issues (appearing in ≥50% of runs)

| Finding | Runs affected | Severity |
|---------|---------------|----------|
| `skipped_test` | 28/37 | medium |
| `requirement_missing: R5 (pagination)` | 22/37 | high |
| `lint_warning: unused import` | 19/37 | low |

R5 (pagination) is either genuinely hard for agents OR the task spec
didn't emphasize it enough. Worth considering for a task-spec revision.
7. Write the report

Use the Output Format below. The report MUST link to individual evaluation.md files so readers can drill in.

Output Format

markdown
# Comparison: {experiment_dir_name}

Generated {timestamp} from {n} evaluated runs across {m} cells.

## Coverage

- Cells with ≥1 evaluation: {n}/{total}
- Runs evaluated: {n}/{total}
- Failed runs: {n}
- Missing evaluations: {list any}

## Per-cell summary

| Cell | n | Req pass | Test pass | Build fails | Shape diversity | Link |
|------|---|----------|-----------|-------------|-----------------|------|
| ... | 3 | 8.7±0.6 | 12.3±0.5 | 0 | 1 | [evals](runs/...) |

## Factor effects

### `{factor}`
...

## Qualitative divergence

...

## Shared issues

...

## Links

- ANOVA / effect sizes: `reports/analysis/`
- Per-run evaluations: `runs/<cell>/rep<N>/evaluation.md`
- Raw findings: `runs/<cell>/rep<N>/findings.jsonl`

Constraints Summary

  • You MUST read evaluation reports, not re-evaluate. This skill is purely aggregation.
  • You MUST NOT produce a single "winner" ranking. Retort is about factor effects, not tournaments.
  • You MUST preserve replicate-level detail at least down to the per-cell table.
  • You MUST cite specific runs by their archive path when surfacing qualitative claims.
  • You MUST cross-reference retort analyze output when making causal claims about factors.
  • Output MUST be markdown that renders in GitHub's viewer.

Troubleshooting

No evaluations yet

  • Exit 0 with a single-line report: No evaluated runs yet. Run evaluate-run on each rep<N>/ first.
  • Do not attempt to evaluate as a side effect — that's the user's / CLI's decision.

Evaluations with different schema versions

  • Read what you can from each, note the schema mismatch in the "Coverage" section.
  • Do not silently drop older evaluations — surface the problem.

retort analyze reports don't exist

  • Proceed without them. Note in "Factor effects" that statistical backing is absent and the direction is qualitative-only.

© adrianco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/compare-runs of adrianco/retort.

Open the folder on GitHubat commit 1f75769

Compare with similar skills

Compare Runs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compare Runs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compare Runs this skilladrianco/retort207—~2.2kAutomated safety check: PassApache-2.0
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.9k3 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone
Statistical Powerspacering-net/codeg3.9k1 repos~3.6kAutomated safety check: NotesMIT

Similar skills

  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.9k GitHub starsUsed in 3 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.9k GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from adrianco/retort

  • Diagnose Failed Run

    adrianco/retort

    Determine the TRUE cause of a failed retort run before attributing it.

    207 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Evaluate Run

    adrianco/retort

    Evaluate a single retort experiment run. An agent skill from adrianco/retort.

    207 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • File Run Issues

    adrianco/retort

    Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.

    207 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Run Summary

    adrianco/retort

    Summarize the architecture of code generated by a single retort run.

    207 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Update Optimal Blog

    adrianco/retort

    Refresh the data tables in optimal-blog.md from master.db. An agent skill from adrianco/retort.

    207 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Compare Runs

What does Compare Runs do?

Compare evaluated runs in a retort experiment along factor dimensions. Compare Runs is an agent skill from adrianco/retort. Compare evaluated runs in a retort experiment along factor dimensions.

When should I use Compare Runs?

Compare Runs fits situations like: tasks that involve Statistics.

How do I install Compare Runs in Claude Code?

Run `npx skills add adrianco/retort --skill compare-runs -a claude-code`. Or copy the skill folder (skills/compare-runs in adrianco/retort) into .claude/skills/compare-runs in your project. Claude Code loads it when a task matches its description.

How do I install Compare Runs in Codex?

Run `npx skills add adrianco/retort --skill compare-runs -a codex`. Or copy the skill folder (skills/compare-runs in adrianco/retort) into .agents/skills/compare-runs in your project. Codex loads it when a task matches its description.

Can I use Compare Runs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adrianco/retort --skill compare-runs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compare-runs, .gemini/skills/compare-runs, .github/skills/compare-runs and .opencode/skills/compare-runs in your project.

What does Compare Runs need to run?

SKILL.md names no scripts, command-line tools or credentials: Compare Runs is instructions for the agent only.

Does Compare Runs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compare Runs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Compare Runs use?

Compare Runs is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compare Runs use?

About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compare Runs?

Skills that share tags, products or a category with Compare Runs: Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.9k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and AI Daily Digest (vigorX777/ai-daily-digest, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compare Runs?

adrianco (a GitHub user) maintains it in adrianco/retort, which has 207 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 9, 2026.

Source: adrianco/retort on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.