---
name: benchmaker
description: Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.
disable-model-invocation: true
---

Apply [library context](../../references/library-context.md). Accept a target: path, callable, command, endpoint, native workflow, skill or capability description. Also accept an optional claim, comparisons, budget and output location. Infer ordinary choices; clarify only material claim, permission or budget decisions. Without an executable target, a named representative stands in for it as the [quality card](../../references/quality-card.md) defines; a description with no executable representative supports only a draft.

A benchmark samples realistic, hard, valuable work and separates systems that differ in the claimed ability. Most candidate tasks should fail admission. A smaller budget admits fewer tasks and reaches a lower stage; it never weakens tasks or skips admission checks. Spend follows earned evidence: a task version that fails a criterion gets no further checks, so costly judgments go to tasks that cheaper evidence has not disqualified.

The coordinator owns every assignment and every target, comparison, known-order check, calibration and adversary launch. Other agents return results and execution requests without delegating. Composing targets run in separate top-level sessions under core `docs/hosts.md#workflow-trials`. Target, comparison, known-order check and calibration attempts run through the package's runner, which enforces their caps; deadlines given to native helpers are requests, so record overruns.

Before launching, declare and show planned research workers, candidates, per-task revision limit, repeats, calibration and known-order check launches for admission and measurement, concurrency, authoring and attempt time, deadlines and estimated spend. When caller limits cannot fit building or solving work at the difficulty target, say so before launching. A wall-clock limit covers the whole build: plan when measurement, review and delivery must start for each to finish inside it. When earlier work runs long or stops converging, reduce scope to fewer tasks or families, or a lower stage, so later steps still run. Caller limits on launches or sessions count every launch they name. When a limit names only target runs, report authoring, auditing, adversary, calibration and known-order check launches separately. Record measurement conditions separately from caller limits; failed attempts consume execution limits. Record each launch and each task's admission state durably before dependent work, so an interrupted build resumes without relaunching completed work and can still deliver a draft.

1. **Claim.** Inspect the target, representative user work, [research lessons](../../references/research.md) and relevant primary precedents. State the solver's goal: the work it does, for whom, and what success looks like. Record the claim entries of the [quality card](../../references/quality-card.md), including the systems the benchmark must measure; ask the user when neither the request nor the target settles them. Choose the comparison set, known-order check, calibration systems and, without an executable target, its representative under benchmarking guidance and the card.
2. **Research the work.** Fan out web research guided by the solver's goal: where people do this work, what real instances look like, how it goes wrong, what makes real cases hard, and the most complex real work people attempt, including work beyond today's strongest systems. Give each fresh research worker a distinct scope with research guidance; workers return sources and findings, not tasks. Gather them into a catalog of realistic scenarios and hard cases, recording sources, capture dates and reuse constraints. Include the real material that could become tasks, such as repositories, issues with their fixes, datasets, incident reports and practitioner accounts. Report uncovered ground as a gap. Without web access, research what the caller supplies and record the gap.
3. **Source.** Apply common and Make benchmarking guidance. Draw more candidate work than the suite needs from the catalog and the target's own failures: real work first, then reconstructions from real material, and synthetic work only as a labeled fallback. Each candidate states its value, why it is hard for the claimed ability and an estimated expert time. Spread candidates across families, independent source groups and the catalog's range of difficulty, including its hardest real work, before expanding any one.
4. **Build and admit.** Authors build each candidate under the [benchmark contract](../../references/benchmark-contract.md): public instruction and interface, environment, reference solution and verifier. Freeze each built task, then gather its admission evidence:
   - run the reference and trivial attempts through the verifier;
   - a fresh reviewer who did not build the task, holding the research catalog, judges whether each reconstructed or synthetic task could occur in practice;
   - a fresh auditor with no authoring history or evaluator material solves from public inputs. The coordinator keeps a copy of each saved outcome; an auditor covering several tasks saves all of them before any disclosure. It then receives the evaluator material and reports unstated requirements, rejected valid outcomes and accepted wrong ones; where the host cannot continue it, a fresh reviewer given the saved outcome does this, recorded as a condition;
   - run labeled outcomes through the verifier: an independent valid alternative, which may be the auditor's when valid, and plausible wrong outcomes, including broken copies of the reference;
   - a fresh adversary with the solver's material and access plus the exploit catalogue in research, and never evaluator material, tries to earn credit without doing the work, and tries again after each repair its findings prompt;
   - the calibration systems and the known-order check attempt the task until more attempts could not change its placement. Measurement waits until every admitted task has this evidence.

   Auditors and adversaries may cover several tasks they did not author. When the suite holds many short items, audit, adversary and realism judgments may cover each family through a recorded sample. Revise or reject by the admission criteria; a revision reruns the checks it affects and counts toward the per-task limit.
5. **Measure.** Freeze the admitted suite. Run the comparison set and the known-order check with predeclared repeats, measuring the actual target or its representative. Compute the card. Classify failures from transcripts as ability, task defect, grading, infrastructure, refusal, cut-off or unknown, and audit passing transcripts, or a recorded sample of them, for unearned credit.
6. **Review and repair.** Apply `shared:review-revise-once` with benchmarking and task guidance, using a fresh reviewer who neither authored, audited nor reviewed the realism of tasks. Its inputs are the frozen suite, card, rejection log, admission evidence and a transcript sample that includes successes. Scope repairs and checks to the package. Review repairs count toward each task's revision limit. Changed tasks receive new identities and rerun affected admission and measurement within remaining allowance. A repaired task that fails re-admission leaves the suite, which then receives a new identity, or stays draft.
7. **Deliver.** Return the suite, commands, card, research catalog, rejection log, per-family results with denominators, measured time and cost, requested and achieved stage, and gaps. State whether measured headroom and separation support the claim. Missing required execution, judgment or review, or an unresolved validity defect, leaves affected claims draft. An interrupted or limit-stopped build delivers these items as a draft from its records. Name the next unmet stage.
