---
name: eval-loop
description: 'Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. You are the conductor: import cases into the file ledger, spawn a blind answerer, then run eval-answer, eval-diagnose, and eval-improve. Persistence is plain files: the set in the model package''s evals/ directory, runs in the set''s workdir; checkpoints are git commits of the model repo. `eval.py` runs each step from the set''s eval.toml. Use to score a model, diagnose failures, improve behind an acceptance check, or roll back a bad direction.'
---

# The Evaluation Loop

You conduct this loop. There is no batch orchestrator to start, no eval API,
and no eval MCP tools. The ledger is plain files: the set in the model
package's git repository, and each run in the set's workdir
(`reference/ledger-schema.md` in `skill:eval-answer` defines every file and
event). `scripts/eval.py` runs each step from the set's `eval.toml`;
`reference/running-a-run.md` has the commands. Scoring is an LLM judge you spawn per case. There is no
scripted scorer, and there will not be one: a script that can pass a wrong
answer is worse than none. The scripts under `scripts/` run the loop -- they
answer, re-execute, spawn the judge, compare runs, and write the ledger -- but
none of them decides whether an answer was right.

```
scrape/run  ->  eval  ->  diagnose  ->  improve  ->  checkpoint
```

**This skill conducts; it does not restate.** Scoring lives in
`skill:eval-answer`. Components and owners live in `skill:eval-diagnose`.
Edit rules live in `skill:eval-improve`.

Do not merge **eval** into **diagnose**. A conductor who scores while
explaining writes the explanation into the score. Do not skip the **acceptance
check** inside improve. The acceptance check decides whether *this* edit
stays. **Checkpoint** decides whether a *sequence* of accepted edits can be
undone.

## Where the rest of this lives

This file is the procedure. The things it used to carry inline are files beside
it now, because each is needed at one moment rather than every run, and loading
all of them for every run is how a skill stops being read.

| When | Read |
|---|---|
| the set does not exist yet, or has never run | `reference/setting-up-a-set.md` |
| about to run one | `reference/running-a-run.md` |
| a golden is wrong, doubted, or out of step with the model | `reference/golden-side-door.md` |
| auditing a key you doubt, or a set you did not author | `reference/auditing-an-answer-key.md` |
| deciding whether an edit stays | `reference/acceptance-check.md` |
| about to quote a number, set the band, or read a set's flips | `reference/measurement.md` |
| the run finished and someone has to read it | `skill:eval-report` |
| you changed judge doctrine or its inputs | `reference/checking-the-judge.md` |

Read the file, do not work from the summary here. The acceptance-check rules and
the golden side door are both places where acting on a half-memory of the rule
produces a confident wrong answer rather than an error.

## The five steps

| Step | Job | Writes |
|---|---|---|
| **a. scrape / run** | Put cases in the ledger; spawn a blind answerer | cases; `attempt`, `tool_call` |
| **b. eval** | Judge the answer; score which required entities retrieval delivered | `score` |
| **c. diagnose** | Why it failed, who owns it | `issue` / `issue_status`. Stop. Do not edit. |
| **d. improve** | One smallest model edit, then the acceptance check | improve writes `candidate`; you write `acceptance_check`. Revert on reject. |
| **e. checkpoint** | Git commit after an accepted acceptance check | `checkpoint` event, then the commit |

Every run then **reports** (below). A run whose result exists only as JSONL and
a scrolled-away console summary has not been delivered.

**scrape** and **run** share a letter but are not the same job. Scrape writes
cases. Run writes attempts. Do not invent questions and score
them in one breath.

**Step d is `skill:eval-improve`, not a hand edit plus another arm.** The
failure mode is specific and it has happened: the edits were made by hand,
published, and validated by re-running the whole arm -- with the answer key
repaired in the same window. No `candidate`, `acceptance_check` or `checkpoint`
event exists in that run's ledger, and the resulting move cannot be attributed
to the model, the key, or the answerer. Re-running a whole arm after changing
two things is the one thing the acceptance check exists to replace: it scores
the edit against dev AND holdout, and it is cheaper than the arm.

### Scrape, minimally

Importing an existing corpus IS the scrape step: copy the set from its home
(for example a benchmarks checkout) into `evals/<set>/` and convert to the
ledger shapes. **`skill:eval-import` is that job in full** -- how to classify
what arrived with each question, why nothing imports as a verified value, and
the seal that makes a later edit to a question detectable. Read it whenever the
questions came from outside, which is most of the time. While importing:

- Freeze each case's `split`: `dev` or `holdout`. Diagnose and improve read
  dev cases only; the acceptance check runs both. A set that is all dev cannot defend an
  accept.
- Later, each diagnosed-and-fixed failure becomes a new frozen dev case, so a
  fixed bug cannot silently return.

Scraping from production logs (chat transcripts, retrieval traces) is the
other supported source, and usually the better one: real traffic asks what
people actually ask. Where your logs physically live is a host concern; look
for a host-specific log-fetching skill. `skill:eval-import` takes over once
you have the text, and its `reference/case-format.md` covers what a log pull
needs that a question list does not.

Prefer variety over volume when you sample, from either source. Cases that
differ in grain, source, filter shape, and phrasing are what move a
measurement; a second sample of the same case is nearly free of new
information.

### Mode aliases

Older mode names still work as aliases for how far one run walks:

| Alias | Steps |
|---|---|
| `measure` | scrape/run + eval (`eval.py package` then needs `--without-diagnosis`) |
| `triage` | plus diagnose |
| `improve` | plus improve + acceptance check + checkpoint on accept |

Say which alias (or which steps) you are running before the first question.
Record it in `run.json`. Do not mix steps in a way that lets the answerer see
gold, issues, or the model file.

Most runs should stop after eval. Diagnose when you need a histogram of
components and owners. Improve only for diagnosed *model* gaps, one batch at
a time. Checkpoint only after the acceptance check **accepts**.

## Roles

| Role | Sees |
|---|---|
| **Answerer** | The question and the Malloy tools. Never the golden, `evals/`, the model file, or any hint it is being evaluated. |
| **Judge** | The golden and the prediction. Never conducts, never answers, never edits. One fresh subagent per verdict (`skill:eval-judge`). |
| **You (conductor / improver)** | Everything, including goldens and traces. |
| **Acceptance check** | The edit and the evidence. Never the improver's self-assessment alone. |

The answerer stays blind. That is not optional. A grader-visible answerer
writes toward the expected answer, and the score is fiction.

There are no eval MCP tools on purpose. The answerer inherits your tools,
including Shell and Read, so any eval convenience surface would also be a
gold path for it. Blindness is prevention plus detection, not a guarantee:
`eval-answer` runs the contamination checklist on every attempt, which is
why you keep a host-side tool-use log per answerer.

## Pick the target first

Both a local model server and a hosted platform expose the same two tools the
answerer needs, `get_context` and `execute_query`, so the loop runs against
either. What differs is which model is answering and whose data it reads, and
those are two separate axes:

| Target | Model under test | Data | Can edit and re-test? |
|---|---|---|---|
| **Local (direct)** | your working files | local (for example duckdb), or a direct warehouse connection | yes |
| **Local (proxied)** | your working files | the platform's connection, through a proxy connection type | yes |
| **Remote** | the published version, through the platform's hosted `get_context`/`execute_query` | the platform's | no, publishing is not an eval action |

The middle row is the one worth knowing about: it decouples the two axes, so you
can evaluate a model you are still editing against the customer's real data. It
is a connection configuration, not a feature.

### Three ways to reach the tools

That table is about which MODEL answers. A second, independent choice is where
the answerer's MCP tools come from, and `--mcp-url` is the whole of it:

| | `--target` | `--mcp-url` | Auth |
|---|---|---|---|
| **1. Local Publisher** | `local` | `http://localhost:4040/mcp` (default) | none |
| **2. Hosted, through an editor extension's local bridge** | `platform` | the localhost URL the extension prints | none: the extension holds the credential |
| **3. Hosted, directly** | `platform` | the host's `https` endpoint, scoped if it offers one | a cached OAuth login, once, interactively |

```bash
# 1. local
--target local            # --mcp-url defaults to the local Publisher

# 2. hosted via the extension's bridge -- no OAuth, but check what it exposes:
#    the same proxy may front a local Publisher instead
--target platform --mcp-url http://localhost:<port-the-extension-prints>/mcp \
  --hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>

# 3. hosted directly -- authenticate first, under the SAME server name
claude mcp add --transport http <name> <scoped-url>
claude   # /mcp -> <name> -> Authenticate
--target platform --mcp-url <scoped-url> \
  --hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>
```

All three hand the answerer the same three capabilities (`get_context`,
`execute_query`, and the docs search), so a comparison between them is between
agents that could do the same things; `test_the_two_arms_hold_the_same_capabilities`
pins it. Modes 2 and 3 take those names from `--hosted-tools`, which defaults to
the bare trio; pass it only if this host names them differently.

A platform run left on the local default `--mcp-url` is refused rather than
probed, because pointing every answerer at a local Publisher measures a
different model over different data than the run claims.

Two rules follow, and both are the kind of mistake that produces confident
nonsense rather than an error:

- **The answerer and the conductor must hit the same target.** If the answerer
  queries the published model and you re-execute its query against your edited
  local copy, the score describes neither. Decide the target before the first
  question and record it.
- **Pin the version the target actually served, not the one you happen to have.**
  A local target pins a commit; a platform target pins the published version.
  Recording a local commit for a run that queried a published model is a pin
  that means nothing.

Which target for which job:

- **Baseline what customers experience:** Remote. It is the deployed model
  through the deployed engine, which is the thing they actually hit. The judge
  sees no re-executed rows on a Remote run (there is no local copy of the
  bytes), so its verdicts rest on the answer text and the golden; say so.
- **Improve and accept:** local, because the acceptance check needs compile,
  reload, and a fresh re-answer between edits. Publishing to a customer
  environment to score an edit is not something this loop does. Where the host
  offers draft execution, that counts as local for this purpose.
- **Measure real data without touching production:** local proxied.

So a measure-only run can use any target; a run that includes **improve** needs
a local one.

Two things to check before a platform run, because neither errors and both make
the run measure something other than what it names:

- **The answerer's skills must be written for THIS host.** A shared skill names
  an MCP tool by its bare name (`get_context`) so it reads correctly anywhere,
  but a host/router skill names its own host's tools directly. Install the
  latter for the wrong host and the answerer is told to call tools it does not
  have. `run_baseline.py` warns when the manifest it loaded names Publisher-only
  tools on a platform target; point `--answerer-manifest`, or `--skills-root`,
  at the checkout that ships this host's manifest.
- **The tool names are configuration.** `--hosted-mcp-server` is both the
  `mcp__<server>__<tool>` prefix and the OAuth cache key, so it has to match the
  name the answerer authenticated under, and `--hosted-tools` lists the bare
  tools that host exposes.
- **Get the hosted tools in front of a headless answerer, one of two ways.**
  A spawned answerer cannot complete an OAuth flow, so the tools have to be
  reachable before the run starts. `run_baseline.py` proves it with one cheap
  probe and refuses to spend an arm otherwise -- a run whose answerers have no
  tools does not error, it reads as a terrible model.

  1. **Authenticate once, interactively.** Works anywhere, including a plain
     CLI install, and is the route to assume unless you know otherwise. The
     token is cached per server NAME, so authenticate under the same name the
     run passes to `--hosted-mcp-server`:

     ```bash
     claude mcp add --transport http <name> <scoped-url>
     claude          # then /mcp -> <name> -> Authenticate
     ```

     Then come back and run. This is a hand-off to a person; there is no
     headless equivalent, so plan for it rather than discovering it mid-run.

  2. **A local proxy that already holds the credential.** Some hosts ship an
     editor extension whose local MCP proxy can expose the hosted
     `get_context` / `execute_query` -- often behind a setting that is off by
     default. Where that exists, point `--mcp-url` at the proxy on localhost
     and no OAuth step is needed, because the extension holds it. Check what
     the proxy actually exposes before relying on it: the same proxy may serve
     a local Publisher's tools instead, and then `--hosted-tools` is
     naming tools that are not there. This route is not available to someone
     running the CLI alone.

- **Prefer a SCOPED endpoint URL over asking for scope.** A hosted MCP is
  usually reachable two ways: a global endpoint where every call carries an
  organization and workspace, and a scoped one where the URL itself is the
  scope. `--scope` and the prompt can only ASK an answerer to stay in one
  package; a scoped URL enforces it. For an agent being measured that is the
  difference between a case answered against the package it names and one
  answered against whatever else the account can see. Authenticate once
  interactively (`claude`, `/mcp`) under the same server name the run will use;
  the token is cached per name, and a spawned headless answerer cannot complete
  an OAuth flow.

- **Pin the VERSION in the scope, not just the package.** Write
  `--scope <env>/<package>@<version>`. Both hosted tools take a version and both
  document the same default for an omitted one: the PINNED version, which is
  whatever the workspace serves at the moment of the call. So an unversioned run
  records `targetVersion` in `run.json` and then answers from whatever is
  current, and the two part company the moment anyone publishes -- including
  mid-run, which measures two builds under one label. `--target-version` fills
  the version in when the scope omits it, so a platform run is pinned without
  opting in; a scope naming a different version is refused rather than taken as
  an override.

## Before you start

1. The model package under evaluation must live in a git repository, with
   `evals/<set>/` in the package, beside the model files. Git is the checkpoint
   mechanism; without it there is no rollback and no run can include improve.

   Keeping the set IN the package is what stops a model edit and its answer key
   drifting apart: they move in one commit, so fixing a measure and forgetting
   the golden that depended on it stops being possible. It is safe -- measured
   on a running server, a `cases.jsonl` inside a package appears in no model
   listing, no notebook listing, no package resource, and 404s over HTTP, so an
   MCP-only answerer has no route to it.

   What it buys differs by target. On a LOCAL Publisher it does not get you free
   versioning -- `sourceContentSha` hashes model paths only, so the set needs
   its own `datasetSha`. On a hosted target that publishes the whole package
   directory as an IMMUTABLE version, the set rides inside that version and
   `targetVersion` pins model and answer key together; nothing can be edited
   under a published version, which is what makes it a pin. Check which you have
   before deciding how much of this you need.

   **Look for a set and for prior runs before you author either.** A minute of
   `find . -name cases.jsonl`, a glance at the set's workdir (`<workdir>/runs/`,
   `~/.malloy-eval/<set>-<hash>/runs/` unless `eval.toml` moves it) and at your host's
   own transcripts for this repo. Two sessions fourteen minutes apart built the
   same 29-case answer key from scratch, because the first had committed
   nothing before it was deleted and the second had no way to know it existed.
   Roughly a working day was spent twice, and three specific things were
   rediscovered at cost: the MCP login flow, the retrieval gate's 401, and a
   judge-rendering bug that had already cost $17 of arm once. The two keys,
   independently authored from the same questions, differ by about ten points
   on comparable answers -- which is the available measure of how much a key
   depends on its author, and a reason to reuse one rather than rebuild it.

   **Audit the set's entity ids before the first arm.** `verify_goldens.py`
   checks that each id names something in the model; `check_findable.py` checks
   that a search of its own kind actually returns it. Both are free of model
   calls. An id that fails either scores a retrieval miss on every run, and the
   miss reads as the model's fault.

   **Commit the set before you spend money on an arm**, and keep durable
   outputs in the repository. A findings document in `~/Downloads` is gone the
   first time somebody tidies up; the set and the write-up belong in git
   beside the model.

   **Runs go in the set's workdir, never inside the model package.** A run
   directory holds a `model.malloy` snapshot, and a built report is a Malloy
   package; inside the package under test, either puts that package into
   `loadErrors`. The default workdir, `~/.malloy-eval/<set>-<hash>/`, is outside git,
   which suits a measure-only run. A run that will improve and checkpoint needs
   its ledger kept: set `[paths] workdir` in `eval.toml` to a directory in the
   model's repository but outside the package, and gitignore its `servers/`
   and `packages/`, which are a server's database and rebuilt reports.

2. The server must be up. Where your host offers retrieval tracing, turn it on
   and confirm a trace lookup answers, so a call's ranked results can be
   recovered afterwards.

   **Open-source Publisher has no trace lookup.** `PUBLISHER_MCP_TRACE=retrieval`
   is set by `serve.py` and makes Publisher write one "Retrieval trace" line per
   ranked call to its log (stage counts and timings, not the ranked results).
   Nothing reads that line yet: there is no trace store and no trace tool, so
   `traceId` is null on every local attempt. This
   does not block a scored run, because attribution never depended on it:
   `rankedSummary` is copied onto the `tool_call` event at capture, precisely so
   the evidence survives without a store. So do not refuse a local run for want
   of tracing -- an earlier version of this rule did, and it refused every local
   run there has ever been. Refuse one whose `tool_call` events carry no
   `rankedSummary`, which is what attribution actually reads.

3. Health-check: your host's status check until it reports serving, and inspect
   `loadErrors`. A dead database that still answers HTTP is an environment
   failure, not a model failure. Stop and fix it. Four consecutive
   environment or no-result attempts means stop the run.

4. Load the set: scrape/import as above, or reuse an existing `evals/<set>/`.
   `eval.py check --set <set>` names every gap in it before anything starts.
   Never keep two live copies of one set; the set directory in the model repo
   is the single source of truth, versioned by `datasetVersion` in
   `set.json`.

5. Review goldens before you score. **Check how many cases can take a verdict
   at all**, not just how many cases there are: a golden the set stamps
   `provisional`, `invalid` or `ambiguous`, or a case with no golden, scores
   `verdict: null` and stays out of the pass rate. An imported set is
   `provisional` throughout by design (`skill:eval-import`), and the only thing
   that changes that is `verify_goldens.py --promote` after a re-derivation
   through the truth package. A run whose every key is underived is refused
   rather than spent; a set of bare questions runs, because its answers are
   what keys get derived from. A verified golden that holds a value but
   no local artifact stays verified by provenance and is not scorable until you
   have rows or a scalar to compare (the judge needs both sides). A `criteria`
   golden is the exception: it holds no value, so its clauses are both sides. If diagnosis later marks
   `BAD-REFERENCE` or `AMBIGUOUS-REFERENCE`, follow
   `reference/golden-side-door.md`. Both are expected in the wild; both
   are the golden side door below, not improve, and not a sixth step.

6. `eval.py run` creates `<workdir>/runs/<runId>/run.json` with the attribution pins
   (`reference/ledger-schema.md`): mode, dataset version, **the target and the
   version it served** (a local target pins a commit, so commit or stash first;
   answering from a dirty tree pins nothing), server version, judge version and
   rubric sha, answerer model, call budget, trace mode. Freeze those for the
   whole run. Raising a call budget mid-run moved mean outcomes on an unchanged
   model.

   The call budget is `--max-turns`, written to `run.json` as `maxTurns`.
   **Size it from a pilot rather than taking the default of 30.** Run the three
   cheapest cases uncapped (`--max-turns 100`) and set the cap at twice their
   maximum. A set whose questions need two or three sources joined does not fit
   a cap sized for single-source lookups: on one such set the completed
   attempts had a median of 17 turns and a 90th percentile of 26 against a cap
   of 30, and four cases died at it. A cap 15% above the 90th percentile of
   completed work is not a safety margin. Note also that `malloy-analysis` tells
   the answerer to persist -- retry a phrasing, let a small query settle whether
   a field exists -- so a tight cap and that instruction are in direct conflict.

   **Do not change the model and the measuring instrument in the same step.**
   Those are two different axes and only one of them is cheap to separate.

   Batching MODEL edits is fine and expected. Clustering exists so that one
   edit closes several cases, and an arm per fix does not survive contact with
   arithmetic: 20 fixes over 100 cases is 2,000 answers, and at the measured
   $0.33 a case on a proxied warehouse that is $660 of answering to attribute
   what the acceptance check attributes for the price of the affected cases
   plus holdout. Budget five arms for a defensible claim (a baseline, two for
   the A/A, two post-edit), not one per edit, and let `skill:eval-improve`
   carry each cluster.

   The instrument is the other axis: the answer key, the judge and its prompt,
   the answerer model, the skills. Move one of those together with the model
   and there is nothing left holding still, so the result measures neither. On
   the run this comes from, nine model commits and a re-derived answer key
   landed between two arms, and the move from 20% to 39% belongs to no one --
   not because two model edits were batched, but because the ruler changed at
   the same time as the thing being measured. A key repair mid-improve is not
   forbidden; it ends that comparison, so re-baseline rather than quoting a
   delta across it.

   The answerer model is instrument too, and it is the one most often left
   unstated. The same 23-case set read 12 match / 11 near / 0 no_match on one
   model and 8 / 4 / 15 on a smaller one. No statement about "the agent's"
   capability means anything until two arms name the same answerer.

7. Generate every answerer prompt from the stored case in `cases.jsonl`.
   Never retype the question. A truncated retype is indistinguishable from a
   real question downstream.

## Per question

1. Health-check again.
2. Spawn a *fresh* blind subagent. Give it only the question text and the
   Malloy analysis tools. Tell it to follow the `malloy-analysis` skill. Do not
   mention eval, gold, scoring, or this skill.
3. Keep a host-side tool-use log for that subagent (name, input path or
   command, MCP tool name). Publisher traces see MCP only; a Read of a gold
   CSV is invisible server-side.
4. `skill:eval-answer`: contamination first, then re-execute, then the judge,
   then events.
5. `skill:eval-diagnose` only when this run includes diagnose, only on dev
   cases, and only after the score event exists.
6. `skill:eval-improve` only when this run includes improve, and only for
   `owner: model`. Then run the acceptance check. On accept, checkpoint.

## Report the run, or nobody can read it

A run directory is JSONL. It is a record, not a result, and the console summary
scrolls away. **Every run ends by producing something a person can open**, and
that job is `skill:eval-report`: it builds the servable run package (the case
matrix app and the aggregate notebook) and gives the template for the write-up.

Read it at the end of every run, including a run that failed. The standing
complaint about this loop is that "a bunch of stuff happens and it is hard to
know the actual results", and a ledger nobody renders is why.

Two rules from it are worth repeating here, because they are the ones a
conductor skips:

- **Separate a MODEL failure from an EVAL failure.** A wrong answer and a
  broken measurement look identical in a pass rate and have nothing else in
  common. An arm holding a truncated, contaminated or environment-failed
  attempt has no rate to quote at all.
- **Say when a step did not run, and why.** `improve` not running because every
  cluster came back `owner: agent-skill` is a RESULT, and it reads identically
  to having forgotten unless it is written down.

Keep the write-up in the repository beside the set, not in a chat log and not
in `~/Downloads`.

## Checkpoint

A checkpoint is a git commit of the model repository, taken after an acceptance check
accepts, so a bad improve direction can be rolled back. It is not a report,
and it is not a remote publish.

1. Commit the model files AND the set's ledger in one commit; put the label
   and the closed issue ids in the message. The run's ledger is in that commit
   only when the workdir is in the repository (Before you start, item 1).
2. Append the `checkpoint` event (`action: created`, label, `modelGitSha`
   from the commit you just made, issueIds). The event line itself rides in
   the next commit; append-only logs trail by one commit and that is fine.
3. Confirm `git status` is clean for the model files.

**Restore**: `git checkout <sha> -- <model files>` (or `git revert` the
checkpoint commits), then reload the package, then append a `checkpoint`
event with `action: restored` and the sha. Readers return to the model that
existed before the bad direction.

Take a checkpoint of the current model *before* the first improve batch if no
commit pins it yet. Rolling back by hand is guesswork.

If reload reports `mode: reinstalled`, the package was re-fetched from its
install location and may have overwritten the restored files. Prefer in-place
/ watch-mounted packages for this loop.

## Out of scope

This loop is local. The ledger is files, the checkpoints are git, you are the
conductor. Do not:

- publish the model to a hosted platform as a "true" checkpoint or learning
  curve
- start a Python orchestrator (`loop.py`, `run_all.py`, `improve_batch.py`)
  that runs the five steps end to end unattended. You conduct; the scripts are
  the steps, not the sequencing. There are more than twenty of them and they
  are not the exception to this: each does one step you invoke and hands back a
  result you read. What is forbidden is a script that decides what to do next,
  because every judgement this loop protects lives in that decision
- score by string-diffing rows instead of judging them, or reintroduce a
  scripted row oracle: one that can pass a wrong answer is worse than none
- wait for a bigger gold set before the loop can run; dev/holdout on what
  exists beats waiting
- register eval MCP tools or stand up an eval API
- encode unsettled goldens into the model

## Prime directives

- The model is the only thing improve edits. No question text, qids, or
  expected values in any name, doc, or comment.
- **You are measuring a model, not reviewing this harness.** Read a script when
  a number you have to report cannot be explained otherwise, and stop there. Do
  not audit the scripts, propose fixes to them, or hand back tooling critique
  in place of a result -- a run that ends in harness feedback has not answered
  the question it was asked. A defect that changed a number gets one sentence
  in the report; a defect that changed nothing gets one line in chat, once.
  Fixing it is a different task, and the user starts it.
- When the environment misbehaves, stop. Never diagnose a sick system. The
  harness is part of the environment: a known-broken measurement does not
  become quotable by being finished. Measured against this directive, an arm
  already paid for gets quoted anyway -- it happened three times in one run,
  once with the confound stated in the same message that started the arm -- so
  the harness now enforces the parts it can. A run with a truncated or
  contaminated attempt prints no pass rate and records `status: incomplete`,
  and `flip_table.py` names what one arm left unscored that the other did not.
  Read `run_error` on the attempts
  and the INCOMPLETE line before quoting any number.
- When a subagent disagrees with you, probe. Do not win by authority.
- When a rule here is wrong, change this file and note it on the run.

## Related skills

- `skill:eval-answer`: contamination, judge protocol, events. Its
  `reference/ledger-schema.md` is the file contract; `skill:eval-judge` is
  the judge.
- `skill:eval-diagnose`: component, owner, issue events. No edit.
- `skill:eval-improve`: smallest model edit, probe receipts, no self-accept.
- `skill:eval-report`: the run package and the write-up a person reads.
- The `malloy-analysis` skill: what the blind answerer follows. It is installed
  from the `analysis` manifest group, not the `eval` group.
