Agent skill

Eval Loop

by malloydata in malloydata/publisher

Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

MITAuto-check passedData & Analytics

Install Eval Loop

skills CLI
$ npx skills add malloydata/publisher --skill eval-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-loop .claude/skills/eval-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-loop
GitHub stars
116
Token cost
~7.8k tokens
SKILL.md length
4,714 words
Files
35 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

  • Works in 2 steps: Authenticate once, interactively. Works… → A local proxy that already holds the…
  • Diagnose failures
  • SKILL.md covers Where the rest of this lives, The five steps, Roles and Pick the target first, plus 7 more sections
  • Runs Python scripts from its folder; calls git and claude

What it does

Eval Loop is an agent skill from malloydata/publisher. Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. You are the conductor: import cases into the file ledger, spawn a blind answerer, then run eval-answer, eval-diagnose, and eval-improve. Persistence is plain files: the set in the model package's evals/ directory, runs in the set's workdir; checkpoints are git commits of the model repo. eval.py runs each step from the set's eval.toml. Use to score a model, diagnose failures, improve behind an acceptance…

Its SKILL.md is about 7.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 36 other files, including scripts (for example `reference/acceptance-check.md`, `reference/auditing-an-answer-key.md` and `reference/checking-the-judge.md`).

It sits in Data & Analytics, covering Web scraping, LLM evaluation and Commit messages. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Diagnose failures
  • Improve behind an acceptance check
  • Roll back a bad direction

Example prompts

  • “s evals/ directory, runs in the set”
  • “/eval-loop”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Authenticate once, interactively. Works anywhere, including a plain
  2. A local proxy that already holds the credential. Some hosts ship an

What it can do on your machine

Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Loop loads about 7.8k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 4,714 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~7.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 4,714 words, ~7,841 tokens.

Download SKILL.mdSave it as .claude/skills/eval-loop/SKILL.md (or your agent's skills folder). This skill also uses 34 other files; get the full folder from GitHub.
name
eval-loop
description
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. You are the conductor: import cases into the file ledger, spawn a blind answerer, then run eval-answer, eval-diagnose, and eval-improve. Persistence is plain files: the set in the model package's evals/ directory, runs in the set's workdir; checkpoints are git commits of the model repo. `eval.py` runs each step from the set's eval.toml. Use to score a model, diagnose failures, improve behind an acceptance check, or roll back a bad direction.

The Evaluation Loop

You conduct this loop. There is no batch orchestrator to start, no eval API, and no eval MCP tools. The ledger is plain files: the set in the model package's git repository, and each run in the set's workdir (reference/ledger-schema.md in skill:eval-answer defines every file and event). scripts/eval.py runs each step from the set's eval.toml; reference/running-a-run.md has the commands. Scoring is an LLM judge you spawn per case. There is no scripted scorer, and there will not be one: a script that can pass a wrong answer is worse than none. The scripts under scripts/ run the loop -- they answer, re-execute, spawn the judge, compare runs, and write the ledger -- but none of them decides whether an answer was right.

scrape/run  ->  eval  ->  diagnose  ->  improve  ->  checkpoint

This skill conducts; it does not restate. Scoring lives in skill:eval-answer. Components and owners live in skill:eval-diagnose. Edit rules live in skill:eval-improve.

Do not merge eval into diagnose. A conductor who scores while explaining writes the explanation into the score. Do not skip the acceptance check inside improve. The acceptance check decides whether this edit stays. Checkpoint decides whether a sequence of accepted edits can be undone.

Where the rest of this lives

This file is the procedure. The things it used to carry inline are files beside it now, because each is needed at one moment rather than every run, and loading all of them for every run is how a skill stops being read.

WhenRead
the set does not exist yet, or has never runreference/setting-up-a-set.md
about to run onereference/running-a-run.md
a golden is wrong, doubted, or out of step with the modelreference/golden-side-door.md
auditing a key you doubt, or a set you did not authorreference/auditing-an-answer-key.md
deciding whether an edit staysreference/acceptance-check.md
about to quote a number, set the band, or read a set's flipsreference/measurement.md
the run finished and someone has to read itskill:eval-report
you changed judge doctrine or its inputsreference/checking-the-judge.md

Read the file, do not work from the summary here. The acceptance-check rules and the golden side door are both places where acting on a half-memory of the rule produces a confident wrong answer rather than an error.

The five steps

StepJobWrites
a. scrape / runPut cases in the ledger; spawn a blind answerercases; attempt, tool_call
b. evalJudge the answer; score which required entities retrieval deliveredscore
c. diagnoseWhy it failed, who owns itissue / issue_status. Stop. Do not edit.
d. improveOne smallest model edit, then the acceptance checkimprove writes candidate; you write acceptance_check. Revert on reject.
e. checkpointGit commit after an accepted acceptance checkcheckpoint event, then the commit

Every run then reports (below). A run whose result exists only as JSONL and a scrolled-away console summary has not been delivered.

scrape and run share a letter but are not the same job. Scrape writes cases. Run writes attempts. Do not invent questions and score them in one breath.

Step d is skill:eval-improve, not a hand edit plus another arm. The failure mode is specific and it has happened: the edits were made by hand, published, and validated by re-running the whole arm -- with the answer key repaired in the same window. No candidate, acceptance_check or checkpoint event exists in that run's ledger, and the resulting move cannot be attributed to the model, the key, or the answerer. Re-running a whole arm after changing two things is the one thing the acceptance check exists to replace: it scores the edit against dev AND holdout, and it is cheaper than the arm.

Scrape, minimally

Importing an existing corpus IS the scrape step: copy the set from its home (for example a benchmarks checkout) into evals/<set>/ and convert to the ledger shapes. skill:eval-import is that job in full -- how to classify what arrived with each question, why nothing imports as a verified value, and the seal that makes a later edit to a question detectable. Read it whenever the questions came from outside, which is most of the time. While importing:

  • Freeze each case's split: dev or holdout. Diagnose and improve read dev cases only; the acceptance check runs both. A set that is all dev cannot defend an accept.
  • Later, each diagnosed-and-fixed failure becomes a new frozen dev case, so a fixed bug cannot silently return.

Scraping from production logs (chat transcripts, retrieval traces) is the other supported source, and usually the better one: real traffic asks what people actually ask. Where your logs physically live is a host concern; look for a host-specific log-fetching skill. skill:eval-import takes over once you have the text, and its reference/case-format.md covers what a log pull needs that a question list does not.

Prefer variety over volume when you sample, from either source. Cases that differ in grain, source, filter shape, and phrasing are what move a measurement; a second sample of the same case is nearly free of new information.

Mode aliases

Older mode names still work as aliases for how far one run walks:

AliasSteps
measurescrape/run + eval (eval.py package then needs --without-diagnosis)
triageplus diagnose
improveplus improve + acceptance check + checkpoint on accept

Say which alias (or which steps) you are running before the first question. Record it in run.json. Do not mix steps in a way that lets the answerer see gold, issues, or the model file.

Most runs should stop after eval. Diagnose when you need a histogram of components and owners. Improve only for diagnosed model gaps, one batch at a time. Checkpoint only after the acceptance check accepts.

Roles

RoleSees
AnswererThe question and the Malloy tools. Never the golden, evals/, the model file, or any hint it is being evaluated.
JudgeThe golden and the prediction. Never conducts, never answers, never edits. One fresh subagent per verdict (skill:eval-judge).
You (conductor / improver)Everything, including goldens and traces.
Acceptance checkThe edit and the evidence. Never the improver's self-assessment alone.

The answerer stays blind. That is not optional. A grader-visible answerer writes toward the expected answer, and the score is fiction.

There are no eval MCP tools on purpose. The answerer inherits your tools, including Shell and Read, so any eval convenience surface would also be a gold path for it. Blindness is prevention plus detection, not a guarantee: eval-answer runs the contamination checklist on every attempt, which is why you keep a host-side tool-use log per answerer.

Pick the target first

Both a local model server and a hosted platform expose the same two tools the answerer needs, get_context and execute_query, so the loop runs against either. What differs is which model is answering and whose data it reads, and those are two separate axes:

TargetModel under testDataCan edit and re-test?
Local (direct)your working fileslocal (for example duckdb), or a direct warehouse connectionyes
Local (proxied)your working filesthe platform's connection, through a proxy connection typeyes
Remotethe published version, through the platform's hosted get_context/execute_querythe platform'sno, publishing is not an eval action

The middle row is the one worth knowing about: it decouples the two axes, so you can evaluate a model you are still editing against the customer's real data. It is a connection configuration, not a feature.

Three ways to reach the tools

That table is about which MODEL answers. A second, independent choice is where the answerer's MCP tools come from, and --mcp-url is the whole of it:

--target--mcp-urlAuth
1. Local Publisherlocalhttp://localhost:4040/mcp (default)none
2. Hosted, through an editor extension's local bridgeplatformthe localhost URL the extension printsnone: the extension holds the credential
3. Hosted, directlyplatformthe host's https endpoint, scoped if it offers onea cached OAuth login, once, interactively
bash
# 1. local
--target local            # --mcp-url defaults to the local Publisher

# 2. hosted via the extension's bridge -- no OAuth, but check what it exposes:
#    the same proxy may front a local Publisher instead
--target platform --mcp-url http://localhost:<port-the-extension-prints>/mcp \
  --hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>

# 3. hosted directly -- authenticate first, under the SAME server name
claude mcp add --transport http <name> <scoped-url>
claude   # /mcp -> <name> -> Authenticate
--target platform --mcp-url <scoped-url> \
  --hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>

All three hand the answerer the same three capabilities (get_context, execute_query, and the docs search), so a comparison between them is between agents that could do the same things; test_the_two_arms_hold_the_same_capabilities pins it. Modes 2 and 3 take those names from --hosted-tools, which defaults to the bare trio; pass it only if this host names them differently.

A platform run left on the local default --mcp-url is refused rather than probed, because pointing every answerer at a local Publisher measures a different model over different data than the run claims.

Two rules follow, and both are the kind of mistake that produces confident nonsense rather than an error:

  • The answerer and the conductor must hit the same target. If the answerer queries the published model and you re-execute its query against your edited local copy, the score describes neither. Decide the target before the first question and record it.
  • Pin the version the target actually served, not the one you happen to have. A local target pins a commit; a platform target pins the published version. Recording a local commit for a run that queried a published model is a pin that means nothing.

Which target for which job:

  • Baseline what customers experience: Remote. It is the deployed model through the deployed engine, which is the thing they actually hit. The judge sees no re-executed rows on a Remote run (there is no local copy of the bytes), so its verdicts rest on the answer text and the golden; say so.
  • Improve and accept: local, because the acceptance check needs compile, reload, and a fresh re-answer between edits. Publishing to a customer environment to score an edit is not something this loop does. Where the host offers draft execution, that counts as local for this purpose.
  • Measure real data without touching production: local proxied.

So a measure-only run can use any target; a run that includes improve needs a local one.

Two things to check before a platform run, because neither errors and both make the run measure something other than what it names:

  • The answerer's skills must be written for THIS host. A shared skill names an MCP tool by its bare name (get_context) so it reads correctly anywhere, but a host/router skill names its own host's tools directly. Install the latter for the wrong host and the answerer is told to call tools it does not have. run_baseline.py warns when the manifest it loaded names Publisher-only tools on a platform target; point --answerer-manifest, or --skills-root, at the checkout that ships this host's manifest.

  • The tool names are configuration. --hosted-mcp-server is both the mcp__<server>__<tool> prefix and the OAuth cache key, so it has to match the name the answerer authenticated under, and --hosted-tools lists the bare tools that host exposes.

  • Get the hosted tools in front of a headless answerer, one of two ways. A spawned answerer cannot complete an OAuth flow, so the tools have to be reachable before the run starts. run_baseline.py proves it with one cheap probe and refuses to spend an arm otherwise -- a run whose answerers have no tools does not error, it reads as a terrible model.

    1. Authenticate once, interactively. Works anywhere, including a plain CLI install, and is the route to assume unless you know otherwise. The token is cached per server NAME, so authenticate under the same name the run passes to --hosted-mcp-server:

      bash
      claude mcp add --transport http <name> <scoped-url>
      claude          # then /mcp -> <name> -> Authenticate

      Then come back and run. This is a hand-off to a person; there is no headless equivalent, so plan for it rather than discovering it mid-run.

    2. A local proxy that already holds the credential. Some hosts ship an editor extension whose local MCP proxy can expose the hosted get_context / execute_query -- often behind a setting that is off by default. Where that exists, point --mcp-url at the proxy on localhost and no OAuth step is needed, because the extension holds it. Check what the proxy actually exposes before relying on it: the same proxy may serve a local Publisher's tools instead, and then --hosted-tools is naming tools that are not there. This route is not available to someone running the CLI alone.

  • Prefer a SCOPED endpoint URL over asking for scope. A hosted MCP is usually reachable two ways: a global endpoint where every call carries an organization and workspace, and a scoped one where the URL itself is the scope. --scope and the prompt can only ASK an answerer to stay in one package; a scoped URL enforces it. For an agent being measured that is the difference between a case answered against the package it names and one answered against whatever else the account can see. Authenticate once interactively (claude, /mcp) under the same server name the run will use; the token is cached per name, and a spawned headless answerer cannot complete an OAuth flow.

  • Pin the VERSION in the scope, not just the package. Write --scope <env>/<package>@<version>. Both hosted tools take a version and both document the same default for an omitted one: the PINNED version, which is whatever the workspace serves at the moment of the call. So an unversioned run records targetVersion in run.json and then answers from whatever is current, and the two part company the moment anyone publishes -- including mid-run, which measures two builds under one label. --target-version fills the version in when the scope omits it, so a platform run is pinned without opting in; a scope naming a different version is refused rather than taken as an override.

Show full SKILL.md (2,507 more words)Show less

Before you start

  1. The model package under evaluation must live in a git repository, with evals/<set>/ in the package, beside the model files. Git is the checkpoint mechanism; without it there is no rollback and no run can include improve.

    Keeping the set IN the package is what stops a model edit and its answer key drifting apart: they move in one commit, so fixing a measure and forgetting the golden that depended on it stops being possible. It is safe -- measured on a running server, a cases.jsonl inside a package appears in no model listing, no notebook listing, no package resource, and 404s over HTTP, so an MCP-only answerer has no route to it.

    What it buys differs by target. On a LOCAL Publisher it does not get you free versioning -- sourceContentSha hashes model paths only, so the set needs its own datasetSha. On a hosted target that publishes the whole package directory as an IMMUTABLE version, the set rides inside that version and targetVersion pins model and answer key together; nothing can be edited under a published version, which is what makes it a pin. Check which you have before deciding how much of this you need.

    Look for a set and for prior runs before you author either. A minute of find . -name cases.jsonl, a glance at the set's workdir (<workdir>/runs/, ~/.malloy-eval/<set>-<hash>/runs/ unless eval.toml moves it) and at your host's own transcripts for this repo. Two sessions fourteen minutes apart built the same 29-case answer key from scratch, because the first had committed nothing before it was deleted and the second had no way to know it existed. Roughly a working day was spent twice, and three specific things were rediscovered at cost: the MCP login flow, the retrieval gate's 401, and a judge-rendering bug that had already cost $17 of arm once. The two keys, independently authored from the same questions, differ by about ten points on comparable answers -- which is the available measure of how much a key depends on its author, and a reason to reuse one rather than rebuild it.

    Audit the set's entity ids before the first arm. verify_goldens.py checks that each id names something in the model; check_findable.py checks that a search of its own kind actually returns it. Both are free of model calls. An id that fails either scores a retrieval miss on every run, and the miss reads as the model's fault.

    Commit the set before you spend money on an arm, and keep durable outputs in the repository. A findings document in ~/Downloads is gone the first time somebody tidies up; the set and the write-up belong in git beside the model.

    Runs go in the set's workdir, never inside the model package. A run directory holds a model.malloy snapshot, and a built report is a Malloy package; inside the package under test, either puts that package into loadErrors. The default workdir, ~/.malloy-eval/<set>-<hash>/, is outside git, which suits a measure-only run. A run that will improve and checkpoint needs its ledger kept: set [paths] workdir in eval.toml to a directory in the model's repository but outside the package, and gitignore its servers/ and packages/, which are a server's database and rebuilt reports.

  2. The server must be up. Where your host offers retrieval tracing, turn it on and confirm a trace lookup answers, so a call's ranked results can be recovered afterwards.

    Open-source Publisher has no trace lookup. PUBLISHER_MCP_TRACE=retrieval is set by serve.py and makes Publisher write one "Retrieval trace" line per ranked call to its log (stage counts and timings, not the ranked results). Nothing reads that line yet: there is no trace store and no trace tool, so traceId is null on every local attempt. This does not block a scored run, because attribution never depended on it: rankedSummary is copied onto the tool_call event at capture, precisely so the evidence survives without a store. So do not refuse a local run for want of tracing -- an earlier version of this rule did, and it refused every local run there has ever been. Refuse one whose tool_call events carry no rankedSummary, which is what attribution actually reads.

  3. Health-check: your host's status check until it reports serving, and inspect loadErrors. A dead database that still answers HTTP is an environment failure, not a model failure. Stop and fix it. Four consecutive environment or no-result attempts means stop the run.

  4. Load the set: scrape/import as above, or reuse an existing evals/<set>/. eval.py check --set <set> names every gap in it before anything starts. Never keep two live copies of one set; the set directory in the model repo is the single source of truth, versioned by datasetVersion in set.json.

  5. Review goldens before you score. Check how many cases can take a verdict at all, not just how many cases there are: a golden the set stamps provisional, invalid or ambiguous, or a case with no golden, scores verdict: null and stays out of the pass rate. An imported set is provisional throughout by design (skill:eval-import), and the only thing that changes that is verify_goldens.py --promote after a re-derivation through the truth package. A run whose every key is underived is refused rather than spent; a set of bare questions runs, because its answers are what keys get derived from. A verified golden that holds a value but no local artifact stays verified by provenance and is not scorable until you have rows or a scalar to compare (the judge needs both sides). A criteria golden is the exception: it holds no value, so its clauses are both sides. If diagnosis later marks BAD-REFERENCE or AMBIGUOUS-REFERENCE, follow reference/golden-side-door.md. Both are expected in the wild; both are the golden side door below, not improve, and not a sixth step.

  6. eval.py run creates <workdir>/runs/<runId>/run.json with the attribution pins (reference/ledger-schema.md): mode, dataset version, the target and the version it served (a local target pins a commit, so commit or stash first; answering from a dirty tree pins nothing), server version, judge version and rubric sha, answerer model, call budget, trace mode. Freeze those for the whole run. Raising a call budget mid-run moved mean outcomes on an unchanged model.

    The call budget is --max-turns, written to run.json as maxTurns. Size it from a pilot rather than taking the default of 30. Run the three cheapest cases uncapped (--max-turns 100) and set the cap at twice their maximum. A set whose questions need two or three sources joined does not fit a cap sized for single-source lookups: on one such set the completed attempts had a median of 17 turns and a 90th percentile of 26 against a cap of 30, and four cases died at it. A cap 15% above the 90th percentile of completed work is not a safety margin. Note also that malloy-analysis tells the answerer to persist -- retry a phrasing, let a small query settle whether a field exists -- so a tight cap and that instruction are in direct conflict.

    Do not change the model and the measuring instrument in the same step. Those are two different axes and only one of them is cheap to separate.

    Batching MODEL edits is fine and expected. Clustering exists so that one edit closes several cases, and an arm per fix does not survive contact with arithmetic: 20 fixes over 100 cases is 2,000 answers, and at the measured $0.33 a case on a proxied warehouse that is $660 of answering to attribute what the acceptance check attributes for the price of the affected cases plus holdout. Budget five arms for a defensible claim (a baseline, two for the A/A, two post-edit), not one per edit, and let skill:eval-improve carry each cluster.

    The instrument is the other axis: the answer key, the judge and its prompt, the answerer model, the skills. Move one of those together with the model and there is nothing left holding still, so the result measures neither. On the run this comes from, nine model commits and a re-derived answer key landed between two arms, and the move from 20% to 39% belongs to no one -- not because two model edits were batched, but because the ruler changed at the same time as the thing being measured. A key repair mid-improve is not forbidden; it ends that comparison, so re-baseline rather than quoting a delta across it.

    The answerer model is instrument too, and it is the one most often left unstated. The same 23-case set read 12 match / 11 near / 0 no_match on one model and 8 / 4 / 15 on a smaller one. No statement about "the agent's" capability means anything until two arms name the same answerer.

  7. Generate every answerer prompt from the stored case in cases.jsonl. Never retype the question. A truncated retype is indistinguishable from a real question downstream.

Per question

  1. Health-check again.
  2. Spawn a fresh blind subagent. Give it only the question text and the Malloy analysis tools. Tell it to follow the malloy-analysis skill. Do not mention eval, gold, scoring, or this skill.
  3. Keep a host-side tool-use log for that subagent (name, input path or command, MCP tool name). Publisher traces see MCP only; a Read of a gold CSV is invisible server-side.
  4. skill:eval-answer: contamination first, then re-execute, then the judge, then events.
  5. skill:eval-diagnose only when this run includes diagnose, only on dev cases, and only after the score event exists.
  6. skill:eval-improve only when this run includes improve, and only for owner: model. Then run the acceptance check. On accept, checkpoint.

Report the run, or nobody can read it

A run directory is JSONL. It is a record, not a result, and the console summary scrolls away. Every run ends by producing something a person can open, and that job is skill:eval-report: it builds the servable run package (the case matrix app and the aggregate notebook) and gives the template for the write-up.

Read it at the end of every run, including a run that failed. The standing complaint about this loop is that "a bunch of stuff happens and it is hard to know the actual results", and a ledger nobody renders is why.

Two rules from it are worth repeating here, because they are the ones a conductor skips:

  • Separate a MODEL failure from an EVAL failure. A wrong answer and a broken measurement look identical in a pass rate and have nothing else in common. An arm holding a truncated, contaminated or environment-failed attempt has no rate to quote at all.
  • Say when a step did not run, and why. improve not running because every cluster came back owner: agent-skill is a RESULT, and it reads identically to having forgotten unless it is written down.

Keep the write-up in the repository beside the set, not in a chat log and not in ~/Downloads.

Checkpoint

A checkpoint is a git commit of the model repository, taken after an acceptance check accepts, so a bad improve direction can be rolled back. It is not a report, and it is not a remote publish.

  1. Commit the model files AND the set's ledger in one commit; put the label and the closed issue ids in the message. The run's ledger is in that commit only when the workdir is in the repository (Before you start, item 1).
  2. Append the checkpoint event (action: created, label, modelGitSha from the commit you just made, issueIds). The event line itself rides in the next commit; append-only logs trail by one commit and that is fine.
  3. Confirm git status is clean for the model files.

Restore: git checkout <sha> -- <model files> (or git revert the checkpoint commits), then reload the package, then append a checkpoint event with action: restored and the sha. Readers return to the model that existed before the bad direction.

Take a checkpoint of the current model before the first improve batch if no commit pins it yet. Rolling back by hand is guesswork.

If reload reports mode: reinstalled, the package was re-fetched from its install location and may have overwritten the restored files. Prefer in-place / watch-mounted packages for this loop.

Out of scope

This loop is local. The ledger is files, the checkpoints are git, you are the conductor. Do not:

  • publish the model to a hosted platform as a "true" checkpoint or learning curve
  • start a Python orchestrator (loop.py, run_all.py, improve_batch.py) that runs the five steps end to end unattended. You conduct; the scripts are the steps, not the sequencing. There are more than twenty of them and they are not the exception to this: each does one step you invoke and hands back a result you read. What is forbidden is a script that decides what to do next, because every judgement this loop protects lives in that decision
  • score by string-diffing rows instead of judging them, or reintroduce a scripted row oracle: one that can pass a wrong answer is worse than none
  • wait for a bigger gold set before the loop can run; dev/holdout on what exists beats waiting
  • register eval MCP tools or stand up an eval API
  • encode unsettled goldens into the model

Prime directives

  • The model is the only thing improve edits. No question text, qids, or expected values in any name, doc, or comment.
  • You are measuring a model, not reviewing this harness. Read a script when a number you have to report cannot be explained otherwise, and stop there. Do not audit the scripts, propose fixes to them, or hand back tooling critique in place of a result -- a run that ends in harness feedback has not answered the question it was asked. A defect that changed a number gets one sentence in the report; a defect that changed nothing gets one line in chat, once. Fixing it is a different task, and the user starts it.
  • When the environment misbehaves, stop. Never diagnose a sick system. The harness is part of the environment: a known-broken measurement does not become quotable by being finished. Measured against this directive, an arm already paid for gets quoted anyway -- it happened three times in one run, once with the confound stated in the same message that started the arm -- so the harness now enforces the parts it can. A run with a truncated or contaminated attempt prints no pass rate and records status: incomplete, and flip_table.py names what one arm left unscored that the other did not. Read run_error on the attempts and the INCOMPLETE line before quoting any number.
  • When a subagent disagrees with you, probe. Do not win by authority.
  • When a rule here is wrong, change this file and note it on the run.
  • skill:eval-answer: contamination, judge protocol, events. Its reference/ledger-schema.md is the file contract; skill:eval-judge is the judge.
  • skill:eval-diagnose: component, owner, issue events. No edit.
  • skill:eval-improve: smallest model edit, probe receipts, no self-accept.
  • skill:eval-report: the run package and the write-up a person reads.
  • The malloy-analysis skill: what the blind answerer follows. It is installed from the analysis manifest group, not the eval group.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 34 other files (scripts) in skills/eval-loop of malloydata/publisher.

  • SKILL.md
  • reference/acceptance-check.md
  • reference/auditing-an-answer-key.md
  • reference/checking-the-judge.md
  • reference/golden-side-door.md
  • reference/measurement.md
  • reference/running-a-run.md
  • reference/setting-up-a-set.md
  • scripts/agent_harness.py
  • scripts/agent_harness_test.py
  • scripts/build_run_package.py
  • scripts/build_run_package_test.py
  • scripts/check_judge.py
  • scripts/check_judge_test.py
  • scripts/check_set.py
  • scripts/check_set_test.py
  • scripts/eval.py
  • scripts/eval_test.py
  • scripts/flip_table.py
  • … and 16 more

Open the folder on GitHubat commit acc1acd

Compare with similar skills

Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Loop this skillmalloydata/publisher116—~7.8kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.8k23 repos~868Automated safety check: PassApache-2.0
Querying Indonesian Gov Datasuryast/indonesia-gov-apis172—~997Automated safety check: PassMIT
Phoenix Release PleaseArize-ai/phoenix12k—~708Automated safety check: PassApache-2.0
Apify Actor Developmentsickn33/agentic-awesome-skills47k2 repos~3.3kAutomated safety check: PassMIT
Daily News Reportaiskillstore/marketplace4306 repos~2.8kAutomated safety check: PassNone

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Querying Indonesian Gov Data

    suryast/indonesia-gov-apis

    Query 57 Indonesian government APIs and data sources — BPJPH halal certification, BPOM food safety, OJK financial legality, BPS statistics, BMKG weather/earthquakes, Bank Indonesia exchange rates…

    172 GitHub stars~997 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Phoenix Release Please

    Arize-ai/phoenix

    Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

    12k GitHub stars~708 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Apify Actor Development

    sickn33/agentic-awesome-skills

    Important: Before you begin, fill in the generatedBy property in the meta section of .actor/actor.json.

    47k GitHub starsUsed in 2 repos~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Daily News Report

    aiskillstore/marketplace

    Scrapes content based on a preset URL list, filters high-quality technical information, and generates daily Markdown reports.

    430 GitHub starsUsed in 6 repos~2.8k tokens
    Data & AnalyticsAuto-check passed
  • Scrapling

    foryourhealth111-pixel/Vibe-Skills

    CLI-first web scraping & content extraction with optional MCP server.

    3.6k GitHub stars~1.1k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Malloy Analysis Report

    malloydata/publisher

    Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.

    116 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Questions about Eval Loop

What does Eval Loop do?

Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. Eval Loop is an agent skill from malloydata/publisher. Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

When should I use Eval Loop?

Eval Loop fits situations like: diagnose failures; improve behind an acceptance check; roll back a bad direction.

How do I install Eval Loop in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-loop -a claude-code`. Or copy the skill folder (skills/eval-loop in malloydata/publisher) into .claude/skills/eval-loop in your project. Claude Code loads it when a task matches its description.

How do I install Eval Loop in Codex?

Run `npx skills add malloydata/publisher --skill eval-loop -a codex`. Or copy the skill folder (skills/eval-loop in malloydata/publisher) into .agents/skills/eval-loop in your project. Codex loads it when a task matches its description.

Can I use Eval Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-loop, .gemini/skills/eval-loop, .github/skills/eval-loop and .opencode/skills/eval-loop in your project.

What does Eval Loop need to run?

Going by SKILL.md and its folder, Eval Loop needs Python for the scripts in its folder and the command-line tools its instructions call (git and claude). Our summary lists: Python 3.

Does Eval Loop access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Loop use?

Eval Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Loop use?

About 7.8k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Loop?

Skills that share tags, products or a category with Eval Loop: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Querying Indonesian Gov Data (suryast/indonesia-gov-apis, 172 stars), Phoenix Release Please (Arize-ai/phoenix, 12k stars) and Apify Actor Development (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Loop?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.