Tmux
trpc-group/trpc-agent-go
Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
$ npx skills add malloydata/publisher --skill eval-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-loop .claude/skills/eval-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .claude/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-loop .agents/skills/eval-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .agents/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-loop .cursor/skills/eval-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .cursor/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-loop .gemini/skills/eval-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .gemini/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-loop .github/skills/eval-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .github/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-loop .opencode/skills/eval-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-loop into .opencode/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-loopConduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
Eval Loop is an agent skill from malloydata/publisher. Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. You are the conductor: import cases into the file ledger, spawn a blind answerer, then run eval-answer, eval-diagnose, and eval-improve. Persistence is plain files: the set in the model package's evals/ directory, runs in the set's workdir; checkpoints are git commits of the model repo. eval.py runs each step from the set's eval.toml. Use to score a model, diagnose failures, improve behind an acceptance…
Its SKILL.md is about 7.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 36 other files, including scripts (for example `reference/acceptance-check.md`, `reference/auditing-an-answer-key.md` and `reference/checking-the-judge.md`).
It sits in Data & Analytics, covering Web scraping, LLM evaluation and Commit messages. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 11 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
gitclaudeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Loop loads about 7.8k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 4,714 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 4,714 words, ~7,841 tokens.
.claude/skills/eval-loop/SKILL.md (or your agent's skills folder). This skill also uses 34 other files; get the full folder from GitHub.You conduct this loop. There is no batch orchestrator to start, no eval API,
and no eval MCP tools. The ledger is plain files: the set in the model
package's git repository, and each run in the set's workdir
(reference/ledger-schema.md in skill:eval-answer defines every file and
event). scripts/eval.py runs each step from the set's eval.toml;
reference/running-a-run.md has the commands. Scoring is an LLM judge you spawn per case. There is no
scripted scorer, and there will not be one: a script that can pass a wrong
answer is worse than none. The scripts under scripts/ run the loop -- they
answer, re-execute, spawn the judge, compare runs, and write the ledger -- but
none of them decides whether an answer was right.
scrape/run -> eval -> diagnose -> improve -> checkpointThis skill conducts; it does not restate. Scoring lives in
skill:eval-answer. Components and owners live in skill:eval-diagnose.
Edit rules live in skill:eval-improve.
Do not merge eval into diagnose. A conductor who scores while explaining writes the explanation into the score. Do not skip the acceptance check inside improve. The acceptance check decides whether this edit stays. Checkpoint decides whether a sequence of accepted edits can be undone.
This file is the procedure. The things it used to carry inline are files beside it now, because each is needed at one moment rather than every run, and loading all of them for every run is how a skill stops being read.
| When | Read |
|---|---|
| the set does not exist yet, or has never run | reference/setting-up-a-set.md |
| about to run one | reference/running-a-run.md |
| a golden is wrong, doubted, or out of step with the model | reference/golden-side-door.md |
| auditing a key you doubt, or a set you did not author | reference/auditing-an-answer-key.md |
| deciding whether an edit stays | reference/acceptance-check.md |
| about to quote a number, set the band, or read a set's flips | reference/measurement.md |
| the run finished and someone has to read it | skill:eval-report |
| you changed judge doctrine or its inputs | reference/checking-the-judge.md |
Read the file, do not work from the summary here. The acceptance-check rules and the golden side door are both places where acting on a half-memory of the rule produces a confident wrong answer rather than an error.
| Step | Job | Writes |
|---|---|---|
| a. scrape / run | Put cases in the ledger; spawn a blind answerer | cases; attempt, tool_call |
| b. eval | Judge the answer; score which required entities retrieval delivered | score |
| c. diagnose | Why it failed, who owns it | issue / issue_status. Stop. Do not edit. |
| d. improve | One smallest model edit, then the acceptance check | improve writes candidate; you write acceptance_check. Revert on reject. |
| e. checkpoint | Git commit after an accepted acceptance check | checkpoint event, then the commit |
Every run then reports (below). A run whose result exists only as JSONL and a scrolled-away console summary has not been delivered.
scrape and run share a letter but are not the same job. Scrape writes cases. Run writes attempts. Do not invent questions and score them in one breath.
Step d is skill:eval-improve, not a hand edit plus another arm. The
failure mode is specific and it has happened: the edits were made by hand,
published, and validated by re-running the whole arm -- with the answer key
repaired in the same window. No candidate, acceptance_check or checkpoint
event exists in that run's ledger, and the resulting move cannot be attributed
to the model, the key, or the answerer. Re-running a whole arm after changing
two things is the one thing the acceptance check exists to replace: it scores
the edit against dev AND holdout, and it is cheaper than the arm.
Importing an existing corpus IS the scrape step: copy the set from its home
(for example a benchmarks checkout) into evals/<set>/ and convert to the
ledger shapes. skill:eval-import is that job in full -- how to classify
what arrived with each question, why nothing imports as a verified value, and
the seal that makes a later edit to a question detectable. Read it whenever the
questions came from outside, which is most of the time. While importing:
split: dev or holdout. Diagnose and improve read
dev cases only; the acceptance check runs both. A set that is all dev cannot defend an
accept.Scraping from production logs (chat transcripts, retrieval traces) is the
other supported source, and usually the better one: real traffic asks what
people actually ask. Where your logs physically live is a host concern; look
for a host-specific log-fetching skill. skill:eval-import takes over once
you have the text, and its reference/case-format.md covers what a log pull
needs that a question list does not.
Prefer variety over volume when you sample, from either source. Cases that differ in grain, source, filter shape, and phrasing are what move a measurement; a second sample of the same case is nearly free of new information.
Older mode names still work as aliases for how far one run walks:
| Alias | Steps |
|---|---|
measure | scrape/run + eval (eval.py package then needs --without-diagnosis) |
triage | plus diagnose |
improve | plus improve + acceptance check + checkpoint on accept |
Say which alias (or which steps) you are running before the first question.
Record it in run.json. Do not mix steps in a way that lets the answerer see
gold, issues, or the model file.
Most runs should stop after eval. Diagnose when you need a histogram of components and owners. Improve only for diagnosed model gaps, one batch at a time. Checkpoint only after the acceptance check accepts.
| Role | Sees |
|---|---|
| Answerer | The question and the Malloy tools. Never the golden, evals/, the model file, or any hint it is being evaluated. |
| Judge | The golden and the prediction. Never conducts, never answers, never edits. One fresh subagent per verdict (skill:eval-judge). |
| You (conductor / improver) | Everything, including goldens and traces. |
| Acceptance check | The edit and the evidence. Never the improver's self-assessment alone. |
The answerer stays blind. That is not optional. A grader-visible answerer writes toward the expected answer, and the score is fiction.
There are no eval MCP tools on purpose. The answerer inherits your tools,
including Shell and Read, so any eval convenience surface would also be a
gold path for it. Blindness is prevention plus detection, not a guarantee:
eval-answer runs the contamination checklist on every attempt, which is
why you keep a host-side tool-use log per answerer.
Both a local model server and a hosted platform expose the same two tools the
answerer needs, get_context and execute_query, so the loop runs against
either. What differs is which model is answering and whose data it reads, and
those are two separate axes:
| Target | Model under test | Data | Can edit and re-test? |
|---|---|---|---|
| Local (direct) | your working files | local (for example duckdb), or a direct warehouse connection | yes |
| Local (proxied) | your working files | the platform's connection, through a proxy connection type | yes |
| Remote | the published version, through the platform's hosted get_context/execute_query | the platform's | no, publishing is not an eval action |
The middle row is the one worth knowing about: it decouples the two axes, so you can evaluate a model you are still editing against the customer's real data. It is a connection configuration, not a feature.
That table is about which MODEL answers. A second, independent choice is where
the answerer's MCP tools come from, and --mcp-url is the whole of it:
--target | --mcp-url | Auth | |
|---|---|---|---|
| 1. Local Publisher | local | http://localhost:4040/mcp (default) | none |
| 2. Hosted, through an editor extension's local bridge | platform | the localhost URL the extension prints | none: the extension holds the credential |
| 3. Hosted, directly | platform | the host's https endpoint, scoped if it offers one | a cached OAuth login, once, interactively |
# 1. local
--target local # --mcp-url defaults to the local Publisher
# 2. hosted via the extension's bridge -- no OAuth, but check what it exposes:
# the same proxy may front a local Publisher instead
--target platform --mcp-url http://localhost:<port-the-extension-prints>/mcp \
--hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>
# 3. hosted directly -- authenticate first, under the SAME server name
claude mcp add --transport http <name> <scoped-url>
claude # /mcp -> <name> -> Authenticate
--target platform --mcp-url <scoped-url> \
--hosted-mcp-server <name> --target-version <v> --scope <env>/<pkg>@<v>All three hand the answerer the same three capabilities (get_context,
execute_query, and the docs search), so a comparison between them is between
agents that could do the same things; test_the_two_arms_hold_the_same_capabilities
pins it. Modes 2 and 3 take those names from --hosted-tools, which defaults to
the bare trio; pass it only if this host names them differently.
A platform run left on the local default --mcp-url is refused rather than
probed, because pointing every answerer at a local Publisher measures a
different model over different data than the run claims.
Two rules follow, and both are the kind of mistake that produces confident nonsense rather than an error:
Which target for which job:
So a measure-only run can use any target; a run that includes improve needs a local one.
Two things to check before a platform run, because neither errors and both make the run measure something other than what it names:
The answerer's skills must be written for THIS host. A shared skill names
an MCP tool by its bare name (get_context) so it reads correctly anywhere,
but a host/router skill names its own host's tools directly. Install the
latter for the wrong host and the answerer is told to call tools it does not
have. run_baseline.py warns when the manifest it loaded names Publisher-only
tools on a platform target; point --answerer-manifest, or --skills-root,
at the checkout that ships this host's manifest.
The tool names are configuration. --hosted-mcp-server is both the
mcp__<server>__<tool> prefix and the OAuth cache key, so it has to match the
name the answerer authenticated under, and --hosted-tools lists the bare
tools that host exposes.
Get the hosted tools in front of a headless answerer, one of two ways.
A spawned answerer cannot complete an OAuth flow, so the tools have to be
reachable before the run starts. run_baseline.py proves it with one cheap
probe and refuses to spend an arm otherwise -- a run whose answerers have no
tools does not error, it reads as a terrible model.
Authenticate once, interactively. Works anywhere, including a plain
CLI install, and is the route to assume unless you know otherwise. The
token is cached per server NAME, so authenticate under the same name the
run passes to --hosted-mcp-server:
claude mcp add --transport http <name> <scoped-url>
claude # then /mcp -> <name> -> AuthenticateThen come back and run. This is a hand-off to a person; there is no headless equivalent, so plan for it rather than discovering it mid-run.
A local proxy that already holds the credential. Some hosts ship an
editor extension whose local MCP proxy can expose the hosted
get_context / execute_query -- often behind a setting that is off by
default. Where that exists, point --mcp-url at the proxy on localhost
and no OAuth step is needed, because the extension holds it. Check what
the proxy actually exposes before relying on it: the same proxy may serve
a local Publisher's tools instead, and then --hosted-tools is
naming tools that are not there. This route is not available to someone
running the CLI alone.
Prefer a SCOPED endpoint URL over asking for scope. A hosted MCP is
usually reachable two ways: a global endpoint where every call carries an
organization and workspace, and a scoped one where the URL itself is the
scope. --scope and the prompt can only ASK an answerer to stay in one
package; a scoped URL enforces it. For an agent being measured that is the
difference between a case answered against the package it names and one
answered against whatever else the account can see. Authenticate once
interactively (claude, /mcp) under the same server name the run will use;
the token is cached per name, and a spawned headless answerer cannot complete
an OAuth flow.
Pin the VERSION in the scope, not just the package. Write
--scope <env>/<package>@<version>. Both hosted tools take a version and both
document the same default for an omitted one: the PINNED version, which is
whatever the workspace serves at the moment of the call. So an unversioned run
records targetVersion in run.json and then answers from whatever is
current, and the two part company the moment anyone publishes -- including
mid-run, which measures two builds under one label. --target-version fills
the version in when the scope omits it, so a platform run is pinned without
opting in; a scope naming a different version is refused rather than taken as
an override.
The model package under evaluation must live in a git repository, with
evals/<set>/ in the package, beside the model files. Git is the checkpoint
mechanism; without it there is no rollback and no run can include improve.
Keeping the set IN the package is what stops a model edit and its answer key
drifting apart: they move in one commit, so fixing a measure and forgetting
the golden that depended on it stops being possible. It is safe -- measured
on a running server, a cases.jsonl inside a package appears in no model
listing, no notebook listing, no package resource, and 404s over HTTP, so an
MCP-only answerer has no route to it.
What it buys differs by target. On a LOCAL Publisher it does not get you free
versioning -- sourceContentSha hashes model paths only, so the set needs
its own datasetSha. On a hosted target that publishes the whole package
directory as an IMMUTABLE version, the set rides inside that version and
targetVersion pins model and answer key together; nothing can be edited
under a published version, which is what makes it a pin. Check which you have
before deciding how much of this you need.
Look for a set and for prior runs before you author either. A minute of
find . -name cases.jsonl, a glance at the set's workdir (<workdir>/runs/,
~/.malloy-eval/<set>-<hash>/runs/ unless eval.toml moves it) and at your host's
own transcripts for this repo. Two sessions fourteen minutes apart built the
same 29-case answer key from scratch, because the first had committed
nothing before it was deleted and the second had no way to know it existed.
Roughly a working day was spent twice, and three specific things were
rediscovered at cost: the MCP login flow, the retrieval gate's 401, and a
judge-rendering bug that had already cost $17 of arm once. The two keys,
independently authored from the same questions, differ by about ten points
on comparable answers -- which is the available measure of how much a key
depends on its author, and a reason to reuse one rather than rebuild it.
Audit the set's entity ids before the first arm. verify_goldens.py
checks that each id names something in the model; check_findable.py checks
that a search of its own kind actually returns it. Both are free of model
calls. An id that fails either scores a retrieval miss on every run, and the
miss reads as the model's fault.
Commit the set before you spend money on an arm, and keep durable
outputs in the repository. A findings document in ~/Downloads is gone the
first time somebody tidies up; the set and the write-up belong in git
beside the model.
Runs go in the set's workdir, never inside the model package. A run
directory holds a model.malloy snapshot, and a built report is a Malloy
package; inside the package under test, either puts that package into
loadErrors. The default workdir, ~/.malloy-eval/<set>-<hash>/, is outside git,
which suits a measure-only run. A run that will improve and checkpoint needs
its ledger kept: set [paths] workdir in eval.toml to a directory in the
model's repository but outside the package, and gitignore its servers/
and packages/, which are a server's database and rebuilt reports.
The server must be up. Where your host offers retrieval tracing, turn it on and confirm a trace lookup answers, so a call's ranked results can be recovered afterwards.
Open-source Publisher has no trace lookup. PUBLISHER_MCP_TRACE=retrieval
is set by serve.py and makes Publisher write one "Retrieval trace" line per
ranked call to its log (stage counts and timings, not the ranked results).
Nothing reads that line yet: there is no trace store and no trace tool, so
traceId is null on every local attempt. This
does not block a scored run, because attribution never depended on it:
rankedSummary is copied onto the tool_call event at capture, precisely so
the evidence survives without a store. So do not refuse a local run for want
of tracing -- an earlier version of this rule did, and it refused every local
run there has ever been. Refuse one whose tool_call events carry no
rankedSummary, which is what attribution actually reads.
Health-check: your host's status check until it reports serving, and inspect
loadErrors. A dead database that still answers HTTP is an environment
failure, not a model failure. Stop and fix it. Four consecutive
environment or no-result attempts means stop the run.
Load the set: scrape/import as above, or reuse an existing evals/<set>/.
eval.py check --set <set> names every gap in it before anything starts.
Never keep two live copies of one set; the set directory in the model repo
is the single source of truth, versioned by datasetVersion in
set.json.
Review goldens before you score. Check how many cases can take a verdict
at all, not just how many cases there are: a golden the set stamps
provisional, invalid or ambiguous, or a case with no golden, scores
verdict: null and stays out of the pass rate. An imported set is
provisional throughout by design (skill:eval-import), and the only thing
that changes that is verify_goldens.py --promote after a re-derivation
through the truth package. A run whose every key is underived is refused
rather than spent; a set of bare questions runs, because its answers are
what keys get derived from. A verified golden that holds a value but
no local artifact stays verified by provenance and is not scorable until you
have rows or a scalar to compare (the judge needs both sides). A criteria
golden is the exception: it holds no value, so its clauses are both sides. If diagnosis later marks
BAD-REFERENCE or AMBIGUOUS-REFERENCE, follow
reference/golden-side-door.md. Both are expected in the wild; both
are the golden side door below, not improve, and not a sixth step.
eval.py run creates <workdir>/runs/<runId>/run.json with the attribution pins
(reference/ledger-schema.md): mode, dataset version, the target and the
version it served (a local target pins a commit, so commit or stash first;
answering from a dirty tree pins nothing), server version, judge version and
rubric sha, answerer model, call budget, trace mode. Freeze those for the
whole run. Raising a call budget mid-run moved mean outcomes on an unchanged
model.
The call budget is --max-turns, written to run.json as maxTurns.
Size it from a pilot rather than taking the default of 30. Run the three
cheapest cases uncapped (--max-turns 100) and set the cap at twice their
maximum. A set whose questions need two or three sources joined does not fit
a cap sized for single-source lookups: on one such set the completed
attempts had a median of 17 turns and a 90th percentile of 26 against a cap
of 30, and four cases died at it. A cap 15% above the 90th percentile of
completed work is not a safety margin. Note also that malloy-analysis tells
the answerer to persist -- retry a phrasing, let a small query settle whether
a field exists -- so a tight cap and that instruction are in direct conflict.
Do not change the model and the measuring instrument in the same step. Those are two different axes and only one of them is cheap to separate.
Batching MODEL edits is fine and expected. Clustering exists so that one
edit closes several cases, and an arm per fix does not survive contact with
arithmetic: 20 fixes over 100 cases is 2,000 answers, and at the measured
$0.33 a case on a proxied warehouse that is $660 of answering to attribute
what the acceptance check attributes for the price of the affected cases
plus holdout. Budget five arms for a defensible claim (a baseline, two for
the A/A, two post-edit), not one per edit, and let skill:eval-improve
carry each cluster.
The instrument is the other axis: the answer key, the judge and its prompt, the answerer model, the skills. Move one of those together with the model and there is nothing left holding still, so the result measures neither. On the run this comes from, nine model commits and a re-derived answer key landed between two arms, and the move from 20% to 39% belongs to no one -- not because two model edits were batched, but because the ruler changed at the same time as the thing being measured. A key repair mid-improve is not forbidden; it ends that comparison, so re-baseline rather than quoting a delta across it.
The answerer model is instrument too, and it is the one most often left unstated. The same 23-case set read 12 match / 11 near / 0 no_match on one model and 8 / 4 / 15 on a smaller one. No statement about "the agent's" capability means anything until two arms name the same answerer.
Generate every answerer prompt from the stored case in cases.jsonl.
Never retype the question. A truncated retype is indistinguishable from a
real question downstream.
malloy-analysis skill. Do not
mention eval, gold, scoring, or this skill.skill:eval-answer: contamination first, then re-execute, then the judge,
then events.skill:eval-diagnose only when this run includes diagnose, only on dev
cases, and only after the score event exists.skill:eval-improve only when this run includes improve, and only for
owner: model. Then run the acceptance check. On accept, checkpoint.A run directory is JSONL. It is a record, not a result, and the console summary
scrolls away. Every run ends by producing something a person can open, and
that job is skill:eval-report: it builds the servable run package (the case
matrix app and the aggregate notebook) and gives the template for the write-up.
Read it at the end of every run, including a run that failed. The standing complaint about this loop is that "a bunch of stuff happens and it is hard to know the actual results", and a ledger nobody renders is why.
Two rules from it are worth repeating here, because they are the ones a conductor skips:
improve not running because every
cluster came back owner: agent-skill is a RESULT, and it reads identically
to having forgotten unless it is written down.Keep the write-up in the repository beside the set, not in a chat log and not
in ~/Downloads.
A checkpoint is a git commit of the model repository, taken after an acceptance check accepts, so a bad improve direction can be rolled back. It is not a report, and it is not a remote publish.
checkpoint event (action: created, label, modelGitSha
from the commit you just made, issueIds). The event line itself rides in
the next commit; append-only logs trail by one commit and that is fine.git status is clean for the model files.Restore: git checkout <sha> -- <model files> (or git revert the
checkpoint commits), then reload the package, then append a checkpoint
event with action: restored and the sha. Readers return to the model that
existed before the bad direction.
Take a checkpoint of the current model before the first improve batch if no commit pins it yet. Rolling back by hand is guesswork.
If reload reports mode: reinstalled, the package was re-fetched from its
install location and may have overwritten the restored files. Prefer in-place
/ watch-mounted packages for this loop.
This loop is local. The ledger is files, the checkpoints are git, you are the conductor. Do not:
loop.py, run_all.py, improve_batch.py)
that runs the five steps end to end unattended. You conduct; the scripts are
the steps, not the sequencing. There are more than twenty of them and they
are not the exception to this: each does one step you invoke and hands back a
result you read. What is forbidden is a script that decides what to do next,
because every judgement this loop protects lives in that decisionstatus: incomplete,
and flip_table.py names what one arm left unscored that the other did not.
Read run_error on the attempts
and the INCOMPLETE line before quoting any number.skill:eval-answer: contamination, judge protocol, events. Its
reference/ledger-schema.md is the file contract; skill:eval-judge is
the judge.skill:eval-diagnose: component, owner, issue events. No edit.skill:eval-improve: smallest model edit, probe receipts, no self-accept.skill:eval-report: the run package and the write-up a person reads.malloy-analysis skill: what the blind answerer follows. It is installed
from the analysis manifest group, not the eval group.© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 34 other files (scripts) in skills/eval-loop of malloydata/publisher.
Open the folder on GitHubat commit acc1acd
Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Loop this skillmalloydata/publisher | 116 | — | ~7.8k | Automated safety check: Pass | MIT | |
| Tmuxtrpc-group/trpc-agent-go | 1.8k | 23 repos | ~868 | Automated safety check: Pass | Apache-2.0 | |
| Querying Indonesian Gov Datasuryast/indonesia-gov-apis | 172 | — | ~997 | Automated safety check: Pass | MIT | |
| Phoenix Release PleaseArize-ai/phoenix | 12k | — | ~708 | Automated safety check: Pass | Apache-2.0 | |
| Apify Actor Developmentsickn33/agentic-awesome-skills | 47k | 2 repos | ~3.3k | Automated safety check: Pass | MIT | |
| Daily News Reportaiskillstore/marketplace | 430 | 6 repos | ~2.8k | Automated safety check: Pass | None |
trpc-group/trpc-agent-go
Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.
suryast/indonesia-gov-apis
Query 57 Indonesian government APIs and data sources — BPJPH halal certification, BPOM food safety, OJK financial legality, BPS statistics, BMKG weather/earthquakes, Bank Indonesia exchange rates…
Arize-ai/phoenix
Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.
sickn33/agentic-awesome-skills
Important: Before you begin, fill in the generatedBy property in the meta section of .actor/actor.json.
aiskillstore/marketplace
Scrapes content based on a preset URL list, filters high-quality technical information, and generates daily Markdown reports.
foryourhealth111-pixel/Vibe-Skills
CLI-first web scraping & content extraction with optional MCP server.
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
malloydata/publisher
Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. Eval Loop is an agent skill from malloydata/publisher. Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
Eval Loop fits situations like: diagnose failures; improve behind an acceptance check; roll back a bad direction.
Run `npx skills add malloydata/publisher --skill eval-loop -a claude-code`. Or copy the skill folder (skills/eval-loop in malloydata/publisher) into .claude/skills/eval-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-loop -a codex`. Or copy the skill folder (skills/eval-loop in malloydata/publisher) into .agents/skills/eval-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-loop, .gemini/skills/eval-loop, .github/skills/eval-loop and .opencode/skills/eval-loop in your project.
Going by SKILL.md and its folder, Eval Loop needs Python for the scripts in its folder and the command-line tools its instructions call (git and claude). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.8k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Loop: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Querying Indonesian Gov Data (suryast/indonesia-gov-apis, 172 stars), Phoenix Release Please (Arize-ai/phoenix, 12k stars) and Apify Actor Development (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.