Eval Harness
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
$ npx skills add malloydata/publisher --skill eval-improve -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-improve --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-improve .claude/skills/eval-improve && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .claude/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-improveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-improve -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-improve --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-improve .agents/skills/eval-improve && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .agents/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-improve -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-improve --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-improve .cursor/skills/eval-improve && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .cursor/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-improve--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-improve -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-improve --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-improve .gemini/skills/eval-improve && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .gemini/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-improveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-improve -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-improve .github/skills/eval-improve && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .github/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-improve -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-improve --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-improve .opencode/skills/eval-improve && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-improve" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-improve into .opencode/skills/eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-improve", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-improveMake the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
Eval Improve is an agent skill from malloydata/publisher. Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim. Use after eval-diagnose, or when asked to fix a model so an agent can discover the right answer. Never accepts its own edit; the acceptance check belongs to eval-loop. Does not decide whether an answer was wrong (eval-answer) or why (eval-diagnose).
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `reference/output-contract.md`, `scripts/improve.py` and `scripts/improve_test.py`).
The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Improve loads about 2.8k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 1,627 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 1,627 words, ~2,799 tokens.
.claude/skills/eval-improve/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Takes an issue with owner: model and produces one smallest edit that closes
the gap. Every factual claim is backed by a query you ran.
Two hard boundaries:
skill:eval-loop admits or reverts. An improver writing the query it
already knows proves the fix is possible, not that the next blind agent
will find it.| Evidence | Edits permitted |
|---|---|
| Verified golden, or a user who states the answer | Any tier. Probes required. Check the golden first. |
| Wrong answer, then a corrected one the user accepted | Prefer docs over structure. The diff between attempts is the missing knowledge. |
| User accepted, later contradicted | Docs, labels, index only. No structural change. |
| Doubt only, or retrieval-only (no verdict) | Docs, labels, index only, and only where the transcript shows a concrete confusion. |
| Silence | No edit. |
Do not edit for BAD-REFERENCE or AMBIGUOUS-REFERENCE. Those are the
golden side door in skill:eval-loop: repair or hold the golden, bump
goldenRevision on the case, and open a new baseline run. Being right and unmatched
beats encoding a defect or an unsettled key. Do not edit for a skill,
retrieval, or dataset owner.
Encode what a domain expert would volunteer unprompted: systems of record, vocabulary to stored codes, what a metric means and at what grain, which relationship is the real one.
The expert test, per edit: would a domain expert have said this about their data with no question in front of them? Reject:
The golden is a hypothesis source. The data is still the verifier.
Every structural claim needs a query you ran: join key, primary key, filter value existence, snapshot assumption, value space, cardinality.
SELECT COUNT(*), COUNT(DISTINCT col) FROM t;
SELECT a.k, COUNT(*) FROM a JOIN b ON … GROUP BY 1 ORDER BY 2 DESC;
SELECT col, COUNT(*) FROM t GROUP BY 1 ORDER BY 2 DESC LIMIT 5;A false primary_key compiles and silently corrupts every aggregate. Of one
pilot's 11 accepted edits, 4 of 5 wrong ones died to a single
COUNT(*) vs COUNT(DISTINCT …) probe that was never run.
Compile-check the edit before saving (scope file for an edit), then reload
the package. Confirm it is not serving a stale model.
This step needs a target you control: a local server, or a host that can
execute a draft. A run whose answerers queried a published model cannot be
improved in place, because publishing to score an edit is not something this
loop does. skill:eval-loop picks the target before the run starts, so if you
have arrived here against a published target, stop and say so rather than
publishing.
Know which copy of the file the server actually reads. Hosts commonly serve a
copy of the package rather than your working tree, so editing the model repo
and reloading recompiles the unchanged copy: the reload succeeds, nothing
changes, and a verification probe quietly tests the old model. Confirm the
edit reached what is served before you trust a probe, and keep the model repo
the source of truth that gets committed. On open-source Publisher the served
copy lives under publisher_data/<env>/<pkg>/ unless the environment is
watch-mounted; other hosts distinguish a draft from a published version.
Prefer edits that add no entities. New sources compete for retrieval and displace answers that already worked.
| Rank | Edit |
|---|---|
| 1 | Disambiguating doc on confusable siblings: "use X for …, Y when …" |
| 2 | Named dimension or measure in user vocabulary |
| 3 | Doc reword, rename, or #(index) annotation |
| 4 | Declared join on a probed key |
| 5 | A new source: last resort, at most one |
Do not put an emphatic claim next to an unsettled one. A dimension doc was changed to read "The metric is the same either way: cumulative reach PERCENT OF TOTAL", to settle the shape of an output. On a question naming no metric the agent read it as settling which metric and silently picked one -- the exact failure its key forbade. The "there is no default, ask" note was on the source; the emphatic sentence was on the dimension, so an agent reading the dimension got the claim without the caveat. Put the caveat adjacent to the claim, not one level up, and say what the claim does not settle.
Do not leave two docs disagreeing about a default. In the same package a
view described itself as "DEFAULT for any unqualified X ask" while its own
source said there is no default and the agent must ask. Whichever the agent
reads first wins. After adding a "no default" note anywhere, grep the package
for the old claim -- ours survived in three places and was found by reading what
get_context returns, not the file.
Do not copy an expression into its own documentation. The tempting fix for
"the agent rebuilt this calculation wrong" is to paste the working expression
into the view's #(doc). That puts it in two places kept in sync by hand and
the doc goes stale silently. State the invariants instead -- what must stay true
at any parameterisation -- and name the view to reuse.
Do not introduce a first-of-its-kind construct for a handful of cases. A parameterized view was considered for three failures; the experimental flag was declared but zero of the package's 33 sources had ever taken a parameter. Generalising an existing view reached the same place without teaching everyone a new shape.
Make the correct thing the default. Guidance phrased as a caveat ("pair with X", "note that Y also includes Z") is retrieved, read, and declined. A source parameter or named measure that is already the safe scope does not invite a judgment call.
You cannot append guidance to every field for free. Doc length trades against the entity's own rank. A declared join is invisible to retrieval; put the rule on the entities agents search for.
Follow the malloy-gotchas-modeling skill so the edit does not introduce a
new modeling mistake. improve.py installs it with the rest of the modeling
manifest group, so it is loaded alongside this skill rather than reached from
here.
An edit that changes what a field means can silently invalidate goldens for
questions you were not working on. The rubric still describes the old meaning,
the stored value is still the old number, and nothing fails -- the case just
starts scoring wrong, against the model, in the direction of your edit. The set's
own value re-derivation will not catch it either, because it re-runs a
canonicalQuery that encodes the same stale definition.
Real instance: fixing lifetime_orders from line items to distinct orders was
correct and targeted. It also silently moved top_customer (defined over it)
from 108 customers to 87, and left two rubrics asserting the pre-fix behaviour.
Two correct answers were marked wrong for a full run before anyone noticed.
So before handing off, for every entity whose meaning you changed -- not every entity you touched; a doc reword changes no meaning:
canonicalQuery or
stated value that mentions it is now in question.improve.py
runs it for you after an edit and needs --truth-publisher to do it: the
goldens are re-derived against the TRUTH package, never against the model you
just changed, which would let an edit certify its own answer key. A verifier
that could not run at all is reported as a harness failure and never as a
suspect golden, because sending someone to settle a key that is fine is the
most expensive wrong turn this step can cause.Report every hit as golden_suspect in the handoff, with the entity, the case,
and the old and new values. Do not repair them yourself. Goldens are the
side door in skill:eval-loop, and an improver that edits the answer key its own
edit is scored against has removed the only independent check on the edit.
A non-empty golden_suspect list blocks the acceptance check until the
conductor settles
each one, because a rerun against stale goldens measures nothing.
Compile, reload, run one trivial query against each source you touched.
Append a candidate event to the run's events.jsonl (shape in
skill:eval-answer reference/ledger-schema.md): the files touched, a
one-line diff summary per file, the issue_ids, probe receipts, and this
report. Every proposal gets its event, accepted or not; a rejected direction
keeps its record. Then stop and wait for the acceptance check in skill:eval-loop. That
skill writes the acceptance_check event, accepts or reverts, and only on accept
checkpoints. This skill never checkpoints and never self-accepts.
COMPONENT / PRIMARY_CODE / OWNER
EVIDENCE: class you worked from, and how it limited the edit
DISAGREEMENT: NONE, or anything in the diagnosis probing showed was wrong
DIAGNOSIS: 2-3 sentences, written before the edit
EDIT: one line, or NONE with why
EXPERT-TEST: the business fact this encodes
PROBES: each probe query and its result
GOLDEN-SUSPECT: NONE, or one line per case: qid, entity, stored -> re-derivedDISAGREEMENT is load-bearing. An improver that cannot push back encodes
its instructions' mistakes. Report from the files on disk, not from memory.
skill:eval-diagnose: the issue this requires.skill:eval-loop: the acceptance check that accepts or reverts, then
checkpoints on accept. Golden hold/repair lives there, not here.skill:eval-answer: scoring after a blind re-answer.malloy-gotchas-modeling skill: mistakes an edit must not introduce.
It arrives with the modeling manifest group, not the eval group.© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/eval-improve of malloydata/publisher.
Open the folder on GitHubat commit acc1acd
Eval Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Improve this skillmalloydata/publisher | 116 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 274k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | 1 repos | ~618 | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 274k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Eval-Driven Development Harnessaffaan-m/ECC | 274k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Skill Eval ImproveArenukvern/mcp_flutter | 385 | — | ~2.4k | Automated safety check: Pass | MIT |
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
Arenukvern/mcp_flutter
Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Never use eval() or unsafe dynamic code execution.
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
malloydata/publisher
Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim. Eval Improve is an agent skill from malloydata/publisher. Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
Run `npx skills add malloydata/publisher --skill eval-improve -a claude-code`. Or copy the skill folder (skills/eval-improve in malloydata/publisher) into .claude/skills/eval-improve in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-improve -a codex`. Or copy the skill folder (skills/eval-improve in malloydata/publisher) into .agents/skills/eval-improve in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-improve, .gemini/skills/eval-improve, .github/skills/eval-improve and .opencode/skills/eval-improve in your project.
Going by SKILL.md and its folder, Eval Improve needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Improve is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Improve: Eval Harness (affaan-m/ECC, 274k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 274k stars) and Eval-Driven Development Harness (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.