Agent skill

Eval Improve

by malloydata in malloydata/publisher

Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

MITAuto-check passed

Install Eval Improve

skills CLI
$ npx skills add malloydata/publisher --skill eval-improve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-improve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-improve .claude/skills/eval-improve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-improve
GitHub stars
116
Token cost
~2.8k tokens
SKILL.md length
1,627 words
Files
4 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

  • Works in 6 steps: What the evidence entitles you to change → What a correct answer may teach → Probe receipts → …
  • SKILL.md covers Step 0: What the evidence…, Step 1: What a correct answer…, Step 2: Probe receipts and Step 3: One smallest edit, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Eval Improve is an agent skill from malloydata/publisher. Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim. Use after eval-diagnose, or when asked to fix a model so an agent can discover the right answer. Never accepts its own edit; the acceptance check belongs to eval-loop. Does not decide whether an answer was wrong (eval-answer) or why (eval-diagnose).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `reference/output-contract.md`, `scripts/improve.py` and `scripts/improve_test.py`).

The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

Example prompts

  • “/eval-improve”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. What the evidence entitles you to change
  2. What a correct answer may teach
  3. Probe receipts
  4. One smallest edit
  5. Check what your edit did to the answer key
  6. Verify, report, hand off

What it can do on your machine

Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Improve loads about 2.8k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 1,627 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 1,627 words, ~2,799 tokens.

Download SKILL.mdSave it as .claude/skills/eval-improve/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
eval-improve
description
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim. Use after eval-diagnose, or when asked to fix a model so an agent can discover the right answer. Never accepts its own edit; the acceptance check belongs to eval-loop. Does not decide whether an answer was wrong (eval-answer) or why (eval-diagnose).

Improve the Model

Takes an issue with owner: model and produces one smallest edit that closes the gap. Every factual claim is backed by a query you ran.

Two hard boundaries:

  1. No diagnosis evidence, no edit. If the issue cannot name a concrete gap with a trace or probe, record that and stop. Edits from an empty diagnosis have been the inert and wrong ones.
  2. This skill never accepts its own edit. You propose and verify. The acceptance check in skill:eval-loop admits or reverts. An improver writing the query it already knows proves the fix is possible, not that the next blind agent will find it.

Step 0: What the evidence entitles you to change

EvidenceEdits permitted
Verified golden, or a user who states the answerAny tier. Probes required. Check the golden first.
Wrong answer, then a corrected one the user acceptedPrefer docs over structure. The diff between attempts is the missing knowledge.
User accepted, later contradictedDocs, labels, index only. No structural change.
Doubt only, or retrieval-only (no verdict)Docs, labels, index only, and only where the transcript shows a concrete confusion.
SilenceNo edit.

Do not edit for BAD-REFERENCE or AMBIGUOUS-REFERENCE. Those are the golden side door in skill:eval-loop: repair or hold the golden, bump goldenRevision on the case, and open a new baseline run. Being right and unmatched beats encoding a defect or an unsettled key. Do not edit for a skill, retrieval, or dataset owner.

Step 1: What a correct answer may teach

Encode what a domain expert would volunteer unprompted: systems of record, vocabulary to stored codes, what a metric means and at what grain, which relationship is the real one.

The expert test, per edit: would a domain expert have said this about their data with no question in front of them? Reject:

  • a field that hard-codes this question's filter and serves no other question
  • this question's text, qid, or expected numbers in a doc, comment, or name
  • a join copied from gold SQL that you have not probed as a real relationship

The golden is a hypothesis source. The data is still the verifier.

Step 2: Probe receipts

Every structural claim needs a query you ran: join key, primary key, filter value existence, snapshot assumption, value space, cardinality.

sql
SELECT COUNT(*), COUNT(DISTINCT col) FROM t;
SELECT a.k, COUNT(*) FROM a JOIN b ON … GROUP BY 1 ORDER BY 2 DESC;
SELECT col, COUNT(*) FROM t GROUP BY 1 ORDER BY 2 DESC LIMIT 5;

A false primary_key compiles and silently corrupts every aggregate. Of one pilot's 11 accepted edits, 4 of 5 wrong ones died to a single COUNT(*) vs COUNT(DISTINCT …) probe that was never run.

Compile-check the edit before saving (scope file for an edit), then reload the package. Confirm it is not serving a stale model.

This step needs a target you control: a local server, or a host that can execute a draft. A run whose answerers queried a published model cannot be improved in place, because publishing to score an edit is not something this loop does. skill:eval-loop picks the target before the run starts, so if you have arrived here against a published target, stop and say so rather than publishing.

Know which copy of the file the server actually reads. Hosts commonly serve a copy of the package rather than your working tree, so editing the model repo and reloading recompiles the unchanged copy: the reload succeeds, nothing changes, and a verification probe quietly tests the old model. Confirm the edit reached what is served before you trust a probe, and keep the model repo the source of truth that gets committed. On open-source Publisher the served copy lives under publisher_data/<env>/<pkg>/ unless the environment is watch-mounted; other hosts distinguish a draft from a published version.

Step 3: One smallest edit

Prefer edits that add no entities. New sources compete for retrieval and displace answers that already worked.

RankEdit
1Disambiguating doc on confusable siblings: "use X for …, Y when …"
2Named dimension or measure in user vocabulary
3Doc reword, rename, or #(index) annotation
4Declared join on a probed key
5A new source: last resort, at most one
Four ways a doc edit backfires, each measured

Do not put an emphatic claim next to an unsettled one. A dimension doc was changed to read "The metric is the same either way: cumulative reach PERCENT OF TOTAL", to settle the shape of an output. On a question naming no metric the agent read it as settling which metric and silently picked one -- the exact failure its key forbade. The "there is no default, ask" note was on the source; the emphatic sentence was on the dimension, so an agent reading the dimension got the claim without the caveat. Put the caveat adjacent to the claim, not one level up, and say what the claim does not settle.

Do not leave two docs disagreeing about a default. In the same package a view described itself as "DEFAULT for any unqualified X ask" while its own source said there is no default and the agent must ask. Whichever the agent reads first wins. After adding a "no default" note anywhere, grep the package for the old claim -- ours survived in three places and was found by reading what get_context returns, not the file.

Do not copy an expression into its own documentation. The tempting fix for "the agent rebuilt this calculation wrong" is to paste the working expression into the view's #(doc). That puts it in two places kept in sync by hand and the doc goes stale silently. State the invariants instead -- what must stay true at any parameterisation -- and name the view to reuse.

Do not introduce a first-of-its-kind construct for a handful of cases. A parameterized view was considered for three failures; the experimental flag was declared but zero of the package's 33 sources had ever taken a parameter. Generalising an existing view reached the same place without teaching everyone a new shape.

Make the correct thing the default. Guidance phrased as a caveat ("pair with X", "note that Y also includes Z") is retrieved, read, and declined. A source parameter or named measure that is already the safe scope does not invite a judgment call.

You cannot append guidance to every field for free. Doc length trades against the entity's own rank. A declared join is invisible to retrieval; put the rule on the entities agents search for.

Follow the malloy-gotchas-modeling skill so the edit does not introduce a new modeling mistake. improve.py installs it with the rest of the modeling manifest group, so it is loaded alongside this skill rather than reached from here.

Show full SKILL.md (546 more words)Show less

Step 4: Check what your edit did to the answer key

An edit that changes what a field means can silently invalidate goldens for questions you were not working on. The rubric still describes the old meaning, the stored value is still the old number, and nothing fails -- the case just starts scoring wrong, against the model, in the direction of your edit. The set's own value re-derivation will not catch it either, because it re-runs a canonicalQuery that encodes the same stale definition.

Real instance: fixing lifetime_orders from line items to distinct orders was correct and targeted. It also silently moved top_customer (defined over it) from 108 customers to 87, and left two rubrics asserting the pre-fix behaviour. Two correct answers were marked wrong for a full run before anyone noticed.

So before handing off, for every entity whose meaning you changed -- not every entity you touched; a doc reword changes no meaning:

  1. Grep the case file for the entity name. Any rubric, canonicalQuery or stated value that mentions it is now in question.
  2. For each hit, re-derive the value under the new definition and compare it to the stored golden. Different means the golden is stale, not that you are wrong.
  3. Run the set's golden verification if it has one, which catches the mechanical subset (drift, and rubric sentences that contradict the model). improve.py runs it for you after an edit and needs --truth-publisher to do it: the goldens are re-derived against the TRUTH package, never against the model you just changed, which would let an edit certify its own answer key. A verifier that could not run at all is reported as a harness failure and never as a suspect golden, because sending someone to settle a key that is fine is the most expensive wrong turn this step can cause.

Report every hit as golden_suspect in the handoff, with the entity, the case, and the old and new values. Do not repair them yourself. Goldens are the side door in skill:eval-loop, and an improver that edits the answer key its own edit is scored against has removed the only independent check on the edit.

A non-empty golden_suspect list blocks the acceptance check until the conductor settles each one, because a rerun against stale goldens measures nothing.

Step 5: Verify, report, hand off

Compile, reload, run one trivial query against each source you touched. Append a candidate event to the run's events.jsonl (shape in skill:eval-answer reference/ledger-schema.md): the files touched, a one-line diff summary per file, the issue_ids, probe receipts, and this report. Every proposal gets its event, accepted or not; a rejected direction keeps its record. Then stop and wait for the acceptance check in skill:eval-loop. That skill writes the acceptance_check event, accepts or reverts, and only on accept checkpoints. This skill never checkpoints and never self-accepts.

COMPONENT / PRIMARY_CODE / OWNER
EVIDENCE:    class you worked from, and how it limited the edit
DISAGREEMENT: NONE, or anything in the diagnosis probing showed was wrong
DIAGNOSIS:   2-3 sentences, written before the edit
EDIT:        one line, or NONE with why
EXPERT-TEST: the business fact this encodes
PROBES:      each probe query and its result
GOLDEN-SUSPECT: NONE, or one line per case: qid, entity, stored -> re-derived

DISAGREEMENT is load-bearing. An improver that cannot push back encodes its instructions' mistakes. Report from the files on disk, not from memory.

  • skill:eval-diagnose: the issue this requires.
  • skill:eval-loop: the acceptance check that accepts or reverts, then checkpoints on accept. Golden hold/repair lives there, not here.
  • skill:eval-answer: scoring after a blind re-answer.
  • The malloy-gotchas-modeling skill: mistakes an edit must not introduce. It arrives with the modeling manifest group, not the eval group.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/eval-improve of malloydata/publisher.

  • SKILL.md
  • reference/output-contract.md
  • scripts/improve.py
  • scripts/improve_test.py

Open the folder on GitHubat commit acc1acd

Compare with similar skills

Eval Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Improve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Improve this skillmalloydata/publisher116—~2.8kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC274k—~2.2kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC274k1 repos~1.7kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC274k—~1.5kAutomated safety check: PassMIT
Skill Eval ImproveArenukvern/mcp_flutter385—~2.4kAutomated safety check: PassMIT

Similar skills

  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    274k GitHub stars~2.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    274k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    274k GitHub stars~1.5k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Skill Eval Improve

    Arenukvern/mcp_flutter

    Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

    385 GitHub stars~2.4k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Avoid Eval

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Never use eval() or unsafe dynamic code execution.

    74k GitHub stars~554 tokensUpdated yesterday
    Auto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Malloy Analysis Report

    malloydata/publisher

    Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.

    116 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Questions about Eval Improve

What does Eval Improve do?

Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim. Eval Improve is an agent skill from malloydata/publisher. Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

How do I install Eval Improve in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-improve -a claude-code`. Or copy the skill folder (skills/eval-improve in malloydata/publisher) into .claude/skills/eval-improve in your project. Claude Code loads it when a task matches its description.

How do I install Eval Improve in Codex?

Run `npx skills add malloydata/publisher --skill eval-improve -a codex`. Or copy the skill folder (skills/eval-improve in malloydata/publisher) into .agents/skills/eval-improve in your project. Codex loads it when a task matches its description.

Can I use Eval Improve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-improve, .gemini/skills/eval-improve, .github/skills/eval-improve and .opencode/skills/eval-improve in your project.

What does Eval Improve need to run?

Going by SKILL.md and its folder, Eval Improve needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Eval Improve access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Improve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Improve use?

Eval Improve is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Improve use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Improve?

Skills that share tags, products or a category with Eval Improve: Eval Harness (affaan-m/ECC, 274k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 274k stars) and Eval-Driven Development Harness (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Improve?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.