Agent skill

Eval Judge

by malloydata in malloydata/publisher

Decide whether ONE answer matches its golden, and say whether you believe the golden.

MITAuto-check passedEducation

Install Eval Judge

skills CLI
$ npx skills add malloydata/publisher --skill eval-judge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-judge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-judge .claude/skills/eval-judge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-judge
GitHub stars
116
Token cost
~3.4k tokens
SKILL.md length
2,132 words
Files
5
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Decide whether ONE answer matches its golden, and say whether you believe the golden.

  • Works in 11 steps: Judge intent, not formatting. The… → Gold-subset containment. The prediction… → Name the column pairing. Pair each gold… → …
  • Scoring an attempt in an evaluation run
  • SKILL.md covers Read one of these before you…, Answer judge and Versioning and regressions
  • Calls git

What it does

Eval Judge is an agent skill from malloydata/publisher. Decide whether ONE answer matches its golden, and say whether you believe the golden. Read this before emitting any verdict. Covers containment, column pairing, nearmatch, refusals, and the goldstatus judgement. Use when scoring an attempt in an evaluation run; never to conduct a run (eval-loop), diagnose a failure (eval-diagnose) or edit a model (eval-improve).

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `reference/refusal.md`, `reference/retrieval-judge.md` and `reference/suspect-goldens.md`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Scoring an attempt in an evaluation run
  • Never to conduct a run (eval-loop)
  • Diagnose a failure (eval-diagnose)
  • Edit a model (eval-improve)

Example prompts

  • “/eval-judge”

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. Judge intent, not formatting. The question defines what counts. A
  2. Gold-subset containment. The prediction must CONTAIN the gold answer.
  3. Name the column pairing. Pair each gold column with the prediction
  4. Rows are a multiset. Order matters only when the question asks for an
  5. Tolerances. Numeric equality within small rounding (relative 1e-6, or
  6. Confidence 1 to 10. 5 or lower means the case needs a human
  7. near_match is not a soft pass, and it is not a soft fail. It is a
  8. On a large row set, compare it as a set rather than scanning pairwise: state
  9. A mustNotUse field is not yours to weigh, unless it is prose. A
  10. Score the data, not the insight. A question that asks for a figure or
  11. Do not demand a grain the question did not fix. When the question names

What it can do on your machine

Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Judge loads about 3.4k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 2,132 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 2,132 words, ~3,448 tokens.

Download SKILL.mdSave it as .claude/skills/eval-judge/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
eval-judge
description
Decide whether ONE answer matches its golden, and say whether you believe the golden. Read this before emitting any verdict. Covers containment, column pairing, near_match, refusals, and the gold_status judgement. Use when scoring an attempt in an evaluation run; never to conduct a run (eval-loop), diagnose a failure (eval-diagnose) or edit a model (eval-improve).

The judge

JUDGE_VERSION: 6

This skill IS the judge. One fresh judge subagent is spawned per attempt, with this skill installed in its workspace and the case materials in its prompt. It is loaded, not pasted -- so the prompt carries the case and this carries the doctrine, and a judge that needs to read a Malloy query can reach for the skills beside it rather than being handed a transcription.

Measured when it stopped being pasted, on the case that had oscillated (a valued golden against a model with no trace of the concept):

pasted into the prompt   match / no_match / match / match
loaded as this skill     no_match x4, and the reasoning cites the rule

It costs about 2.5x per verdict, which is the price of the judge actually reading its own rules.

Record judge_version and this file's git blob sha (git rev-parse HEAD:skills/eval-judge/SKILL.md, or the model repo's copy) on every verdict, so a rubric change never silently rewrites what old scores meant.

The judge is not blind. It sees the golden. It must never be the same subagent that answered, and it never edits anything: it returns a verdict object and stops.

Read one of these before you decide

This file is the decision procedure. Four situations have their own rules, and each is a file beside this one. Read the file BEFORE emitting a verdict, not after -- these are the cases where judging from the general rubric alone gets it wrong, which is why they are called out rather than summarised.

IfRead
the answer declines, or gives no value at allreference/refusal.md
the golden itself looks wrong to youreference/suspect-goldens.md
you are judging retrieval, not an answerreference/retrieval-judge.md
you are AUTHORING a case rather than judging onereference/writing-rubrics.md

The first row is the one that catches people. A refusal is only exempt from containment when golden.kind is unanswerable; against a golden that holds a value, an answer containing none of it is no_match however well it reasons. reference/refusal.md is the whole rule.

A third kind holds no value: criteria, where the case's rubric IS the key and there is no number to contain. Grade the clauses and nothing else. Do not manufacture a figure to check the answer against, and do not read the absence of a value as a missing golden: a criteria golden is complete. Report gold_status on it the same way, on the criteria rather than on a number, so a clause that contradicts the model still surfaces.

Answer judge

Input, all of it (a judge with only two row sets grades formatting, not intent):

  • the question, exactly as the answerer saw it
  • the golden: rows or scalar, plus canonicalQuery when present
  • the prediction: the rows the CONDUCTOR re-executed from the answerer's final_query (never the answerer's self-reported rows)
  • the relevant source and field definitions from the model (docs, join list)

Output, exactly this shape:

json
{
  "verdict": "match | near_match | no_match",
  "confidence": 7,
  "why": "one short paragraph",
  "column_pairing": { "gold_col": "pred_col", ... },
  "gold_status": "verified | verified_benign | suspect | verified_wrong",
  "gold_note": "why, when not verified"
}
Rubric
  1. Judge intent, not formatting. The question defines what counts. A result that answers the question in a different but faithful shape is a match.

  2. Gold-subset containment. The prediction must CONTAIN the gold answer. Extra columns or benign extra context downgrade to near_match at worst; they never make a containing answer no_match.

  3. Name the column pairing. Pair each gold column with the prediction column that carries the same meaning, using names, the question's role for the value, and the values together. Never pair numeric columns by value overlap alone: a year column is not a count column even when magnitudes overlap. If a gold column has no counterpart, say which.

  4. Rows are a multiset. Order matters only when the question asks for an order. For a "top N" with possible ties, check that the boundary value is right and every returned row legitimately qualifies; any valid tie-break is a match.

  5. Tolerances. Numeric equality within small rounding (relative 1e-6, or the display precision the golden uses). A percentage and its fraction (50 and 0.5) are the same value in different units when the pairing says the column is a rate.

  6. Confidence 1 to 10. 5 or lower means the case needs a human: the conductor records needs_human, which is neither a pass nor a fail. Do not inflate confidence to be helpful; a wrong confident verdict is worse than an abstention.

  7. near_match is not a soft pass, and it is not a soft fail. It is a third outcome meaning defensibly different: the answer took a reading the rubric allows but did not prefer, broke a tie the other way, or buried a caveat that should have been plain. It is excluded from the pass rate and from the acceptance check, exactly like needs_human.

    So do not reach for it to avoid a hard call. If the prediction contains the gold answer, that is match -- extra columns and benign extra context never reduce it (rule 2). If it does not, and the rubric does not sanction the reading that produced it, that is no_match. Use near_match only when you can name the rubric clause that makes the difference defensible.

    A rubric clause cannot make a wrong VALUE defensible, and a clause that tries is a defect in the rubric rather than a licence to you. near_match turns on the answer being right under a reading the QUESTION allows -- a tie broken the other way, a grain the question left open, a basis the question never fixed. It does not turn on the answer being transparent about how it got a figure the question did not ask for. Those two look alike in a rubric and are opposites in a report: one is a number a reader can act on, the other is a number a reader would act on wrongly. Naming the method makes a wrong figure DIAGNOSABLE, which is worth having, and it is not partial credit.

    The test, before you write near_match on a case with a value: would a reader who acted on this figure be wrong? If yes, it is no_match however plainly the answer explained itself, and however the rubric is worded. Say in why that you are overriding a rubric clause, so the clause gets fixed. This rule exists because a set shipped one: a question asked for sales over the company's season, the answer gave the meteorological window 25% lower and said which window it used, and a clause granting near_match for a stated window kept a materially wrong answer out of the pass rate entirely -- the arm reported 100%.

    It is a third outcome because as a pass it was a large share of the measured noise: the same unchanged answer reads match in one run and near_match in the next, and the pass rate moves although nothing did. A verdict whose content is "this is arguable" cannot be allowed to decide anything. Its count is still reported, and a rising one means the rubrics are going vague. (What that share was for a given set is in that set's CALIBRATION.md.)

    A near_match that lands the same way in two arms is a different animal from one that flickers. Stable across a pair, it is not judge noise: the model cannot distinguish two readings the question does, which is a coverage finding, and softening the rubric will not close it. flip_table.py lists the stable ones and diagnose.py --verdicts near_match takes them.

  8. On a large row set, compare it as a set rather than scanning pairwise: state how many gold rows you located in the prediction, name the ones you could not, and say what the mismatched values look like (uniformly scaled, off in one column, a different population). "I checked all 76" without that breakdown is not a comparison.

  9. A mustNotUse field is not yours to weigh, unless it is prose. A script checks the final query for the field names golden.mustNotUse lists and forces no_match on a hit before you are asked, so a case that reaches you with a MUST NOT USE line is carrying only what a text check could not decide: a reading described in words, an objection to a USE of a field rather than to the field (X as ..., X through ...), or a bare field name that may or may not be the forbidden one. Apply those as the rubric's own clauses. Do not soften a verdict because a veto might have caught it, and do not invent a veto the rubric did not ask for.

  10. Score the data, not the insight. A question that asks for a figure or a series is judged on the figure or the series. Where the question also asks for an interpretation -- "when did it flatten out", "what drove the change" -- that interpretation is not scored unless the rubric marks it REQUIRED with a criterion that resolves from the data alone. Two analysts reading the same exact curve name different weeks; an eval that scores which week they named is measuring taste, and a run that lost a case that way (13 of 13 weekly values exact, plateau named one week outside a window) was measuring nothing. Exact data with a different reading of it is match.

  11. Do not demand a grain the question did not fix. When the question names no grain -- by medium, by week, campaign total -- a figure that is correct at the grain the answer states is correct. The golden's grain is PREFERRED, not the only one: an answer at another grain is match when the grain is stated and the figures are right at it; near_match when the grain is left unstated; no_match only when the figures are wrong at the grain claimed. An answer that named the right segment and showed the index split by medium, every number right, was once scored down for not showing the campaign total; the question had never asked for one. A rubric that means "campaign total only" must say so as REQUIRED, and the question should say so too.

Show full SKILL.md (472 more words)Show less
Rubric markers

A case rubric marks its alternate readings and disclosures with the words below, and each word fixes the verdict. Apply them as written; do not re-weigh a reading the rubric has already classified. (reference/writing-rubrics.md is where authors are told to use them; this table is the judge's half.)

MarkerVerdictMeaning
PREFERREDmatchThe reading the golden encodes.
ACCEPTmatchEqually right: a different but faithful route to the same claim. Check the figure against the golden the way the clause says to.
DIVERGENTnear_matchDefensible and not what was asked for. Never no_match, however clearly the answer committed to it.
WRONGno_matchPlausible and incorrect; the clause usually names the trap.
REQUIREDomitted: no_matchA disclosure without which the number misleads.
CREDITEDomitted: matchContext a good analyst adds; its absence costs nothing.

Measured on an unchanged answer, rubric and golden: an answer whose recommended figure the rubric marked DIVERGENT scored near_match under one judge and no_match under the next, because the judge had never been told what the word meant and weighed the commitment instead. An unmarked clause is CREDITED (the author's bug, not yours to repair by inventing a requirement).

Anchors
  • match: question "total sales by category"; golden 8 rows (category, revenue); prediction 8 rows (product_category, gross_revenue, order_count). Same categories, revenues equal within rounding; the extra count column does not change what the answer says. Verdict: match, confidence 9.
  • near_match: question "top 5 states by returns"; golden and prediction agree on 4 of 5 states, and the disagreement is at rank 5 where two states tie exactly; the prediction chose the other tie-break. The boundary value is right, the membership defensible, but the golden pinned one tie-break. Verdict: near_match, confidence 7, why names the tie.
  • no_match: question "revenue in 2024, completed orders only"; golden 1.2M; prediction 1.9M and the pairing shows the prediction summed all statuses. Same shape, wrong population. Verdict: no_match, confidence 9.

Keep the anchor set balanced. A judge shown only matches learns a base rate, not a rubric.

Coverage

A case may be labelled coverage: derivable: the model has no entity for the concept and the answer had to be built from the parts that exist. Judge the result exactly as the rubric says -- a derived answer that matches the golden is a match, and the absence of a named measure is not a deduction. But when the answer states what it built, say so in the why. That sentence is what tells diagnosis the gap is real and lets coverage_note become a model edit rather than a guess.

Versioning and regressions

Any change to this file is a judge change: bump JUDGE_VERSION, commit, and re-run evals/<set>/judge-regressions.jsonl (the human-overruled verdicts) before trusting new scores. Runs record judge_version and rubric_sha, so a delta across a rubric change is attributable to the rubric, not the model.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/eval-judge of malloydata/publisher.

  • SKILL.md
  • reference/refusal.md
  • reference/retrieval-judge.md
  • reference/suspect-goldens.md
  • reference/writing-rubrics.md

Open the folder on GitHubat commit 39a546f

Compare with similar skills

Eval Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Judge this skillmalloydata/publisher116—~3.4kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated yesterday
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Malloy Analysis Report

    malloydata/publisher

    Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.

    116 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Eval Judge

What does Eval Judge do?

Decide whether ONE answer matches its golden, and say whether you believe the golden. Eval Judge is an agent skill from malloydata/publisher. Decide whether ONE answer matches its golden, and say whether you believe the golden.

When should I use Eval Judge?

Eval Judge fits situations like: scoring an attempt in an evaluation run; never to conduct a run (eval-loop); diagnose a failure (eval-diagnose); edit a model (eval-improve).

How do I install Eval Judge in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-judge -a claude-code`. Or copy the skill folder (skills/eval-judge in malloydata/publisher) into .claude/skills/eval-judge in your project. Claude Code loads it when a task matches its description.

How do I install Eval Judge in Codex?

Run `npx skills add malloydata/publisher --skill eval-judge -a codex`. Or copy the skill folder (skills/eval-judge in malloydata/publisher) into .agents/skills/eval-judge in your project. Codex loads it when a task matches its description.

Can I use Eval Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-judge, .gemini/skills/eval-judge, .github/skills/eval-judge and .opencode/skills/eval-judge in your project.

What does Eval Judge need to run?

Going by SKILL.md and its folder, Eval Judge needs the command-line tools its instructions call (git).

Does Eval Judge access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Judge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Judge use?

Eval Judge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Judge use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Judge?

Skills that share tags, products or a category with Eval Judge: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Judge?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.