DeepTutor CLI
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
Decide whether ONE answer matches its golden, and say whether you believe the golden.
$ npx skills add malloydata/publisher --skill eval-judge -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-judge --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-judge .claude/skills/eval-judge && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .claude/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-judgeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-judge -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-judge --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-judge .agents/skills/eval-judge && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .agents/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-judge -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-judge --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-judge .cursor/skills/eval-judge && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .cursor/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-judge--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-judge -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-judge --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-judge .gemini/skills/eval-judge && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .gemini/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-judgeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-judge -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-judge .github/skills/eval-judge && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .github/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-judge -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-judge --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-judge .opencode/skills/eval-judge && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-judge" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-judge into .opencode/skills/eval-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-judge", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-judgeDecide whether ONE answer matches its golden, and say whether you believe the golden.
Eval Judge is an agent skill from malloydata/publisher. Decide whether ONE answer matches its golden, and say whether you believe the golden. Read this before emitting any verdict. Covers containment, column pairing, nearmatch, refusals, and the goldstatus judgement. Use when scoring an attempt in an evaluation run; never to conduct a run (eval-loop), diagnose a failure (eval-diagnose) or edit a model (eval-improve).
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `reference/refusal.md`, `reference/retrieval-judge.md` and `reference/suspect-goldens.md`).
It sits in Education, covering Quizzes and assessments. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
11 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Judge loads about 3.4k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 2,132 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 2,132 words, ~3,448 tokens.
.claude/skills/eval-judge/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.JUDGE_VERSION: 6
This skill IS the judge. One fresh judge subagent is spawned per attempt, with this skill installed in its workspace and the case materials in its prompt. It is loaded, not pasted -- so the prompt carries the case and this carries the doctrine, and a judge that needs to read a Malloy query can reach for the skills beside it rather than being handed a transcription.
Measured when it stopped being pasted, on the case that had oscillated (a valued golden against a model with no trace of the concept):
pasted into the prompt match / no_match / match / match
loaded as this skill no_match x4, and the reasoning cites the ruleIt costs about 2.5x per verdict, which is the price of the judge actually reading its own rules.
Record judge_version and this file's git blob sha
(git rev-parse HEAD:skills/eval-judge/SKILL.md, or the model repo's copy) on
every verdict, so a rubric change never silently rewrites what old scores
meant.
The judge is not blind. It sees the golden. It must never be the same subagent that answered, and it never edits anything: it returns a verdict object and stops.
This file is the decision procedure. Four situations have their own rules, and each is a file beside this one. Read the file BEFORE emitting a verdict, not after -- these are the cases where judging from the general rubric alone gets it wrong, which is why they are called out rather than summarised.
| If | Read |
|---|---|
| the answer declines, or gives no value at all | reference/refusal.md |
| the golden itself looks wrong to you | reference/suspect-goldens.md |
| you are judging retrieval, not an answer | reference/retrieval-judge.md |
| you are AUTHORING a case rather than judging one | reference/writing-rubrics.md |
The first row is the one that catches people. A refusal is only exempt from
containment when golden.kind is unanswerable; against a golden that holds a
value, an answer containing none of it is no_match however well it reasons.
reference/refusal.md is the whole rule.
A third kind holds no value: criteria, where the case's rubric IS the key
and there is no number to contain. Grade the clauses and nothing else. Do not
manufacture a figure to check the answer against, and do not read the absence
of a value as a missing golden: a criteria golden is complete. Report
gold_status on it the same way, on the criteria rather than on a number, so
a clause that contradicts the model still surfaces.
Input, all of it (a judge with only two row sets grades formatting, not intent):
canonicalQuery when presentfinal_query (never the answerer's self-reported rows)Output, exactly this shape:
{
"verdict": "match | near_match | no_match",
"confidence": 7,
"why": "one short paragraph",
"column_pairing": { "gold_col": "pred_col", ... },
"gold_status": "verified | verified_benign | suspect | verified_wrong",
"gold_note": "why, when not verified"
}Judge intent, not formatting. The question defines what counts. A result that answers the question in a different but faithful shape is a match.
Gold-subset containment. The prediction must CONTAIN the gold answer.
Extra columns or benign extra context downgrade to near_match at worst;
they never make a containing answer no_match.
Name the column pairing. Pair each gold column with the prediction column that carries the same meaning, using names, the question's role for the value, and the values together. Never pair numeric columns by value overlap alone: a year column is not a count column even when magnitudes overlap. If a gold column has no counterpart, say which.
Rows are a multiset. Order matters only when the question asks for an order. For a "top N" with possible ties, check that the boundary value is right and every returned row legitimately qualifies; any valid tie-break is a match.
Tolerances. Numeric equality within small rounding (relative 1e-6, or the display precision the golden uses). A percentage and its fraction (50 and 0.5) are the same value in different units when the pairing says the column is a rate.
Confidence 1 to 10. 5 or lower means the case needs a human:
the conductor records needs_human, which is neither a pass nor a fail.
Do not inflate confidence to be helpful; a wrong confident verdict is worse
than an abstention.
near_match is not a soft pass, and it is not a soft fail. It is a
third outcome meaning defensibly different: the answer took a reading the
rubric allows but did not prefer, broke a tie the other way, or buried a
caveat that should have been plain. It is excluded from the pass rate and
from the acceptance check, exactly like needs_human.
So do not reach for it to avoid a hard call. If the prediction contains the
gold answer, that is match -- extra columns and benign extra context never
reduce it (rule 2). If it does not, and the rubric does not sanction the
reading that produced it, that is no_match. Use near_match only when you
can name the rubric clause that makes the difference defensible.
A rubric clause cannot make a wrong VALUE defensible, and a clause that
tries is a defect in the rubric rather than a licence to you. near_match
turns on the answer being right under a reading the QUESTION allows -- a tie
broken the other way, a grain the question left open, a basis the question
never fixed. It does not turn on the answer being transparent about how it
got a figure the question did not ask for. Those two look alike in a rubric
and are opposites in a report: one is a number a reader can act on, the
other is a number a reader would act on wrongly. Naming the method makes a
wrong figure DIAGNOSABLE, which is worth having, and it is not partial
credit.
The test, before you write near_match on a case with a value: would a
reader who acted on this figure be wrong? If yes, it is no_match however
plainly the answer explained itself, and however the rubric is worded. Say
in why that you are overriding a rubric clause, so the clause gets fixed.
This rule exists because a set shipped one: a question asked for sales over
the company's season, the answer gave the meteorological window 25% lower
and said which window it used, and a clause granting near_match for a
stated window kept a materially wrong answer out of the pass rate
entirely -- the arm reported 100%.
It is a third outcome because as a pass it was a large share of the measured
noise: the same unchanged answer reads match in one run and near_match
in the next, and the pass rate moves although nothing did. A verdict whose
content is "this is arguable" cannot be allowed to decide anything. Its
count is still reported, and a rising one means the rubrics are going vague.
(What that share was for a given set is in that set's CALIBRATION.md.)
A near_match that lands the same way in two arms is a different animal
from one that flickers. Stable across a pair, it is not judge noise: the
model cannot distinguish two readings the question does, which is a coverage
finding, and softening the rubric will not close it. flip_table.py lists
the stable ones and diagnose.py --verdicts near_match takes them.
On a large row set, compare it as a set rather than scanning pairwise: state how many gold rows you located in the prediction, name the ones you could not, and say what the mismatched values look like (uniformly scaled, off in one column, a different population). "I checked all 76" without that breakdown is not a comparison.
A mustNotUse field is not yours to weigh, unless it is prose. A
script checks the final query for the field names golden.mustNotUse
lists and forces no_match on a hit before you are asked, so a case that
reaches you with a MUST NOT USE line is carrying only what a text check
could not decide: a reading described in words, an objection to a USE of a
field rather than to the field (X as ..., X through ...), or a bare field
name that may or may not be the forbidden one. Apply those as the rubric's own
clauses. Do not soften a verdict because a veto might have caught it, and do
not invent a veto the rubric did not ask for.
Score the data, not the insight. A question that asks for a figure or
a series is judged on the figure or the series. Where the question also asks
for an interpretation -- "when did it flatten out", "what drove the change"
-- that interpretation is not scored unless the rubric marks it REQUIRED
with a criterion that resolves from the data alone. Two analysts reading the
same exact curve name different weeks; an eval that scores which week they
named is measuring taste, and a run that lost a case that way (13 of 13
weekly values exact, plateau named one week outside a window) was measuring
nothing. Exact data with a different reading of it is match.
Do not demand a grain the question did not fix. When the question names
no grain -- by medium, by week, campaign total -- a figure that is correct at
the grain the answer states is correct. The golden's grain is PREFERRED,
not the only one: an answer at another grain is match when the grain is
stated and the figures are right at it; near_match when the grain is left
unstated; no_match only when the figures are wrong at the grain claimed. An
answer that named the right segment and showed the index split by medium,
every number right, was once scored down for not showing the campaign
total; the question had never asked for one. A rubric that means "campaign
total only" must say so as REQUIRED, and the question should say so too.
A case rubric marks its alternate readings and disclosures with the words
below, and each word fixes the verdict. Apply them as written; do not re-weigh
a reading the rubric has already classified. (reference/writing-rubrics.md
is where authors are told to use them; this table is the judge's half.)
| Marker | Verdict | Meaning |
|---|---|---|
PREFERRED | match | The reading the golden encodes. |
ACCEPT | match | Equally right: a different but faithful route to the same claim. Check the figure against the golden the way the clause says to. |
DIVERGENT | near_match | Defensible and not what was asked for. Never no_match, however clearly the answer committed to it. |
WRONG | no_match | Plausible and incorrect; the clause usually names the trap. |
REQUIRED | omitted: no_match | A disclosure without which the number misleads. |
CREDITED | omitted: match | Context a good analyst adds; its absence costs nothing. |
Measured on an unchanged answer, rubric and golden: an answer whose
recommended figure the rubric marked DIVERGENT scored near_match under
one judge and no_match under the next, because the judge had never been
told what the word meant and weighed the commitment instead. An unmarked
clause is CREDITED (the author's bug, not yours to repair by inventing a
requirement).
(category, revenue); prediction 8 rows (product_category, gross_revenue, order_count). Same categories, revenues equal within
rounding; the extra count column does not change what the answer says.
Verdict: match, confidence 9.Keep the anchor set balanced. A judge shown only matches learns a base rate, not a rubric.
A case may be labelled coverage: derivable: the model has no entity for the
concept and the answer had to be built from the parts that exist. Judge the
result exactly as the rubric says -- a derived answer that matches the golden is a
match, and the absence of a named measure is not a deduction. But when the
answer states what it built, say so in the why. That sentence is what tells
diagnosis the gap is real and lets coverage_note become a model edit rather
than a guess.
Any change to this file is a judge change: bump JUDGE_VERSION, commit, and
re-run evals/<set>/judge-regressions.jsonl (the human-overruled verdicts)
before trusting new scores. Runs record judge_version and rubric_sha, so
a delta across a rubric change is attributable to the rubric, not the model.
© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/eval-judge of malloydata/publisher.
Open the folder on GitHubat commit 39a546f
Eval Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Judge this skillmalloydata/publisher | 116 | — | ~3.4k | Automated safety check: Pass | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2k | Automated safety check: Pass | MIT | |
| Codebase to Coursezarazhangrui/codebase-to-course | 5.7k | — | ~4.4k | Automated safety check: Pass | None | |
| AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Scholar EvaluationK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~2.9k | Automated safety check: Notes | MIT |
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
zarazhangrui/codebase-to-course
Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.
rohitg00/ai-engineering-from-scratch
Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.
K-Dense-AI/claude-scientific-writer
Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.
Categories
Decide whether ONE answer matches its golden, and say whether you believe the golden. Eval Judge is an agent skill from malloydata/publisher. Decide whether ONE answer matches its golden, and say whether you believe the golden.
Eval Judge fits situations like: scoring an attempt in an evaluation run; never to conduct a run (eval-loop); diagnose a failure (eval-diagnose); edit a model (eval-improve).
Run `npx skills add malloydata/publisher --skill eval-judge -a claude-code`. Or copy the skill folder (skills/eval-judge in malloydata/publisher) into .claude/skills/eval-judge in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-judge -a codex`. Or copy the skill folder (skills/eval-judge in malloydata/publisher) into .agents/skills/eval-judge in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-judge, .gemini/skills/eval-judge, .github/skills/eval-judge and .opencode/skills/eval-judge in your project.
Going by SKILL.md and its folder, Eval Judge needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Judge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Judge: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.