MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
$ npx skills add malloydata/publisher --skill eval-answer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher eval-answer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-answer .claude/skills/eval-answer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .claude/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/eval-answerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill eval-answer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher eval-answer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-answer .agents/skills/eval-answer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .agents/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-answer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher eval-answer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-answer .cursor/skills/eval-answer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .cursor/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/eval-answer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill eval-answer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher eval-answer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-answer .gemini/skills/eval-answer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .gemini/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher eval-answerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill eval-answer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-answer .github/skills/eval-answer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .github/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill eval-answer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher eval-answer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-answer .opencode/skills/eval-answer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-answer" agent skill from https://github.com/malloydata/publisher/tree/main/skills/eval-answer into .opencode/skills/eval-answer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-answer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-answerScore one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
Eval Answer is an agent skill from malloydata/publisher. Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. Run the contamination checklist, re-execute the submitted query yourself, then spawn a judge subagent per skill:eval-judge. Append attempt, toolcall, score, and retrievalscore events to the file ledger (reference/ledger-schema.md). Never explain the failure (eval-diagnose) or edit the model (eval-improve). Use when asked whether an answer was correct, to score a run, or…
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 32 other files, including scripts (for example `reference/coverage-limits.md`, `reference/definition-ledger.md` and `reference/ledger-schema.md`).
It sits in Agent Workflows, covering Agent evaluation and testing and Subagents. It works with Model Context Protocol. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 15 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Answer loads about 4.3k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 2,631 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 2,631 words, ~4,272 tokens.
.claude/skills/eval-answer/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.One user intent, answered once. This skill decides whether that answer was correct, records the evidence, and stops.
Scope boundary: verdict and events only. No diagnosis, no model edit.
A chat is not the unit. Segment by user intent. Feedback ("break it out by region") is a revision inside the same answer; grade the final accepted revision.
Take the question from the stored case (evals/<set>/cases.jsonl), never from
memory or a truncated console line. Record question_sha of the exact text
the answerer saw. Record servedRevision from get_context or reload, not
the package name: a same-named decoy has been measured for hours.
The answerer can Read or Shell its way to gold. Publisher traces do not see that, so the check runs on the HOST-side tool-use log you kept for the answerer subagent (every tool name and its path or command), plus the MCP call counts the answerer reported.
The checklist. An attempt is contaminated when its log shows any of:
evals/ or a gold artifact path;modelPath argument on an MCP execute_query is NOT contamination; the
server resolves it, the answerer never reads the file);reported_calls greater than host_tool_uses (the detectable
under-report floor is reported at most total tool uses). host_tool_uses
is EVERY tool use the host logged, MCP calls included; while it counted only
the non-MCP ones this comparison was true of almost every clean attempt, so
a run whose attempts predate the split cannot be checked this way.skills/eval-answer/scripts/check_contamination.py is a reference aid that
mechanizes the same checklist over a JSON log; your reading of the transcript
is the check, the script is a second pair of eyes.
Contaminated attempts get verdict: null and contaminated: true. They are
excluded from the run aggregates. They are not "wrong answers."
If you cannot produce a host log, mark contaminated: "unknown" on both the
attempt and its score event, and do not treat the attempt as a clean pass.
Never score the agent's reported rows. Take its final query, execute it with
execute_query, and write a prediction CSV under the run's
artifacts/ directory.
A named view is a submitted query. execute_query takes either ad-hoc Malloy
or a queryName plus sourceName, and its own tool description steers an
answerer to the named form; record it as the Malloy it stands for,
run: <source> -> <view>, which re-executes and reads the same as any other.
Capturing only the ad-hoc form recorded an attempt that did query as
submitted: false with no query to re-execute, and the judge then graded prose.
submitted: false when there is no final query. That is not a wrong answer, and
it is not by itself a reason to withhold a verdict. No verdict can be issued
(verdict: null, with the reason) when the attempt produced neither a query nor
any answer text, when the golden is missing, provisional, invalid, or ambiguous,
or when a verified golden that holds a value has no local artifact to compare.
A criteria golden holds no value and needs no artifact: its clauses are the
whole comparison, and withholding a verdict for a missing artifact there drops
a scorable case out of the pass rate. An attempt that
wrote prose and ran nothing IS judged: against a golden holding a value, an
answer containing none of it is no_match however well it reasons.
Spawn one fresh judge subagent per attempt, following
skill:eval-judge (the rubric, the anchors, and the output shape live
there; this skill does not restate them). Give it the question, the golden,
your re-executed prediction rows, the canonical query when present, and the
relevant source and field definitions from the model. It returns
{verdict, confidence, why, column_pairing}.
needs_human: neither a pass nor a fail,
excluded from acceptance arithmetic, queued for a human look.near_match is also neither. It means defensibly different, not "nearly a
pass", and it stays out of the pass rate and the acceptance check for the same reason
needs_human does. Report the count; do not fold it into either column.evals/<set>/judge-regressions.jsonl.unanswerable, a refusal that names the gap is the pass; a confident
numeric answer is the fail.golden.mustNotUse is the exception, and it is not the judge's. It names the
similar-but-wrong field, and using one is a failure however good the number
looks, which is a question about query TEXT. Run
scripts/check_must_not_use.py over the final query: a named field found
there forces no_match and records must_not_use_hits, keeping the judge's
own verdict beside it as judge_verdict. Only a BARE name vetoes. An entry
with a connective (weekly_active_users as a cumulative series, product.cost through the order_items join) objects to a use of the field rather than to
the field, so it goes to the judge as prose, as does an entry that names no
field at all and a path's bare leaf. A veto that fires on a correct answer is
worse than one that misses, and reading the head of X as ... as "ban X"
failed an answer whose only sin was showing X as an extra column.Per attempt, mechanically, from the ledger -- scripts/score_retrieval.py. Each
case names the entities its answer depends on (expectedEntities.required, and
requiredAnyOf groups where the model offers more than one route). An entity
was delivered if the attempt's get_context calls returned it as a ranked entity
under its id, under the same type and name on a sibling source, or by name inside
a returned source's documentation -- text the answerer reads and acts on. Only
missing is a retrieval miss; the route per entity is recorded so the strict
count is still there.
Recall 1.0 with a wrong answer exonerates retrieval: everything arrived. Whether
the agent misused it or the docs never said how to use it is eval-diagnose's
call, sufficiency first, so the row reads delivered, wrong and names no owner.
Recall below 1.0 and coverage: covered means the entity existed and search did
not surface it -- a documentation finding, because the retrieval algorithm is
fixed (semantic search over doc strings) and an entity that exists and does not
come back is one whose docs do not say what people ask; eval-diagnose calls it
NOT-RETURNED, owner model. derivable or absent means there was nothing to
surface. Those look identical in an answer score and have different owners,
which is what makes this number worth having. It uses the search terms the answerer chose, so it
attributes a failure within an arm and does not compare retrieval across arms
-- that is the engine-side eval-retrieval skill, which does not ship here.
scripts/check_coverage.py measures the other half, and it is worth knowing
which question each answers. Recall asks whether search surfaced the entities a
case names. Coverage asks whether the model holds the concepts at all, decided
by READING the model: no answerer, no judge, no goldens, no warehouse. A version
can score full recall on a case the model could never have answered, and then
recall is scoring search against an expectation that was never satisfiable.
Because it runs from the model text alone it is cheap enough to point at every
published version and read as a trend, which is what it is for. Its verdicts are
eval-diagnose's codes verbatim, validated against that table at startup, so
MISSING, AMBIGUOUS, RULE_UNWRITTEN and UNDERSPECIFIED mean there exactly
what they mean here. A pass is MODELLED. It writes nothing: not a golden, not the case-level coverage
field, not a run directory. Where its verdict and that field disagree is where a
version regressed, and the field is the standing judgement about the question
while this is a measurement against one build.
Its --out report records compiledSurface: read, or why the compiled field
list was not read. Without that list the judge calls a column a source exposes
implicitly absent, so check the field before comparing two runs. A --model
run never has it.
How to invoke it. Either --model <file-or-dir> for a local package or
--publisher <url> --package <pkg> for a served one, plus --set <dir>, and
--version <label> to stamp the report so a trend has an x-axis:
python3 check_coverage.py --set evals/ecommerce --model model.malloy --version 0.0.58Then hand the report back to the run: run_baseline.py --coverage <report>.
Its per-case verdict beats the case's authored coverage label in retrieval
attribution, run.json records which report was read, and each retrieval row
says whether measured, authored or none charged the failure. Without a
report, a case with no label is attributed to nobody rather than to the model,
which is what used to happen.
Two flags change what the number means, so choose them rather than inheriting
them. --repeat N samples each case N times and takes the majority; it defaults
to 1 for a set, because the score is a trend over many cases rather than a
verdict on one, and to 3 for --self-check, where a single fixture is the whole
measurement. A case that flips between samples is arguable rather than covered,
so a tie goes to the gap, and two DIFFERENT gaps tying leaves the case undecided
and out of the denominator. --self-check runs the shipped fixture instead of a
set, which is how to confirm the checker still detects a gap it is known to
detect; run it after editing the prompt, the verdict vocabulary or the fixture
itself. The tests pin the fixture's shape and cannot pin its verdict, because
the judge needs a live model, so a fixture swap that goes unmeasured leaves the
whole metric unguarded.
Undecided cases are excluded from the percentage and reported on their own line. Read that line: a coverage number over a handful of decided cases is not a measurement, it is a sample size.
Read reference/coverage-limits.md before quoting a coverage number. Run
against all 49 hand-labelled cases of the evals/ecommerce set it disagreed
with the authored label on 17, while the two headline percentages landed two
points apart because the disagreements cancel. Checking each one against the
model found the LABEL was the stale side more often than the verdict was: four
notes call a measure missing that the model declares, and one prescribes a
filter on a field the model documents as "not evidence of a sale". So read a
coverage number as a rough signal over many cases, never as a verdict on one and
never as a small movement between versions, and use --compare-labels to print
the disagreements and triage them by reading the model. Agreement with the
labels is not a score for the checker; the two answer different questions on
purpose. Separately, the whole model goes into one prompt per case, so past a
few thousand lines it stops running on Linux and above about 100 KB the verdict
stops being stable. Do not take the over-size message's advice to narrow
--model to one file, which drops every imported source and manufactures
MISSING verdicts.
A golden must not be derived through the definitions the question TESTS. It may reuse everything the question does not turn on -- joins, base sources, date handling -- because circularity only bites where the key and the answer share the step under measurement. That is what makes this affordable on a model too large to reimplement.
scripts/verify_definitions.py checks each definition once against the layer
directly beneath it, and every case depending on it inherits the result. A golden
is trustworthy if it was derived independently, OR if every definition it tests
has itself been validated.
Two things it will not do, and both matter more than what it does. It never reports a definition reaching through a join as validated, because fanout inflates the measure and the control expression equally and the comparison stays green on a broken join. And building the ledger without a server exits 3, not 0: nothing was checked, and a caller must not read that as a pass.
A raw check -- a definition that reaches through a join -- is an authored
control: a person writes the population down as a query on the ledger record and
the tool re-runs it every time. It is not found unaided, and the docs say so.
verify_goldens.py --definitions <ledger> then applies the composition rule at
the gate: a set with no truth package exits 0 when every value-bearing case's
tested definitions are validated, and names the cases that are not.
Read reference/definition-ledger.md before building or quoting one.
A reference answer can be wrong (parent-column fanout, a join on a shared
non-identifying key, or a rubric describing a model that has since been fixed).
Fanout is not automatically a defect: AVG / STDDEV / MIN / MAX survive
uniform duplication. Classify verified_wrong (exclude from scoring) vs
verified_benign (keep).
The judge produces this, not you. It is the only station holding the golden,
the re-executed rows and the model source at once, so it is the only one that can
see the key contradict any of them; the rules and the four values are in
skill:eval-judge. Carry its gold_status and gold_note onto the score
event unchanged, and where it says nothing, fall back to the case's standing
golden.status.
Do not encode a rewrite of a bad golden into the model. A suspect or
verified_wrong, or a no_match whose why indicts the golden rather than the
prediction, routes to the golden side door in skill:eval-loop as a dataset
issue, which is where the audit procedure for settling it lives. It is never a model failure, and it must be settled before improve runs --
otherwise a modelling agent is dispatched to fix a model that is already right.
Append to <workdir>/runs/<runId>/events.jsonl with caseId set (the run
directory eval.py run printed; the workdir is the set's, from eval.toml). Shapes
live in reference/ledger-schema.md.
attempt: qid, sample, phase, question_sha, submitted, final_query,
served revision, call counts, contamination verdict, transcript path.tool_call: one per MCP get_context / execute_query, with traceId
and the rankedSummary copied from the trace (per-target ranks included).
Do not copy full traces into the event; the trace store holds the body.score: the judge's verdict object plus judge_version, rubric_sha,
golden_revision, contaminated, gold_status, and the judge output's
artifact path.A stage never rewrites another stage's fields. End-of-run numbers come from counting events, not from your arithmetic in prose.
Sample each case once. Breadth across cases beats repeats of one case; the
comparison rule for a before/after is the flip count in skill:eval-loop
Measurement, not a mean over samples.
When eval-loop has repaired a golden and opened a new run, this skill runs
again without a new answerer: same stored final_query (or its saved
prediction CSV), new gold artifact, fresh judge, new golden_revision on the
score event. Contamination does not need to be re-litigated if the attempt
was already clean. If you must re-execute, do it yourself; do not ask the
original answerer to "try again" with the new key in context.
skill:eval-diagnose: why it failed, after this record exists.skill:eval-improve: smallest model edit, model-owned issues only.malloy-analysis skill: checks before you trust a result you ran.© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 30 other files (scripts) in skills/eval-answer of malloydata/publisher.
Open the folder on GitHubat commit acc1acd
Eval Answer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Answer this skillmalloydata/publisher | 116 | — | ~4.3k | Automated safety check: Pass | MIT | |
| MCP Server Builderanthropics/skills | 180k | 62 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Claude Automation Recommenderanthropics/claude-plugins-official | 37k | 3 repos | ~2.7k | Automated safety check: Notes | Apache-2.0 | |
| Agent Deckasheshgoplani/agent-deck | 1k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Autocontext for Hermesgreyhaven-ai/autocontext | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Kayba Pipelinekayba-ai/agentic-context-engine | 2.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
anthropics/claude-plugins-official
Scans a codebase and suggests which Claude Code hooks, subagents, skills, plugins and MCP servers fit its stack, without changing any files.
asheshgoplani/agent-deck
agent-deck, the terminal session manager for AI coding agents.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
kayba-ai/agentic-context-engine
End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.
breaking-brake/cc-wf-studio
Creates and edits visual agent workflows in CC Workflow Studio through conversation, with the agent reading and writing the canvas over MCP.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
malloydata/publisher
Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.
Works with
Categories
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. Eval Answer is an agent skill from malloydata/publisher. Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
Eval Answer fits situations like: asked whether an answer was correct; baseline a model.
Run `npx skills add malloydata/publisher --skill eval-answer -a claude-code`. Or copy the skill folder (skills/eval-answer in malloydata/publisher) into .claude/skills/eval-answer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill eval-answer -a codex`. Or copy the skill folder (skills/eval-answer in malloydata/publisher) into .agents/skills/eval-answer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-answer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-answer, .gemini/skills/eval-answer, .github/skills/eval-answer and .opencode/skills/eval-answer in your project.
Going by SKILL.md and its folder, Eval Answer needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Answer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Answer: MCP Server Builder (anthropics/skills, 180k stars), Claude Automation Recommender (anthropics/claude-plugins-official, 37k stars), Agent Deck (asheshgoplani/agent-deck, 1k stars) and Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.