Agent skill

Eval Answer

by malloydata in malloydata/publisher

Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

MITAuto-check passedAgent Workflows

Install Eval Answer

skills CLI
$ npx skills add malloydata/publisher --skill eval-answer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-answer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-answer .claude/skills/eval-answer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-answer
GitHub stars
116
Token cost
~4.3k tokens
SKILL.md length
2,631 words
Files
31 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

  • Works in 6 steps: Contamination check, before any score → Re-run the submitted query yourself → Judge the answer → …
  • Asked whether an answer was correct
  • SKILL.md covers The unit, Step 1: Contamination check,…, Step 2: Re-run the submitted… and Step 3: Judge the answer, plus 6 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Eval Answer is an agent skill from malloydata/publisher. Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. Run the contamination checklist, re-execute the submitted query yourself, then spawn a judge subagent per skill:eval-judge. Append attempt, toolcall, score, and retrievalscore events to the file ledger (reference/ledger-schema.md). Never explain the failure (eval-diagnose) or edit the model (eval-improve). Use when asked whether an answer was correct, to score a run, or…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 32 other files, including scripts (for example `reference/coverage-limits.md`, `reference/definition-ledger.md` and `reference/ledger-schema.md`).

It sits in Agent Workflows, covering Agent evaluation and testing and Subagents. It works with Model Context Protocol. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Asked whether an answer was correct
  • Baseline a model

Example prompts

  • “/eval-answer”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Contamination check, before any score
  2. Re-run the submitted query yourself
  3. Judge the answer
  4. Score what retrieval delivered
  5. Distrust the golden
  6. Append events, then stop

What it can do on your machine

Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 15 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Answer loads about 4.3k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 2,631 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 2,631 words, ~4,272 tokens.

Download SKILL.mdSave it as .claude/skills/eval-answer/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.
name
eval-answer
description
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. Run the contamination checklist, re-execute the submitted query yourself, then spawn a judge subagent per skill:eval-judge. Append attempt, tool_call, score, and retrieval_score events to the file ledger (reference/ledger-schema.md). Never explain the failure (eval-diagnose) or edit the model (eval-improve). Use when asked whether an answer was correct, to score a run, or to baseline a model.

Evaluate One Answer

One user intent, answered once. This skill decides whether that answer was correct, records the evidence, and stops.

Scope boundary: verdict and events only. No diagnosis, no model edit.

The unit

A chat is not the unit. Segment by user intent. Feedback ("break it out by region") is a revision inside the same answer; grade the final accepted revision.

Take the question from the stored case (evals/<set>/cases.jsonl), never from memory or a truncated console line. Record question_sha of the exact text the answerer saw. Record servedRevision from get_context or reload, not the package name: a same-named decoy has been measured for hours.

Step 1: Contamination check, before any score

The answerer can Read or Shell its way to gold. Publisher traces do not see that, so the check runs on the HOST-side tool-use log you kept for the answerer subagent (every tool name and its path or command), plus the MCP call counts the answerer reported.

The checklist. An attempt is contaminated when its log shows any of:

  1. a Read, Shell, or any file tool touching evals/ or a gold artifact path;
  2. any access to the model file under test through a file tool (the modelPath argument on an MCP execute_query is NOT contamination; the server resolves it, the answerer never reads the file);
  3. reported_calls greater than host_tool_uses (the detectable under-report floor is reported at most total tool uses). host_tool_uses is EVERY tool use the host logged, MCP calls included; while it counted only the non-MCP ones this comparison was true of almost every clean attempt, so a run whose attempts predate the split cannot be checked this way.

skills/eval-answer/scripts/check_contamination.py is a reference aid that mechanizes the same checklist over a JSON log; your reading of the transcript is the check, the script is a second pair of eyes.

Contaminated attempts get verdict: null and contaminated: true. They are excluded from the run aggregates. They are not "wrong answers."

If you cannot produce a host log, mark contaminated: "unknown" on both the attempt and its score event, and do not treat the attempt as a clean pass.

Step 2: Re-run the submitted query yourself

Never score the agent's reported rows. Take its final query, execute it with execute_query, and write a prediction CSV under the run's artifacts/ directory.

A named view is a submitted query. execute_query takes either ad-hoc Malloy or a queryName plus sourceName, and its own tool description steers an answerer to the named form; record it as the Malloy it stands for, run: <source> -> <view>, which re-executes and reads the same as any other. Capturing only the ad-hoc form recorded an attempt that did query as submitted: false with no query to re-execute, and the judge then graded prose.

submitted: false when there is no final query. That is not a wrong answer, and it is not by itself a reason to withhold a verdict. No verdict can be issued (verdict: null, with the reason) when the attempt produced neither a query nor any answer text, when the golden is missing, provisional, invalid, or ambiguous, or when a verified golden that holds a value has no local artifact to compare. A criteria golden holds no value and needs no artifact: its clauses are the whole comparison, and withholding a verdict for a missing artifact there drops a scorable case out of the pass rate. An attempt that wrote prose and ran nothing IS judged: against a golden holding a value, an answer containing none of it is no_match however well it reasons.

Step 3: Judge the answer

Spawn one fresh judge subagent per attempt, following skill:eval-judge (the rubric, the anchors, and the output shape live there; this skill does not restate them). Give it the question, the golden, your re-executed prediction rows, the canonical query when present, and the relevant source and field definitions from the model. It returns {verdict, confidence, why, column_pairing}.

  • The judge sees gold. It is therefore never the answerer, and its verdict never leaks back to any answerer.
  • Confidence 5 or lower records as needs_human: neither a pass nor a fail, excluded from acceptance arithmetic, queued for a human look.
  • near_match is also neither. It means defensibly different, not "nearly a pass", and it stays out of the pass rate and the acceptance check for the same reason needs_human does. Report the count; do not fold it into either column.
  • When a human overrules a verdict, append the case to evals/<set>/judge-regressions.jsonl.
  • For a scalar golden, the same protocol applies to a one-value prediction. For unanswerable, a refusal that names the gap is the pass; a confident numeric answer is the fail.
  • Large row sets are still the judge's job. There is no scripted row oracle: a script that can pass a wrong answer is worse than none, and the rubric's containment and column-pairing rules are what the comparison needs.
  • golden.mustNotUse is the exception, and it is not the judge's. It names the similar-but-wrong field, and using one is a failure however good the number looks, which is a question about query TEXT. Run scripts/check_must_not_use.py over the final query: a named field found there forces no_match and records must_not_use_hits, keeping the judge's own verdict beside it as judge_verdict. Only a BARE name vetoes. An entry with a connective (weekly_active_users as a cumulative series, product.cost through the order_items join) objects to a use of the field rather than to the field, so it goes to the judge as prose, as does an entry that names no field at all and a path's bare leaf. A veto that fires on a correct answer is worse than one that misses, and reading the head of X as ... as "ban X" failed an answer whose only sin was showing X as an extra column.

Step 4: Score what retrieval delivered

Per attempt, mechanically, from the ledger -- scripts/score_retrieval.py. Each case names the entities its answer depends on (expectedEntities.required, and requiredAnyOf groups where the model offers more than one route). An entity was delivered if the attempt's get_context calls returned it as a ranked entity under its id, under the same type and name on a sibling source, or by name inside a returned source's documentation -- text the answerer reads and acts on. Only missing is a retrieval miss; the route per entity is recorded so the strict count is still there.

Recall 1.0 with a wrong answer exonerates retrieval: everything arrived. Whether the agent misused it or the docs never said how to use it is eval-diagnose's call, sufficiency first, so the row reads delivered, wrong and names no owner. Recall below 1.0 and coverage: covered means the entity existed and search did not surface it -- a documentation finding, because the retrieval algorithm is fixed (semantic search over doc strings) and an entity that exists and does not come back is one whose docs do not say what people ask; eval-diagnose calls it NOT-RETURNED, owner model. derivable or absent means there was nothing to surface. Those look identical in an answer score and have different owners, which is what makes this number worth having. It uses the search terms the answerer chose, so it attributes a failure within an arm and does not compare retrieval across arms -- that is the engine-side eval-retrieval skill, which does not ship here.

scripts/check_coverage.py measures the other half, and it is worth knowing which question each answers. Recall asks whether search surfaced the entities a case names. Coverage asks whether the model holds the concepts at all, decided by READING the model: no answerer, no judge, no goldens, no warehouse. A version can score full recall on a case the model could never have answered, and then recall is scoring search against an expectation that was never satisfiable.

Because it runs from the model text alone it is cheap enough to point at every published version and read as a trend, which is what it is for. Its verdicts are eval-diagnose's codes verbatim, validated against that table at startup, so MISSING, AMBIGUOUS, RULE_UNWRITTEN and UNDERSPECIFIED mean there exactly what they mean here. A pass is MODELLED. It writes nothing: not a golden, not the case-level coverage field, not a run directory. Where its verdict and that field disagree is where a version regressed, and the field is the standing judgement about the question while this is a measurement against one build.

Its --out report records compiledSurface: read, or why the compiled field list was not read. Without that list the judge calls a column a source exposes implicitly absent, so check the field before comparing two runs. A --model run never has it.

How to invoke it. Either --model <file-or-dir> for a local package or --publisher <url> --package <pkg> for a served one, plus --set <dir>, and --version <label> to stamp the report so a trend has an x-axis:

python3 check_coverage.py --set evals/ecommerce --model model.malloy --version 0.0.58

Then hand the report back to the run: run_baseline.py --coverage <report>. Its per-case verdict beats the case's authored coverage label in retrieval attribution, run.json records which report was read, and each retrieval row says whether measured, authored or none charged the failure. Without a report, a case with no label is attributed to nobody rather than to the model, which is what used to happen.

Two flags change what the number means, so choose them rather than inheriting them. --repeat N samples each case N times and takes the majority; it defaults to 1 for a set, because the score is a trend over many cases rather than a verdict on one, and to 3 for --self-check, where a single fixture is the whole measurement. A case that flips between samples is arguable rather than covered, so a tie goes to the gap, and two DIFFERENT gaps tying leaves the case undecided and out of the denominator. --self-check runs the shipped fixture instead of a set, which is how to confirm the checker still detects a gap it is known to detect; run it after editing the prompt, the verdict vocabulary or the fixture itself. The tests pin the fixture's shape and cannot pin its verdict, because the judge needs a live model, so a fixture swap that goes unmeasured leaves the whole metric unguarded.

Undecided cases are excluded from the percentage and reported on their own line. Read that line: a coverage number over a handful of decided cases is not a measurement, it is a sample size.

Read reference/coverage-limits.md before quoting a coverage number. Run against all 49 hand-labelled cases of the evals/ecommerce set it disagreed with the authored label on 17, while the two headline percentages landed two points apart because the disagreements cancel. Checking each one against the model found the LABEL was the stale side more often than the verdict was: four notes call a measure missing that the model declares, and one prescribes a filter on a field the model documents as "not evidence of a sale". So read a coverage number as a rough signal over many cases, never as a verdict on one and never as a small movement between versions, and use --compare-labels to print the disagreements and triage them by reading the model. Agreement with the labels is not a score for the checker; the two answer different questions on purpose. Separately, the whole model goes into one prompt per case, so past a few thousand lines it stops running on Linux and above about 100 KB the verdict stops being stable. Do not take the over-size message's advice to narrow --model to one file, which drops every imported source and manufactures MISSING verdicts.

Show full SKILL.md (706 more words)Show less

Validate the definitions, not every answer

A golden must not be derived through the definitions the question TESTS. It may reuse everything the question does not turn on -- joins, base sources, date handling -- because circularity only bites where the key and the answer share the step under measurement. That is what makes this affordable on a model too large to reimplement.

scripts/verify_definitions.py checks each definition once against the layer directly beneath it, and every case depending on it inherits the result. A golden is trustworthy if it was derived independently, OR if every definition it tests has itself been validated.

Two things it will not do, and both matter more than what it does. It never reports a definition reaching through a join as validated, because fanout inflates the measure and the control expression equally and the comparison stays green on a broken join. And building the ledger without a server exits 3, not 0: nothing was checked, and a caller must not read that as a pass.

A raw check -- a definition that reaches through a join -- is an authored control: a person writes the population down as a query on the ledger record and the tool re-runs it every time. It is not found unaided, and the docs say so. verify_goldens.py --definitions <ledger> then applies the composition rule at the gate: a set with no truth package exits 0 when every value-bearing case's tested definitions are validated, and names the cases that are not.

Read reference/definition-ledger.md before building or quoting one.

Step 5: Distrust the golden

A reference answer can be wrong (parent-column fanout, a join on a shared non-identifying key, or a rubric describing a model that has since been fixed). Fanout is not automatically a defect: AVG / STDDEV / MIN / MAX survive uniform duplication. Classify verified_wrong (exclude from scoring) vs verified_benign (keep).

The judge produces this, not you. It is the only station holding the golden, the re-executed rows and the model source at once, so it is the only one that can see the key contradict any of them; the rules and the four values are in skill:eval-judge. Carry its gold_status and gold_note onto the score event unchanged, and where it says nothing, fall back to the case's standing golden.status.

Do not encode a rewrite of a bad golden into the model. A suspect or verified_wrong, or a no_match whose why indicts the golden rather than the prediction, routes to the golden side door in skill:eval-loop as a dataset issue, which is where the audit procedure for settling it lives. It is never a model failure, and it must be settled before improve runs -- otherwise a modelling agent is dispatched to fix a model that is already right.

Step 6: Append events, then stop

Append to <workdir>/runs/<runId>/events.jsonl with caseId set (the run directory eval.py run printed; the workdir is the set's, from eval.toml). Shapes live in reference/ledger-schema.md.

  1. attempt: qid, sample, phase, question_sha, submitted, final_query, served revision, call counts, contamination verdict, transcript path.
  2. tool_call: one per MCP get_context / execute_query, with traceId and the rankedSummary copied from the trace (per-target ranks included). Do not copy full traces into the event; the trace store holds the body.
  3. score: the judge's verdict object plus judge_version, rubric_sha, golden_revision, contaminated, gold_status, and the judge output's artifact path.

A stage never rewrites another stage's fields. End-of-run numbers come from counting events, not from your arithmetic in prose.

Sample each case once. Breadth across cases beats repeats of one case; the comparison rule for a before/after is the flip count in skill:eval-loop Measurement, not a mean over samples.

Re-score after a golden repair

When eval-loop has repaired a golden and opened a new run, this skill runs again without a new answerer: same stored final_query (or its saved prediction CSV), new gold artifact, fresh judge, new golden_revision on the score event. Contamination does not need to be re-litigated if the attempt was already clean. If you must re-execute, do it yourself; do not ask the original answerer to "try again" with the new key in context.

  • skill:eval-diagnose: why it failed, after this record exists.
  • skill:eval-improve: smallest model edit, model-owned issues only.
  • Step 5 of the malloy-analysis skill: checks before you trust a result you ran.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 30 other files (scripts) in skills/eval-answer of malloydata/publisher.

  • SKILL.md
  • reference/coverage-limits.md
  • reference/definition-ledger.md
  • reference/ledger-schema.md
  • scripts/check_contamination.py
  • scripts/check_contamination_test.py
  • scripts/check_coverage.py
  • scripts/check_coverage_test.py
  • scripts/check_findable.py
  • scripts/check_findable_test.py
  • scripts/check_must_not_use.py
  • scripts/check_must_not_use_test.py
  • scripts/config.py
  • scripts/config_test.py
  • scripts/golden_rows.py
  • scripts/golden_rows_test.py
  • scripts/init_truth_package.py
  • scripts/init_truth_package_test.py
  • scripts/json_scan.py
  • … and 12 more

Open the folder on GitHubat commit acc1acd

Compare with similar skills

Eval Answer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Answer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Answer this skillmalloydata/publisher116—~4.3kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Claude Automation Recommenderanthropics/claude-plugins-official37k3 repos~2.7kAutomated safety check: NotesApache-2.0
Agent Deckasheshgoplani/agent-deck1k—~1.7kAutomated safety check: PassMIT
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Kayba Pipelinekayba-ai/agentic-context-engine2.6k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Claude Automation Recommender

    anthropics/claude-plugins-official

    Official

    Scans a codebase and suggests which Claude Code hooks, subagents, skills, plugins and MCP servers fit its stack, without changing any files.

    37k GitHub starsUsed in 3 repos~2.7k tokens
    Agent WorkflowsAuto-check: notes
  • Agent Deck

    asheshgoplani/agent-deck

    agent-deck, the terminal session manager for AI coding agents.

    1k GitHub stars~1.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Kayba Pipeline

    kayba-ai/agentic-context-engine

    End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

    2.6k GitHub stars~1.4k tokensUpdated 13 days ago
    Agent WorkflowsAuto-check passed
  • CC Workflow Studio AI Editor

    breaking-brake/cc-wf-studio

    Creates and edits visual agent workflows in CC Workflow Studio through conversation, with the agent reading and writing the canvas over MCP.

    5.4k GitHub stars~561 tokensUpdated today
    Agent WorkflowsAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Malloy Analysis Report

    malloydata/publisher

    Combine validated Malloy queries into a notebook report. An agent skill from malloydata/publisher.

    116 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Eval Answer

What does Eval Answer do?

Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. Eval Answer is an agent skill from malloydata/publisher. Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

When should I use Eval Answer?

Eval Answer fits situations like: asked whether an answer was correct; baseline a model.

How do I install Eval Answer in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-answer -a claude-code`. Or copy the skill folder (skills/eval-answer in malloydata/publisher) into .claude/skills/eval-answer in your project. Claude Code loads it when a task matches its description.

How do I install Eval Answer in Codex?

Run `npx skills add malloydata/publisher --skill eval-answer -a codex`. Or copy the skill folder (skills/eval-answer in malloydata/publisher) into .agents/skills/eval-answer in your project. Codex loads it when a task matches its description.

Can I use Eval Answer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-answer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-answer, .gemini/skills/eval-answer, .github/skills/eval-answer and .opencode/skills/eval-answer in your project.

What does Eval Answer need to run?

Going by SKILL.md and its folder, Eval Answer needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Eval Answer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Answer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Answer use?

Eval Answer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Answer use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Answer?

Skills that share tags, products or a category with Eval Answer: MCP Server Builder (anthropics/skills, 180k stars), Claude Automation Recommender (anthropics/claude-plugins-official, 37k stars), Agent Deck (asheshgoplani/agent-deck, 1k stars) and Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Answer?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.