Agent skill

Eval Report

by malloydata in malloydata/publisher

Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden…

MITAuto-check passed

Install Eval Report

skills CLI
$ npx skills add malloydata/publisher --skill eval-report -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher eval-report --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-report .claude/skills/eval-report && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-report
GitHub stars
116
Token cost
~6.4k tokens
SKILL.md length
3,926 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden…

  • Works in 2 steps: build the artifacts, before writing a word → write the report
  • SKILL.md covers Step 1: build the artifacts,…, Step 2: write the report, The two failure sections, and… and The four measurements, and why…, plus 5 more sections
  • Calls python3

What it does

Eval Report is an agent skill from malloydata/publisher. Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden, unreadable judge), what was learned, and what to do next. Every claim links to the artifact behind it. Use at the end of any run, before quoting a number to anyone. Never scores an answer (eval-answer) or decides who owns a failure (eval-diagnose).

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

Example prompts

  • “/eval-report”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. build the artifacts, before writing a word
  2. write the report

What it can do on your machine

Read from SKILL.md and the folder at commit acc1acd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Report loads about 6.4k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 3,926 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from malloydata/publisher at commit acc1acd, republished under its MIT licence (© malloydata). 3,926 words, ~6,385 tokens.

Download SKILL.mdSave it as .claude/skills/eval-report/SKILL.md (or your agent's skills folder).
name
eval-report
description
Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden, unreadable judge), what was learned, and what to do next. Every claim links to the artifact behind it. Use at the end of any run, before quoting a number to anyone. Never scores an answer (eval-answer) or decides who owns a failure (eval-diagnose).

Report a Run

A run directory is JSONL and a console summary that scrolls away. Neither is a result. This skill turns one into something a person can open, and it is the last step of every run, including a run that failed.

The complaint this exists to answer, in a reviewer's words: "eval runs and a bunch of stuff happens and it is hard to know the actual results."

Scope boundary: this skill reports what the ledger already says. It does not score an answer (skill:eval-answer), decide who owns a failure (skill:eval-diagnose), or edit a model (skill:eval-improve). If a number is not in the ledger, do not put it in the report.

The report is about THIS RUN, not about the harness. Defects you find in the eval tooling while running it are real and worth filing, and they do not belong here: the reader wants to know what their model scored and why, not what is wrong with the thing that measured it. File those against the harness. The one exception is anything that qualifies THIS run's number -- a truncated attempt, a contaminated one, an unestablished answer key -- which the "Eval failures" section exists for.

Do not audit the harness. You were asked to measure a model. This is the most common way this job goes wrong, and it does not look like going wrong: a run turns up something odd in a script, the odd thing is genuinely a bug, and the reply comes back as a critique of the tooling with the model's score somewhere underneath. The reader asked what their model scored. Answer that.

So, unless the user asked you to work on the harness:

  • Do not read harness source to satisfy your own curiosity about a number. Read it when a number you must report cannot be explained any other way, and stop when it can.
  • Do not propose harness fixes, refactors, flags or "while I was in there" improvements. Not in the report, not in the chat reply.
  • When a harness defect DID change this run's number, the report gets one sentence: what the number should be and why. Not the mechanism, not the file, not the fix.
  • Keep a defect that changed nothing out of the report entirely. Mention it once in chat, in a line, and let the user decide whether they want it chased.

A harness bug you found and did not chase is not a loose end. It is the job being done. If the user wants it fixed they will say so, and then it is a different task with its own turn.

Step 1: build the artifacts, before writing a word

bash
python3 skills/eval-loop/scripts/eval.py package --set <set-dir> --label <label>

That builds a Malloy package over the run's own CSVs, registers it with no restart, and prints the two URLs below. It registers on the TRUTH server, because the package holds the answer key and the answerer must not reach it. If the truth server is not running it says so and prints the curl to run once it is. A set with no truth server gets no registration: the only Publisher is the answerer's. Add a [truth] section and eval.py serve truth, or pass --on-model-server, which prints the curl and the DELETE to run before the next run.

It refuses a run with no diagnosis, since the report's cluster views would be empty. Run eval.py diagnose first; when this run skipped diagnosis on purpose, pass --without-diagnosis and say in the report that it has no clusters.

It gives you two things to link:

ArtifactWhat it isURL
The case matrix appEvery question, its verdict, which needed entities retrieval delivered, and a drawer per case holding the reference answer, the judge's reasoning, the re-executed rows and every query the answerer ran<publisher>/environments/<env>/packages/eval-<label>/
notebooks/eval_run.malloyThe aggregate tables: pass rate, effort, cost, most-missed entities, the backlog<publisher>/<env>/eval-<label>/notebooks/eval_run

Those two URLs are in different path spaces, and guessing costs a 404. The app is served by the in-package public/ handler, which owns /environments/<env>/packages/<pkg>/<file>. The notebook is a MODEL, rendered by the Console, whose routes are /<env>/<package>/<model path> with no environments or packages segment at all. Putting the notebook under the app's prefix 404s, because the public-file handler answers there and the notebook is not in public/. Verified both, on a real package, by opening them.

Pass --run more than once to put two arms side by side.

Step 2: write the report

Put it in the repository beside the set, as RESULTS.md. Not in a chat log, not in ~/Downloads, both of which have lost a findings document before.

Derive every number from events.jsonl and run.json, not from memory or from the console. A figure retyped from scrollback is a figure nobody can check, and the console rounds.

Keep it under a page and a half

Budget: about 80 lines, and 120 is the ceiling for a run with several distinct failures. Reports have shipped at 200 and the length is not thoroughness -- it is the artifact links restated as prose, the cascade explained twice, and a paragraph apologising for a figure nobody disputed. A reader who wants the detail opens the case matrix, which is why step 1 builds it. What cannot be recovered from the artifacts is your judgement: what broke, why, and what to do. Spend the lines there.

What earns its place: the headline, one row per case, one short entry per failure, the retrieval numbers, and the next steps. What does not: restating a number you already gave, explaining what a cascade is before showing it, defending a decision nobody questioned, or any section whose content is that nothing happened -- except "Eval failures", where "None." is the point.

Being brief is not being terse with the vocabulary. The reader does not know what near_match, a cluster, recall or a holdout is, so the first time one appears, say what it means in the same sentence -- "near_match, which is excluded from the pass rate" -- and then use it. Cutting the explanations is how a short report becomes an unreadable one; cutting the restatements is how it becomes a good one.

The template
markdown
# <set> -- <label>

<one sentence: what was being measured, against what, and why now>

## What ran

- Set `<name>` v`<datasetVersion>`, N cases (N dev, N holdout)
- Target `<env>/<package>`, model `<modelPath>`, pinned `<modelSha or targetVersion>`
- Answerer `<model>`, judge `<model>`, cap `<maxTurns>` turns
- Steps run: scrape/run, eval<, diagnose><, improve>
- Cost $X answerer + $Y judge, N turns median / N p90, N entities delivered per answer

## The result

**N of M decided (P%).**   <-- or: **No pass rate.** See "Eval failures" below.

| qid | question | verdict |
|---|---|---|

[case matrix](<url>) - [aggregate tables](<url>)

## Coverage, retrieval, accuracy, cost

**Coverage N of M** -- how many questions the model can express an answer to at
all. **Entity recall N%** -- one number, first, before the cascade. It is the share
of the entities an answer needed that retrieval actually handed the agent, and
it is the headline of this section for the same reason the pass rate is the
headline of the last one. Then the cascade: Covered? -> Retrieved? -> Correct?,
with the per-arm numbers under each. Say in one clause what recall counts, then
give the number.

## Model failures

One entry per wrong answer. What it got wrong in plain words, then the
mechanism, then the link.

## Eval failures

What went wrong with the MEASUREMENT rather than the model. Empty is a real
and good answer; say "none" rather than dropping the section.

## Query errors

Malloy the answerer wrote that would not run, and whether it recovered.

## What this run taught us

## What to do next
"What to do next" is a checklist, and most of it is derived

Do not invent this section. Most of it falls out of what the run did NOT do, and a reader should be able to see that nothing was skipped silently. Walk these in order and put every one that fires into the list, with its command:

If the run showsThen the next step is
coverage: unmeasuredthe run was given --no-coverage, or the measurement failed and said so. It is measured by default, so say in "Eval failures" that this score cannot tell a model gap from a bad answer, and re-run without the flag
goldenCheck: skipped or the set names no truthPackagebuild one with init_truth_package.py; until then the goldens were derived through the model under test and certify themselves
any golden still provisionalre-derive and verify_goldens.py --promote
a stale entity name warningfix expectedEntities; it scores as a retrieval miss on every run until you do. A next step, not a section: it goes in this list and nowhere else in the report
a passing case with recall below 1.0check whether required over-specifies one path
truncated non-emptyre-run those cases at a higher cap with --from
diagnose did not runrun it, or say the failures have no owner yet
a cluster with owner: modelskill:eval-improve, then the acceptance check
a cluster with owner: agent-skilledit that skill. This is NOT a dead end
a cluster with owner: datasetthe golden side door in skill:eval-loop
only one arm existsnote that no noise band has been measured for this set yet

An agent-skill cluster is work, not an absence of work. eval-improve may not touch it, and writing "nothing to do, the model is fine" there is how a real defect gets closed as a non-finding. Name the skill, name the rule to add or change, and say who owns it. The model being innocent is a statement about the model, never about the run.

Write it for someone who was not there

The reader is a colleague who did not run this and does not know the loop's vocabulary. Two rules, both learned by handing a report to one:

  • Never make a bare count carry the meaning. "The only cluster is owner: agent-skill" tells a reader nothing: what is a cluster, and what follows from it? Write what happened and what to do: "The one wrong answer came from the agent picking the wrong kind of field, so the fix is in the analysis skill, not in the model."
  • Do not open with a negation of something you just reported. A section that says "nothing to do" directly under a section reporting a wrong answer reads as a contradiction, and the reader stops trusting both. Lead with the wrong answer and what closes it; put anything genuinely needing no action after that, and say why.

The two failure sections, and why they are separate

This is the part most reports get wrong. A wrong answer and a broken measurement look identical in a pass rate and have nothing else in common: one is work for the model owner, the other is work for whoever runs the harness. Mixing them sends a modelling agent to fix a model that is fine.

Model failures are cases that were fairly asked, fairly answered and got the wrong answer. One entry each: what the answer said, what was right, and the mechanism in one sentence. Do not write a cluster id here; write what it got wrong.

Eval failures are the things that stopped the run measuring the model, and only those. The test is one question: did this cost a verdict, or make one untrustworthy? If no, it does not appear in the report at any length.

That rules out most of what is tempting to put here, and all of it has been put here on a real run: a -dirty model pin, a stale-entity-name warning that cost no case, a defect you hit in the harness and worked around, a setup step that took two tries, anything you would open with "worth knowing, though it changes no verdict". A reader wants to know what their model scored. Harness defects are real and belong in a harness issue, filed against the harness -- that is the skill's opening rule, and this section is where it gets broken.

The section is usually two words. "None." is a complete and good answer, and a reader who sees it learns exactly what they need to. Do not pad it into a paragraph explaining the absence.

Report every one that DID occur, with its count and its qids, and say plainly that these are NOT evidence about the model:

What happenedHow it reads in the ledgerWho fixes it
Answerer ran out of turnsreason: answerer_truncated, and run.json truncatedraise --max-turns, re-run those cases with --from
Isolation breachedreason: contaminated, contamination_reasons on the attemptharness config; the answerer held a tool it should not
Harness or server failedreason: environment_failure: <what>fix the environment, re-run the arm
Answer key not establishedreason: golden_provisional / _invalid / _ambiguous / _missingderive the key (skill:eval-import), or the golden side door
Judge reply unreadablereason: judge_unparseableit already retried; read artifacts/<qid>/judge.md
Judge doubted the keygold_status suspect / verified_wrong, run.json doubtedGoldensthe golden side door in skill:eval-loop, NOT improve
Rubric quoted a stale figureverify_goldens.py check 2 review itemsrepair the rubric against its own golden rows
Arm stopped earlyrun.json status: abortedfour consecutive dead attempts; fix and re-run

If any of the first three occurred, there is no pass rate to quote. The harness prints INCOMPLETE and withholds the percentage; the report does the same. Give the counts and the re-run command instead of a number with a caveat, because the number is what gets repeated and the caveat is what gets dropped.

Show full SKILL.md (1,927 more words)Show less

The four measurements, and why a report carries all of them

A pass rate says an answer was wrong. It does not say where, and on its own it makes every failure look like the model's. A run measures four things, in this order, and each one tells you whether the next is even a fair question:

MeasuresA failure here means
a. Coveragecan the data and the model express an answer at allnothing downstream was winnable. Fix the model
b. Retrievaldid get_context hand the agent the entities it neededthe model has it and the agent never saw it: the docs, or the search wording
c. Accuracydid the agent get the answer rightit had what it needed and still missed: the agent, or a doc that misleads
d. Costdollars, turns, wall-clock, entities per answerit works and cannot be afforded, which is its own kind of not working

Coverage first, and a report without it is incomplete. A model that cannot express an answer can never succeed at that question, so a pass rate quoted without coverage cannot distinguish a bad model from an unanswerable set. run_baseline.py measures it by default; if a run skipped it, say so in "Eval failures" and treat the score as provisional.

Cost is a result, not an aside. Report dollars, the turn and call counts, and the entity payload per answer. An agent that answers correctly in 40 turns and 75 delivered entities per question is a different product from one that does it in 8 and 5, and only the report says which you have.

Report them all, because a report that omits them makes every failure look like the model's:

Copy the cascade the run printed. Do not re-derive it. cascade_lines() in run_baseline.py already renders every rung, and it carries one field a hand-made table keeps losing: how many cases answered correctly anyway after stopping on that rung. Paste its block into the report.

A finer measurement must never revise the headline result. This is the rule the re-derivation breaks. Written as a funnel with a shrinking denominator -- 5 of 10 covered, then 3 of the 5, then 3 of the 3 -- the last rung reads as "only 3 of 10 succeeded" on a run where all ten answers were correct. That report went to a reviewer, and the objection was the right one: adding detail about retrieval cannot turn a 10 of 10 into a 3.

Two things stop it:

  • Report every rung over all N cases, never over the survivors of the rung above. The rungs are three measurements of the same cases, not a narrowing of them.
  • delivered, right is NOT the pass rate. The passing cases are delivered, right plus passed_not_covered plus passed_not_retrieved. A report that omits those last two has dropped the only numbers that reconcile the cascade with the score.

Label a rung for what it measures, not for what a reader will assume. "Can the model express an answer?" is wrong for most of what lands on the no side of the first rung. RULE_UNWRITTEN means the data is present and the model does not encode the rule for combining or filtering it; AMBIGUOUS means several candidates and no doc saying which the question means. The model can express an answer in both -- it does not say WHICH answer is meant. So the rung asks whether the model NAMES what the question needs.

Give the per-entity number too, and label both.

  • per CASE, the cascade: did this case get everything it needed
  • per ENTITY, recall: of the N entities the answers depended on, how many were delivered

They answer different questions, and the entity number is the more natural read of "how did search do". One run reported 3 of 5 cases and buried 16 required entities, 80% delivered, and the case number was the only thing a reader saw.

  • Covered? Can the model express a correct answer at all? Not computed by the run: check_coverage.py reads the MODEL rather than the answers, and the run consumes its report through --coverage. If it was not run, say unmeasured rather than leaving the row out -- an unmeasured coverage label is not evidence that the model covers the question.
  • Retrieved? Of the entities the golden answer depends on, how many did get_context hand back. Quote required_count, delivered_count and how many were ranked rather than merely mentioned in a returned source's docs.
  • Correct? The pass rate, which is the rung the other two qualify.

Two numbers here are routinely misread, so qualify them in the report or leave them out:

  • Entity precision is a PAYLOAD number, not a retrieval-quality one. Report it as what it measures: how much context the agent was handed against how much it needed. 52 entities returned per attempt for the 2 an answer used is a real cost, in tokens and in attention, and it is worth tracking. What it cannot tell you is whether retrieval worked, because the denominator is everything returned and it ignores rank: an entity that came back at rank 3 of 51 scores identically to one at rank 51.
  • For quality, report rank. "Required entities came back at median rank 3, 9 of 16 in the top 5" is a statement about retrieval. Precision alone reads as an indictment of it and is mostly a statement about how many fields the package has.
  • Report the misses split by cause, not as one recall figure. The run attributes each to never asked (no target of a type that could return it -- deterministic, because target_type is a hard filter on the server, so no documentation could have delivered it) or not retrieved (a compatible target was issued and it still did not come back, which is the docs or the wording, and diagnose separates them). Those have different owners and one recall number hides which you have.
  • A miss on a case that PASSED is still a miss, and the report says so. The answer was right by another route; name the route. Measured on one run, 3 of 4 findings were on passing cases and the route was always the same: the agent rebuilding the model's own measure inline, which held for count() and failed on the one measure carrying a grain rule.
  • Read the entities that did NOT come back, and what was asked for. That is where the retrieval signal actually is, and the misses are rarely independent. Measured on one run: 5 required entities never came back as ranked results, and 4 of the 5 had the same cause -- the agent sent only source and dimension targets, and target_type is a hard filter, so no measure could be returned however well the model documents it. One defect, five symptoms, and the same one that produced the run's only wrong answer. A per-case list of misses beside what was asked for would have shown it in a glance.
  • A PASSING case with recall below 1.0 is usually an expectation defect, not a retrieval miss: the set named one path to an answer the agent reached by another. Report the count and read it as a prompt to check required.

Query errors are worth a section of their own

Malloy the answerer wrote that would not compile or run, whether or not the case passed. These sit on tool_call events as error, and nothing else surfaces them: a case that errored twice and then recovered scores exactly like one that got it right first time, so the effort disappears.

They are the sharpest available evidence on whether the skills and the docs are leading agents astray, because each one names a specific thing the agent believed and the language does not support. Report the count, the distinct error kinds, and for each whether an existing skill already covers it. An error whose fix IS documented, in a skill the answerer did not open, is a finding about the skills rather than about the agent.

eval-diagnose does not see these today: it reads failing cases only, so an error inside a passing case is invisible to it. Say so rather than implying they were triaged.

Say what did not run, and why

A step that did not run is a result, and it is indistinguishable from having forgotten unless it is written down. The common one: improve does not run when every cluster came back owner: agent-skill or owner: dataset, because skill:eval-improve may only touch owner: model. That means the model is not at fault, which is worth a sentence rather than a silence.

Same for diagnose. If it did not run, no failure in this report has an owner yet, and the report says so rather than implying the failures are the model's.

Findings and improvement ideas

Two different things, kept apart.

  • Findings are what this run established, each with the evidence beside it. A finding with no artifact behind it is an opinion.
  • Improvement ideas are what somebody might do about them, tagged with who owns each: the model, the answer key, the analysis skills, or the harness. An idea is a proposal, not a decision, and the report does not pretend a cluster has been triaged when it has not.

Rank by how many cases each would move, and say when the answer is "nothing": a 90% run whose one failure is a known fan-out trap may need no action at all.

Report facts, do not instruct the reader. "This set has had one arm, so no noise band exists for it" is a fact they can act on. "Do not quote 90% until X" is a lecture about their own number, and they did not ask for one. State what was and was not measured; what to do with it is theirs.

A report whose evidence cannot be opened is a report nobody checks.

  • Absolute paths, never relative ones. A reader is not in your working directory, and a terminal only makes an absolute path clickable.
  • The served URL for the app and the notebook, with the real host and port the run used, not a placeholder.
  • Per case, link the FILES, not the directory: artifacts/<qid>/answer.md, artifacts/<qid>/judge.md, artifacts/<qid>/answerer.jsonl. A file:// link to a directory opens nothing in most editors, which is a dead link that looks live.
  • Move the run somewhere durable before you link it. A run written under a session scratch directory (/tmp/..., a host's per-session sandbox) is deleted when the session ends, and many editors will not linkify a path there even while it exists. Copy the run directory, artifacts included, next to the report under the set, then link that. A report whose evidence is in a temp directory is a report with no evidence by next week. This was caught twice by a reader clicking and getting nothing.
  • Write a local path bare, as /abs/path/to/file, not as [text](file:///...). Terminals and editors linkify a bare absolute path; a file:// markdown link is frequently inert, and an inert link is worse than the path in plain text because the reader has nothing to copy.
  • Never link a run LABEL or a qid as though it were a path. faa-v1-baseline-01 is a label; the path is the run directory. A backticked label that looks like a link and resolves to nothing is worse than plain text, and it has already wasted a reader's time.
  • skill:eval-loop: conducts the run this reports on.
  • skill:eval-answer: the verdicts and the ledger schema this reads.
  • skill:eval-diagnose: owners and clusters, when it ran.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-report of malloydata/publisher.

Open the folder on GitHubat commit acc1acd

Compare with similar skills

Eval Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Report compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Report this skillmalloydata/publisher116—~6.4kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC274k—~2.2kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC274k1 repos~1.7kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC274k—~1.5kAutomated safety check: PassMIT
Avoid Evalthedaviddias/Front-End-Checklist74k—~554Automated safety check: PassMIT

Similar skills

  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    274k GitHub stars~2.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    274k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    274k GitHub stars~1.5k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Avoid Eval

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Never use eval() or unsafe dynamic code execution.

    74k GitHub stars~554 tokensUpdated yesterday
    Auto-check passed
  • Nrepl Eval

    penpot/penpot

    Evaluate Clojure code via nREPL using the standalone scripts/nrepl-eval.mjs CLI tool.

    61k GitHub stars~328 tokensUpdated today
    Auto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Eval Report

What does Eval Report do?

Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden…. Eval Report is an agent skill from malloydata/publisher. Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden, unreadable judge), what was learned, and what to do next.

How do I install Eval Report in Claude Code?

Run `npx skills add malloydata/publisher --skill eval-report -a claude-code`. Or copy the skill folder (skills/eval-report in malloydata/publisher) into .claude/skills/eval-report in your project. Claude Code loads it when a task matches its description.

How do I install Eval Report in Codex?

Run `npx skills add malloydata/publisher --skill eval-report -a codex`. Or copy the skill folder (skills/eval-report in malloydata/publisher) into .agents/skills/eval-report in your project. Codex loads it when a task matches its description.

Can I use Eval Report in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill eval-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-report, .gemini/skills/eval-report, .github/skills/eval-report and .opencode/skills/eval-report in your project.

What does Eval Report need to run?

Going by SKILL.md and its folder, Eval Report needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Eval Report access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Report safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Report use?

Eval Report is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Report use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Report?

Skills that share tags, products or a category with Eval Report: Eval Harness (affaan-m/ECC, 274k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 274k stars) and Eval-Driven Development Harness (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Report?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.