Agent skill

Eval Genius

by davila7 in davila7/claude-code-templates

Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.

MITAuto-check passedTesting & QA

Install Eval Genius

skills CLI
$ npx skills add davila7/claude-code-templates --skill eval-genius -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates eval-genius --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .claude/skills/eval-genius && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-genius
GitHub stars
32k
Token cost
~2.2k tokens
SKILL.md length
1,199 words
Files
1
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.

  • Works in 7 steps: Does this need an eval, and where does… → Foundation (five minutes, never skipped) → Choose the grader (deterministic first) → …
  • Tasks that involve Unit testing
  • SKILL.md covers Step 0: Does this need an…, Route the request, Step 1: Foundation (five… and Step 2: Choose the grader…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Genius is an agent skill from davila7/claude-code-templates. Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Not ordinary unit tests.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Unit testing

Example prompts

  • “/eval-genius”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Does this need an eval, and where does it go?
  2. Foundation (five minutes, never skipped)
  3. Choose the grader (deterministic first)
  4. Build or adopt
  5. Run under hard rules (a run that breaks one is not a result)
  6. Read the result (in this order)
  7. Report honestly

What it can do on your machine

Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Genius loads about 2.2k tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 1,199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 1,199 words, ~2,182 tokens.

Download SKILL.mdSave it as .claude/skills/eval-genius/SKILL.md (or your agent's skills folder).
name
eval-genius
description
Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Not ordinary unit tests.

Eval Genius

An eval is a claim you are willing to defend under hostile audit. You measure to earn the right to say "this is better" and have it hold when someone sharp pushes back. Behave like a measurement engineer: state the promise, fix the bar before looking, hold everything else constant, distrust the instrument first, report the number that hurts. Assume the user may be starting from zero; plain language first, jargon when it earns it.


Step 0: Does this need an eval, and where does it go?

Three questions decide it (references/00-start-here.md): does the output vary (model, prompt, retriever)? will it change again, and would a quiet regression cost something? is a decision or a public claim coming? No to all: a hand spot check, stop. Yes to any: an eval, sized to the stage the project is in.

Stage the user is atInstrumentSmallest useful version
Exploring prompts and modelsSpot check10 inputs, eyeball
First working versionSmoke eval20 to 50 real inputs, code-checked; this run is the baseline
Changing one thingPaired eval vs baselineSame items both arms, per-item diff, bar written first
Merging or shippingCI gateHeld-out items, three-way outcome, a known-bad item that must fail
Comparing or claiming publiclyBenchmarkVersioned dataset and harness, intervals, report
In productionMonitorSame scorer on sampled live traffic

Build the first eval at "first working version", never before, rarely after. For a first-timer, run the one-afternoon recipe in 00-start-here.md and touch nothing else.

Route the request

Identify the job, then load only that reference. Every job still passes through Step 1.

User needs to...Load
Know if they need an eval, where it fits, which one, or how to startreferences/00-start-here.md
Decide what to measure at all, or the ask is "make it better"references/01-foundation.md
Pick a grader or metric for a taskreferences/02-grading-and-metrics.md
Use, prompt, or trust an LLM judgereferences/03-judge-calibration.md
Choose between an existing benchmark and a custom onereferences/04-search-vs-build.md
Assemble items, labels, negatives, splits; contamination, overfittingreferences/05-dataset-construction.md
Write or fix the runner, scorer, or reporterreferences/06-harness-design.md
Put an eval in CI or a release gatereferences/07-gates-and-ci.md
Say whether a delta is realreferences/08-statistics.md
Write results up, or retire a benchmarkreferences/09-reporting.md
Evaluate an agent, tool use, or multi-turn taskreferences/10-agentic-evals.md
Read a result file with no prior experiencereferences/11-reading-results.md

Templates in templates/ get copied into the project, never edited in place. Scripts in scripts/ are stdlib-only CLIs with --help, exiting nonzero with a readable message: check_gate.py (per-item diff of treatment vs baseline; exits 0 PASS, 1 FAIL, 2 CANNOT-MEASURE; refuses a comparison across mismatched fixture or judge fingerprints), paired_bootstrap.py (paired bootstrap interval on the delta, cluster-aware), judge_agreement.py (Cohen's kappa and PASS precision/recall of a judge vs human labels), hash_fixture.py (the canonical fixture content hash the manifest's fixture_hash wants, so two runs hash the same fixture to the same string). Match effort to stakes: a spot check needs Step 0 and little else; a release gate needs the whole chain. Load references on demand, not all at once.


Step 1: Foundation (five minutes, never skipped)

Copy templates/preregistration.md next to the fixture and fill it before touching data or code. A first-timer fills promise, lever, baseline, and bar; the rest follows.

  1. Promise. One plain sentence: what does the system promise, what is "better"?
  2. Variables. Levers (what changes), outcomes (what is watched), controls (what is frozen). One lever per comparison; an unclassifiable variable means stop.
  3. Placement. A decision point, a risky seam, or a public claim; elsewhere, a spot check or nothing.
  4. Weight. Spot check, eval, benchmark, or monitor (Step 0 table).
  5. Bar. The threshold in numbers, the falsifier (what proves the change useless), and the outlier rule. Written before any run.

Done when all five fields are filled and a baseline is named.

Step 2: Choose the grader (deterministic first)

Push every check that can be code-graded down to code: exact match, regex, schema, test suite, threshold. Free text gets decomposed (required facts present, forbidden content absent, format) before any judge sees it. Reserve a model or human judge for the edge no assertion captures. 80% deterministic / 20% judged is trusted; 100% judged is an opinion with error bars. Report layers separately, never one blended number. A judge is an instrument: calibrate against human labels, blind it, randomize order, pin model and prompt hash (references/03-judge-calibration.md).

Show full SKILL.md (470 more words)Show less

Step 3: Build or adopt

Search before building; an established benchmark buys ground truth nobody in the room cooked. Build custom the moment the public one rewards a proxy the system does not target, reusing public plumbing. Score candidates on templates/benchmark-assessment-scorecard.md: what it rewards, contamination, label-error ceiling, whether it exercises this mechanism, whether it is maintained.

Step 4: Run under hard rules (a run that breaks one is not a result)

  • Freeze the fixture. Same items, corpus, snapshot, seeds; only the lever varies. Refuse comparisons across mismatched fingerprints.
  • Three-way outcome. PASS, FAIL, or CANNOT-MEASURE. A crash, missing baseline, or fingerprint mismatch is CANNOT-MEASURE, never FAIL and never PASS (check_gate.py).
  • Prove the treatment arm is live before the run.
  • Cover the negative space: items where the right answer is to refuse or return nothing.
  • Repeat and show spread. Noise wider than the effect means no result yet.
  • Verify the verifier. A known-bad case must go red before green is trusted.

Step 5: Read the result (in this order)

Full walk in references/11-reading-results.md; each check gates the next.

  1. Did it run? Exit code before score; negative control failed.
  2. Against the written bar, not against hope. Above: candidate win. Below the falsifier: rejected. Between: not established.
  3. Bigger than noise? Paired interval on the delta (paired_bootstrap.py); interval includes zero means "not established", never "no effect".
  4. Items, not averages. Read regressions first; reproduce one flip by hand.
  5. Surprised? A 0%, a 99%, a thirty-point jump is a harness bug until proven otherwise. Never tune the system against a suspect gauge; fix the gauge or stop.
  6. Layers and cost separately. A win that doubled cost is a trade.

Step 6: Report honestly

Numbers are claims with tiers, measured / estimated / aspirational, never summed across tiers. Three sentences minimum: the bar and whether it was met; the delta with interval, n, and flip counts; the caveat that most weakens the claim. Say when a benchmark is self-run. Retire what fails its bar, in writing. Template: templates/eval-report.md.


Anti-patterns (named so they can be refused)

  • Metric-first. A dataset and scale chosen before the promise is written.
  • Eval too early or too late. Measuring a prototype still in flux, or a shipped system with twenty unattributed changes behind it.
  • Blended score. One number hiding which layer moved.
  • Post-hoc bar. Threshold decided after the result is known.
  • Fixture drift. Comparing across corpora, caches, or snapshots.
  • Cache blindness. Reading a pre-built cache never exercises the write path.
  • Silent crash. Harness error recorded as a score.
  • Oracle judge. Uncalibrated model judge treated as ground truth.
  • Rubric-author bias. Whoever built the system also wrote the rubric, alone.
  • Run until green. Repeating a noisy eval until one run passes.
  • Gate-set tuning / overfitting. Iterating on the held-out items the gate uses.
  • Run-completion checklist before calling anything done: templates/quality-checklist.md.

Source: alexgreensh/eval-genius, Apache-2.0.

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-tool/components/skills/development/eval-genius of davila7/claude-code-templates.

Open the folder on GitHubat commit 14680ec

Compare with similar skills

Eval Genius next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Genius compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Genius this skilldavila7/claude-code-templates32k—~2.2kAutomated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Testing OpenLogi UIAprilNEA/OpenLogi23k—~1.1kAutomated safety check: PassApache-2.0
Go Testingcxuu/golang-skills1701 repos~1.3kAutomated safety check: PassApache-2.0
Contractssamchon/nestia2.2k—~1.3kAutomated safety check: PassMIT
Cohesion Over TestabilityEpicenterHQ/epicenter4.8k—~2kAutomated safety check: PassCustom licence

Similar skills

  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Testing OpenLogi UI

    AprilNEA/OpenLogi

    Verifies OpenLogi's native GPUI interface with focused tests, the component gallery and a mock agent, choosing the evidence that fits each change.

    23k GitHub stars~1.1k tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Go Testing

    cxuu/golang-skills

    A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.

    170 GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Contracts

    samchon/nestia

    Defines self-acknowledgments for production declarations and tests.

    2.2k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Cohesion Over Testability

    EpicenterHQ/epicenter

    Collapse test-shaped production boundaries while preserving behavior and coverage.

    4.8k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • JS-in-HTML Testing

    liaohch3/claude-tap

    Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.

    3.3k GitHub stars~924 tokensUpdated 16 days ago
    Testing & QAAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Categories

Questions about Eval Genius

What does Eval Genius do?

Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Eval Genius is an agent skill from davila7/claude-code-templates. Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.

When should I use Eval Genius?

Eval Genius fits situations like: tasks that involve Unit testing.

How do I install Eval Genius in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill eval-genius -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/eval-genius in davila7/claude-code-templates) into .claude/skills/eval-genius in your project. Claude Code loads it when a task matches its description.

How do I install Eval Genius in Codex?

Run `npx skills add davila7/claude-code-templates --skill eval-genius -a codex`. Or copy the skill folder (cli-tool/components/skills/development/eval-genius in davila7/claude-code-templates) into .agents/skills/eval-genius in your project. Codex loads it when a task matches its description.

Can I use Eval Genius in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill eval-genius -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-genius, .gemini/skills/eval-genius, .github/skills/eval-genius and .opencode/skills/eval-genius in your project.

What does Eval Genius need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Genius is instructions for the agent only.

Does Eval Genius access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Eval Genius safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Genius use?

Eval Genius is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Genius use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Genius?

Skills that share tags, products or a category with Eval Genius: TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Testing OpenLogi UI (AprilNEA/OpenLogi, 23k stars), Go Testing (cxuu/golang-skills, 170 stars) and Contracts (samchon/nestia, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Genius?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.