TDD Workflow
hellangleZ/burn-in-cceverywhere-ralph
A skill your agent uses when writing new features, fixing bugs, or refactoring code.
Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.
$ npx skills add davila7/claude-code-templates --skill eval-genius -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install davila7/claude-code-templates eval-genius --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .claude/skills/eval-genius && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .claude/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-geniusType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add davila7/claude-code-templates --skill eval-genius -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install davila7/claude-code-templates eval-genius --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .agents/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .agents/skills/eval-genius && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .agents/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill eval-genius -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install davila7/claude-code-templates eval-genius --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .cursor/skills/eval-genius && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .cursor/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/davila7/claude-code-templates.git --path cli-tool/components/skills/development/eval-genius--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add davila7/claude-code-templates --skill eval-genius -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install davila7/claude-code-templates eval-genius --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .gemini/skills/eval-genius && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .gemini/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install davila7/claude-code-templates eval-geniusInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add davila7/claude-code-templates --skill eval-genius -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .github/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .github/skills/eval-genius && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .github/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill eval-genius -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install davila7/claude-code-templates eval-genius --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/cli-tool/components/skills/development/eval-genius .opencode/skills/eval-genius && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-genius" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/eval-genius into .opencode/skills/eval-genius/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-genius", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-geniusDecide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.
Eval Genius is an agent skill from davila7/claude-code-templates. Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Not ordinary unit tests.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Unit testing. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Genius loads about 2.2k tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 1,199 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 1,199 words, ~2,182 tokens.
.claude/skills/eval-genius/SKILL.md (or your agent's skills folder).An eval is a claim you are willing to defend under hostile audit. You measure to earn the right to say "this is better" and have it hold when someone sharp pushes back. Behave like a measurement engineer: state the promise, fix the bar before looking, hold everything else constant, distrust the instrument first, report the number that hurts. Assume the user may be starting from zero; plain language first, jargon when it earns it.
Three questions decide it (references/00-start-here.md): does the output vary (model,
prompt, retriever)? will it change again, and would a quiet regression cost something?
is a decision or a public claim coming? No to all: a hand spot check, stop. Yes to any:
an eval, sized to the stage the project is in.
| Stage the user is at | Instrument | Smallest useful version |
|---|---|---|
| Exploring prompts and models | Spot check | 10 inputs, eyeball |
| First working version | Smoke eval | 20 to 50 real inputs, code-checked; this run is the baseline |
| Changing one thing | Paired eval vs baseline | Same items both arms, per-item diff, bar written first |
| Merging or shipping | CI gate | Held-out items, three-way outcome, a known-bad item that must fail |
| Comparing or claiming publicly | Benchmark | Versioned dataset and harness, intervals, report |
| In production | Monitor | Same scorer on sampled live traffic |
Build the first eval at "first working version", never before, rarely after. For a
first-timer, run the one-afternoon recipe in 00-start-here.md and touch nothing else.
Identify the job, then load only that reference. Every job still passes through Step 1.
| User needs to... | Load |
|---|---|
| Know if they need an eval, where it fits, which one, or how to start | references/00-start-here.md |
| Decide what to measure at all, or the ask is "make it better" | references/01-foundation.md |
| Pick a grader or metric for a task | references/02-grading-and-metrics.md |
| Use, prompt, or trust an LLM judge | references/03-judge-calibration.md |
| Choose between an existing benchmark and a custom one | references/04-search-vs-build.md |
| Assemble items, labels, negatives, splits; contamination, overfitting | references/05-dataset-construction.md |
| Write or fix the runner, scorer, or reporter | references/06-harness-design.md |
| Put an eval in CI or a release gate | references/07-gates-and-ci.md |
| Say whether a delta is real | references/08-statistics.md |
| Write results up, or retire a benchmark | references/09-reporting.md |
| Evaluate an agent, tool use, or multi-turn task | references/10-agentic-evals.md |
| Read a result file with no prior experience | references/11-reading-results.md |
Templates in templates/ get copied into the project, never edited in place. Scripts in
scripts/ are stdlib-only CLIs with --help, exiting nonzero with a readable message:
check_gate.py (per-item diff of treatment vs baseline; exits 0 PASS, 1 FAIL,
2 CANNOT-MEASURE; refuses a comparison across mismatched fixture or judge fingerprints),
paired_bootstrap.py (paired bootstrap interval on the delta, cluster-aware),
judge_agreement.py (Cohen's kappa and PASS precision/recall of a judge vs human labels),
hash_fixture.py (the canonical fixture content hash the manifest's fixture_hash wants,
so two runs hash the same fixture to the same string). Match effort to stakes: a spot check needs Step 0 and little else; a
release gate needs the whole chain. Load references on demand, not all at once.
Copy templates/preregistration.md next to the fixture and fill it before touching
data or code. A first-timer fills promise, lever, baseline, and bar; the rest follows.
Done when all five fields are filled and a baseline is named.
Push every check that can be code-graded down to code: exact match, regex, schema,
test suite, threshold. Free text gets decomposed (required facts present, forbidden
content absent, format) before any judge sees it. Reserve a model or human judge for
the edge no assertion captures. 80% deterministic / 20% judged is trusted; 100% judged
is an opinion with error bars. Report layers separately, never one blended number. A
judge is an instrument: calibrate against human labels, blind it, randomize order, pin
model and prompt hash (references/03-judge-calibration.md).
Search before building; an established benchmark buys ground truth nobody in the room
cooked. Build custom the moment the public one rewards a proxy the system does not
target, reusing public plumbing. Score candidates on templates/benchmark-assessment-scorecard.md:
what it rewards, contamination, label-error ceiling, whether it exercises this
mechanism, whether it is maintained.
check_gate.py).Full walk in references/11-reading-results.md; each check gates the next.
paired_bootstrap.py);
interval includes zero means "not established", never "no effect".Numbers are claims with tiers, measured / estimated / aspirational, never summed across
tiers. Three sentences minimum: the bar and whether it was met; the delta with interval,
n, and flip counts; the caveat that most weakens the claim. Say when a benchmark is
self-run. Retire what fails its bar, in writing. Template: templates/eval-report.md.
templates/quality-checklist.md.Source: alexgreensh/eval-genius, Apache-2.0.
© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in cli-tool/components/skills/development/eval-genius of davila7/claude-code-templates.
Open the folder on GitHubat commit 14680ec
Eval Genius next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Genius this skilldavila7/claude-code-templates | 32k | — | ~2.2k | Automated safety check: Pass | MIT | |
| TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph | 112 | 11 repos | ~2.4k | Automated safety check: Pass | None | |
| Testing OpenLogi UIAprilNEA/OpenLogi | 23k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Go Testingcxuu/golang-skills | 170 | 1 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Contractssamchon/nestia | 2.2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Cohesion Over TestabilityEpicenterHQ/epicenter | 4.8k | — | ~2k | Automated safety check: Pass | Custom licence |
hellangleZ/burn-in-cceverywhere-ralph
A skill your agent uses when writing new features, fixing bugs, or refactoring code.
AprilNEA/OpenLogi
Verifies OpenLogi's native GPUI interface with focused tests, the component gallery and a mock agent, choosing the evidence that fits each change.
cxuu/golang-skills
A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.
samchon/nestia
Defines self-acknowledgments for production declarations and tests.
EpicenterHQ/epicenter
Collapse test-shaped production boundaries while preserving behavior and coverage.
liaohch3/claude-tap
Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.
davila7/claude-code-templates
Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.
davila7/claude-code-templates
Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.
davila7/claude-code-templates
Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.
davila7/claude-code-templates
Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.
davila7/claude-code-templates
Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.
davila7/claude-code-templates
Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.
Categories
Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Eval Genius is an agent skill from davila7/claude-code-templates. Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it.
Eval Genius fits situations like: tasks that involve Unit testing.
Run `npx skills add davila7/claude-code-templates --skill eval-genius -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/eval-genius in davila7/claude-code-templates) into .claude/skills/eval-genius in your project. Claude Code loads it when a task matches its description.
Run `npx skills add davila7/claude-code-templates --skill eval-genius -a codex`. Or copy the skill folder (cli-tool/components/skills/development/eval-genius in davila7/claude-code-templates) into .agents/skills/eval-genius in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill eval-genius -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-genius, .gemini/skills/eval-genius, .github/skills/eval-genius and .opencode/skills/eval-genius in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Genius is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Genius is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Genius: TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Testing OpenLogi UI (AprilNEA/OpenLogi, 23k stars), Go Testing (cxuu/golang-skills, 170 stars) and Contracts (samchon/nestia, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.
Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.