Agent skill

Eval Harness First

by wshobson in wshobson/agents

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.

MITAuto-check passedAI & LLM Engineering

Install Eval Harness First

skills CLI
$ npx skills add wshobson/agents --skill eval-harness-first -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents eval-harness-first --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .claude/skills/eval-harness-first && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-harness-first
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
938 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.

  • Works in 8 steps: Collect traces — production/agent spans,… → Error analysis — open coding on ≥100… → One grader per bucket — deterministic… → …
  • Starting a fine-tuning effort
  • SKILL.md covers The Gate, Building Goldens, Graders and Judge Calibration Is a…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Harness First is an agent skill from wshobson/agents. Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/grader-templates.md` and `references/judge-calibration.md`).

It sits in AI & LLM Engineering, covering LLM evaluation, Fine-tuning and Performance reviews. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Starting a fine-tuning effort
  • Converting traces into an eval set
  • Calibrating a judge against human labels

Example prompts

  • “/eval-harness-first”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Collect traces — production/agent spans, or
  2. Error analysis — open coding on ≥100 traces,
  3. One grader per bucket — deterministic first;
  4. Prioritize by frequency × severity × value.
  5. **The labeled traces feed dataset curation, minus
  6. Train.
  7. Re-run the same harness on the checkpoint —
  8. Drift detection feeds back to step 2 — new

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Harness First loads about 2k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 938 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 938 words, ~1,958 tokens.

Download SKILL.mdSave it as .claude/skills/eval-harness-first/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
eval-harness-first
description
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

Eval Harness First

The Phase 0 gate for the whole plugin: finetuning-method-selection and every downstream skill assume this harness exists before a training config gets written. The harness is not a run-end side artifact — it is the data-curation engine. The same labeled traces that build the goldens feed training data, minus an explicit holdout.

Input: production/agent traces if they exist, or a task spec if they don't, plus labelers willing to grade ≥100 examples. Output format: the eval/ directory below — goldens, graders, drift suite, and the base-model baseline that later phases gate on.

The Gate

No eval harness, no fine-tune. Skip to a training config and there is nothing to measure against, nothing to catch regressions, and no labeled data to train on. The flywheel:

  1. Collect traces — production/agent spans, or synthetic tasks if none exist yet.
  2. Error analysis — open coding on ≥100 traces, axial coding into 4–8 failure buckets.
  3. One grader per bucket — deterministic first; calibrated LLM-judge only for genuinely subjective criteria.
  4. Prioritize by frequency × severity × value.
  5. The labeled traces feed dataset curation, minus an explicit holdout. Every eval/goldens.jsonl ID stays excluded from training data by ID.
  6. Train.
  7. Re-run the same harness on the checkpoint — not a different, looser one.
  8. Drift detection feeds back to step 2 — new production failure modes re-open error analysis.

Steps 2–4 build the harness; steps 5–8 are why it must exist first — it is both the training data source and the checkpoint's exit gate.

Building Goldens

  • From traces, when they exist: run error analysis — open coding on ≥100 real traces (read them, tag failures in your own words, no fixed taxonomy yet), then axial coding to collapse those tags into 4–8 named failure buckets. Fewer than 4 means the coding pass was too shallow; more than 8 means buckets need merging. Exception: single-failure-surface tasks (e.g. strict-schema extraction) may land at 1–2 buckets with per-field sub-metrics inside one grader — don't invent artificial splits with no evidence behind them.
  • Synthetic, when traces don't exist yet: dimension-based generation — enumerate the axes that matter (task type, difficulty, edge case, persona) and sample the cross-product; free- generated prompts cluster around whatever's easiest to write.
  • Goldens are versioned like code — commit eval/goldens.jsonl, diff it in review, tag it per release. It doubles as the CI regression suite.

Graders

One grader per failure bucket from error analysis — not one for the whole eval set. A single blended score hides which bucket regressed.

  • Deterministic first. Regex, schema validation, or execution checks are cheaper, reproducible, and need no calibration.
  • LLM-judge only for genuinely subjective criteria — tone, faithfulness, "which response is better" — where no deterministic check can express it.
  • Binary pass/fail over Likert. A 1–5 or 1–10 scale is noisier to calibrate and harder to apply consistently; collapse to pass/fail.
  • Drift-suite MMLU-style scoring: prefer logprob over generate-and-extract — a tight token budget makes generate-and-extract parse-brittle for models that preamble, conflating format compliance with the knowledge being measured. Templates for all four grader shapes and this scoring note: references/grader-templates.md.
Show full SKILL.md (438 more words)Show less

Judge Calibration Is a Prerequisite

Any bucket routed to an LLM-judge needs calibration before its verdicts count for anything beyond exploration — a hard prerequisite, not a nice-to-have. N/A when no bucket routes to a judge — an all-deterministic harness has nothing to calibrate; state that rather than leaving this section unaddressed.

  • Label ≥100 items, split train/dev/sealed test (report once, no re-touching after).
  • Report TPR and TNR, not one blended accuracy number — a judge can hit 90% by always saying "pass" on a skewed set.
  • Pin the judge to a fixed model snapshot and recalibrate on judge-model change, quarterly regardless.
  • The judge must come from a different model family than the model under test.
  • A judge that misses the agreed TPR/TNR bar ships advisory-only — flags for human review, never gates a promotion. Full protocol, bias correction, and recalibration checklist: references/judge-calibration.md.

The Baseline

Before Phase 1 (method selection) starts, run the full harness — goldens plus the capability-drift suite — against the unmodified base model. This is the number every later checkpoint gets compared against.

eval/baseline-<model>.json is the gate token. No baseline file, no comparison basis for checkpoint-promotion — a checkpoint that "looks better" against nothing measured isn't a finding.

Directory Contract

eval/
├── goldens.jsonl          # labeled traces + synthetic goldens, versioned
├── graders/                # one module per failure bucket
│   ├── schema_compliance.py
│   ├── exact_match.py
│   └── rubric_judge.py
├── drift-suite.yaml        # frozen benchmarks + 200-500 domain-adjacent items
└── baseline-<model>.json   # gate token: harness + drift suite vs the base model
runs/
└── <run-id>/
    └── results.json         # per-run harness output, one per checkpoint

eval/ persists across runs and lives outside runs/ — the fixed measuring stick, not a run artifact. runs/ is disposable; eval/ is not. Never let a run script write into eval/. Canonical location: every per-trace results.json — the Phase 0 baseline included — lives at runs/<run-id>/results.json, never under eval/runs/...; an instruction requesting the latter is wrong, not this contract.

Phase 0 Exit Checklist

Before finetuning-method-selection, confirm:

  1. ≥100 traces open-coded; 4–8 failure buckets (N/A floor for synthetic goldens on a single-failure- surface task — see the Building Goldens exception; bucket count then comes from post-baseline error analysis instead).
  2. eval/goldens.jsonl committed and versioned.
  3. One grader per bucket, deterministic first.
  4. Judges calibrated — TPR/TNR, snapshot pinned, different family (N/A when no bucket routes to an LLM-judge; state that explicitly).
  5. eval/drift-suite.yaml frozen.
  6. eval/baseline-<model>.json written.

Missing any of the six (or its stated N/A)? Not Phase 0 complete — /finetune checks the baseline file before a run.

General-purpose evaluation guidance (dashboards, A/B testing, non-fine-tuning harnesses) lives in the llm-application-dev plugin's llm-evaluation skill — this skill covers only the fine-tuning coupling: goldens that double as training data, and the baseline that gates a checkpoint.

  • finetuning-method-selection — routes here first.
  • dataset-curation — formats these traces into training rows.
  • trace-to-training-data — turns graded traces into training examples.
  • checkpoint-promotion — consumes baseline-<model>.json, re-runs this harness on each candidate checkpoint.

References

  • references/grader-templates.md — runnable grader examples per shape, plus a drift-suite.yaml example and MMLU logprob-scoring note.
  • references/judge-calibration.md — the calibration protocol, including the all- deterministic N/A path.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/eval-harness-first of wshobson/agents.

  • SKILL.md
  • references/grader-templates.md
  • references/judge-calibration.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Eval Harness First next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Harness First compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Harness First this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Jd Gap Analysisstarkyru/learn-ai107—~1.9kAutomated safety check: PassMIT
ML Training Run VerifierLeeroo-AI/superml195—~3.8kAutomated safety check: PassApache-2.0
Advanced Evaluationguanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMIT
Audit Sft Data Qualitytokenbender/agent-guides367—~2.7kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT

Similar skills

  • Jd Gap Analysis

    starkyru/learn-ai

    Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

    107 GitHub stars~1.9k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • ML Training Run Verifier

    Leeroo-AI/superml

    Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

    195 GitHub stars~3.8k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Audit Sft Data Quality

    tokenbender/agent-guides

    Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

    367 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub stars~1.7k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Evaluating With Leakage Gates

    maziyarpanahi/openmed

    Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1.

    5.5k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed

Questions about Eval Harness First

What does Eval Harness First do?

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Eval Harness First is an agent skill from wshobson/agents. Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.

When should I use Eval Harness First?

Eval Harness First fits situations like: starting a fine-tuning effort; converting traces into an eval set; calibrating a judge against human labels.

How do I install Eval Harness First in Claude Code?

Run `npx skills add wshobson/agents --skill eval-harness-first -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/eval-harness-first in wshobson/agents) into .claude/skills/eval-harness-first in your project. Claude Code loads it when a task matches its description.

How do I install Eval Harness First in Codex?

Run `npx skills add wshobson/agents --skill eval-harness-first -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/eval-harness-first in wshobson/agents) into .agents/skills/eval-harness-first in your project. Codex loads it when a task matches its description.

Can I use Eval Harness First in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill eval-harness-first -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-harness-first, .gemini/skills/eval-harness-first, .github/skills/eval-harness-first and .opencode/skills/eval-harness-first in your project.

What does Eval Harness First need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Harness First is instructions for the agent only.

Does Eval Harness First access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Harness First safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Harness First use?

Eval Harness First is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Harness First use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Eval Harness First?

Skills that share tags, products or a category with Eval Harness First: Jd Gap Analysis (starkyru/learn-ai, 107 stars), ML Training Run Verifier (Leeroo-AI/superml, 195 stars), Advanced Evaluation (guanyang/open-agent-hub, 977 stars) and Audit Sft Data Quality (tokenbender/agent-guides, 367 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Harness First?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,305 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.