Agent skill

Eval Harness

by Archive228 in Archive228/loopkit

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

MITAuto-check passedAI & LLM Engineering

Install Eval Harness

skills CLI
$ npx skills add Archive228/loopkit --skill eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Archive228/loopkit eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-harness .claude/skills/eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-harness
GitHub stars
756
Token cost
~876 tokens
SKILL.md length
459 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

  • Works in 3 steps: inputs.jsonl → runner → verifier
  • Tasks that involve LLM evaluation
  • SKILL.md covers The three-stage loop, Stage 1 — inputs.jsonl, Stage 2 — runner and Stage 3 — verifier, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Harness is an agent skill from Archive228/loopkit. Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.

Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Subagents. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Subagents

Example prompts

  • “/eval-harness”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. inputs.jsonl
  2. runner
  3. verifier

What it can do on your machine

Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Harness loads about 876 tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 459 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~876

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 459 words, ~876 tokens.

Download SKILL.mdSave it as .claude/skills/eval-harness/SKILL.md (or your agent's skills folder).
name
eval-harness
description
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.
when_to_use
tuning a prompt, changing a skill, comparing two models, regression-testing a workflow, "does this actually work better?"

Eval Harness

Every prompt tweak in a long-running agent looks like an improvement in the moment. The only way to know is a graded run against fixed inputs. Loopkit already ships .claude/agents/verifier.md — that is your grader. Do not rebuild it.

The three-stage loop

inputs.jsonl  →  runner  →  outputs.jsonl  →  verifier (per row)  →  verdicts.jsonl  →  diff vs baseline

Each stage writes to disk. No stage holds the whole run in context.

Stage 1 — inputs.jsonl

One JSON object per row: {"id": "case-01", "input": "...", "expected": "..."}.

  • 20-100 cases is enough for a signal. More is nice, not required.
  • Include known-hard cases, edge cases, and a couple of trivial ones as sanity anchors.
  • Freeze the file. Rev the eval with a suffix (inputs-v2.jsonl) when you change it. Never edit in place — you lose the baseline.

Stage 2 — runner

A dumb loop: for each row, call the model with the current prompt/skill, capture output, write {"id": ..., "output": ...} to outputs.jsonl. No grading here — just capture.

  • Same temperature every run (usually 0 for evals).
  • Same seed / model version.
  • Log the git SHA of the prompt/skill under test in the file header.

If the runner is smart it will bias the eval. Keep it dumb.

Stage 3 — verifier

Fan out one subagent per row (see subagent-fanout). Each gets:

  • The input.
  • The expected output (or spec).
  • The actual output.
  • The verifier system prompt from .claude/agents/verifier.md.

Verifier returns strict JSON: {"pass": bool, "why": "..."}. Collect into verdicts.jsonl.

Diff vs baseline

Two runs of the same eval on two prompt versions → compare pass rates per case. What matters:

  • Overall pass rate — the headline.
  • Regressions — cases that were green and went red. These block ship.
  • New passes — cases that were red and went green. These justify ship.
  • Flappy cases — inconsistent across reruns. Investigate; may be genuine model nondeterminism or a bad case.

A change that raises the mean but adds regressions is usually a loss — the new failures are cases you already knew worked.

Show full SKILL.md (153 more words)Show less

Red flags

  • Grader is the same model that produced the output, with the same prompt. Self-grading is lenient. Use a different persona at minimum; ideally a different model tier.
  • Eval passes 100% on day one. The cases are too easy, or the grader is a rubber stamp. Add adversarial cases.
  • Eval takes >30 minutes. Fan out. A serial 100-case eval is a serial 100-case bottleneck.
  • Grader sees your prompt under test. It will grade what you wanted, not what happened. Feed it only spec + input + output.
  • Baseline lost. Without baseline, "improvement" is vibes. Commit verdicts.jsonl to git.

When NOT to do this

  • One-off script — build a checklist, not a harness.
  • Prompt that changes daily and won't stabilize — evals need a fixed target.
  • Task where "correct" isn't checkable (open-ended creative writing) — use human eval or a rubric-based grader, not pass/fail.

The verifier is already yours. The harness is 100 lines of glue around it.

© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-harness of Archive228/loopkit.

Open the folder on GitHubat commit 5ae033e

Compare with similar skills

Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Harness this skillArchive228/loopkit756—~876Automated safety check: PassMIT
Analyze Runget-convex/convex-evals129—~2.1kAutomated safety check: PassApache-2.0
Wjs Evaling Voicedrop Promptsjianshuo/claude-skills130—~475Automated safety check: PassMIT
Cxas Agent FoundryGoogleCloudPlatform/cxas-scrapi107—~2.4kAutomated safety check: PassApache-2.0
Woo AI Smokewoocommerce/woocommerce-ios3581 repos~7.4kAutomated safety check: NotesGPL-2.0
Evevercel/vercel-plugin3015 repos~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Analyze Run

    get-convex/convex-evals

    Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

    129 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Wjs Evaling Voicedrop Prompts

    jianshuo/claude-skills

    A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…

    130 GitHub stars~475 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cxas Agent Foundry

    GoogleCloudPlatform/cxas-scrapi

    End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and…

    107 GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub starsUsed in 1 repo~7.4k tokens
    EducationAuto-check: notes
  • Eve

    vercel/vercel-plugin

    Official

    eve framework guidance for durable AI agents and agent-powered applications.

    301 GitHub starsUsed in 5 repos~1.2k tokens
    DevOps & CloudAuto-check passed
  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed

More from Archive228/loopkit

All 43 skills in this repo
  • Hitl Escalate

    Archive228/loopkit

    Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.

    756 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Structured Output

    Archive228/loopkit

    Get JSON out of the model reliably. An agent skill from Archive228/loopkit.

    756 GitHub stars~830 tokensUpdated 2 mo ago
    Auto-check passed
  • Using Loopkit

    Archive228/loopkit

    A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…

    756 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Active Memory Reminder

    Archive228/loopkit

    Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).

    756 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Feature List JSON

    Archive228/loopkit

    Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.

    756 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Prompt Caching

    Archive228/loopkit

    Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.

    756 GitHub stars~735 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Eval Harness

What does Eval Harness do?

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Eval Harness is an agent skill from Archive228/loopkit. Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

When should I use Eval Harness?

Eval Harness fits situations like: tasks that involve LLM evaluation; tasks that involve Subagents.

How do I install Eval Harness in Claude Code?

Run `npx skills add Archive228/loopkit --skill eval-harness -a claude-code`. Or copy the skill folder (skills/eval-harness in Archive228/loopkit) into .claude/skills/eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install Eval Harness in Codex?

Run `npx skills add Archive228/loopkit --skill eval-harness -a codex`. Or copy the skill folder (skills/eval-harness in Archive228/loopkit) into .agents/skills/eval-harness in your project. Codex loads it when a task matches its description.

Can I use Eval Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-harness, .gemini/skills/eval-harness, .github/skills/eval-harness and .opencode/skills/eval-harness in your project.

What does Eval Harness need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Harness is instructions for the agent only.

Does Eval Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Harness use?

Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Harness use?

About 876 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Harness?

Skills that share tags, products or a category with Eval Harness: Analyze Run (get-convex/convex-evals, 129 stars), Wjs Evaling Voicedrop Prompts (jianshuo/claude-skills, 130 stars), Cxas Agent Foundry (GoogleCloudPlatform/cxas-scrapi, 107 stars) and Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Harness?

Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 756 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.

Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.