Analyze Run
get-convex/convex-evals
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
$ npx skills add Archive228/loopkit --skill eval-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Archive228/loopkit eval-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-harness .claude/skills/eval-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .claude/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Archive228/loopkit/tree/main/skills/eval-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Archive228/loopkit --skill eval-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Archive228/loopkit eval-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-harness .agents/skills/eval-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .agents/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Archive228/loopkit --skill eval-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Archive228/loopkit eval-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-harness .cursor/skills/eval-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .cursor/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Archive228/loopkit.git --path skills/eval-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Archive228/loopkit --skill eval-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Archive228/loopkit eval-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-harness .gemini/skills/eval-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .gemini/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Archive228/loopkit eval-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Archive228/loopkit --skill eval-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-harness .github/skills/eval-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .github/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Archive228/loopkit --skill eval-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Archive228/loopkit eval-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-harness .opencode/skills/eval-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-harness" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/eval-harness into .opencode/skills/eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-harnessBuild a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
Eval Harness is an agent skill from Archive228/loopkit. Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.
Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Subagents. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Harness loads about 876 tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 459 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 459 words, ~876 tokens.
.claude/skills/eval-harness/SKILL.md (or your agent's skills folder).Every prompt tweak in a long-running agent looks like an improvement in the moment. The only way to know is a graded run against fixed inputs. Loopkit already ships .claude/agents/verifier.md — that is your grader. Do not rebuild it.
inputs.jsonl → runner → outputs.jsonl → verifier (per row) → verdicts.jsonl → diff vs baselineEach stage writes to disk. No stage holds the whole run in context.
One JSON object per row: {"id": "case-01", "input": "...", "expected": "..."}.
inputs-v2.jsonl) when you change it. Never edit in place — you lose the baseline.A dumb loop: for each row, call the model with the current prompt/skill, capture output, write {"id": ..., "output": ...} to outputs.jsonl. No grading here — just capture.
If the runner is smart it will bias the eval. Keep it dumb.
Fan out one subagent per row (see subagent-fanout). Each gets:
.claude/agents/verifier.md.Verifier returns strict JSON: {"pass": bool, "why": "..."}. Collect into verdicts.jsonl.
Two runs of the same eval on two prompt versions → compare pass rates per case. What matters:
A change that raises the mean but adds regressions is usually a loss — the new failures are cases you already knew worked.
verdicts.jsonl to git.The verifier is already yours. The harness is 100 lines of glue around it.
© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval-harness of Archive228/loopkit.
Open the folder on GitHubat commit 5ae033e
Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Harness this skillArchive228/loopkit | 756 | — | ~876 | Automated safety check: Pass | MIT | |
| Analyze Runget-convex/convex-evals | 129 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Wjs Evaling Voicedrop Promptsjianshuo/claude-skills | 130 | — | ~475 | Automated safety check: Pass | MIT | |
| Cxas Agent FoundryGoogleCloudPlatform/cxas-scrapi | 107 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Woo AI Smokewoocommerce/woocommerce-ios | 358 | 1 repos | ~7.4k | Automated safety check: Notes | GPL-2.0 | |
| Evevercel/vercel-plugin | 301 | 5 repos | ~1.2k | Automated safety check: Pass | Custom licence |
get-convex/convex-evals
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
jianshuo/claude-skills
A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…
GoogleCloudPlatform/cxas-scrapi
End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and…
woocommerce/woocommerce-ios
Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.
vercel/vercel-plugin
eve framework guidance for durable AI agents and agent-powered applications.
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Archive228/loopkit
Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.
Archive228/loopkit
Get JSON out of the model reliably. An agent skill from Archive228/loopkit.
Archive228/loopkit
A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…
Archive228/loopkit
Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).
Archive228/loopkit
Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.
Archive228/loopkit
Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.
Categories
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Eval Harness is an agent skill from Archive228/loopkit. Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
Eval Harness fits situations like: tasks that involve LLM evaluation; tasks that involve Subagents.
Run `npx skills add Archive228/loopkit --skill eval-harness -a claude-code`. Or copy the skill folder (skills/eval-harness in Archive228/loopkit) into .claude/skills/eval-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Archive228/loopkit --skill eval-harness -a codex`. Or copy the skill folder (skills/eval-harness in Archive228/loopkit) into .agents/skills/eval-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-harness, .gemini/skills/eval-harness, .github/skills/eval-harness and .opencode/skills/eval-harness in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Harness is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 876 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Harness: Analyze Run (get-convex/convex-evals, 129 stars), Wjs Evaling Voicedrop Prompts (jianshuo/claude-skills, 130 stars), Cxas Agent Foundry (GoogleCloudPlatform/cxas-scrapi, 107 stars) and Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 756 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.
Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.