Eval Harness
Archive228/loopkit
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
$ npx skills add get-convex/convex-evals --skill analyze-run -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install get-convex/convex-evals analyze-run --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/analyze-run .claude/skills/analyze-run && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .claude/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-runType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add get-convex/convex-evals --skill analyze-run -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install get-convex/convex-evals analyze-run --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/analyze-run .agents/skills/analyze-run && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .agents/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill analyze-run -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install get-convex/convex-evals analyze-run --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/analyze-run .cursor/skills/analyze-run && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .cursor/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/get-convex/convex-evals.git --path .cursor/skills/analyze-run--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add get-convex/convex-evals --skill analyze-run -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install get-convex/convex-evals analyze-run --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/analyze-run .gemini/skills/analyze-run && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .gemini/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install get-convex/convex-evals analyze-runInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add get-convex/convex-evals --skill analyze-run -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/analyze-run .github/skills/analyze-run && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .github/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill analyze-run -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install get-convex/convex-evals analyze-run --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/analyze-run .opencode/skills/analyze-run && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-run" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/analyze-run into .opencode/skills/analyze-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-run", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-runAnalyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
Analyze Run is an agent skill from get-convex/convex-evals. Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Use when the user asks to analyze an entire run, review all failures in a run, or wants to understand why a model scored poorly.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Subagents. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit aa7b0ab. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curljqnpxgitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
fabulous-panther-525.convex.cloudconvex-evals.netlify.appFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Run loads about 2.1k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 764 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from get-convex/convex-evals at commit aa7b0ab, republished under its Apache-2.0 licence (© get-convex). 764 words, ~2,150 tokens.
.claude/skills/analyze-run/SKILL.md (or your agent's skills folder).https://convex-evals.netlify.app/experiment/.../run/$runId/...Extract the run ID from the visualizer URL. The URL pattern is:
/experiment/$experimentId/run/$runId/...The $runId is the Convex document ID (e.g. jn7922j1w29pdxm76bj9ps0enx80mg9e).
This skill covers coding runs only. For a decision run (visualizer URLs under /decision/run/), the queries below throw "This operation requires a coding run", which the public API reports as a bare "Server Error".
Reports are stored in reports/{provider}/{model}/, where {provider}/{model} is the run's model slug (the model field from Step 3), e.g. reports/anthropic/claude-opus-4.8/ for anthropic/claude-opus-4.8. Don't use the run's provider field. It stores the OpenRouter endpoint provider (or openrouter when discovery failed), which can differ from the slug's vendor.
List the directory for the model being analyzed and read the most recent report(s). This gives you:
Reference prior findings when the same eval fails again — note whether it's a repeat and whether any prior fix should have resolved it.
Use the public production query runs:getRunDetails over HTTP. It needs no login:
URL=https://fabulous-panther-525.convex.cloud
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d '{"path":"runs:getRunDetails","args":{"runId":"<runId>"}}' > /tmp/run-<runId>.json
jq '.value | {model, provider, experiment, status: .status.kind,
totalEvals: (.evals | length),
passedCount: ([.evals[] | select(.status.kind == "passed")] | length),
failedEvals: [.evals[] | select(.status.kind == "failed") | {_id, evalPath,
failureReason: .status.failureReason,
failedStep: ([.steps[] | select(.status.kind == "failed") | {name, failureReason: .status.failureReason}] | first)}]}' /tmp/run-<runId>.jsonThis returns:
model (the slug), provider, experiment (null means default), status -- run metadatatotalEvals, passedCount -- overall statsfailedEvals -- array of failed evals, each with _id, evalPath, failureReason, and failedStep (which step failed and its error)If there are no failures, report that all evals passed and stop.
Don't use npx convex run --prod for this. The debug functions (debugQueries:getFailedEvalsForRun, debug:getEvalDebugInfo) are internal, and agents usually hit team SSO ("Single-sign on login is required").
For each failed eval, spawn a sub-agent (up to 4 in parallel) with this prompt template. Fill in <REPO_ROOT> with the absolute path from git rev-parse --show-toplevel. If a sub-agent can't fetch its eval, give Mike this command to run and paste back: cd evalScores && npx convex run --prod debug:getEvalDebugInfo '{"evalId": "<EVAL_ID>"}'.
You are investigating a failing eval from the convex-evals system.
The repo root is <REPO_ROOT>. Work in a new temp directory, not the repo.
Fetch the eval from the public production API (no login needed):
URL=https://fabulous-panther-525.convex.cloud
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d '{"path":"runs:getRunDetails","args":{"runId":"<RUN_ID>"}}' \
| jq '.value.evals[] | select(._id == "<EVAL_ID>")' > eval.json
eval.json has the task text (task), status (failureReason, outputStorageId),
evalSourceStorageId, and steps. Get a download URL for each storage ID:
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d '{"path":"runs:getOutputUrl","args":{"storageId":"<STORAGE_ID>"}}' | jq -r .value
Download status.outputStorageId to output.zip and evalSourceStorageId to
source.zip with curl -s -o, then unzip each into output/ and source/.
Don't use npx convex run --prod. If a request fails, stop and report the error.
Then analyze the result:
1. Which step failed and what was the exact error?
2. Look at the model's generated code in output/.
3. Look at the expected answer and grader in source/.
4. Look at the task description in eval.json's task field.
5. Is this a genuine model mistake, or is the test/lint/task unfair?
Classify the failure as one of:
- MODEL_FAULT: The model genuinely got it wrong
- OVERLY_STRICT: The eval/lint/test requirements are unreasonable for what was asked
- AMBIGUOUS_TASK: The task description is unclear and the model's interpretation was reasonable
- KNOWN_GAP: A known limitation of this eval that affects all models (e.g. the Convex API returns fields the model can't predict without being told)
Return a structured summary:
- Eval: <name> (<category>)
- Failed step: <step name>
- Error: <one-line error summary>
- Classification: <one of the above>
- Reasoning: <2-3 sentences explaining your classification>
- Model output snippet: <the relevant problematic code, if applicable>
- Expected code snippet: <what the answer looks like, if applicable>Once all sub-agents return, build the analysis:
For each failure, list: eval name, failed step, classification, one-line reasoning.
Look for patterns across failures:
Group recommendations by type:
Always create a report file at:
reports/{provider}/{model}/{runIdPrefix}_{date}.mdFor example: reports/anthropic/claude-opus-4.8/jn72t14a_2026-09-29.md for a run of anthropic/claude-opus-4.8.
{provider}/{model} is the run's model slug, as in Step 2. The runIdPrefix is the first 8 characters of the run ID.
The report should contain:
Present the full analysis to the user. End with:
"These are my findings. Would you like me to implement any of these recommendations, or would you like to discuss specific failures in more detail?"
Do NOT make any code/config changes until the user explicitly asks.
If the user asks you to implement any recommendations, update the report file's "Actions taken" section after making the changes. Record:
This ensures future analysis sessions can see which recommendations were already acted on and avoid re-recommending changes that have already been made.
© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/analyze-run of get-convex/convex-evals.
Open the folder on GitHubat commit aa7b0ab
Analyze Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Run this skillget-convex/convex-evals | 130 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Eval HarnessArchive228/loopkit | 755 | — | ~876 | Automated safety check: Pass | MIT | |
| Wjs Evaling Voicedrop Promptsjianshuo/claude-skills | 131 | — | ~475 | Automated safety check: Pass | MIT | |
| Woo AI Smokewoocommerce/woocommerce-ios | 358 | — | ~7.4k | Automated safety check: Notes | GPL-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 5 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT |
Archive228/loopkit
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
jianshuo/claude-skills
A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…
woocommerce/woocommerce-ios
Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
nicobailon/pi-prompt-template-model
Write and run custom Pi prompt templates (slash commands) for this extension.
get-convex/convex-evals
Design, implement, validate, and calibrate a new eval for the convex-evals suite.
get-convex/convex-evals
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
get-convex/convex-evals
Investigate a single failing eval from the convex-evals system.
get-convex/convex-evals
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
Categories
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Analyze Run is an agent skill from get-convex/convex-evals. Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
Analyze Run fits situations like: the user asks to analyze an entire run; review all failures in a run; wants to understand why a model scored poorly.
Run `npx skills add get-convex/convex-evals --skill analyze-run -a claude-code`. Or copy the skill folder (.cursor/skills/analyze-run in get-convex/convex-evals) into .claude/skills/analyze-run in your project. Claude Code loads it when a task matches its description.
Run `npx skills add get-convex/convex-evals --skill analyze-run -a codex`. Or copy the skill folder (.cursor/skills/analyze-run in get-convex/convex-evals) into .agents/skills/analyze-run in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill analyze-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-run, .gemini/skills/analyze-run, .github/skills/analyze-run and .opencode/skills/analyze-run in your project.
Going by SKILL.md and its folder, Analyze Run needs the command-line tools its instructions call (curl, jq, npx and git). Our summary lists: Node.js.
SKILL.md names 2 domains. In commands or code: fabulous-panther-525.convex.cloud and convex-evals.netlify.app; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyze Run: Eval Harness (Archive228/loopkit, 755 stars), Wjs Evaling Voicedrop Prompts (jianshuo/claude-skills, 131 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Agent Builder (shareAI-lab/learn-claude-code, 78k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 130 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 9, 2026.
Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.