Agent skill

Analyze Run

by get-convex in get-convex/convex-evals

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Analyze Run

skills CLI
$ npx skills add get-convex/convex-evals --skill analyze-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install get-convex/convex-evals analyze-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/analyze-run .claude/skills/analyze-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-run
GitHub stars
130
Token cost
~2.1k tokens
SKILL.md length
764 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

  • Works in 5 steps: Get the run ID → Check previous reports for this model → Fetch the failure summary → …
  • The user asks to analyze an entire run
  • SKILL.md covers When to use, Step 1: Get the run ID, Step 2: Check previous reports… and Step 3: Fetch the failure…, plus 2 more sections
  • Calls curl, jq and npx; reaches fabulous-panther-525.convex.cloud and convex-evals.netlify.app

What it does

Analyze Run is an agent skill from get-convex/convex-evals. Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Use when the user asks to analyze an entire run, review all failures in a run, or wants to understand why a model scored poorly.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Subagents. The licence is Apache-2.0.

When your agent uses it

  • The user asks to analyze an entire run
  • Review all failures in a run
  • Wants to understand why a model scored poorly

Example prompts

  • “/analyze-run”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Get the run ID
  2. Check previous reports for this model
  3. Fetch the failure summary
  4. Fan out sub-agents to analyze each failure
  5. Collate, present, and create report

What it can do on your machine

Read from SKILL.md and the folder at commit aa7b0ab. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq
    • npx
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • fabulous-panther-525.convex.cloud
    • convex-evals.netlify.app

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Run loads about 2.1k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 764 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from get-convex/convex-evals at commit aa7b0ab, republished under its Apache-2.0 licence (© get-convex). 764 words, ~2,150 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-run/SKILL.md (or your agent's skills folder).
name
analyze-run
description
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Use when the user asks to analyze an entire run, review all failures in a run, or wants to understand why a model scored poorly.

Analyze Run

When to use

  • User asks "analyze this run" or "why did this model score poorly?"
  • User shares a run URL like https://convex-evals.netlify.app/experiment/.../run/$runId/...
  • User wants to review all failures across an entire eval run

Step 1: Get the run ID

Extract the run ID from the visualizer URL. The URL pattern is:

/experiment/$experimentId/run/$runId/...

The $runId is the Convex document ID (e.g. jn7922j1w29pdxm76bj9ps0enx80mg9e).

This skill covers coding runs only. For a decision run (visualizer URLs under /decision/run/), the queries below throw "This operation requires a coding run", which the public API reports as a bare "Server Error".

Step 2: Check previous reports for this model

Reports are stored in reports/{provider}/{model}/, where {provider}/{model} is the run's model slug (the model field from Step 3), e.g. reports/anthropic/claude-opus-4.8/ for anthropic/claude-opus-4.8. Don't use the run's provider field. It stores the OpenRouter endpoint provider (or openrouter when discovery failed), which can differ from the slug's vendor. List the directory for the model being analyzed and read the most recent report(s). This gives you:

  • Known recurring failures for this model
  • Actions already taken (lint config changes, grader fixes, task updates)
  • Classifications from prior analysis that may still apply

Reference prior findings when the same eval fails again — note whether it's a repeat and whether any prior fix should have resolved it.

Step 3: Fetch the failure summary

Use the public production query runs:getRunDetails over HTTP. It needs no login:

bash
URL=https://fabulous-panther-525.convex.cloud
curl -s $URL/api/query -H 'Content-Type: application/json' \
  -d '{"path":"runs:getRunDetails","args":{"runId":"<runId>"}}' > /tmp/run-<runId>.json
jq '.value | {model, provider, experiment, status: .status.kind,
  totalEvals: (.evals | length),
  passedCount: ([.evals[] | select(.status.kind == "passed")] | length),
  failedEvals: [.evals[] | select(.status.kind == "failed") | {_id, evalPath,
    failureReason: .status.failureReason,
    failedStep: ([.steps[] | select(.status.kind == "failed") | {name, failureReason: .status.failureReason}] | first)}]}' /tmp/run-<runId>.json

This returns:

  • model (the slug), provider, experiment (null means default), status -- run metadata
  • totalEvals, passedCount -- overall stats
  • failedEvals -- array of failed evals, each with _id, evalPath, failureReason, and failedStep (which step failed and its error)

If there are no failures, report that all evals passed and stop.

Don't use npx convex run --prod for this. The debug functions (debugQueries:getFailedEvalsForRun, debug:getEvalDebugInfo) are internal, and agents usually hit team SSO ("Single-sign on login is required").

Step 4: Fan out sub-agents to analyze each failure

For each failed eval, spawn a sub-agent (up to 4 in parallel) with this prompt template. Fill in <REPO_ROOT> with the absolute path from git rev-parse --show-toplevel. If a sub-agent can't fetch its eval, give Mike this command to run and paste back: cd evalScores && npx convex run --prod debug:getEvalDebugInfo '{"evalId": "<EVAL_ID>"}'.

You are investigating a failing eval from the convex-evals system.

The repo root is <REPO_ROOT>. Work in a new temp directory, not the repo.
Fetch the eval from the public production API (no login needed):

URL=https://fabulous-panther-525.convex.cloud
curl -s $URL/api/query -H 'Content-Type: application/json' \
  -d '{"path":"runs:getRunDetails","args":{"runId":"<RUN_ID>"}}' \
  | jq '.value.evals[] | select(._id == "<EVAL_ID>")' > eval.json

eval.json has the task text (task), status (failureReason, outputStorageId),
evalSourceStorageId, and steps. Get a download URL for each storage ID:

curl -s $URL/api/query -H 'Content-Type: application/json' \
  -d '{"path":"runs:getOutputUrl","args":{"storageId":"<STORAGE_ID>"}}' | jq -r .value

Download status.outputStorageId to output.zip and evalSourceStorageId to
source.zip with curl -s -o, then unzip each into output/ and source/.
Don't use npx convex run --prod. If a request fails, stop and report the error.

Then analyze the result:
1. Which step failed and what was the exact error?
2. Look at the model's generated code in output/.
3. Look at the expected answer and grader in source/.
4. Look at the task description in eval.json's task field.
5. Is this a genuine model mistake, or is the test/lint/task unfair?

Classify the failure as one of:
- MODEL_FAULT: The model genuinely got it wrong
- OVERLY_STRICT: The eval/lint/test requirements are unreasonable for what was asked
- AMBIGUOUS_TASK: The task description is unclear and the model's interpretation was reasonable
- KNOWN_GAP: A known limitation of this eval that affects all models (e.g. the Convex API returns fields the model can't predict without being told)

Return a structured summary:
- Eval: <name> (<category>)
- Failed step: <step name>
- Error: <one-line error summary>
- Classification: <one of the above>
- Reasoning: <2-3 sentences explaining your classification>
- Model output snippet: <the relevant problematic code, if applicable>
- Expected code snippet: <what the answer looks like, if applicable>

Step 5: Collate, present, and create report

Once all sub-agents return, build the analysis:

5a. Overall summary
  • Model, experiment, pass rate (X/Y evals passed)
  • Breakdown by failure type: how many eslint, tsc, deploy, test failures
5b. Failure classification table

For each failure, list: eval name, failed step, classification, one-line reasoning.

5c. Cross-cutting patterns

Look for patterns across failures:

  • Are multiple failures caused by the same root issue? (e.g. same lint rule, same API misunderstanding, same missing pattern)
  • Are there categories of evals that are systematically harder?
  • Do prior reports for this model already document these issues?
Show full SKILL.md (302 more words)Show less
5d. Recommendations

Group recommendations by type:

  • Eval improvements: Tasks that should be clarified, tests that should be relaxed
  • Lint/config changes: Rules that are too strict for what we're testing
  • Model-specific notes: Patterns this model struggles with that other models might not
  • No action needed: Failures that are genuinely the model's fault
5e. Create report file

Always create a report file at:

reports/{provider}/{model}/{runIdPrefix}_{date}.md

For example: reports/anthropic/claude-opus-4.8/jn72t14a_2026-09-29.md for a run of anthropic/claude-opus-4.8.

{provider}/{model} is the run's model slug, as in Step 2. The runIdPrefix is the first 8 characters of the run ID.

The report should contain:

  • Run metadata (ID, model, experiment, date, pass rate)
  • Failure summary table (by step type)
  • Per-failure analysis with classification, reasoning, and code snippets
  • Cross-cutting patterns (especially recurring failures from prior reports)
  • Recommendations (eval improvements, lint/config changes, model-specific notes)
  • Net impact assessment (how many failures are actionable vs genuine model faults)
  • Actions taken: List any changes made as a result of this analysis (e.g. "Updated TASK.txt for 007-http_action_routing to clarify getSiteURL placement"). Default to "None" if no changes were made — this makes it explicit that recommendations were reviewed and deliberately not acted on, rather than simply forgotten.
5f. Present to user

Present the full analysis to the user. End with:

"These are my findings. Would you like me to implement any of these recommendations, or would you like to discuss specific failures in more detail?"

Do NOT make any code/config changes until the user explicitly asks.

5g. Update report after implementing changes

If the user asks you to implement any recommendations, update the report file's "Actions taken" section after making the changes. Record:

  • What was changed (file path + brief description)
  • Which failure(s) it addresses
  • Date of the change

This ensures future analysis sessions can see which recommendations were already acted on and avoid re-recommending changes that have already been made.

© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/analyze-run of get-convex/convex-evals.

Open the folder on GitHubat commit aa7b0ab

Compare with similar skills

Analyze Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Run this skillget-convex/convex-evals130—~2.1kAutomated safety check: PassApache-2.0
Eval HarnessArchive228/loopkit755—~876Automated safety check: PassMIT
Wjs Evaling Voicedrop Promptsjianshuo/claude-skills131—~475Automated safety check: PassMIT
Woo AI Smokewoocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0
Agent BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT

Similar skills

  • Eval Harness

    Archive228/loopkit

    Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

    755 GitHub stars~876 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Wjs Evaling Voicedrop Prompts

    jianshuo/claude-skills

    A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…

    131 GitHub stars~475 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub stars~7.4k tokensUpdated today
    EducationAuto-check: notes
  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Prompt Template Authoring

    nicobailon/pi-prompt-template-model

    Write and run custom Pi prompt templates (slash commands) for this extension.

    321 GitHub stars~1.3k tokensUpdated 17 days ago
    AI & LLM EngineeringAuto-check passed

More from get-convex/convex-evals

  • Add Eval

    get-convex/convex-evals

    Design, implement, validate, and calibrate a new eval for the convex-evals suite.

    130 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    130 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Analyze Eval

    get-convex/convex-evals

    Investigate a single failing eval from the convex-evals system.

    130 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Validate Guidelines

    get-convex/convex-evals

    Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

    130 GitHub stars~2.2k tokensUpdated today
    Auto-check: notes

Questions about Analyze Run

What does Analyze Run do?

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Analyze Run is an agent skill from get-convex/convex-evals. Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

When should I use Analyze Run?

Analyze Run fits situations like: the user asks to analyze an entire run; review all failures in a run; wants to understand why a model scored poorly.

How do I install Analyze Run in Claude Code?

Run `npx skills add get-convex/convex-evals --skill analyze-run -a claude-code`. Or copy the skill folder (.cursor/skills/analyze-run in get-convex/convex-evals) into .claude/skills/analyze-run in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Run in Codex?

Run `npx skills add get-convex/convex-evals --skill analyze-run -a codex`. Or copy the skill folder (.cursor/skills/analyze-run in get-convex/convex-evals) into .agents/skills/analyze-run in your project. Codex loads it when a task matches its description.

Can I use Analyze Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill analyze-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-run, .gemini/skills/analyze-run, .github/skills/analyze-run and .opencode/skills/analyze-run in your project.

What does Analyze Run need to run?

Going by SKILL.md and its folder, Analyze Run needs the command-line tools its instructions call (curl, jq, npx and git). Our summary lists: Node.js.

Does Analyze Run access the network?

SKILL.md names 2 domains. In commands or code: fabulous-panther-525.convex.cloud and convex-evals.netlify.app; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Analyze Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Run use?

Analyze Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Run use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Run?

Skills that share tags, products or a category with Analyze Run: Eval Harness (Archive228/loopkit, 755 stars), Wjs Evaling Voicedrop Prompts (jianshuo/claude-skills, 131 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Agent Builder (shareAI-lab/learn-claude-code, 78k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Run?

get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 130 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 9, 2026.

Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.