LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Design, implement, validate, and calibrate a new eval for the convex-evals suite.
$ npx skills add get-convex/convex-evals --skill add-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install get-convex/convex-evals add-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-eval .claude/skills/add-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .claude/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add get-convex/convex-evals --skill add-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install get-convex/convex-evals add-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/add-eval .agents/skills/add-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .agents/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill add-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install get-convex/convex-evals add-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/add-eval .cursor/skills/add-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .cursor/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/get-convex/convex-evals.git --path .cursor/skills/add-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add get-convex/convex-evals --skill add-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install get-convex/convex-evals add-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/add-eval .gemini/skills/add-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .gemini/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install get-convex/convex-evals add-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add get-convex/convex-evals --skill add-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/add-eval .github/skills/add-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .github/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill add-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install get-convex/convex-evals add-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/add-eval .opencode/skills/add-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-eval" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-eval into .opencode/skills/add-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-evalDesign, implement, validate, and calibrate a new eval for the convex-evals suite.
Add Eval is an agent skill from get-convex/convex-evals. Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
buncurljqbunxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
docs.convex.devfabulous-panther-525.convex.cloudFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Eval loads about 4.6k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 2,280 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 2,280 words, ~4,555 tokens.
.claude/skills/add-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Follow these steps whenever the user asks to create a new eval. Read .cursor/skills/add-eval/reference.md for grader helpers, test patterns, and conventions. The eval-writing rules in AGENTS.md ("Authoring New Evals") and README.md ("Writing evals") take precedence over this skill.
Steps 0-2 (gather info, research, design) are collaborative and read-only. Don't create or edit any files until the user has seen the research findings and approved the eval design. Implementation starts at Step 3.
Determine the following (ask the user if not provided):
List the existing categories by scanning evals/ top-level directories, then propose the best-fit category. The current categories are:
| Category | Scope |
|---|---|
000-fundamentals | Basic Convex concepts (empty functions, schema definition, crons, scheduling) |
001-data_modeling | Schema design, indexes, relationships, unions, optional fields |
002-queries | Reading data, joins, pagination, aggregation, filtering |
003-mutations | Writing data, inserts, patches, deletes, cascades |
004-actions | HTTP fetch, file storage, node runtime, HTTP action routing |
005-idioms | File organization, internal functions, batch patterns |
006-clients | useQuery, useMutation, usePaginatedQuery |
007-components | Components: usage evals wiring a named component, selection evals choosing one, local components, function handles |
Determine the eval number by listing existing evals in the chosen category and picking the next sequential number.
Run these four research tracks. Use sub-agents or parallel tool calls where possible.
https://docs.convex.dev/llms.txt to get the docs table of contents.runner/models/guidelines.md for the concept being tested. It is the source. runner/models/guidelines.ts only loads it.reference.md for the full catalog.evals/ to get a high-level view of coverage.Present the full eval design to the user for review. Stay read-only until they approve it.
State what Convex-specific knowledge the eval measures and what silently breaks in production when a model lacks it. Every eval issue and PR needs this section (AGENTS.md). If you can't write it, the eval probably tests trivia.
Write the complete TASK.txt content. Follow these principles:
internalMutation or how pagination works, that's a meaningful signal, not a task problem.Describe the files that will be created and the key implementation approach. Don't write the full code yet, just the structure and important decisions.
Describe how the eval will be graded:
reference.md.reference.md for the catalog and decision tree).If unit tests cannot fully verify the concept (e.g. testing that a model uses an index rather than a filter, or testing code organization patterns), STOP and discuss with the user. Present the options:
getSchema, hasIndexForFields)createAIGraderTest, currently disabled and requires a repo change to re-enable)Let the user decide before proceeding.
Summarize the guideline context before implementation:
If a new guideline is likely needed, design it minimal-first. Every token in the guidelines is sent with every prompt, so bloat costs real money. Start with the smallest guideline that teaches the critical pattern (usually one code example), test it, and only expand if models still fail. Avoid pinning specific dependency versions in guidelines as they age quickly. Prefer "always install the latest version" instead. Follow the guideline rules in AGENTS.md "Authoring New Evals": prove the line helps a specific eval, and prefer improving an existing line over adding one.
Before presenting the design, critically evaluate it. Warn the user if:
.withIndex() calls). We're testing knowledge, not instruction-following.After the user approves the design, implement:
Create directory: evals/<category>/<eval_slug>/
Write TASK.txt with the approved content. For the static or module pipeline, also add eval.json (see "Eval Pipelines" in reference.md).
Create answer directory:
answer/package.json:{
"name": "convexbot",
"version": "1.0.0",
"dependencies": {
"convex": "^1.31.2"
}
}convex and the component instead (e.g. "convex": "1.41.0", "@convex-dev/aggregate": "0.2.2"). Copy the pins from a sibling in evals/007-components/. Usage-eval tasks state the same pins. Selection-eval tasks can't name the component, so they pin only convex.answer/convex/schema.ts (if applicable)answer/convex/index.ts)returns: validators unless the task tests them (AGENTS.md). Answers are likely training data.Add generated types: bunx convex codegen fails without a deployment ("No CONVEX_DEPLOYMENT set"). Copy _generated and tsconfig.json from a sibling single-module eval with the same files, e.g. convex/index.ts and convex/schema.ts:
cp -R evals/<sibling>/answer/convex/_generated evals/<sibling>/answer/convex/tsconfig.json evals/<category>/<eval_slug>/answer/convex/See "Generated types" in reference.md for when to edit the copied imports.
Write grader.test.ts using the approved test approach. Backend-pipeline graders import from grader/index.ts with relative paths:
import { responseClient, responseAdminClient, addDocuments } from "../../../grader";
import { api } from "./answer/convex/_generated/api";Adjust the depth of ../ based on the eval's nesting level. Static and module graders import from ../../../grader/outputDir instead. grader/index.ts throws "CONVEX_PORT is not set" without a backend.
Typecheck:
bun run typecheckFirst run canonical answer validation for the new eval:
TEST_FILTER=<category>/<eval_slug> bun run validate:answersThe smoke test below and Step 5 make paid model calls. Before running either, state a dollar estimate and wait for the user's approval. Take each model's average full-suite cost from the public production query modelScores:getSchedulingStats, as in the add-model skill:
URL=https://fabulous-panther-525.convex.cloud
for slug in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 openai/gpt-5.5 deepseek/deepseek-v4-pro; do
ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
-d "{\"path\":\"models:getBySlug\",\"args\":{\"slug\":\"$slug\"}}" | jq -r .value._id)
for exp in '' ',"experiment":"no_guidelines"'; do
printf "%s%s " "$slug" "$exp"
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"$exp}}" | jq .value.averageRunCostUsd
done
doneDivide each average by the eval count (ls -d evals/*/*/ | wc -l) for a per-eval cost, sum over the models and conditions you'll run, and multiply by the repeats. On 2026-09-29 these four models averaged about $22.60 per default suite and $14 per no_guidelines suite over 112 evals, so one eval in both conditions cost about $0.33 per repeat, or $1 for three repeats. Evals with long tasks cost more.
Run the eval for one model. This validates model generation against the new eval:
DISABLE_CONVEX_REPORTING=1 MODELS=anthropic/claude-sonnet-5 TEST_FILTER=<category>/<eval_slug> bun run local:runDISABLE_CONVEX_REPORTING=1 keeps results local.
If the smoke test fails:
run.log to understand what happened.Start with a smaller representative set of models to calibrate difficulty. If the result is unclear, expand to a broader sweep. Use full OpenRouter slugs from ALL_MODELS in runner/models/index.ts. Run each model in both conditions, default and EVALS_EXPERIMENT=no_guidelines. The difference is the guidelines' measured contribution (AGENTS.md). Launch one background process per model:
# Suggested first-pass set
for m in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 openai/gpt-5.5 deepseek/deepseek-v4-pro; do
(
DISABLE_CONVEX_REPORTING=1 MODELS=$m TEST_FILTER=<category>/<eval_slug> bun run local:run
DISABLE_CONVEX_REPORTING=1 EVALS_EXPERIMENT=no_guidelines MODELS=$m TEST_FILTER=<category>/<eval_slug> bun run local:run
) > "/tmp/calibrate-${m//\//_}.log" 2>&1 &
done
waitIf those results are too noisy or too uniform, expand to a broader sweep across providers and tiers from ALL_MODELS. The user can override the list. Any expansion needs a new estimate and approval.
Monitor progress by reading the log files. Each process runs one eval in two conditions, so they should complete in a few minutes.
Run each model at least twice (ideally three times) to distinguish systematic failures from flaky ones. Model output is non-deterministic, so a single pass or fail doesn't tell the whole story. Common variance sources:
Collect pass/fail from all model runs in both conditions and present a summary table:
Model default no_guidelines
--------------------------- ------- -------------
anthropic/claude-sonnet-5 3/3 1/3
anthropic/claude-opus-4.8 3/3 2/3
openai/gpt-5.5 2/3 0/3
...Then assess the results:
Then explicitly ask: is this primarily an eval/task gap, a model gap, or a guideline gap?
no_guidelines match, the current guidelines aren't moving this eval. After the eval is settled, recommend a change to runner/models/guidelines.md that follows AGENTS.md, then a regression check with the validate-guidelines skill filtered to this eval's category. That run needs its own cost estimate and approval.Model output is preserved in the temp directory printed at the start of each run (Using tempdir: ...). These directories are NOT cleaned up automatically. For each model, look at:
<tempdir>/output/<provider>/<model>/<category>/<eval>/ for the generated source filesrun.log in that directory for step-by-step output (install, deploy, tsc, eslint, vitest)node_modules/ for the actual resolved dependency versionspackage.json for what the model requested vs what was installedThis is essential for distinguishing "model wrote bad code" from "model's code is fine but the grader is too strict" from "dependency version bug".
Push back with specific recommendations if calibration looks off. Suggest concrete changes to the task, answer, or tests.
_generated copied from a sibling evalbun run typecheck passesbun run validate:answers passes for the new eval© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .cursor/skills/add-eval of get-convex/convex-evals.
Open the folder on GitHubat commit 68f5c0e
Add Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Eval this skillget-convex/convex-evals | 129 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
get-convex/convex-evals
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
get-convex/convex-evals
Investigate a single failing eval from the convex-evals system.
get-convex/convex-evals
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
get-convex/convex-evals
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
Categories
Design, implement, validate, and calibrate a new eval for the convex-evals suite. Add Eval is an agent skill from get-convex/convex-evals. Design, implement, validate, and calibrate a new eval for the convex-evals suite.
Add Eval fits situations like: the user wants to add a new eval; test a new Convex concept; expand eval coverage.
Run `npx skills add get-convex/convex-evals --skill add-eval -a claude-code`. Or copy the skill folder (.cursor/skills/add-eval in get-convex/convex-evals) into .claude/skills/add-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add get-convex/convex-evals --skill add-eval -a codex`. Or copy the skill folder (.cursor/skills/add-eval in get-convex/convex-evals) into .agents/skills/add-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill add-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-eval, .gemini/skills/add-eval, .github/skills/add-eval and .opencode/skills/add-eval in your project.
Going by SKILL.md and its folder, Add Eval needs the command-line tools its instructions call (bun, curl, jq and bunx).
SKILL.md names 2 domains. In commands or code: docs.convex.dev and fabulous-panther-525.convex.cloud; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add Eval: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.
Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.