Evals Run On Demand
elastic/kibana
Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English.
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
$ npx skills add get-convex/convex-evals --skill add-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install get-convex/convex-evals add-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-model .claude/skills/add-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .claude/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add get-convex/convex-evals --skill add-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install get-convex/convex-evals add-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/add-model .agents/skills/add-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .agents/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill add-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install get-convex/convex-evals add-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/add-model .cursor/skills/add-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .cursor/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/get-convex/convex-evals.git --path .cursor/skills/add-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add get-convex/convex-evals --skill add-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install get-convex/convex-evals add-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/add-model .gemini/skills/add-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .gemini/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install get-convex/convex-evals add-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add get-convex/convex-evals --skill add-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/add-model .github/skills/add-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .github/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill add-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install get-convex/convex-evals add-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/add-model .opencode/skills/add-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/add-model into .opencode/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-modelAdd a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
Add Model is an agent skill from get-convex/convex-evals. Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. Use when the user wants to add or onboard a model, or names a new model or OpenRouter link for the leaderboard.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Model routing and gateways. It works with OpenRouter. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghcurljqbunFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
openrouter.aifabulous-panther-525.convex.cloudFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENROUTER_API_KEYEXA_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Model loads about 1.5k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 677 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
eads `OPENROUTER_API_KEY` from the root `.env`. In a new worktree, run `bun run setup` first. It installs all dependenciAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 677 words, ~1,464 tokens.
.claude/skills/add-model/SKILL.md (or your agent's skills folder).Every model runs through OpenRouter. The runner takes any OpenRouter slug and looks up the display name and API kind at run time (runner/models/openRouterDiscovery.ts). Adding a model is a list change in a PR plus paid baseline runs.
curl -s https://openrouter.ai/api/v1/models | jq -r '.data[] | "\(.id)\t\(.name)"' | grep -i sonnetUse the plain slug, e.g. anthropic/claude-sonnet-5.5, not a :batch variant or a ~...-latest alias. If the model isn't listed, the runner can't run it yet.
ALL_MODELS is the curated list. It feeds the periodic schedule, the curated cohort of the next benchmark mint, and the ablation and guideline-validation scripts.
runner/models/index.ts: add the slug to ALL_MODELS next to its family. Entries are plain strings. Keep older siblings unless the maintainer says to drop them.runner/models.test.ts: add expect(ALL_MODELS).toContain("<slug>"); to "contains the current curated models".The decision benchmark is optional, so ask if unsure. Pick a short key without dots or dashes (fable51, gemini38flash) and add it in three places, next to its family each time:
DECISION_CI_MODELS in runner/decisions/ci.ts, e.g. fable51: { provider: "openrouter", model: "anthropic/claude-fable-5.1" },model input's options in .github/workflows/decision_evals.ymlall JSON array on the matrix.model line of the same filerunner/decisions/ci.test.ts fails if the three drift apart. PR #338 is a complete example.
manual_evals.yml needs no edit. It takes models as a dispatch input.
DISABLE_CONVEX_REPORTING=1 MODELS=<slug> TEST_FILTER="000-fundamentals/000-empty_functions|000-fundamentals/003-crons" bun run local:runThis reads OPENROUTER_API_KEY from the root .env. In a new worktree, run bun run setup first. It installs all dependencies and copies .env from another worktree.
Expect [preflight] Endpoint is available and a score for both evals. A low score is fine. not supported and not found on OpenRouter means a wrong slug. A model-specific request error, like Kimi K3 rejecting temperature (PR #225), needs a runner change, so stop and raise it.
bun run typecheck
bun run testOpen it against main with Why, What, Validation (smoke result, typecheck, test counts) and, if you added a decision key, After merge. PR #340 is the template. Never merge it or enable auto-merge. The maintainer merges.
These don't wait for the merge. The runner discovers the slug, and dispatches on main report to the production leaderboard. Other refs run with reporting off.
State a dollar estimate and get approval first. Take per-run averages for a model at a similar price from the public production query modelScores:getSchedulingStats:
URL=https://fabulous-panther-525.convex.cloud
ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
-d '{"path":"models:getBySlug","args":{"slug":"anthropic/claude-sonnet-5"}}' | jq -r .value._id)
for exp in '' ',"experiment":"no_guidelines"' ',"experiment":"no_guidelines_with_web"'; do
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"$exp}}" | jq .value.averageRunCostUsd
doneThat prints the default, no_guidelines and no_guidelines_with_web averages. Web costs include Exa charges. One dispatch runs all three conditions, so three dispatches cost three times the sum. Sonnet 5 on 2026-09-29: $4.17 + $1.82 + $2.55 = $8.54 per dispatch, about $26 total.
Then dispatch three times:
gh workflow run manual_evals.yml --ref main -f models=<slug> -f run_guidelines=true -f run_no_guidelines=true -f run_no_guidelines_with_web=trueENABLE_CLIENT_WEB_PRODUCTION is true (gh variable list). The workflow supplies EXA_API_KEY and CLIENT_WEB_TOOLS=1.-f max_concurrency=2 or 1.false.Find the runs with gh run list --workflow=manual_evals.yml --limit 3, follow one with gh run watch <id>, and read failures with gh run view <id> --log-failed.
Skip this if you didn't add a decision key. The workflow only runs on main, and ci.ts rejects keys it doesn't know, so nothing works before the merge.
Each job stops at $5 of known cost, which isn't a hard cap when a provider omits cost. Quote up to $15 and get approval, then dispatch three times, like every current decision model:
gh workflow run decision_evals.yml --ref main -f model=<key>Always pass model. Its default, all, runs every decision model. condition defaults to no_guidelines, the only condition used so far.
gh workflow list --allIf Periodic Evaluations is disabled_manually, the model only gets the manual runs above until the maintainer re-enables it. Don't re-enable it yourself. New models never need a benchmark mint. They score under the current version.
© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/add-model of get-convex/convex-evals.
Open the folder on GitHubat commit 68f5c0e
Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Model this skillget-convex/convex-evals | 129 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | |
| Evals Run On Demandelastic/kibana | 21k | — | ~2.6k | Automated safety check: Notes | Custom licence | |
| Legal AI Model Router Stephane Boghossianlawve-ai/awesome-legal-skills | 836 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0-or-later | |
| Openrouter Text2musicQinghongLin/data2story-skill | 156 | 1 repos | ~407 | Automated safety check: Pass | MIT | |
| OmniRoute CLI Evalsdiegosouzapw/OmniRoute | 74k | — | ~1.3k | Automated safety check: Pass | MIT | |
| To Imgrtadewald/skills | 180 | — | ~673 | Automated safety check: Notes | None |
elastic/kibana
Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English.
lawve-ai/awesome-legal-skills
Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.
QinghongLin/data2story-skill
Generate music (NOT speech) via OpenRouter using Google Lyria 3 Pro.
diegosouzapw/OmniRoute
Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.
rtadewald/skills
Gera imagens em paralelo via OpenRouter a partir de um pedido simples ou de um batch com prompts e filenames definidos.
davidondrej/skills
Design, build, debug, and optimize OpenRouter API integrations.
get-convex/convex-evals
Design, implement, validate, and calibrate a new eval for the convex-evals suite.
get-convex/convex-evals
Investigate a single failing eval from the convex-evals system.
get-convex/convex-evals
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
get-convex/convex-evals
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
Works with
Categories
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. Add Model is an agent skill from get-convex/convex-evals. Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
Add Model fits situations like: the user wants to add; onboard a model; names a new model; openRouter link for the leaderboard.
Run `npx skills add get-convex/convex-evals --skill add-model -a claude-code`. Or copy the skill folder (.cursor/skills/add-model in get-convex/convex-evals) into .claude/skills/add-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add get-convex/convex-evals --skill add-model -a codex`. Or copy the skill folder (.cursor/skills/add-model in get-convex/convex-evals) into .agents/skills/add-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-model, .gemini/skills/add-model, .github/skills/add-model and .opencode/skills/add-model in your project.
Going by SKILL.md and its folder, Add Model needs the command-line tools its instructions call (gh, curl, jq and bun) and credentials named OPENROUTER_API_KEY and EXA_API_KEY. Our summary lists: A credential in OPENROUTER_API_KEY; A credential in EXA_API_KEY.
SKILL.md names 2 domains. In commands or code: openrouter.ai and fabulous-panther-525.convex.cloud; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Add Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add Model: Evals Run On Demand (elastic/kibana, 21k stars), Legal AI Model Router Stephane Boghossian (lawve-ai/awesome-legal-skills, 836 stars), Openrouter Text2music (QinghongLin/data2story-skill, 156 stars) and OmniRoute CLI Evals (diegosouzapw/OmniRoute, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.
Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.