Agent skill

Add Model

by get-convex in get-convex/convex-evals

Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Add Model

skills CLI
$ npx skills add get-convex/convex-evals --skill add-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install get-convex/convex-evals add-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-model .claude/skills/add-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-model
GitHub stars
129
Token cost
~1.5k tokens
SKILL.md length
677 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

  • Works in 8 steps: Find the OpenRouter slug → Add it to the lists → Smoke test locally → …
  • The user wants to add
  • SKILL.md covers 1. Find the OpenRouter slug, 2. Add it to the lists, 3. Smoke test locally and 4. Typecheck and test, plus 4 more sections
  • Calls gh, curl and jq; reaches openrouter.ai and fabulous-panther-525.convex.cloud; needs OPENROUTER_API_KEY and EXA_API_KEY

What it does

Add Model is an agent skill from get-convex/convex-evals. Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. Use when the user wants to add or onboard a model, or names a new model or OpenRouter link for the leaderboard.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Model routing and gateways. It works with OpenRouter. The licence is Apache-2.0.

When your agent uses it

  • The user wants to add
  • Onboard a model
  • Names a new model
  • OpenRouter link for the leaderboard

Example prompts

  • “/add-model”

Requirements

  • A credential in OPENROUTER_API_KEY
  • A credential in EXA_API_KEY

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Find the OpenRouter slug
  2. Add it to the lists
  3. Smoke test locally
  4. Typecheck and test
  5. Open a PR
  6. Dispatch baseline coding runs
  7. Dispatch decision runs after the merge
  8. Check the periodic schedule

What it can do on your machine

Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • curl
    • jq
    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai
    • fabulous-panther-525.convex.cloud

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • EXA_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Model loads about 1.5k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 677 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:41
    eads `OPENROUTER_API_KEY` from the root `.env`. In a new worktree, run `bun run setup` first. It installs all dependenci

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 677 words, ~1,464 tokens.

Download SKILL.mdSave it as .claude/skills/add-model/SKILL.md (or your agent's skills folder).
name
add-model
description
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. Use when the user wants to add or onboard a model, or names a new model or OpenRouter link for the leaderboard.

Add a model

Every model runs through OpenRouter. The runner takes any OpenRouter slug and looks up the display name and API kind at run time (runner/models/openRouterDiscovery.ts). Adding a model is a list change in a PR plus paid baseline runs.

1. Find the OpenRouter slug

bash
curl -s https://openrouter.ai/api/v1/models | jq -r '.data[] | "\(.id)\t\(.name)"' | grep -i sonnet

Use the plain slug, e.g. anthropic/claude-sonnet-5.5, not a :batch variant or a ~...-latest alias. If the model isn't listed, the runner can't run it yet.

2. Add it to the lists

ALL_MODELS is the curated list. It feeds the periodic schedule, the curated cohort of the next benchmark mint, and the ablation and guideline-validation scripts.

  • runner/models/index.ts: add the slug to ALL_MODELS next to its family. Entries are plain strings. Keep older siblings unless the maintainer says to drop them.
  • runner/models.test.ts: add expect(ALL_MODELS).toContain("<slug>"); to "contains the current curated models".

The decision benchmark is optional, so ask if unsure. Pick a short key without dots or dashes (fable51, gemini38flash) and add it in three places, next to its family each time:

  • DECISION_CI_MODELS in runner/decisions/ci.ts, e.g. fable51: { provider: "openrouter", model: "anthropic/claude-fable-5.1" },
  • the model input's options in .github/workflows/decision_evals.yml
  • the all JSON array on the matrix.model line of the same file

runner/decisions/ci.test.ts fails if the three drift apart. PR #338 is a complete example.

manual_evals.yml needs no edit. It takes models as a dispatch input.

3. Smoke test locally

bash
DISABLE_CONVEX_REPORTING=1 MODELS=<slug> TEST_FILTER="000-fundamentals/000-empty_functions|000-fundamentals/003-crons" bun run local:run

This reads OPENROUTER_API_KEY from the root .env. In a new worktree, run bun run setup first. It installs all dependencies and copies .env from another worktree.

Expect [preflight] Endpoint is available and a score for both evals. A low score is fine. not supported and not found on OpenRouter means a wrong slug. A model-specific request error, like Kimi K3 rejecting temperature (PR #225), needs a runner change, so stop and raise it.

4. Typecheck and test

bash
bun run typecheck
bun run test

5. Open a PR

Open it against main with Why, What, Validation (smoke result, typecheck, test counts) and, if you added a decision key, After merge. PR #340 is the template. Never merge it or enable auto-merge. The maintainer merges.

Show full SKILL.md (340 more words)Show less

6. Dispatch baseline coding runs

These don't wait for the merge. The runner discovers the slug, and dispatches on main report to the production leaderboard. Other refs run with reporting off.

State a dollar estimate and get approval first. Take per-run averages for a model at a similar price from the public production query modelScores:getSchedulingStats:

bash
URL=https://fabulous-panther-525.convex.cloud
ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
  -d '{"path":"models:getBySlug","args":{"slug":"anthropic/claude-sonnet-5"}}' | jq -r .value._id)
for exp in '' ',"experiment":"no_guidelines"' ',"experiment":"no_guidelines_with_web"'; do
  curl -s $URL/api/query -H 'Content-Type: application/json' \
    -d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"$exp}}" | jq .value.averageRunCostUsd
done

That prints the default, no_guidelines and no_guidelines_with_web averages. Web costs include Exa charges. One dispatch runs all three conditions, so three dispatches cost three times the sum. Sonnet 5 on 2026-09-29: $4.17 + $1.82 + $2.55 = $8.54 per dispatch, about $26 total.

Then dispatch three times:

bash
gh workflow run manual_evals.yml --ref main -f models=<slug> -f run_guidelines=true -f run_no_guidelines=true -f run_no_guidelines_with_web=true
  • Web runs fail unless the repo variable ENABLE_CLIENT_WEB_PRODUCTION is true (gh variable list). The workflow supplies EXA_API_KEY and CLIENT_WEB_TOOLS=1.
  • If the provider rate-limits, pass -f max_concurrency=2 or 1.
  • The conditions are sequential steps in one job, and a failed step skips the rest. Re-dispatch just the missing ones with the others set to false.
  • In a web run, a provider or Exa error that survives its retries aborts the whole run, which is marked failed rather than scored. Read the log, re-dispatch, and don't count it.

Find the runs with gh run list --workflow=manual_evals.yml --limit 3, follow one with gh run watch <id>, and read failures with gh run view <id> --log-failed.

7. Dispatch decision runs after the merge

Skip this if you didn't add a decision key. The workflow only runs on main, and ci.ts rejects keys it doesn't know, so nothing works before the merge.

Each job stops at $5 of known cost, which isn't a hard cap when a provider omits cost. Quote up to $15 and get approval, then dispatch three times, like every current decision model:

bash
gh workflow run decision_evals.yml --ref main -f model=<key>

Always pass model. Its default, all, runs every decision model. condition defaults to no_guidelines, the only condition used so far.

8. Check the periodic schedule

bash
gh workflow list --all

If Periodic Evaluations is disabled_manually, the model only gets the manual runs above until the maintainer re-enables it. Don't re-enable it yourself. New models never need a benchmark mint. They score under the current version.

© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/add-model of get-convex/convex-evals.

Open the folder on GitHubat commit 68f5c0e

Compare with similar skills

Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Model this skillget-convex/convex-evals129—~1.5kAutomated safety check: NotesApache-2.0
Evals Run On Demandelastic/kibana21k—~2.6kAutomated safety check: NotesCustom licence
Legal AI Model Router Stephane Boghossianlawve-ai/awesome-legal-skills836—~1.5kAutomated safety check: PassAGPL-3.0-or-later
Openrouter Text2musicQinghongLin/data2story-skill1561 repos~407Automated safety check: PassMIT
OmniRoute CLI Evalsdiegosouzapw/OmniRoute74k—~1.3kAutomated safety check: PassMIT
To Imgrtadewald/skills180—~673Automated safety check: NotesNone

Similar skills

  • Evals Run On Demand

    elastic/kibana

    Official

    Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English.

    21k GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Legal AI Model Router Stephane Boghossian

    lawve-ai/awesome-legal-skills

    Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.

    836 GitHub stars~1.5k tokensUpdated 5 days ago
    Legal & ComplianceAuto-check passed
  • Openrouter Text2music

    QinghongLin/data2story-skill

    Generate music (NOT speech) via OpenRouter using Google Lyria 3 Pro.

    156 GitHub starsUsed in 1 repo~407 tokens
    AI & LLM EngineeringAuto-check passed
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    74k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • To Img

    rtadewald/skills

    Gera imagens em paralelo via OpenRouter a partir de um pedido simples ou de um batch com prompts e filenames definidos.

    180 GitHub stars~673 tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check: notes
  • Openrouter

    davidondrej/skills

    Design, build, debug, and optimize OpenRouter API integrations.

    4.1k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from get-convex/convex-evals

  • Add Eval

    get-convex/convex-evals

    Design, implement, validate, and calibrate a new eval for the convex-evals suite.

    129 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Analyze Eval

    get-convex/convex-evals

    Investigate a single failing eval from the convex-evals system.

    129 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Analyze Run

    get-convex/convex-evals

    Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

    129 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Validate Guidelines

    get-convex/convex-evals

    Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

    129 GitHub stars~2.2k tokensUpdated today
    Auto-check: notes

Works with

Questions about Add Model

What does Add Model do?

Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. Add Model is an agent skill from get-convex/convex-evals. Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

When should I use Add Model?

Add Model fits situations like: the user wants to add; onboard a model; names a new model; openRouter link for the leaderboard.

How do I install Add Model in Claude Code?

Run `npx skills add get-convex/convex-evals --skill add-model -a claude-code`. Or copy the skill folder (.cursor/skills/add-model in get-convex/convex-evals) into .claude/skills/add-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Model in Codex?

Run `npx skills add get-convex/convex-evals --skill add-model -a codex`. Or copy the skill folder (.cursor/skills/add-model in get-convex/convex-evals) into .agents/skills/add-model in your project. Codex loads it when a task matches its description.

Can I use Add Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-model, .gemini/skills/add-model, .github/skills/add-model and .opencode/skills/add-model in your project.

What does Add Model need to run?

Going by SKILL.md and its folder, Add Model needs the command-line tools its instructions call (gh, curl, jq and bun) and credentials named OPENROUTER_API_KEY and EXA_API_KEY. Our summary lists: A credential in OPENROUTER_API_KEY; A credential in EXA_API_KEY.

Does Add Model access the network?

SKILL.md names 2 domains. In commands or code: openrouter.ai and fabulous-panther-525.convex.cloud; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Add Model safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Add Model use?

Add Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Model use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Model?

Skills that share tags, products or a category with Add Model: Evals Run On Demand (elastic/kibana, 21k stars), Legal AI Model Router Stephane Boghossian (lawve-ai/awesome-legal-skills, 836 stars), Openrouter Text2music (QinghongLin/data2story-skill, 156 stars) and OmniRoute CLI Evals (diegosouzapw/OmniRoute, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Model?

get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.