Agent skill

Add Eval

by get-convex in get-convex/convex-evals

Design, implement, validate, and calibrate a new eval for the convex-evals suite.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Eval

skills CLI
$ npx skills add get-convex/convex-evals --skill add-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install get-convex/convex-evals add-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-eval .claude/skills/add-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-eval
GitHub stars
129
Token cost
~4.6k tokens
SKILL.md length
2,280 words
Files
2
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design, implement, validate, and calibrate a new eval for the convex-evals suite.

  • Works in 7 steps: Gather Information → Research → Design the Eval → …
  • The user wants to add a new eval
  • SKILL.md covers Step 0: Gather Information, Step 1: Research, Step 2: Design the Eval and Step 3: Implement the Eval, plus 4 more sections
  • Calls bun, curl and jq; reaches docs.convex.dev and fabulous-panther-525.convex.cloud

What it does

Add Eval is an agent skill from get-convex/convex-evals. Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. The licence is Apache-2.0.

When your agent uses it

  • The user wants to add a new eval
  • Test a new Convex concept
  • Expand eval coverage

Example prompts

  • “/add-eval”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Gather Information
  2. Research
  3. Design the Eval
  4. Implement the Eval
  5. Validate the Answer
  6. Run Against Multiple Models
  7. Review Results and Calibrate

What it can do on your machine

Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • curl
    • jq
    • bunx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.convex.dev
    • fabulous-panther-525.convex.cloud

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Eval loads about 4.6k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 2,280 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 2,280 words, ~4,555 tokens.

Download SKILL.mdSave it as .claude/skills/add-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
add-eval
description
Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.

Add a New Eval

Follow these steps whenever the user asks to create a new eval. Read .cursor/skills/add-eval/reference.md for grader helpers, test patterns, and conventions. The eval-writing rules in AGENTS.md ("Authoring New Evals") and README.md ("Writing evals") take precedence over this skill.

Steps 0-2 (gather info, research, design) are collaborative and read-only. Don't create or edit any files until the user has seen the research findings and approved the eval design. Implementation starts at Step 3.

Step 0: Gather Information

Determine the following (ask the user if not provided):

  1. Concept to test - what Convex feature or pattern should this eval exercise? (e.g. "vector search", "pagination with joins", "cascade deletes")
  2. Specific focus - any edge cases, constraints, or behaviors the user wants to emphasize?
Category Selection

List the existing categories by scanning evals/ top-level directories, then propose the best-fit category. The current categories are:

CategoryScope
000-fundamentalsBasic Convex concepts (empty functions, schema definition, crons, scheduling)
001-data_modelingSchema design, indexes, relationships, unions, optional fields
002-queriesReading data, joins, pagination, aggregation, filtering
003-mutationsWriting data, inserts, patches, deletes, cascades
004-actionsHTTP fetch, file storage, node runtime, HTTP action routing
005-idiomsFile organization, internal functions, batch patterns
006-clientsuseQuery, useMutation, usePaginatedQuery
007-componentsComponents: usage evals wiring a named component, selection evals choosing one, local components, function handles
  • If the concept clearly fits one category, propose it with a brief justification.
  • If it's ambiguous (e.g. "scheduled mutations" could be fundamentals or mutations), stop and ask the user. Present the candidate categories with reasoning for each.
  • If it doesn't fit any existing category, propose creating a new category and ask the user to confirm.

Determine the eval number by listing existing evals in the chosen category and picking the next sequential number.

Step 1: Research

Run these four research tracks. Use sub-agents or parallel tool calls where possible.

A. Convex Docs (source of truth)
  1. Fetch https://docs.convex.dev/llms.txt to get the docs table of contents.
  2. Identify the 1-3 most relevant doc pages for the concept being tested.
  3. Use WebFetch to retrieve those specific pages.
  4. Extract the correct API patterns, constraints, and best practices. These are ground truth for designing the eval and answer.
B. Existing Guidelines
  1. Read the relevant sections of runner/models/guidelines.md for the concept being tested. It is the source. runner/models/guidelines.ts only loads it.
  2. Record:
    • Which existing guidelines are relevant
    • What behavior those guidelines would lead a model to produce
    • Whether the current guidelines would already be expected to make strong models pass
    • Whether any guideline seems to conflict with what the proposed eval is trying to reward
  3. Do not treat this as a reason to block the eval automatically. This is context for design and later calibration.
  4. If an existing guideline appears to directly contradict the proposed eval, STOP and discuss with the user before proceeding.
C. Existing Eval Patterns
  1. Read 2-3 evals in the same or a similar category.
  2. Study: TASK.txt style, answer structure, test approach, schema design.
  3. Note which grader helpers and test patterns they use. See reference.md for the full catalog.
D. Overlap Check
  1. List ALL eval directory names under evals/ to get a high-level view of coverage.
  2. For any eval whose name suggests overlap with the proposed concept, read its TASK.txt.
  3. If an existing eval tests the same or very similar concept, STOP and warn the user:
    • Name the overlapping eval(s) and explain what they already cover.
    • Ask whether to: differentiate the new eval (narrow its scope), adjust the existing eval, or abandon.
  4. Flag even partial overlap, e.g. "002-queries/012-index_and_filter already tests index usage but doesn't cover compound indexes."

Step 2: Design the Eval

Present the full eval design to the user for review. Stay read-only until they approve it.

Why this matters

State what Convex-specific knowledge the eval measures and what silently breaks in production when a model lacks it. Every eval issue and PR needs this section (AGENTS.md). If you can't write it, the eval probably tests trivia.

TASK.txt Draft

Write the complete TASK.txt content. Follow these principles:

  • Laser-focused on the concept being tested. Control other variables (keep schema simple, minimize unrelated code).
  • Explicit about schema, function names, argument types, return shapes, and which files to create.
  • Don't over-specify Convex implementation details that are covered by the guidelines. If a model needs the task to spell out how to use internalMutation or how pagination works, that's a meaningful signal, not a task problem.
  • Don't under-specify the problem domain. The model should not need to guess what the feature does, only how to implement it in Convex.
  • Include schema as a TypeScript code block when applicable.
  • Specify edge cases (empty results, error messages, missing data).
Answer Outline

Describe the files that will be created and the key implementation approach. Don't write the full code yet, just the structure and important decisions.

Test Approach

Describe how the eval will be graded:

  • Pick the primary grading primitive first: behavior tests, schema inspection, function-spec comparison, HTTP testing, AI grading, AST analysis, or some combination.
  • Pick the pipeline: backend (default), static for selection evals, or module. See "Eval Pipelines" in reference.md.
  • Which grader helpers to use (see reference.md for the catalog and decision tree).
  • What behaviors to assert on.
  • Whether standard unit tests are sufficient, or if you need schema inspection, HTTP testing, AI grading, or something else.

If unit tests cannot fully verify the concept (e.g. testing that a model uses an index rather than a filter, or testing code organization patterns), STOP and discuss with the user. Present the options:

  • Schema/index inspection (using getSchema, hasIndexForFields)
  • AI grading (createAIGraderTest, currently disabled and requires a repo change to re-enable)
  • AST analysis (parse the generated TypeScript files)
  • Restructure the eval so the concept can be tested via behavior
  • Accept the limitation and test what we can

Let the user decide before proceeding.

Guidelines Hypothesis

Summarize the guideline context before implementation:

  • Which existing guidelines are relevant to this eval
  • What you would expect guideline-following models to do
  • Whether failures on this eval would likely indicate a model gap, an eval/task problem, or a missing/weak guideline
  • Any existing guideline that might need to be revised if calibration shows an unexpected result

If a new guideline is likely needed, design it minimal-first. Every token in the guidelines is sent with every prompt, so bloat costs real money. Start with the smallest guideline that teaches the critical pattern (usually one code example), test it, and only expand if models still fail. Avoid pinning specific dependency versions in guidelines as they age quickly. Prefer "always install the latest version" instead. Follow the guideline rules in AGENTS.md "Authoring New Evals": prove the line helps a specific eval, and prefer improving an existing line over adding one.

Push Back

Before presenting the design, critically evaluate it. Warn the user if:

  • The eval overlaps heavily with an existing eval (should have been caught in Step 1C, but re-check).
  • The task tests too many concepts at once. Each eval should focus on one thing.
  • The task is so explicit that it tells the model exactly how to solve it (e.g. specifying .withIndex() calls). We're testing knowledge, not instruction-following.
  • The current guidelines suggest a different approach than the eval is rewarding, or make the expected signal unclear.
  • The eval seems too easy (every model will pass) or too hard (no model will pass).

Step 3: Implement the Eval

After the user approves the design, implement:

  1. Create directory: evals/<category>/<eval_slug>/

  2. Write TASK.txt with the approved content. For the static or module pipeline, also add eval.json (see "Eval Pipelines" in reference.md).

  3. Create answer directory:

    • answer/package.json:
      json
      {
        "name": "convexbot",
        "version": "1.0.0",
        "dependencies": {
          "convex": "^1.31.2"
        }
      }
      Component evals pin exact versions of convex and the component instead (e.g. "convex": "1.41.0", "@convex-dev/aggregate": "0.2.2"). Copy the pins from a sibling in evals/007-components/. Usage-eval tasks state the same pins. Selection-eval tasks can't name the component, so they pin only convex.
    • answer/convex/schema.ts (if applicable)
    • Implementation files (e.g. answer/convex/index.ts)
    • No returns: validators unless the task tests them (AGENTS.md). Answers are likely training data.
  4. Add generated types: bunx convex codegen fails without a deployment ("No CONVEX_DEPLOYMENT set"). Copy _generated and tsconfig.json from a sibling single-module eval with the same files, e.g. convex/index.ts and convex/schema.ts:

    bash
    cp -R evals/<sibling>/answer/convex/_generated evals/<sibling>/answer/convex/tsconfig.json evals/<category>/<eval_slug>/answer/convex/

    See "Generated types" in reference.md for when to edit the copied imports.

  5. Write grader.test.ts using the approved test approach. Backend-pipeline graders import from grader/index.ts with relative paths:

    typescript
    import { responseClient, responseAdminClient, addDocuments } from "../../../grader";
    import { api } from "./answer/convex/_generated/api";

    Adjust the depth of ../ based on the eval's nesting level. Static and module graders import from ../../../grader/outputDir instead. grader/index.ts throws "CONVEX_PORT is not set" without a backend.

  6. Typecheck:

    bash
    bun run typecheck
Show full SKILL.md (857 more words)Show less

Step 4: Validate the Answer

First run canonical answer validation for the new eval:

bash
TEST_FILTER=<category>/<eval_slug> bun run validate:answers
Cost approval

The smoke test below and Step 5 make paid model calls. Before running either, state a dollar estimate and wait for the user's approval. Take each model's average full-suite cost from the public production query modelScores:getSchedulingStats, as in the add-model skill:

bash
URL=https://fabulous-panther-525.convex.cloud
for slug in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 openai/gpt-5.5 deepseek/deepseek-v4-pro; do
  ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
    -d "{\"path\":\"models:getBySlug\",\"args\":{\"slug\":\"$slug\"}}" | jq -r .value._id)
  for exp in '' ',"experiment":"no_guidelines"'; do
    printf "%s%s " "$slug" "$exp"
    curl -s $URL/api/query -H 'Content-Type: application/json' \
      -d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"$exp}}" | jq .value.averageRunCostUsd
  done
done

Divide each average by the eval count (ls -d evals/*/*/ | wc -l) for a per-eval cost, sum over the models and conditions you'll run, and multiply by the repeats. On 2026-09-29 these four models averaged about $22.60 per default suite and $14 per no_guidelines suite over 112 evals, so one eval in both conditions cost about $0.33 per repeat, or $1 for three repeats. Evals with long tasks cost more.

Smoke test

Run the eval for one model. This validates model generation against the new eval:

bash
DISABLE_CONVEX_REPORTING=1 MODELS=anthropic/claude-sonnet-5 TEST_FILTER=<category>/<eval_slug> bun run local:run

DISABLE_CONVEX_REPORTING=1 keeps results local.

If the smoke test fails:

  1. Read the output directory and run.log to understand what happened.
  2. Determine whether it's a test problem (fix the grader) or a model problem (expected, move on).
  3. If the test itself is broken, fix it and re-run before proceeding.

Step 5: Run Against Multiple Models

Start with a smaller representative set of models to calibrate difficulty. If the result is unclear, expand to a broader sweep. Use full OpenRouter slugs from ALL_MODELS in runner/models/index.ts. Run each model in both conditions, default and EVALS_EXPERIMENT=no_guidelines. The difference is the guidelines' measured contribution (AGENTS.md). Launch one background process per model:

bash
# Suggested first-pass set
for m in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 openai/gpt-5.5 deepseek/deepseek-v4-pro; do
  (
    DISABLE_CONVEX_REPORTING=1 MODELS=$m TEST_FILTER=<category>/<eval_slug> bun run local:run
    DISABLE_CONVEX_REPORTING=1 EVALS_EXPERIMENT=no_guidelines MODELS=$m TEST_FILTER=<category>/<eval_slug> bun run local:run
  ) > "/tmp/calibrate-${m//\//_}.log" 2>&1 &
done
wait

If those results are too noisy or too uniform, expand to a broader sweep across providers and tiers from ALL_MODELS. The user can override the list. Any expansion needs a new estimate and approval.

Monitor progress by reading the log files. Each process runs one eval in two conditions, so they should complete in a few minutes.

Run-to-run variance

Run each model at least twice (ideally three times) to distinguish systematic failures from flaky ones. Model output is non-deterministic, so a single pass or fail doesn't tell the whole story. Common variance sources:

  • Hallucinated dependency versions that sometimes resolve and sometimes don't
  • Different code styles across runs (e.g. one integration test vs many unit tests)
  • Different library version choices from training data, where some versions have bugs

Step 6: Review Results and Calibrate

Collect pass/fail from all model runs in both conditions and present a summary table:

Model                        default  no_guidelines
---------------------------  -------  -------------
anthropic/claude-sonnet-5    3/3      1/3
anthropic/claude-opus-4.8    3/3      2/3
openai/gpt-5.5               2/3      0/3
...

Then assess the results:

  • All pass - The eval is likely too easy or the task is too explicit. Recommend tightening the task (remove implementation hints) or adding harder edge cases.
  • All fail - The eval might be too hard, poorly specified, or testing something not covered by the guidelines. Investigate the failures. Common causes: ambiguous requirements, missing context, concept beyond current model capabilities.
  • Mixed results (ideal) - The eval discriminates between model capabilities. Note whether the pass/fail pattern makes sense given model tiers.
  • Unexpected pattern (e.g. only one provider's models fail) - Might indicate a provider-specific quirk rather than a meaningful eval. Investigate before keeping.

Then explicitly ask: is this primarily an eval/task gap, a model gap, or a guideline gap?

  • Eval/task gap - The task is ambiguous, over-specified, under-specified, or the grader is not testing the right thing. Fix the eval first.
  • Model gap - The task is sound, the grading is sound, and failures are what we would expect. Keep the eval.
  • Guideline gap - The failures suggest there should be a guideline that helps here, or an existing guideline is weak/confusing/contradictory. If default and no_guidelines match, the current guidelines aren't moving this eval. After the eval is settled, recommend a change to runner/models/guidelines.md that follows AGENTS.md, then a regression check with the validate-guidelines skill filtered to this eval's category. That run needs its own cost estimate and approval.
Debugging failures

Model output is preserved in the temp directory printed at the start of each run (Using tempdir: ...). These directories are NOT cleaned up automatically. For each model, look at:

  • <tempdir>/output/<provider>/<model>/<category>/<eval>/ for the generated source files
  • run.log in that directory for step-by-step output (install, deploy, tsc, eslint, vitest)
  • node_modules/ for the actual resolved dependency versions
  • package.json for what the model requested vs what was installed

This is essential for distinguishing "model wrote bad code" from "model's code is fine but the grader is too strict" from "dependency version bug".

Push back with specific recommendations if calibration looks off. Suggest concrete changes to the task, answer, or tests.

Summary Checklist

  • Concept and category confirmed with user
  • Convex docs consulted for the feature being tested
  • Relevant existing guidelines checked, with expected implications noted
  • No significant overlap with existing evals (or overlap discussed with user)
  • "Why this matters" written
  • TASK.txt reviewed and approved by user before any files were created
  • Test approach and pipeline discussed, especially if non-standard grading is needed
  • Answer implemented and _generated copied from a sibling eval
  • grader.test.ts written
  • bun run typecheck passes
  • bun run validate:answers passes for the new eval
  • Smoke test passes for at least one model
  • Cost estimate approved before any model run
  • Calibrated on a representative set of models in both conditions, expanded if needed
  • Results reviewed, including eval gap vs model gap vs guideline gap
  • Difficulty is appropriate

© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .cursor/skills/add-eval of get-convex/convex-evals.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 68f5c0e

Compare with similar skills

Add Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Eval this skillget-convex/convex-evals129—~4.6kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from get-convex/convex-evals

  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    129 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Analyze Eval

    get-convex/convex-evals

    Investigate a single failing eval from the convex-evals system.

    129 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Analyze Run

    get-convex/convex-evals

    Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

    129 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Validate Guidelines

    get-convex/convex-evals

    Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

    129 GitHub stars~2.2k tokensUpdated today
    Auto-check: notes

Questions about Add Eval

What does Add Eval do?

Design, implement, validate, and calibrate a new eval for the convex-evals suite. Add Eval is an agent skill from get-convex/convex-evals. Design, implement, validate, and calibrate a new eval for the convex-evals suite.

When should I use Add Eval?

Add Eval fits situations like: the user wants to add a new eval; test a new Convex concept; expand eval coverage.

How do I install Add Eval in Claude Code?

Run `npx skills add get-convex/convex-evals --skill add-eval -a claude-code`. Or copy the skill folder (.cursor/skills/add-eval in get-convex/convex-evals) into .claude/skills/add-eval in your project. Claude Code loads it when a task matches its description.

How do I install Add Eval in Codex?

Run `npx skills add get-convex/convex-evals --skill add-eval -a codex`. Or copy the skill folder (.cursor/skills/add-eval in get-convex/convex-evals) into .agents/skills/add-eval in your project. Codex loads it when a task matches its description.

Can I use Add Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill add-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-eval, .gemini/skills/add-eval, .github/skills/add-eval and .opencode/skills/add-eval in your project.

What does Add Eval need to run?

Going by SKILL.md and its folder, Add Eval needs the command-line tools its instructions call (bun, curl, jq and bunx).

Does Add Eval access the network?

SKILL.md names 2 domains. In commands or code: docs.convex.dev and fabulous-panther-525.convex.cloud; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Add Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Eval use?

Add Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Eval use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Eval?

Skills that share tags, products or a category with Add Eval: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Eval?

get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.