Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install get-convex/convex-evals validate-guidelines --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/validate-guidelines .claude/skills/validate-guidelines && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .claude/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelinesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install get-convex/convex-evals validate-guidelines --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/validate-guidelines .agents/skills/validate-guidelines && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .agents/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install get-convex/convex-evals validate-guidelines --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/validate-guidelines .cursor/skills/validate-guidelines && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .cursor/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/get-convex/convex-evals.git --path .cursor/skills/validate-guidelines--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install get-convex/convex-evals validate-guidelines --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/validate-guidelines .gemini/skills/validate-guidelines && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .gemini/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install get-convex/convex-evals validate-guidelinesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/validate-guidelines .github/skills/validate-guidelines && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .github/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install get-convex/convex-evals validate-guidelines --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/validate-guidelines .opencode/skills/validate-guidelines && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "validate-guidelines" agent skill from https://github.com/get-convex/convex-evals/tree/main/.cursor/skills/validate-guidelines into .opencode/skills/validate-guidelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-guidelines", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
validate-guidelinesEmpirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
Validate Guidelines is an agent skill from get-convex/convex-evals. Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Use when proposing or reviewing changes to runner/models/guidelines.md, or when the user asks to validate guidelines.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitcurljqbunghFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
fabulous-panther-525.convex.cloudFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENROUTER_API_KEYCONVEX_AUTH_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Validate Guidelines loads about 2.2k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 1,022 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
s only `OPENROUTER_API_KEY` in the root `.env`. Without it the script skips every model and still prints "Safe to commit`OPENROUTER_API_KEY` is loaded from `.env` via dotenv (see AGENTS.md). The script does not report to Convex.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 1,022 words, ~2,223 tokens.
.claude/skills/validate-guidelines/SKILL.md (or your agent's skills folder).runner/models/guidelines.md and wants to ensure they don't regress other modelsGuideline changes are validated by running evals twice per model: once with the current (before) guidelines and once with the proposed (after) guidelines. Results are compared; any eval that passed before and fails after is a regression. The goal is to ensure changes improve or at least do not regress scores across multiple models.
Determine which sections of runner/models/guidelines.md were modified (its ## headings, e.g. ## Function guidelines, ## Query guidelines, ## File storage guidelines) and the intent (new rule, clarification, token compaction). runner/models/guidelines.md is the source. runner/models/guidelines.ts only loads it.
Do what the Validate Guideline Changes workflow (.github/workflows/validate_guidelines.yml) does. Write both files to a temp directory outside the repo:
git fetch origin
git show origin/main:runner/models/guidelines.md > <tmp>/before.mdrunner/models/guidelines.md, then cp runner/models/guidelines.md <tmp>/after.md.The script sends each file as-is as the guidelines. Ensure both paths are absolute or relative to the repo root and that the script can read them.
Choose a --filter regex on eval category/name from this list, or omit it for the full suite:
## Function guidelines (http, validators, registration, calling, pagination): 000-fundamentals|006-clients, or full
## Schema guidelines: 001-data_modeling
## Authentication guidelines: 005-idioms/003|005-idioms/004
## Typescript guidelines: omit (run all)
## Full text search guidelines: 002-queries/009|002-queries/020
## Vector search guidelines: 004-actions/008
## Component guidelines: 007-components
## Query guidelines: 002-queries
## Mutation guidelines: 003-mutations
## Action guidelines: 004-actions
## Scheduling guidelines: 000-fundamentals/003|000-fundamentals/004
## Testing guidelines: 005-idioms/005
## File storage guidelines: 000-fundamentals/007|004-actions/004|004-actions/005
Targeted change (e.g. one section): use a filter that matches the evals most likely affected.
Broad change (e.g. wording across many sections): omit --filter to run all evals.
Default set (preferred for validation, and the workflow's default): anthropic/claude-sonnet-5, anthropic/claude-opus-4.8, deepseek/deepseek-v4-pro, openai/gpt-5.5. Each model must be a full OpenRouter slug in ALL_MODELS (runner/models/index.ts). The script exits if one isn't.
Every model runs through OpenRouter, so the script needs only OPENROUTER_API_KEY in the root .env. Without it the script skips every model and still prints "Safe to commit", so check that the summary has a row per model. Use a subset to cut cost; at least two models are recommended.
Each model runs the selected evals twice, before and after, in the default condition. A run without --filter is 2 full suites per model. Take each model's average full-suite cost from the public production query modelScores:getSchedulingStats, as in the add-model skill:
URL=https://fabulous-panther-525.convex.cloud
for slug in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 deepseek/deepseek-v4-pro openai/gpt-5.5; do
ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
-d "{\"path\":\"models:getBySlug\",\"args\":{\"slug\":\"$slug\"}}" | jq -r .value._id)
printf "%s " "$slug"
curl -s $URL/api/query -H 'Content-Type: application/json' \
-d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"}}" | jq .value.averageRunCostUsd
doneEstimate = 2 x the sum of the averages x the share of evals the filter matches. Count with ls -d evals/*/*/ | wc -l and ls -d evals/*/*/ | sed 's|^evals/||; s|/$||' | grep -cE '<filter>'. On 2026-09-29 the four defaults averaged about $22.60 per full suite, so an unfiltered run cost about $45 and --filter "002-queries" (28 of 112 evals) about $11.
Recommend --filter to the categories the change affects. State the estimate and wait for approval before running. Every re-run needs its own estimate and approval.
Do not set CONVEX_EVAL_URL or CONVEX_AUTH_TOKEN so results stay local.
bun run validate:guidelines --before <tmp>/before.md --after <tmp>/after.md --models anthropic/claude-sonnet-5,anthropic/claude-opus-4.8,deepseek/deepseek-v4-pro,openai/gpt-5.5 --filter "002-queries"Optional: --output <path> to write the JSON summary to a specific file. By default it is written to guideline-validation/results/<timestamp>.json.
The script runs each model sequentially: first all evals with "before" guidelines, then all evals with "after" guidelines. Pass/fail is collected and deltas are computed.
IMPORTANT: You must orchestrate the entire run end-to-end. Start the command in the background with its output redirected to a log file, then read the log periodically until the run finishes (look for the GUIDELINE VALIDATION SUMMARY banner and the process exiting). Use exponential backoff for polling (e.g. 30s, 60s, 120s). Do NOT return to the user until the run is fully complete and you have read and analyzed the results. The user expects a complete report, not a "check back later" handoff.
Validate Guideline Changes (.github/workflows/validate_guidelines.yml) runs the same script in CI. It is manual dispatch only. It compares origin/main (input base_ref) against runner/models/guidelines.md on the dispatched branch, and each model (input models, default the four above) runs the suite twice, before and after. That makes it a paid run. Estimate the cost as in Step 5 and get approval before dispatching, the same as a local run. Push the branch first, then:
gh workflow run validate_guidelines.yml --ref <branch> -f filter='002-queries'Read the summary in the job log (gh run view <id> --log). The JSON summary is uploaded as the guideline-validation-<run id> artifact.
The script prints:
Read the script output and present the full summary table and verdict to the user.
bun run validate:guidelines --before <path> --after <path> --models <m1,m2,...> [--filter <regex>] [--output <path>]--before, --after: Paths to guideline markdown files (current vs proposed).--models: Comma-separated OpenRouter slugs from ALL_MODELS in runner/models/index.ts (e.g. openai/gpt-5.5, anthropic/claude-sonnet-5).--filter: Optional regex on eval category/name (e.g. 005-idioms or 002-queries/015).--output: Optional path for the JSON summary file.OPENROUTER_API_KEY is loaded from .env via dotenv (see AGENTS.md). The script does not report to Convex.
© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/validate-guidelines of get-convex/convex-evals.
Open the folder on GitHubat commit 68f5c0e
Validate Guidelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Validate Guidelines this skillget-convex/convex-evals | 129 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 15 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
get-convex/convex-evals
Design, implement, validate, and calibrate a new eval for the convex-evals suite.
get-convex/convex-evals
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
get-convex/convex-evals
Investigate a single failing eval from the convex-evals system.
get-convex/convex-evals
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.
Categories
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Validate Guidelines is an agent skill from get-convex/convex-evals. Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.
Validate Guidelines fits situations like: reviewing changes to runner/models/guidelines.md; the user asks to validate guidelines.
Run `npx skills add get-convex/convex-evals --skill validate-guidelines -a claude-code`. Or copy the skill folder (.cursor/skills/validate-guidelines in get-convex/convex-evals) into .claude/skills/validate-guidelines in your project. Claude Code loads it when a task matches its description.
Run `npx skills add get-convex/convex-evals --skill validate-guidelines -a codex`. Or copy the skill folder (.cursor/skills/validate-guidelines in get-convex/convex-evals) into .agents/skills/validate-guidelines in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill validate-guidelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-guidelines, .gemini/skills/validate-guidelines, .github/skills/validate-guidelines and .opencode/skills/validate-guidelines in your project.
Going by SKILL.md and its folder, Validate Guidelines needs the command-line tools its instructions call (git, curl, jq, bun and gh) and credentials named OPENROUTER_API_KEY and CONVEX_AUTH_TOKEN. Our summary lists: A credential in OPENROUTER_API_KEY; A credential in CONVEX_AUTH_TOKEN.
SKILL.md names 1 domain. In commands or code: fabulous-panther-525.convex.cloud; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Validate Guidelines is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Validate Guidelines: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.
Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.