Agent skill

Validate Guidelines

by get-convex in get-convex/convex-evals

Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Validate Guidelines

skills CLI
$ npx skills add get-convex/convex-evals --skill validate-guidelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install get-convex/convex-evals validate-guidelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/get-convex/convex-evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/validate-guidelines .claude/skills/validate-guidelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-guidelines
GitHub stars
129
Token cost
~2.2k tokens
SKILL.md length
1,022 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

  • Works in 8 steps: Identify the change → Build before and after guideline files → Select target evals → …
  • Reviewing changes to runner/models/guidelines.md
  • SKILL.md covers When to use, Overview, Step 1: Identify the change and Step 2: Build before and after…, plus 7 more sections
  • Calls git, curl and jq; reaches fabulous-panther-525.convex.cloud; needs OPENROUTER_API_KEY and CONVEX_AUTH_TOKEN

What it does

Validate Guidelines is an agent skill from get-convex/convex-evals. Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Use when proposing or reviewing changes to runner/models/guidelines.md, or when the user asks to validate guidelines.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The licence is Apache-2.0.

When your agent uses it

  • Reviewing changes to runner/models/guidelines.md
  • The user asks to validate guidelines

Example prompts

  • “/validate-guidelines”

Requirements

  • A credential in OPENROUTER_API_KEY
  • A credential in CONVEX_AUTH_TOKEN

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Identify the change
  2. Build before and after guideline files
  3. Select target evals
  4. Select models
  5. Estimate the cost and get approval
  6. Run the validation script and monitor to completion
  7. Parse and report results
  8. Recommend next steps

What it can do on your machine

Read from SKILL.md and the folder at commit 68f5c0e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • curl
    • jq
    • bun
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • fabulous-panther-525.convex.cloud

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • CONVEX_AUTH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validate Guidelines loads about 2.2k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 1,022 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:60
    s only `OPENROUTER_API_KEY` in the root `.env`. Without it the script skips every model and still prints "Safe to commit
  • NoteMentions a .env fileSKILL.md:136
    `OPENROUTER_API_KEY` is loaded from `.env` via dotenv (see AGENTS.md). The script does not report to Convex.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from get-convex/convex-evals at commit 68f5c0e, republished under its Apache-2.0 licence (© get-convex). 1,022 words, ~2,223 tokens.

Download SKILL.mdSave it as .claude/skills/validate-guidelines/SKILL.md (or your agent's skills folder).
name
validate-guidelines
description
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Use when proposing or reviewing changes to runner/models/guidelines.md, or when the user asks to validate guidelines.

Validate Guidelines

When to use

  • User proposes or has made changes to runner/models/guidelines.md and wants to ensure they don't regress other models
  • User says "validate the guideline changes" or "run the guideline validation"
  • Before committing guideline edits, to confirm improvements or no-regression across the default models (or a subset)

Overview

Guideline changes are validated by running evals twice per model: once with the current (before) guidelines and once with the proposed (after) guidelines. Results are compared; any eval that passed before and fails after is a regression. The goal is to ensure changes improve or at least do not regress scores across multiple models.

Step 1: Identify the change

Determine which sections of runner/models/guidelines.md were modified (its ## headings, e.g. ## Function guidelines, ## Query guidelines, ## File storage guidelines) and the intent (new rule, clarification, token compaction). runner/models/guidelines.md is the source. runner/models/guidelines.ts only loads it.

Step 2: Build before and after guideline files

Do what the Validate Guideline Changes workflow (.github/workflows/validate_guidelines.yml) does. Write both files to a temp directory outside the repo:

  • Before: the committed guidelines on main.
    bash
    git fetch origin
    git show origin/main:runner/models/guidelines.md > <tmp>/before.md
  • After: apply the proposed edits to runner/models/guidelines.md, then cp runner/models/guidelines.md <tmp>/after.md.

The script sends each file as-is as the guidelines. Ensure both paths are absolute or relative to the repo root and that the script can read them.

Step 3: Select target evals

Choose a --filter regex on eval category/name from this list, or omit it for the full suite:

  • ## Function guidelines (http, validators, registration, calling, pagination): 000-fundamentals|006-clients, or full

  • ## Schema guidelines: 001-data_modeling

  • ## Authentication guidelines: 005-idioms/003|005-idioms/004

  • ## Typescript guidelines: omit (run all)

  • ## Full text search guidelines: 002-queries/009|002-queries/020

  • ## Vector search guidelines: 004-actions/008

  • ## Component guidelines: 007-components

  • ## Query guidelines: 002-queries

  • ## Mutation guidelines: 003-mutations

  • ## Action guidelines: 004-actions

  • ## Scheduling guidelines: 000-fundamentals/003|000-fundamentals/004

  • ## Testing guidelines: 005-idioms/005

  • ## File storage guidelines: 000-fundamentals/007|004-actions/004|004-actions/005

  • Targeted change (e.g. one section): use a filter that matches the evals most likely affected.

  • Broad change (e.g. wording across many sections): omit --filter to run all evals.

Step 4: Select models

Default set (preferred for validation, and the workflow's default): anthropic/claude-sonnet-5, anthropic/claude-opus-4.8, deepseek/deepseek-v4-pro, openai/gpt-5.5. Each model must be a full OpenRouter slug in ALL_MODELS (runner/models/index.ts). The script exits if one isn't.

Every model runs through OpenRouter, so the script needs only OPENROUTER_API_KEY in the root .env. Without it the script skips every model and still prints "Safe to commit", so check that the summary has a row per model. Use a subset to cut cost; at least two models are recommended.

Step 5: Estimate the cost and get approval

Each model runs the selected evals twice, before and after, in the default condition. A run without --filter is 2 full suites per model. Take each model's average full-suite cost from the public production query modelScores:getSchedulingStats, as in the add-model skill:

bash
URL=https://fabulous-panther-525.convex.cloud
for slug in anthropic/claude-sonnet-5 anthropic/claude-opus-4.8 deepseek/deepseek-v4-pro openai/gpt-5.5; do
  ID=$(curl -s $URL/api/query -H 'Content-Type: application/json' \
    -d "{\"path\":\"models:getBySlug\",\"args\":{\"slug\":\"$slug\"}}" | jq -r .value._id)
  printf "%s " "$slug"
  curl -s $URL/api/query -H 'Content-Type: application/json' \
    -d "{\"path\":\"modelScores:getSchedulingStats\",\"args\":{\"modelId\":\"$ID\"}}" | jq .value.averageRunCostUsd
done

Estimate = 2 x the sum of the averages x the share of evals the filter matches. Count with ls -d evals/*/*/ | wc -l and ls -d evals/*/*/ | sed 's|^evals/||; s|/$||' | grep -cE '<filter>'. On 2026-09-29 the four defaults averaged about $22.60 per full suite, so an unfiltered run cost about $45 and --filter "002-queries" (28 of 112 evals) about $11.

Recommend --filter to the categories the change affects. State the estimate and wait for approval before running. Every re-run needs its own estimate and approval.

Step 6: Run the validation script and monitor to completion

Do not set CONVEX_EVAL_URL or CONVEX_AUTH_TOKEN so results stay local.

bash
bun run validate:guidelines --before <tmp>/before.md --after <tmp>/after.md --models anthropic/claude-sonnet-5,anthropic/claude-opus-4.8,deepseek/deepseek-v4-pro,openai/gpt-5.5 --filter "002-queries"

Optional: --output <path> to write the JSON summary to a specific file. By default it is written to guideline-validation/results/<timestamp>.json.

The script runs each model sequentially: first all evals with "before" guidelines, then all evals with "after" guidelines. Pass/fail is collected and deltas are computed.

IMPORTANT: You must orchestrate the entire run end-to-end. Start the command in the background with its output redirected to a log file, then read the log periodically until the run finishes (look for the GUIDELINE VALIDATION SUMMARY banner and the process exiting). Use exponential backoff for polling (e.g. 30s, 60s, 120s). Do NOT return to the user until the run is fully complete and you have read and analyzed the results. The user expects a complete report, not a "check back later" handoff.

Show full SKILL.md (342 more words)Show less
Or dispatch the GitHub workflow

Validate Guideline Changes (.github/workflows/validate_guidelines.yml) runs the same script in CI. It is manual dispatch only. It compares origin/main (input base_ref) against runner/models/guidelines.md on the dispatched branch, and each model (input models, default the four above) runs the suite twice, before and after. That makes it a paid run. Estimate the cost as in Step 5 and get approval before dispatching, the same as a local run. Push the branch first, then:

bash
gh workflow run validate_guidelines.yml --ref <branch> -f filter='002-queries'

Read the summary in the job log (gh run view <id> --log). The JSON summary is uploaded as the guideline-validation-<run id> artifact.

Step 7: Parse and report results

The script prints:

  1. A comparison table: per-model before pass count, after pass count, delta, number of regressions, number of improvements.
  2. Regressions: evals that passed before and failed after (by model).
  3. Improvements: evals that failed before and passed after (by model).
  4. A verdict line: either "REGRESSIONS DETECTED" or "Safe to commit."

Read the script output and present the full summary table and verdict to the user.

  • If there are regressions: list them and recommend reverting or narrowing the guideline change; optionally run analyze-eval on a regression to see why it failed.
  • If there are no regressions: recommend committing the guideline change; mention any improvements.

Step 8: Recommend next steps

  • No regressions, with or without improvements: Safe to commit the guideline changes.
  • Any regressions: Do not commit. Suggest reverting the change or narrowing it (e.g. only add the new rule to a subsection that doesn’t affect the regressed eval). Re-run validation after adjusting.
  • Unclear or noisy: If only one model regresses one eval, consider re-running that model to check for flakiness, or run the full suite once more.

Reference: Script usage

bun run validate:guidelines --before <path> --after <path> --models <m1,m2,...> [--filter <regex>] [--output <path>]
  • --before, --after: Paths to guideline markdown files (current vs proposed).
  • --models: Comma-separated OpenRouter slugs from ALL_MODELS in runner/models/index.ts (e.g. openai/gpt-5.5, anthropic/claude-sonnet-5).
  • --filter: Optional regex on eval category/name (e.g. 005-idioms or 002-queries/015).
  • --output: Optional path for the JSON summary file.

OPENROUTER_API_KEY is loaded from .env via dotenv (see AGENTS.md). The script does not report to Convex.

© get-convex, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/validate-guidelines of get-convex/convex-evals.

Open the folder on GitHubat commit 68f5c0e

Compare with similar skills

Validate Guidelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Guidelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Guidelines this skillget-convex/convex-evals129—~2.2kAutomated safety check: NotesApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k15 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 15 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from get-convex/convex-evals

  • Add Eval

    get-convex/convex-evals

    Design, implement, validate, and calibrate a new eval for the convex-evals suite.

    129 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    129 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Analyze Eval

    get-convex/convex-evals

    Investigate a single failing eval from the convex-evals system.

    129 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Analyze Run

    get-convex/convex-evals

    Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

    129 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Questions about Validate Guidelines

What does Validate Guidelines do?

Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Validate Guidelines is an agent skill from get-convex/convex-evals. Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions.

When should I use Validate Guidelines?

Validate Guidelines fits situations like: reviewing changes to runner/models/guidelines.md; the user asks to validate guidelines.

How do I install Validate Guidelines in Claude Code?

Run `npx skills add get-convex/convex-evals --skill validate-guidelines -a claude-code`. Or copy the skill folder (.cursor/skills/validate-guidelines in get-convex/convex-evals) into .claude/skills/validate-guidelines in your project. Claude Code loads it when a task matches its description.

How do I install Validate Guidelines in Codex?

Run `npx skills add get-convex/convex-evals --skill validate-guidelines -a codex`. Or copy the skill folder (.cursor/skills/validate-guidelines in get-convex/convex-evals) into .agents/skills/validate-guidelines in your project. Codex loads it when a task matches its description.

Can I use Validate Guidelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-convex/convex-evals --skill validate-guidelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-guidelines, .gemini/skills/validate-guidelines, .github/skills/validate-guidelines and .opencode/skills/validate-guidelines in your project.

What does Validate Guidelines need to run?

Going by SKILL.md and its folder, Validate Guidelines needs the command-line tools its instructions call (git, curl, jq, bun and gh) and credentials named OPENROUTER_API_KEY and CONVEX_AUTH_TOKEN. Our summary lists: A credential in OPENROUTER_API_KEY; A credential in CONVEX_AUTH_TOKEN.

Does Validate Guidelines access the network?

SKILL.md names 1 domain. In commands or code: fabulous-panther-525.convex.cloud; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Validate Guidelines safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Validate Guidelines use?

Validate Guidelines is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Guidelines use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validate Guidelines?

Skills that share tags, products or a category with Validate Guidelines: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Guidelines?

get-convex (a GitHub organization) maintains it in get-convex/convex-evals, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: get-convex/convex-evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.