Evals Context
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.
$ npx skills add supabase/evals --skill docs-eval-planning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install supabase/evals docs-eval-planning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .claude/skills/docs-eval-planning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .claude/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add supabase/evals --skill docs-eval-planning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install supabase/evals docs-eval-planning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .agents/skills/docs-eval-planning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .agents/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add supabase/evals --skill docs-eval-planning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install supabase/evals docs-eval-planning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .cursor/skills/docs-eval-planning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .cursor/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/supabase/evals.git --path evals/docs/.claude/skills/docs-eval-planning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add supabase/evals --skill docs-eval-planning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install supabase/evals docs-eval-planning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .gemini/skills/docs-eval-planning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .gemini/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install supabase/evals docs-eval-planningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add supabase/evals --skill docs-eval-planning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .github/skills/docs-eval-planning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .github/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add supabase/evals --skill docs-eval-planning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install supabase/evals docs-eval-planning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .opencode/skills/docs-eval-planning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "docs-eval-planning" agent skill from https://github.com/supabase/evals/tree/main/evals/docs/.claude/skills/docs-eval-planning into .opencode/skills/docs-eval-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "docs-eval-planning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
docs-eval-planningPlan a documentation eval for supabase/evals, where a docs guide is the subject under test.
Docs Eval Planning is an agent skill from supabase/evals, published by the product's own GitHub organization. Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. Use when asked to write, add, or design an eval for a Supabase docs guide, when a ticket asks for a deterministic eval on a page, or when deciding what a guide-under-test eval should check. Produces a plan, not files. Not for debugging a scorer, running an eval, or writing example solutions.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/check-rules.md`, `references/flakiness.md` and `references/prompt-rules.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. It works with Supabase. The repository describes itself as: Evaluating agents across Supabase. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6d8b6bd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Docs Eval Planning loads about 2.9k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,781 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from supabase/evals at commit 6d8b6bd, republished under its Apache-2.0 licence (© supabase). 1,781 words, ~2,884 tokens.
.claude/skills/docs-eval-planning/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.A documentation eval measures a page, not an agent. The prompt sends a user's request plus the page's url, and the checks say whether an agent that read the page produced working code. A gap in the page counts as a failure.
That framing is the whole difficulty. The failure modes are collected in references/flakiness.md. Read it before designing checks, not after.
A plan. It does not write PROMPT.md, EVAL.ts, or a seed. Implementation and local scoring are a
separate job.
Model a new eval on build-docs-003-api-keys-guide or later.
build-docs-NNN-<subject>, named after the doc rather than the feature. The id is permanent. The
published results series is keyed on it, so renaming breaks history. So does renaming a check.
evals/docs/build-docs-NNN-<subject>/
PROMPT.md frontmatter and the task the agent sees
EVAL.ts the scorer
README.md the design rationale, addressed to the next editor
<helper>.ts flat beside EVAL.ts, never in a subdirectory
local/ the seed workspaceCONTRIBUTING.md makes the README.md optional. For a page under test it is required: it is the only
place the stripped-word list and the constraints that must not be edited away live. Its sections:
EVAL.ts is orchestration. Import named check* functions, assemble one flat array, return
{ passed: checks.every(c => c.passed), checks }, and wrap it in a try/catch whose catch returns a
single self-named failure check. Split the implementations into modules and keep the full check list
declared in EVAL.ts. Comment why the phase order is what it is.
Every page-under-test eval carries a check that the referenced page was read with content. Import
checkDocsGuideRead from core if it has landed there; otherwise copy the most recent eval's. It
resolves the url from the harness's docs result rather than the raw tool call, because a search_docs
hit carries the guide's url in its result rather than its request.
Write nothing in this phase or the next four. Gathering, choosing the claim, drafting the prompt, and designing the checks are all reading. The output is a plan.
Where the harness offers a planning mode that enforces read-only, enter it first and let it hold the
line. In Claude Code that is EnterPlanMode. Where it does not, the discipline is yours to keep.
Then read three things:
CONTRIBUTING.md, which carries the repo's rules for any eval: suite selection, the folder shape,
the motivation: requirement, prompt discipline, scorer discipline, and how to refresh results in
CI. This skill defers to it and does not restate it. Where this skill goes further, it says so..md variant.Work through references/sources.md. Run independent sources in parallel.
The checks come from outside the page. A check list derived from the guide inherits the guide's blind spots, so every check passes and the score reports that the guide is fine when nobody looked. Build the list from the sources, then diff it against the page. Anything the sources treat as essential and the page omits is a candidate check, and it should fail a solution written from the page. That failure is the finding the paired docs ticket acts on.
Several sources need Supabase employee access. Settle which ones you can reach before you start, plan from those, and ignore the rest. Missing access is never a blocker. Record in the plan which sources the inventory came from and which were out of reach.
This phase ends with a failure-point inventory: one bullet per finding, each carrying an issue id, a thread, or a url. A finding with no source does not go in.
Name the single thing the guide has to transmit, in the user's terms rather than the product's.
One claim, not three. An eval that measures several unrelated things reports which of them an agent got, and a docs ticket cannot act on that. It also means every check hangs off one subject, which is what makes a saturated check obvious later.
A dense page yields ten candidates, all evidenced. Pick on consequence: the failure that costs a user the most, and among equals the one the page itself warns about. A page that carries a danger callout has already named its own worst case.
Then confirm the harness can measure it, before designing a single check. A claim the harness cannot observe is not the claim, however well evidenced. Three questions, and read the code for the answers rather than assuming:
ctx.exec, not on the host: the host-side build helpers relink the workspace's
node_modules to the framework's, which does not carry an eval's own dependencies. Check whether any
existing eval runs the same toolchain, and treat a first as a risk to state in the plan.Write the answers into the plan. A feasibility problem found here is a design change; found in Phase 4 it is a rewrite.
Work through references/prompt-rules.md.
The prompt is a product request plus the page's url, written the way the user would write it, with the vocabulary that gives away the answer removed.
The seed is where the detail goes. Write seed comments in product vocabulary, never the vocabulary the prompt strips: the endpoint the rest of the team builds against, the shapes a handler must return, which table the API serves. Pre-solve what belongs to another scenario and say so in the comment, so a mistake there cannot fail this eval for the wrong reason. Choose fixture values that a memorized answer gets wrong, so a whitelist check measures something. Seed a conflict where the subject allows one, so a single approach cannot satisfy every case.
Where the task contract lives is a real choice. Put it in seed comments and gain a positive control
the scorer can prove, at the cost of a discovery question. Leave it out and measure whether the agent
derives it, which is harder and less observable. Decide deliberately and record which in the
README.md.
This phase ends with the prompt body, the list of stripped words, the seed, and the motivation:
frontmatter.
Work through references/check-rules.md, in order. The first three rules matter most.
Use a judge only when the artifact class is unbounded, meaning free prose or files whose shape you cannot predict. Use a deterministic check for anything a query or a file read settles. Scope the rubric to what the check is worth, state what not to grade, give the judge a tie-break, and make one claim per check. An eval passes only when every check passes, so an ambitious rubric on a secondary check fails the whole page.
This phase ends with a table: the check name, what it proves, what it does when the object is absent, whether it reads files or runs code, and the evidence it came from.
Then name the checks you expect to saturate, and write the prediction down. A check whose answer the seed labels costs an agent nothing and carries no signal. A plan that says which checks are cheap is honest about how much the eval measures.
Sections: context, failure-point inventory, the seed, the prompt, the check table, what is not scored, the expected-failure table, verification, the paired docs ticket, and Phase 6 as the closing step.
Everything in the prompt is either measured or named as out of scope. A reviewer finds the rule that is neither.
Verification means example solutions: one you believe is correct, plus a few carrying a single
deliberate flaw each. Write down which checks you expect each to fail before running anything; that
list is the test. A flaw usually trips several checks, so do not aim for exactly one failure each.
Never commit them. evals/*/*/solutions/ is a git exclusion, and you add it to .git/info/exclude
before the first git add of an eval directory, because that file is local to your clone. A
transcript-anchored check cannot be exercised by a solution, because no agent ran.
The plan's last step is to append whatever broke to references/flakiness.md once the baseline lands, with the PR number.
This is a plan step rather than advice, because the file is only worth having if it grows. Ask reviewers to write findings into it directly rather than leaving them in a review thread, where the next author does not look.
Two kinds of check, and a plan needs both.
A check that reads files asserts content shape. It is cheap, stable, and has no opinion about whether anything runs. Every one of them passes for an agent that edited a config file and stopped.
A check that runs the code asserts the outcome. It is what stops a decorative pass, and it is slower and more exposed to the environment.
The table says which each check is, because the two fail for different reasons and a reviewer reading a failure needs to know which kind they are looking at.
Some claims are not this skill's to make. Whether the page carries runnable commands in fenced blocks, declares its environment variables and prerequisites, and ends with a verification step is a claim about the page's own markdown. That belongs to the page and to the project's authoring rubric. No tooling in this repo asserts it; the docs end-to-end suite is rendering, links, and accessibility.
What belongs here is the claim about what an agent produces after reading the page. Mixing the two produces an eval that fails when someone reformats a code block.
© supabase, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in evals/docs/.claude/skills/docs-eval-planning of supabase/evals.
Open the folder on GitHubat commit 6d8b6bd
Docs Eval Planning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Docs Eval Planning this skillsupabase/evals | 143 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Evals Contextzgsm-ai/costrict | 4.4k | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| AI Project Copilotsun461941-hub/ai-project-copilot | 100 | — | ~3k | Automated safety check: Pass | MIT | |
| Eee Dataset Conversionevaleval/every_eval_ever | 133 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Analyze Id Eval Rankingopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 |
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
sun461941-hub/ai-project-copilot
A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…
evaleval/every_eval_ever
Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
wso2/agent-manager
Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.
supabase/evals
Manages shadcn components and projects — adding, searching, fixing, debugging, styling, and composing UI.
Works with
Categories
Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. Docs Eval Planning is an agent skill from supabase/evals, published by the product's own GitHub organization. Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.
Docs Eval Planning fits situations like: design an eval for a Supabase docs guide; A ticket asks for a deterministic eval on a page; deciding what a guide-under-test eval should check.
Run `npx skills add supabase/evals --skill docs-eval-planning -a claude-code`. Or copy the skill folder (evals/docs/.claude/skills/docs-eval-planning in supabase/evals) into .claude/skills/docs-eval-planning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add supabase/evals --skill docs-eval-planning -a codex`. Or copy the skill folder (evals/docs/.claude/skills/docs-eval-planning in supabase/evals) into .agents/skills/docs-eval-planning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add supabase/evals --skill docs-eval-planning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/docs-eval-planning, .gemini/skills/docs-eval-planning, .github/skills/docs-eval-planning and .opencode/skills/docs-eval-planning in your project.
Going by SKILL.md and its folder, Docs Eval Planning needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Docs Eval Planning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Docs Eval Planning: Evals Context (zgsm-ai/costrict, 4.4k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 100 stars) and Eee Dataset Conversion (evaleval/every_eval_ever, 133 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
supabase (a GitHub organization, an official publisher) maintains it in supabase/evals, which has 143 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.
Source: supabase/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.