Official agent skill

Docs Eval Planning

by supabase in supabase/evals

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Docs Eval Planning

skills CLI
$ npx skills add supabase/evals --skill docs-eval-planning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install supabase/evals docs-eval-planning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/supabase/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evals/docs/.claude/skills/docs-eval-planning .claude/skills/docs-eval-planning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
docs-eval-planning
GitHub stars
143
Token cost
~2.9k tokens
SKILL.md length
1,781 words
Files
5 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

  • Works in 7 steps: read-only, and the repo rules → gather the evidence → pick the one claim, and confirm it can… → …
  • Design an eval for a Supabase docs guide
  • SKILL.md covers What this skill produces, The shape it takes, Phase 0: read-only, and the… and Phase 1: gather the evidence, plus 7 more sections
  • Calls git

What it does

Docs Eval Planning is an agent skill from supabase/evals, published by the product's own GitHub organization. Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. Use when asked to write, add, or design an eval for a Supabase docs guide, when a ticket asks for a deterministic eval on a page, or when deciding what a guide-under-test eval should check. Produces a plan, not files. Not for debugging a scorer, running an eval, or writing example solutions.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/check-rules.md`, `references/flakiness.md` and `references/prompt-rules.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. It works with Supabase. The repository describes itself as: Evaluating agents across Supabase. The licence is Apache-2.0.

When your agent uses it

  • Design an eval for a Supabase docs guide
  • A ticket asks for a deterministic eval on a page
  • Deciding what a guide-under-test eval should check

Example prompts

  • “/docs-eval-planning”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. read-only, and the repo rules
  2. gather the evidence
  3. pick the one claim, and confirm it can be measured
  4. draft the prompt and the seed
  5. design the checks
  6. write the plan
  7. feed the catalog

What it can do on your machine

Read from SKILL.md and the folder at commit 6d8b6bd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Docs Eval Planning loads about 2.9k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,781 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from supabase/evals at commit 6d8b6bd, republished under its Apache-2.0 licence (© supabase). 1,781 words, ~2,884 tokens.

Download SKILL.mdSave it as .claude/skills/docs-eval-planning/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
docs-eval-planning
description
Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. Use when asked to write, add, or design an eval for a Supabase docs guide, when a ticket asks for a deterministic eval on a page, or when deciding what a guide-under-test eval should check. Produces a plan, not files. Not for debugging a scorer, running an eval, or writing example solutions.

Planning a documentation eval

A documentation eval measures a page, not an agent. The prompt sends a user's request plus the page's url, and the checks say whether an agent that read the page produced working code. A gap in the page counts as a failure.

That framing is the whole difficulty. The failure modes are collected in references/flakiness.md. Read it before designing checks, not after.

What this skill produces

A plan. It does not write PROMPT.md, EVAL.ts, or a seed. Implementation and local scoring are a separate job.

The shape it takes

Model a new eval on build-docs-003-api-keys-guide or later.

build-docs-NNN-<subject>, named after the doc rather than the feature. The id is permanent. The published results series is keyed on it, so renaming breaks history. So does renaming a check.

evals/docs/build-docs-NNN-<subject>/
  PROMPT.md      frontmatter and the task the agent sees
  EVAL.ts        the scorer
  README.md      the design rationale, addressed to the next editor
  <helper>.ts    flat beside EVAL.ts, never in a subdirectory
  local/         the seed workspace

CONTRIBUTING.md makes the README.md optional. For a page under test it is required: it is the only place the stripped-word list and the constraints that must not be edited away live. Its sections:

  • What this eval measures. The page is the subject, not the agent.
  • Do not reintroduce the vocabulary. The stripped words, listed.
  • The seed carries the contract. What the seed fixes, and what each fixed thing buys and costs.
  • Do not drop the positive controls. Which checks pass for an agent that built nothing, and which ones make them mean something.
  • The guide has to actually be read.
  • What this eval does not score. Each entry with its reason. An unmeasured rule from the prompt is named here.

EVAL.ts is orchestration. Import named check* functions, assemble one flat array, return { passed: checks.every(c => c.passed), checks }, and wrap it in a try/catch whose catch returns a single self-named failure check. Split the implementations into modules and keep the full check list declared in EVAL.ts. Comment why the phase order is what it is.

Every page-under-test eval carries a check that the referenced page was read with content. Import checkDocsGuideRead from core if it has landed there; otherwise copy the most recent eval's. It resolves the url from the harness's docs result rather than the raw tool call, because a search_docs hit carries the guide's url in its result rather than its request.

Phase 0: read-only, and the repo rules

Write nothing in this phase or the next four. Gathering, choosing the claim, drafting the prompt, and designing the checks are all reading. The output is a plan.

Where the harness offers a planning mode that enforces read-only, enter it first and let it hold the line. In Claude Code that is EnterPlanMode. Where it does not, the discipline is yours to keep.

Then read three things:

  1. CONTRIBUTING.md, which carries the repo's rules for any eval: suite selection, the folder shape, the motivation: requirement, prompt discipline, scorer discipline, and how to refresh results in CI. This skill defers to it and does not restate it. Where this skill goes further, it says so.
  2. The guide under test, as its .md variant.
  3. references/flakiness.md, before designing anything rather than after.

Phase 1: gather the evidence

Work through references/sources.md. Run independent sources in parallel.

The checks come from outside the page. A check list derived from the guide inherits the guide's blind spots, so every check passes and the score reports that the guide is fine when nobody looked. Build the list from the sources, then diff it against the page. Anything the sources treat as essential and the page omits is a candidate check, and it should fail a solution written from the page. That failure is the finding the paired docs ticket acts on.

Several sources need Supabase employee access. Settle which ones you can reach before you start, plan from those, and ignore the rest. Missing access is never a blocker. Record in the plan which sources the inventory came from and which were out of reach.

This phase ends with a failure-point inventory: one bullet per finding, each carrying an issue id, a thread, or a url. A finding with no source does not go in.

Phase 2: pick the one claim, and confirm it can be measured

Name the single thing the guide has to transmit, in the user's terms rather than the product's.

One claim, not three. An eval that measures several unrelated things reports which of them an agent got, and a docs ticket cannot act on that. It also means every check hangs off one subject, which is what makes a saturated check obvious later.

A dense page yields ten candidates, all evidenced. Pick on consequence: the failure that costs a user the most, and among equals the one the page itself warns about. A page that carries a danger callout has already named its own worst case.

Then confirm the harness can measure it, before designing a single check. A claim the harness cannot observe is not the claim, however well evidenced. Three questions, and read the code for the answers rather than assuming:

  • Does the end state exist somewhere a scorer can reach? A file in the workspace, a row in the database, a response from a running process. If the correct answer is something the platform supplies at deploy time, decide now where the seed lets it land.
  • Can the scorer run whatever has to run? A framework the eval seeds is installed and built inside the sandbox with ctx.exec, not on the host: the host-side build helpers relink the workspace's node_modules to the framework's, which does not carry an eval's own dependencies. Check whether any existing eval runs the same toolchain, and treat a first as a risk to state in the plan.
  • What is the time budget? The agent's timeout does not bound the scorer, so a long install and build costs wall-clock rather than the agent's budget. Per-command limits are yours to set.

Write the answers into the plan. A feasibility problem found here is a design change; found in Phase 4 it is a rewrite.

Show full SKILL.md (791 more words)Show less

Phase 3: draft the prompt and the seed

Work through references/prompt-rules.md.

The prompt is a product request plus the page's url, written the way the user would write it, with the vocabulary that gives away the answer removed.

The seed is where the detail goes. Write seed comments in product vocabulary, never the vocabulary the prompt strips: the endpoint the rest of the team builds against, the shapes a handler must return, which table the API serves. Pre-solve what belongs to another scenario and say so in the comment, so a mistake there cannot fail this eval for the wrong reason. Choose fixture values that a memorized answer gets wrong, so a whitelist check measures something. Seed a conflict where the subject allows one, so a single approach cannot satisfy every case.

Where the task contract lives is a real choice. Put it in seed comments and gain a positive control the scorer can prove, at the cost of a discovery question. Leave it out and measure whether the agent derives it, which is harder and less observable. Decide deliberately and record which in the README.md.

This phase ends with the prompt body, the list of stripped words, the seed, and the motivation: frontmatter.

Phase 4: design the checks

Work through references/check-rules.md, in order. The first three rules matter most.

Use a judge only when the artifact class is unbounded, meaning free prose or files whose shape you cannot predict. Use a deterministic check for anything a query or a file read settles. Scope the rubric to what the check is worth, state what not to grade, give the judge a tie-break, and make one claim per check. An eval passes only when every check passes, so an ambitious rubric on a secondary check fails the whole page.

This phase ends with a table: the check name, what it proves, what it does when the object is absent, whether it reads files or runs code, and the evidence it came from.

Then name the checks you expect to saturate, and write the prediction down. A check whose answer the seed labels costs an agent nothing and carries no signal. A plan that says which checks are cheap is honest about how much the eval measures.

Phase 5: write the plan

Sections: context, failure-point inventory, the seed, the prompt, the check table, what is not scored, the expected-failure table, verification, the paired docs ticket, and Phase 6 as the closing step.

Everything in the prompt is either measured or named as out of scope. A reviewer finds the rule that is neither.

Verification means example solutions: one you believe is correct, plus a few carrying a single deliberate flaw each. Write down which checks you expect each to fail before running anything; that list is the test. A flaw usually trips several checks, so do not aim for exactly one failure each. Never commit them. evals/*/*/solutions/ is a git exclusion, and you add it to .git/info/exclude before the first git add of an eval directory, because that file is local to your clone. A transcript-anchored check cannot be exercised by a solution, because no agent ran.

Phase 6: feed the catalog

The plan's last step is to append whatever broke to references/flakiness.md once the baseline lands, with the PR number.

This is a plan step rather than advice, because the file is only worth having if it grows. Ask reviewers to write findings into it directly rather than leaving them in a review thread, where the next author does not look.

Content shape against end-to-end proof

Two kinds of check, and a plan needs both.

A check that reads files asserts content shape. It is cheap, stable, and has no opinion about whether anything runs. Every one of them passes for an agent that edited a config file and stopped.

A check that runs the code asserts the outcome. It is what stops a decorative pass, and it is slower and more exposed to the environment.

The table says which each check is, because the two fail for different reasons and a reviewer reading a failure needs to know which kind they are looking at.

Where the boundary sits

Some claims are not this skill's to make. Whether the page carries runnable commands in fenced blocks, declares its environment variables and prerequisites, and ends with a verification step is a claim about the page's own markdown. That belongs to the page and to the project's authoring rubric. No tooling in this repo asserts it; the docs end-to-end suite is rendering, links, and accessibility.

What belongs here is the claim about what an agent produces after reading the page. Mixing the two produces an eval that fails when someone reformats a code block.

© supabase, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in evals/docs/.claude/skills/docs-eval-planning of supabase/evals.

  • SKILL.md
  • references/check-rules.md
  • references/flakiness.md
  • references/prompt-rules.md
  • references/sources.md

Open the folder on GitHubat commit 6d8b6bd

Compare with similar skills

Docs Eval Planning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Docs Eval Planning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Docs Eval Planning this skillsupabase/evals143—~2.9kAutomated safety check: PassApache-2.0
Evals Contextzgsm-ai/costrict4.4k1 repos~1.9kAutomated safety check: PassApache-2.0
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
AI Project Copilotsun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT
Eee Dataset Conversionevaleval/every_eval_ever133—~2.5kAutomated safety check: PassMIT
Analyze Id Eval Rankingopen-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Evals Context

    zgsm-ai/costrict

    Provides context about the CoStrict evals system structure in this monorepo.

    4.4k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    100 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Eee Dataset Conversion

    evaleval/every_eval_ever

    Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

    133 GitHub stars~2.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 8 days ago
    Research & ScienceAuto-check passed
  • Add Evaluator

    wso2/agent-manager

    Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

    105 GitHub stars~710 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from supabase/evals

  • Shadcn

    supabase/evals

    Official

    Manages shadcn components and projects — adding, searching, fixing, debugging, styling, and composing UI.

    143 GitHub starsUsed in 42 repos~4.5k tokens
    Auto-check passed

Works with

Questions about Docs Eval Planning

What does Docs Eval Planning do?

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. Docs Eval Planning is an agent skill from supabase/evals, published by the product's own GitHub organization. Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

When should I use Docs Eval Planning?

Docs Eval Planning fits situations like: design an eval for a Supabase docs guide; A ticket asks for a deterministic eval on a page; deciding what a guide-under-test eval should check.

How do I install Docs Eval Planning in Claude Code?

Run `npx skills add supabase/evals --skill docs-eval-planning -a claude-code`. Or copy the skill folder (evals/docs/.claude/skills/docs-eval-planning in supabase/evals) into .claude/skills/docs-eval-planning in your project. Claude Code loads it when a task matches its description.

How do I install Docs Eval Planning in Codex?

Run `npx skills add supabase/evals --skill docs-eval-planning -a codex`. Or copy the skill folder (evals/docs/.claude/skills/docs-eval-planning in supabase/evals) into .agents/skills/docs-eval-planning in your project. Codex loads it when a task matches its description.

Can I use Docs Eval Planning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add supabase/evals --skill docs-eval-planning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/docs-eval-planning, .gemini/skills/docs-eval-planning, .github/skills/docs-eval-planning and .opencode/skills/docs-eval-planning in your project.

What does Docs Eval Planning need to run?

Going by SKILL.md and its folder, Docs Eval Planning needs the command-line tools its instructions call (git).

Does Docs Eval Planning access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Docs Eval Planning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Docs Eval Planning use?

Docs Eval Planning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Docs Eval Planning use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.5k tokens, read only when the agent opens those files.

What are the alternatives to Docs Eval Planning?

Skills that share tags, products or a category with Docs Eval Planning: Evals Context (zgsm-ai/costrict, 4.4k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 100 stars) and Eee Dataset Conversion (evaleval/every_eval_ever, 133 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Docs Eval Planning?

supabase (a GitHub organization, an official publisher) maintains it in supabase/evals, which has 143 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.

Source: supabase/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.