Official agent skill

Sanity Bench

by sanity-io in sanity-io/sanity

Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…

OfficialMITAuto-check passed

Install Sanity Bench

skills CLI
$ npx skills add sanity-io/sanity --skill sanity-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sanity-io/sanity sanity-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sanity-bench .claude/skills/sanity-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sanity-bench
GitHub stars
6.4k
Token cost
~2.1k tokens
SKILL.md length
616 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…

  • Asked to benchmark a PR
  • SKILL.md covers Local, CI: .github/workflows/bench.yml, Reading a result and Changing the suite
  • Calls pnpm, gh and node
  • Compare two commits

What it does

Sanity Bench is an agent skill from sanity-io/sanity, published by the product's own GitHub organization. Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite, selftest, backfillsha, releasetag, abfrom/abto) and which combine, and how to read 🔴🟢✅⚪ verdicts. Use when asked to benchmark a PR or commit, compare two commits, backfill or re-run the daily main series, measure a release, add a bench scenario, extend the mock API, or interpret a bench PR comment or a failed Studio…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Sanity Studio – Rapidly configure content workspaces powered by structured content. The licence is MIT.

When your agent uses it

  • Asked to benchmark a PR
  • Compare two commits
  • Re-run the daily main series
  • Measure a release

Example prompts

  • “/sanity-bench”

What it can do on your machine

Read from SKILL.md and the folder at commit efa15fb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • gh
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sanity Bench loads about 2.1k tokens when it runs. Until then it costs about 136 tokens; SKILL.md has 616 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sanity-io/sanity at commit efa15fb, republished under its MIT licence (© sanity-io). 616 words, ~2,135 tokens.

Download SKILL.mdSave it as .claude/skills/sanity-bench/SKILL.md (or your agent's skills folder).
name
sanity-bench
description
Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflow_dispatch inputs (run_suite, self_test, backfill_sha, release_tag, ab_from/ab_to) and which combine, and how to read 🔴🟢✅⚪ verdicts. Use when asked to benchmark a PR or commit, compare two commits, backfill or re-run the daily main series, measure a release, add a bench scenario, extend the mock API, or interpret a bench PR comment or a failed Studio Bench run.

Studio bench (perf/bench)

A built studio measured in Chromium against an in-process mock of the Sanity API: no tokens, no network, one clock. The interesting output is an A/B verdict per metric from interleaved reference/experiment sessions with a cluster bootstrap — 🔴 regression, 🟢 improvement, ✅ neutral, ⚪ inconclusive (CI too wide within budget; never a coin flip). Absolute numbers are host-relative and only comparable with the run's calibration score alongside. perf/bench/README.md is the full reference; this skill is the operating card. Stored results are read through sanity-radar.

Local

bash
pnpm build:bench                                              # required first: packages + bench studio
pnpm bench help                                               # commands; `pnpm bench run --help` for flags
pnpm bench scenarios                                          # singleString, arrayI18n, article, recipe, synthetic, syntheticLarge, loginReady, loginToTool, toolReady, emptyToolReady, loginToEmptyTool (+ customization scenarios)
pnpm bench run --scenario singleString --sessions 4           # absolute interaction run
pnpm bench run --mode pageload --scenario singleString        # load vitals + bundle size
pnpm bench run --mode inp --scenario singleString --sessions 3
pnpm bench run --mode soak --scenario singleString --minutes 5
pnpm --filter bench build:reference-config \
  && pnpm bench run --scenario singleString --reference-dist perf/bench/.reference/dist   # self-test: must be all-neutral
pnpm bench prepare-backfill --sha <sha>                       # build a historical commit into perf/bench/dist (tarball recipe)
pnpm bench dev                                                # mock + sanity dev, type into a seeded doc
pnpm bench:unit                                               # mock contract + stats tests

Flags worth knowing: --headed, --throttle 1 (interaction runs default to 4× CPU throttle), --seed N, --budget <s>, --json-out <file>, --fail-on-verdict. In a sandbox where the tsx CLI cannot open its IPC pipe (listen EPERM … tsx-*/…pipe), run the CLI directly: node --conditions=monorepo --import tsx perf/bench/cli/index.ts <command>.

CI: .github/workflows/bench.yml

TriggerWhat runsResult lands in
Label a PR trigger:perf-bench (maintainers)PR head vs merge-base A/B, one shard per scenario. The label persists: every later push re-runs the suite until the label is removed.Sticky PR comment bench-report; absolute side stored under the branch (?branches=main,<branch> in Radar)
Cron 05:00 UTC dailyAbsolute suite on main + soak + INPbenchRun mode: "absolute" in Radar (Trends)
Cron Monday 06:00 UTCSelf-test (build vs itself, --fail-on-verdict)Red job = harness drift
-f run_suite=trueSame as the daily cron, on demandRadar
-f run_suite=true -f backfill_sha=<40-char sha>Replay a historical main commit, store under that sha and its commit dateRadar (fills a hole in the series)
--ref v6.x.y -f run_suite=true -f release_tag=v6.x.yMeasure a release commit (release-latest.yml does this automatically)Radar, trigger: "release" → anchors the release marker
-f ab_from=<sha> -f ab_to=<sha>A/B of two arbitrary commits (reference → experiment)Run summary page + benchRun mode: "ab" (Radar Comparisons tool)
-f self_test=trueSelf-test on demandJob status
bash
gh workflow run bench.yml -R sanity-io/sanity -f ab_from=<sha> -f ab_to=<sha>
gh run list -R sanity-io/sanity --workflow bench.yml --event workflow_dispatch --limit 5
gh run watch -R sanity-io/sanity <run-id>; gh run view -R sanity-io/sanity <run-id> --web

Input rules (the workflow validates them): ab_from/ab_to go together and exclude backfill_sha and self_test; backfill_sha requires run_suite and excludes self_test; release_tag must point at the measured commit (dispatch at the tag, or pass the tag's sha as backfill_sha from main). Shas are full 40 characters. Historical builds (backfill, A/B, PR reference) install the old commit with --frozen-lockfile, pack sanity and its workspace deps as tarballs, and build HEAD's perf/bench harness against them — harness and scenarios are always HEAD's; only product code differs. Backfill and A/B fail loudly, the PR reference falls back to absolute mode with a warning in the comment.

Radar's run popover copies the A/B command for "this point vs the previous run" (Copy A/B vs previous run) so the shas never need typing.

Show full SKILL.md (200 more words)Show less

Reading a result

  • The PR comment / run summary is deliberately minimal: verdicts, then a link to Radar. A missing scenario is named (missing-scenarios.json fails the run) — silence is never "fine".
  • Gates: interaction median with absMs 16ms floor (two Event Timing quantisation steps), pageload time-to-editable; INP, vitals, resources, bundle are report-only. Thresholds live in perf/bench/stats/gate.ts and are shared with Radar's drift feed.
  • ⚪ on a heavy scenario usually means the budget ran out; re-dispatch or look at stoppedBy and failures in the stored run.
  • A hermeticity violation or unexpected endpoint session failure means the studio made a request the mock does not know: extend perf/bench/mock-api in the same PR as the studio change.

Changing the suite

  • New scenario: schema in perf/bench/studio/schemas, defineScenario in perf/bench/scenarios (seeded PRNG, never Math.random), register in scenarios/index.ts, add it to the bench-interaction matrix and BENCH_EXPECTED_INTERACTION_SCENARIOS in bench.yml, verify one absolute session passes.
  • Stored document shape: perf/bench/report/types.ts → storeShape.ts, mirrored by dev/radar/schemaTypes/benchRun.ts; bump all three together.
  • Flake rules are in the README (no fixed sleeps, noise widens sampling, seeded PRNG, failed sessions retried and counted, environment drift fails fast). A change that makes a verdict flip on identical builds is a harness bug — the weekly self-test exists to catch it.

© sanity-io, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/sanity-bench of sanity-io/sanity.

Open the folder on GitHubat commit efa15fb

Compare with similar skills

Sanity Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sanity Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sanity Bench this skillsanity-io/sanity6.4k—~2.1kAutomated safety check: PassMIT
Agent Benchmark Suiteruvnet/ruflo74k2 repos~4.9kAutomated safety check: PassMIT
Benchmarking Kubernetes With Kube Benchmukul975/Anthropic-Cybersecurity-Skills34k—~2.3kAutomated safety check: NotesApache-2.0
Performing Kubernetes Cis Benchmark With Kube Benchmukul975/Anthropic-Cybersecurity-Skills34k—~1.7kAutomated safety check: NotesApache-2.0
Benchmarkaffaan-m/ECC276k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~412Automated safety check: PassMIT

Similar skills

  • Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

    74k GitHub starsUsed in 2 repos~4.9k tokens
    Agent WorkflowsAuto-check passed
  • Benchmarking Kubernetes With Kube Bench

    mukul975/Anthropic-Cybersecurity-Skills

    Installs and runs the kube-bench tool against a Kubernetes cluster as a Job, DaemonSet, or standalone binary, selecting the correct benchmark version and targets (control plane, etcd, kubelet…

    34k GitHub stars~2.3k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Performing Kubernetes Cis Benchmark With Kube Bench

    mukul975/Anthropic-Cybersecurity-Skills

    Turns kube-bench output into a finished CIS Kubernetes Benchmark audit: interpreting PASS/FAIL/WARN per control, judging which failures are genuine on a managed cluster, writing remediation, and…

    34k GitHub stars~1.7k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    276k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    276k GitHub stars~412 tokensUpdated yesterday
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    276k GitHub stars~330 tokensUpdated yesterday
    Auto-check passed

More from sanity-io/sanity

All 34 skills in this repo
  • Playwright CLI

    sanity-io/sanity

    Official

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    6.4k GitHub starsUsed in 18 repos~1.9k tokens
    Auto-check passed
  • Find Skills

    sanity-io/sanity

    Official

    Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities.

    6.4k GitHub starsUsed in 67 repos~1.2k tokens
    Auto-check passed
  • Official

    React and Next.js performance optimization guidelines from Vercel Engineering.

    6.4k GitHub starsUsed in 129 repos~1.6k tokens
    Auto-check passed
  • Before And After

    sanity-io/sanity

    Official

    Add existing screenshots or screen recordings to a GitHub pull request as a before/after or preview block.

    6.4k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • React Devtools

    sanity-io/sanity

    Official

    React DevTools CLI for AI agents. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Auto-check passed

Questions about Sanity Bench

What does Sanity Bench do?

Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…. Sanity Bench is an agent skill from sanity-io/sanity, published by the product's own GitHub organization. Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite, selftest, backfillsha, releasetag, abfrom/abto) and which combine, and how to read 🔴🟢✅⚪ verdicts.

When should I use Sanity Bench?

Sanity Bench fits situations like: asked to benchmark a PR; compare two commits; re-run the daily main series; measure a release.

How do I install Sanity Bench in Claude Code?

Run `npx skills add sanity-io/sanity --skill sanity-bench -a claude-code`. Or copy the skill folder (.agents/skills/sanity-bench in sanity-io/sanity) into .claude/skills/sanity-bench in your project. Claude Code loads it when a task matches its description.

How do I install Sanity Bench in Codex?

Run `npx skills add sanity-io/sanity --skill sanity-bench -a codex`. Or copy the skill folder (.agents/skills/sanity-bench in sanity-io/sanity) into .agents/skills/sanity-bench in your project. Codex loads it when a task matches its description.

Can I use Sanity Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sanity-io/sanity --skill sanity-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sanity-bench, .gemini/skills/sanity-bench, .github/skills/sanity-bench and .opencode/skills/sanity-bench in your project.

What does Sanity Bench need to run?

Going by SKILL.md and its folder, Sanity Bench needs the command-line tools its instructions call (pnpm, gh and node).

Does Sanity Bench access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sanity Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sanity Bench use?

Sanity Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sanity Bench use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sanity Bench?

Skills that share tags, products or a category with Sanity Bench: Agent Benchmark Suite (ruvnet/ruflo, 74k stars), Benchmarking Kubernetes With Kube Bench (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Performing Kubernetes Cis Benchmark With Kube Bench (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Benchmark (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sanity Bench?

sanity-io (a GitHub organization, an official publisher) maintains it in sanity-io/sanity, which has 6,352 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 9, 2026.

Source: sanity-io/sanity on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.