Agent Benchmark Suite
ruvnet/ruflo
Agent skill for benchmark-suite - invoke with $agent-benchmark-suite
Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…
$ npx skills add sanity-io/sanity --skill sanity-bench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sanity-io/sanity sanity-bench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sanity-bench .claude/skills/sanity-bench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .claude/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-benchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sanity-io/sanity --skill sanity-bench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sanity-io/sanity sanity-bench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/sanity-bench .agents/skills/sanity-bench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .agents/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sanity-io/sanity --skill sanity-bench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sanity-io/sanity sanity-bench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/sanity-bench .cursor/skills/sanity-bench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .cursor/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sanity-io/sanity.git --path .agents/skills/sanity-bench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sanity-io/sanity --skill sanity-bench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sanity-io/sanity sanity-bench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/sanity-bench .gemini/skills/sanity-bench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .gemini/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sanity-io/sanity sanity-benchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sanity-io/sanity --skill sanity-bench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/sanity-bench .github/skills/sanity-bench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .github/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sanity-io/sanity --skill sanity-bench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sanity-io/sanity sanity-bench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sanity-io/sanity.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/sanity-bench .opencode/skills/sanity-bench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sanity-bench" agent skill from https://github.com/sanity-io/sanity/tree/main/.agents/skills/sanity-bench into .opencode/skills/sanity-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sanity-bench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sanity-benchRun, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…
Sanity Bench is an agent skill from sanity-io/sanity, published by the product's own GitHub organization. Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite, selftest, backfillsha, releasetag, abfrom/abto) and which combine, and how to read 🔴🟢✅⚪ verdicts. Use when asked to benchmark a PR or commit, compare two commits, backfill or re-run the daily main series, measure a release, add a bench scenario, extend the mock API, or interpret a bench PR comment or a failed Studio…
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Sanity Studio – Rapidly configure content workspaces powered by structured content. The licence is MIT.
Read from SKILL.md and the folder at commit efa15fb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pnpmghnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pnpm and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sanity Bench loads about 2.1k tokens when it runs. Until then it costs about 136 tokens; SKILL.md has 616 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sanity-io/sanity at commit efa15fb, republished under its MIT licence (© sanity-io). 616 words, ~2,135 tokens.
.claude/skills/sanity-bench/SKILL.md (or your agent's skills folder).A built studio measured in Chromium against an in-process mock of the Sanity API: no tokens, no
network, one clock. The interesting output is an A/B verdict per metric from interleaved
reference/experiment sessions with a cluster bootstrap — 🔴 regression, 🟢 improvement, ✅ neutral,
⚪ inconclusive (CI too wide within budget; never a coin flip). Absolute numbers are host-relative
and only comparable with the run's calibration score alongside. perf/bench/README.md is the
full reference; this skill is the operating card. Stored results are read through sanity-radar.
pnpm build:bench # required first: packages + bench studio
pnpm bench help # commands; `pnpm bench run --help` for flags
pnpm bench scenarios # singleString, arrayI18n, article, recipe, synthetic, syntheticLarge, loginReady, loginToTool, toolReady, emptyToolReady, loginToEmptyTool (+ customization scenarios)
pnpm bench run --scenario singleString --sessions 4 # absolute interaction run
pnpm bench run --mode pageload --scenario singleString # load vitals + bundle size
pnpm bench run --mode inp --scenario singleString --sessions 3
pnpm bench run --mode soak --scenario singleString --minutes 5
pnpm --filter bench build:reference-config \
&& pnpm bench run --scenario singleString --reference-dist perf/bench/.reference/dist # self-test: must be all-neutral
pnpm bench prepare-backfill --sha <sha> # build a historical commit into perf/bench/dist (tarball recipe)
pnpm bench dev # mock + sanity dev, type into a seeded doc
pnpm bench:unit # mock contract + stats testsFlags worth knowing: --headed, --throttle 1 (interaction runs default to 4× CPU throttle),
--seed N, --budget <s>, --json-out <file>, --fail-on-verdict. In a sandbox where the tsx
CLI cannot open its IPC pipe (listen EPERM … tsx-*/…pipe), run the CLI directly:
node --conditions=monorepo --import tsx perf/bench/cli/index.ts <command>.
.github/workflows/bench.yml| Trigger | What runs | Result lands in |
|---|---|---|
Label a PR trigger:perf-bench (maintainers) | PR head vs merge-base A/B, one shard per scenario. The label persists: every later push re-runs the suite until the label is removed. | Sticky PR comment bench-report; absolute side stored under the branch (?branches=main,<branch> in Radar) |
| Cron 05:00 UTC daily | Absolute suite on main + soak + INP | benchRun mode: "absolute" in Radar (Trends) |
| Cron Monday 06:00 UTC | Self-test (build vs itself, --fail-on-verdict) | Red job = harness drift |
-f run_suite=true | Same as the daily cron, on demand | Radar |
-f run_suite=true -f backfill_sha=<40-char sha> | Replay a historical main commit, store under that sha and its commit date | Radar (fills a hole in the series) |
--ref v6.x.y -f run_suite=true -f release_tag=v6.x.y | Measure a release commit (release-latest.yml does this automatically) | Radar, trigger: "release" → anchors the release marker |
-f ab_from=<sha> -f ab_to=<sha> | A/B of two arbitrary commits (reference → experiment) | Run summary page + benchRun mode: "ab" (Radar Comparisons tool) |
-f self_test=true | Self-test on demand | Job status |
gh workflow run bench.yml -R sanity-io/sanity -f ab_from=<sha> -f ab_to=<sha>
gh run list -R sanity-io/sanity --workflow bench.yml --event workflow_dispatch --limit 5
gh run watch -R sanity-io/sanity <run-id>; gh run view -R sanity-io/sanity <run-id> --webInput rules (the workflow validates them): ab_from/ab_to go together and exclude
backfill_sha and self_test; backfill_sha requires run_suite and excludes self_test;
release_tag must point at the measured commit (dispatch at the tag, or pass the tag's sha as
backfill_sha from main). Shas are full 40 characters. Historical builds (backfill, A/B, PR
reference) install the old commit with --frozen-lockfile, pack sanity and its workspace deps
as tarballs, and build HEAD's perf/bench harness against them — harness and scenarios are
always HEAD's; only product code differs. Backfill and A/B fail loudly, the PR reference falls
back to absolute mode with a warning in the comment.
Radar's run popover copies the A/B command for "this point vs the previous run" (Copy A/B vs previous run) so the shas never need typing.
missing-scenarios.json fails the run) — silence is never "fine".absMs 16ms floor (two Event Timing quantisation steps), pageload time-to-editable; INP, vitals, resources, bundle are report-only. Thresholds live in perf/bench/stats/gate.ts and are shared with Radar's drift feed.stoppedBy and failures in the stored run.hermeticity violation or unexpected endpoint session failure means the studio made a request the mock does not know: extend perf/bench/mock-api in the same PR as the studio change.perf/bench/studio/schemas, defineScenario in perf/bench/scenarios (seeded PRNG, never Math.random), register in scenarios/index.ts, add it to the bench-interaction matrix and BENCH_EXPECTED_INTERACTION_SCENARIOS in bench.yml, verify one absolute session passes.perf/bench/report/types.ts → storeShape.ts, mirrored by dev/radar/schemaTypes/benchRun.ts; bump all three together.© sanity-io, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/sanity-bench of sanity-io/sanity.
Open the folder on GitHubat commit efa15fb
Sanity Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sanity Bench this skillsanity-io/sanity | 6.4k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Agent Benchmark Suiteruvnet/ruflo | 74k | 2 repos | ~4.9k | Automated safety check: Pass | MIT | |
| Benchmarking Kubernetes With Kube Benchmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | |
| Performing Kubernetes Cis Benchmark With Kube Benchmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~1.7k | Automated safety check: Notes | Apache-2.0 | |
| Benchmarkaffaan-m/ECC | 276k | 3 repos | ~654 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | — | ~412 | Automated safety check: Pass | MIT |
ruvnet/ruflo
Agent skill for benchmark-suite - invoke with $agent-benchmark-suite
mukul975/Anthropic-Cybersecurity-Skills
Installs and runs the kube-bench tool against a Kubernetes cluster as a Job, DaemonSet, or standalone binary, selecting the correct benchmark version and targets (control plane, etcd, kubelet…
mukul975/Anthropic-Cybersecurity-Skills
Turns kube-bench output into a finished CIS Kubernetes Benchmark audit: interpreting PASS/FAIL/WARN per control, judging which failures are genuine on a managed cluster, writing remediation, and…
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
affaan-m/ECC
このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.
affaan-m/ECC
使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。
sanity-io/sanity
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
sanity-io/sanity
Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities.
sanity-io/sanity
React and Next.js performance optimization guidelines from Vercel Engineering.
sanity-io/sanity
Add existing screenshots or screen recordings to a GitHub pull request as a before/after or preview block.
sanity-io/sanity
React DevTools CLI for AI agents. An agent skill from sanity-io/sanity.
sanity-io/sanity
Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.
Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite…. Sanity Bench is an agent skill from sanity-io/sanity, published by the product's own GitHub organization. Run, dispatch and read the hermetic Studio performance benchmark suite in perf/bench — local runs and modes, the trigger:perf-bench PR label, the Studio Bench workflowdispatch inputs (runsuite, selftest, backfillsha, releasetag, abfrom/abto) and which combine, and how to read 🔴🟢✅⚪ verdicts.
Sanity Bench fits situations like: asked to benchmark a PR; compare two commits; re-run the daily main series; measure a release.
Run `npx skills add sanity-io/sanity --skill sanity-bench -a claude-code`. Or copy the skill folder (.agents/skills/sanity-bench in sanity-io/sanity) into .claude/skills/sanity-bench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sanity-io/sanity --skill sanity-bench -a codex`. Or copy the skill folder (.agents/skills/sanity-bench in sanity-io/sanity) into .agents/skills/sanity-bench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sanity-io/sanity --skill sanity-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sanity-bench, .gemini/skills/sanity-bench, .github/skills/sanity-bench and .opencode/skills/sanity-bench in your project.
Going by SKILL.md and its folder, Sanity Bench needs the command-line tools its instructions call (pnpm, gh and node).
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Sanity Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Sanity Bench: Agent Benchmark Suite (ruvnet/ruflo, 74k stars), Benchmarking Kubernetes With Kube Bench (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Performing Kubernetes Cis Benchmark With Kube Bench (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Benchmark (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sanity-io (a GitHub organization, an official publisher) maintains it in sanity-io/sanity, which has 6,352 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 9, 2026.
Source: sanity-io/sanity on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.