Official agent skill

Sandbox Bench

by vercel in vercel/next.js

Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

OfficialMITAuto-check passedData & Analytics

Install Sandbox Bench

skills CLI
$ npx skills add vercel/next.js --skill sandbox-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vercel/next.js sandbox-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vercel/next.js.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sandbox-bench .claude/skills/sandbox-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sandbox-bench
GitHub stars
143k
Token cost
~4.1k tokens
SKILL.md length
2,202 words
Files
13 (incl. scripts, references)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

  • Works in 5 steps: Resolve what's being compared → Gate correctness before spending bench… → Launch the bench (background,… → …
  • The user asks to bench
  • SKILL.md covers One-time setup, Workflow, Results database and Keeping the user informed, plus 2 more sections
  • Runs JavaScript scripts from its folder; calls node, vercel and bash

What it does

Sandbox Bench is an agent skill from vercel/next.js, published by the product's own GitHub organization. Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app (rps, latency, p95; TTFB, RSS and document/Flight bytes when the Next side captures them) and, for React changes, through the react repo's flight-ssr-bench fixture (Node AND Edge web-streams paths, Fizz and Flight+Fizz). Use whenever the user asks to bench, perf test, or A/B a React PR, a react-server-dom / Flight /…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `references/methodology.md`).

It sits in Data & Analytics, covering Statistics and Backend development. It works with Next.js, Vercel and React. The licence is MIT.

When your agent uses it

  • The user asks to bench
  • A react-server-dom / Flight / vendored React change
  • A Next.js PR (is this PR faster
  • Does this regress RSC?

Example prompts

  • “is this PR faster”
  • “does this regress RSC?”
  • “measure the perf impact of <commit”
  • “/sandbox-bench”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Resolve what's being compared
  2. Gate correctness before spending bench compute
  3. Launch the bench (background, non-blocking)
  4. Read the result like a skeptical data scientist
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit a32ddfd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • vercel
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use vercel, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sandbox Bench loads about 4.1k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 227 tokens; SKILL.md has 2,202 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~227
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from vercel/next.js at commit a32ddfd, republished under its MIT licence (© vercel). 2,202 words, ~4,129 tokens.

Download SKILL.mdSave it as .claude/skills/sandbox-bench/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
sandbox-bench
description
Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app (rps, latency, p95; TTFB, RSS and document/Flight bytes when the Next side captures them) and, for React changes, through the react repo's flight-ssr-bench fixture (Node AND Edge web-streams paths, Fizz and Flight+Fizz). Use whenever the user asks to bench, perf test, or A/B a React PR, a react-server-dom / Flight / vendored React change, or a Next.js PR ("is this PR faster", "does this regress RSC?", "measure the perf impact of <commit>"), even if they don't say "benchmark" — any request to quantify a server-side performance difference between two revisions belongs here. Runs remotely (laptop-free), applies correctness gates before measuring, and reports boot-level confidence intervals.
metadata.internal
true

Sandbox bench: paired A/B perf runs for React and Next.js changes

Measures what a change is actually worth, end to end: two revisions ("arms") built into otherwise-identical Next.js apps, exercised by the bench/render-pipeline harness on Vercel Sandbox VMs, compared with paired statistics that treat the VM boot as the unit of replication. All heavy work happens on sandbox VMs; the laptop only orchestrates.

Scripts live in scripts/ next to this file and run from anywhere. Arms are git refs, resolved in cached clones of react and next.js; the Next side defaults to canary. Everything is cached content-addressed: first use of a new pair builds caches (~45-60 min extra, once); later runs boot straight into measurement.

One-time setup

  1. node scripts/config.mjs show — if it reports NOT CONFIGURED, ask the user which Vercel team and project the sandbox VMs should run under (these are billed resources; never guess, never default), then node scripts/config.mjs set team=<slug> project=<name>. Config lives in ~/.config/sandbox-bench/config.json — never commit team/project names into the repo.
  2. The Vercel CLI session must have access to that team. On a 403, stop launching (don't retry through it) and check whether access is already back: vercel whoami --scope <team-slug> plus one scoped read call (e.g. vercel sandbox ls) — grants drop and recover on their own, and a transient 403 needs no login at all. If verification still fails, run vercel login <team-slug> yourself as a background task (the token lives with the CLI session, not with the user). It opens a browser/device confirmation — relay the URL if one is printed — but keep re-running the verification pair every minute or two while it waits: access often returns before the login flow reports success, and once verification passes, kill the pending login and resume. After a 403 outage, expect in-flight runs to have died: run node scripts/bench-status.mjs and follow its recovery actions (measurement VMs will have hit their ~5h timeout if the outage was long — those cells need relaunching, not collecting).
  3. react and next.js clones land in the cache on first use (or point reactRepo/nextRepo in the config at existing checkouts).

Before the first real run with a new configuration, sanity-check the plan with --dry-run (prints what would happen, touches nothing).

Workflow

1. Resolve what's being compared
  • React PR: --pr <url|number> — base is computed automatically (merge-base of the PR head with react main).
  • React refs: --arms base=<ref>,cand=<ref> — base FIRST. For a multi-commit branch, base is the merge-base with main, not cand^.
  • Next.js PR: --next-pr <url|number>. The React side defaults to whatever each Next ref vendors (that's what would ship); pass --react-ref only to pin both arms to one specific React build.
  • Next refs: --next-arms base=<ref>,cand=<ref>.

Exactly one side varies; the other is identical in both arms. That isolation is what makes the numbers attributable — never vary both.

2. Gate correctness before spending bench compute

A bench number from an arm that fails its own tests is meaningless. For any arm that is not already CI-green upstream (hand-assembled branches, cherry-picks with resolved conflicts, local commits):

sh
node scripts/sandbox-gate.mjs --arms cand=<ref>

The bench itself enforces the primary gate: every react arm's commit must have green CI on the react repo, checked automatically before any build or VM is spent. PRs and main-history commits normally satisfy this with no extra work. For local or unpushed refs (no CI exists), gate on a VM with sandbox-gate.mjs and then pass --allow-ungated to the bench. The VM gate runs the full test suite in prod mode (the channel that gets benched). PASS requires seeing the actual test counts in the output. If a gate fails, report the failures and stop — do not bench a broken arm. Each arm is gated in its own lockfile's environment. Bench the exact sha the gate prints (a branch ref can move between gate and bench).

3. Launch the bench (background, non-blocking)
sh
bash -c 'node scripts/sandbox-e2e.mjs --pr <url> --label <slug> \
  2>&1 | grep --line-buffered -v "^live "; exit ${PIPESTATUS[0]}'

For React PRs, launch BOTH suites (separate background tasks; they share arm builds and caches):

sh
bash -c 'node scripts/sandbox-ssr.mjs --pr <url> --label <slug>-ssr \
  2>&1 | grep --line-buffered -v "^live "; exit ${PIPESTATUS[0]}'

The e2e suite measures the Node path through a real Next.js app; the ssr suite measures the react repo's flight-ssr-bench fixture — 8 variants (Fizz and Flight+Fizz, Node and Edge web streams, sync and async), each sequentially with Flight script injection and behind an HTTP server at c=1/c=10. Edge cells are the ssr suite's headline (the e2e suite cannot see that path); its Node and Fizz-only cells attribute an effect to the Flight layer, the Fizz layer, or the stream plumbing. The fixture (the workload) is pinned to one ref for both arms — react main by default — so only the React builds differ; if the PR itself edits the fixture, the launcher says so and the run does not measure those edits. Next PRs run the e2e suite only.

  • Run it as a background task and proceed on its completion notification. Never hold a foreground wait; never poll in a loop — the rule is about control flow, not status relay: reading the output tail to answer "how's it going" is always fine.
  • The harness handles the invariants internally: both arms in the same VM, interleaved ABBA, paired per (vm, run); detached remote execution (transport drops don't kill runs); build fingerprints recorded in every result row.
  • live ... lines are streaming estimates for progress display only. Never stop a run early because a live p-value looks good, and never report a live number — sequential peeking manufactures false positives. Only the final analysis counts.
  • Defaults (16 VMs × 2 paired runs) implement the methodology; don't reduce VM count to save time — boots are the unit of inference, and fewer boots means wider intervals, not faster answers.
  • Sandbox compute is internal capacity, not a budget: launch, relaunch, and confirm runs without asking about cost or shrinking them to save it.
  • The Next side's default, canary, is the latest published canary release (the launcher prints its version and sha), so repeat benches reuse the built snapshot until a new canary ships.
  • Useful flags: --bench-env KEY=VALUE (runtime-only env for the bench process — it does NOT affect the snapshot's app build), --isolate-routes (tail investigations), --no-profile (skip the CPU capture that runs by default after the timed runs), --prepare (build caches only — use when two cells will share an arm, to avoid duplicate builds racing).
  • CPU profiles are captured by default: one profile pass per arm runs strictly AFTER the timed runs (it cannot touch the numbers), costs ~45-60 min extra VM wall-clock, and lands in <runDir>/prof-vm<N>/ as standard V8 .cpuprofile files. Cross-VM profile diffs are highly stable (observed 16/16 sign agreement on real movers), so one profiled cell suffices to rank hot paths. Analysis caveats: aggregate by (functionName, line, column) — bare minified names collide across the bundle — and never diff arms by minified name (the minifier renames between builds); match positions or code snippets instead.
  • The bench exercises Next's node-streams path (__NEXT_USE_NODE_STREAMS is inlined as true for the node runtime at build time). React changes that only touch the EDGE stream configs are not exercised end-to-end and will (correctly) bench as no detected difference.
4. Read the result like a skeptical data scientist

The goal is the truth about the change, not making its author feel good. The final analysis prints, per route/phase/metric, the boot-level mean, ±95% CI, and p across boots. Apply the policy in references/methodology.md:

  • Claim only boot-level p < 0.01, with the CI, on an A/A-validated team/config (see methodology).
  • The PR is a hypothesis, not an explanation. Claims come from the analysis output alone. When the numbers agree with the PR's story, check whether the captured data actually discriminates that mechanism from alternatives — a latency win attributed to smaller payloads should come with a document-bytes delta; if the bytes didn't move, the story doesn't hold and the report says so.
  • Use every captured metric, and voice anything that does not add up: one metric family moving against the others, effects with no byte-level or RSS trace, throughput moving without latency, sign flips across boots. An inconsistency you cannot explain belongs in the report, not in the drawer.
  • The within-run p shown in brackets is a diagnostic, never a claim.
  • Check the fingerprint header first: two distinct fingerprints = valid A/B; "inconsistent fingerprints" = invalid, report no numbers. The fingerprint hashes both bundlers' compiled server files — arms touching only client files can still legitimately show identical fingerprints with different version strings.
  • Per-boot values are printed; if boots disagree in sign, say so.
  • Any claim that will drive a decision gets one independent confirmation run before it's stated as fact.

Re-analyze any past run without re-running it: node scripts/bench-analyze.mjs <runDir>.

Show full SKILL.md (796 more words)Show less
5. Report

Name what was measured with links: the PR title (printed in the analysis header, stored in meta.json) linking to the PR; for ref arms, the commit title. Lead with a table of the significant cells, each row carrying the effect with its unit, the CI, and p:

## [<PR title>](<PR url>) — e2e, Vercel Sandbox (x86 Xeon), <n> boots

Significant (boot-level p < 0.01, A/A-validated):
| cell | effect | 95% CI | p |
|---|---|---|---|
| /dashboard under load | +14.4% throughput (req/s) | ±3.2% | <0.0001 |
| /dashboard serial | −10.7% median latency (ms) | ±0.6% | <0.0001 |

No detected difference: <every cell not in the table, by name>.
Flags: <cells at 0.01 ≤ p < 0.05, sign disagreements across boots,
fingerprint caveats, anything that does not add up>

One row per cell: rps and median restate each other, so report the throughput number (add a p95 row only when the tail moves differently from the median). Document metrics (raw/gzip/Flight KB) get their own rows when they differ — they are the mechanism evidence. When the Next side predates the document-metrics harness (vercel/next.js#95828) those cells are absent; say so instead of silently reporting less. State the platform next to the numbers. Magnitudes are platform-dependent (GC share differs by CPU); direction and mechanism transfer, percentages do not. Never present a noise-compatible delta as a small win or loss — it is "no detected difference".

Results database

Every collected run lands in one SQLite file, ~/.cache/sandbox-bench/results.db — raw measurements and artifacts (CPU profiles, logs) only, written exclusively by the importer, never by hand. The launcher imports and verifies automatically at collection; bench-analyze reads the db and nothing else, so every statistic is a pure function of it. Numbers in reports come from the analysis output verbatim — never retype, recompute, or aggregate them yourself.

  • node scripts/bench-db.mjs ls — all runs with sample/artifact counts.
  • node scripts/bench-db.mjs verify [runId] — integrity checks: sqlite-level, referential, one fingerprint per arm, paired sample counts, artifact sha256. Run it before drawing on old data.
  • node scripts/bench-db.mjs export out.db <runId...> — cut a self-contained db of specific runs (with their profiles) to send to someone. It opens in any SQLite tool.
  • node scripts/bench-analyze.mjs <runId> — re-analyze anything in the db; a run-dir argument imports it first.

Keeping the user informed

The launcher narrates itself on stdout: launch facts first (run dir, arms, CI verdicts), then a progress line every ~2 minutes with rows collected and interim per-route effects with confidence. Relay to the user: the run dir and expected duration right after launching, notable interim shifts if they ask how it's going, and the full verdict from the final analysis when the completion notification arrives. The analysis names metrics that were not captured on this run — repeat that in the verdict when it limits what the data can say (document metrics absent means the payload mechanism is unverified, not verified-identical).

While a run is active, open any reply with a one-line status per run: read the tail of the launcher's output and quote its latest progress line. If the session supports timed wakeups or reminders, schedule a check at each expected transition (arm builds -> experiment snapshot -> measuring, then every ~15 minutes of measurement) and post the progress line; if not, say when the next update will arrive so silence is never ambiguous. Interim effects in progress lines are streaming estimates — share them as progress, never as claims.

If a launcher process dies (session teardown, crash), the remote VMs keep executing their measurement loops — the data is not lost. node scripts/bench-collect.mjs <runDir> reconnects, waits for the loops, downloads the results, cleans up, and analyzes. Run it before the VMs hit their ~5h timeout.

Failure recovery

  • First move, always: node scripts/bench-status.mjs. Session restarts silently kill background launchers while their detached VMs keep measuring, and a dead launcher's log still ends with a healthy-looking progress line — never infer liveness from log tails or task output files. bench-status checks each run's recorded launcher pid and prints the per-run recovery action (running / collect now / relaunch). Run it at the start of any session that expects work in flight, after any crash, and before telling the user what is or isn't running. Launcher crashes are also recorded in the run's status.json (phase: "failed" plus the error).
  • Interrupted local process: remote VMs keep running detached. vercel sandbox list (with the configured team/project) to find them; poll each VM's /vercel/sandbox/loop.done, cp its results.jsonl down when done, then remove the VM and analyze with bench-analyze.mjs.
  • Leaked VMs after any crash: node scripts/sandbox-sweep.mjs lists this skill's VMs (matched by sbench-* name AND the purpose=sandbox-bench tag, and only when older than --min-age-hours, default 3, so healthy in-flight runs are never touched); --yes removes them by exact listed name.
  • Flaky uploads/transports: the harnesses size-check artifacts and abort on truncation. A failed cell is safe to relaunch; caches make the retry cheap. Don't relaunch two cells that need the same uncached arm at the same moment — they'll race to build it; use --prepare first instead.

Cost expectations (set these with the user before big runs)

Per cell at defaults: ~18 VMs (8 measurement + build/snapshot VMs), ~1-2h wall-clock cold, ~30-60 min warm. A/A calibration and confirmation runs are extra cells. VMs are billed to the configured team — for anything beyond a single PR check, confirm scope first.

© vercel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in .agents/skills/sandbox-bench of vercel/next.js.

  • SKILL.md
  • references/methodology.md
  • scripts/bench-analyze.mjs
  • scripts/bench-collect.mjs
  • scripts/bench-common.mjs
  • scripts/bench-db.mjs
  • scripts/bench-stats.mjs
  • scripts/bench-status.mjs
  • scripts/config.mjs
  • scripts/sandbox-e2e.mjs
  • scripts/sandbox-gate.mjs
  • scripts/sandbox-ssr.mjs
  • scripts/sandbox-sweep.mjs

Open the folder on GitHubat commit a32ddfd

Compare with similar skills

Sandbox Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sandbox Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sandbox Bench this skillvercel/next.js143k—~4.1kAutomated safety check: PassMIT
React Performanceaffaan-m/ECC275k1 repos~4.5kAutomated safety check: PassMIT
AI SDKvercel-labs/ai-facts16821 repos~1.2kAutomated safety check: PassNone
Vercel React Best Practicessanity-io/sanity6.4k130 repos~1.6kAutomated safety check: PassMIT
Next Best Practicesvercel-labs/openreview1.7k18 repos~1kAutomated safety check: PassNone
Cosscrafter-station/petdex4.2k—~1.4kAutomated safety check: PassMIT

Similar skills

  • React Performance

    affaan-m/ECC

    React and Next.js performance optimization patterns adapted from Vercel Engineering's React Best Practices (https://github.com/vercel-labs/agent-skills).

    275k GitHub starsUsed in 1 repo~4.5k tokens
    DevelopmentAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 21 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    React and Next.js performance optimization guidelines from Vercel Engineering.

    6.4k GitHub starsUsed in 130 repos~1.6k tokens
    Frontend & DesignAuto-check passed
  • Next Best Practices

    vercel-labs/openreview

    Official

    Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling

    1.7k GitHub starsUsed in 18 repos~1k tokens
    DevelopmentAuto-check passed
  • Coss

    crafter-station/petdex

    Helps implement coss UI components correctly. An agent skill from crafter-station/petdex.

    4.2k GitHub stars~1.4k tokensUpdated 9 days ago
    Frontend & DesignAuto-check passed
  • Petdex

    crafter-station/petdex

    Browse, install, submit, and edit pixel-art pets with the Petdex CLI.

    4.2k GitHub stars~614 tokensUpdated 9 days ago
    Game DevelopmentAuto-check passed

More from vercel/next.js

All 27 skills in this repo
  • Gh Stack

    vercel/next.js

    Official

    Manages stacked PRs and splits multi-part work into reviewable branches with gh-stack.

    143k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • Next Dev Loop

    vercel/next.js

    Official

    Verify Next.js runtime behavior after editing app code. An agent skill from vercel/next.js.

    143k GitHub starsUsed in 9 repos~2.3k tokens
    Auto-check passed
  • Docs Diagrams

    vercel/next.js

    Official

    Draw diagrams for the Next.js docs in the style of the ones already published there: the light/dark PNGs an mdx references with <Image srcLight="/docs/light/<name.png" srcDark="/docs/dark/<name.png".

    143k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces.

    143k GitHub starsUsed in 6 repos~8.3k tokens
    Auto-check passed
  • React Sync

    vercel/next.js

    Official

    Build local React changes in the bundle variants consumed by Next.js, sync them into a local Next.js checkout, and test the resulting integration.

    143k GitHub stars~486 tokensUpdated today
    Auto-check passed
  • Next Bundle Optimizer

    vercel/next.js

    Official

    Audit and reduce Next.js browser initial-load work. An agent skill from vercel/next.js.

    143k GitHub stars~2.4k tokensUpdated today
    Auto-check passed

Questions about Sandbox Bench

What does Sandbox Bench do?

Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…. js, published by the product's own GitHub organization.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app (rps, latency, p95; TTFB, RSS and document/Flight bytes when the Next side captures them) and, for React changes, through the react repo's flight-ssr-bench fixture (Node AND Edge web-streams paths, Fizz and Flight+Fizz).

When should I use Sandbox Bench?

Sandbox Bench fits situations like: the user asks to bench; A react-server-dom / Flight / vendored React change; A Next.js PR (is this PR faster; does this regress RSC?.

How do I install Sandbox Bench in Claude Code?

Run `npx skills add vercel/next.js --skill sandbox-bench -a claude-code`. Or copy the skill folder (.agents/skills/sandbox-bench in vercel/next.js) into .claude/skills/sandbox-bench in your project. Claude Code loads it when a task matches its description.

How do I install Sandbox Bench in Codex?

Run `npx skills add vercel/next.js --skill sandbox-bench -a codex`. Or copy the skill folder (.agents/skills/sandbox-bench in vercel/next.js) into .agents/skills/sandbox-bench in your project. Codex loads it when a task matches its description.

Can I use Sandbox Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vercel/next.js --skill sandbox-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sandbox-bench, .gemini/skills/sandbox-bench, .github/skills/sandbox-bench and .opencode/skills/sandbox-bench in your project.

What does Sandbox Bench need to run?

Going by SKILL.md and its folder, Sandbox Bench needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node, vercel and bash). Our summary lists: Node.js.

Does Sandbox Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sandbox Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sandbox Bench use?

Sandbox Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sandbox Bench use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Sandbox Bench?

Skills that share tags, products or a category with Sandbox Bench: React Performance (affaan-m/ECC, 275k stars), AI SDK (vercel-labs/ai-facts, 168 stars), Vercel React Best Practices (sanity-io/sanity, 6.4k stars) and Next Best Practices (vercel-labs/openreview, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sandbox Bench?

vercel (a GitHub organization, an official publisher) maintains it in vercel/next.js, which has 143,241 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 8, 2026.

Source: vercel/next.js on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.