Official agent skill

Benchmark Sandbox

by vercel in vercel/vercel-plugin

Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.

OfficialCustom licenceAuto-check: notesProduct & Project Management

Install Benchmark Sandbox

skills CLI
$ npx skills add vercel/vercel-plugin --skill benchmark-sandbox -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vercel/vercel-plugin benchmark-sandbox --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vercel/vercel-plugin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-sandbox .claude/skills/benchmark-sandbox && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-sandbox
GitHub stars
301
Token cost
~5.4k tokens
SKILL.md length
1,946 words
Files
9
Skills in repo
38
Repo updated
First seen
Licence
Custom licence

At a glance

Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.

  • Works in 12 steps: Snapshots work: sandbox.snapshot()… → Plugin install: Use npx add-plugin -s… → File uploads: Use sandbox.writeFiles([{… → …
  • Tasks that involve Test coverage
  • SKILL.md covers Proven Working Script, Dynamic Scenarios (Recommended…, Structured Scoring (Haiku) and Critical Sandbox Environment…, plus 12 more sections
  • Runs TypeScript scripts from its folder; calls bun, npx and vercel; reaches sb-xxx.vercel.run and xxx.vercel.app; needs ANTHROPIC_API_KEY and ANTHROPIC_AUTH_TOKEN

What it does

Benchmark Sandbox is an agent skill from vercel/vercel-plugin, published by the product's own GitHub organization. Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels. Provisions ephemeral microVMs with Claude Code + plugin pre-installed, runs benchmark prompts, extracts hook artifacts, and produces coverage reports.

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files (for example `logger.ts`, `run-eval.ts` and `sandbox-analyze.ts`).

It sits in Product & Project Management, covering Test coverage. It works with Vercel. The repository describes itself as: Comprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel…

When your agent uses it

  • Tasks that involve Test coverage

Example prompts

  • “/benchmark-sandbox”

Requirements

  • Node.js
  • A credential in ANTHROPIC_AUTH_TOKEN
  • A credential in ANTHROPIC_API_KEY

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Snapshots work: sandbox.snapshot() preserves files AND npm globals. Use it after build to create a restore point before verify/deploy…
  2. Plugin install: Use npx add-plugin -s project -y --target claude-code — works because claude is in PATH after npm install -g. The --target…
  3. File uploads: Use sandbox.writeFiles([{ path, content: Buffer }]) — NOT runCommand heredocs. Heredocs with special characters cause 400…
  4. Claude flags: Always use --dangerously-skip-permissions --debug. The --debug flag writes to ~/.claude/debug/.
  5. Auth: API key from macOS Keychain (ANTHROPIC_AUTH_TOKEN — a vck_* Vercel Claude Key for AI Gateway), Vercel token from…
  6. OIDC for sandbox SDK: Run npx vercel link --scope vercel-labs -y + npx vercel env pull once before first use.
  7. Port exposure: Pass ports: [3000] in Sandbox.create() to get a public URL immediately via sandbox.domain(3000). Works on v1.8.0 — URL is…
  8. extendTimeout: Use sandbox.extendTimeout(ms) to keep sandboxes alive past their initial timeout. Verified working — extends by the…
  9. Background commands: runCommand with backgrounded processes (& or nohup) may throw ZodError on v1. Write a script file first, then execute…
  10. Session cleanup race: The session-end-cleanup.mjs hook deletes /tmp/vercel-plugin-*-seen-skills.d/ on session end. Extract artifacts…
  11. agent-browser works in sandboxes: Install via npm install -g agent-browser. Claude Code can use it for browser-based verification inside…
  12. No hobby tier cap: Early 301s timeouts were from lower default timeout values in earlier script iterations, not a tier limitation…

What it can do on your machine

Read from SKILL.md and the folder at commit 82fa491. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • bun
    • npx
    • vercel
    • claude
    • npm
    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • sb-xxx.vercel.run
    • xxx.vercel.app
    • ai-gateway.vercel.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • ANTHROPIC_AUTH_TOKEN
    • VERCEL_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Sandbox loads about 5.4k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,946 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:278
    npx vercel env pull .env.local

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,946 words (~5,403 tokens).

“Run benchmark scenarios inside Vercel Sandboxes — ephemeral Firecracker microVMs with node24. Each sandbox gets a fresh Claude Code + Vercel CLI + agent-browser install, the local vercel-plugin uploaded, and runs a 3-phase eval pipeline:”

— opening of SKILL.md by vercel, Custom licence
name
benchmark-sandbox

Read the full SKILL.md on GitHub

Files

SKILL.md and 8 other files in .claude/skills/benchmark-sandbox of vercel/vercel-plugin.

  • SKILL.md
  • logger.ts
  • run-eval.ts
  • sandbox-analyze.ts
  • sandbox-runner.ts
  • spike/create-snapshot.ts
  • spike/interactive-proof.ts
  • spike/provision.ts
  • types.ts

Open the folder on GitHubat commit 82fa491

Compare with similar skills

Benchmark Sandbox next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Sandbox compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Sandbox this skillvercel/vercel-plugin301—~5.4kAutomated safety check: NotesCustom licence
Enrich Release Roadmaptrickle-labs/pg-trickle146—~1.9kAutomated safety check: PassApache-2.0
AlignmentOvid/paad131—~4.9kAutomated safety check: PassMIT
Nextjs Supabase VercelTechNomadCode/AI-Product-Development-Toolkit1k1 repos~159Automated safety check: PassMIT
Test Planningpetrkindlmann/qa-skills163—~4.2kAutomated safety check: PassMIT
Sveltekit WebappLeoYeAI/openclaw-master-skills2.2k—~5.2kAutomated safety check: NotesMIT

Similar skills

  • Enrich Release Roadmap

    trickle-labs/pg-trickle

    Enrich a pgtrickle release roadmap with prioritised items across six quality pillars: correctness, stability, performance, scalability, ease-of-use, and test coverage.

    146 GitHub stars~1.9k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • Alignment

    Ovid/paad

    A skill your agent uses when verifying that requirements/specs/PRDs and their implementation plans match — before starting work, after a spec or plan update, or when suspecting coverage gaps, scope…

    131 GitHub stars~4.9k tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • Nextjs Supabase Vercel

    TechNomadCode/AI-Product-Development-Toolkit

    Set up the current empty folder with the AI App Starter, so the agent builds a Next.js, Supabase and Vercel app with the owner step by step, from idea to launch.

    1k GitHub starsUsed in 1 repo~159 tokens
    Product & Project ManagementAuto-check passed
  • Test Planning

    petrkindlmann/qa-skills

    Build a single sprint or release test plan. An agent skill from petrkindlmann/qa-skills.

    163 GitHub stars~4.2k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Sveltekit Webapp

    LeoYeAI/openclaw-master-skills

    Scaffold and configure a production-ready SvelteKit PWA with opinionated defaults.

    2.2k GitHub stars~5.2k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check: notes
  • Vibe Scenario Matrix

    ash1794/vibe-engineering

    Generates a behavioral scenario acceptance matrix for comprehensive test coverage planning.

    162 GitHub stars~663 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from vercel/vercel-plugin

All 38 skills in this repo
  • Deployments Cicd

    vercel/vercel-plugin

    Official

    Vercel deployment and CI/CD expert guidance. An agent skill from vercel/vercel-plugin.

    301 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Plugin Audit

    vercel/vercel-plugin

    Official

    Audit vercel-plugin performance on real-world projects. An agent skill from vercel/vercel-plugin.

    301 GitHub stars~739 tokensUpdated today
    Auto-check passed
  • Vercel CLI

    vercel/vercel-plugin

    Official

    Vercel CLI expert guidance. An agent skill from vercel/vercel-plugin.

    301 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • AI Gateway

    vercel/vercel-plugin

    Official

    Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and…

    301 GitHub stars~4.4k tokensUpdated today
    Auto-check: notes
  • Vercel Services

    vercel/vercel-plugin

    Official

    Configure and troubleshoot Vercel Services for multiple frontends and backends in one project.

    301 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Official

    Access and test Vercel deployments protected by Vercel Authentication, SSO, or Deployment Protection.

    301 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes

Works with

Questions about Benchmark Sandbox

What does Benchmark Sandbox do?

Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels. Benchmark Sandbox is an agent skill from vercel/vercel-plugin, published by the product's own GitHub organization. Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.

When should I use Benchmark Sandbox?

Benchmark Sandbox fits situations like: tasks that involve Test coverage.

How do I install Benchmark Sandbox in Claude Code?

Run `npx skills add vercel/vercel-plugin --skill benchmark-sandbox -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-sandbox in vercel/vercel-plugin) into .claude/skills/benchmark-sandbox in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Sandbox in Codex?

Run `npx skills add vercel/vercel-plugin --skill benchmark-sandbox -a codex`. Or copy the skill folder (.claude/skills/benchmark-sandbox in vercel/vercel-plugin) into .agents/skills/benchmark-sandbox in your project. Codex loads it when a task matches its description.

Can I use Benchmark Sandbox in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vercel/vercel-plugin --skill benchmark-sandbox -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-sandbox, .gemini/skills/benchmark-sandbox, .github/skills/benchmark-sandbox and .opencode/skills/benchmark-sandbox in your project.

What does Benchmark Sandbox need to run?

Going by SKILL.md and its folder, Benchmark Sandbox needs TypeScript for the scripts in its folder, the command-line tools its instructions call (bun, npx, vercel, claude, npm and sh) and credentials named ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN and VERCEL_TOKEN. Our summary lists: Node.js; A credential in ANTHROPIC_AUTH_TOKEN; A credential in ANTHROPIC_API_KEY.

Does Benchmark Sandbox access the network?

SKILL.md names 3 domains. In commands or code: sb-xxx.vercel.run, xxx.vercel.app and ai-gateway.vercel.sh; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Benchmark Sandbox safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Sandbox use?

Benchmark Sandbox has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Benchmark Sandbox use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Sandbox?

Skills that share tags, products or a category with Benchmark Sandbox: Enrich Release Roadmap (trickle-labs/pg-trickle, 146 stars), Alignment (Ovid/paad, 131 stars), Nextjs Supabase Vercel (TechNomadCode/AI-Product-Development-Toolkit, 1k stars) and Test Planning (petrkindlmann/qa-skills, 163 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Sandbox?

vercel (a GitHub organization, an official publisher) maintains it in vercel/vercel-plugin, which has 301 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 6, 2026.

Source: vercel/vercel-plugin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.