Official agent skill

Benchmark Agents

by vercel in vercel/vercel-plugin

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

OfficialCustom licenceAuto-check passedAgent Workflows

Install Benchmark Agents

skills CLI
$ npx skills add vercel/vercel-plugin --skill benchmark-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vercel/vercel-plugin benchmark-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vercel/vercel-plugin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-agents .claude/skills/benchmark-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-agents
GitHub stars
301
Token cost
~3.6k tokens
SKILL.md length
1,304 words
Files
2
Skills in repo
38
Repo updated
First seen
Licence
Custom licence

At a glance

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

  • Works in 4 steps: Create test directory and install plugin → Launch session via WezTerm → Find the debug log (wait ~25s for… → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers How Evals Work (The Only…, DO NOT (Hard Rules), Setup & Launch (Exact Commands) and Monitoring, plus 8 more sections
  • Calls npx, bun and node

What it does

Benchmark Agents is an agent skill from vercel/vercel-plugin, published by the product's own GitHub organization. Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `prompts.md`).

It sits in Agent Workflows, covering LLM evaluation, Load testing and Multi-agent orchestration. It works with Vercel, Model Context Protocol and Bash. The repository describes itself as: Comprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel…

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Load testing
  • Tasks that involve Multi-agent orchestration

Example prompts

  • “/benchmark-agents”

Requirements

  • Node.js

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Create test directory and install plugin
  2. Launch session via WezTerm
  3. Find the debug log (wait ~25s for SessionStart hooks)
  4. Launch multiple sessions in parallel

What it can do on your machine

Read from SKILL.md and the folder at commit 82fa491. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • bun
    • node
    • claude
    • bash
    • vercel
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, vercel and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Agents loads about 3.6k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 1,304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,304 words (~3,617 tokens).

“Launch real Claude Code sessions with the plugin installed, verify skill injection, monitor PostToolUse validation catches, and produce a coverage report. This skill covers the full eval loop: setup → launch → monitor → verify → fix → release →…”

— opening of SKILL.md by vercel, Custom licence
name
benchmark-agents

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in .claude/skills/benchmark-agents of vercel/vercel-plugin.

  • SKILL.md
  • prompts.md

Open the folder on GitHubat commit 82fa491

Compare with similar skills

Benchmark Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Agents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Agents this skillvercel/vercel-plugin301—~3.6kAutomated safety check: PassCustom licence
Octocode Graph Eval Loopbgauryy/octocode946—~1.6kAutomated safety check: PassMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
Octocode Benchmark Runnerbgauryy/octocode946—~2.1kAutomated safety check: PassMIT
Agent Observability Eval Bootstrapdatadog-labs/agent-skills177—~25kAutomated safety check: PassMIT
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    946 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    946 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Agent Observability Eval Bootstrap

    datadog-labs/agent-skills

    Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

    177 GitHub stars~25k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Swarmclaw

    swarmclawai/swarmclaw

    Manage your SwarmClaw agent fleet — agents, tasks, chats, chatrooms, goals, schedules, memory, wallets, connectors, autonomy, and 40+ more command groups.

    687 GitHub stars~4.2k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed

More from vercel/vercel-plugin

All 38 skills in this repo
  • Deployments Cicd

    vercel/vercel-plugin

    Official

    Vercel deployment and CI/CD expert guidance. An agent skill from vercel/vercel-plugin.

    301 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Plugin Audit

    vercel/vercel-plugin

    Official

    Audit vercel-plugin performance on real-world projects. An agent skill from vercel/vercel-plugin.

    301 GitHub stars~739 tokensUpdated yesterday
    Auto-check passed
  • Vercel CLI

    vercel/vercel-plugin

    Official

    Vercel CLI expert guidance. An agent skill from vercel/vercel-plugin.

    301 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • AI Gateway

    vercel/vercel-plugin

    Official

    Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and…

    301 GitHub stars~4.4k tokensUpdated yesterday
    Auto-check: notes
  • Vercel Services

    vercel/vercel-plugin

    Official

    Configure and troubleshoot Vercel Services for multiple frontends and backends in one project.

    301 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Official

    Access and test Vercel deployments protected by Vercel Authentication, SSO, or Deployment Protection.

    301 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes

Questions about Benchmark Agents

What does Benchmark Agents do?

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Benchmark Agents is an agent skill from vercel/vercel-plugin, published by the product's own GitHub organization. Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

When should I use Benchmark Agents?

Benchmark Agents fits situations like: tasks that involve LLM evaluation; tasks that involve Load testing; tasks that involve Multi-agent orchestration.

How do I install Benchmark Agents in Claude Code?

Run `npx skills add vercel/vercel-plugin --skill benchmark-agents -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-agents in vercel/vercel-plugin) into .claude/skills/benchmark-agents in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Agents in Codex?

Run `npx skills add vercel/vercel-plugin --skill benchmark-agents -a codex`. Or copy the skill folder (.claude/skills/benchmark-agents in vercel/vercel-plugin) into .agents/skills/benchmark-agents in your project. Codex loads it when a task matches its description.

Can I use Benchmark Agents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vercel/vercel-plugin --skill benchmark-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-agents, .gemini/skills/benchmark-agents, .github/skills/benchmark-agents and .opencode/skills/benchmark-agents in your project.

What does Benchmark Agents need to run?

Going by SKILL.md and its folder, Benchmark Agents needs the command-line tools its instructions call (npx, bun, node, claude, bash and vercel). Our summary lists: Node.js.

Does Benchmark Agents access the network?

SKILL.md contains no URLs. Its commands use npx and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Agents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Agents use?

Benchmark Agents has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Benchmark Agents use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Agents?

Skills that share tags, products or a category with Benchmark Agents: Octocode Graph Eval Loop (bgauryy/octocode, 946 stars), Waza Interactive (microsoft/waza, 1.4k stars), Octocode Benchmark Runner (bgauryy/octocode, 946 stars) and Agent Observability Eval Bootstrap (datadog-labs/agent-skills, 177 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Agents?

vercel (a GitHub organization, an official publisher) maintains it in vercel/vercel-plugin, which has 301 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 6, 2026.

Source: vercel/vercel-plugin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.