Agent skill

Writing Benchmarks

by computesdk in computesdk/benchmarks

Author or improve a ComputeSDK benchmark. An agent skill from computesdk/benchmarks.

MITAuto-check: notes

Install Writing Benchmarks

skills CLI
$ npx skills add computesdk/benchmarks --skill writing-benchmarks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install computesdk/benchmarks writing-benchmarks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/computesdk/benchmarks.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/writing-benchmarks .claude/skills/writing-benchmarks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-benchmarks
GitHub stars
126
Token cost
~2.1k tokens
SKILL.md length
909 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Author or improve a ComputeSDK benchmark. An agent skill from computesdk/benchmarks.

  • Works in 8 steps: Define the unit of work → Horizontal load — iterations × concurrency → Decompose the work into steps → …
  • Asked to create
  • SKILL.md covers 1. Define the unit of work, 2. Horizontal load —…, 3. Decompose the work into steps and 4. Write the file, plus 5 more sections
  • Calls pnpm; needs BENCHMARKS_PLATFORM_API_KEY

What it does

Writing Benchmarks is an agent skill from computesdk/benchmarks. Author or improve a ComputeSDK benchmark. Turn a unit of work into a declarative .bench.ts file (config + task with benchsdk steps), set up participants, scoring, and display, verify with a dry run, and wire it into package.json scripts and CI. Use when asked to create, extend, or improve a benchmark.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with npm. The repository describes itself as: Compare performance across sandbox, storage, browsers, and ai gateway providers. The licence is MIT.

When your agent uses it

  • Asked to create
  • Improve a benchmark

Example prompts

  • “/writing-benchmarks”

Requirements

  • A credential in BENCHMARKS_PLATFORM_API_KEY

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Define the unit of work
  2. Horizontal load — iterations × concurrency
  3. Decompose the work into steps
  4. Write the file
  5. Participants
  6. Scoring and display
  7. Verify
  8. Wire it in

What it can do on your machine

Read from SKILL.md and the folder at commit bcf2052. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BENCHMARKS_PLATFORM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Writing Benchmarks loads about 2.1k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 909 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:101
    ../src/env.js'` first (loads `benchmarks/.env`).

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from computesdk/benchmarks at commit bcf2052, republished under its MIT licence (© computesdk). 909 words, ~2,087 tokens.

Download SKILL.mdSave it as .claude/skills/writing-benchmarks/SKILL.md (or your agent's skills folder).
name
writing-benchmarks
description
Author or improve a ComputeSDK benchmark. Turn a unit of work into a declarative *.bench.ts file (config + task with benchsdk steps), set up participants, scoring, and display, verify with a dry run, and wire it into package.json scripts and CI. Use when asked to create, extend, or improve a benchmark.

Writing a benchmark

A benchmark is a unit of work (what one iteration measures) expressed as a declarative *.bench.ts file that exports exactly config + task. The bench run CLI owns orchestration — the file never calls the runner itself.

Canonical API reference: WRITING_BENCHMARKS.md (repo root) — read it for the full field tables. Runnable templates: examples/01-hello.bench.ts through 11-logging.bench.ts. Reference implementations:

  • benchmarks/sandbox/tti.bench.ts — minimal sandbox benchmark (create → exec → destroy, one metric).
  • benchmarks/sandbox/dax.bench.ts — rich benchmark (work measured inside the sandbox, pre-measured phase steps, TaskError, legacy results writer).

1. Define the unit of work

Before writing code, pin down:

  • Timed boundary — what starts the clock and what ends it. For TTI that's sandbox.create() → first command succeeding; destroy is excluded. Write the boundary in the file's header comment.
  • Uniformity — the task body must run identically for every participant. Push provider differences into participant config (sandboxOptions, timeout, destroyTimeoutMs), never into if (provider === ...) branches inside the task.
  • Headline metric(s) — the numbers the task reports via measure() and what counts as a success (exit-code checks, scoring.success.requireData).
  • Load profile — which of iterations, concurrency, staggerDelayMs define the benchmark vs. stay CLI-overridable. A shape may set slug/name/staggerDelayMs, never iterations/concurrency. See § Horizontal load.

2. Horizontal load — iterations × concurrency

Horizontal load (how many units of work run at once, and how fast they ramp) is controlled entirely by three config/CLI knobs — the task function itself never manages parallelism:

  • iterations — total tasks per participant. Each iteration is one full unit of work (one sandbox create → work → destroy). This is the horizontal volume knob: iterations: 100 launches 100 sandboxes per participant.
  • concurrency — max tasks in flight at once. 1 = sequential (one-at-a-time cold starts), N = up to N units of work running simultaneously. iterations: N, concurrency: N is a burst: all N sandboxes launch at once.
  • staggerDelayMs — task i launches at i * staggerDelayMs. Combined with concurrency: N this is a staggered ramp (e.g. 100 sandboxes, one every 200ms) — measures how TTI degrades as load ramps.

The named run shapes are just presets over these knobs: sequential = {concurrency: 1}, burst = {concurrency: N}, staggered = {concurrency: N, staggerDelayMs: D}.

Two related knobs that change ordering/fan-out, not load:

  • groupBy — 'participant' (default: each participant runs to completion before the next) vs 'round' (participants take turns per iteration, so every participant's Nth task sees the same external conditions).
  • phases — split iterations into named segments (cold/warm). Total iterations = sum of phase iterations; --iterations scales each phase.

In CI, horizontal scale comes from the matrix: one job per provider, each passing --provider <name> plus the same --run-key so they share one platform run, and --iterations/--concurrency from workflow inputs.

3. Decompose the work into steps

Wrap each phase in step('name', fn) — each step becomes a timed, labeled row on the platform run page.

Sandbox-benchmark conventions (follow tti.bench.ts):

  • Steps: create → the work → destroy. Keep names short; display.steps maps them to labels.
  • Always destroy in finally: step('destroy', () => withTimeout(sandbox.destroy(), timeout), { reportConcurrency: false }).catch(err => log('destroy failed', { level: 'warn', ... })).
  • Guard every external call with withTimeout (benchmarks/src/util/timeout.ts) and a provider overridable timeout.
  • log() with { level, meta } for diagnosis — the worker log uploads as a run artifact. Log the sandbox id after destroy (sandboxId() util).
  • Step return values shaped like { stdout, stderr, error, exitCode, ... } are captured as step output, not returned — pass captureOutput: false or wrap.
  • Work measured inside the sandbox (a script printing BENCH_* lines): return { data, steps, latencyMs } from the task with pre-measured TaskStepRecord[], and set groupBy: 'round' so the runner honors them (participant mode ignores pre-measured steps). See dax.bench.ts.
  • Errors: throw TaskError (code + data) for domain failures; let plain Errors bubble for unexpected ones.
Show full SKILL.md (326 more words)Show less

4. Write the file

  • Place under benchmarks/<area>/<name>.bench.ts. Start from the closest existing benchmark, not from scratch.
  • import '../src/env.js' first (loads benchmarks/.env).
  • config essentials: benchmarkSlug, benchmarkName, iterations, concurrency, participants, display, scoring. Optional: shapes, phases, groupBy, defaultProviders, dimensions, onComplete, customCliFlags — see WRITING_BENCHMARKS.md before reaching for them.

5. Participants

  • Sandbox benchmarks share benchmarks/sandbox/providers.ts (ProviderConfig: name, requiredEnvVars, createCompute, sandboxOptions, timeout, destroyTimeoutMs). Other areas have their own providers.ts; new domains should copy the pattern.
  • Missing requiredEnvVars → participant is skipped with a log; all skipped → clean NoAvailableParticipantsError exit. --provider filters further.
  • Vercel always gets sandboxOptions: { persistent: false } — the default auto-snapshots on every stop() and accrues Snapshot Storage.
  • defaultProviders limits what runs without --provider (e.g. dax runs a subset by default).

6. Scoring and display

  • scoring.metrics: weights.median + p95 + p99 across all metrics must sum to 1.0; ceiling = worst acceptable value; higherIsBetter: true + floor for throughput-style metrics; trim (default 0.05) drops outliers before median/p95/p99.
  • Prefer scoring (validated at startup); use onScore only when a metric needs a function extractor.
  • display.steps[].key must exactly match step() names; display.metrics[].key matches measure() keys. Undeclared keys still render — declarations only improve labels/units/direction.
  • scoring.groupBy renders a breakout per distinct data value (only with 2–12 values) — different from top-level groupBy (execution ordering).

7. Verify

bash
# packages/*/dist is not committed — build once first
pnpm -r --filter "./packages/**" build
pnpm typecheck

# Dry run still requires platform auth, but skips ingestion
BENCHMARKS_PLATFORM_API_KEY=bp-... pnpm exec bench run \
  benchmarks/<area>/<name>.bench.ts --provider <name> --iterations 2 --dry-run
  • No provider creds? Exercise the file with a stub participant (requiredEnvVars: []), or use the local-platform-e2e skill to run against a local benchmarks-platform.
  • Check the run page URL printed by the runner when not dry-running.

8. Wire it in

  • package.json: add "bench:<name>" (and bench:<name>:<provider> variants when useful) following the existing script shape.
  • CI: copy the closest workflow in .github/workflows/ — provider matrix, load-vault-secrets.sh regex (add any new env var names), shared --run-key "$GITHUB_RUN_ID-$GITHUB_RUN_ATTEMPT", --no-ingest on push, SHOULD_RUN guard for single-provider dispatch.
  • results/ legacy JSON + generate-*.svg only if the README/results pipeline consumes the new benchmark — otherwise skip the legacy writers.
  • Draft PR by default.

Gotchas

Read WRITING_BENCHMARKS.md § "Common gotchas" before debugging: stale dist/, --iterations scaling per-phase, env-filtered participants, the weight-sum rule, outcome-shaped step returns, round-vs-participant step ownership, Vercel persistent: false.

© computesdk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/writing-benchmarks of computesdk/benchmarks.

Open the folder on GitHubat commit bcf2052

Compare with similar skills

Writing Benchmarks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Writing Benchmarks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Writing Benchmarks this skillcomputesdk/benchmarks126—~2.1kAutomated safety check: NotesMIT
Defuddlekepano/obsidian-skills49k12 repos~208Automated safety check: PassMIT
Vercel Deploybytedance/deer-flow83k10 repos~797Automated safety check: PassMIT
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Knap Markdown Templateskepano/obsidian-skills49k2 repos~986Automated safety check: PassMIT
Install Anti-Slop Oxlint Rulesdmmulroy/anti-slop5.3k1 repos~2.2kAutomated safety check: PassMIT

Similar skills

  • Defuddle

    kepano/obsidian-skills

    Uses the Defuddle CLI to pull clean, readable Markdown, JSON or metadata from web pages, stripping navigation, ads and clutter to save tokens.

    49k GitHub starsUsed in 12 repos~208 tokens
    Knowledge ManagementAuto-check passed
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    83k GitHub starsUsed in 10 repos~797 tokens
    DevOps & CloudAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Knap Markdown Templates

    kepano/obsidian-skills

    Renders Markdown notes from Knap templates and JSON data on the command line, including notes built from Defuddle web page output.

    49k GitHub starsUsed in 2 repos~986 tokens
    Documents & OfficeAuto-check passed
  • Installs, updates or migrates the vendored anti-slop Oxlint plugin in a repository, keeping local rule changes and the plugin's license and provenance files.

    5.3k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Nx Run Tasks

    nomcopter/react-mosaic

    Helps with running tasks in an Nx workspace. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 8 repos~613 tokens
    DevelopmentAuto-check passed

More from computesdk/benchmarks

  • Add Sandbox Provider

    computesdk/benchmarks

    Prep a new sandbox provider for the ComputeSDK benchmarks (wire up the dependency, env var, providers list, and CI workflow).

    126 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Local Platform E2E

    computesdk/benchmarks

    Stand up benchmarks-platform locally (Postgres + MinIO + ClickHouse in docker) and run a real @benchsdk/runner benchmark against it, with no cloud or provider credentials.

    126 GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • Benchsdk CLI

    computesdk/benchmarks

    Install, authenticate, and use the unified bench CLI to query and run ComputeSDK benchmarks.

    126 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Writing Benchmarks

What does Writing Benchmarks do?

Author or improve a ComputeSDK benchmark. An agent skill from computesdk/benchmarks. Writing Benchmarks is an agent skill from computesdk/benchmarks. Author or improve a ComputeSDK benchmark.

When should I use Writing Benchmarks?

Writing Benchmarks fits situations like: asked to create; improve a benchmark.

How do I install Writing Benchmarks in Claude Code?

Run `npx skills add computesdk/benchmarks --skill writing-benchmarks -a claude-code`. Or copy the skill folder (.agents/skills/writing-benchmarks in computesdk/benchmarks) into .claude/skills/writing-benchmarks in your project. Claude Code loads it when a task matches its description.

How do I install Writing Benchmarks in Codex?

Run `npx skills add computesdk/benchmarks --skill writing-benchmarks -a codex`. Or copy the skill folder (.agents/skills/writing-benchmarks in computesdk/benchmarks) into .agents/skills/writing-benchmarks in your project. Codex loads it when a task matches its description.

Can I use Writing Benchmarks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add computesdk/benchmarks --skill writing-benchmarks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-benchmarks, .gemini/skills/writing-benchmarks, .github/skills/writing-benchmarks and .opencode/skills/writing-benchmarks in your project.

What does Writing Benchmarks need to run?

Going by SKILL.md and its folder, Writing Benchmarks needs the command-line tools its instructions call (pnpm) and credentials named BENCHMARKS_PLATFORM_API_KEY. Our summary lists: A credential in BENCHMARKS_PLATFORM_API_KEY.

Does Writing Benchmarks access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Writing Benchmarks safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Writing Benchmarks use?

Writing Benchmarks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Writing Benchmarks use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Writing Benchmarks?

Skills that share tags, products or a category with Writing Benchmarks: Defuddle (kepano/obsidian-skills, 49k stars), Vercel Deploy (bytedance/deer-flow, 83k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars) and Knap Markdown Templates (kepano/obsidian-skills, 49k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Writing Benchmarks?

computesdk (a GitHub organization) maintains it in computesdk/benchmarks, which has 126 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: computesdk/benchmarks on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.