Agent skill

Benchmarking

by wellwelwel in wellwelwel/poku

Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…

MITAuto-check passedTesting & QA

Install Benchmarking

skills CLI
$ npx skills add wellwelwel/poku --skill benchmarking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wellwelwel/poku benchmarking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmarking .claude/skills/benchmarking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmarking
GitHub stars
1.2k
Token cost
~3k tokens
SKILL.md length
1,471 words
Files
4
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…

  • Works in 9 steps: Scope → Hypotheses (parallel) → Harness → …
  • Proving the performance of any project logic
  • SKILL.md covers Ground rules, Pipeline, Findings already validated and Checklist
  • Runs TypeScript and Python scripts from its folder; calls npm, bun and node

What it does

Benchmarking is an agent skill from wellwelwel/poku. Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement, composition re-testing, and anomaly forensics. Use when measuring, comparing, or proving the performance of any project logic, when the user asks for benchmarks or optimization validation, or when a micro-optimization claim needs evidence before touching src/.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `templates/fixtures.ts`, `templates/gen_variants.py` and `templates/run.ts`).

It sits in Testing & QA. The repository describes itself as: 🐷 Poku makes testing easy for Node.js, Bun, Deno, and you at the same time. The licence is MIT.

When your agent uses it

  • Proving the performance of any project logic
  • The user asks for benchmarks
  • Optimization validation
  • A micro-optimization claim needs evidence before touching src/

Example prompts

  • “/benchmarking”

Requirements

  • Python 3
  • Node.js

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Scope
  2. Hypotheses (parallel)
  3. Harness
  4. Variants
  5. Measurement
  6. Interpretation
  7. Composition and focused workloads
  8. Anomaly forensics
  9. Application and final proof

What it can do on your machine

Read from SKILL.md and the folder at commit e02df0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • bun
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmarking loads about 3k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 1,471 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wellwelwel/poku at commit e02df0b, republished under its MIT licence (© wellwelwel). 1,471 words, ~3,029 tokens.

Download SKILL.mdSave it as .claude/skills/benchmarking/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
benchmarking
description
Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement, composition re-testing, and anomaly forensics. Use when measuring, comparing, or proving the performance of any project logic, when the user asks for benchmarks or optimization validation, or when a micro-optimization claim needs evidence before touching src/.
user-invocable
true

Benchmarking

Target-agnostic methodology to measure and prove micro-optimizations in any hot path of this project. Theory does not count: every claim must survive a byte-identical correctness gate and a statistically clean benchmark before touching src/.

Ground rules

  1. Behavior-identical or explicitly flagged. A variant must produce byte-identical output (side effects included) versus the baseline. Intentional behavior changes are flagged experiments, and the final decision belongs to the user.
  2. Realistic workloads only. Fixtures replicate what the real call sites feed the target, in production-like ratios.
  3. Sequential measurement, never parallel. One hyperfine comparison at a time, on an otherwise idle machine.
  4. Noise floor discipline. Differences under ~2%, or within one stddev overlap, are neutral. Neutral results do not justify code changes.
  5. Composition is not additive. Re-measure winners combined (see the JIT cliff in Findings).

Pipeline

Phase 0: Scope
  • Map the target functions, every consumer (grep -rn "fnName" src/), and the call frequency of each. Frequency decides the verdict: a 5x win on a function called once per run has zero real impact, and a lookup table built at module load never amortizes for it. The final report must weigh gains against real call counts.
  • Extract the exact runtime constants the target depends on (formatted prefixes, environment values, shared state) so the harness replicates them byte by byte.
Phase 1: Hypotheses (parallel)

Three independent lenses, each producing structured hypotheses:

  • V8/JIT: inline caches and megamorphic property loads, instanceof vs Array.isArray vs getPrototypeOf, typeof chain ordering, inlining budgets (function size), builtin fast paths, deopt triggers.
  • Allocation/GC: per-call allocations, intermediate strings, methods called only as a boolean test, lookup tables for closed domains, pre-allocation when sizes are known.
  • Strings/algorithms: rope (ConsString) vs array+join, repeated scans over the same data, scan order by selectivity, hoisting invariant checks out of loops.

Each hypothesis states: claim, concrete replacement code, the workload where the effect appears, and a behaviorChange boolean. With subagents, run the lenses in parallel (independent and read-only).

Phase 2: Harness

One directory per investigation under the gitignored temp/:

sh
mkdir -p ./temp/bench/<target>/impl ./temp/bench/<target>/results
cp .claude/skills/benchmarking/templates/* ./temp/bench/<target>/
./temp/bench/<target>/
  run.ts                static: generic driver, runs unmodified
  gen_variants.py       static skeleton: fill only the `variants` dict
  fixtures.ts           per target: fill every TODO(target) marker
  impl/baseline.ts      per target: faithful port of the current code, zero project imports
  impl/<variant>.ts     per target: generated, one file per hypothesis, minimal diff over baseline
  results/*.json        per comparison: hyperfine exports

Runner: bun --bun run.ts ... (~11ms startup). Fallback only when Bun is not available: node --import=tsx run.ts ... (~125ms). Never mix runners between the two sides of a comparison. The runner picks the engine: Bun measures JavaScriptCore, Node + tsx measures V8.

run.ts is generic because fixtures.ts exports a fixed contract: TargetModule (surface of the module under test), verifyCases (string per case), benchModes (iterations plus a checksum-returning run), and resetSideEffects/sideEffectsSnapshot for shared counters.

Requirements:

  • Faithful baseline port: copy the target nearly verbatim (types stay, the runner executes TS), inlining constants and stubbing shared state so the module is self-contained.
  • Hot/cold fixture split: common cases weighted ~5x over exotic ones, matching what production sees.
  • Checksum output: every bench mode accumulates and prints a checksum. Prevents dead-code elimination and doubles as a cross-variant sanity check.
  • Calibration: tune iterations until each process does 500-800ms of real work, so startup does not dominate. Check with time bun --bun run.ts bench-x impl/baseline.ts.
  • Verify is the gate: compares baseline vs variant over all fixtures (outputs and side-effect counters), exits non-zero on mismatch. No variant is benchmarked before passing it, except flagged behavior-change experiments whose mismatch is inspected and confirmed as only the intended difference.
  • Domain sweep for enumerable inputs: fixtures sample the input space, which is not enough for numeric kernels with enumerable domains (a byte, a bounded integer, a nanoseconds field). Brute-force compare baseline vs variant over the domain (costs seconds). A variant here passed an 8-sample gate while diverging by 1 ULP on 29% of the domain.
Phase 3: Variants

gen_variants.py generates each variant from the baseline via exact-match replacements with occurrence-count assertions. Fill only the variants dict, copying old blocks verbatim from impl/baseline.ts (indentation included). A not-found or miscounted block aborts generation: that is the protection against silently benchmarking a baseline-identical variant. Never hand-edit near-identical variant files.

Phase 4: Measurement

One comparison at a time. hyperfine already runs its warmups and runs sequentially, so isolation only breaks at the orchestration level: never start a comparison while another is running.

sh
hyperfine --warmup 3 --runs 10 \
  -n baseline 'bun --bun run.ts bench-x impl/baseline.ts' \
  -n variant  'bun --bun run.ts bench-x impl/variant.ts' \
  --export-json results/variant.json
  • --warmup 3 --runs 10 is the project standard. Same runner on both sides, always.
  • Export JSON, record mean, stddev, min, max per side, and hyperfine outlier warnings verbatim (they change interpretation).
  • With subagents: one per comparison, in a sequential loop (for ... await), never parallel().
Phase 5: Interpretation
  • Under ~2%, or within one stddev overlap: neutral, discard.
  • Outlier warnings with inflated mean: compare medians, or re-run.
  • A regression is as valuable as a win: it proves the current code right and blocks future "improvements". Document it.
  • Behavior-change variants that measure neutral are dead (all cost, no gain).
Phase 6: Composition and focused workloads
  • Combine winners into a combined variant and re-measure against baseline: gains do not add up.
  • Wins diluted in the mixed workload (a branch only some inputs reach) get a focused bench mode exercising only that path, measured baseline vs variant vs combined. A 0.998x "neutral" here measured 2.9x focused.
Phase 7: Anomaly forensics

When a combination behaves unexpectedly, isolate before discarding:

  1. Micro-mutation bisection: variants differing by exactly one ingredient until the toxic pair appears.
  2. Deopt check: node --trace-deopt ... | grep -c deoptimizing against baseline.
  3. JIT on/off: run both shapes with --max-opt=0. A tie interpreted plus divergence with JIT on means an optimization-quality cliff, not algorithmic cost. Document the toxic shape.
Show full SKILL.md (595 more words)Show less
Phase 8: Application and final proof
  1. Port winners to the real source, preserving project style.
  2. Re-verify the real source: an adapter module re-exports the actual src/**/*.ts symbols under the TargetModule contract, then:
    sh
    bun --bun run.ts verify impl/real-src.ts
    This catches transcription errors between the benchmarked port and the applied code.
  3. Confirmation benchmark of the real source against impl/baseline.ts. The canonical numbers remain the Phase 4-6 ones.
  4. npm run typecheck && npm run lint:fix && npm run build && npm test.
  5. Internal contract changes update all callers in the same change. Tests pinning the old internal shape may be updated only when the change is intentional and the validated content stays the same.

Findings already validated

Transferable heuristics, measured on Apple Silicon (engine noted when it matters). Re-validate before relying on them in another engine.

Confirmed wins:

  • Functions beyond the JIT inlining budget benefit from a small dispatch entry that handles cheap common cases inline and delegates the heavy body to a second function (-8%, V8).
  • getPrototypeOf(x) === Object.prototype check before instanceof chains lets a dominant plain-object path skip them (-3.7%, V8).
  • Needle selectivity for substring search: when a multi-character needle starts with a prefix that is common in the haystack, prefilter with the needle's rarest single character first. Single-char includes uses a vectorized path and discards most lines in one pass (-19.5%, V8).
  • Rope accumulator over array+join when the consumer wants one string: acc += part beats parts.push(part) + join (-15%, V8).
  • Lookup tables for closed domains on the affected path (2.9x focused on V8, 5.8x on JSC, invisible in mixed workloads). Only worth applying when call counts amortize the table built at module load.
  • Ordering scans by selectivity: the cheapest and most selective test first, the long needle only on survivors.

Confirmed non-wins (do not "fix" these):

  • localeCompare() with no arguments has a V8 fast path. A cached Intl.Collator().compare is 2.8x slower. Default sort() was neutral, so dropping locale order buys nothing.
  • split('\n') loses to a manual indexOf/substring loop (+18%, V8).
  • /\S/.test(line) loses to line.trim().length > 0 as a blank-line test (+6%, V8).
  • Pre-gating searches with a shared-prefix scan loses (+9%, V8): the extra scan does not pay for itself.
  • A native API replacing a tiny hand-written formatter lost 2x (JSC): Date.prototype.toTimeString().slice() builds a far larger string than three padded getters.
  • Replacing / 1eN division with * 1e-N multiplication is a correctness bug, not an optimization: results diverge by 1 ULP on a large share of the domain (engine-independent IEEE behavior).
  • Manual toFixed replacements change rounding on binary-float edges (0.015 formats as 0.01, round-half-up gives 0.02).
  • typeof chain order (primitives first vs object first) is a measured tie (1.00 ± 0.09). Pick by readability.
  • Precomputed separator tables, first-iteration peeling, lazy Set allocation, manual key quoting: all within noise (V8).

Documented JIT cliff (V8): memoized indexOf(needle, offset) positions combined with a rope accumulator in the same loop measured 2.1x slower, while each ingredient alone was faster. Identical when interpreted (--max-opt=0), so the combined function shape pessimizes under TurboFan. Re-measure every composition.

Checklist

  • Map call sites, data shapes, and call frequency of the target
  • Generate hypotheses through the three lenses (parallel subagents work well)
  • Copy templates to ./temp/bench/<target>/, fill baseline, fixtures, and calibration
  • Generate variants via the variants dict, never by hand
  • Gate every variant with verify, plus a domain sweep when inputs are enumerable
  • hyperfine --warmup 3 --runs 10, sequential, same runner both sides, export JSON
  • Discard anything within ±2% or one stddev
  • Re-measure winners combined, add focused modes for diluted wins
  • Bisect anomalies (micro-mutations, --trace-deopt, --max-opt=0)
  • Apply only when call frequency justifies it, re-verify the real source, run typecheck/lint/build/tests

© wellwelwel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .claude/skills/benchmarking of wellwelwel/poku.

  • SKILL.md
  • templates/fixtures.ts
  • templates/gen_variants.py
  • templates/run.ts

Open the folder on GitHubat commit e02df0b

Compare with similar skills

Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmarking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmarking this skillwellwelwel/poku1.2k—~3kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k32 repos~2.1kAutomated safety check: PassApache-2.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 32 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    381 GitHub starsUsed in 9 repos~2.9k tokens
    Testing & QAAuto-check passed

More from wellwelwel/poku

  • Architecture

    wellwelwel/poku

    Architecture deep-dive for poku covering project structure, execution flow, config discovery, runtime detection, and competitive context.

    1.2k GitHub stars~867 tokensUpdated 3 mo ago
    Auto-check passed
  • Documentation

    wellwelwel/poku

    Documentation deep-dive for the poku website covering Docusaurus setup, language policy, feature-to-docs mapping, conventions, JSDoc, and the changelog.

    1.2k GitHub stars~982 tokensUpdated 3 mo ago
    Auto-check passed
  • Engineering

    wellwelwel/poku

    Engineering deep-dive for poku covering performance, code, and security patterns, TypeScript config, build pipeline, developer experience, and CI/CD.

    1.2k GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Pre Commit

    wellwelwel/poku

    Run the mandatory pre-commit checks for the poku repository before staging a commit.

    1.2k GitHub stars~728 tokensUpdated 3 mo ago
    Auto-check passed
  • Testing

    wellwelwel/poku

    Testing deep-dive for poku covering test structure, commands, patterns, fixtures, utils, Docker compatibility, and coverage.

    1.2k GitHub stars~1k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Benchmarking

What does Benchmarking do?

Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…. Benchmarking is an agent skill from wellwelwel/poku. Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement, composition re-testing, and anomaly forensics.

When should I use Benchmarking?

Benchmarking fits situations like: proving the performance of any project logic; the user asks for benchmarks; optimization validation; A micro-optimization claim needs evidence before touching src/.

How do I install Benchmarking in Claude Code?

Run `npx skills add wellwelwel/poku --skill benchmarking -a claude-code`. Or copy the skill folder (.claude/skills/benchmarking in wellwelwel/poku) into .claude/skills/benchmarking in your project. Claude Code loads it when a task matches its description.

How do I install Benchmarking in Codex?

Run `npx skills add wellwelwel/poku --skill benchmarking -a codex`. Or copy the skill folder (.claude/skills/benchmarking in wellwelwel/poku) into .agents/skills/benchmarking in your project. Codex loads it when a task matches its description.

Can I use Benchmarking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wellwelwel/poku --skill benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking, .gemini/skills/benchmarking, .github/skills/benchmarking and .opencode/skills/benchmarking in your project.

What does Benchmarking need to run?

Going by SKILL.md and its folder, Benchmarking needs TypeScript and Python for the scripts in its folder and the command-line tools its instructions call (npm, bun and node). Our summary lists: Python 3; Node.js.

Does Benchmarking access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmarking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmarking use?

Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmarking use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmarking?

Skills that share tags, products or a category with Benchmarking: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmarking?

wellwelwel (a GitHub user) maintains it in wellwelwel/poku, which has 1,183 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 18, 2026.

Source: wellwelwel/poku on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.