Agent skill

Benchmarking

by nubjs in nubjs/nub

Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar.

MITAuto-check passed

Install Benchmarking

skills CLI
$ npx skills add nubjs/nub --skill benchmarking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nubjs/nub benchmarking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nubjs/nub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmarking .claude/skills/benchmarking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmarking
GitHub stars
4.4k
Token cost
~1.8k tokens
SKILL.md length
783 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar.

  • SKILL.md covers The tool: hyperfine, The cold / warm protocol, GENUINE-COLD per tool — the… and VERIFY each cold is genuine —…, plus 7 more sections
  • Calls bun, pnpm and docker

What it does

Benchmarking is an agent skill from nubjs/nub. Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar. Invoke (via the Skill tool) whenever you need to benchmark nub install against another package manager, produce or update the homepage/blog install numbers, or verify a perf claim before it ships. Encodes the hard-won gotchas: time setup OUTSIDE the measurement (hyperfine --prepare), the cache lives on DISK so env-var isolation is NOT…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with pnpm, npm, Bun and Rust. The repository describes itself as: The fast all-in-one Node.js toolkit. The licence is MIT.

Example prompts

  • “/benchmarking”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 568e73a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • pnpm
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmarking loads about 1.8k tokens when it runs. Until then it costs about 215 tokens; SKILL.md has 783 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~215
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nubjs/nub at commit 568e73a, republished under its MIT licence (© nubjs). 783 words, ~1,844 tokens.

Download SKILL.mdSave it as .claude/skills/benchmarking/SKILL.md (or your agent's skills folder).
name
benchmarking
description
Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar. Invoke (via the Skill tool) whenever you need to benchmark `nub install` against another package manager, produce or update the homepage/blog install numbers, or verify a perf claim before it ships. Encodes the hard-won gotchas: time setup OUTSIDE the measurement (hyperfine `--prepare`), the cache lives on DISK so env-var isolation is NOT trustworthy (bun ignores `BUN_INSTALL_CACHE_DIR`/`$HOME` — wipe the real path), VERIFY every cold is genuine via an offline-fails check, and only measure wall-clock on a quiet machine (gate on low load) with file counts as a load-independent cross-check. Pairs with `pm-perf-tracing` for the internal Rust phase decomposition.
metadata
internal: true

benchmarking

Honest, reproducible comparative install benchmarks of nub install against npm / pnpm / bun — the external, wall-clock-and-file-count method. For decomposing where the time goes INSIDE a single nub install, use pm-perf-tracing instead.

Applies to: benchmarking against another PM, refreshing homepage/blog install numbers, verifying a perf claim before it ships. A single non-genuine cell discredits the whole table.

The tool: hyperfine

/opt/homebrew/bin/hyperfine is the canonical timer. The load-bearing flag is --prepare, which runs setup BEFORE each timed run, untimed. Never time setup.

sh
# COLD: empty the tool's REAL cache + wipe node_modules before each run (untimed), then time the install.
hyperfine --warmup 0 --runs 5 \
  --prepare 'rm -rf node_modules && d="$(bun pm cache)" && rm -rf "${d:?}"' \
  'bun install --ignore-scripts'

# WARM-RELINK: cache populated, wipe ONLY node_modules before each run.
hyperfine --warmup 1 --runs 5 \
  --prepare 'rm -rf node_modules' \
  'nub install --ignore-scripts'

# WARM-SAT: node_modules already present (idempotency path) — no prepare wipe.
hyperfine --warmup 1 --runs 5 'nub install --ignore-scripts'

Prefer median + spread (min–max) over the mean — contention skews the mean.

The cold / warm protocol

node_modules is deleted before EVERY timed run, cold and warm, always in --prepare. The cold/warm axis differs only in global-cache state:

  • cold — the tool's real cache is EMPTY (genuine download).
  • warm-relink — cache populated, node_modules wiped (the link-from-store path; the homepage number).
  • warm-sat — node_modules already present; a separate, clearly-labeled scenario, not the headline.

GENUINE-COLD per tool — the cache is on DISK

A "cold" run is only cold if the tool's real on-disk cache is gone. Setting a cache-dir env var does not guarantee that — bun ignores BUN_INSTALL_CACHE_DIR and $HOME, resolving its cache via the OS passwd home. You must wipe the disk path.

toolreal cache pathclear command for cold
nubits store (NUB_CACHE_DIR + XDG_DATA_HOME/XDG_CACHE_HOME)rm -rf "$NUB_CACHE_DIR" "$XDG_DATA_HOME" "$XDG_CACHE_HOME"
npm~/.npm/_cacache (or --cache <dir>)rm -rf <cache> (the --cache dir you pass)
pnpmpnpm store path (or --store-dir <dir>)rm -rf <store> (or pnpm store prune)
bunbun pm cache = the real ~/.bun/install/cacherm -rf "$(bun pm cache)" — env vars do NOT relocate it

VERIFY each cold is genuine — the offline-fails check

After wiping a cache, an --offline install MUST FAIL. A pass means the cache wasn't cleared (wrong disk path), and any "cold" number from it is a warm-link artifact. Run this for every tool before trusting a cold number:

toolexpected after cache wipegenuine?
nubrc≠0, "not available in the local cache"FAIL → genuine
pnpmrc=1, ERR_PNPM_NO_OFFLINE_TARBALLFAIL → genuine
npmrc≠0, cache-miss errorFAIL → genuine
bunif bun install --offline SUCCEEDS → cache NOT cleared → bun cold is NOT genuine

A true bun cold needs a clean container, or wiping the user's real cache (destructive — don't, unless in Docker).

Apples-to-apples isolation

  • Each tool gets its OWN cache dir.
  • --ignore-scripts for ALL tools (or --allow-scripts for all — same on both sides).
  • Identical fixture, identical lockfile throughout (use a --frozen-lockfile/--frozen equivalent where the tool offers one).
  • Interleave tool order round-robin (nub → npm → bun → pnpm, repeat) — never all-of-one-then-the-other — so drift in host load hits every tool equally.
Show full SKILL.md (361 more words)Show less

Load discipline — measure only on a quiet machine

  • Check the load average before measuring and only proceed when the machine is quiet. A dedicated quiet box or a CI runner near zero load is ideal for anything that will be published.
  • On a shared dev host, WAIT for load to fall below a threshold — don't measure above it. A practical gate is a 1-minute load average under ~40 (pick a threshold the machine actually reaches; below ~5 is ideal on a dedicated box). Poll, wait, then run; abort and retry if it spikes mid-run.
  • Report the load that held during the run alongside the numbers.
  • File counts and store-entry counts are exact and load-independent — a robust cross-check, and the clearest way to tell the dedup story. They complement a properly-measured time, not replace it.

File-count forensics (the load-independent crux)

sh
find node_modules -type f -o -type l | wc -l          # total materialized entries
ls -d node_modules/**/core-js node_modules/core-js    # physical copies of a duplicated dep (dedup story)

Reach for this first — exact, reproducible, immune to load.

The honesty bar (anti-juicing)

  • VERIFY every cold is genuine (offline-fails check) BEFORE citing it.
  • NEVER compare one tool's genuine-cold to another's warm-link — the exact misleading comparison the offline check exists to prevent.
  • Report what was actually measured, caveats included (which colds are genuine, host load, sample size).
  • The homepage cites the WARM number because it is the honest, reproducible one.
  • A single non-genuine cell discredits the whole table. When in doubt, exclude the cell and say why.

Process hygiene (this runs on the maintainer's machine)

Short-lived install/measurement processes reap themselves. The hazard is a long-lived process a bench starts and forgets:

  • Never leave a dev server running. If a bench starts one, tear it down in the same run (trap '…kill…' EXIT INT TERM).
  • Docker: docker run --rm, and confirm docker ps is empty when done.

Reference template

/tmp/cs-bench-final.sh is a working 4-tool harness (nub / npm / bun / pnpm) — per-tool isolated cache dirs, --ignore-scripts for all, interleaved round-robin order, node_modules wiped before every run, median/spread helpers. Read its structure before hand-rolling a new one; adapt the cache paths and fixture, keep the protocol.

Internal decomposition

When a comparative number raises "WHY is nub's phase X slow?", switch to pm-perf-tracing: RUST_LOG=debug nub install for the phase:resolve/fetch/link split, and the gated AUBE_DIAG_FILE per-file linker strategy tally.

© nubjs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/benchmarking of nubjs/nub.

Open the folder on GitHubat commit 568e73a

Compare with similar skills

Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmarking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmarking this skillnubjs/nub4.4k—~1.8kAutomated safety check: PassMIT
Testing Changespnpm/pnpm37k—~1.1kAutomated safety check: PassMIT
Monorepo Tooling and Dependenciespierrecomputer/pierre6.2k—~1.1kAutomated safety check: PassApache-2.0
Building Farm From Sourcefarm-fe/farm5.6k—~857Automated safety check: PassMIT
Farm Ready Gatefarm-fe/farm5.6k—~875Automated safety check: PassMIT
Interlinked Supply ChainQuentinCody/interlinked-cli178—~2.8kAutomated safety check: PassMIT

Similar skills

  • Run the tests that cover a change in the pnpm repository, in the Rust workspace (pnpm/, pnpr/) or the TypeScript CLI (pnpm11/), and recognize the cases where a scoped run passes without testing…

    37k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Sets one monorepo's rules for toolchain pins, pnpm package operations, the shared dependency catalog and moon tasks, so the agent adds versions and scripts the right way.

    6.2k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • A skill your agent uses when building Farm locally from source, native bindings or Rust plugin artifacts are missing, or an example/E2E needs local core and plugin artifacts.

    5.6k GitHub stars~857 tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • Farm Ready Gate

    farm-fe/farm

    Run Farm verification with build-first constraints. An agent skill from farm-fe/farm.

    5.6k GitHub stars~875 tokensUpdated 15 days ago
    DevelopmentAuto-check passed
  • Interlinked Supply Chain

    QuentinCody/interlinked-cli

    Respond to blocked package installs and manage the Interlinked supply-chain allowlist.

    178 GitHub stars~2.8k tokensUpdated 5 days ago
    SecurityAuto-check passed
  • Uv Workflow

    Aedelon/claude-code-blueprint

    Master uv package manager for Python: project setup, dependency management, virtual environments, lockfiles, CI/CD integration, Docker builds, and migration from pip/poetry.

    120 GitHub stars~1.2k tokensUpdated 7 mo ago
    DevOps & CloudAuto-check: notes

More from nubjs/nub

All 31 skills in this repo
  • Cpu Reduction

    nubjs/nub

    Diagnose and clear CPU, memory, and disk contention on the maintainer's dev host.

    4.4k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Reclaim disk on the maintainer's Mac when the volume is full or filling — ENOSPC, "no space left on device", a failed build or agent harness, or a routine sweep of Rust build residue.

    4.4k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Nub Charts

    nubjs/nub

    Build a performance chart for nubjs.com — the SVG bar figures in blog posts, docs pages and social posts (a runtime augmentation against plain node, an install or dispatch comparison, a cross-tool…

    4.4k GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Audit Thread

    nubjs/nub

    A skill your agent uses when running a compatibility/parity AUDIT — enumerating where nub diverges from a reference it claims parity with (pnpm CLI grammar, a lockfile format, a Node behavior, a…

    4.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Linux Vm Test

    nubjs/nub

    Run ad-hoc Nub tests and debugging probes on real local Linux guests.

    4.4k GitHub stars~986 tokensUpdated today
    Auto-check passed
  • Performance-trace Nub package-manager installs using the existing phase timings, structured diagnostics, and sampling-profiler workflow.

    4.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Benchmarking

What does Benchmarking do?

Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar. Benchmarking is an agent skill from nubjs/nub. Comparative install-benchmarking methodology for nub vs npm/pnpm/bun — cold/warm protocol, genuine-cold cache isolation, load-robust measurement, and the anti-juicing honesty bar.

How do I install Benchmarking in Claude Code?

Run `npx skills add nubjs/nub --skill benchmarking -a claude-code`. Or copy the skill folder (.claude/skills/benchmarking in nubjs/nub) into .claude/skills/benchmarking in your project. Claude Code loads it when a task matches its description.

How do I install Benchmarking in Codex?

Run `npx skills add nubjs/nub --skill benchmarking -a codex`. Or copy the skill folder (.claude/skills/benchmarking in nubjs/nub) into .agents/skills/benchmarking in your project. Codex loads it when a task matches its description.

Can I use Benchmarking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nubjs/nub --skill benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking, .gemini/skills/benchmarking, .github/skills/benchmarking and .opencode/skills/benchmarking in your project.

What does Benchmarking need to run?

Going by SKILL.md and its folder, Benchmarking needs the command-line tools its instructions call (bun, pnpm and docker). Our summary lists: Docker.

Does Benchmarking access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmarking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmarking use?

Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmarking use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmarking?

Skills that share tags, products or a category with Benchmarking: Testing Changes (pnpm/pnpm, 37k stars), Monorepo Tooling and Dependencies (pierrecomputer/pierre, 6.2k stars), Building Farm From Source (farm-fe/farm, 5.6k stars) and Farm Ready Gate (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmarking?

nubjs (a GitHub organization) maintains it in nubjs/nub, which has 4,370 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: nubjs/nub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.