Agent skill

Bench Compare

by ewhauser in ewhauser/shuck

Compare benchmark performance between two git worktrees (or the current worktree vs main).

MITAuto-check passedDevelopment

Install Bench Compare

skills CLI
$ npx skills add ewhauser/shuck --skill bench-compare -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ewhauser/shuck bench-compare --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bench-compare .claude/skills/bench-compare && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bench-compare
GitHub stars
136
Token cost
~2.4k tokens
SKILL.md length
753 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Compare benchmark performance between two git worktrees (or the current worktree vs main).

  • Works in 7 steps: Create a shared scratch directory → Run Criterion baseline on main → Run Criterion comparison on current… → …
  • The user asks to compare benchmarks
  • SKILL.md covers Before you start, Step 1: Create a shared…, Step 2: Run Criterion baseline… and Step 3: Run Criterion…, plus 5 more sections
  • Calls cargo, nix and git

What it does

Bench Compare is an agent skill from ewhauser/shuck. Compare benchmark performance between two git worktrees (or the current worktree vs main). Runs Criterion microbenchmarks and hyperfine macrobenchmarks, extracts deltas, and drills down into regressions. Use this skill whenever the user asks to compare benchmarks, check for performance regressions, benchmark their branch against main, run a perf comparison, or says things like "bench compare", "any regressions?", "compare perf", "how does this branch perform", "run benchmarks against main". Even if the user just…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Git worktrees. The repository describes itself as: A lightning fast shell linter/formatter/LSP server with zsh support. The licence is MIT.

When your agent uses it

  • The user asks to compare benchmarks
  • Check for performance regressions
  • Benchmark their branch against main
  • Run a perf comparison

Example prompts

  • “bench compare”
  • “any regressions?”
  • “compare perf”
  • “/bench-compare”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Create a shared scratch directory
  2. Run Criterion baseline on main
  3. Run Criterion comparison on current worktree
  4. Run macrobenchmarks on both worktrees
  5. Extract and report deltas
  6. Interpret the results
  7. Drill down into regressions

What it can do on your machine

Read from SKILL.md and the folder at commit 904974e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • nix
    • git
    • shellcheck

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bench Compare loads about 2.4k tokens when it runs. Until then it costs about 154 tokens; SKILL.md has 753 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~154
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ewhauser/shuck at commit 904974e, republished under its MIT licence (© ewhauser). 753 words, ~2,432 tokens.

Download SKILL.mdSave it as .claude/skills/bench-compare/SKILL.md (or your agent's skills folder).
name
bench-compare
description
Compare benchmark performance between two git worktrees (or the current worktree vs main). Runs Criterion microbenchmarks and hyperfine macrobenchmarks, extracts deltas, and drills down into regressions. Use this skill whenever the user asks to compare benchmarks, check for performance regressions, benchmark their branch against main, run a perf comparison, or says things like "bench compare", "any regressions?", "compare perf", "how does this branch perform", "run benchmarks against main". Even if the user just says "benchmarks" in the context of a feature branch or worktree, this skill applies.

Benchmark Comparison

Compare shuck's benchmark performance between two worktrees (typically the current feature worktree and the main worktree). Produces a delta report covering both Criterion microbenchmarks (lexer, parser, semantic, linter) and hyperfine macrobenchmarks (wall-clock comparison against ShellCheck).

Before you start

Identify the two worktrees to compare:

bash
git worktree list

This gives you:

  • Main worktree — the first entry, typically /Users/.../shuck
  • Current worktree — the one you're running in (may be the same as main)

If both are the same directory, you'll need to stash or commit changes and use --save-baseline / --baseline on the same tree. But the typical case is two separate worktrees.

Also check what benchmark targets exist — the registered Criterion benches may change over time:

bash
grep -A1 '^\[\[bench\]\]' crates/shuck-benchmark/Cargo.toml | grep '^name'

Step 1: Create a shared scratch directory

Both worktrees must write Criterion artifacts to the same target/criterion/ directory for baseline comparison to work. Create a temp directory and point CARGO_TARGET_DIR at it.

bash
SCRATCH=$(mktemp -d "${TMPDIR:-/tmp}/shuck-bench-compare.XXXXXX")
export CARGO_TARGET_DIR="$SCRATCH/target"
echo "Scratch: $SCRATCH"

Also create subdirectories for macrobenchmark exports:

bash
mkdir -p "$SCRATCH/main-macro" "$SCRATCH/current-macro"

Step 2: Run Criterion baseline on main

cd into the main worktree and run Criterion with --save-baseline=main.

Important: Do NOT use bare cargo bench -p shuck-benchmark -- --save-baseline=main. That trips over the crate lib harness and fails. You must enumerate each bench target explicitly with --bench:

bash
cd /path/to/main/worktree
env CARGO_TARGET_DIR="$SCRATCH/target" \
  cargo bench -p shuck-benchmark \
    --bench lexer --bench lexer_hot_path --bench parser --bench semantic --bench linter \
    -- --save-baseline=main --noplot \
  > "$SCRATCH/main-criterion.log" 2>&1

Monitor progress by tailing the log. This typically takes 3-8 minutes depending on the machine.

If additional bench targets exist (check Cargo.toml), add them to the --bench list.

Step 3: Run Criterion comparison on current worktree

cd into the current worktree and run Criterion with --baseline=main (note: not --save-baseline). This compares against the baseline saved in Step 2.

bash
cd /path/to/current/worktree
env CARGO_TARGET_DIR="$SCRATCH/target" \
  cargo bench -p shuck-benchmark \
    --bench lexer --bench lexer_hot_path --bench parser --bench semantic --bench linter \
    -- --baseline=main --noplot \
  > "$SCRATCH/current-criterion.log" 2>&1

Step 4: Run macrobenchmarks on both worktrees

Macrobenchmarks use hyperfine via scripts/benchmarks/run.sh, which requires hyperfine and shellcheck — tools only available inside the nix dev shell.

For each worktree (main first, then current):

bash
cd /path/to/worktree

# Clear stale bench exports to avoid mixing results
rm -f .cache/bench-*.json .cache/bench-*.md 2>/dev/null || true

# Build and verify deps
nix --extra-experimental-features 'nix-command flakes' develop --command \
  ./scripts/benchmarks/setup.sh > "$SCRATCH/{side}-macro-setup.log" 2>&1

# Run hyperfine comparisons
nix --extra-experimental-features 'nix-command flakes' develop --command \
  ./scripts/benchmarks/run.sh > "$SCRATCH/{side}-macro.log" 2>&1

# Copy exports to scratch
cp .cache/bench-*.json "$SCRATCH/{side}-macro/"

Replace {side} with main or current as appropriate.

Step 5: Extract and report deltas

Criterion deltas

Criterion stores change estimates at: $CARGO_TARGET_DIR/criterion/{group}/{case}/change/estimates.json

Extract them with a Python script:

python
import json, pathlib

ROOT = pathlib.Path("$SCRATCH")
crit = ROOT / "target" / "criterion"
rows = []

for group in sorted(p for p in crit.iterdir() if p.is_dir()):
    for case in sorted(p for p in group.iterdir() if p.is_dir()):
        try:
            base = json.loads((case / "main" / "estimates.json").read_text())["mean"]["point_estimate"]
            new = json.loads((case / "new" / "estimates.json").read_text())["mean"]["point_estimate"]
            change = json.loads((case / "change" / "estimates.json").read_text())["mean"]["point_estimate"] * 100
        except FileNotFoundError:
            continue
        rows.append((group.name, case.name, base, new, change))

print("=== Criterion Summary (group-level) ===")
for group, case, base, new, change in rows:
    if case == "all":
        sign = "+" if change > 0 else ""
        print(f"  {group:30s}  {base:12.0f} → {new:12.0f} ns  ({sign}{change:.2f}%)")

print("\n=== Top Regressions ===")
for group, case, base, new, change in sorted(rows, key=lambda r: r[4], reverse=True)[:10]:
    print(f"  {group}/{case:20s}  {change:+.2f}%")

print("\n=== Top Improvements ===")
for group, case, base, new, change in sorted(rows, key=lambda r: r[4])[:10]:
    print(f"  {group}/{case:20s}  {change:+.2f}%")
Macrobenchmark deltas

Hyperfine exports JSON with a results array. Each result has command and mean (in seconds). Compare the shuck entries between main and current:

python
for p in sorted((ROOT / "main-macro").glob("bench-*.json")):
    name = p.stem.removeprefix("bench-")
    main_results = json.loads(p.read_text())["results"]
    cur_results = json.loads((ROOT / "current-macro" / p.name).read_text())["results"]
    m_shuck = next(r for r in main_results if r["command"].startswith("shuck/"))
    c_shuck = next(r for r in cur_results if r["command"].startswith("shuck/"))
    change = (c_shuck["mean"] / m_shuck["mean"] - 1) * 100
    print(f"  {name:30s}  {m_shuck['mean']*1000:.1f} → {c_shuck['mean']*1000:.1f} ms  ({change:+.2f}%)")

Step 6: Interpret the results

Present a summary table to the user covering:

  • Each Criterion benchmark group with % change
  • Each macrobenchmark fixture with % change
  • Highlight anything above +3% as a potential regression (Criterion has noise, so small changes are usually not meaningful)

If no regressions are found, report that and stop.

Step 7: Drill down into regressions

If a significant regression is found (>3% in Criterion or >5% in macrobenchmarks), investigate the root cause:

7a. Identify which component regressed

The Criterion benchmarks isolate stages: lexer, parser, semantic, linter. If only linter regressed but parser and semantic are flat, the regression is in the linting pipeline (checker, suppression, directives).

Show full SKILL.md (307 more words)Show less
7b. Build a breakdown harness (if needed)

For finer-grained measurement, create a temporary crates/shuck-benchmark/examples/lint_breakdown.rs that measures individual pipeline stages:

  1. Parse only
  2. Indexer only
  3. Directives/suppression only
  4. Semantic model build only
  5. Lint with empty rules
  6. Lint with default rules

Run it on both worktrees with cargo run -p shuck-benchmark --release --example lint_breakdown and compare the per-stage timings. This isolates which stage accounts for the regression.

7c. Diff the suspected code

Once you've narrowed to a stage, diff the relevant source files between worktrees:

bash
git diff --no-index /path/to/main/crates/... /path/to/current/crates/...

Look for:

  • New .clone() calls on hot paths
  • Recursive traversals that weren't there before
  • Changed data structures (Vec → HashMap, etc.)
  • New allocations in tight loops
7d. Report findings

Tell the user:

  • Which benchmark regressed and by how much
  • Which code change caused it (with file path and description)
  • Whether you have a fix suggestion

Remove any temporary breakdown harness files after the investigation.

Gotchas

  • Explicit --bench flags: Always enumerate bench targets individually (--bench lexer --bench parser ...). The bare cargo bench -p shuck-benchmark form trips over the crate lib harness when passing Criterion flags like --save-baseline.

  • Macrobenchmarks need nix: hyperfine and shellcheck are only available inside nix develop. Always run macro setup and benchmarks through nix --extra-experimental-features 'nix-command flakes' develop --command ....

  • Shared CARGO_TARGET_DIR: Both worktrees must use the same target directory for Criterion baseline comparison to work. This is the whole reason for the scratch directory.

  • Suppression-heavy fixtures: ruby-build.sh and nvm.sh have many shellcheck disable= comments. Regressions in the directive/suppression parsing path show up disproportionately on these fixtures.

  • zsh glob errors: When clearing .cache/bench-*.json, use rm -f ... 2>/dev/null || true because zsh complains when no files match a glob.

  • Benchmark noise: Criterion microbenchmarks on a laptop can have 2-3% noise. Don't chase regressions under 3% unless they're consistent across all fixtures. Macrobenchmarks (hyperfine) are more stable since they measure full CLI invocations.

© ewhauser, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/bench-compare of ewhauser/shuck.

Open the folder on GitHubat commit 904974e

Compare with similar skills

Bench Compare next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bench Compare compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bench Compare this skillewhauser/shuck136—~2.4kAutomated safety check: PassMIT
Finishing a Development Branchobra/superpowers296k5 repos~1.9kAutomated safety check: PassMIT
Migrate Core Code to Submodulestinyhumansai/openhuman41k—~2.6kAutomated safety check: PassGPL-3.0
Finishing A Development Branchfarm-fe/farm5.6k33 repos~1.8kAutomated safety check: PassMIT
Git Worktree Cleanuplobehub/lobehub83k—~2.8kAutomated safety check: PassCustom licence
Keep Codex Fastvibeforge1111/keep-codex-fast1.6k—~3.1kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    296k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Migrate Core Code to Submodules

    tinyhumansai/openhuman

    Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.

    41k GitHub stars~2.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • A skill your agent uses when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for…

    5.6k GitHub starsUsed in 33 repos~1.8k tokens
    DevelopmentAuto-check passed
  • Git Worktree Cleanup

    lobehub/lobehub

    Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.

    83k GitHub stars~2.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Keep Codex Fast

    vibeforge1111/keep-codex-fast

    A skill your agent uses when Codex feels slow or bloated, when local sessions/logs/worktrees/config have grown over time, or when a user wants safe maintenance for Codex Desktop/CLI state.

    1.6k GitHub stars~3.1k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Pre-Release PR Triage

    jamiepine/voicebox

    Sorts a backlog of open pull requests into must-merge, candidate, superseded and deferred, writes a triage doc and works the merge loop before a release.

    57k GitHub stars~3.1k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from ewhauser/shuck

All 8 skills in this repo
  • Profile Shuck Script

    ewhauser/shuck

    Profile shuck scripts and large-corpus fixtures, especially requests to profile a corpus script/fixture, reprofile after a shuck performance change, or produce a hotspot table from a samply profile.

    136 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Conformance Check

    ewhauser/shuck

    Verify ShellCheck conformance for a shuck rule by running the large corpus test, analyzing deltas, and producing a structured bug document in docs/bugs/.

    136 GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Fix Rule

    ewhauser/shuck

    Fix a shuck lint rule that has conformance deltas against ShellCheck.

    136 GitHub stars~3.2k tokensUpdated 2 days ago
    Auto-check passed
  • Implement Fix

    ewhauser/shuck

    Implement an autofix for an existing shuck-rs lint rule. An agent skill from ewhauser/shuck.

    136 GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Implement Rule

    ewhauser/shuck

    Implement a shuck-rs lint rule from its YAML definition in docs/rules/.

    136 GitHub stars~4.8k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Writer

    ewhauser/shuck

    Write and update technical design specifications. An agent skill from ewhauser/shuck.

    136 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Bench Compare

What does Bench Compare do?

Compare benchmark performance between two git worktrees (or the current worktree vs main). Bench Compare is an agent skill from ewhauser/shuck. Compare benchmark performance between two git worktrees (or the current worktree vs main).

When should I use Bench Compare?

Bench Compare fits situations like: the user asks to compare benchmarks; check for performance regressions; benchmark their branch against main; run a perf comparison.

How do I install Bench Compare in Claude Code?

Run `npx skills add ewhauser/shuck --skill bench-compare -a claude-code`. Or copy the skill folder (.claude/skills/bench-compare in ewhauser/shuck) into .claude/skills/bench-compare in your project. Claude Code loads it when a task matches its description.

How do I install Bench Compare in Codex?

Run `npx skills add ewhauser/shuck --skill bench-compare -a codex`. Or copy the skill folder (.claude/skills/bench-compare in ewhauser/shuck) into .agents/skills/bench-compare in your project. Codex loads it when a task matches its description.

Can I use Bench Compare in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ewhauser/shuck --skill bench-compare -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bench-compare, .gemini/skills/bench-compare, .github/skills/bench-compare and .opencode/skills/bench-compare in your project.

What does Bench Compare need to run?

Going by SKILL.md and its folder, Bench Compare needs the command-line tools its instructions call (cargo, nix, git and shellcheck). Our summary lists: Python 3.

Does Bench Compare access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bench Compare safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bench Compare use?

Bench Compare is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bench Compare use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bench Compare?

Skills that share tags, products or a category with Bench Compare: Finishing a Development Branch (obra/superpowers, 296k stars), Migrate Core Code to Submodules (tinyhumansai/openhuman, 41k stars), Finishing A Development Branch (farm-fe/farm, 5.6k stars) and Git Worktree Cleanup (lobehub/lobehub, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bench Compare?

ewhauser (a GitHub user) maintains it in ewhauser/shuck, which has 136 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 5, 2026.

Source: ewhauser/shuck on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.