Finishing a Development Branch
obra/superpowers
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
Compare benchmark performance between two git worktrees (or the current worktree vs main).
$ npx skills add ewhauser/shuck --skill bench-compare -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ewhauser/shuck bench-compare --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bench-compare .claude/skills/bench-compare && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .claude/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compareType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ewhauser/shuck --skill bench-compare -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ewhauser/shuck bench-compare --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/bench-compare .agents/skills/bench-compare && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .agents/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ewhauser/shuck --skill bench-compare -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ewhauser/shuck bench-compare --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/bench-compare .cursor/skills/bench-compare && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .cursor/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ewhauser/shuck.git --path .claude/skills/bench-compare--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ewhauser/shuck --skill bench-compare -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ewhauser/shuck bench-compare --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/bench-compare .gemini/skills/bench-compare && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .gemini/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ewhauser/shuck bench-compareInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ewhauser/shuck --skill bench-compare -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/bench-compare .github/skills/bench-compare && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .github/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ewhauser/shuck --skill bench-compare -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ewhauser/shuck bench-compare --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ewhauser/shuck.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/bench-compare .opencode/skills/bench-compare && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bench-compare" agent skill from https://github.com/ewhauser/shuck/tree/main/.claude/skills/bench-compare into .opencode/skills/bench-compare/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bench-compare", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bench-compareCompare benchmark performance between two git worktrees (or the current worktree vs main).
Bench Compare is an agent skill from ewhauser/shuck. Compare benchmark performance between two git worktrees (or the current worktree vs main). Runs Criterion microbenchmarks and hyperfine macrobenchmarks, extracts deltas, and drills down into regressions. Use this skill whenever the user asks to compare benchmarks, check for performance regressions, benchmark their branch against main, run a perf comparison, or says things like "bench compare", "any regressions?", "compare perf", "how does this branch perform", "run benchmarks against main". Even if the user just…
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Git worktrees. The repository describes itself as: A lightning fast shell linter/formatter/LSP server with zsh support. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 904974e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cargonixgitshellcheckFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bench Compare loads about 2.4k tokens when it runs. Until then it costs about 154 tokens; SKILL.md has 753 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ewhauser/shuck at commit 904974e, republished under its MIT licence (© ewhauser). 753 words, ~2,432 tokens.
.claude/skills/bench-compare/SKILL.md (or your agent's skills folder).Compare shuck's benchmark performance between two worktrees (typically the current feature worktree and the main worktree). Produces a delta report covering both Criterion microbenchmarks (lexer, parser, semantic, linter) and hyperfine macrobenchmarks (wall-clock comparison against ShellCheck).
Identify the two worktrees to compare:
git worktree listThis gives you:
/Users/.../shuckIf both are the same directory, you'll need to stash or commit changes and use
--save-baseline / --baseline on the same tree. But the typical case is two
separate worktrees.
Also check what benchmark targets exist — the registered Criterion benches may change over time:
grep -A1 '^\[\[bench\]\]' crates/shuck-benchmark/Cargo.toml | grep '^name'Both worktrees must write Criterion artifacts to the same target/criterion/
directory for baseline comparison to work. Create a temp directory and point
CARGO_TARGET_DIR at it.
SCRATCH=$(mktemp -d "${TMPDIR:-/tmp}/shuck-bench-compare.XXXXXX")
export CARGO_TARGET_DIR="$SCRATCH/target"
echo "Scratch: $SCRATCH"Also create subdirectories for macrobenchmark exports:
mkdir -p "$SCRATCH/main-macro" "$SCRATCH/current-macro"cd into the main worktree and run Criterion with --save-baseline=main.
Important: Do NOT use bare cargo bench -p shuck-benchmark -- --save-baseline=main.
That trips over the crate lib harness and fails. You must enumerate each bench target
explicitly with --bench:
cd /path/to/main/worktree
env CARGO_TARGET_DIR="$SCRATCH/target" \
cargo bench -p shuck-benchmark \
--bench lexer --bench lexer_hot_path --bench parser --bench semantic --bench linter \
-- --save-baseline=main --noplot \
> "$SCRATCH/main-criterion.log" 2>&1Monitor progress by tailing the log. This typically takes 3-8 minutes depending on the machine.
If additional bench targets exist (check Cargo.toml), add them to the --bench list.
cd into the current worktree and run Criterion with --baseline=main (note: not
--save-baseline). This compares against the baseline saved in Step 2.
cd /path/to/current/worktree
env CARGO_TARGET_DIR="$SCRATCH/target" \
cargo bench -p shuck-benchmark \
--bench lexer --bench lexer_hot_path --bench parser --bench semantic --bench linter \
-- --baseline=main --noplot \
> "$SCRATCH/current-criterion.log" 2>&1Macrobenchmarks use hyperfine via scripts/benchmarks/run.sh, which requires
hyperfine and shellcheck — tools only available inside the nix dev shell.
For each worktree (main first, then current):
cd /path/to/worktree
# Clear stale bench exports to avoid mixing results
rm -f .cache/bench-*.json .cache/bench-*.md 2>/dev/null || true
# Build and verify deps
nix --extra-experimental-features 'nix-command flakes' develop --command \
./scripts/benchmarks/setup.sh > "$SCRATCH/{side}-macro-setup.log" 2>&1
# Run hyperfine comparisons
nix --extra-experimental-features 'nix-command flakes' develop --command \
./scripts/benchmarks/run.sh > "$SCRATCH/{side}-macro.log" 2>&1
# Copy exports to scratch
cp .cache/bench-*.json "$SCRATCH/{side}-macro/"Replace {side} with main or current as appropriate.
Criterion stores change estimates at:
$CARGO_TARGET_DIR/criterion/{group}/{case}/change/estimates.json
Extract them with a Python script:
import json, pathlib
ROOT = pathlib.Path("$SCRATCH")
crit = ROOT / "target" / "criterion"
rows = []
for group in sorted(p for p in crit.iterdir() if p.is_dir()):
for case in sorted(p for p in group.iterdir() if p.is_dir()):
try:
base = json.loads((case / "main" / "estimates.json").read_text())["mean"]["point_estimate"]
new = json.loads((case / "new" / "estimates.json").read_text())["mean"]["point_estimate"]
change = json.loads((case / "change" / "estimates.json").read_text())["mean"]["point_estimate"] * 100
except FileNotFoundError:
continue
rows.append((group.name, case.name, base, new, change))
print("=== Criterion Summary (group-level) ===")
for group, case, base, new, change in rows:
if case == "all":
sign = "+" if change > 0 else ""
print(f" {group:30s} {base:12.0f} → {new:12.0f} ns ({sign}{change:.2f}%)")
print("\n=== Top Regressions ===")
for group, case, base, new, change in sorted(rows, key=lambda r: r[4], reverse=True)[:10]:
print(f" {group}/{case:20s} {change:+.2f}%")
print("\n=== Top Improvements ===")
for group, case, base, new, change in sorted(rows, key=lambda r: r[4])[:10]:
print(f" {group}/{case:20s} {change:+.2f}%")Hyperfine exports JSON with a results array. Each result has command and mean
(in seconds). Compare the shuck entries between main and current:
for p in sorted((ROOT / "main-macro").glob("bench-*.json")):
name = p.stem.removeprefix("bench-")
main_results = json.loads(p.read_text())["results"]
cur_results = json.loads((ROOT / "current-macro" / p.name).read_text())["results"]
m_shuck = next(r for r in main_results if r["command"].startswith("shuck/"))
c_shuck = next(r for r in cur_results if r["command"].startswith("shuck/"))
change = (c_shuck["mean"] / m_shuck["mean"] - 1) * 100
print(f" {name:30s} {m_shuck['mean']*1000:.1f} → {c_shuck['mean']*1000:.1f} ms ({change:+.2f}%)")Present a summary table to the user covering:
If no regressions are found, report that and stop.
If a significant regression is found (>3% in Criterion or >5% in macrobenchmarks), investigate the root cause:
The Criterion benchmarks isolate stages: lexer, parser, semantic, linter. If only
linter regressed but parser and semantic are flat, the regression is in the
linting pipeline (checker, suppression, directives).
For finer-grained measurement, create a temporary
crates/shuck-benchmark/examples/lint_breakdown.rs that measures individual pipeline
stages:
Run it on both worktrees with cargo run -p shuck-benchmark --release --example lint_breakdown
and compare the per-stage timings. This isolates which stage accounts for the
regression.
Once you've narrowed to a stage, diff the relevant source files between worktrees:
git diff --no-index /path/to/main/crates/... /path/to/current/crates/...Look for:
.clone() calls on hot pathsTell the user:
Remove any temporary breakdown harness files after the investigation.
Explicit --bench flags: Always enumerate bench targets individually
(--bench lexer --bench parser ...). The bare cargo bench -p shuck-benchmark
form trips over the crate lib harness when passing Criterion flags like
--save-baseline.
Macrobenchmarks need nix: hyperfine and shellcheck are only available
inside nix develop. Always run macro setup and benchmarks through
nix --extra-experimental-features 'nix-command flakes' develop --command ....
Shared CARGO_TARGET_DIR: Both worktrees must use the same target directory for Criterion baseline comparison to work. This is the whole reason for the scratch directory.
Suppression-heavy fixtures: ruby-build.sh and nvm.sh have many
shellcheck disable= comments. Regressions in the directive/suppression parsing
path show up disproportionately on these fixtures.
zsh glob errors: When clearing .cache/bench-*.json, use
rm -f ... 2>/dev/null || true because zsh complains when no files match a glob.
Benchmark noise: Criterion microbenchmarks on a laptop can have 2-3% noise. Don't chase regressions under 3% unless they're consistent across all fixtures. Macrobenchmarks (hyperfine) are more stable since they measure full CLI invocations.
© ewhauser, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/bench-compare of ewhauser/shuck.
Open the folder on GitHubat commit 904974e
Bench Compare next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bench Compare this skillewhauser/shuck | 136 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Finishing a Development Branchobra/superpowers | 296k | 5 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Migrate Core Code to Submodulestinyhumansai/openhuman | 41k | — | ~2.6k | Automated safety check: Pass | GPL-3.0 | |
| Finishing A Development Branchfarm-fe/farm | 5.6k | 33 repos | ~1.8k | Automated safety check: Pass | MIT | |
| Git Worktree Cleanuplobehub/lobehub | 83k | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| Keep Codex Fastvibeforge1111/keep-codex-fast | 1.6k | — | ~3.1k | Automated safety check: Pass | MIT |
obra/superpowers
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
tinyhumansai/openhuman
Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.
farm-fe/farm
A skill your agent uses when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for…
lobehub/lobehub
Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.
vibeforge1111/keep-codex-fast
A skill your agent uses when Codex feels slow or bloated, when local sessions/logs/worktrees/config have grown over time, or when a user wants safe maintenance for Codex Desktop/CLI state.
jamiepine/voicebox
Sorts a backlog of open pull requests into must-merge, candidate, superseded and deferred, writes a triage doc and works the merge loop before a release.
ewhauser/shuck
Profile shuck scripts and large-corpus fixtures, especially requests to profile a corpus script/fixture, reprofile after a shuck performance change, or produce a hotspot table from a samply profile.
ewhauser/shuck
Verify ShellCheck conformance for a shuck rule by running the large corpus test, analyzing deltas, and producing a structured bug document in docs/bugs/.
ewhauser/shuck
Fix a shuck lint rule that has conformance deltas against ShellCheck.
ewhauser/shuck
Implement an autofix for an existing shuck-rs lint rule. An agent skill from ewhauser/shuck.
ewhauser/shuck
Implement a shuck-rs lint rule from its YAML definition in docs/rules/.
ewhauser/shuck
Write and update technical design specifications. An agent skill from ewhauser/shuck.
Categories
Compare benchmark performance between two git worktrees (or the current worktree vs main). Bench Compare is an agent skill from ewhauser/shuck. Compare benchmark performance between two git worktrees (or the current worktree vs main).
Bench Compare fits situations like: the user asks to compare benchmarks; check for performance regressions; benchmark their branch against main; run a perf comparison.
Run `npx skills add ewhauser/shuck --skill bench-compare -a claude-code`. Or copy the skill folder (.claude/skills/bench-compare in ewhauser/shuck) into .claude/skills/bench-compare in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ewhauser/shuck --skill bench-compare -a codex`. Or copy the skill folder (.claude/skills/bench-compare in ewhauser/shuck) into .agents/skills/bench-compare in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ewhauser/shuck --skill bench-compare -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bench-compare, .gemini/skills/bench-compare, .github/skills/bench-compare and .opencode/skills/bench-compare in your project.
Going by SKILL.md and its folder, Bench Compare needs the command-line tools its instructions call (cargo, nix, git and shellcheck). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bench Compare is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bench Compare: Finishing a Development Branch (obra/superpowers, 296k stars), Migrate Core Code to Submodules (tinyhumansai/openhuman, 41k stars), Finishing A Development Branch (farm-fe/farm, 5.6k stars) and Git Worktree Cleanup (lobehub/lobehub, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ewhauser (a GitHub user) maintains it in ewhauser/shuck, which has 136 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 5, 2026.
Source: ewhauser/shuck on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.