Agent skill

Climber Step Minimization

by ben-manes in ben-manes/caffeine

Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

Apache-2.0Auto-check: notesDevelopment

Install Climber Step Minimization

skills CLI
$ npx skills add ben-manes/caffeine --skill climber-minimize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ben-manes/caffeine climber-minimize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ben-manes/caffeine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/climber-minimize .claude/skills/climber-minimize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
climber-minimize
GitHub stars
18k
Token cost
~3k tokens
SKILL.md length
1,870 words
Files
2
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

  • Works in 4 steps: Count the firings, then prove the state… → Name the shape it was for.… → Check the graveyard. §5 records what a… → …
  • Checking before a release whether any climber step has stopped paying for itself
  • SKILL.md covers How to run, The arms, Reading the result, and the… and The 2026-08-23 baseline…, plus 1 more section
  • Runs Python scripts from its folder; calls python3 and git

What it does

Other climber skills judge whether a change is good; this one asks whether a step should exist at all. It removes one algorithmic step at a time through a named arm (for example nostarve, noladder, noscale, nocommit, norepeat, nowedge, nofollow or noshield) and reports what the system loses. The reason is that the climber grows by repair, so recorded step prices go stale and branches can become inert without anything failing. The usual result is priced and kept, with removal the rare case, and it is meant to run before a release or after several repairs in a row.

Arms live in climber-gate/harness.py, so you first create a detached git worktree for a commit and apply the harness to it. ablate.py then runs against a directory of traces with a preset: quick is a smoke run of about 130 runs, and standard, which covers the constructed families plus the whole real corpus, is the preset a keep-or-prune decision needs. Arms are rotated inside each seed so all see the same machine state and the same admission draws. The skill sits in the Caffeine caching library's repo and may use Read, Grep, Glob, Bash and Write.

When your agent uses it

  • Checking before a release whether any climber step has stopped paying for itself
  • Re-pricing steps after several repairs in a row have changed the terrain
  • Finding branches that never fire anymore

Example prompts

  • “Run the climber minimization quick preset on the current commit and tell me which steps look inert.”
  • “Price the noscale and noladder arms with the standard preset before the release.”
  • “Which climber steps look stale after the last several repairs?”

Requirements

  • A git worktree wired with climber-gate/harness.py
  • Python 3
  • A directory of generated traces
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Count the firings, then prove the state unreachable. ablate.py does the counting for
  2. Name the shape it was for. hill-climber.md §3 lists every family and what defeats it. A
  3. Check the graveyard. §5 records what a step replaced. A step that is inert because a later
  4. Spend a holdout on the prune, not on the decision. The battery is what the decision was

What it can do on your machine

Read from SKILL.md and the folder at commit e972fb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Climber Step Minimization loads about 3k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 1,870 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ben-manes/caffeine at commit e972fb0, republished under its Apache-2.0 licence (© ben-manes). 1,870 words, ~3,032 tokens.

Download SKILL.mdSave it as .claude/skills/climber-minimize/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
climber-minimize
description
Price each algorithmic step of the window climber by removing it, to find steps that no longer earn their keep and branches that no longer fire
allowed-tools
Read, Grep, Glob, Bash, Write
argument-hint
[quick|standard|full|<cells>]
context
fork

Climber minimization

The other climber skills ask whether a change is good. This one asks whether a step should exist. It removes one algorithmic step at a time and reports what the machine loses, so a mechanism that has stopped paying for itself can be found rather than waited for.

It exists because the machine grows by repair. Every round adds a rule that fixes a workload, and nothing in the process asks the older rules to re-justify themselves. Two things follow, and both have been observed: a step's recorded price goes stale as later repairs change the terrain it acted on (the guard rail's veto was 11:1 in 2026-08-18 and 1.86:1 on that same cell set in 2026-08-23), and a step can become inert without anyone noticing, because nothing fails when a branch stops mattering.

Its usual output is "priced, kept", not a deletion. Pricing a mechanism is the point; removing one is the rare case. Run it before a release, or after the machine has taken several repairs in a row.

How to run

The arms live in climber-gate/harness.py, so a wired worktree is required:

bash
git worktree add --detach <wt> <commit>
python3 .claude/skills/climber-gate/harness.py apply <wt>
CAF_TREE=<wt> python3 ablate.py <traces-dir> all quick 1 data/quick.csv       # smoke, ~130 runs
CAF_TREE=<wt> python3 ablate.py <traces-dir> all standard 1,2 data/std.csv    # the decision pass
CAF_TREE=<wt> python3 ablate.py <traces-dir> all full 1,2 data/full.csv       # hours

standard is the preset a keep-or-prune decision needs: the constructed families plus the whole real corpus. Traces come from climber-gate/SKILL.md's generation block. Arms are rotated inside each seed, so every arm sees the same machine state and the same admission draws.

The arms

Each is a single disable at the step's own site. noaudit and the tier arms predate this skill; most of the rest were wired 2026-08-18, noreturncover and nowidecover in 2026-08.

armthe step it removes
cornerproberestores the upper-corner probe deleted 2026-08-21, so its sign reads inverted
nostarveany blind corner arms a starvation probe
noladdera completed experiment deepens its rung
noscaledeep rungs walk 2x/4x the flat stride
nocommitdeep rungs commit the walk past the stray zone
norepeata confirm that re-finds lost ground escalates instead of rewarding
nowedgea confirm the density arm reverses escalates instead of rewarding
nofollowa park's first audit follows the walk that confirmed it
noshielda fresh park is shielded from crash-scale weather
novetothe guard rail returns the window to the anchor
noretesta veto's return re-tests the claim that sent it, on arrival
noreturncovera veto's return's landing and settle samples wait for that retest instead of standing the anchor down
nowidecovera retreat's cover runs without a held park (the 2026-08-21 widening)
nofreezean up-probe is judged against probation frozen at the arm
noauditthe whole equilibrium-audit layer

Add an arm whenever a step lands: a rule that arrives without one cannot be re-priced later, which is how the set decayed to two before 2026-08-18. harness.py's FLAGS block and the table above are the two places to edit.

Reading the result, and the trap in it

Never adjudicate on a battery mean. The mean is the one summary that hides what makes a mechanism worth keeping: a step that buys a great deal on a few workloads and costs a little on many is insurance, and insurance always looks bad on average. ablate.py reports the asymmetry instead — total gain across the cells an arm helps, total cost across the cells it hurts, the worst single row, and how many cells the arm leaves bit-identical. The audit layer's own case is the worked example: +268pp across the rows it helps against −13pp across the rows it hurts, about 21:1, on a mean of +4.91 that says much less (hill-climber.md §5).

The verdicts:

  • LOAD-BEARING — gains nowhere, losses somewhere. Keep, and update the recorded price.
  • PRICED — a trade. Report the ratio and the worst row. Most steps land here, and "priced; it stays" is a complete answer. A ratio under 1:1 is a candidate, not a verdict: the guard rail's veto reads 0.3:1 on the battery and buys 6.4pp on a planted cell.
  • NEGATIVE — removing the step helps everywhere. That is a defect report, not a simplification. Hand it to /audit-adaptivity.
  • INERT — bit-identical on every cell. This is the interesting one, and it is not a licence to delete.

Where the battery is blind: the start. Every cell in it starts the cache where the product does, at 1%. A step that defends a window the machine has been driven away from therefore has almost nothing to act on, and its ratio collapses without the mechanism changing. Before treating a sub-1:1 ratio as a candidate, re-run the step's own cells with the window planted (CAF_EXTRA=-Dcaffeine.climber.startwin=0.55, or climber-gate/startwin.py for the sweep). The guard rail's veto is the worked case: 0.3:1 on the battery, +6.4pp on mainsat planted at 55%.

Why inert is not dead. The corpus is a sample of workloads that were interesting enough for someone to capture, plus constructions aimed at defects already imagined. A step that changes nothing on that sample may still be the only thing standing between a real deployment and a 5pp hole, and nobody would ever file that bug, because a user cannot see a hit rate they did not get. So an inert arm earns a second pass, not a patch:

  1. Count the firings, then prove the state unreachable. ablate.py does the counting for you: it runs ship once per cell with -Dcaffeine.climber.counts, which dumps STEPFIRE <step>=<n> at exit, and folds the totals into the verdict. A branch that executes and changes no outcome is inert, which is weak; a branch that never executes on a narrow set is unexercised, which is weaker still, and the tool says so. norepeat never fires on demoflood and is worth 18.61 on absolve_p8, which is why the DEAD verdict is gated on the full preset and everything short of it prints PROVISIONAL.

    A zero count on the full set is still a fact about a sample, though, and a delete needs a fact about the machine. Close that gap with an argument from the gates themselves: read the conditions that guard the site and show they cannot hold together, or probe the machine's own readings for the state rather than the cells for the outcome. The worked example is external — the WaveCounter port's ADR-0055 rejected an input by probing 15,693 governor readings for the target state, finding it zero times, and then showing its two gates mutually exclusive by construction. The probe alone would have licensed the same delete on much thinner evidence.

  2. Name the shape it was for. hill-climber.md §3 lists every family and what defeats it. A step whose family is still in the list and still passes is doing its job on a cell that the preset skipped; widen the cells before concluding anything.

  3. Check the graveyard. §5 records what a step replaced. A step that is inert because a later rule subsumed it is a genuine prune; a step that is inert because its trap was retired is a prune plus a note that the trap should come back.

  4. Spend a holdout on the prune, not on the decision. The battery is what the decision was made on, so it cannot also verify it. climber-gate/SKILL.md records which holdouts are unspent.

One more asymmetry worth stating. A wrong keep costs complexity, which is visible and recoverable. A wrong prune costs hit rate on a workload nobody is measuring, which is neither. The bar for removing a step should be higher than the bar for keeping one, and this skill is built to be run often and to delete rarely.

Show full SKILL.md (631 more words)Show less

The 2026-08-23 baseline (f3fad1bdb, full preset, seeds 1 and 2)

91 cells, the whole gate battery plus the whole real corpus, at two seeds against all fourteen arms, with the firing counts from a ship run of each cell. buys is what the step is worth where it acts; spends is what it costs elsewhere.

stepfiresbuysspendsratio
the return's retest1310.000.3033:1
a repeat confirm escalates1639.311.7622:1
a park's first audit follows its walk2660.713.0020:1
the frozen probation baseline104191.2111.3917:1
the audit layer—921.9368.3713.5:1
the return and retreat cover—44.124.739.3:1
a reversed confirm escalates4371.699.417.6:1
deep rungs commit the walk8612.482.345.3:1
a blind corner arms a probe551300.3761.074.9:1
the fresh-park shield10360.7714.714.1:1
the refractory ladder192193.8249.353.9:1
deep rungs stride wider67876.0520.563.7:1
the guard rail's veto140.943.360.3:1
the deleted upper corner's probe04.329.610.4:1, restored

Nothing is dead, and nothing is inert. corner reads 0 because that step no longer exists in the tree; every other site fires and every arm moves at least one cell, so there are no free deletions to argue about.

Both 2026-08-18 candidates are closed, in opposite directions. Restoring the upper corner's probe now costs 2.2 for every 1 it returns, and 27:1 against on the 2026-08-18 cell set itself, so the 2026-08-21 deletion holds on the evidence that flagged it. The walk's commitment depth, 0.23:1 at N=8 then and the other candidate for deletion, reads 5.3:1 here.

The guard rail's veto is the one price that inverted, and it is the worked example of why the battery is not the whole answer. Re-priced on the 2026-08-18 cell set at the same seed it fell from 11:1 to 1.86:1, and over the full battery it reads 0.3:1. At N=8 across the eight cells either rail arm moves it reads 0.33:1, with 10.95 of its 12.87pp cost on sidecliff alone (−1.34 to −1.42 on all eight seeds); cp_w015 splits by basin rather than pricing anything, +0.26 on five seeds against −0.26 to −0.37 on three. But every battery cell starts where the product starts it, and a rail whose job is to return the window to an anchor has little to defend from a 1% start. Planted, it is worth +6.3 to +6.6pp on four of four seeds on mainsat at a 55% window (32.42–32.68 against noveto's 25.90–26.65) and +0.15 to +0.52 at 70%. It stays, and the finding is about the instrument: add a planted cell before running this skill again, or the rail reads deletable on a battery that never asks it to work.

A step's price can also be split. noreturncover removes the whole of isReturnTest, which is two things: the held-park retreat cover that predates 2026-08-21 and the widening that let it run without a held park. nowidecover scopes the cover back to a held park with the return half kept, and it is bit-identical on 66 of 68 rows, costing 0.07 and 0.08 on cp_w081. So the 44.12pp is the older cover and the return half, which is why the moat and hazefloor rows move under noreturncover while the commit that landed the return half read them bit-identical. When an arm removes more than one thing, split it before quoting its ratio as one step's price.

Reporting

A per-arm table with the four statistics and a verdict, the cells each arm moved, and — for anything inert — the firing count and which of the four passes above was applied. Record a mechanism's price in hill-climber.md §5 when it is priced and kept, so the next run can see whether it has gone stale.

© ben-manes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/climber-minimize of ben-manes/caffeine.

  • SKILL.md
  • ablate.py

Open the folder on GitHubat commit e972fb0

Compare with similar skills

Climber Step Minimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Climber Step Minimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Climber Step Minimization this skillben-manes/caffeine18k—~3kAutomated safety check: NotesApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT
Keybase RPC Log Analysiskeybase/client9.3k—~3kAutomated safety check: PassBSD-3-Clause
Content-Hash File Cache Patternaffaan-m/ECC276k5 repos~1.4kAutomated safety check: PassMIT
Memory Leak InterviewerPrepLabsAI/InterviewMentor112—~2.9kAutomated safety check: PassMIT
Interval Profiling Performance AnalyzerArabelaTso/Skills-4-SE253—~1.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Captures a clean Keybase service log and analyzes it for redundant, duplicated or looping RPCs, then checks whether a caching fix reduced the calls.

    9.3k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • Caches slow file processing results in Python keyed by a SHA-256 hash of the file content, so renames still hit the cache and edits invalidate it automatically.

    276k GitHub starsUsed in 5 repos~1.4k tokens
    DevelopmentAuto-check passed
  • Memory Leak Interviewer

    PrepLabsAI/InterviewMentor

    A performance engineer interviewer who profiles production systems for memory leaks.

    112 GitHub stars~2.9k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Profile programs at the function/method level to identify performance hotspots, bottlenecks, and optimization opportunities.

    253 GitHub stars~1.7k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Phy Memory Leak Detector

    LeoYeAI/openclaw-master-skills

    Static memory leak pattern scanner for Node.js, Python, Go, and Java.

    2.2k GitHub stars~4.3k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed

More from ben-manes/caffeine

All 33 skills in this repo
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Codebase Audit

    ben-manes/caffeine

    Runs a hostile review of the Caffeine Java caching library with parallel subagents that get no design docs, then challenges and consolidates their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Caffeine Performance Audit

    ben-manes/caffeine

    Audits the Caffeine cache source for hot-path costs such as allocations, contention and memory layout, reporting only findings tied to specific lines.

    18k GitHub stars~559 tokensUpdated yesterday
    Auto-check passed
  • Audit Sibling Divergence

    ben-manes/caffeine

    Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

    18k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check: notes
  • Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes

Works with

Categories

Questions about Climber Step Minimization

What does Climber Step Minimization do?

Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire. Other climber skills judge whether a change is good; this one asks whether a step should exist at all. It removes one algorithmic step at a time through a named arm (for example nostarve, noladder, noscale, nocommit, norepeat, nowedge, nofollow or noshield) and reports what the system loses.

When should I use Climber Step Minimization?

Climber Step Minimization fits situations like: checking before a release whether any climber step has stopped paying for itself; re-pricing steps after several repairs in a row have changed the terrain; finding branches that never fire anymore.

How do I install Climber Step Minimization in Claude Code?

Run `npx skills add ben-manes/caffeine --skill climber-minimize -a claude-code`. Or copy the skill folder (.claude/skills/climber-minimize in ben-manes/caffeine) into .claude/skills/climber-minimize in your project. Claude Code loads it when a task matches its description.

How do I install Climber Step Minimization in Codex?

Run `npx skills add ben-manes/caffeine --skill climber-minimize -a codex`. Or copy the skill folder (.claude/skills/climber-minimize in ben-manes/caffeine) into .agents/skills/climber-minimize in your project. Codex loads it when a task matches its description.

Can I use Climber Step Minimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ben-manes/caffeine --skill climber-minimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/climber-minimize, .gemini/skills/climber-minimize, .github/skills/climber-minimize and .opencode/skills/climber-minimize in your project.

What does Climber Step Minimization need to run?

Going by SKILL.md and its folder, Climber Step Minimization needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and git). Our summary lists: A git worktree wired with climber-gate/harness.py; Python 3; A directory of generated traces. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write.

Does Climber Step Minimization access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Climber Step Minimization safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Climber Step Minimization use?

Climber Step Minimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Climber Step Minimization use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Climber Step Minimization?

Skills that share tags, products or a category with Climber Step Minimization: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Keybase RPC Log Analysis (keybase/client, 9.3k stars), Content-Hash File Cache Pattern (affaan-m/ECC, 276k stars) and Memory Leak Interviewer (PrepLabsAI/InterviewMentor, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Climber Step Minimization?

ben-manes (a GitHub user) maintains it in ben-manes/caffeine, which has 17,882 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 9, 2026.

Source: ben-manes/caffeine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.