Agent skill

Audit Adaptivity

by ben-manes in ben-manes/caffeine

Audit the adaptive window hill-climber and region-resize logic for implementation defects (not algorithm quality)

Apache-2.0Auto-check passed

Install Audit Adaptivity

skills CLI
$ npx skills add ben-manes/caffeine --skill audit-adaptivity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ben-manes/caffeine audit-adaptivity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ben-manes/caffeine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/audit-adaptivity .claude/skills/audit-adaptivity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-adaptivity
GitHub stars
18k
Token cost
~3.1k tokens
SKILL.md length
1,391 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit the adaptive window hill-climber and region-resize logic for implementation defects (not algorithm quality)

  • Works in 7 steps: Region partition sum: windowMaximum +… → Non-negative maxima: can windowMaximum… → Quota accounting: in… → …
  • SKILL.md covers Scope boundary —…, Methods in scope, Structural invariants to… and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audit Adaptivity is an agent skill from ben-manes/caffeine. Audit the adaptive window hill-climber and region-resize logic for implementation defects (not algorithm quality)

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: A high performance caching library for Java. The licence is Apache-2.0.

Example prompts

  • “/audit-adaptivity”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Region partition sum: windowMaximum + mainMaximum (probation + protected)
  2. Non-negative maxima: can windowMaximum or mainProtectedMaximum go negative
  3. Quota accounting: in increaseWindow/decreaseWindow the quota is decremented
  4. determineAdjustment math
  5. setMaximumSize at boundaries: the step-sign flip at max <= SLOW_ADAPT_THRESHOLD
  6. Stale adjustment consumption: climb calls determineAdjustment then
  7. Layer ownership of the ladders, streaks, and schedules — the highest-yield row, because

What it can do on your machine

Read from SKILL.md and the folder at commit e972fb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Adaptivity loads about 3.1k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 1,391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ben-manes/caffeine at commit e972fb0, republished under its Apache-2.0 licence (© ben-manes). 1,391 words, ~3,060 tokens.

Download SKILL.mdSave it as .claude/skills/audit-adaptivity/SKILL.md (or your agent's skills folder).
name
audit-adaptivity
description
Audit the adaptive window hill-climber and region-resize logic for implementation defects (not algorithm quality)
context
fork
agent
auditor
disable-model-invocation
true

Audit the adaptive W-TinyLFU control loop and the admission-window / main-region resize logic for IMPLEMENTATION defects. This is the one eviction subsystem not in /audit-subsystem-safety's scope, and it was recently rebuilt — the climber now lives in the package-private WindowClimber class (reached through the generated climber field), with a new density-signal + probe-machine large tier (>4096) — so it has elevated bug density.

Scope boundary — implementation correctness, NOT algorithm quality

The adaptation policy itself is at its tuned frontier. Out of scope: convergence rate, hit-rate, oscillation-as-a-design-tradeoff, and the choice of the tuning constants. (Behavioral hit-rate regression against the adversarial trap suite is /climber-gate's job, not this audit's.) Do NOT report "the climber could converge faster / oscillates / constant X should be Y." Report only defects: arithmetic that yields a wrong value, a sign error, a state that violates a structural invariant, a race, a NaN, or an overflow.

Methods in scope

  • climb, determineAdjustment — the feedback step (reactive: hit-rate delta; density: within-sample densities → step → adjustment). determineAdjustment carries mainProtectedMaximum to size the probation capacity for the probe verdict
  • densityClimb, armProbe, walkStep, probeEnding, undoProbe, DensityClimber.steer — the large tier's signal, probe machine, and verdict internals (all on WindowClimber)
  • increaseWindow, decreaseWindow, demoteFromMainProtected — region transfer
  • setMaximumSize — initial window/main split, plus WindowClimber.resized (the SLOW_ADAPT_THRESHOLD step-sign flip and the probe-machine reset)
  • evictFromWindow / evictFromMain — they consume the region maxima the climber sets

Probe machine (>4096, in WindowClimber): starved samples (region hits below max(4, requestCount >> MIN_SIGNAL_SHIFT)) at a blind corner launch a bold-driver walk. Endings: crash-abort (full undo, refractory re-armed WITHOUT doubling), reversal-through-base and budget expiry (failed experiments: full undo + ladder x2), adjudication at >=4x the bar under the committed depth (confirm keeps the position + ladder resets to 1; anything else fails with a full undo). The verdict is asymmetric BY DESIGN: an up-probe confirms iff ln((windowDensity+eps)/(walk.baseProbationDensity+eps)) > 0 — the probation-marginal baseline FROZEN in armProbe — while a down-probe uses the average-density sign test (error*dir > 0). Ladder: PROBE_BACKOFF_INITIAL (16) doubling to _MAX (64). Invariants to audit: 1 <= starvation.rung <= 64; 0 <= refractoryLeft <= starvation.rung, which both oracles assert directly — the walk budget is a separate field (Walk.samples, bounded by PROBE_WALK_BUDGET) on an object that exists only while walking, so the two can no longer alias; probe state fully reset by resized; an undo returns exactly to walk.baseWindow; the below-floor lift cannot exceed the floor; sample.probationHits <= sample.hits - sample.windowHits; walk.baseProbationDensity is written ONLY in armProbe (each re-arm re-snapshots) and is non-negative and finite; the probation capacity denominator is max(1, maximum - windowMaximum - mainProtectedMaximum) — capacities, never occupancy.

NOT bugs (adjudicated design; read hill-climber.md §4 before flagging): the up/down verdict asymmetry (the window has no marginal substructure to price against); the frozen — hence stale-looking — baseline (judging against the LIVE probation rate is an absorbing false-veto: the walk's own demotions enrich it — the demoflood gate row; a cold-start-transient baseline is bounded and self-heals because every re-arm re-snapshots); probation attribution captured BEFORE reorderProbation can promote the entry in onAccess (the promoting access counts as a probation hit, the next as protected); a zero baseline auto-confirming any >=4x-bar earnings (the opportunity cost of a dead boundary is ~zero — nullchurn stays harmless); and the lowmix named trade (a gate sentinel). A walk's bases (down, baseHitRate, baseWindow, baseProbationDensity) are final on the Walk object, which exists only while one is in flight — "dead state while not probing" is the absent object, so there is nothing to go stale.

Cache-side fields: windowMaximum, mainProtectedMaximum, windowWeightedSize, mainProtectedWeightedSize; QUEUE_TRANSFER_THRESHOLD.

WindowClimber owns adjustment, sample (Sample: hits, misses, windowHits, probationHits, previousHitRate), and tier (a ReactiveClimber or DensityClimber). The Climber base class owns step (Step: size), initialized from the maximum and the strategy's chosen direction at construction. DensityClimber owns refractoryLeft, retreatLeft, undoRemaining, and its helper objects: walk (Walk, null while none is in flight: ladder, isAudit, down, baseWindow, baseHitRate, baseSmoothedRate, baseProbationDensity, samples, belowBarStreak, aboveStreak, beatBase), starvation and audit (a Ladder each: rung, crashStreak), auditClock (AuditClock: down, waitSamples, stillSamples, lastWindow), anchor (Anchor: window, rate, held, freshLeft, returning, returnLeft, shortfallStreak), and rates (Rates: smoothed, deviation). Reading is the per-sample derived view, computed once and read-only.

The 38 constants live with the mechanism each tunes, so a knob names its owner: WindowClimber — RESTART_THRESHOLD; DensityClimber — DENSITY_THRESHOLD, DENSITY_GAIN, SAMPLE_MULTIPLIER; Step — STEP_PERCENT, STEP_DECAY_RATE, MIN_INITIAL_STEP; ReactiveClimber — SLOW_ADAPT_THRESHOLD, SLOW_ADAPT_RATIO_CAP, SLOW_ADAPT_DECAY_RATE; Reading — STABLE_BAND_FRACTION, MAX_STEP_FRACTION, WINDOW_FLOOR_FRACTION, DENSITY_EPSILON, MIN_SIGNAL_SHIFT, MIN_STARVATION_BAR; Rates — VETO_MARGIN_MIN, RATE_SMOOTHING, DEVIATION_SEED, VETO_MARGIN_SCALE; AuditClock — AUDIT_WAIT_INITIAL, AUDIT_WAIT_FIRST, AUDIT_WAIT_MAX; Anchor — VETO_STREAK, VETO_RETURN_BUDGET; Walk — PROBE_WALK_BUDGET, PROBE_BAR_CAP, PROBE_EXIT_BAR_MULTIPLE, AUDIT_CRASH_PERSISTENCE, AUDIT_COMMITMENT, AUDIT_CONFIRM_STREAK; Ladder — PROBE_BACKOFF_INITIAL, PROBE_BACKOFF_MAX, PROBE_CRASH_ESCALATION, PROBE_STRIDE_SCALE_MID, PROBE_STRIDE_SCALE_DEEP, PROBE_COMMITMENT_MID, PROBE_COMMITMENT_DEEP.

Show full SKILL.md (671 more words)Show less

Structural invariants to attack (violations are real bugs)

  1. Region partition sum: windowMaximum + mainMaximum (probation + protected) must equal maximum after every climb and every resize. Can any single transfer, or a sequence capped by QUEUE_TRANSFER_THRESHOLD, drift the sum?
  2. Non-negative maxima: can windowMaximum or mainProtectedMaximum go negative — a quota larger than the donor region, or repeated decreaseWindow at the floor? EXCEPTION (adjudicated 2026-07, F1; duration priced 2026-08-22, M1): a transiently negative policyWeight — the sanctioned telescoping race, an out-of-order UpdateTask drain — can over-shift the caps beyond the commanded adjustment, even past these bounds. Tolerated by design: the caps are policy targets (eviction is driven by the telescoping weightedSize/maximum), and they walk back only by the weight each later transfer moves, so a swing larger than a cycle's transfer suspends the split for many cycles with the total still bounded. Report cap drift only with a mechanism that never walks back.
  3. Quota accounting: in increaseWindow/decreaseWindow the quota is decremented per transferred node by policyWeight. With weighted entries, can quota underflow, skip/over-run the loop, or transfer the wrong count? Does the QUEUE_TRANSFER_THRESHOLD cap leave the regions half-adjusted such that the next climb mis-reads them? (Same F1 exception as #2 for the negative transient.)
  4. determineAdjustment math:
    • requestCount = hits + misses; the early return guards requestCount < effectiveSampleSize. Is the hitRate division ever reachable with requestCount == 0?
    • small-cache branch (ReactiveClimber.samplePeriod): (long) (sketchSampleSize * ratio), where ratio = clamp(initialStep / magnitude). Can initialStep be 0 (maximum 0 or tiny) making magnitude 0 → division by zero? Can the (long) cast truncate ratio so it defeats the intended sample-period growth?
    • ReactiveClimber.climb uses Math.copySign(Step.restartMagnitude(max), amount). For amount == 0.0 / -0.0, does copySign choose the intended direction? Can step.size become NaN or 0 and permanently stall adaptation (a stuck-window bug, distinct from slow convergence)?
  5. setMaximumSize at boundaries: the step-sign flip at max <= SLOW_ADAPT_THRESHOLD plus a runtime maximum change via Policy.eviction.setMaximum — when maximum crosses SLOW_ADAPT_THRESHOLD in either direction, do the window/main split, the step.size sign, and the sample state stay mutually consistent?
  6. Stale adjustment consumption: climb calls determineAdjustment then increaseWindow/decreaseWindow off the climber's adjustment. When determineAdjustment early-returns (uninitialized sketch, sub-sample request count), can a stale adjustment from a prior cycle be re-applied?
  7. Layer ownership of the ladders, streaks, and schedules — the highest-yield row, because it has produced two real bugs and both were invisible from a single code path. hill-climber.md §4 carries a write-owner table (observation / active walk / starvation retry / audit retry+schedule / goal guard / motion out). Enumerate every write to starvation.rung, starvation.crashStreak, audit.rung, audit.crashStreak, auditClock.waitSamples, auditClock.stillSamples, anchor.held, anchor.freshLeft and check each against its owner — including endings that are not crashes. Both landed defects were exactly this shape: the shared crash streak let exogenous pulses pair separate audit crashes into a rung ratchet (H4-C1), and later a non-crash ending (budget expiry, reversal-through-base) still cleared the other layer's streak, which disarmed that layer's escalation and its AUDIT_CRASH_PERSISTENCE tolerance — so an interleaved blind corner reached around the first fix. One cross-write is sanctioned, journaled: an audit confirm rewards the starvation ladder (starvation.reward) and zeroes its refractory. An audit's undo leaves the refractory alone (since 2026-08-16, pinned by undoProbe_auditRetreat_leavesTheStarvationRefractoryAlone). Anything else is a finding. Checkable products the oracles already assert, and which a new invariant should join: anchor.freshLeft > 0 ⇒ anchor.held, anchor.held ⇒ anchor.isPlanted, walk != null ⇒ undoRemaining == 0, and auditClock.waitSamples > PROBE_BACKOFF_MAX ⇒ audit.rung == PROBE_BACKOFF_MAX (the ratchet as an invariant). Note the one legitimate coupling so it is not reported: the audit clock is a function of window position, so any layer that moves the window decays stillSamples — that is the intended semantics, and it is the remaining path by which frequent blind corners defer audits.

Output

For each defect: give concrete maximum/weight/access values, trace the arithmetic step by step, show the resulting invariant violation or wrong region size, and a Verification (a BoundedLocalCacheTest white-box method plus the required -P flags).

Everything here runs under evictionLock (single-writer), so most findings will be arithmetic / state-corruption, not races — but explicitly check whether any climber-written field (adjustment, step.size, the region maxima) is also read off-lock by a concurrent reader before concluding "single-writer, cannot race."

© ben-manes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/audit-adaptivity of ben-manes/caffeine.

Open the folder on GitHubat commit e972fb0

Compare with similar skills

Audit Adaptivity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Adaptivity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Adaptivity this skillben-manes/caffeine18k—~3.1kAutomated safety check: PassApache-2.0
Windows Desktop E2Eaffaan-m/ECC276k1 repos~7.6kAutomated safety check: PassMIT
Windows Desktop E2Eaffaan-m/ECC276k—~5.5kAutomated safety check: PassMIT
Agent Adaptive Coordinatorruvnet/ruflo74k2 repos~4kAutomated safety check: PassMIT
Text Resizingthedaviddias/Front-End-Checklist74k—~509Automated safety check: PassMIT
Landmark Regionsthedaviddias/Front-End-Checklist74k—~449Automated safety check: PassMIT

Similar skills

  • E2E testing for Windows native desktop apps (WPF, WinForms, Win32/MFC, Qt) using pywinauto and Windows UI Automation.

    276k GitHub starsUsed in 1 repo~7.6k tokens
    Testing & QAAuto-check passed
  • E2E testing for Windows native desktop apps (WPF, WinForms, Win32/MFC, Qt) using pywinauto and Windows UI Automation.

    276k GitHub stars~5.5k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Agent skill for adaptive-coordinator - invoke with $agent-adaptive-coordinator

    74k GitHub starsUsed in 2 repos~4k tokens
    Auto-check passed
  • Text Resizing

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Support text resizing to 200%.

    74k GitHub stars~509 tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed
  • Landmark Regions

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Use landmark regions correctly.

    74k GitHub stars~449 tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed
  • Logical Properties

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing stylesheets, component styles, and responsive behavior related to Use CSS logical properties for i18n and RTL support.

    74k GitHub stars~526 tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed

More from ben-manes/caffeine

All 33 skills in this repo
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Adversarial Codebase Audit

    ben-manes/caffeine

    Runs a hostile review of the Caffeine Java caching library with parallel subagents that get no design docs, then challenges and consolidates their findings.

    18k GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Caffeine Performance Audit

    ben-manes/caffeine

    Audits the Caffeine cache source for hot-path costs such as allocations, contention and memory layout, reporting only findings tied to specific lines.

    18k GitHub stars~559 tokensUpdated today
    Auto-check passed
  • Audit Sibling Divergence

    ben-manes/caffeine

    Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

    18k GitHub stars~4.3k tokensUpdated today
    Auto-check: notes
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated today
    Auto-check: notes

Questions about Audit Adaptivity

What does Audit Adaptivity do?

Audit the adaptive window hill-climber and region-resize logic for implementation defects (not algorithm quality). Audit Adaptivity is an agent skill from ben-manes/caffeine.

How do I install Audit Adaptivity in Claude Code?

Run `npx skills add ben-manes/caffeine --skill audit-adaptivity -a claude-code`. Or copy the skill folder (.claude/skills/audit-adaptivity in ben-manes/caffeine) into .claude/skills/audit-adaptivity in your project. Claude Code loads it when a task matches its description.

How do I install Audit Adaptivity in Codex?

Run `npx skills add ben-manes/caffeine --skill audit-adaptivity -a codex`. Or copy the skill folder (.claude/skills/audit-adaptivity in ben-manes/caffeine) into .agents/skills/audit-adaptivity in your project. Codex loads it when a task matches its description.

Can I use Audit Adaptivity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ben-manes/caffeine --skill audit-adaptivity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-adaptivity, .gemini/skills/audit-adaptivity, .github/skills/audit-adaptivity and .opencode/skills/audit-adaptivity in your project.

What does Audit Adaptivity need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Adaptivity is instructions for the agent only.

Does Audit Adaptivity access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Adaptivity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Adaptivity use?

Audit Adaptivity is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Adaptivity use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Adaptivity?

Skills that share tags, products or a category with Audit Adaptivity: Windows Desktop E2E (affaan-m/ECC, 276k stars), Windows Desktop E2E (affaan-m/ECC, 276k stars), Agent Adaptive Coordinator (ruvnet/ruflo, 74k stars) and Text Resizing (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Adaptivity?

ben-manes (a GitHub user) maintains it in ben-manes/caffeine, which has 17,881 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 9, 2026.

Source: ben-manes/caffeine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.