Agent skill

Caffeine Cache Optimization Experiments

by ben-manes in ben-manes/caffeine

Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

Apache-2.0Auto-check: notesDevelopment

Install Caffeine Cache Optimization Experiments

skills CLI
$ npx skills add ben-manes/caffeine --skill optimize-cache -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ben-manes/caffeine optimize-cache --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ben-manes/caffeine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/optimize-cache .claude/skills/optimize-cache && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimize-cache
GitHub stars
18k
Token cost
~2.6k tokens
SKILL.md length
1,258 words
Files
3 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

  • Works in 5 steps: Prepare the comparison → Run and profile the baseline → Form one candidate → …
  • Looking for contention in the cache's read or write path
  • SKILL.md covers Inputs and defaults, 1. Prepare the comparison, 2. Run and profile the baseline and 3. Form one candidate, plus 2 more sections
  • Calls git

What it does

The skill investigates the Caffeine caching library with controlled JMH experiments to find demonstrated unnecessary work or shared contention worth removing, then reviews correctness and returns a patch from an isolated workspace. Microbenchmarks are treated as stress and diagnostic tests, not estimates of application speed, and a higher score on one hot key is not by itself a reason to change the cache.

Priority goes to costs shared across unrelated keys: the read and write buffers, frequency sketch, eviction lock, drain status and timer wheel. The default budget is two hours, enough for about two screened hypotheses and one confirmation. Differences of 2 to 3 percent count as small, 4 percent regressions as concerning and 6 percent or more as worth deep investigation, with 3 percent as the default tolerated loss. The agent keeps separate baseline and candidate worktrees plus a ledger of commands and raw results, and applies, commits or publishes only if you have authorized it.

When your agent uses it

  • Looking for contention in the cache's read or write path
  • Checking whether a suspected hot-path cost is worth removing
  • Confirming a candidate optimization with a baseline and candidate comparison

Example prompts

  • “Look for contention in Caffeine's read buffer and run JMH experiments to confirm any fix.”
  • “Investigate whether the frequency sketch adds unnecessary work on the read path, with a four-hour budget.”
  • “Run a confirmation benchmark on the candidate patch and report any stress cell that regresses by more than 3%.”

Requirements

  • A checkout of the Caffeine repository with its JMH benchmarks and Gradle build
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, Edit, Agent, WebSearch, WebFetch, AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prepare the comparison
  2. Run and profile the baseline
  3. Form one candidate
  4. Measure, review, and decide
  5. Preserve the result

What it can do on your machine

Read from SKILL.md and the folder at commit e972fb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • Edit
    • Agent
    • WebSearch
    • WebFetch
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Caffeine Cache Optimization Experiments loads about 2.6k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 40 tokens; SKILL.md has 1,258 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, Edit, Agent, WebSearch, WebFetch, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ben-manes/caffeine at commit e972fb0, republished under its Apache-2.0 licence (© ben-manes). 1,258 words, ~2,598 tokens.

Download SKILL.mdSave it as .claude/skills/optimize-cache/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
optimize-cache
description
Investigate Caffeine shared-structure costs and read/write hot paths with controlled JMH experiments, correctness review, and a reviewable patch.
allowed-tools
Read, Grep, Glob, Bash, Write, Edit, Agent, WebSearch, WebFetch, AskUserQuestion
argument-hint
[scope; JDK; hypothesis limit; time budget; priorities]
context
fork
disable-model-invocation
true

Cache performance experiments

Find demonstrated unnecessary work or shared contention worth removing. Use microbenchmarks as stress and diagnostic tests, not estimates of application performance. A higher score on one hot key is not, by itself, a reason to change the cache.

Inputs and defaults

$ARGUMENTS

Read .claude/CLAUDE.md, matching rules, and the relevant design/synchronization references. The default budget is two hours including setup and validation, which holds about two screened hypotheses and one confirmation; confirmation alone costs at least 50 minutes at the measurement reference's settings. More hypotheses need more budget. Honor supplied scope and budgets without a requirements questionnaire. Return patches from an isolated workspace; apply, commit, or publish only when already authorized.

Prioritize costs shared across unrelated keys: read/write buffers, frequency sketch, eviction lock, drain status, and timer wheel. Also consider concrete per-operation waste, such as an unnecessary field read or allocation. A win confined to contention on a single node or hash bin needs evidence of broader relevance before it merits an implementation experiment. Policy retuning and weaker semantics are outside the default scope.

For stress-result triage, regard 2–3% differences as small, 4% regressions as concerning, and 6% or larger effects as worth substantial investigation. These bands guide effort, not statistical certainty or automatic acceptance. Use 3% as the default tolerated loss in required stress cells; record user overrides and the primary practical gain threshold before candidate measurements. A small, concrete cleanup can remain useful even when its throughput benefit is unresolved.

1. Prepare the comparison

Inspect current source and read the Performance section of ruled-out.md. Reuse a current /audit-performance report as the hypothesis queue, rechecking its locations and mechanisms at the starting version. If none exists, use bounded source/profiling questions; run that explicit-only audit when the user requests it rather than silently starting an exhaustive audit inside this budget.

Create .local/experiments/<run>/ for source snapshots, commands, raw results, and LEDGER.md. Prepare separate baseline and candidate worktrees from the same snapshot, including relevant staged/unstaged changes and required untracked source/build inputs. A clean HEAD is not the dirty caller tree. Preserve the caller's checkout. Record source, build, generator, harness, and artifact hashes so the starting state is reconstructible without committing it.

Name the immutable starting version and the last independently confirmed version. They initially match. Screen candidates against the latter; use the former only for cumulative reporting. Freeze source and build inputs during each build/measurement interval, including edits by other agents. Verify hashes afterward and invalidate a run if inputs changed while it was running.

Read measurement.md and choose the exact cells, runtime, fork order, practical thresholds, and confirmation sample count. Pin both javaVersion and javaTestVersion for all commands. Check for other active JVM work with the selected JDK's jps -l and lsof. On macOS, confirm sleep prevention with pmset -g assertions before a long batch; caffeinate cannot create one from the agent shell here, and exits 0 after printing the failure to stderr. Keep timed runs separate from builds, tests, profiles, and other benchmarks on the host.

2. Run and profile the baseline

Use the bundled init script to set real task properties at configuration time: one fork, at least five warmup and five measurement iterations, and unique JSON/report/profile destinations. It leaves the convention plugin unchanged. Read its controls in the measurement reference.

The following commands run inside a prepared isolated arm. Choose perf_run as a fresh absolute output directory outside disposable worktrees; perf_init is an absolute path to the same frozen copy of this skill's scripts/jmh-session.init.gradle for both arms. These macOS examples use JDK 27:

bash
./gradlew :caffeine:jmh --rerun --no-configuration-cache -I "$perf_init" -PjavaVersion=27 -PjavaTestVersion=27 -PoptimizeRunDir="$perf_run/timing-01" '-PincludePattern=^com.github.benmanes.caffeine.cache.GetPutBenchmark.(read_only|readwrite|write_only)$' '-PbenchmarkParameters=cacheType=Caffeine'
./gradlew :caffeine:jmh --rerun --no-configuration-cache -I "$perf_init" -PjavaVersion=27 -PjavaTestVersion=27 -PoptimizeRunDir="$perf_run/cpu-01" '-PincludePattern=^com.github.benmanes.caffeine.cache.GetPutBenchmark.write_only$' '-PbenchmarkParameters=cacheType=Caffeine' -Pasync=tree -PasyncEvent=cpu

Prefer call trees for inspection; use JFR when event detail is needed. Keep profiled timing out of acceptance comparisons. Verify fresh execution and actual JVM/settings/cells in the JSON. Trace the measured path: resident unchanged-weight puts may feed the read buffer, not the write queue. Grouped reader/writer counts specify threads, not the completed operation ratio. Track read-only, write-only, mixed reads, and mixed writes separately. Add insertion, churn, or other configurations only where they test the proposed shared mechanism or a plausible regression.

Show full SKILL.md (590 more words)Show less

3. Form one candidate

For each hypothesis record the source site, observed cost, operation removed, JIT evidence, contract boundaries, and a falsifying observation. Keep this queue short; an empty queue is valid. Source-only relocation of constant-foldable conditions is not a mechanism. A compiler-directed idea needs evidence from the hot native code and its compilation/inlining context; javap alone cannot establish removed runtime work.

Branch from the confirmed snapshot and make one conceptual change. Preserve publication and lifecycle ordering, callbacks, statistics, and access recording. Use generator inputs for generated classes. Trace all relevant writers/readers/subclasses before changing memory ordering. Diagnostic removals that weaken behavior may price a cost, but are not candidate optimizations.

Build immutable artifacts. Select focused methods using the Test Discovery Guide in testing.md. Keep gradlew :caffeine:test and its --tests selectors on one physical line: the current test-scope hook can miss continued selectors. For example, when the change concerns put/refresh invalidation:

bash
./gradlew :caffeine:test -PjavaVersion=27 -PjavaTestVersion=27 --tests 'LoadingCacheTest.refresh_discard_put' --rerun

Inspect fresh XML discovery and loaded artifacts. Zero tests, failed discovery containers, or reported skips are not a pass. Add a public contract witness when required; passing ordinary tests alone cannot prove a concurrent change safe.

4. Measure, review, and decide

Screen the affected cells first and label apparent wins provisional. Nominate one final patch for fresh confirmation by default. Five A/B fork pairs mean ten separate invocations with fork=1; use the balanced loop in the measurement reference. Confirm against the current confirmed version, with identical inputs and no editing during runs. Combining edits creates a new candidate whose whole diff needs measurement; percentages do not add.

Use /review-change for the independent final diff review when explicitly requested. Otherwise give independent reviewers its applicable contract checks and the candidate-only diff, excluding the starting user edits. Consume the review's findings rather than duplicating a completed review. Read an invoked sibling's instructions in full and preserve its isolation; batch reviewers if concurrency slots are limited. The coordinator checks the actual patch, raw results, and discovered tests rather than trusting a delegate's keep/drop flag.

Changes to WindowClimber or window resizing require /climber-gate, plus other tests required by the matching project rules. Throughput does not establish eviction quality. If the candidate affects policy work, validate the associated quality contract rather than winning by recording or maintaining less.

Promote only when mechanism, correctness, and confirmed performance support it and every required cell meets its declared regression limit. Keep a small mechanism-confirmed cleanup separate from a claim of measurable speedup. Restore rejected candidates only inside the disposable workspace, checking the restored diff/hash. git checkout alone does not undo an already committed candidate.

5. Preserve the result

Stop at the budget, a fully confirmed goal, or an exhausted hypothesis queue. Reserve confirmation and review time before starting more screens; a full three-group confirmation at the supplied 5+5/10-second settings already costs at least 50 minutes. Start a batch only when its planned duration fits the remaining budget. An unfinished candidate stays unconfirmed.

Return the best confirmed candidate-only diff, separate results and uncertainty for every cell, mechanism evidence, validation limits, and the original checkout's status. Preserve commands, artifacts, rejected patches, and unresolved questions in LEDGER.md.

Also distill adjudicated negative mechanisms into ruled-out.md Performance, as a reviewable documentation diff. Record the mechanism, affected configurations, reason/evidence, and what would justify reopening it. Make the entry understandable without local artifacts; never cite a .local/ path as its evidence. Preserve scope and counterarguments. A noisy or incomplete experiment is inconclusive, not a durable ruling that the mechanism is dead. This closes the loop for later runs without turning machine-specific measurements into universal claims.

© ben-manes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in .claude/skills/optimize-cache of ben-manes/caffeine.

  • SKILL.md
  • references/measurement.md
  • scripts/jmh-session.init.gradle

Open the folder on GitHubat commit e972fb0

Compare with similar skills

Caffeine Cache Optimization Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Caffeine Cache Optimization Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Caffeine Cache Optimization Experiments this skillben-manes/caffeine18k—~2.6kAutomated safety check: NotesApache-2.0
Groovy 5 Developer Guideapache/grails-core2.9k—~3kAutomated safety check: PassApache-2.0
Keybase RPC Log Analysiskeybase/client9.3k—~3kAutomated safety check: PassBSD-3-Clause
Performance CheckZeroDeng01/sublinkPro1.7k—~1.8kAutomated safety check: PassMIT
Minecraft CI ReleaseJahrome907/minecraft-agent-skills166—~3.6kAutomated safety check: PassMIT
Releasesol4k/sol4k135—~949Automated safety check: PassApache-2.0

Similar skills

  • Groovy 5 Developer Guide

    apache/grails-core

    Guidance for Groovy 5 work in Grails projects: syntax, closures, traits, DSLs, metaprogramming, Spock tests, static compilation and Java 21 integration.

    2.9k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • Captures a clean Keybase service log and analyzes it for redundant, duplicated or looping RPCs, then checks whether a caching fix reduced the calls.

    9.3k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • Performance Check

    ZeroDeng01/sublinkPro

    Checklist for reviewing code changes that touch queries, APIs, rendering, caching or algorithms for performance, scalability and resource-usage problems.

    1.7k GitHub stars~1.8k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Minecraft CI Release

    Jahrome907/minecraft-agent-skills

    Set up and review CI, artifact publishing, versioning, and release management for Minecraft 26.x or legacy 1.21.x mods and Paper plugins.

    166 GitHub stars~3.6k tokensUpdated 24 days ago
    DevelopmentAuto-check passed
  • Release

    sol4k/sol4k

    Bump the sol4k library version everywhere, open a release PR, and draft GitHub release notes.

    135 GitHub stars~949 tokensUpdated 9 days ago
    DevelopmentAuto-check passed
  • Guide for writing modern Java 21 in a Grails and Groovy codebase: records, sealed classes, pattern matching, text blocks and how Java works alongside Groovy.

    2.9k GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed

More from ben-manes/caffeine

All 33 skills in this repo
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Codebase Audit

    ben-manes/caffeine

    Runs a hostile review of the Caffeine Java caching library with parallel subagents that get no design docs, then challenges and consolidates their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Caffeine Performance Audit

    ben-manes/caffeine

    Audits the Caffeine cache source for hot-path costs such as allocations, contention and memory layout, reporting only findings tied to specific lines.

    18k GitHub stars~559 tokensUpdated yesterday
    Auto-check passed
  • Audit Sibling Divergence

    ben-manes/caffeine

    Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

    18k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check: notes
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes

Works with

Categories

Questions about Caffeine Cache Optimization Experiments

What does Caffeine Cache Optimization Experiments do?

Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch. The skill investigates the Caffeine caching library with controlled JMH experiments to find demonstrated unnecessary work or shared contention worth removing, then reviews correctness and returns a patch from an isolated workspace. Microbenchmarks are treated as stress and diagnostic tests, not estimates of application speed, and a higher score on one hot key is not by itself a reason to change the cache.

When should I use Caffeine Cache Optimization Experiments?

Caffeine Cache Optimization Experiments fits situations like: looking for contention in the cache's read or write path; checking whether a suspected hot-path cost is worth removing; confirming a candidate optimization with a baseline and candidate comparison.

How do I install Caffeine Cache Optimization Experiments in Claude Code?

Run `npx skills add ben-manes/caffeine --skill optimize-cache -a claude-code`. Or copy the skill folder (.claude/skills/optimize-cache in ben-manes/caffeine) into .claude/skills/optimize-cache in your project. Claude Code loads it when a task matches its description.

How do I install Caffeine Cache Optimization Experiments in Codex?

Run `npx skills add ben-manes/caffeine --skill optimize-cache -a codex`. Or copy the skill folder (.claude/skills/optimize-cache in ben-manes/caffeine) into .agents/skills/optimize-cache in your project. Codex loads it when a task matches its description.

Can I use Caffeine Cache Optimization Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ben-manes/caffeine --skill optimize-cache -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-cache, .gemini/skills/optimize-cache, .github/skills/optimize-cache and .opencode/skills/optimize-cache in your project.

What does Caffeine Cache Optimization Experiments need to run?

Going by SKILL.md and its folder, Caffeine Cache Optimization Experiments needs the command-line tools its instructions call (git). Our summary lists: A checkout of the Caffeine repository with its JMH benchmarks and Gradle build. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, Edit, Agent, WebSearch, WebFetch, AskUserQuestion.

Does Caffeine Cache Optimization Experiments access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Caffeine Cache Optimization Experiments safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Caffeine Cache Optimization Experiments use?

Caffeine Cache Optimization Experiments is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Caffeine Cache Optimization Experiments use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Caffeine Cache Optimization Experiments?

Skills that share tags, products or a category with Caffeine Cache Optimization Experiments: Groovy 5 Developer Guide (apache/grails-core, 2.9k stars), Keybase RPC Log Analysis (keybase/client, 9.3k stars), Performance Check (ZeroDeng01/sublinkPro, 1.7k stars) and Minecraft CI Release (Jahrome907/minecraft-agent-skills, 166 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Caffeine Cache Optimization Experiments?

ben-manes (a GitHub user) maintains it in ben-manes/caffeine, which has 17,880 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 6, 2026.

Source: ben-manes/caffeine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.