Agent skill

Audit Sibling Divergence

by ben-manes in ben-manes/caffeine

Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

Apache-2.0Auto-check: notesDevelopment

Install Audit Sibling Divergence

skills CLI
$ npx skills add ben-manes/caffeine --skill audit-sibling-divergence -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ben-manes/caffeine audit-sibling-divergence --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ben-manes/caffeine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/audit-sibling-divergence .claude/skills/audit-sibling-divergence && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-sibling-divergence
GitHub stars
18k
Token cost
~4.3k tokens
SKILL.md length
1,289 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

  • Works in 5 steps: Inventory matched pairs → Spawn parallel differential auditors → Evaluator challenge → …
  • After significant changes to a library with paired sync and async code paths
  • SKILL.md covers When to run, Step 1: Inventory matched pairs, Step 2: Spawn parallel… and Step 3: Evaluator challenge, plus 3 more sections
  • Reaches raw.githubusercontent.com

What it does

Instead of asking whether one code path is correct, this audit asks whether two paths that should give the same observable result actually agree. It was written for the Caffeine Java caching library, where mismatches between paired paths have caused real bugs, and it starts by inventorying matched pairs: sync and async caches, bounded and unbounded storage, generated node variants, view consistency, bulk versus single operations, fast versus slow read paths and adapter conformance.

One auditor is spawned per sibling pair, and a finding counts only when the auditor produces a witness scenario in which the paths diverge. It is meant for after major changes to the core cache classes or the jcache and guava adapters, before a major release and once per quarter as a baseline. The skill itself calls the run heavyweight, taking roughly four to six hours of agent time and many tokens, and points to a separate review command for routine pre-commit checks.

When your agent uses it

  • After significant changes to a library with paired sync and async code paths
  • Before a major release, as a check for silent drift between sibling implementations
  • Auditing adapters against the specification they claim to match

Example prompts

  • “Run the sibling divergence audit on the async cache changes from this branch.”
  • “Check whether Cache.get and AsyncCache.get handle listeners and exceptions the same way.”
  • “Do the quarterly baseline audit for the jcache and guava adapters.”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Agent, Write

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Inventory matched pairs
  2. Spawn parallel differential auditors
  3. Evaluator challenge
  4. Adjudicate against design docs
  5. Triage and report

What it can do on your machine

Read from SKILL.md and the folder at commit e972fb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Agent
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • raw.githubusercontent.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Sibling Divergence loads about 4.3k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,289 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Agent, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ben-manes/caffeine at commit e972fb0, republished under its Apache-2.0 licence (© ben-manes). 1,289 words, ~4,311 tokens.

Download SKILL.mdSave it as .claude/skills/audit-sibling-divergence/SKILL.md (or your agent's skills folder).
name
audit-sibling-divergence
description
Differential audit comparing matched code paths that should behave identically. Spawns one auditor per sibling pair (sync/async, bounded/unbounded, view consistency, bulk vs single, generated node variants, read fast vs slow, adapter conformance) and requires a concrete witness scenario where the two paths diverge observably.
allowed-tools
Read, Grep, Glob, Bash, Agent, Write
context
fork
disable-model-invocation
true

Audit: Sibling Divergence

The existing snapshot audits look at one code path and ask "is it correct?" This audit looks at TWO paths that should produce the same observable result and asks "do they actually agree?" Divergence between matched paths is a confirmed historical bug pattern in Caffeine (refresh+expiration sync/async asymmetry, weak/strong publication differences, fast-path/slow-path disagreement under contention).

When to run

  • After significant changes to BoundedLocalCache, LocalAsyncCache, or any generator under caffeine/src/javaPoet/java/.
  • After changes to the user-facing adapters (jcache/src/main, guava/src/main): their write/expiry/exception-translation paths are sibling families with each other AND with the external spec/reference they claim to match (Group G). This is where the audit's highest-yield gap was — adapters are user-facing yet were never in any audit's scope.
  • Before a major release as a defense against silent sync/async drift.
  • Once per quarter as a baseline audit.

Heavyweight (4-6h of agent time). Token-intensive. Do not run for routine pre-commit review (use /review-change for that).

Step 1: Inventory matched pairs

The priority matched pairs are listed below. Before launching, glance at the codebase to confirm each pair is still present and add any new pairs (e.g., a new view, a new feature with sync+async variants):

Group A — sync vs async cache (highest historical bug yield)

  • A1: Cache.get(k, Function) vs AsyncCache.get(k, BiFunction) — load semantics, listener delivery, refresh hand-off
  • A2: LoadingCache.refresh(k) vs AsyncLoadingCache.refresh(k) — completion, exception, stale-detection
  • A3: Cache.asMap().compute* vs AsyncCache.asMap().compute* — atomicity, listener cause, weight delta
  • A4: Cache.invalidate* vs AsyncCache.synchronous().invalidate* — listener cause, in-flight value handling

Group B — storage variants

  • B1: BoundedLocalCache vs UnboundedLocalCache for shared map operations (most ops are inherited; check overridden ones and the inherited ones that reference eviction-only fields)
  • B2: Generated Node variants (PS, FS, PW, FW, PSAW, PSWMS, etc.) — for each method declared in Node.java that has multiple subclass implementations, verify all subclasses use consistent access modes, lifecycle transitions, and weight discipline

Group C — view consistency

  • C1: keySet().contains(k) vs containsKey(k) vs asMap().get(k) != null
  • C2: values().contains(v) vs containsValue(v)
  • C3: entrySet() iteration vs forEach() vs keySet() + get(k) per key
  • C4: size() vs entrySet().size() vs counting via iterator
  • C5: keySet().remove(k) / entrySet().remove(entry) vs the corresponding map removal — mapping outcome, blocking, and pending-load disposition, accounting for different return contracts

Group D — bulk vs single-key

  • D1: getAllPresent(keys) vs N×getIfPresent(k)
  • D2: getAll(keys, loader) vs N×get(k, loader)
  • D3: putAll(m) vs N×put(k, v)
  • D4: invalidateAll(keys) vs N×invalidate(k)
  • D5: invalidateAll() vs invalidateAll(allKeys())

Group E — equivalent-by-construction

  • E1: LoadingCache.get(k) vs Cache.get(k, cacheLoader::load)
  • E2: getOrDefault(k, d) vs (get(k) == null ? d : get(k)) (ignoring atomicity)
  • E3: putIfAbsent(k, v) vs compute(k, (key, val) -> val == null ? v : val)
  • E4: AsyncCache.synchronous() view vs an equivalent sync Cache built with the same configuration

Group F — internal paths to the same outcome

  • F1: Read fast path (getIfPresent optimistic) vs slow path (under synchronized(node)) — same key, same logical time, must agree on present-ness and value identity
  • F2: Eviction listener (sync) vs Removal listener (async) for an evicted entry — both should fire, with consistent key/value/cause and the same set of entries
  • F3: Maintenance task variants (AddTask, UpdateTask, RemovalTask, RemovedTask) — weight delta sign convention, telescoping sum preservation across task orderings

Group G — adapter conformance (user-facing: jcache/, guava/)

  • G1: JCache write-path family — every path that builds an Expirable<V> and gates on ExpiryPolicy must agree on the create/update guard, the put statistic, and the CREATED/UPDATED/EXPIRED events: put/putAll (putNoCopyOrAwait), putIfAbsent (putIfAbsentNoAwait), getAndPut, replace, and the EntryProcessor postProcess CREATED/UPDATED/LOADED cases. (A zero-creation-expiry guard present in the put helpers but missing from postProcess was a real bug — a phantom CREATED event + put stat.)
  • G2: JCache adapter vs reference — for each spec-ambiguous edge the TCK does NOT pin (zero/eternal/null creation & update expiry, loader-exception wrapping, read-through statistics), compare the adapter's observable outcome against the JSR-107 RI and ≥3 ecosystem impls (Ehcache3, Infinispan, Hazelcast, Coherence, cache2k). The spec is the contract; the RI + majority resolve ambiguity. A green TCK is necessary, not sufficient — it is silent on exactly the paths that drift.
  • G3: Guava facade vs underlying Caffeine — each CaffeinatedGuavaCache / facade-view method vs the Caffeine method it delegates to: null-query tolerance, exception translation (InvalidCacheLoadException / ExecutionException / UncheckedExecutionException / ExecutionError by checked-ness), bulk partial results.
  • G4: Guava facade vs real Guava — the executable oracle for G3: run the SAME guava-testlib suite (matched feature flags) and operation sequences against the facade and a real Guava CacheBuilder cache; adjudicate differences against Guava's contract and the accepted facade limits.

If a new feature has a sync and async variant not listed above, add it as group H before launching. When auditing the simulator, add reader-vs-sibling-reader (shared binary/text formats, e.g. the libCacheSim family) and climber-vs-climber (shared gradient/timestep convention) port pairs as a group — lower priority, since simulator divergence yields misleading benchmark numbers rather than user-facing bugs. Sim-internal pairs only: the production WindowClimber deliberately has NO faithful simulator reference (product.Caffeine, the real cache, is the arbiter) — do not flag sim-reference-vs-production divergence as a finding.

Show full SKILL.md (505 more words)Show less

Step 2: Spawn parallel differential auditors

Launch one subagent per group (seven groups → seven parallel agents). Each agent gets the prompt below, adapted to its group. Run them in a single message so they execute in parallel.

Tell each subagent to write its report to a group-suffixed path (.local/audits/<model>/audit-sibling-divergence-group<letter>.md), never to the canonical audit-sibling-divergence.md — that path is reserved for the orchestrator's consolidated report (Step 5), and parallel groups writing it clobber each other. If a group returns its report inline instead, persist it to the group-suffixed path before launching that group's evaluator.

You are auditing the Caffeine cache for sibling divergence: cases where two
code paths that should produce identical observable behavior do not.

YOUR GROUP: <group letter and pairs from the inventory>

The Caffeine source code is at:
- Core: caffeine/src/main/java/com/github/benmanes/caffeine/cache/
- Generated: caffeine/build/generated/sources/ (run `./gradlew :caffeine:generateNodes
  :caffeine:generateLocalCaches` if empty)
- Generators: caffeine/src/javaPoet/java/com/github/benmanes/caffeine/cache/
- Adapters (Group G): guava/src/main/java/com/github/benmanes/caffeine/guava/ and
  jcache/src/main/java/com/github/benmanes/caffeine/jcache/. For G2/G4 the reference
  contract is external — the JSR-107 1.1.1 spec/TCK and a real Guava `CacheBuilder`
  cache; construct the witness as a differential test (run both sides), not a read alone.
  To read Guava's actual behavior (no clone needed), WebFetch a specific method from
  `https://raw.githubusercontent.com/google/guava/master/guava/src/com/google/common/cache/LocalCache.java`
  — the source-level complement to G4's executable oracle; the executable side must use
  the pinned Guava version (`libs.versions.toml`), not master.

# Phase 0: Plan
For each pair in your group:
- State the contract the two paths jointly promise (what does an observer
  see when they call A vs B?).
- Predict the 2-3 most likely categories of divergence (different access
  mode, different listener cause, different exception handling, ordering of
  notifications, in-flight value visibility, weight accounting, etc.).
- For blocking differences, identify which thread completes the awaited work. If executor
  capacity affects progress, compare one explicitly sized executor with a spare-worker control;
  distinguish application dependency cycles from a violated cache progress guarantee.

# Phase 1: Trace each side
For each pair:
1. Locate both implementations. Read each end to end (not just the diff).
2. Build a side-by-side table of the observable steps each path takes:
   field reads/writes (with access mode), lock acquisitions, listener
   invocations, exceptions thrown, return values.
3. Identify every step where the two paths differ. For each difference:
   - Is the difference observable to a caller?
   - Is it explained by an intentional design decision? (read
     .claude/docs/design-decisions.md and .claude/rules/design-decisions.md
     ONLY AFTER you have recorded the difference — design context causes
     premature dismissal)
   - If observable and unexplained, this is a candidate finding.

# Phase 2: Construct a witness
For each candidate finding, construct a CONCRETE WITNESS:
- Cache configuration (size, weigher, listener, expiry, ...)
- Exact sequence of method calls
- Expected observation if the paths agreed
- Actual observation given the divergence
- Strong enough that a developer could write a failing unit test from it
  without further investigation.

A finding without a concrete witness is NOT acceptable — drop it.

# Phase 3: Self-challenge
For each finding, attempt to refute it by re-reading the source. If you
can construct a path through the code that resolves the divergence (e.g.,
the second path also fires the listener through a different route you
missed), drop the finding. Be ruthless — the user wants high-precision
findings, not volume.

For differing guards, compare reachable callers and preconditions on both paths. One path may
establish the condition elsewhere or never reach the guarded state; its safety does not make
the other path's guard redundant, and the difference need not imply a bug on either side.

# Phase 4: Output
For each surviving finding, output:
- PAIR: <pair label, e.g., A1>
- PATH-A: file:method
- PATH-B: file:method
- DIVERGENCE: <one-line description of what differs>
- WITNESS: <concrete scenario>
- OBSERVABLE: <what the caller sees that contradicts the joint contract>
- DESIGN-MATCH: <design-decisions.md item it partially matches, or "none">
- SEVERITY: critical | high | medium | low
- CONFIDENCE: high | medium

If a group has zero findings, output a coverage summary listing every pair
inspected, every method traced, and every difference dismissed (with the
reason for dismissal). Zero findings with thorough coverage is acceptable.
Zero findings with shallow coverage is not — keep looking.

DO NOT report:
- Performance differences (covered by /audit-performance)
- Style differences
- Comment/javadoc differences unless they constitute the contract drift
- Differences explicitly documented as intentional in design-decisions.md
  (note them in the "explained" section instead)

Step 3: Evaluator challenge

For each agent that returned findings OR a zero-findings coverage proof, spawn ONE evaluator subagent (general-purpose). The evaluator sees ONLY the reviewer's report — no source code.

You are challenging a differential audit report. Your job is to find what
the auditor MISSED.

For each finding:
1. Is the witness scenario actually reproducible? Identify any unstated
   precondition (specific config, timing, prior state) that the witness
   does not enumerate but requires.
2. Is the divergence actually observable to a caller? Or is it an internal
   difference that produces the same external result?
3. Is the auditor's "joint contract" the actual contract, or did they
   assume a stronger contract than the documentation promises?

For each zero-findings claim:
4. Given the auditor's stated coverage, what categories of divergence
   might they have under-weighted? (E.g., focused on synchronous control
   flow, missed exception paths; focused on happy path, missed in-flight
   transitions.)

Output: prioritized list of challenges. Be specific about which finding
and which gap.

Have the original reviewer address each challenge by re-reading source. Drop findings the reviewer cannot defend with concrete evidence. Add new findings the reviewer confirms.

Step 4: Adjudicate against design docs

Read .claude/docs/design-decisions.md and .claude/rules/design-decisions.md. For each surviving finding, classify:

  • confirmed-divergence — a concrete supported witness violates the stated joint contract or invariant; an unexplained difference alone is insufficient
  • intentional-divergence — documented design decision (e.g., async listener delivery is intentionally different from sync; weight=0 entries are pinned across all paths); keep in the report under "explained"
  • documentation-gap — intentional behavior with a useful, supported wording clarification; check existing guidance and respect prior decisions declining that same disclosure or wording

When several explained differences affect the same entry lifetime, check at most one concrete sequence combining them per group. Compare its observable outcome with the joint contract and the ruling's mechanism, consequence, trigger, and scope. Combining accepted behavior does not itself reopen a ruling. If reachability or the contract remains unresolved, retain the precise question under Residual risk rather than forcing a bug or documentation-gap label.

Step 5: Triage and report

Triage confirmed findings by severity per .claude/docs/finding-taxonomy.md.

Tag each confirmed finding with the divergence axis:

  • sync-async — the two paths represent the sync and async variants
  • storage-variant — bounded vs unbounded or generated node variant
  • view-consistency — view methods disagreeing with each other
  • bulk-vs-single — aggregate operation diverging from per-element
  • equivalent-by-construction — two APIs that should be interchangeable
  • internal-path — fast vs slow path or maintenance task variants
  • adapter-conformance — an adapter path diverging from a sibling adapter path, the underlying Caffeine method it delegates to, or the external spec/reference it claims to match

Write the full report to .local/audits/<model>/audit-sibling-divergence.md (see .claude/docs/audit-output.md).

Format:

# Sibling Divergence Audit

[N] differential auditors compared [M] sibling pairs across [G] groups.
[K] findings survived self-challenge and evaluator challenge.

## Confirmed divergences (likely bugs)

#1 [severity] [axis] PAIR — one-line summary
- PATH-A: file:method
- PATH-B: file:method
- DIVERGENCE: ...
- WITNESS: ...
- OBSERVABLE: ...

## Intentional divergences (documented)
- [pair] — link to design doc that explains the difference

## Documentation clarifications
- [pair] — the useful clarification and the existing guidance or prior wording decision checked

## Coverage summary
- Group A: [pairs inspected, methods traced, dismissals]
- Group B: ...
- ...

## Evaluator challenges
- [N] challenges received across [M] groups; [K] led to new findings;
  [J] confirmed the original conclusion with additional evidence.

## Residual risk
What was not inspected and why.

Notes

  • The auditor agent's 4-phase methodology applies here too, but per-pair rather than per-method. Each subagent runs Phase 0 through Phase 3 for its assigned pairs.
  • Generator pairs (B2) are special: read both the generator and the generated output. A divergence might be in the generator's emit logic (one feature combination emits the wrong access mode) or in a missing generator case (a feature combination emits no method at all).
  • The synchronous() view of an async cache (E4) is the highest-yield pair in group E because it goes through the async code path internally but promises sync semantics. Historical bugs here include synchronous() exposing in-flight CompletableFuture state in unexpected ways.

© ben-manes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/audit-sibling-divergence of ben-manes/caffeine.

Open the folder on GitHubat commit e972fb0

Compare with similar skills

Audit Sibling Divergence next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Sibling Divergence compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Sibling Divergence this skillben-manes/caffeine18k—~4.3kAutomated safety check: NotesApache-2.0
Code Review Skillawesome-skills/code-review-skill2.1k—~2.8kAutomated safety check: NotesMIT
Cross-Language Coding Standardszereight/gitlab-mcp2k1 repos~1.4kAutomated safety check: PassMIT
Code Qualitypiomin/claude-ai-spring-boot1.3k—~2.2kAutomated safety check: PassApache-2.0
Code Review Excellenceandrew-yangy/gru-ai155—~1.7kAutomated safety check: NotesMIT
Code Revieweralirezarezvani/claude-code-tresor777—~1.8kAutomated safety check: PassMIT

Similar skills

  • Code Review Skill

    awesome-skills/code-review-skill

    Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, Java 8, PHP, Ruby, Rails, Python, Django, FastAPI, Go, C/.NET, Kotlin, Swift, Dart…

    2.1k GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Shared reference for naming, function size, complexity and error handling rules that reviewer agents apply across TypeScript, Python, Go, Rust, Java, C# and Swift.

    2k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • Code Quality

    piomin/claude-ai-spring-boot

    Comprehensive code review for Java - clean code principles, API contracts, null safety, exception handling, and performance.

    1.3k GitHub stars~2.2k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Code Review Excellence

    andrew-yangy/gru-ai

    Provides comprehensive code review guidance for React 19, Vue 3, Rust, TypeScript, Java, Python, and C/C++.

    155 GitHub stars~1.7k tokensUpdated 7 mo ago
    DevelopmentAuto-check: notes
  • Code Reviewer

    alirezarezvani/claude-code-tresor

    Automatic code quality and best practices analysis. An agent skill from alirezarezvani/claude-code-tresor.

    777 GitHub stars~1.8k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Code Review

    dotnet/maui

    Official

    Deep code review of PR or materialized candidate-patch changes for correctness, safety, and MAUI conventions.

    23k GitHub stars~8.2k tokensUpdated today
    DevelopmentAuto-check passed

More from ben-manes/caffeine

All 33 skills in this repo
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Codebase Audit

    ben-manes/caffeine

    Runs a hostile review of the Caffeine Java caching library with parallel subagents that get no design docs, then challenges and consolidates their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Caffeine Performance Audit

    ben-manes/caffeine

    Audits the Caffeine cache source for hot-path costs such as allocations, contention and memory layout, reporting only findings tied to specific lines.

    18k GitHub stars~559 tokensUpdated yesterday
    Auto-check passed
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

    18k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Audit Sibling Divergence

What does Audit Sibling Divergence do?

Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree. Instead of asking whether one code path is correct, this audit asks whether two paths that should give the same observable result actually agree. It was written for the Caffeine Java caching library, where mismatches between paired paths have caused real bugs, and it starts by inventorying matched pairs: sync and async caches, bounded and unbounded storage, generated node variants, view consistency, bulk versus single operations, fast versus slow read paths and adapter conformance.

When should I use Audit Sibling Divergence?

Audit Sibling Divergence fits situations like: after significant changes to a library with paired sync and async code paths; before a major release, as a check for silent drift between sibling implementations; auditing adapters against the specification they claim to match.

How do I install Audit Sibling Divergence in Claude Code?

Run `npx skills add ben-manes/caffeine --skill audit-sibling-divergence -a claude-code`. Or copy the skill folder (.claude/skills/audit-sibling-divergence in ben-manes/caffeine) into .claude/skills/audit-sibling-divergence in your project. Claude Code loads it when a task matches its description.

How do I install Audit Sibling Divergence in Codex?

Run `npx skills add ben-manes/caffeine --skill audit-sibling-divergence -a codex`. Or copy the skill folder (.claude/skills/audit-sibling-divergence in ben-manes/caffeine) into .agents/skills/audit-sibling-divergence in your project. Codex loads it when a task matches its description.

Can I use Audit Sibling Divergence in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ben-manes/caffeine --skill audit-sibling-divergence -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-sibling-divergence, .gemini/skills/audit-sibling-divergence, .github/skills/audit-sibling-divergence and .opencode/skills/audit-sibling-divergence in your project.

What does Audit Sibling Divergence need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Sibling Divergence is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Agent, Write.

Does Audit Sibling Divergence access the network?

SKILL.md names 1 domain. In commands or code: raw.githubusercontent.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Audit Sibling Divergence safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Audit Sibling Divergence use?

Audit Sibling Divergence is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Sibling Divergence use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Sibling Divergence?

Skills that share tags, products or a category with Audit Sibling Divergence: Code Review Skill (awesome-skills/code-review-skill, 2.1k stars), Cross-Language Coding Standards (zereight/gitlab-mcp, 2k stars), Code Quality (piomin/claude-ai-spring-boot, 1.3k stars) and Code Review Excellence (andrew-yangy/gru-ai, 155 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Sibling Divergence?

ben-manes (a GitHub user) maintains it in ben-manes/caffeine, which has 17,882 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 9, 2026.

Source: ben-manes/caffeine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.