Agent skill

Parallel Adversarial Change Review

by ben-manes in ben-manes/caffeine

Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

Apache-2.0Auto-check: notesDevelopment

Install Parallel Adversarial Change Review

skills CLI
$ npx skills add ben-manes/caffeine --skill review-change -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ben-manes/caffeine review-change --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ben-manes/caffeine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/review-change .claude/skills/review-change && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-change
GitHub stars
18k
Token cost
~1.9k tokens
SKILL.md length
324 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

  • Works in 5 steps: Gather the diff → Launch three parallel review subagents → Large diff handling → …
  • Reviewing a Java change that touches concurrent code before it merges
  • SKILL.md covers Input, Step 1: Gather the diff, Step 2: Launch three parallel… and Step 3: Large diff handling, plus 2 more sections
  • Calls git

What it does

Given a branch name, a commit range or nothing, the skill reviews a diff: no argument means uncommitted changes, and a branch is compared with main. An empty diff ends the run with a note that there is nothing to review, and a diff over about 3000 lines triggers a warning and an offer to chunk it by file group.

Three subagents run at the same time. A blind concurrency reviewer sees only the diff and judges it from first principles as a Java concurrency expert. A design-aware reviewer reads the project's design decisions first and reviews against them. A regression pattern matcher checks whether the change resembles historical bug patterns in the Caffeine cache library, the repository it comes from.

A triage step then merges duplicate findings, treating agreement between layers as higher confidence, drops blind-reviewer findings that the design-aware reviewer identifies as intentional, and classifies each survivor as patch, defer or reject. Its prompts and file paths are written for Caffeine, so other projects need to adapt them.

When your agent uses it

  • Reviewing a Java change that touches concurrent code before it merges
  • Getting independent opinions on a branch with and without project context
  • Sorting review findings into fix now, pre-existing and noise

Example prompts

  • “Review my uncommitted changes with the three-layer reviewers.”
  • “Run the multi-layer review on the branch feature/async-refresh against main.”
  • “Review this commit range and triage what the reviewers find.”

Requirements

  • A git repository with the change as a diff, branch or commit range
  • Project design notes for the design-aware reviewer, as in the Caffeine repo
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, Agent, WebSearch, WebFetch, AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Gather the diff
  2. Launch three parallel review subagents
  3. Large diff handling
  4. Triage
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 998978c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • Agent
    • WebSearch
    • WebFetch
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Parallel Adversarial Change Review loads about 1.9k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 324 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, Agent, WebSearch, WebFetch, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ben-manes/caffeine at commit 998978c, republished under its Apache-2.0 licence (© ben-manes). 324 words, ~1,853 tokens.

Download SKILL.mdSave it as .claude/skills/review-change/SKILL.md (or your agent's skills folder).
name
review-change
description
Multi-layer adversarial code review of a diff or branch using parallel specialized reviewers with triage
allowed-tools
Read, Grep, Glob, Bash, Write, Agent, WebSearch, WebFetch, AskUserQuestion
argument-hint
[branch or commit range, default: uncommitted changes]
context
fork
disable-model-invocation
true

Review code changes using three parallel specialized reviewers, then triage and deduplicate findings.

Input

$ARGUMENTS

If no argument given, review uncommitted changes (git diff). If a branch name, review git diff main...$ARGUMENTS. If a commit range, use it directly.

Step 1: Gather the diff

bash
# Uncommitted changes
git diff

# Or branch comparison
git diff main...<branch>

Capture the full diff output as DIFF.

If the diff is empty, report "No changes to review" and stop.

Step 2: Launch three parallel review subagents

Launch all three simultaneously. Each runs independently with different context.

Layer 1: Blind Concurrency Reviewer

Spawn a subagent with ONLY the diff — no project files, no design docs, no context.

Prompt:

You are a Java concurrency expert reviewing a diff. You have NO context about
this project's design decisions or conventions. Analyze purely from first
principles:

1. For every shared mutable field touched: is the access mode sufficient?
2. For every lock acquisition: is the ordering safe? Can it deadlock?
3. For every state transition: is it atomic? Can it be observed partially?
4. For every callback/listener invocation: what locks are held?
5. For every exception path: is cleanup complete?

Output findings as a JSON array:
[{"location": "file:line", "issue": "description", "severity": "critical|high|medium|low", "evidence": "the specific code pattern"}]

Return [] if no issues found.

THE DIFF:
<diff content>
Layer 2: Design-Aware Reviewer

Spawn a subagent WITH project read access. It reads the design docs first.

Prompt:

You are reviewing a code change to the Caffeine cache library. Before reviewing,
read these files to understand intentional design decisions:
- .claude/docs/design-decisions.md
- .claude/docs/synchronization.md
- .claude/rules/design-decisions.md
If the diff touches jcache/, also read .claude/rules/jcache-adapter.md and the
divergence catalogue in .claude/docs/jsr107-conformance.md; guava/ →
.claude/rules/guava-adapter.md; simulator/ → .claude/rules/simulator.md;
examples/ → verify third-party API usage against the upstream contract.

Now review this diff. For each change:
1. Is it consistent with the documented invariants?
2. Does it maintain the lock ordering (evictionLock → CHM bin → synchronized(node))?
3. Does it preserve weight accounting convergence?
4. Does it preserve node lifecycle monotonicity (alive → retired → dead)?
5. Are VarHandle access modes correct for the fields touched?
6. If it touches callback invocation points, are they at the documented positions?

ONLY report issues that are NOT explained by design-decisions.md. If a pattern
looks wrong but is documented as intentional, skip it.

Output findings as a JSON array:
[{"location": "file:line", "issue": "description", "severity": "critical|high|medium|low", "design_ref": "which invariant is violated"}]

Return [] if no issues found.

THE DIFF:
<diff content>
Layer 3: Regression Pattern Matcher

Spawn a subagent WITH project read access. It reads historical bug patterns.

Prompt:

You are checking whether a code change resembles historical bug patterns in
the Caffeine cache. Read these first:
- .claude/docs/design-decisions.md (the "References" section about keyReference visibility)

Then check the diff against these known bug patterns:

1. REFRESH + EXPIRATION RACE: Does the change touch refresh or expiration paths?
   Could it allow a dead key to be passed to a loader, or a refresh to prevent
   expiration indefinitely?

2. VALUE REFERENCE VISIBILITY: Does the change modify how weak/soft values are
   published? Is setRelease + storeStoreFence used (not just setRelease)?

3. ASYNC COMPLETION RACE: Does the change touch async value handling? Could a
   future complete between a check and an action?

4. WRITE BUFFER TASK LOSS: Does the change modify afterWrite or task creation?
   Could a task be lost without the inline fallback firing?

5. NOTIFICATION ORDERING: Does the change move notifyEviction or notifyRemoval
   relative to user code? notifyEviction MUST be before user code.

6. WEIGHT ACCOUNTING: Does the change modify weight tracking? Does it preserve
   the telescoping sum property?

7. BUILD CACHE RELOCATABILITY: Does the change add inputs.files() without
   .withPathSensitivity(PathSensitivity.RELATIVE)?

8. OBLIGATION PAIRING (jcache): Does the change publish events via the
   EventDispatcher? Every publishing thread must drain via awaitSynchronous()/
   ignoreSynchronous() or use the Quietly variants — otherwise pending
   synchronous-listener futures accumulate in the ThreadLocal.

For each match, explain which historical bug it resembles and why.

Output findings as a JSON array:
[{"location": "file:line", "issue": "description", "severity": "critical|high|medium|low", "historical_bug": "issue number or description of the original bug"}]

Return [] if no patterns match.

THE DIFF:
<diff content>

Step 3: Large diff handling

If the diff exceeds ~3000 lines, warn and offer to chunk by file group. If chunked, run steps 2-5 per chunk and consolidate at the end.

Step 4: Triage

Collect findings from all three layers. Then:

  1. Deduplicate: If two layers report the same issue, merge them. Keep the most specific location and combine evidence. Note which layers agreed (agreement = higher confidence).

  2. Filter design-explained findings: If Layer 1 (blind) reports something that Layer 2 (design-aware) explicitly identifies as intentional, drop it and note it was filtered.

  3. Classify each surviving finding:

    • patch — real issue, fixable in this change
    • defer — real issue but pre-existing, not introduced by this change
    • reject — noise, false positive, or handled elsewhere
  4. Track rejections: Count rejected findings and note any where agents disagreed (blind flagged, design-aware cleared). These are worth mentioning to show review thoroughness.

  5. Rank patches by: critical > high > medium > low, then by layer agreement count.

Step 5: Report

Present findings grouped by classification:

══════════════════════════════════════════════════════════
REVIEW: [N] findings ([M] patch, [K] defer, [J] rejected)
══════════════════════════════════════════════════════════

## Patch (actionable in this change)
#1 [critical] [blind+design] file:line — description
   Source: blind + design (2 layers agree)
   Evidence: ...
   Historical: ... (if regression match)

## Defer (pre-existing, not from this change)
...

## Rejected ([J] findings classified as noise)
Notable rejections where agents disagreed:
- [description] (blind: flagged; design-aware: verified correct) — [reason]
- ...

## Filtered (explained by design docs)
- [N] findings from blind review explained by design-decisions.md
══════════════════════════════════════════════════════════

If zero findings survive triage, state how many were raised and rejected, rather than just saying "clean review." Show the work.

© ben-manes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/review-change of ben-manes/caffeine.

Open the folder on GitHubat commit 998978c

Compare with similar skills

Parallel Adversarial Change Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Parallel Adversarial Change Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Parallel Adversarial Change Review this skillben-manes/caffeine18k—~1.9kAutomated safety check: NotesApache-2.0
Cursor Composer Task DelegateChachamaru127/claude-code-harness3.2k—~4.4kAutomated safety check: NotesMIT
Java Code Reviewsivaprasadreddy/sivalabs-agent-skills188—~1.2kAutomated safety check: PassMIT
Parallel Specialist PR Reviewposhan0126/dotclaude870—~1.7kAutomated safety check: PassMIT
Code Revieweralirezarezvani/claude-code-tresor777—~1.8kAutomated safety check: PassMIT
Review And Simplify ChangesDimillian/Skills4k—~2kAutomated safety check: PassMIT

Similar skills

  • Cursor Composer Task Delegate

    Chachamaru127/claude-code-harness

    Hands one implementation task to Cursor Composer in an isolated git worktree, then reviews its diff and cherry-picks the result into the main branch.

    3.2k GitHub stars~4.4k tokensUpdated 6 days ago
    DevelopmentAuto-check: notes
  • Java Code Review

    sivaprasadreddy/sivalabs-agent-skills

    Review Java code for bugs, duplicate code, correctness risks, maintainability improvements, and missing tests.

    188 GitHub stars~1.2k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Parallel Specialist PR Review

    poshan0126/dotclaude

    Reviews a pull request, staged changes or a file by sending the diff to specialist reviewer agents in parallel, then merges their findings into one compact report.

    870 GitHub stars~1.7k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Code Reviewer

    alirezarezvani/claude-code-tresor

    Automatic code quality and best practices analysis. An agent skill from alirezarezvani/claude-code-tresor.

    777 GitHub stars~1.8k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Review a git diff or explicit file scope for reuse, code quality, efficiency, clarity, and standards issues, then optionally apply safe Codex-driven fixes.

    4k GitHub stars~2k tokensUpdated 6 mo ago
    DevelopmentAuto-check passed
  • Ad Review

    CorridorTech/PoseCap

    Two-axis fresh-context code review per WORKFLOW §10. An agent skill from CorridorTech/PoseCap.

    224 GitHub stars~2.4k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes

More from ben-manes/caffeine

All 33 skills in this repo
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Adversarial Codebase Audit

    ben-manes/caffeine

    Runs a hostile review of the Caffeine Java caching library with parallel subagents that get no design docs, then challenges and consolidates their findings.

    18k GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Caffeine Performance Audit

    ben-manes/caffeine

    Audits the Caffeine cache source for hot-path costs such as allocations, contention and memory layout, reporting only findings tied to specific lines.

    18k GitHub stars~855 tokensUpdated today
    Auto-check passed
  • Audit Sibling Divergence

    ben-manes/caffeine

    Compares code paths that should behave the same, such as sync and async cache methods, and requires a concrete scenario where the two observably disagree.

    18k GitHub stars~4.6k tokensUpdated today
    Auto-check: notes
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated today
    Auto-check: notes

Works with

Categories

Questions about Parallel Adversarial Change Review

What does Parallel Adversarial Change Review do?

Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings. Given a branch name, a commit range or nothing, the skill reviews a diff: no argument means uncommitted changes, and a branch is compared with main. An empty diff ends the run with a note that there is nothing to review, and a diff over about 3000 lines triggers a warning and an offer to chunk it by file group.

When should I use Parallel Adversarial Change Review?

Parallel Adversarial Change Review fits situations like: reviewing a Java change that touches concurrent code before it merges; getting independent opinions on a branch with and without project context; sorting review findings into fix now, pre-existing and noise.

How do I install Parallel Adversarial Change Review in Claude Code?

Run `npx skills add ben-manes/caffeine --skill review-change -a claude-code`. Or copy the skill folder (.claude/skills/review-change in ben-manes/caffeine) into .claude/skills/review-change in your project. Claude Code loads it when a task matches its description.

How do I install Parallel Adversarial Change Review in Codex?

Run `npx skills add ben-manes/caffeine --skill review-change -a codex`. Or copy the skill folder (.claude/skills/review-change in ben-manes/caffeine) into .agents/skills/review-change in your project. Codex loads it when a task matches its description.

Can I use Parallel Adversarial Change Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ben-manes/caffeine --skill review-change -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-change, .gemini/skills/review-change, .github/skills/review-change and .opencode/skills/review-change in your project.

What does Parallel Adversarial Change Review need to run?

Going by SKILL.md and its folder, Parallel Adversarial Change Review needs the command-line tools its instructions call (git). Our summary lists: A git repository with the change as a diff, branch or commit range; Project design notes for the design-aware reviewer, as in the Caffeine repo. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, Agent, WebSearch, WebFetch, AskUserQuestion.

Does Parallel Adversarial Change Review access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Parallel Adversarial Change Review safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Parallel Adversarial Change Review use?

Parallel Adversarial Change Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Parallel Adversarial Change Review use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Parallel Adversarial Change Review?

Skills that share tags, products or a category with Parallel Adversarial Change Review: Cursor Composer Task Delegate (Chachamaru127/claude-code-harness, 3.2k stars), Java Code Review (sivaprasadreddy/sivalabs-agent-skills, 188 stars), Parallel Specialist PR Review (poshan0126/dotclaude, 870 stars) and Code Reviewer (alirezarezvani/claude-code-tresor, 777 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Parallel Adversarial Change Review?

ben-manes (a GitHub user) maintains it in ben-manes/caffeine, which has 17,881 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 11, 2026.

Source: ben-manes/caffeine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.