Agent skill

Cross-Model Adversarial Code Review

by basicmachines-co in basicmachines-co/basic-memory

Reviews the current branch's diff with two different model families, has each try to refute the other's findings and reports the survivors by confidence.

MITAuto-check passedDevelopment

Install Cross-Model Adversarial Code Review

skills CLI
$ npx skills add basicmachines-co/basic-memory --skill adversarial-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install basicmachines-co/basic-memory adversarial-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/basicmachines-co/basic-memory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/adversarial-review .claude/skills/adversarial-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adversarial-review
GitHub stars
4.1k
Token cost
~2k tokens
SKILL.md length
850 words
Files
5
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Reviews the current branch's diff with two different model families, has each try to refute the other's findings and reports the survivors by confidence.

  • Works in 4 steps: Deterministic gates (before the models) → Independent review (you + the other… → Cross-refute → …
  • Asking for a second-opinion review from a different model before merging
  • SKILL.md covers You are the orchestrator — and…, Inputs, Preflight and Phase 0 — Deterministic gates…, plus 4 more sections
  • Calls codex, claude and just

What it does

This skill sets up a cross-vendor code review of your current branch. The running agent reviews the diff itself, then launches the other model family as a fresh subprocess, `codex exec` when you are in Claude Code and `claude -p` when you are in Codex, so the second pass shares no context. Each reviewer then tries to refute the other's findings, and a finding's confidence depends on whether it survives that cross-examination.

The aim is to avoid two weaknesses of solo LLM review: a model going easy on its own work, and confident false alarms. You can set `BASE`, the ref to diff against (default main), and `SCOPE`, a pathspec that narrows the review. The diff command is built once as an argument array so paths with spaces survive. Prompts and JSON schemas for findings and verdicts ship with the skill, and the review is report-only: it never applies fixes.

When your agent uses it

  • Asking for a second-opinion review from a different model before merging
  • Wanting high-confidence findings from a branch diff
  • Reviewing only one directory of a large change

Example prompts

  • “Run an adversarial review of this branch against main.”
  • “Do a cross-model review of my changes limited to src/billing and list the findings by confidence.”
  • “I want a second opinion before I merge. Have both models review the diff and challenge each other.”

Requirements

  • Both the Claude Code and Codex command line tools
  • A git repository with a branch to compare against main

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Deterministic gates (before the models)
  2. Independent review (you + the other model, concurrently)
  3. Cross-refute
  4. Synthesize and report (no auto-fix)

What it can do on your machine

Read from SKILL.md and the folder at commit cb7407f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • codex
    • claude
    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cross-Model Adversarial Code Review loads about 2k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 850 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from basicmachines-co/basic-memory at commit cb7407f, republished under its MIT licence (© basicmachines-co). 850 words, ~1,962 tokens.

Download SKILL.mdSave it as .claude/skills/adversarial-review/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
adversarial-review
description
Cross-vendor adversarial code review of the current branch. Two different model families (Claude + Codex/GPT) review the diff independently, then try to refute each other's findings; survivors are reported by confidence. Runs from either Claude Code or Codex. Use when the user asks for an adversarial review, a cross-model / second-opinion review, or wants high-confidence findings before merging. Report-only — never auto-applies fixes.
license
MIT

Adversarial code review

Two reviewers from different model families — Claude and Codex/GPT — review the same diff independently, then each tries to refute the other's findings. A finding's confidence comes from whether it survives that cross-examination. This kills the two failure modes of solo LLM review: self-ratification (a model won't critique its own work) and confident false positives.

You are the orchestrator — and one of the two reviewers

This skill runs from either Claude Code or Codex. First, identify which model family you are (Claude or Codex/GPT). Then:

  • You are reviewer #1. You review natively, in this session, using your own tools.
  • The other family is reviewer #2. You invoke it as a subprocess CLI for an independent pass: a fresh process, no shared context — that independence is the point.

The CLI for "the other model":

If you are…Invoke the other via…
Claudecodex exec (GPT)
Codexclaude -p (Claude)

Everything else in the flow is symmetric. Resolve the prompts/ and schemas/ paths below relative to this skill's own directory (where this SKILL.md lives).

Inputs

Two independent, optional inputs:

  • BASE — the ref to diff against. Default main.
  • SCOPE — a pathspec to narrow the review (e.g. src/basic_memory). Default: none (whole diff).

These are separate: a ref and a pathspec are not interchangeable. Build the canonical diff command once in preflight and reuse it everywhere below — never re-spell the diff inline (the scattered, inconsistent spelling is what broke earlier). Build it as an argv array, not a string, so a $SCOPE containing spaces or glob characters survives intact:

bash
BASE="${BASE:-main}"
DIFF=(git diff "$BASE...HEAD")          # argv array — never a scalar string
[ -n "$SCOPE" ] && DIFF+=(-- "$SCOPE")  # pathspec stays one argument even with spaces
DIFF_STR=$(printf '%q ' "${DIFF[@]}")   # shell-quoted rendering, for embedding in a prompt

To run it, use "${DIFF[@]}" (quoted, no word-splitting). To embed it as text inside a subprocess prompt, use $DIFF_STR.

Preflight

  1. Set SKILL_DIR to the directory this SKILL.md lives in. Canonical location is .agents/skills/adversarial-review (the shared agent-skills store); Claude Code reaches it via the .claude/skills/adversarial-review symlink, Codex via its own skills path. The prompts/ and schemas/ subdirs are siblings of this file in every case.
  2. Confirm the other model's CLI is on PATH (codex if you're Claude, claude if you're Codex). If it's missing, tell the user the panel falls back to single-model (which loses the cross-vendor benefit) and ask whether to proceed or stop.
  3. Run "${DIFF[@]}". If it prints nothing, report "nothing to review against $BASE" (mention $SCOPE if set) and stop.
  4. RUN=$(mktemp -d) — scratch dir for the other model's output. Transient, never committed. No persisted artifacts, no state file.

Phase 0 — Deterministic gates (before the models)

Models are statistically blind to negation ("never do X"). Enforce mechanical house rules with tools, not prompts, and treat hits as high-confidence facts (reported separately from model findings):

  • just lint and just typecheck if the diff touches src/.
  • Grep the diff for catchable house-rule violations: getattr(.*,.*, defaults, bare except: / except Exception: pass, function-scope imports.
Show full SKILL.md (389 more words)Show less

Phase 1 — Independent review (you + the other model, concurrently)

Both reviewers get the same brief: prompts/review.md + the repo's CLAUDE.md house rules, reviewing the diff from "${DIFF[@]}". Both emit findings matching schemas/findings.schema.json.

Your native pass: review as yourself, following prompts/review.md. Hold your findings as that JSON shape.

The other model's pass — run, from the repo root, the row that matches you:

Always redirect codex stdin from /dev/null — if stdin is a pipe (e.g. the call gets backgrounded), codex exec blocks "Reading additional input from stdin..." and fails.

bash
# You are Claude → run Codex:
codex exec -s read-only \
  --output-schema "$SKILL_DIR/schemas/findings.schema.json" \
  -o "$RUN/other_findings.json" \
  "$(cat "$SKILL_DIR/prompts/review.md")

Review the diff: $DIFF_STR" </dev/null

# You are Codex → run Claude (read-only via plan mode; parse the JSON block it returns):
claude -p --permission-mode plan --output-format json \
  "$(cat "$SKILL_DIR/prompts/review.md")

Review the diff: $DIFF_STR
Return ONLY a JSON object matching this schema:
$(cat "$SKILL_DIR/schemas/findings.schema.json")" </dev/null > "$RUN/other_raw.json"
# claude --output-format json output shape varies by CLI version: it may be a JSON ARRAY
# of event objects, OR a single result object. Normalize before reading: if it's an array,
# take the element with type=='result'; otherwise use the object as-is. Then read its
# .result string, strip the ```json fence if present, and parse that.
# (Verified empirically: the CLI in this environment emits the array form.)

Runtime note for Codex orchestrating: claude -p needs network access, which Codex's default sandbox blocks. Run it from a Codex session whose project is trusted with network allowed (or approve the claude call when prompted). Keep Codex's own sandbox on — do not bypass it just to reach the network.

Tag each finding with its origin (claude / codex).

Phase 2 — Cross-refute

Each model tries to refute the other's findings, per prompts/refute.md (verdicts match schemas/verdicts.schema.json).

  • You refute the other model's findings natively.
  • The other model refutes your findings — invoke it again the same way (swap prompts/review.md for prompts/refute.md, append your findings JSON and $DIFF_STR so it judges against the right base and scope, and for Codex use --output-schema "$SKILL_DIR/schemas/verdicts.schema.json").

Match verdicts to findings by id.

Phase 3 — Synthesize and report (no auto-fix)

Merge, dedupe (same file + overlapping lines + same root cause = one finding), assign confidence from provenance:

  • High — both models raised it independently, OR one raised it and the other upheld it.
  • Medium — one raised it; the other could not refute it but did not independently find it.
  • Low / contested — one raised it and the other refuted it. Keep it, show both sides, let the human judge. Never silently drop a contested finding.
  • Deterministic-gate hits are reported as facts, separate from the model panel.

Rank by severity × confidence. Present a compact table: severity | confidence | file:line | claim | found-by / upheld-or-refuted-by. Expand the high-confidence ones with why and any suggested fix.

End by asking which findings, if any, to fix. Do not edit code until the user picks. Convergence between the models is not correctness — your job is to surface a ranked, cross-examined list, not to declare the branch clean.

Deliberately NOT done

  • No loop-until-both-agree (models converge by going silent, not by being right).
  • No persisted artifacts / state machine — the scratch dir is thrown away.
  • No auto-applying fixes.

© basicmachines-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in .agents/skills/adversarial-review of basicmachines-co/basic-memory.

  • SKILL.md
  • prompts/refute.md
  • prompts/review.md
  • schemas/findings.schema.json
  • schemas/verdicts.schema.json

Open the folder on GitHubat commit cb7407f

Compare with similar skills

Cross-Model Adversarial Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cross-Model Adversarial Code Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cross-Model Adversarial Code Review this skillbasicmachines-co/basic-memory4.1k—~2kAutomated safety check: PassMIT
Tbdjlevy/strif131—~3.5kAutomated safety check: PassMIT
O2 Review Loopopenobserve/openobserve22k—~3.7kAutomated safety check: PassAGPL-3.0
Reviewatelier-fashion/adlc-toolkit171—~1.9kAutomated safety check: PassMIT
Clawteamwin4r/ClawTeam-OpenClaw1.5k—~3.1kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Tbd

    jlevy/strif

    Git-native issue tracking (beads), coding guidelines, knowledge injection, and spec-driven planning for AI agents.

    131 GitHub stars~3.5k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Review

    atelier-fashion/adlc-toolkit

    Multi-agent code review covering correctness, quality, architecture, test coverage, and security

    171 GitHub stars~1.9k tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Clawteam

    win4r/ClawTeam-OpenClaw

    Multi-agent swarm orchestration. An agent skill from win4r/ClawTeam-OpenClaw.

    1.5k GitHub stars~3.1k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Understand Diff Analysis

    Egonex-AI/Understand-Anything

    Reads your git changes or a pull request against a prebuilt knowledge graph of the project to explain what changed, which components are affected and what is risky.

    86k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed

More from basicmachines-co/basic-memory

All 49 skills in this repo
  • cmux Settings Editor

    basicmachines-co/basic-memory

    Views, sets, unsets and validates cmux settings in ~/.config/cmux/cmux.json with a helper script that checks keys against the schema.

    4.1k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • cmux Window and Pane Control

    basicmachines-co/basic-memory

    End-user control of cmux topology and routing (windows, workspaces, panes/surfaces, focus, moves, reorder, identify, trigger flash). Use when automation needs…

    4.1k GitHub starsUsed in 2 repos~842 tokens
    Auto-check passed
  • Cmux Markdown Viewer Panel

    basicmachines-co/basic-memory

    Opens markdown files in a formatted cmux panel beside the terminal that re-renders on every change, handy for plans and task lists.

    4.1k GitHub starsUsed in 2 repos~527 tokens
    Auto-check passed
  • cmux Workspace Scoping

    basicmachines-co/basic-memory

    Keeps agent actions scoped to the cmux workspace and terminal that invoked it, and lays out pane and surface commands that avoid disrupting the user's own focus.

    4.1k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Basic Memory Repo Images

    basicmachines-co/basic-memory

    Produces PR, changelog and two-week retro images for the Basic Memory repository from evidence in PR bodies, saved to fixed paths under docs/assets/infographics.

    4.1k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Logfire Instrumentation

    basicmachines-co/basic-memory

    Adds Pydantic Logfire tracing, logging and metrics to Python, JavaScript or TypeScript and Rust projects, with the correct setup order and library extras.

    4.1k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Cross-Model Adversarial Code Review

What does Cross-Model Adversarial Code Review do?

Reviews the current branch's diff with two different model families, has each try to refute the other's findings and reports the survivors by confidence. This skill sets up a cross-vendor code review of your current branch. The running agent reviews the diff itself, then launches the other model family as a fresh subprocess, `codex exec` when you are in Claude Code and `claude -p` when you are in Codex, so the second pass shares no context.

When should I use Cross-Model Adversarial Code Review?

Cross-Model Adversarial Code Review fits situations like: asking for a second-opinion review from a different model before merging; wanting high-confidence findings from a branch diff; reviewing only one directory of a large change.

How do I install Cross-Model Adversarial Code Review in Claude Code?

Run `npx skills add basicmachines-co/basic-memory --skill adversarial-review -a claude-code`. Or copy the skill folder (.agents/skills/adversarial-review in basicmachines-co/basic-memory) into .claude/skills/adversarial-review in your project. Claude Code loads it when a task matches its description.

How do I install Cross-Model Adversarial Code Review in Codex?

Run `npx skills add basicmachines-co/basic-memory --skill adversarial-review -a codex`. Or copy the skill folder (.agents/skills/adversarial-review in basicmachines-co/basic-memory) into .agents/skills/adversarial-review in your project. Codex loads it when a task matches its description.

Can I use Cross-Model Adversarial Code Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add basicmachines-co/basic-memory --skill adversarial-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adversarial-review, .gemini/skills/adversarial-review, .github/skills/adversarial-review and .opencode/skills/adversarial-review in your project.

What does Cross-Model Adversarial Code Review need to run?

Going by SKILL.md and its folder, Cross-Model Adversarial Code Review needs the command-line tools its instructions call (codex, claude and just). Our summary lists: Both the Claude Code and Codex command line tools; A git repository with a branch to compare against main.

Does Cross-Model Adversarial Code Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cross-Model Adversarial Code Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cross-Model Adversarial Code Review use?

Cross-Model Adversarial Code Review is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cross-Model Adversarial Code Review use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cross-Model Adversarial Code Review?

Skills that share tags, products or a category with Cross-Model Adversarial Code Review: Tbd (jlevy/strif, 131 stars), O2 Review Loop (openobserve/openobserve, 22k stars), Review (atelier-fashion/adlc-toolkit, 171 stars) and Clawteam (win4r/ClawTeam-OpenClaw, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cross-Model Adversarial Code Review?

basicmachines-co (a GitHub organization) maintains it in basicmachines-co/basic-memory, which has 4,115 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 7, 2026.

Source: basicmachines-co/basic-memory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.