Agent skill

Adversarial Spec Review

by JuliusBrussee in JuliusBrussee/cavekit

Builds a skeptical reviewer grounded in the codebase and research notes to try to refute a spec before any code is written, citing file:line evidence and ending in a go or no-go gate.

MITAuto-check passedDevelopment

Install Adversarial Spec Review

skills CLI
$ npx skills add JuliusBrussee/cavekit --skill review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JuliusBrussee/cavekit review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JuliusBrussee/cavekit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/review .claude/skills/review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review
GitHub stars
1.2k
Token cost
~959 tokens
SKILL.md length
467 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Builds a skeptical reviewer grounded in the codebase and research notes to try to refute a spec before any code is written, citing file:line evidence and ending in a go or no-go gate.

  • Works in 5 steps: CAPTURE → CONSTRUCT THE SENIOR → REFUTE → …
  • Reviewing a spec before building a change to a shared module, auth or payments
  • SKILL.md covers WHEN TO REVIEW, PHASE 0 — CAPTURE, PHASE 1 — CONSTRUCT THE SENIOR and PHASE 2 — REFUTE, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill treats a review as an attempt at refutation, not approval: every finding must cite evidence, a file and line or a source, and a flaw that cannot be proven is downgraded to an unverified note rather than waved through or inflated into a blocker. It applies before building a high-blast-radius change, such as one touching a shared module, auth, money or a public API, or when a spec's interface or invariants section affects dependent code, and it is explicitly skipped for a trivial, reversible, well-understood change, since adversarial review on a typo wastes the effort budget.

Phase 0 reads the whole spec rather than relying on memory of the conversation. Phase 1 builds the reviewer's authority from three sources: the actual codebase patterns and invariants, what the research section already established, and a live best-practice check for any claim that seems out of date. Phase 2 attacks the spec on specific axes: whether the goal addresses the real problem, missing invariants, interface drift against existing callers, contradicting constraints, unowned edge cases and wrong altitude. Phase 3 classifies each finding as BLOCK, HARDEN or NOTE, and phase 4 hardens the spec's verification section and ends in an explicit go or no-go decision.

When your agent uses it

  • Reviewing a spec before building a change to a shared module, auth or payments
  • Red-teaming a plan before committing engineering time to it
  • Deciding whether a spec is sound enough to move to a go or no-go gate

Example prompts

  • “Review the spec for this payments refactor before we build it.”
  • “Red-team this plan and tell me what it misses.”
  • “Is this plan sound? Run the senior review and give me a go or no-go.”

Requirements

  • A written spec with the sections this review reads

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. CAPTURE
  2. CONSTRUCT THE SENIOR
  3. REFUTE
  4. CLASSIFY
  5. HARDEN §V & GATE

What it can do on your machine

Read from SKILL.md and the folder at commit 7421e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adversarial Spec Review loads about 959 tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 467 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~959

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JuliusBrussee/cavekit at commit 7421e87, republished under its MIT licence (© JuliusBrussee). 467 words, ~959 tokens.

Download SKILL.mdSave it as .claude/skills/review/SKILL.md (or your agent's skills folder).
name
review
description
Adversarial senior review of the spec before any code is written. Constructs a skeptical reviewer whose authority comes from the codebase, §R research, and live best-practice — then tries to REFUTE the spec, not rubber-stamp it. Every finding cites evidence (file:line or source); unverifiable ones are flagged. Survivors harden §V; the run ends in an explicit go / no-go gate. Triggers before building anything high-blast-radius, when the user says "review the spec", "red-team this", "is this plan sound", "senior review", or invokes /ck:review.

review — refute the spec before build

Every finding cites evidence — file:line or a source. No evidence → flag [unverified]. Default to refuted: a flaw you cannot prove is a flaw you note, not one you wave through.

An LLM cannot self-correct on its own judgment — left alone it drifts or degrades. Review fixes that the only way that works: a separate skeptic anchored to an external oracle — the code, §R, the test suite, the docs. "Looks good" is not a review. A refutation attempt is.

WHEN TO REVIEW

  • Before /build on a high-blast-radius change (shared module, auth, data, money, public API).
  • Spec touched §I or §V that other code depends on.
  • Right-sizing says the cost of a wrong build > the cost of one review pass.

Skip for a trivial, reversible, well-understood change. Adversarial review on a typo hallucinates flaws & wastes the budget — the self-critique paradox is real.

PHASE 0 — CAPTURE

Read the spec: §G §C §I §R §V §T. Hold the whole thing. You review the spec, not your memory of the conversation.

PHASE 1 — CONSTRUCT THE SENIOR

Build a reviewer with real authority, not a generic critic:

  • Codebase — grep/read the modules this spec touches. What patterns, what invariants already hold?
  • §R — what did research establish? A spec decision that contradicts §R is a finding.
  • Live — for any best-practice claim you are unsure of, fetch it. An out-of-date assumption is a flaw.

A reviewer with no evidence is just an opinion. Earn the authority first.

Show full SKILL.md (223 more words)Show less

PHASE 2 — REFUTE

Attack the spec on these axes. For each, try to find the case where it breaks:

  • Goal vs reality — does §G solve the actual problem, or a proxy?
  • Missing invariant — what can go wrong that no §V catches? (most findings live here)
  • Interface drift — does §I match what callers already expect? (cite the caller, file:line)
  • Constraint conflict — do two §C bullets contradict? does one fight §R?
  • Unowned edge — the input, ordering, failure, or concurrency case no §T covers.
  • Altitude — §T too vague to act on, or so granular it is just typing?

PHASE 3 — CLASSIFY

Each finding: evidence → claim → severity.

  • BLOCK — build on this spec ships a real defect. Must fix first.
  • HARDEN — add/sharpen a §V so the build cannot regress it.
  • NOTE — worth knowing, not blocking.

No evidence? Down-rank to NOTE & tag [unverified]. ⊥ inflate a hunch to BLOCK.

PHASE 4 — HARDEN §V & GATE

  • Each HARDEN finding → a draft §V line (testable, cites the §I/behavior it guards). Hand to spec to write.
  • End on an explicit gate:
## review verdict
BLOCK: 1 — §I.api shape ≠ caller src/client.ts:40. fix §I before build.
HARDEN: 2 — drafted V8 (idempotent refund), V9 (tx around dual write).
NOTE: 1 — §T4 vague, split before /build.
gate: NO-GO until BLOCK cleared. then /build §T after spec writes V8,V9.

GO or NO-GO, never a shrug. Review is the checkpoint that stops a confident wrong build.

BOUNDARIES

  • ⊥ write SPEC.md. Draft §V & hand to spec.
  • ⊥ pass a finding with no evidence as fact. Flag [unverified].
  • ⊥ review trivia. Right-size or skip.
  • ⊥ rewrite the user's intent. You harden the spec, you do not replace its goal.

© JuliusBrussee, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/review of JuliusBrussee/cavekit.

Open the folder on GitHubat commit 7421e87

Compare with similar skills

Adversarial Spec Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adversarial Spec Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adversarial Spec Review this skillJuliusBrussee/cavekit1.2k—~959Automated safety check: PassMIT
Backend Code Reviewlanggenius/dify158k—~676Automated safety check: PassCustom licence
Code Review Skillawesome-skills/code-review-skill2.1k—~2.8kAutomated safety check: NotesMIT
Pascal Architecture PR Reviewpascalorg/editor25k—~7.5kAutomated safety check: PassMIT
Brooks Audithyhmrright/brooks-lint1.5k1 repos~537Automated safety check: PassMIT
Code Review SkillRain-kl/OpenFlare289—~2.3kAutomated safety check: NotesMIT

Similar skills

  • Backend Code Review

    langgenius/dify

    Reviews backend code under api/ for concrete, reproducible defects, routes to rule packs for architecture, schema, repositories and SQLAlchemy, and ranks findings from P0 to P3.

    158k GitHub stars~676 tokensUpdated today
    DevelopmentAuto-check passed
  • Code Review Skill

    awesome-skills/code-review-skill

    Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, Java 8, PHP, Ruby, Rails, Python, Django, FastAPI, Go, C/.NET, Kotlin, Swift, Dart…

    2.1k GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Reviews a pull request against the Pascal editor's architectural rules: package boundaries, registry-driven node composition, hook hygiene and selector performance.

    25k GitHub stars~7.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Brooks Audit

    hyhmrright/brooks-lint

    Architecture audit that maps module dependencies, checks layering integrity, and flags structural decay across a codebase, drawing on twelve classic engineering books.

    1.5k GitHub starsUsed in 1 repo~537 tokens
    DevelopmentAuto-check passed
  • Code Review Skill

    Rain-kl/OpenFlare

    Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, PHP, Python, Django, Go, C/.NET, Kotlin, Swift, NestJS, C/C++, and more.

    289 GitHub stars~2.3k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Pattern Conformance Audit

    Totoro-jam/battle-tested-patterns

    Audits a codebase's existing patterns, such as rate limiters, circuit breakers and caches, against canonical invariants and flags mislabeled or divergent ones.

    345 GitHub stars~1.6k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from JuliusBrussee/cavekit

All 8 skills in this repo
  • Backprop: Bug-to-Spec Protocol

    JuliusBrussee/cavekit

    After a bug is found, traces its root cause and feeds a new testable invariant back into the project spec so the bug class can't recur.

    1.2k GitHub stars~653 tokensUpdated 1 mo ago
    Auto-check passed
  • Caveman Spec Compression

    JuliusBrussee/cavekit

    Compresses SPEC.md writes and spec-referencing prose into terse, symbol-heavy fragments that drop articles, filler and hedging while keeping facts intact.

    1.2k GitHub stars~721 tokensUpdated 1 mo ago
    Auto-check passed
  • Spec Drift Check

    JuliusBrussee/cavekit

    Read-only detector that compares SPEC.md with the code and reports invariant, interface and task drift grouped by severity, without changing anything.

    1.2k GitHub stars~666 tokensUpdated 1 mo ago
    Auto-check passed
  • Deepen Module Design

    JuliusBrussee/cavekit

    Scans the code a spec touches for its shallowest module, then proposes a refactor that hides more behind a smaller interface without changing behavior.

    1.2k GitHub stars~1k tokensUpdated 1 mo ago
    Auto-check passed
  • Grill Before Spec

    JuliusBrussee/cavekit

    Interrogates a vague idea one question at a time, recommending an answer each round and recording results as goals and constraints before a spec is written.

    1.2k GitHub stars~812 tokensUpdated 1 mo ago
    Auto-check passed
  • Research

    JuliusBrussee/cavekit

    Gather external knowledge the spec needs and distill it into §R — the durable research log — so build grounds in facts instead of hallucinating library behavior.

    1.2k GitHub stars~782 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Adversarial Spec Review

What does Adversarial Spec Review do?

Builds a skeptical reviewer grounded in the codebase and research notes to try to refute a spec before any code is written, citing file:line evidence and ending in a go or no-go gate. The skill treats a review as an attempt at refutation, not approval: every finding must cite evidence, a file and line or a source, and a flaw that cannot be proven is downgraded to an unverified note rather than waved through or inflated into a blocker. It applies before building a high-blast-radius change, such as one touching a shared module, auth, money or a public API, or when a spec's interface or invariants section affects dependent code, and it is explicitly skipped for a trivial, reversible, well-understood change, since adversarial review on a typo wastes the effort budget.

When should I use Adversarial Spec Review?

Adversarial Spec Review fits situations like: reviewing a spec before building a change to a shared module, auth or payments; red-teaming a plan before committing engineering time to it; deciding whether a spec is sound enough to move to a go or no-go gate.

How do I install Adversarial Spec Review in Claude Code?

Run `npx skills add JuliusBrussee/cavekit --skill review -a claude-code`. Or copy the skill folder (skills/review in JuliusBrussee/cavekit) into .claude/skills/review in your project. Claude Code loads it when a task matches its description.

How do I install Adversarial Spec Review in Codex?

Run `npx skills add JuliusBrussee/cavekit --skill review -a codex`. Or copy the skill folder (skills/review in JuliusBrussee/cavekit) into .agents/skills/review in your project. Codex loads it when a task matches its description.

Can I use Adversarial Spec Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JuliusBrussee/cavekit --skill review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review, .gemini/skills/review, .github/skills/review and .opencode/skills/review in your project.

What does Adversarial Spec Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Adversarial Spec Review is instructions for the agent only. Our summary lists: A written spec with the sections this review reads.

Does Adversarial Spec Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Adversarial Spec Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adversarial Spec Review use?

Adversarial Spec Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adversarial Spec Review use?

About 959 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adversarial Spec Review?

Skills that share tags, products or a category with Adversarial Spec Review: Backend Code Review (langgenius/dify, 158k stars), Code Review Skill (awesome-skills/code-review-skill, 2.1k stars), Pascal Architecture PR Review (pascalorg/editor, 25k stars) and Brooks Audit (hyhmrright/brooks-lint, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adversarial Spec Review?

JuliusBrussee (a GitHub user) maintains it in JuliusBrussee/cavekit, which has 1,151 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on August 14, 2026.

Source: JuliusBrussee/cavekit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.