Agent skill

Adversarial Review

by boraoztunc in boraoztunc/skills

Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass.

Apache-2.0Auto-check passedDevelopment

Install Adversarial Review

skills CLI
$ npx skills add boraoztunc/skills --skill adversarial-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install boraoztunc/skills adversarial-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/boraoztunc/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/adversarial-review .claude/skills/adversarial-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adversarial-review
GitHub stars
398
Token cost
~2.3k tokens
SKILL.md length
1,255 words
Files
2
Skills in repo
52
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass.

  • Works in 5 steps: Stabilize the target → Discover the invariants → Discover the domains → …
  • The user wants a deep
  • SKILL.md covers The stance (this is the whole…, Method, Scaling and Anti-patterns (don't do these)
  • Runs JavaScript scripts from its folder; calls git

What it does

Adversarial Review is an agent skill from boraoztunc/skills. Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass. Use when the user wants a deep, skeptical review of a diff, PR, or branch ("adversarial review", "try to break this", "find what's wrong before I ship", "review this like a skeptic", "what will break in production"). Scales from a single inline pass to a multi-agent fan-out (one skeptic per domain → adversarial verify → synthesis). Distinct from a normal code review: it…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `workflow.js`).

It sits in Development. The repository describes itself as: Claude Code skills for copywriting, SEO, design, and more. The licence is Apache-2.0.

When your agent uses it

  • The user wants a deep
  • Skeptical review of a diff
  • Branch (adversarial review
  • Try to break this

Example prompts

  • “adversarial review”
  • “try to break this”
  • “find what”
  • “/adversarial-review”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Stabilize the target
  2. Discover the invariants
  3. Discover the domains
  4. Verify before you believe (the adversarial pass)
  5. Synthesize honestly

What it can do on your machine

Read from SKILL.md and the folder at commit 645553c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adversarial Review loads about 2.3k tokens when it runs. Until then it costs about 163 tokens; SKILL.md has 1,255 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from boraoztunc/skills at commit 645553c, republished under its Apache-2.0 licence (© boraoztunc). 1,255 words, ~2,259 tokens.

Download SKILL.mdSave it as .claude/skills/adversarial-review/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
adversarial-review
description
Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass. Use when the user wants a deep, skeptical review of a diff, PR, or branch ("adversarial review", "try to break this", "find what's wrong before I ship", "review this like a skeptic", "what will break in production"). Scales from a single inline pass to a multi-agent fan-out (one skeptic per domain → adversarial verify → synthesis). Distinct from a normal code review: it leads with concrete failure scenarios over style, distrusts the tests, and is honest about what it could not verify.

Adversarial Review

A normal review confirms a change looks right. An adversarial review assumes it is broken and tries to prove it — then refuses to trust its own findings until each one survives an independent attempt to refute it. The output is not "looks good" or a list of nitpicks; it is a ranked set of concrete failure scenarios that withstood scrutiny, plus an honest account of what could not be checked.

Use this when the cost of a missed bug is high: before shipping, on a large or half-finished diff, on a new code path with no test history, or any time the user says "try to break this."

The stance (this is the whole skill)

Adopt these as hard rules for the duration of the review:

  1. Assume it's broken until proven otherwise. A review that finds nothing is a failed review unless you can show you tried hard to break it. Your job is to find the input that crashes it, the sequence that corrupts data, the change that silently regresses a fixed bug.
  2. Distrust the tests. Passing tests prove the tested paths work — not that the change is correct. Ask what isn't tested. Look for assertions that would still pass if the code were wrong, or tests weakened to make the diff go green.
  3. Distrust comments and commit messages. Review what the code does, not what it claims. If a comment says "killed on exit," verify the kill actually fires.
  4. Every finding is a failure scenario, not an opinion. "This could be cleaner" is not a finding. "This panics on a 0-byte file because extractMeta unwraps the page count" is. Each finding needs: concrete inputs/sequence → observed bad outcome, and a file:line.
  5. Lead with correctness. Data loss, crashes, regressions, security, race conditions first. Quality/clarity findings are a separate, lower section.
  6. Be honest about coverage. Never imply you checked more than you did. End every review with what you could not verify (paths you couldn't reproduce, anything that needs the app running, cited lines you couldn't confirm).

Method

1. Stabilize the target

Pin down exactly what you're reviewing so findings cite a stable state and the user can revert cleanly.

  • Working-tree changes: snapshot them onto a branch (git checkout -b review/<topic> && git add -A && git commit), or note the exact git diff range. This prevents findings from drifting as the tree changes and lets the user restore easily.
  • A PR: review git diff <base>...<head> (three-dot — the changes the PR introduces).
  • Read the full diff plus enough surrounding code to know the invariant each hunk touches. Never review a hunk in isolation.
2. Discover the invariants

The highest-value bugs violate a rule the codebase already established. Before reviewing, spend a few minutes learning the project's rules — do not invent generic ones:

  • Read CLAUDE.md / AGENTS.md / README / docs/ for documented invariants, "this was the bug, keep it" notes, architecture rules, and gotchas.
  • Skim recent commit messages and the tests for behavior the project treats as load-bearing.
  • A diff that violates a documented invariant is a bug, even if it compiles and tests pass. Cite the invariant in the finding.
3. Discover the domains

Split the diff into independent failure domains so each gets focused, deep attention instead of one shallow context-switching pass. Derive the split from this diff — common cuts:

  • By language / layer (backend vs frontend vs schema/migration).
  • By brand-new code (a fresh module/importer/parser with no test history is its own domain — it's where invariants are most likely unhonored).
  • By cross-cutting boundary (API/IPC contracts: a signature changed on one side but not the other — compiles in each language, breaks at the boundary; a new command not registered; a missing capability/permission; docs claiming behavior the code lacks).

For each domain, hunt for the failure modes that domain is prone to. A starting checklist (extend per project):

  • Data integrity: idempotency (re-run duplicates rows?), dedup keys that don't compose, partial writes on error, schema/contract drift, nullable fields assumed non-null.
  • Resource & lifecycle: leaked processes/handles/listeners, per-request clients that should be shared, work that should stop on cancel/unmount but doesn't, O(n²) hidden in a loop, locks held across I/O or parsing.
  • Network & trust: SSRF (guard the initial host and every redirect hop), missing size/time caps, unvalidated input reaching a query (injection), secrets cleared on a transient failure when they should be kept.
  • Concurrency / state: stale closures, missing effect deps, optimistic UI desyncing from the persisted write, double-invocation under StrictMode, races between batches.
  • Input robustness: empty / huge / malformed / Unicode / multibyte input; offsets computed on one string and applied to another; unwraps/panics on bad input.
Show full SKILL.md (494 more words)Show less
4. Verify before you believe (the adversarial pass)

This is what separates an adversarial review from a thorough one. Do not report a finding until you've tried to refute it. For each candidate finding, take the opposite side: read the cited code and surrounding context and ask —

  • Does the failure actually reproduce? Construct the concrete repro if you can.
  • Is there an existing guard the first pass missed? Is the cited line even reachable?
  • Is the severity inflated?

Default to refuted unless you can confirm the exact failure path. When the domains are large or correctness matters a lot, give each finding to a fresh perspective (a separate agent, or a deliberate second pass) prompted to refute it — independence catches the plausible-but-wrong findings that confirmation bias keeps. For findings that can fail in more than one way, verify along distinct lenses (does-it-reproduce / security / does-the-guard-hold) rather than re-checking the same way N times.

Only findings that survive refutation go in the report. Track refuted ones too — "I checked X and it's actually guarded at line N" is useful signal.

5. Synthesize honestly

Produce a tight report:

  1. One-line verdict on the change's overall risk.
  2. Findings by severity (Critical → Low). Each: title, file:line, the concrete failure scenario, the violated invariant/expectation, and a one-line fix direction. Flag low-confidence survivors as "needs a human check."
  3. Single most likely thing to break in production — the one finding that fires on the most common path, even if it's not the most severe.
  4. What could not be verified — paths you couldn't reproduce, anything needing the app running, cited lines you couldn't confirm, areas out of scope. Be specific; never imply coverage you didn't achieve.

Scaling

Match effort to the request and the stakes:

  • Quick / small diff: one inline adversarial pass — find, then a deliberate refute pass on each candidate, then synthesize. No orchestration needed.
  • Larger diff or "be thorough / audit this": spawn one skeptic agent per discovered domain (use the Agent tool, or a Workflow if the user has opted into multi-agent orchestration), each loaded with the discovered invariants; verify each finding with an independent refuter; synthesize the survivors. A parameterized Workflow template that encodes exactly this — discover → per-domain skeptics → adversarial verify → synthesis — is in workflow.js. Treat it as a starting point: fill in the discovered domains and invariants, don't ship the placeholders.

Note: the Workflow path requires the user to have opted into multi-agent orchestration (e.g. "ultracode", "use a workflow", or an explicit ask). Without that, run the inline path or a handful of Agent calls — never silently fan out dozens of agents.

Anti-patterns (don't do these)

  • Reporting a finding you didn't try to refute.
  • Padding the report with style nits to look thorough — they bury the real bugs.
  • "Looks good to me" with no account of what you actually probed.
  • Trusting a green test suite as proof of correctness.
  • Inventing generic invariants instead of reading the project's actual rules.
  • Claiming coverage of areas you never opened.

© boraoztunc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in adversarial-review of boraoztunc/skills.

  • SKILL.md
  • workflow.js

Open the folder on GitHubat commit 645553c

Compare with similar skills

Adversarial Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adversarial Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adversarial Review this skillboraoztunc/skills398—~2.3kAutomated safety check: PassApache-2.0
Trellis Session Insightmindfold-ai/Trellis15k4 repos~1.7kAutomated safety check: PassAGPL-3.0
Warp Factory Fileswarpdotdev/warp65k1 repos~2.5kAutomated safety check: PassAGPL-3.0
Migrate Core Code to Submodulestinyhumansai/openhuman42k—~2.6kAutomated safety check: PassGPL-3.0
Analyze Logsactivepieces/activepieces25k1 repos~1.6kAutomated safety check: PassMIT
GitHub Review Iterationprisma/orm48k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Trellis Session Insight

    mindfold-ai/Trellis

    Reach into past AI conversation history through the trellis mem CLI.

    15k GitHub starsUsed in 4 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Warp Factory Files

    warpdotdev/warp

    Authors and edits file-based Warp software factory definitions rooted at factory.yaml, covering agents, automations, scorers and webhooks, and validates them before a pull request.

    65k GitHub starsUsed in 1 repo~2.5k tokens
    DevelopmentAuto-check passed
  • Migrate Core Code to Submodules

    tinyhumansai/openhuman

    Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.

    42k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Analyze Logs

    activepieces/activepieces

    Analyze application logs from the .evlog/logs/ directory. An agent skill from activepieces/activepieces.

    25k GitHub starsUsed in 1 repo~1.6k tokens
    DevelopmentAuto-check passed
  • Official

    Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.

    48k GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Release

    PrefectHQ/fastmcp

    Cut a FastMCP release end to end. An agent skill from PrefectHQ/fastmcp.

    28k GitHub stars~2.9k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from boraoztunc/skills

All 52 skills in this repo
  • Gsap

    boraoztunc/skills

    GSAP animation reference for HyperFrames. An agent skill from boraoztunc/skills.

    398 GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Hyperframes

    boraoztunc/skills

    Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.

    398 GitHub starsUsed in 9 repos~7.6k tokens
    Auto-check passed
  • Remotion To Hyperframes

    boraoztunc/skills

    Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition.

    398 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Animejs

    boraoztunc/skills

    Anime.js adapter patterns for HyperFrames. An agent skill from boraoztunc/skills.

    398 GitHub starsUsed in 2 repos~828 tokens
    Auto-check passed
  • Beam Glow States

    boraoztunc/skills

    Create React loading, processing, selected, current, focus, and pressed states with the border-beam package's animated edge glow.

    398 GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Minimal Zine Poster

    boraoztunc/skills

    Compile a theme, sentence, object, mood, article idea, or photo into a quiet Japanese/Korean zine-style editorial poster — tall aged paper, large negative space, one small image anchor, experimental…

    398 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Adversarial Review

What does Adversarial Review do?

Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass. Adversarial Review is an agent skill from boraoztunc/skills. Adversarially review a code change — assume it's broken, try to break it, and only report findings that survive an independent refutation pass.

When should I use Adversarial Review?

Adversarial Review fits situations like: the user wants a deep; skeptical review of a diff; branch (adversarial review; try to break this.

How do I install Adversarial Review in Claude Code?

Run `npx skills add boraoztunc/skills --skill adversarial-review -a claude-code`. Or copy the skill folder (adversarial-review in boraoztunc/skills) into .claude/skills/adversarial-review in your project. Claude Code loads it when a task matches its description.

How do I install Adversarial Review in Codex?

Run `npx skills add boraoztunc/skills --skill adversarial-review -a codex`. Or copy the skill folder (adversarial-review in boraoztunc/skills) into .agents/skills/adversarial-review in your project. Codex loads it when a task matches its description.

Can I use Adversarial Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add boraoztunc/skills --skill adversarial-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adversarial-review, .gemini/skills/adversarial-review, .github/skills/adversarial-review and .opencode/skills/adversarial-review in your project.

What does Adversarial Review need to run?

Going by SKILL.md and its folder, Adversarial Review needs JavaScript for the scripts in its folder and the command-line tools its instructions call (git). Our summary lists: Node.js.

Does Adversarial Review access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Adversarial Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adversarial Review use?

Adversarial Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adversarial Review use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adversarial Review?

Skills that share tags, products or a category with Adversarial Review: Trellis Session Insight (mindfold-ai/Trellis, 15k stars), Warp Factory Files (warpdotdev/warp, 65k stars), Migrate Core Code to Submodules (tinyhumansai/openhuman, 42k stars) and Analyze Logs (activepieces/activepieces, 25k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adversarial Review?

boraoztunc (a GitHub user) maintains it in boraoztunc/skills, which has 398 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on August 15, 2026.

Source: boraoztunc/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.