Agent skill

Agreement-Based Code Review

by phuryn in phuryn/pm-skills

Reviews a diff by anchoring on agreements between two sides of a boundary, forcing a concrete violating execution, and refuting each finding before reporting it.

MITAuto-check passedDevelopment

Install Agreement-Based Code Review

skills CLI
$ npx skills add phuryn/pm-skills --skill code-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install phuryn/pm-skills code-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/phuryn/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pm-ai-shipping/skills/code-review .claude/skills/code-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review
GitHub stars
27k
Token cost
~3.6k tokens
SKILL.md length
1,887 words
Files
4 (incl. references)
Skills in repo
62
Repo updated
First seen
Licence
MIT

At a glance

Reviews a diff by anchoring on agreements between two sides of a boundary, forcing a concrete violating execution, and refuting each finding before reporting it.

  • Works in 2 steps: Authority reconciliation. Follow a… → Identity and correlation. Trace how an…
  • Reviewing a diff or pull request for correctness defects before merge
  • SKILL.md covers Purpose, Structure: one engine, three…, Invocation and Shared engine, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Treats the review's basic unit as an agreement between two participants, such as a caller and a callee or a writer and a later reader, rather than a single file, on the reasoning that file-by-file checklists miss the defects that only show up when both sides of that agreement are compared at once. Correctness is the core, default dimension and is described in full; performance and security are run as sub-cases of the same engine, each reading its own short reference file only when selected, since the core discipline and report format carry over unchanged.

Every reported finding must survive a strong counterargument that the review already checked, and must carry a required behavior, a feasible way to trigger the violation, a concrete contradiction, and an observable consequence, so the output stays small and every item is something the reader can act on rather than a pile of things that merely look wrong. One root cause touching more than one dimension, say correctness and security together, is reported once with both impacts rather than duplicated.

Invocation accepts an explicit list of dimensions or `all` for every sub-case, while a bare request to review or find bugs defaults to correctness alone; Claude Code's own bundled review command is a separate thing, so the plugin-qualified invocation is used when this skill specifically is meant.

When your agent uses it

  • Reviewing a diff or pull request for correctness defects before merge
  • Auditing a change for performance regressions under growth or contention
  • Checking a fix for a security hole with an attacker-controlled source
  • Finding defects that only show up when two sides of an agreement disagree

Example prompts

  • “/pm-ai-shipping:code-review this diff for correctness before I merge it.”
  • “Run a security-focused review on the new file upload endpoint.”
  • “Review both correctness and performance on this caching change.”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Authority reconciliation. Follow a proposed value through validation, normalisation,
  2. Identity and correlation. Trace how an operation's result finds its originating entity, then

What it can do on your machine

Read from SKILL.md and the folder at commit 8607e3b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agreement-Based Code Review loads about 3.6k tokens when it runs, and up to ~7.6k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 1,887 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from phuryn/pm-skills at commit 8607e3b, republished under its MIT licence (© phuryn). 1,887 words, ~3,599 tokens.

Download SKILL.mdSave it as .claude/skills/code-review/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
code-review
description
Review code for actionable defects. Correctness is the core; performance and security are optional sub-cases of the same engine. Anchors on agreements between participants across a boundary, forces a violating execution, and refutes every candidate before reporting. Use when asked to review changes, find bugs, audit a codebase, or check whether a fix is safe.

Code Review

Purpose

Most review output is noise: a list of things that look wrong, unranked, unrefuted, and impossible to act on. This skill produces the opposite — a small number of findings, each with a required behaviour, a feasible trigger, a concrete contradiction, an observable consequence, and the strongest counterargument already checked.

Its central bet: the defects reviewers miss are rarely visible inside one file. They are disagreements between two participants that each look reasonable alone — a caller and a callee, a producer and a consumer, a writer and a later reader, two branches that should establish the same state. A checklist applied file-by-file cannot see those, because the two halves are never in view at the same time. So the unit of review here is the agreement, not the file.

Structure: one engine, three anchors

Code review is the skill. Correctness is its core — the dimension generic tooling covers worst, and the one described in full below. Performance and security are sub-cases: the same engine, the same refutation discipline, the same report contract, with a different anchor and one or two extra rules each.

Sub-caseAnchorWhere its rules live
Correctness (core, default)Agreements between participants across a boundaryThis file + references/correctness-taxonomy.md
PerformanceWorkload → resource demand → growth or contention → consequencereferences/performance-review.md
SecuritySource → trust boundary → sink, with an attacker controlling the sourcereferences/security-review.md

Read a sub-case's file only when that sub-case is selected. Each is short on purpose: it states what differs, and the rest of this file still applies.

Sub-cases are independently activated, not mutually exclusive. One root cause can carry correctness and security impact — report it once, with both impacts.

Invocation

/pm-ai-shipping:code-review
/pm-ai-shipping:code-review dimensions=correctness scope=changes
/pm-ai-shipping:code-review dimensions=performance,security
/pm-ai-shipping:code-review dimensions=all

Claude Code ships its own bundled /code-review. Use the plugin-qualified form above when you mean this one.

These are instruction arguments, not shell flags.

  • Default: correctness. Bare "review this" or "find bugs" means correctness only.
  • An explicit list selects exactly those sub-cases; all selects three. Never silently reinterpret an unknown or empty selection — ask.
  • Scope: use what was asked. Otherwise review working changes if present, else the repository.
  • State the selected dimensions, the scope and the comparison baseline before investigating.
  • Reviewing changes means following dependencies beyond the changed lines, and distinguishing defects the change introduced from defects it merely revealed.
  • Review and report. Apply fixes only when asked.

Shared engine

Every sub-case uses one skeleton. Only the anchor and the refutation rules differ.

Map a flow → identify an obligation → inspect every participant → construct a violating execution → trace the consequence → attempt refutation → report.

Build one minimal map first: inputs, major execution flows, who owns which state, external dependencies, observable effects. Each selected sub-case enriches it — do not build three maps, and do not make a security-only run wait on correctness mapping.

Correctness: the agreement engine

A boundary is semantic, not a file split. It separates a caller and a callee, two callbacks, two executions of the same function, a producer and a consumer, or a value written now and read later.

For each consequential agreement, hold these in working notes — not in the report:

Participants:
Value, entity or effect exchanged:
Authority (who decides the real answer):
Identity and lifetime/version:
Required relationship:
Evidence for that relationship:
Relevant transitions or orderings:
Observable consumer or consequence:

Establish the obligation without inventing intent. Evidence comes from specifications, documented contracts, language or protocol semantics, tests that encode an expectation, or a necessary producer/consumer relationship. A consumer's implementation alone does not prove the consumer is right. Where participants disagree, say why the disagreement produces a wrong outcome — sometimes the contradiction is certain while which side should change is genuinely open. Missing documentation is a limitation, not automatically a finding.

Start where agreements are most likely to break: values transformed or negotiated, identities reassigned, work becoming asynchronous, state persisted and reloaded, several effects that must agree. Then do a local pass over ordinary decisions, arithmetic, boundaries and error branches — the anchor must not become a filter that discards plain bugs.

Force a violating execution

A suspicion is not a finding until you construct the execution that breaks it. Where the implementation permits:

  • make a requested value differ from the accepted or effective one;
  • keep two operations live at once and vary their completion order;
  • change the relevant identity or generation between observation and use;
  • compare distinct transitions that should end in equivalent state;
  • inject failure between effects, and interruption before completion;
  • exercise empty, exact-boundary and adjacent-boundary inputs.

Establish that each case is actually reachable. Do not assume it.

Two lenses that need a forced probe, not a mention

Across a large evaluation of planted runtime defects in real codebases, two classes were almost never even reported by strong agents — not missed at the fix, missed at the look. Naming them in a checklist will not help; each needs an explicit probe:

  1. Authority reconciliation. Follow a proposed value through validation, normalisation, negotiation or commit, and find downstream state still derived from the proposal where the authority can return something different. A requested value is not an applied value. Probe: force them apart and ask what still reads the request.
  2. Identity and correlation. Trace how an operation's result finds its originating entity, then establish that the key is unique, stable and live for long enough — under overlap, reordering, removal and reuse. A label, a position or arrival order is suspicious exactly when those properties can fail. Probe: run two operations concurrently and complete them out of order.

The full set of thirteen diagnostic lenses, each with a detection tell, is in references/correctness-taxonomy.md. They are overlapping lenses, not a quota to fill.

Refutation: the discipline that makes this worth running

A candidate becomes a finding only with all five:

  1. A supported obligation — what must hold, and on what evidence.
  2. A feasible execution — inputs, state and ordering the real system permits.
  3. A concrete contradiction — where the obligation fails.
  4. An observable consequence — wrong output, state, effect, completion or progress.
  5. An examined counterargument — the strongest mechanism that would prevent or repair it.

Actively hunt for the refutation: an enclosing guarantee that makes the execution impossible; synchronisation excluding the interleaving; reconciliation before any consequential read; an intentional contract; a precondition excluding the input; a different owner responsible for it.

OutcomeRule
KeepEvidence establishes the defect; the counterargument checked does not prevent it.
DropCited evidence defeats the execution, the obligation or the consequence.
UnresolvedAn essential contract or runtime fact is unknown. List it separately from findings.

Do not import the security sub-case's attacker/victim test into correctness. A correctness defect can harm only the person who triggered it and still be serious. Equally, "keep unless disproved" is too permissive here — an ungrounded suspicion with no constructed execution is not a finding. When both sub-cases are active, apply each test only to its own dimension.

Absorption is not prevention. The most expensive refutation mistake is finding something downstream that happens to hide the defect - a cache that usually holds the value, a retry that usually succeeds, a default that is usually right - and dropping the finding. That is not a guarantee, it is a coincidence with good odds, and it fails the day the absorber is cold, evicted or reconfigured. Drop only on a mechanism that makes the execution impossible, and say which mechanism it was. For the same reason, "it works nearly always" describes a race, not a refutation - a timing window that usually resolves correctly is a finding, and the fact that you had to reason about which side usually wins is the evidence.

Passing tests, unfamiliar code, a suspicious name, a missing test and a sibling difference are evidence to investigate — none of them is proof, and none is refutation. Deduplicate by violated agreement and root cause, never by file. There is no findings quota; zero supported findings is a valid result.

Show full SKILL.md (615 more words)Show less

Parallelism

Fan out over complete flows or connected groups of agreements — never over files, and never one agent per taxonomy class. Partitioning by file is precisely the split that hides cross-boundary defects, which are the ones worth finding.

  1. The coordinator builds the initial map and identifies shared state.
  2. Each worker gets a bounded flow, its participants, the selected sub-cases and open questions.
  3. Workers inspect both sides of their agreements and may follow dependencies outside their list.
  4. Workers return candidates, cited evidence, completed refutations and unresolved relationships.
  5. The coordinator reconciles assumptions and any relationship that crosses assignments.
  6. Strong candidates get a separate verification pass before they are reported.

Allow overlapping reads. Two workers reading the same authority is far cheaper than either one holding half its contract. Keep integration capacity in reserve: an unresolved relationship spanning two assignments stays unexamined until someone closes it. One level of fan-out is the target; if delegation is unavailable or the scope is small, run the same procedure sequentially.

Run workers on the strongest model available, and match the current session's effort level. This is recall-first work: a missed cross-boundary flow is the costly failure, and a worker that silently drops to a cheaper model or a lower effort is the cheapest way to lose one. If any worker is rerouted or downgraded, say which in the report — a reader who assumes one model saw everything will misjudge the coverage.

One model. Name it on every worker. Fan-out here buys coverage, not a second opinion. Pass the coordinator's own model explicitly on each spawn — "inherit" is not a routing decision, and a worker that quietly lands on a cheaper model is the easiest way to lose a finding. Do not bring in a different model, to review or to cross-check, unless you are explicitly asked: mixing models makes the result unattributable, and when this skill is being measured or compared across models, one foreign worker invalidates the number. The independent second-model pass is a separate, explicitly-invoked step (/ship-check Step 6), never something this skill reaches for on its own.

Tell workers they are reading, not editing. A review worker needs to read, search and navigate; it must not modify the tree. State that in the worker's instructions — a worker that starts editing drifts from reviewing into "helpfully" fixing and stops reporting what it silently repaired, and the findings can no longer be checked against the code they describe. Say it in the prompt rather than assuming the host will enforce it, and confirm the tree is unchanged when the run ends.

Report

Lead with supported findings, ordered by impact. Keep severity separate from evidential strength.

Review scope:
Comparison baseline:
Selected dimensions:

[Severity] [Dimension] Concrete consequence
  Expectation:  required behaviour, and the evidence for it
  Trigger:      feasible preconditions and execution
  Defect:       the violated relationship
  Evidence:     source locations for EVERY participant
  Impact:       observable consequence and affected scope
  Refutation:   strongest counterargument checked, and why it fails
  Remedy:       minimal correction to the violated relationship
  Verification: what was executed, versus established from source

Coverage:
Unexamined areas and essential unknowns:

Cite both participants for a cross-boundary defect, and do not group findings only by file — that hides the relationship the review exists to find.

Coverage means work performed, not boxes ticked. For each selected sub-case report: examined with supported findings · examined, none supported · not applicable, with reason · not examined, with reason. Zero findings in a category does not mean "not covered", and a table of ticks is not evidence of completeness. Say "no supported findings in the examined scope" — never that the code is bug-free.

Notes

  • Say explicitly what is well built. A review that only accuses is easy to dismiss.
  • The two sub-cases have mature commands behind them: /security-audit-static (trust boundaries, sinks, OWASP backstop) and /performance-audit-static (over-fetching, indexes, caching). Run the command when the sub-case is the whole job; use the reference file when it is one dimension of a broader review. This skill does not restate either.
  • For the doc-vs-code axis use the intended-vs-implemented skill.
  • A static review produces code-review findings, not confirmed exploits or measured regressions.

© phuryn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in pm-ai-shipping/skills/code-review of phuryn/pm-skills.

  • SKILL.md
  • references/correctness-taxonomy.md
  • references/performance-review.md
  • references/security-review.md

Open the folder on GitHubat commit 8607e3b

Compare with similar skills

Agreement-Based Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agreement-Based Code Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agreement-Based Code Review this skillphuryn/pm-skills27k—~3.6kAutomated safety check: PassMIT
Fix The Classjoetawil7/first-pass112—~1.5kAutomated safety check: PassMIT
Devnpc-live/clawfirm156—~642Automated safety check: PassNone
Git History Bug Auditben-manes/caffeine18k—~3.3kAutomated safety check: PassApache-2.0
Code Review Graph Navigatorhandsontable/handsontable22k—~939Automated safety check: PassCustom licence
RoamCranot/roam-code517—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Fix The Class

    joetawil7/first-pass

    Bug-fix routine that fixes the whole class of bug, not just the reported instance.

    112 GitHub stars~1.5k tokensUpdated 5 days ago
    DevelopmentAuto-check passed
  • Dev

    npc-live/clawfirm

    Software development workflow dispatcher. An agent skill from npc-live/clawfirm.

    156 GitHub stars~642 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Code Review Graph Navigator

    handsontable/handsontable

    Queries a pre-built, Tree-sitter-based code graph of the whole monorepo instead of grepping call chains, for exploring, debugging, refactoring or reviewing code.

    22k GitHub stars~939 tokensUpdated today
    DevelopmentAuto-check passed
  • Roam

    Cranot/roam-code

    Codebase comprehension via roam-code CLI. An agent skill from Cranot/roam-code.

    517 GitHub stars~2.4k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Code Review

    aide-family/moon

    Reviews code for correctness and potential bugs, pinpoints bug locations by file and line, and suggests concrete fixes.

    253 GitHub stars~815 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed

More from phuryn/pm-skills

All 62 skills in this repo
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated 22 days ago
    Auto-check passed
  • Team OKR Brainstorm

    phuryn/pm-skills

    Drafts three alternative sets of team OKRs, each with an inspiring objective and measurable key results, tied to the company strategy you provide.

    27k GitHub stars~1.1k tokensUpdated 22 days ago
    Auto-check passed
  • Analyzes uploaded cohort data to compute retention curves and feature adoption trends, builds heatmaps and charts, and suggests qualitative follow-up research.

    27k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed
  • Dummy Dataset Generator

    phuryn/pm-skills

    Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script.

    27k GitHub stars~983 tokensUpdated 22 days ago
    Auto-check passed
  • Grammar and Flow Checker

    phuryn/pm-skills

    Reviews a draft for grammar, logic and flow problems and returns located, prioritized fix suggestions without rewriting the whole text.

    27k GitHub stars~2.4k tokensUpdated 22 days ago
    Auto-check passed
  • Builds a structured customer interview script with opening, warm-up, jobs-to-be-done exploration and wrap-up sections, following Mom Test rules against leading questions.

    27k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed

Questions about Agreement-Based Code Review

What does Agreement-Based Code Review do?

Reviews a diff by anchoring on agreements between two sides of a boundary, forcing a concrete violating execution, and refuting each finding before reporting it. Treats the review's basic unit as an agreement between two participants, such as a caller and a callee or a writer and a later reader, rather than a single file, on the reasoning that file-by-file checklists miss the defects that only show up when both sides of that agreement are compared at once. Correctness is the core, default dimension and is described in full; performance and security are run as sub-cases of the same engine, each reading its own short reference file only when selected, since the core discipline and report format carry over unchanged.

When should I use Agreement-Based Code Review?

Agreement-Based Code Review fits situations like: reviewing a diff or pull request for correctness defects before merge; auditing a change for performance regressions under growth or contention; checking a fix for a security hole with an attacker-controlled source; finding defects that only show up when two sides of an agreement disagree.

How do I install Agreement-Based Code Review in Claude Code?

Run `npx skills add phuryn/pm-skills --skill code-review -a claude-code`. Or copy the skill folder (pm-ai-shipping/skills/code-review in phuryn/pm-skills) into .claude/skills/code-review in your project. Claude Code loads it when a task matches its description.

How do I install Agreement-Based Code Review in Codex?

Run `npx skills add phuryn/pm-skills --skill code-review -a codex`. Or copy the skill folder (pm-ai-shipping/skills/code-review in phuryn/pm-skills) into .agents/skills/code-review in your project. Codex loads it when a task matches its description.

Can I use Agreement-Based Code Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add phuryn/pm-skills --skill code-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review, .gemini/skills/code-review, .github/skills/code-review and .opencode/skills/code-review in your project.

What does Agreement-Based Code Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Agreement-Based Code Review is instructions for the agent only.

Does Agreement-Based Code Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agreement-Based Code Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agreement-Based Code Review use?

Agreement-Based Code Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agreement-Based Code Review use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Agreement-Based Code Review?

Skills that share tags, products or a category with Agreement-Based Code Review: Fix The Class (joetawil7/first-pass, 112 stars), Dev (npc-live/clawfirm, 156 stars), Git History Bug Audit (ben-manes/caffeine, 18k stars) and Code Review Graph Navigator (handsontable/handsontable, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agreement-Based Code Review?

phuryn (a GitHub user) maintains it in phuryn/pm-skills, which has 26,809 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on September 14, 2026.

Source: phuryn/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.