Agent skill

Code Review Discipline

by awebai in awebai/aweb

Practices to apply when taking a code-review assignment. An agent skill from awebai/aweb.

MITAuto-check passedDevelopment

Install Code Review Discipline

skills CLI
$ npx skills add awebai/aweb --skill code-review-discipline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awebai/aweb code-review-discipline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awebai/aweb.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/code-review-discipline .claude/skills/code-review-discipline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review-discipline
GitHub stars
115
Token cost
~2.2k tokens
SKILL.md length
1,196 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Practices to apply when taking a code-review assignment. An agent skill from awebai/aweb.

  • Tasks that involve Code review
  • SKILL.md covers Subagent confidence is…, Shim/helper adoption requires…, First-contact vs continuation… and Gate-shape sanity check, plus 1 more section
  • Calls uv
  • Tasks that involve Subagents

What it does

Code Review Discipline is an agent skill from awebai/aweb. Practices to apply when taking a code-review assignment. Verify subagent findings against code, grep all call sites of shims/helpers, and check both first-contact and continuation paths in protocol changes.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Code review and Subagents. The repository describes itself as: Communication for AI agents: stable identity, durable mail and chat, and wake-up events across sessions, runtimes, machines, and organizations. MIT, self-hostable. The licence is MIT.

When your agent uses it

  • Tasks that involve Code review
  • Tasks that involve Subagents

Example prompts

  • “Use the code-review-discipline skill to practice to apply when taking a code-review assignment. An agent skill from awebai/aweb”
  • “/code-review-discipline”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 272d187. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Review Discipline loads about 2.2k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 1,196 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from awebai/aweb at commit 272d187, republished under its MIT licence (© awebai). 1,196 words, ~2,204 tokens.

Download SKILL.mdSave it as .claude/skills/code-review-discipline/SKILL.md (or your agent's skills folder).
name
code-review-discipline
description
Practices to apply when taking a code-review assignment. Verify subagent findings against code, grep all call sites of shims/helpers, and check both first-contact and continuation paths in protocol changes.

Code review discipline

Load this skill at the start of any code review assignment. The lessons here come from real misses that shipped close to deploy and were caught by either another reviewer or a subagent's second pass.

Subagent confidence is uncorrelated with correctness

Running a code-reviewer subagent for a "second set of eyes" is cheap and sometimes catches what the primary reviewer missed. It also sometimes produces wrong-premise findings with confident framing. Banked data points from the aweb-aaou federation review chain:

  • One subagent pass returned 0 valid findings out of 5 — every finding was a wrong-premise about tool names or registration patterns that didn't exist in source.
  • Another subagent pass returned 2 valid findings — both were real misses the primary reviewer hadn't caught.

The subagent's confidence in BOTH cases sounded the same. Confidence is not a signal of correctness.

Practice: when a subagent surfaces a finding, treat it as a LEAD, not a conclusion. Open the actual file at the claimed line, verify the claim against code, then decide whether to act. If the subagent claims "X is missing," grep -n X first. If the subagent claims "Y is wrong," read Y in source first. Subagent output gets verified-against-code with the same scrutiny as any other review claim — including findings that match your own intuition.

This applies symmetrically: the subagent's wrong findings AND its valid findings both need code-level verification before they shape your review output.

Shim/helper adoption requires grep-all-call-sites

When a code change introduces a helper that other code SHOULD route through (an authentication wrapper, a federation shim, a fail-closed gate, a logging hook), reviewing the helper's factoring is necessary but not sufficient. The factoring being correct DOES NOT prove that every site needing it actually uses it.

Concrete miss from the aweb-aaou review chain:

  • mcp_federation_request shim introduced to let MCP tools call into route helpers for federation outbound. Shim factoring reviewed and signed off as correct — the abstraction shape, the dependency-injection seam, the call conventions all checked out.
  • BUT: only the FIRST-CONTACT branches of mcp/tools/mail.py and mcp/tools/chat.py were updated to use the shim. The CONTINUATION branches (conversation_id / session_id paths in the same files) still called deliver_message and send_in_session directly, bypassing federation entirely.
  • Result: a remote first-contact federated correctly via MCP, but the reply on the same conversation silently dropped to the sender's local DB. Recipient never saw it.

Practice: when reviewing a change that introduces a helper, grep ALL call sites that PLAUSIBLY need it, not just the call sites visible in the diff. A diff showing one branch updated does not prove the other branches were considered:

bash
grep -rn '<helper_name>' <relevant tree>     # all uses of the helper
grep -rn '<old_direct_call>' <relevant tree> # all direct calls the helper
                                              # was supposed to replace

The diff shows what changed; grep shows what should have changed. The difference is the review gap.

First-contact vs continuation as a review checkpoint

When reviewing a protocol or wire-format change, almost every flow has two stages:

  • Establish stage: first-contact, session-create, connect, handshake.
  • Extend stage: continuation, reply, in-session message, refresh, retry.

Reviewers naturally focus on the establish stage because that's where the new contract is most visible — first-contact carries the full envelope, the expensive checks (signature, cert, address resolution) all fire there. The extend stage gets less attention because it usually reuses state established earlier ("the conversation is already authorized; we just append to it").

This asymmetry is where wire-protocol changes most often leak. A federation contract that handles first-contact correctly but drops continuation isn't "federation that works minus a small edge case" — it's federation that LOOKS like it works in development but fails the first time a real user replies.

Practice: explicitly check BOTH stages when reviewing any change involving sessions, conversations, or stateful protocols. Use a two-column checklist mentally:

concernfirst-contact pathcontinuation path
signed_payload bindingverified?verified?
target re-resolutionverified?uses stored route?
auth / cert presentationverified?verified or scoped-not-required?
fail-closed on missing inputverified?verified?
dispatched via shim/helperverified?verified?

If either column has a blank, the review isn't done yet.

Show full SKILL.md (550 more words)Show less

Gate-shape sanity check

A passing CI gate is evidence of source-code correctness only when the gate actually exercises the code-under-test. The gate path can silently decouple the test environment from current source via Docker image builds, lockfile- pinned dependencies, cached test fixtures, or stale container layers.

Test before treating a gate signal as evidence: would this gate fail if I broke X in the source? If not, the gate isn't measuring X. The signal is decorative for the question "did my change work."

Specific failure modes seen:

  • Docker build + uv sync from lockfile → tests against pinned PyPI version, not local source. The aweb 1.23.0 federation work surfaced this: the ac release image copied sibling aweb sources, but uv sync still installed PyPI aweb==1.22.0 from uv.lock. The Docker user-journey gate ran against the stale package for the duration of the federation work; all "Docker e2e green" signals were false-evidence for source-level correctness.
  • Test fixtures hard-coding mocked behavior in places that should be live → test passes against stub, not implementation. (Related to the subagent-confidence pattern above: confidence in a validation signal needs the validation to actually measure the thing.)
  • Container build cache reuse → tests against pre-build state when the build step is the change-under-test.
  • Privileged fixture setup that bypasses the supported user path → test exercises the wire-level / library surface but not the supported entry point. The aweb-aaou federation v1 ship surfaced this: the OSS 2-server federation e2e passed 25/25 because the fixture direct-SQL-updated awid default_delivery_origin, which the supported namespace-controller CLI/API path did not yet expose. The test validated envelope flow once the route was set, but proved nothing about whether real operators (or hosted) could set the route through the supported surface. "Federation works" was true at the library layer and false at the operator surface the feature ships as.

General principle: a test that bypasses the supported user path with privileged fixture setup is measuring something narrower than the feature ships as. The library can be correct while the feature is operationally inert. Sanity-check by asking: if I removed the privileged fixture step, does the test still pass via the supported path? If not, the supported path is untested even when the test is green.

When loading this section: any review of CI gate failures or successes where the change touches a dependency-resolution path (Docker, lockfiles, container builds, npm registry resolution, cached fixtures). Also when a gate has been green for N runs and the change-under-test is non-trivial — ask: "what would have made it red?"

The discipline applies symmetrically to red and green signals. A red gate isn't evidence of broken code if the gate doesn't actually exercise the broken path; a green gate isn't evidence of working code if the gate doesn't actually exercise the changed code.

When to load this skill

  • Taking a code-review assignment that involves a wire protocol, an authentication change, or a new dispatch shape.
  • Running a code-reviewer subagent and considering whether to act on its findings.
  • Doing a "fresh eyes" pass on your own work because a customer-blocker or near-miss surfaced.
  • Reviewing a multi-commit bundle where a helper was introduced in one commit and consumed in subsequent commits.
  • Interpreting a CI gate signal where the change touches Docker images, lockfiles, container builds, package-manager resolution, or cached fixtures — verify the gate actually exercises the changed code.

© awebai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/code-review-discipline of awebai/aweb.

Open the folder on GitHubat commit 272d187

Compare with similar skills

Code Review Discipline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Review Discipline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Review Discipline this skillawebai/aweb115—~2.2kAutomated safety check: PassMIT
GitHub Review Iterationprisma/orm48k—~2.2kAutomated safety check: PassApache-2.0
Cherry Studio PR ReviewCherryHQ/cherry-studio52k—~3.9kAutomated safety check: PassAGPL-3.0
PR Reviewjaemk/self_update961—~1.5kAutomated safety check: NotesMIT
Cursor Composer Task DelegateChachamaru127/claude-code-harness3.2k—~4.4kAutomated safety check: NotesMIT
Local PR Reviewwindmill-labs/windmill18k—~995Automated safety check: PassCustom licence

Similar skills

  • Official

    Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.

    48k GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Cherry Studio PR Review

    CherryHQ/cherry-studio

    Reviews Cherry Studio branches, pull requests, commits, files and docs against the project's own architecture, naming, API-boundary and UI rules, report-only by default.

    52k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed
  • PR Review

    jaemk/self_update

    Targeted, read-only review of a PR or checked-out branch. An agent skill from jaemk/self_update.

    961 GitHub stars~1.5k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Cursor Composer Task Delegate

    Chachamaru127/claude-code-harness

    Hands one implementation task to Cursor Composer in an isolated git worktree, then reviews its diff and cherry-picks the result into the main branch.

    3.2k GitHub stars~4.4k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Local PR Review

    windmill-labs/windmill

    Runs the same code review locally that GitHub's auto-review actions run on a PR, delegating to a fresh-context subagent so the review isn't biased by the main session's own reasoning.

    18k GitHub stars~995 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Auto Devflow

    HuangPuStar/FenixAgent

    A skill your agent uses when starting an issue, bugfix, feature, or refactor that should be driven by multiple coordinated subagents: explore → plan → code → review, with the main agent acting as…

    552 GitHub stars~3k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from awebai/aweb

All 17 skills in this repo
  • Recognizes old aweb bootstrap-era `agents/` directories and migrates them to current team and identity primitives, since the old command family is retired.

    115 GitHub stars~701 tokensUpdated yesterday
    Auto-check passed
  • Guides decisions for agents working in an aweb team: when to check shared state, claim tasks, take locks, read team roles and instructions, and open separate worktrees.

    115 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • aweb Messaging

    awebai/aweb

    Guides how an agent reads and responds to aweb mail and chat events, choosing between asynchronous mail and synchronous chat and respecting sender verification and encryption boundaries.

    115 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • This skill should be used when joining or being added to an aweb team, picking the correct invite/add-member path for the team's authority model (hosted vs BYOT), accepting invites, fetching team…

    115 GitHub stars~5.4k tokensUpdated yesterday
    Auto-check passed
  • Creates or appends a folio document from the built-in pitch, memo or metrics templates by sending schema-checked slots that folio renders to Markdown.

    115 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Present To Human

    awebai/aweb

    A skill your agent uses when an agent needs to show a human an folio document: mint a document-bound capability link with POST /v1/present, open the returned URL for the human, print it as fallback…

    115 GitHub stars~590 tokensUpdated yesterday
    Auto-check passed

Questions about Code Review Discipline

What does Code Review Discipline do?

Practices to apply when taking a code-review assignment. An agent skill from awebai/aweb. Code Review Discipline is an agent skill from awebai/aweb. Practices to apply when taking a code-review assignment.

When should I use Code Review Discipline?

Code Review Discipline fits situations like: tasks that involve Code review; tasks that involve Subagents.

How do I install Code Review Discipline in Claude Code?

Run `npx skills add awebai/aweb --skill code-review-discipline -a claude-code`. Or copy the skill folder (.claude/skills/code-review-discipline in awebai/aweb) into .claude/skills/code-review-discipline in your project. Claude Code loads it when a task matches its description.

How do I install Code Review Discipline in Codex?

Run `npx skills add awebai/aweb --skill code-review-discipline -a codex`. Or copy the skill folder (.claude/skills/code-review-discipline in awebai/aweb) into .agents/skills/code-review-discipline in your project. Codex loads it when a task matches its description.

Can I use Code Review Discipline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awebai/aweb --skill code-review-discipline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review-discipline, .gemini/skills/code-review-discipline, .github/skills/code-review-discipline and .opencode/skills/code-review-discipline in your project.

What does Code Review Discipline need to run?

Going by SKILL.md and its folder, Code Review Discipline needs the command-line tools its instructions call (uv). Our summary lists: Docker.

Does Code Review Discipline access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Code Review Discipline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Review Discipline use?

Code Review Discipline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Review Discipline use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Review Discipline?

Skills that share tags, products or a category with Code Review Discipline: GitHub Review Iteration (prisma/orm, 48k stars), Cherry Studio PR Review (CherryHQ/cherry-studio, 52k stars), PR Review (jaemk/self_update, 961 stars) and Cursor Composer Task Delegate (Chachamaru127/claude-code-harness, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Review Discipline?

awebai (a GitHub organization) maintains it in awebai/aweb, which has 115 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.

Source: awebai/aweb on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.