Agent skill

Code Review

by ericrisco in ericrisco/rsc-harness

A skill your agent uses to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you…

MITAuto-check passedDevelopment

Install Code Review

skills CLI
$ npx skills add ericrisco/rsc-harness --skill code-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness code-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/code-review .claude/skills/code-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review
GitHub stars
156
Token cost
~2.3k tokens
SKILL.md length
1,123 words
Files
4 (incl. references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you…

  • Judge a concrete diff
  • SKILL.md covers Get the change and its intent…, The pass order, Confidence floor and the… and Severity and finding format, plus 5 more sections
  • Calls gh and git
  • Read-only unless --comment

What it does

Code Review is an agent skill from ericrisco/rsc-harness. Use to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you can defend, one verdict, read-only unless --comment or --fix. NOT the SDD gate keyed to 02-DOCS/wiki/sdd/ that also processes incoming review comments (that is review).

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/pr-workflow.md`).

It sits in Development, covering Code review. It works with GitHub. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Judge a concrete diff
  • Read-only unless --comment

Example prompts

  • “/code-review”

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Review loads about 2.3k tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,123 words, ~2,300 tokens.

Download SKILL.mdSave it as .claude/skills/code-review/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
code-review
description
Use to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you can defend, one verdict, read-only unless --comment or --fix. NOT the SDD gate keyed to 02-DOCS/wiki/sdd/ that also processes incoming review comments (that is `review`).
tags
code-review, pr-review, quality, correctness
recommends
review, secure-coding, verify
origin
risco

Code review — standalone, spec-less diff judgment

You are reviewing a concrete change — a git diff, a branch, a GitHub PR, a pasted patch — on its own merits. No rsc-SDD spec/plan/constitution chain is required and you should not pretend one exists. This is the doctrine behind the executable /code-review slash command: same evidence bar, written as a discipline you run by hand. If the user is mid-SDD and wants to process incoming comments against 02-DOCS/wiki/sdd/, that is ../review/SKILL.md; a naked diff or an inbound third-party PR is this skill.

The north star is signal-to-noise. Report only findings you would stake your name on. A clean diff is APPROVE, not a manufactured nit. High-false-positive review gets tuned out by humans in about two weeks; the bar to aim for is the logic-error review where under 1% of findings come back marked wrong. Padding does not make you look thorough — it trains the reader to ignore you.

Get the change and its intent first

Three inputs, in this order: the diff, its stated purpose, and the touched surface (the files around the hunks, not just the hunks).

bash
# A GitHub PR
gh pr diff 1432
gh pr view 1432 --json title,body,files,additions,deletions

# A local branch against the trunk
git diff main...HEAD
git diff --stat main...HEAD   # see blast radius before reading

# A pasted patch — read it as given

A review with no notion of intent is a review of vibes. If no purpose is stated, infer it from the diff and say what you assumed ("Assuming this is meant to add idempotency to the webhook handler…") so the reader can correct a wrong premise — a silent wrong premise produces a confidently wrong review. Then read the whole changed file, not just the green/red lines: the structural failure of standalone review is judging a hunk without its context and shipping generic pattern-matched suggestions.

The pass order

Run these in order. Passes 1–5 are correctness/safety and are blocking-eligible; pass 6 is cleanup and is usually [should-fix] or [nit]. A clean pass is a reportable result ("contracts: nothing changed shape, no finding"), not a pass you silently skip.

#PassThe questionTypical defects
1Intent fidelityDoes it do what it claims?Wrong behaviour, missing case from the stated goal, scope creep
2Correctness & boundariesRight on the edges?Off-by-one, null/empty/unicode, overflow, timezone, concurrency, swallowed errors
3Contracts & dataDo callers/data still hold?Broken API shape, migration without backfill, nullable made non-null, enum drift
4Security boundaryUntrusted input → dangerous sink?Unsanitized input to query/shell/template, authz gap, secret in code/log
5Tests as evidenceDo the tests prove the change?Tests assert nothing, test the mock, miss the new branch, were deleted to go green
6Reuse / simplification / efficiencyCould existing code do this?Reimplemented helper, copy-paste divergence, N+1, needless allocation in a hot loop

Pass 4 is a boundary pass — trace untrusted input to its sink and flag the reachable ones. For a real STRIDE/OWASP threat model with exploitability ranking and vulnerable→fixed diffs, hand off to ../secure-coding/SKILL.md.

Adjacent jobs, delegated by name: running the lint/type/test gates until they are green is ../verify/SKILL.md (this skill judges whether green is correct); root-causing one confirmed failure is ../debug/SKILL.md; cross-checking spec/plan/tasks before code exists is ../analyze/SKILL.md.

Confidence floor and the false-positive skip-list

The 80% rule: if you are not at least ~80% sure a finding is real, you have two moves — trace the code until you are sure, or downgrade it to [question]. Never ship a guess dressed as a defect.

Skip these common false positives outright (or demote to [question]/[nit]):

  • Guarded upstream — the "missing" check happens in the caller you can see; trace before you flag.
  • Framework-enforced — the framework already does it (e.g. an ORM that parameterizes, a router that validates).
  • Behind an off-everywhere flag — real but unreachable in any deployed config → [nit], not blocking.
  • Test-only / generated code held to the prod bar — don't demand prod-grade error handling in a fixture or a generated client.
  • Style the linter owns — quotes, import order, line length. If a tool enforces it, don't spend a finding on it.

No severity inflation. Rank by blast radius × reachability, not by how clever the catch was. A typo in a log string is a nit even if it took effort to spot.

Show full SKILL.md (456 more words)Show less

Severity and finding format

  • [blocking] — wrong/unsafe; merging causes a real defect. Must be fixed.
  • [should-fix] — a real problem with bounded blast radius; fix it or consciously accept it.
  • [nit] — minor; the reader may ignore it without consequence.
  • [question] — you suspect an issue but cannot prove reachability; asking, not asserting.

Every finding carries where / why / repro / fix:

text
[should-fix] api/orders.py:88 — duplicated total logic
  where:  `subtotal = sum(i.price * i.qty for i in items)` re-implements
          `cart.compute_subtotal()` (cart/totals.py:14), which also applies
          per-item discounts this copy silently drops.
  why:    discounted items now bill at full price on this path only;
          the two implementations will drift on the next discount change.
  repro:  order containing any item with `discount_pct > 0` → charged the
          undiscounted amount; covered by no test.
  fix:    call `cart.compute_subtotal(items)` instead of inlining the sum.

Rule: no repro or stated mechanism → it is a [question], not a blocker. "This could overflow" with no path is a question; "n*1000 with n up to 3M exceeds int32 at orders.py:51" is a finding.

Verify before you flag

Read the surrounding code, trace the value, confirm the path is reachable.

  • Bad: "Looks like SQL injection." (pattern-match)
  • Good: "search() interpolates req.query.q straight into db.execute(\… WHERE name='${q}'`)at search.ts:22;q` is unvalidated user input → injection." (traced)

If you cannot trace it to a concrete value and a reachable sink, you do not yet have a finding.

The verdict

End every review with exactly one, plainly — no mushy middle:

  • APPROVE — no blockers, no should-fix. Point the user to ../ship/SKILL.md to merge.
  • APPROVE WITH NITS — mergeable; nits listed but none gate the merge.
  • CHANGES REQUESTED — at least one [blocking]. List precisely what unblocks it, so the author knows when they are done.

Effort dial

Mirror the slash command's effort level: low/medium → fewer, high-confidence findings (raise the confidence floor, focus on passes 1–4). high/max → broader coverage; uncertain findings are allowed but must be labelled [question], never inflated into blockers. This is coverage vs precision, not the user's register — it changes what you look at, not how you word it.

Emitting comments and applying fixes

Read-only by default. You produce findings + a verdict and stop there. Two opt-in modes:

  • --comment → post the findings as an inline-anchored review on the PR.
  • --fix → apply the agreed findings to the working tree.
bash
# Summary review (the verdict)
gh pr review 1432 --request-changes -b "CHANGES REQUESTED — see inline. Blocker: orders.py:88 …"
gh pr review 1432 --approve -b "APPROVE — correctness and contracts clean."

Inline line-anchored comments go through the GitHub REST API — see references/pr-workflow.md for the JSON shape, fork-PR handling, and large-diff strategy. If --fix puts you on the default branch, branch first; commit or push only when the user asks; git authorship is Eric (no Claude co-author or generated footer).

Anti-patterns

Failure modeReality
"It compiles and the tests pass, so it's correct."Tests prove green, not correct. Pass 5 asks whether the tests actually exercise the new branch — green for the wrong reason is a finding.
Listing everything you would have done differently.That is noise. Report defects and reuse wins you can defend; preference is not a finding.
"It's just a dependency bump, skim it."Bumps carry supply-chain and transitive risk and behaviour changes. Check the changelog/lockfile diff, not just the version string.
Applying every nit "to be safe" under --fix.Each unrequested edit is scope creep and a regression surface. Apply the agreed findings only.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/code-review of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/pr-workflow.md

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Review this skillericrisco/rsc-harness156—~2.3kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
GitHub Review Iterationprisma/orm48k—~2.2kAutomated safety check: PassApache-2.0
PR Finalize Reviewmicrosoft/garnet12k—~3.1kAutomated safety check: PassMIT
PR Review State Fetchprisma/orm48k—~767Automated safety check: PassApache-2.0
Fastlane Pull Request Reviewfastlane/fastlane42k—~550Automated safety check: PassMIT

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Official

    Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.

    48k GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • PR Finalize Review

    microsoft/garnet

    Official

    Checks that a pull request's title and description match its implementation and reviews the code for Garnet best practices, reporting findings without posting them.

    12k GitHub stars~3.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Fetches a pull request's canonical review state as JSON, validates it, and renders markdown, a text summary and triage target files from it using bundled scripts.

    48k GitHub stars~767 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Reviews a fastlane pull request against its linked issue and the project guides, separating blocking from non-blocking findings and handling vulnerabilities privately.

    42k GitHub stars~550 tokensUpdated today
    DevelopmentAuto-check passed
  • Reviews open pull requests in the daisyUI repository using read-only GitHub data and isolated base-versus-PR checks, then writes a merge verdict report.

    43k GitHub stars~766 tokensUpdated 7 days ago
    DevelopmentAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Code Review

What does Code Review do?

A skill your agent uses to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you…. Code Review is an agent skill from ericrisco/rsc-harness. Use to judge a concrete diff, branch, or GitHub PR on its own merits with no rsc-SDD spec/plan chain to key off — the spec-less giving pass behind /code-review: only findings you can defend, one verdict, read-only unless --comment or --fix.

When should I use Code Review?

Code Review fits situations like: judge a concrete diff; read-only unless --comment.

How do I install Code Review in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill code-review -a claude-code`. Or copy the skill folder (skills/code-review in ericrisco/rsc-harness) into .claude/skills/code-review in your project. Claude Code loads it when a task matches its description.

How do I install Code Review in Codex?

Run `npx skills add ericrisco/rsc-harness --skill code-review -a codex`. Or copy the skill folder (skills/code-review in ericrisco/rsc-harness) into .agents/skills/code-review in your project. Codex loads it when a task matches its description.

Can I use Code Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill code-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review, .gemini/skills/code-review, .github/skills/code-review and .opencode/skills/code-review in your project.

What does Code Review need to run?

Going by SKILL.md and its folder, Code Review needs the command-line tools its instructions call (gh and git).

Does Code Review access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Code Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Review use?

Code Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Review use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 897 tokens, read only when the agent opens those files.

What are the alternatives to Code Review?

Skills that share tags, products or a category with Code Review: PR Babysitter (openinterpreter/openinterpreter, 69k stars), GitHub Review Iteration (prisma/orm, 48k stars), PR Finalize Review (microsoft/garnet, 12k stars) and PR Review State Fetch (prisma/orm, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Review?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.