Agent skill

Code Review Criteria

by penpot in penpot/penpot

Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes.

MPL-2.0Auto-check passedDevelopment

Install Code Review Criteria

skills CLI
$ npx skills add penpot/penpot --skill code-review-criteria -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install penpot/penpot code-review-criteria --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/penpot/penpot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/code-review-criteria .claude/skills/code-review-criteria && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review-criteria
GitHub stars
61k
Token cost
~3.6k tokens
SKILL.md length
1,959 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
MPL-2.0

At a glance

Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes.

  • Works in 5 steps: Correctness → Readability & Simplicity → Architecture → …
  • Tasks that involve Code review
  • SKILL.md covers Overview, When to Use, Core Principles and The Five-Axis Review, plus 9 more sections
  • Calls npm

What it does

Code Review Criteria is an agent skill from penpot/penpot. Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes. Loaded by the reviewer subagent of the review-code flow. Not a user-facing flow — to review code, use the review-code flow.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Code review and Accessibility. The repository describes itself as: Penpot: The open-source design platform for Product teams that need scalable collaboration. The licence is MPL-2.0.

When your agent uses it

  • Tasks that involve Code review
  • Tasks that involve Accessibility

Example prompts

  • “/code-review-criteria”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Correctness
  2. Readability & Simplicity
  3. Architecture
  4. Security
  5. Performance

What it can do on your machine

Read from SKILL.md and the folder at commit bcb7a83. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Review Criteria loads about 3.6k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 1,959 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from penpot/penpot at commit bcb7a83, republished under its MPL-2.0 licence (© penpot). 1,959 words, ~3,567 tokens.

Download SKILL.mdSave it as .claude/skills/code-review-criteria/SKILL.md (or your agent's skills folder).
name
code-review-criteria
description
Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes. Loaded by the reviewer subagent of the review-code flow. Not a user-facing flow — to review code, use the review-code flow.

Code Review Criteria and Quality

Overview

Multi-dimensional code review with quality gates. Every change gets reviewed before merge — no exceptions. Review covers five axes: correctness, readability, architecture, security, and performance.

The approval standard: Approve a change when it definitely improves overall code health, even if it isn't perfect. Perfect code doesn't exist — the goal is continuous improvement. Don't block a change because it isn't exactly how you would have written it. If it improves the codebase and follows the project's conventions, approve it.

When to Use

  • The reviewer subagent of the review-code flow loads this skill to perform the review of a code change.
  • To review code, always go through the review-code flow — never load this skill directly for that. This is the criteria reference, not the flow.

Core Principles

These principles underpin every axis. When in doubt, default to them.

  • DRY (Don't Repeat Yourself): Every piece of knowledge has one authoritative representation. If the same logic appears in two places, extract it into a shared helper, model, or type. Reviewers: flag duplicated logic as a required change — it's not "just similar," it's drift that will diverge.
  • KISS (Keep It Simple, Stupid): The simplest solution that works is the best solution. Complexity must earn its place. Reviewers: if you need more than one sentence to explain what a piece of code does, it's too complex — push for simplification before merge.
  • YAGNI (You Aren't Gonna Need It): Don't add abstractions, hooks, or generalizations for hypothetical future use cases. Generalize on the third occurrence, not the first. Reviewers: delete speculative generality.
  • Don't invent problems: Do not manufacture issues to produce more feedback. Every finding must be a real risk, a real readability barrier, or a real architectural concern — not a hypothetical or a stylistic preference disguised as a problem.

The Five-Axis Review

Every review evaluates code across these dimensions.

1. Correctness

Does the code do what it claims to do?

  • Does it match the spec or task requirements?
  • Are edge cases handled (null, empty, boundary values)?
  • Are error paths handled (not just the happy path)?
  • Does it pass all tests? Are the tests actually testing the right things?
  • Are there off-by-one errors, race conditions, or state inconsistencies?
2. Readability & Simplicity

Can another engineer (or agent) understand this code without the author explaining it?

  • Are names descriptive and consistent with project conventions? (No temp, data, result without context)
  • Is the control flow straightforward (avoid nested ternaries, deep callbacks)?
  • Are there any "clever" tricks that should be simplified?
  • KISS check: Is this the simplest approach that solves the problem? A 20-line straightforward function beats a 5-line clever one that requires a comment to explain.
  • Could this be done in fewer lines? (1000 lines where 100 suffice is a failure)
  • Are abstractions earning their complexity? (Don't generalize until the third use case)
  • Is a new conditional bolted onto an unrelated flow? Push the logic into its own helper, state, or policy.
  • Do repeated conditionals on the same shape appear? They signal a missing model or dispatcher.
  • Are there dead code artifacts: no-op variables, backwards-compat shims, or // removed comments?
3. Architecture

Does the change fit the system's design?

  • Does it follow existing patterns or introduce a new one? If new, is it justified?
  • Does it maintain clean module boundaries?
  • DRY check: Is there existing code that does the same thing? Reuse the canonical helper instead of writing a near-duplicate. If two branches do nearly the same thing, collapse them.
  • Are dependencies flowing in the right direction (no circular dependencies)?
  • Is the abstraction level appropriate (not over-engineered, not too coupled)?
  • Does this refactor reduce complexity or just relocate it? Count the concepts a reader must hold. Prefer the restructuring that makes whole branches disappear over one that re-centralizes the same logic. Prefer deleting an abstraction to polishing it.
  • Is feature-specific logic leaking into a shared or general-purpose module?
  • Are type boundaries explicit? Question gratuitous any/unknown/optional/casts and silent fallbacks.
  • Structural remedies: When you flag a problem, propose the move — not just the problem. Replace conditionals with dispatchers, collapse duplicate branches, separate orchestration from business logic, extract helpers, split large files. Prefer the remedy that removes moving pieces over one that spreads the same complexity around.
4. Security

For detailed security guidance, see security-and-hardening.

  • Is user input validated and sanitized?
  • Are secrets kept out of code, logs, and version control?
  • Is authentication/authorization checked where needed?
  • Are SQL queries parameterized (no string concatenation)?
  • Are outputs encoded to prevent XSS?
  • Are dependencies from trusted sources with no known vulnerabilities?
  • Is data from external sources (APIs, logs, user content, config files) treated as untrusted?
5. Performance
  • Any N+1 query patterns?
  • Any unbounded loops or unconstrained data fetching?
  • Any synchronous operations that should be async?
  • Any unnecessary re-renders in UI components?
  • Any missing pagination on list endpoints?
  • Any large objects created in hot paths?

Review Process

  1. Understand the intent — What is this change trying to accomplish? What spec or task does it implement?
  2. Review tests first — Tests reveal intent and coverage. Do they test behavior, not implementation details? Are edge cases covered?
  3. Review the implementation — Walk through each file with the five axes in mind.
  4. Categorize findings — Label every comment with its severity:
PrefixMeaningAuthor Action
Critical:Blocks mergeSecurity vulnerability, data loss, broken functionality
High:Required changeMust address before merge
Medium:Should fixStrongly recommended, not a blocker
Low:Minor, optionalAuthor may ignore — formatting, style preferences
Suggestion:Worth consideringNot required, but improves the code

Unique finding IDs. Assign every finding a stable identifier: F1, F2, F3, … numbered in order of severity (Critical first, then High, Medium, Low, Suggestion). Use the ID everywhere the finding is mentioned — in section headers, in the verdict, in follow-up discussion. Never renumber within a review. Example: **F3 (High)** — app/validate.cljs:42 — duplicate branch logic….

For each finding, describe the circumstances under which it could fail: specific inputs, load conditions, timing, or user actions that trigger the problem. "This crashes when input is null" is actionable; "this might crash" is not.

Lead with what matters: correctness and security first, then structural issues, then everything else. A few high-conviction comments beat a long list.

  1. Verify the verification — What tests were run? Did the build pass? Was the change tested manually? Screenshots for UI changes?

Review Output

Structure every review using this format:

Summary

Briefly explain what the code does and give an overall assessment.

Critical and High-Priority Issues

List problems that could cause security incidents, data loss, crashes, incorrect behavior, or major performance degradation. Each finding gets its unique ID (F1, F2, …). For each: state the severity, identify the file/function/code section, explain why it's a problem, describe failure circumstances, and provide a concrete improvement with corrected code when useful.

Other Findings

List medium- and low-priority issues, including maintainability and design concerns. Continue the ID sequence started above (F3, F4, …).

Suggested Refactoring

Provide focused code changes or revised snippets. Preserve existing behavior unless a behavior change is explicitly justified.

Testing Recommendations

Identify missing tests and describe specific test cases, including edge cases and failure scenarios.

Show full SKILL.md (787 more words)Show less
Positive Observations

Mention implementation choices that are clear, safe, efficient, or well designed. This is not fluff — it reinforces good patterns and tells the author what to keep doing.

Final Verdict

Choose one:

  • Approve — Ready to merge
  • Approve with minor changes — Good to merge after addressing low/medium issues
  • Request changes — Critical or high issues must be resolved before merge

List the finding IDs the verdict depends on (e.g. "Request changes: F1, F4").

Change Sizing

Small, focused changes are easier to review, faster to merge, and safer to deploy.

~100 lines changed   → Good. Reviewable in one sitting.
~300 lines changed   → Acceptable if it's a single logical change.
~1000 lines changed  → Too large. Split it.

Watch file size, not just diff size. Around 1000 total lines in a single file is a common inspection signal. When a change materially grows an already-large file, decompose first.

Splitting strategies:

StrategyHowWhen
StackSubmit a small change, start the next one based on itSequential dependencies
By file groupSeparate changes for groups needing different reviewersCross-cutting concerns
HorizontalCreate shared code/stubs first, then consumersLayered architecture
VerticalBreak into smaller full-stack slices of the featureFeature work

Separate refactoring from feature work. A change that refactors and adds new behavior is two changes — submit them separately.

Change Descriptions

  • First line: Short, imperative, standalone. "Delete the FizzBuzz RPC" not "Deleting the FizzBuzz RPC."
  • Body: What is changing and why. Include context and reasoning not visible in the code itself.
  • Anti-patterns: "Fix bug," "Fix build," "Add patch," "Phase 1."

Dependencies

Before adding any dependency:

  1. Does the existing stack solve this? (Often it does.)
  2. How large is the dependency? (Check bundle impact.)
  3. Is it actively maintained? (Check last commit, open issues.)
  4. Does it have known vulnerabilities? (npm audit)
  5. What's the license? (Must be compatible with the project.)

Rule: Prefer standard library and existing utilities over new dependencies. Every dependency is a liability.

Upgrading dependencies:

  • Read the changelog, not just the version number. Semver is a promise the maintainer may not have kept.
  • One dependency per change. When a bulk bump breaks the build, you've lost which package did it.
  • Let the tests decide — a green suite before and after, not just "it installed."
  • Review the lockfile diff, not just package.json. Commit it and never hand-edit it.

For supply-chain risk triage, follow the security-and-hardening skill.

Common Rationalizations

RationalizationReality
"It works, that's good enough"Working code that's unreadable, insecure, or architecturally wrong creates debt that compounds.
"I wrote it, so I know it's correct"Authors are blind to their own assumptions. Every change benefits from another set of eyes.
"We'll clean it up later"Later never comes. The review is the quality gate — use it.
"AI-generated code is probably fine"AI code needs more scrutiny, not less. It's confident and plausible, even when wrong.
"The tests pass, so it's good"Tests are necessary but not sufficient. They don't catch architecture, security, or readability problems.
"The refactor makes it cleaner"Relocating complexity isn't reducing it. If the reader still holds the same number of concepts, the structure didn't improve.
"It's only a small addition to this file"Small diffs still push files past healthy size and bolt branches onto unrelated flows.
"It's just a version bump"A bump is a behavior change you didn't write. Read the changelog.
"I'll upgrade everything in one PR"A bulk bump hides which package broke the build. One per change.
"It's duplicated but it's only two places"Two becomes three becomes five. Extract now, before the copies diverge.
"The abstraction is future-proof"YAGNI. Delete speculative generality — generalize on the third occurrence, not the first.
"It's clever but efficient"Cleverness is a readability tax. If it needs a comment to understand, simplify it.

Red Flags

  • PRs merged without any review
  • Review that only checks if tests pass (ignoring other axes)
  • "LGTM" without evidence of actual review
  • Security-sensitive changes without security-focused review
  • Large PRs that are "too big to review properly" (split them)
  • No regression tests with bug fix PRs
  • Accepting "I'll fix it later" — it never happens
  • A refactor that moves code around without reducing the number of concepts a reader must hold
  • New conditionals scattered into unrelated code paths (a missing abstraction)
  • A bespoke helper that duplicates an existing canonical one
  • A bulk "bump dependencies" PR with no changelog review

Verification

Before emitting the verdict, verify the change as it stands. This is the reviewer's own due diligence — it covers the state of the code at review time, not the later resolution of findings (fixing findings is the author's job; confirming them is a new review):

  • Tests pass — run them yourself, don't trust the claim
  • Build succeeds
  • The verification story is documented (what changed, how it was verified)
  • Dependency upgrades reviewed against changelog, isolated per package, verified by green suite

See Also

  • For detailed security review guidance, see security-and-hardening

© penpot, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/code-review-criteria of penpot/penpot.

Open the folder on GitHubat commit bcb7a83

Compare with similar skills

Code Review Criteria next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Review Criteria compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Review Criteria this skillpenpot/penpot61k—~3.6kAutomated safety check: PassMPL-2.0
Frontend Code Reviewlanggenius/dify158k—~938Automated safety check: PassCustom licence
Core Components Code Reviewcore-ds/core-components137—~5.4kAutomated safety check: PassMIT
Codebase Review SwarmZaxbyHub/opencode-swarm493—~2.8kAutomated safety check: PassMIT
Code Review And QualityHangYu8123/mini-harness199—~1.8kAutomated safety check: PassNone
Code ReviewOrange-OpenSource/Orange-Boosted-Bootstrap222—~846Automated safety check: PassMIT

Similar skills

  • Frontend Code Review

    langgenius/dify

    Reviews frontend changes under `web/` or `packages/dify-ui/` for concrete defects and broken project contracts, using routed rule packs and a severity scale for findings.

    158k GitHub stars~938 tokensUpdated today
    DevelopmentAuto-check passed
  • Core Components Code Review

    core-ds/core-components

    Review a Pull Request or diff in the @alfalab/core-components UI library — correctness bugs, public API/breaking changes, accessibility, keyboard/focus/pointer interaction, component states…

    137 GitHub stars~5.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Codebase Review Swarm

    ZaxbyHub/opencode-swarm

    Runs an evidence-gated, quote-grounded audit of a codebase for security, QA, accessibility, performance and more, and writes a verified report without changing source files.

    493 GitHub stars~2.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Code Review And Quality

    HangYu8123/mini-harness

    Multi-axis, review-only code review of a diff across six axes — request achievement, correctness, readability/simplicity, architecture, security, and performance — with severity-labelled findings.

    199 GitHub stars~1.8k tokensUpdated 16 days ago
    DevelopmentAuto-check passed
  • Code Review

    Orange-OpenSource/Orange-Boosted-Bootstrap

    OUDS Web compliance code review. An agent skill from Orange-OpenSource/Orange-Boosted-Bootstrap.

    222 GitHub stars~846 tokensUpdated today
    DevelopmentAuto-check passed
  • Handsontable Code Review

    handsontable/handsontable

    Reviews Handsontable monorepo changes across architecture, code quality, performance and accessibility, and tests, with a confidence rating on every finding.

    22k GitHub stars~999 tokensUpdated today
    DevelopmentAuto-check passed

More from penpot/penpot

All 24 skills in this repo
  • Hardens code against vulnerabilities. An agent skill from penpot/penpot.

    61k GitHub starsUsed in 6 repos~4.7k tokens
    Auto-check: notes
  • Create PR

    penpot/penpot

    PR flow — open a new PR for the current task branch (validates base branch, commits, issue and push state) or update an existing PR's title or description to match Penpot conventions.

    61k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Bat Cat

    penpot/penpot

    A cat clone with syntax highlighting, line numbers, and Git integration - a modern replacement for cat.

    61k GitHub starsUsed in 2 repos~1.1k tokens
    Auto-check passed
  • Local CI

    penpot/penpot

    Run local CI-style checks with ./scripts/ci (lint, tests, format) per monorepo module.

    61k GitHub stars~822 tokensUpdated today
    Auto-check passed
  • Ste

    penpot/penpot

    Write or rewrite text in ASD-STE100 Simplified Technical English.

    61k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Create Commit

    penpot/penpot

    Stage, review, and commit files following Penpot commit conventions.

    61k GitHub stars~690 tokensUpdated today
    Auto-check: notes

Categories

Questions about Code Review Criteria

What does Code Review Criteria do?

Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes. Code Review Criteria is an agent skill from penpot/penpot. Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes.

When should I use Code Review Criteria?

Code Review Criteria fits situations like: tasks that involve Code review; tasks that involve Accessibility.

How do I install Code Review Criteria in Claude Code?

Run `npx skills add penpot/penpot --skill code-review-criteria -a claude-code`. Or copy the skill folder (.agents/skills/code-review-criteria in penpot/penpot) into .claude/skills/code-review-criteria in your project. Claude Code loads it when a task matches its description.

How do I install Code Review Criteria in Codex?

Run `npx skills add penpot/penpot --skill code-review-criteria -a codex`. Or copy the skill folder (.agents/skills/code-review-criteria in penpot/penpot) into .agents/skills/code-review-criteria in your project. Codex loads it when a task matches its description.

Can I use Code Review Criteria in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add penpot/penpot --skill code-review-criteria -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review-criteria, .gemini/skills/code-review-criteria, .github/skills/code-review-criteria and .opencode/skills/code-review-criteria in your project.

What does Code Review Criteria need to run?

Going by SKILL.md and its folder, Code Review Criteria needs the command-line tools its instructions call (npm).

Does Code Review Criteria access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Code Review Criteria safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Review Criteria use?

Code Review Criteria is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Review Criteria use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Review Criteria?

Skills that share tags, products or a category with Code Review Criteria: Frontend Code Review (langgenius/dify, 158k stars), Core Components Code Review (core-ds/core-components, 137 stars), Codebase Review Swarm (ZaxbyHub/opencode-swarm, 493 stars) and Code Review And Quality (HangYu8123/mini-harness, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Review Criteria?

penpot (a GitHub organization) maintains it in penpot/penpot, which has 60,833 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: penpot/penpot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.