Agent skill

Plan Review Criteria

by penpot in penpot/penpot

Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans.

MPL-2.0Auto-check passedAgent Workflows

Install Plan Review Criteria

skills CLI
$ npx skills add penpot/penpot --skill plan-review-criteria -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install penpot/penpot plan-review-criteria --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/penpot/penpot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/plan-review-criteria .claude/skills/plan-review-criteria && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plan-review-criteria
GitHub stars
61k
Token cost
~3.3k tokens
SKILL.md length
1,296 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
MPL-2.0

At a glance

Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans.

  • Works in 12 steps: Completeness → Task Quality → Architecture & Sequencing → …
  • Tasks that involve Accessibility
  • SKILL.md covers Overview, When to Use, The Six-Axis Review and Structural Remedies, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Plan Review Criteria is an agent skill from penpot/penpot. Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans. Loaded by the reviewer subagent of the review-plan flow. Not a user-facing flow — to review a plan, use the review-plan flow.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Accessibility, Planning and Quizzes and assessments. The repository describes itself as: Penpot: The open-source design platform for Product teams that need scalable collaboration. The licence is MPL-2.0.

When your agent uses it

  • Tasks that involve Accessibility
  • Tasks that involve Planning
  • Tasks that involve Quizzes and assessments

Example prompts

  • “/plan-review-criteria”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Completeness
  2. Task Quality
  3. Architecture & Sequencing
  4. Risk Coverage
  5. Actionability
  6. Proposed Code Quality (when the plan includes implementation details)
  7. Understand the Goal
  8. Check Completeness First
  9. Review Task Quality
  10. Validate Sequencing
  11. Assess Actionability
  12. Verify the Verification Story

What it can do on your machine

Read from SKILL.md and the folder at commit 10955f1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Plan Review Criteria loads about 3.3k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 1,296 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from penpot/penpot at commit 10955f1, republished under its MPL-2.0 licence (© penpot). 1,296 words, ~3,278 tokens.

Download SKILL.mdSave it as .claude/skills/plan-review-criteria/SKILL.md (or your agent's skills folder).
name
plan-review-criteria
description
Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans. Loaded by the reviewer subagent of the review-plan flow. Not a user-facing flow — to review a plan, use the review-plan flow.

Plan Review Criteria

Overview

Multi-dimensional plan review with quality gates. Every plan gets reviewed before implementation starts — no exceptions. Review covers six axes: completeness, task quality, architecture & sequencing, risk coverage, actionability, and proposed code quality.

The approval standard: Approve a plan when it is specific enough that a skilled implementer could execute it without guessing, the task ordering is sound, and risks are acknowledged. Perfect plans don't exist — the goal is confidence that implementation won't derail. Don't block a plan because it isn't exactly how you would have structured it. If it's executable and well-organized, approve it.

When to Use

  • The reviewer subagent of the review-plan flow loads this skill to perform the review of a plan.
  • To review a plan, always go through the review-plan flow — never load this skill directly for that. This is the criteria reference, not the flow.

Do NOT use for: Single-file changes with obvious scope, or when the task is trivial enough to just do.

The Six-Axis Review

Every plan gets evaluated across these dimensions:

1. Completeness

Does the plan cover everything needed to implement successfully?

  • Is the context clear? (What problem, why now, what's the goal?)
  • Are affected modules identified with paths?
  • Are architecture decisions documented with rationale?
  • Is there a testing strategy?
  • Are verification commands explicit (not "run the tests")?
  • Are open questions listed (not buried in someone's head)?
  • Is there a parallelization assessment for multi-task plans?

Missing any of these is a gap, not a nit.

2. Task Quality

Are the tasks well-defined and independently executable?

  • Does every task have acceptance criteria? (Testable, not vague)
  • Does every task have verification steps?
  • Are tasks sized appropriately? (XS–M is ideal, L is acceptable, XL must be split)
  • Are dependencies between tasks explicitly stated?
  • Are files likely touched listed?
  • Is each task a single, self-contained change? (Not "implement the whole feature")
  • Could a skilled implementer pick up any task and execute it without asking clarifying questions?
3. Architecture & Sequencing

Is the plan structured so implementation flows correctly?

  • Does implementation order follow the dependency graph (foundations first)?
  • Are tasks vertically sliced (feature paths) rather than horizontally layered?
  • Does each task leave the system in a working state?
  • Are there checkpoints between major phases?
  • Are high-risk tasks early (fail fast)?
  • Is the total plan a reasonable number of tasks? (More than ~15 tasks suggests the scope should be split into multiple plans)
4. Risk Coverage

Are the hard parts acknowledged and mitigated?

  • Are edge cases identified?
  • Are breaking changes or migration concerns noted?
  • Are security implications considered?
  • Are performance implications considered?
  • Are external dependencies or integration risks flagged?
  • Is there a plan for rollback if something goes wrong?
  • Are data integrity risks addressed (what happens if a migration fails mid-way)?
5. Actionability

Can an implementer actually execute this?

  • Are file paths specific (not "update the relevant files")?
  • Are function/method names mentioned where applicable?
  • Are verification commands copy-pasteable (not "run the linter")?
  • Are test commands project-specific (not generic)?
  • Is the code shape described where the implementation isn't obvious?
  • Are conventions referenced (naming, patterns, existing utilities to reuse)?
  • Does the plan reference existing code the implementer should read first?
6. Proposed Code Quality (when the plan includes implementation details)

If the plan proposes code shapes, function signatures, data structures, or API designs, evaluate those proposals against code-review-criteria:

  • Correctness: Do the proposed types/signatures handle edge cases (null, empty, boundaries)?
  • Readability: Are proposed names descriptive and consistent with project conventions?
  • Architecture: Do proposed abstractions follow existing patterns? Are they justified (not over-engineered)?
  • Security: Do proposed APIs validate input at boundaries? Any injection/XSS vectors in the design?
  • Performance: Do proposed data structures avoid N+1 patterns? Any unbounded operations in the design?

When to apply: Only when the plan includes specific code snippets, type definitions, API contracts, or function signatures. Plans that only describe "what" without showing "how" skip this axis.

Structural Remedies

When you flag a structural problem in a plan, propose the fix — not just the problem:

  • A task is too large (XL): Split it into vertical slices. Each slice should be independently testable.
  • Missing acceptance criteria: Draft 2–3 specific, testable conditions for the task.
  • Wrong sequencing: Identify the dependency and propose the correct order.
  • No checkpoints: Suggest where checkpoints should go (typically after every 2–3 tasks).
  • Vague verification: Replace "run tests" with the actual project command.
  • Horizontal slicing: Restructure into vertical feature paths.
  • Missing risk section: Draft the risks you can identify from the plan content.

Prefer the remedy that makes the plan immediately actionable over one that just flags the gap.

Plan Sizing

Plans should be scoped to a single deliverable:

1–5 tasks    → Good. A focused feature or bug fix.
6–10 tasks   → Acceptable for a moderate feature.
11–15 tasks  → Large. Consider splitting into phases.
15+ tasks    → Too large. Split into multiple plans.

What counts as "one plan": A self-contained set of changes that delivers a single coherent capability. If you can describe the goal in one sentence, it's one plan.

Show full SKILL.md (505 more words)Show less

Categorize Findings

Label every comment with its severity so the author knows what's required vs optional:

PrefixMeaningAuthor Action
(no prefix)Required changeMust address before implementation starts
Critical:Blocks implementationMissing security consideration, data integrity risk, fundamentally wrong approach
Nit:Minor, optionalAuthor may ignore — wording, formatting
Optional: / Consider:SuggestionWorth considering but not required
FYIInformational onlyNo action needed — context for future reference

Lead with what matters. Order findings by leverage: missing risks and wrong sequencing first, then task quality gaps, then completeness, then nits. If you have one critical sequencing problem and ten nits, the sequencing problem is the review.

Review Process

Step 1: Understand the Goal

Before evaluating structure, understand intent:

- What is this plan trying to accomplish?
- What problem does it solve?
- What does "done" look like?
Step 2: Check Completeness First

Scan for missing sections before diving into content:

- Context present?
- Affected modules listed?
- Architecture decisions documented?
- Risks acknowledged?
- Testing strategy defined?
- Verification commands explicit?
Step 3: Review Task Quality

Walk through each task:

For each task:
1. Can I tell exactly what to build?
2. Are acceptance criteria specific and testable?
3. Is the size reasonable (not XL)?
4. Are dependencies clear?
5. Would I know which files to touch?
Step 4: Validate Sequencing

Check the dependency graph:

- Are foundations built first?
- Does each task leave the system working?
- Are checkpoints placed correctly?
- Are high-risk items early?
- Is it vertically sliced?
Step 5: Assess Actionability

Put yourself in the implementer's shoes:

- Could I pick up task 1 and start coding without asking any questions?
- Are the verification commands copy-pasteable?
- Are file paths and function names specific?
- Is existing code referenced where I'd need to read it?
Step 6: Verify the Verification Story

Check that the plan can actually confirm it worked:

- What tests should pass after implementation?
- What build/compile commands are relevant?
- What manual checks are needed?
- How do we know the feature works end-to-end?
Step 7: Evaluate Proposed Code Quality (if applicable)

If the plan includes code snippets, types, or API designs:

- Load code-review-criteria skill for criteria
- Check proposed signatures for edge cases
- Verify naming follows project conventions
- Confirm abstractions follow existing patterns
- Scan for security vectors in proposed APIs
- Check for performance issues in proposed data structures

Review Checklist

markdown
## Review: [Plan title]

### Completeness
- [ ] Context explains the problem and goal
- [ ] Affected modules are listed with paths
- [ ] Architecture decisions have rationale
- [ ] Testing strategy is defined
- [ ] Verification commands are explicit and project-specific
- [ ] Open questions are listed

### Task Quality
- [ ] Every task has acceptance criteria
- [ ] Every task has verification steps
- [ ] Tasks are sized XS–M (L acceptable, XL must be split)
- [ ] Task dependencies are stated
- [ ] Files likely touched are listed

### Architecture & Sequencing
- [ ] Order follows dependency graph (foundations first)
- [ ] Vertically sliced (not horizontal layers)
- [ ] Each task leaves system working
- [ ] Checkpoints exist between phases
- [ ] High-risk tasks are early

### Risk Coverage
- [ ] Edge cases identified
- [ ] Breaking changes / migrations noted
- [ ] Security implications considered
- [ ] Performance implications considered
- [ ] Rollback strategy exists (if applicable)

### Actionability
- [ ] File paths are specific
- [ ] Verification commands are copy-pasteable
- [ ] Existing code to read is referenced
- [ ] Conventions and patterns are noted

### Proposed Code Quality *(if plan includes implementation details)*
- [ ] Proposed types/signatures handle edge cases
- [ ] Proposed names follow project conventions
- [ ] Proposed abstractions follow existing patterns
- [ ] No security vectors in proposed APIs
- [ ] No performance issues in proposed structures

### Verdict
- [ ] **Approve** — Ready to implement
- [ ] **Request changes** — Gaps must be addressed

Common Rationalizations

RationalizationReality
"I'll figure out the details during implementation"That's how you discover blocking dependencies mid-task. Surface them now.
"The tasks are obvious, no need for criteria"Write them anyway. Explicit criteria surface hidden assumptions.
"It's just a small feature, it doesn't need a plan"Small features have edge cases too. 3 tasks with criteria takes 5 minutes.
"The plan is good enough""Good enough" without acceptance criteria means the implementer defines "done" — and they might define it differently.
"I'll add verification steps later"Later never comes. The plan is the contract — define verification now.
"Risks are minimal"Every change has risks. If you can't name them, you haven't thought about them.
"The file paths are obvious"They're obvious to the author. The implementer might not know the codebase.
"The code in the plan is fine, it'll get reviewed later"Plan-level code review catches design problems before implementation — fixing them after coding is more expensive.

Red Flags

  • No acceptance criteria on any task
  • Tasks that say "implement the feature" without specifics
  • No verification steps anywhere in the plan
  • All tasks are XL-sized
  • No checkpoints between phases
  • Dependency order isn't considered (e.g., API handler before domain model)
  • No testing strategy
  • Verification commands are generic ("run tests") instead of project-specific
  • Plan has 20+ tasks (scope too large for one plan)
  • No risk section on a plan with migrations, breaking changes, or security implications
  • Horizontal slicing (all domain, then all services, then all API)
  • File paths are vague ("update the relevant files")
  • Missing open questions section despite stated unknowns
  • Proposed code ignores project conventions or existing patterns
  • Proposed types use gratuitous any/unknown/optional without justification
  • Proposed APIs don't validate input at boundaries

See Also

  • For producing plans, use the planner skill
  • For reviewing implemented code, use code-review-criteria — also the criteria source for axis 6
  • For security-specific concerns, see security-and-hardening
  • For testing strategy guidance, see testing

© penpot, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/plan-review-criteria of penpot/penpot.

Open the folder on GitHubat commit 10955f1

Compare with similar skills

Plan Review Criteria next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plan Review Criteria compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plan Review Criteria this skillpenpot/penpot61k—~3.3kAutomated safety check: PassMPL-2.0
ArenaJakeschincariol/arena-skill402—~4.8kAutomated safety check: PassMIT
Dynamic WorkflowsHabitat-Thinking/ai-literacy-superpowers114—~1.9kAutomated safety check: PassCustom licence
Woo AI Smokewoocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0
Subagent Driven DevelopmentAsvarox/allkaraoke26138 repos~1.2kAutomated safety check: PassNone
Executing PlansGanyuanRan/Aegis1.3k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Arena

    Jakeschincariol/arena-skill

    Make 100 versions of Claude fight to the death over one task.

    402 GitHub stars~4.8k tokensUpdated 14 days ago
    Agent WorkflowsAuto-check passed
  • Dynamic Workflows

    Habitat-Thinking/ai-literacy-superpowers

    This skill should be used when an agent is deciding whether to author a dynamic workflow — a self-authored, ephemeral multi-agent harness — for a task.

    114 GitHub stars~1.9k tokensUpdated 20 days ago
    Agent WorkflowsAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub stars~7.4k tokensUpdated yesterday
    EducationAuto-check: notes
  • Subagent Driven Development

    Asvarox/allkaraoke

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    261 GitHub starsUsed in 38 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans

    GanyuanRan/Aegis

    A skill your agent uses when executing a written implementation plan across sessions or with review checkpoints.

    1.3k GitHub starsUsed in 1 repo~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Exec

    umputun/cc-thingz

    Execute plan tasks sequentially using subagents. An agent skill from umputun/cc-thingz.

    485 GitHub stars~8k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed

More from penpot/penpot

All 24 skills in this repo
  • Hardens code against vulnerabilities. An agent skill from penpot/penpot.

    61k GitHub starsUsed in 6 repos~4.7k tokens
    Auto-check: notes
  • Create PR

    penpot/penpot

    PR flow — open a new PR for the current task branch (validates base branch, commits, issue and push state) or update an existing PR's title or description to match Penpot conventions.

    61k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Bat Cat

    penpot/penpot

    A cat clone with syntax highlighting, line numbers, and Git integration - a modern replacement for cat.

    61k GitHub starsUsed in 2 repos~1.1k tokens
    Auto-check passed
  • Local CI

    penpot/penpot

    Run local CI-style checks with ./scripts/ci (lint, tests, format) per monorepo module.

    61k GitHub stars~822 tokensUpdated yesterday
    Auto-check passed
  • Ste

    penpot/penpot

    Write or rewrite text in ASD-STE100 Simplified Technical English.

    61k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Code review criteria — the five review axes, core principles, severity format, and verdict for reviewing code changes.

    61k GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed

Questions about Plan Review Criteria

What does Plan Review Criteria do?

Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans. Plan Review Criteria is an agent skill from penpot/penpot. Plan review criteria — the six review axes, severity rubric, approval standard, and output format for reviewing implementation plans.

When should I use Plan Review Criteria?

Plan Review Criteria fits situations like: tasks that involve Accessibility; tasks that involve Planning; tasks that involve Quizzes and assessments.

How do I install Plan Review Criteria in Claude Code?

Run `npx skills add penpot/penpot --skill plan-review-criteria -a claude-code`. Or copy the skill folder (.agents/skills/plan-review-criteria in penpot/penpot) into .claude/skills/plan-review-criteria in your project. Claude Code loads it when a task matches its description.

How do I install Plan Review Criteria in Codex?

Run `npx skills add penpot/penpot --skill plan-review-criteria -a codex`. Or copy the skill folder (.agents/skills/plan-review-criteria in penpot/penpot) into .agents/skills/plan-review-criteria in your project. Codex loads it when a task matches its description.

Can I use Plan Review Criteria in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add penpot/penpot --skill plan-review-criteria -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plan-review-criteria, .gemini/skills/plan-review-criteria, .github/skills/plan-review-criteria and .opencode/skills/plan-review-criteria in your project.

What does Plan Review Criteria need to run?

SKILL.md names no scripts, command-line tools or credentials: Plan Review Criteria is instructions for the agent only.

Does Plan Review Criteria access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plan Review Criteria safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Plan Review Criteria use?

Plan Review Criteria is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plan Review Criteria use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Plan Review Criteria?

Skills that share tags, products or a category with Plan Review Criteria: Arena (Jakeschincariol/arena-skill, 402 stars), Dynamic Workflows (Habitat-Thinking/ai-literacy-superpowers, 114 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Subagent Driven Development (Asvarox/allkaraoke, 261 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plan Review Criteria?

penpot (a GitHub organization) maintains it in penpot/penpot, which has 60,869 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: penpot/penpot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.