Agent skill

Review Quality

by jpicklyk in jpicklyk/task-orchestrator

Review quality framework for the work-to-review transition gate.

MITAuto-check passedDevelopment

Install Review Quality

skills CLI
$ npx skills add jpicklyk/task-orchestrator --skill review-quality -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jpicklyk/task-orchestrator review-quality --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jpicklyk/task-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/review-quality .claude/skills/review-quality && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-quality
GitHub stars
207
Token cost
~3.7k tokens
SKILL.md length
2,092 words
Files
1
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Review quality framework for the work-to-review transition gate.

  • Works in 4 steps: Test Suite Verification → Plan Alignment → Test Quality → …
  • Filling review-checklist notes
  • SKILL.md covers Getting Started, Review Areas and Review Output
  • Calls git

What it does

Review Quality is an agent skill from jpicklyk/task-orchestrator. Review quality framework for the work-to-review transition gate. Guides verification of plan alignment, test quality, and code simplification before marking implementation complete. Referenced by schema guidance fields during review-phase note filling. Use when filling review-checklist notes or when asked to review completed implementation work.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Code simplification. The repository describes itself as: Server-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what… The licence is MIT.

When your agent uses it

  • Filling review-checklist notes
  • Asked to review completed implementation work

Example prompts

  • “/review-quality”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Test Suite Verification
  2. Plan Alignment
  3. Test Quality
  4. Simplification

What it can do on your machine

Read from SKILL.md and the folder at commit d2d362a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review Quality loads about 3.7k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 2,092 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jpicklyk/task-orchestrator at commit d2d362a, republished under its MIT licence (© jpicklyk). 2,092 words, ~3,742 tokens.

Download SKILL.mdSave it as .claude/skills/review-quality/SKILL.md (or your agent's skills folder).
name
review-quality
description
Review quality framework for the work-to-review transition gate. Guides verification of plan alignment, test quality, and code simplification before marking implementation complete. Referenced by schema guidance fields during review-phase note filling. Use when filling review-checklist notes or when asked to review completed implementation work.
user-invocable
false

Review Quality Framework

This skill defines what a reviewer must verify before implementation work advances to completion. It applies whether the reviewer is the orchestrator directly or a delegated subagent.

The review gate exists because implementation agents optimize for getting things working, not for verifying they built the right thing. Without a structured review checkpoint, planned work gets silently dropped, tests get written to pass rather than to verify, and unnecessary complexity accumulates. The review is where these failure modes get caught.

Critical separation of concerns: The reviewer must not be the same agent that wrote the code or the tests. An agent reviewing its own work will rationalize rather than evaluate. The reviewer reads, runs, and reports — it never fixes. If issues are found, they go back to the implementation agent for resolution.

This is the same principle the needs-test-author trait applies one step earlier, to test authorship itself — the test author must not be the implementer, for the same rationalize-not-evaluate reason. Verifying that separation actually held (see "Independence verification" in Area 3) is this rule applied to test authorship, not an optional extra.


Getting Started

The reviewer is given an MCP item ID. Use MCP tools and codebase access to gather what you need — do not expect context to be pre-loaded for you.

  1. Load the item's notes — query_notes(itemId=..., includeBody=true) to retrieve the planning note and implementation-notes. The planning note's key depends on the item's schema: feature-summary (feature-implementation), task-scope (feature-task), or diagnosis (bug-fix).
  2. Read the changed files — use the implementation notes to identify which files were modified, then read them directly. Review the actual code, not just summaries.
  3. Run the test suite — execute the project's test command and capture the results. Do not assume tests pass because the implementation agent said they did.
  4. For items carrying the needs-test-author trait — load the trait's test-plan and test-manifest notes via query_notes(operation="list", itemId=..., includeBody=true), which carries the test author's own commit SHA range field, and obtain the orchestrator-provided per-child SHA table (Pre-SHA/Post-SHA/Test-Pre-SHA/ Test-Post-SHA) plus any declared orchestrator fixture-repair commit SHAs from the review handoff, for cross-checking. A trait-bearing item with no test-plan note is a blocking issue on its own — do not proceed to the Area 3 independence verification until the note exists.

If the planning note (feature-summary / task-scope / diagnosis) or implementation notes are missing, the review cannot proceed. Report this as a blocking issue.

Before reporting any file or artifact as missing, or any behavior as broken, verify with a direct check — Read the exact expected path, an exact-path Glob, or a reproduction — rather than inferring absence from one plausible directory or from a prior bug's pattern. Two review false-positives reached verdicts this way before being caught downstream.


Review Areas

These four areas form the minimum review. Each one catches a different class of failure. If the review surfaces additional concerns, include them — this is a floor, not a ceiling.

1. Test Suite Verification

Run the test suite before anything else. Everything downstream depends on knowing the actual state of the tests.

Run ./gradlew :current:test and capture the output. Record the total test count and the pass/fail breakdown.

If tests fail: Document every failure — test name, assertion message, and the file where the test lives. Do not attempt to fix failures. Do not speculate about whether failures are pre-existing or new. Report what you observe. Test failures are a blocking issue — the item cannot advance with a failing test suite.

If tests pass: Record the count and move on.

2. Plan Alignment

Compare what was built against the planning note (feature-summary / task-scope / diagnosis). The goal is to catch drift in both directions — work that was planned but not done, and work that was done but not planned.

Check each acceptance criterion. Walk through the acceptance criteria from the planning note one by one. For each criterion, identify the specific code change that satisfies it. If a criterion has no corresponding implementation, flag it — either the work is incomplete or the criterion was intentionally descoped (which should appear in the implementation notes).

Check for unplanned changes. Review the changed files for modifications that don't trace back to any acceptance criterion. Unplanned changes aren't automatically wrong — sometimes implementation reveals necessary adjacent work. But they should be acknowledged and justified in the implementation notes, not silent.

Check non-goals weren't violated. Review the planning note's non-goals list. If the implementation touched areas that were explicitly scoped out, flag it.

Check semantic-claim citations. Every sentence in the planning note, the changed docs, the CHANGELOG bullet and any new header or KDoc comment that states nullability, redaction, ownership or an honest limit should cite a file:line. Open each cited line and confirm it establishes the claim as worded (watch for inverted conditions such as ADMIN-AND vs ADMIN-OR). An uncited claim of this kind is an observation. A claim the cited code contradicts is blocking. For a NEW-SURFACE claim that cites a spec decision, confirm it against the diff. Prose that is not a semantic claim needs no citation.

3. Test Quality

The planning note's test strategy defined what should be tested — happy paths, failure paths, and edge cases. The reviewer verifies that the tests actually deliver on that strategy, not just that they exist and pass.

This is where the separation of concerns matters most. The agent that wrote the tests has an inherent bias toward believing they're correct. An independent reviewer can evaluate whether the tests verify real behavior or just confirm that code runs.

Map tests to the test strategy. For each scenario in the planning note's test strategy, identify the corresponding test. Missing coverage is a gap to report.

Evaluate test substance. Watch for these patterns that produce green results without catching real bugs:

  • Tautological assertions — asserting something equals itself, or that a non-null value is not null, without verifying the actual value is correct.
  • Mock-heavy tests that verify nothing real — every dependency mocked, test only confirms mocks were called in order. Mocks are fine for isolation, but the test must still assert something meaningful about the unit's output or state change.
  • Happy-path-only coverage — if the test strategy called for failure paths, those tests need to exist and need to verify the failure behavior is correct (right exception type, right error message, right fallback behavior).
  • Overly broad assertions — result != null or list.isNotEmpty() when specific values, sizes, or contents should be checked. These pass even when the implementation is wrong.
  • Assumption escapes — assumeTrue (or an equivalent guard) gating out a real failure instead of asserting against it, so the test silently skips rather than reporting the bug it was written to catch.
  • Implementation-derived oracles — an expected value computed from the implementation's own formula or output rather than from an independent oracle source. These pass by construction and verify nothing.

Check edge cases. Verify each boundary condition from the test strategy has a corresponding test. If implementation notes documented new edge cases discovered during development, check whether tests were added for those too.

Independence verification (items with the needs-test-author trait). When the item carries this trait, verify the test-author/implementer separation actually held before trusting anything else found in this area — a compromised separation undermines every other finding above it:

  • Separation held. Confirm the test files were introduced within the test author's commit range, not the implementer's — check git log over both ranges using the test-manifest's commit SHA range field together with the orchestrator's per-child SHA table. Record the result as independent when actor and commit-range separation both hold, independent-degraded (temporal-only) when only ordering separates them (for example a Direct-tier bug-fix run in single-actor mode), or not-independent when separation did not hold. Any silent implementer edit to a test file after the author's range is a blocking issue, regardless of whether the edit looks benign — except orchestrator fixture-repair commits that were declared in the review handoff (listed with their SHAs). Those are not silent edits; verify they touched only construction/ setup code, never assertions, before excluding them from the blocking rule.
  • Manifest maps to spec scenarios. Cross-check the test-manifest's S-id-to-test mapping against the test-plan's numbered scenarios (S1…): every scenario must be marked covered or explicitly not-covered with a reason. Do not take "covered" on faith — open at least two of the claimed tests and read their bodies to verify the claimed coverage is real.
  • Oracle spot-check. Pick one non-trivial scenario and trace its expected value back to the oracle source recorded in the test-plan (spec clause, stated algorithm, or external reference). If the expected value matches what the implementation produces but not what the named oracle source specifies, that is a blocking issue — it means the oracle was derived from the implementation rather than the spec.
  • Arbitration record. Every ambiguity the test author flagged must have a named resolver in the manifest's arbitration record. Give oracle-degraded scenarios — ones where the oracle source itself was uncertain — a second, closer look.
  • Forbidden-pattern declaration is true. Fetch the served rules forbidden-test-patterns and test-assertion-vacuity (query_rules(operation="get", rootId, key); fall back to the test-author skill's §7 if either is unserved) and read the test bodies behind the test-manifest's forbidden-pattern declaration — do not take "none used" on faith. Every instance of a listed pattern (skip-guards on behavioral conditions, disjunctive escapes, not-null-only or assert.ok-only assertions, assertions that cannot fail given their fixture or harness) must appear in the declaration with a justification. An undeclared instance, or a declaration claiming none when instances exist, is a blocking issue regardless of whether the suite is green — a false declaration carries the same weight as not-independent, because the manifest is the audit trail everything above relies on.
Show full SKILL.md (500 more words)Show less
4. Simplification

The reviewer does not run /simplify — that pass belongs at the feature level, not per-task review. Check the implementation notes for whether /simplify was run during implementation.

If /simplify made changes, verify those changes have test coverage. This is the one thing the reviewer checks here — not the simplification itself, just whether the resulting code is tested. Report coverage gaps; do not re-run simplify or evaluate its judgment calls.

If /simplify was not run or made no changes, there is nothing to check in this area — move on.

Coverage gaps found here are not blocking unless they leave a structural change entirely unverified.


Review Output

The review produces a review-checklist note on the MCP item. Structure the note around findings, not process.

Verdict

Every review must end with a clear verdict:

  • Pass — all acceptance criteria met, tests pass and have substance, no blocking issues. The item can advance. An independence-verification result of independent-degraded (temporal-only) maps to Pass — record it in the audit note as the observed mode, not as a finding; it is the sanctioned outcome for Direct-tier single-actor runs, not a degradation to flag.
  • Fail — blocking issues — test failures, missing acceptance criteria, critical gaps in test coverage, a nullability/redaction/ownership/honest-limit claim that its cited line contradicts, or (for items with the needs-test-author trait) a not-independent independence-verification result, an unexplained implementer edit to test files, or an undeclared or falsely declared forbidden test pattern. These fail the item even when the test suite is green — a compromised separation or a false manifest makes a passing suite untrustworthy. The item must go back for fixes before it can advance. List every blocking issue.
  • Pass with observations — no blocking issues, but simplification findings or minor test quality concerns worth addressing. The item can advance, but the observations should be tracked for follow-up.
Findings Format

Report every finding you observe, at every severity. Do not withhold minor findings and do not apply a high-severity-only bar — current models follow severity filters literally, which suppresses real findings. Mark each finding blocking or observation and state your confidence; the orchestrator's verdict handling (see Verdict, above) is the downstream filter, not your own judgment about what is worth mentioning.

For each finding, state:

  • What was expected (from the planning note or test strategy)
  • What was found (in the code or test output)
  • Severity (blocking or observation)
  • Confidence (how sure you are this is a real issue, not a false positive)

Be specific. "Tests could be better" is not actionable. "Test testCreateItem asserts only that the result is not null — it should verify the item's title and status match the input parameters" is actionable.

Gate Enforcement

The reviewer does not advance the item. It fills the review-checklist note and reports the verdict. The orchestrator reads the verdict and decides whether to:

  • Advance the item (pass or pass-with-observations)
  • Send the item back to the implementation agent with the blocking issues list (fail)

A failing verdict with clear findings gives the implementation agent exactly what to fix without ambiguity.

© jpicklyk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/review-quality of jpicklyk/task-orchestrator.

Open the folder on GitHubat commit d2d362a

Compare with similar skills

Review Quality next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Quality compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Quality this skilljpicklyk/task-orchestrator207—~3.7kAutomated safety check: PassMIT
Find Slopletehaha/moneymatter162—~3kAutomated safety check: PassAGPL-3.0
Subagent ReviewNikiforovAll/claude-code-rules141—~927Automated safety check: PassApache-2.0
Ponytail Lazy Developer ModeDietrichGebert/ponytail160k1 repos~873Automated safety check: PassMIT
Ponytail Reviewkortix-ai/suna20k4 repos~593Automated safety check: PassCustom licence
PonytailDavidObando/gsharp5657 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Find Slop

    letehaha/moneymatter

    Hunt "AI slop" in this codebase — duplication, reinvented wheels, over-engineering, defensive cruft, dead code, comment slop, performance antipatterns.

    162 GitHub stars~3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Subagent Review

    NikiforovAll/claude-code-rules

    Review changed code for reuse, quality, and efficiency using three parallel disposable subagents.

    141 GitHub stars~927 tokensUpdated 8 days ago
    DevelopmentAuto-check passed
  • Ponytail Lazy Developer Mode

    DietrichGebert/ponytail

    Makes the agent pick the laziest solution that works: skip unneeded work, reuse what exists, prefer the standard library and platform features, and keep diffs small.

    160k GitHub starsUsed in 1 repo~873 tokens
    DevelopmentAuto-check passed
  • Ponytail Review

    kortix-ai/suna

    Code review focused exclusively on over-engineering. An agent skill from kortix-ai/suna.

    20k GitHub starsUsed in 4 repos~593 tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    565 GitHub starsUsed in 7 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Finds and implements evidence-backed simplifications in the ego-lite repository, such as dead code, duplicated state and speculative abstractions, without hiding behavior changes.

    17k GitHub stars~1.2k tokensUpdated today
    DevelopmentAuto-check passed

More from jpicklyk/task-orchestrator

All 28 skills in this repo
  • Task Orchestrator Server Setup

    jpicklyk/task-orchestrator

    Walks through how to launch and reach the MCP Task Orchestrator server container: transport, REST API, port publishing, config mounts and config-sync.

    207 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Run Wave

    jpicklyk/task-orchestrator

    Resolves ready MCP work items into a run plan, shows it to you, then executes it through the Workflow tool or direct subagent dispatch, with post-run verification.

    207 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Adopt Project Scope Migration

    jpicklyk/task-orchestrator

    Migrates an existing unscoped Task Orchestrator database to the project-scoping convention in place, creating one project anchor root and re-parenting work trees under it after a mandatory dry run.

    207 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Bulk Task Completion

    jpicklyk/task-orchestrator

    Completes or cancels a whole feature subtree, a named list of items, or a batch of stale work items at once, previewing the impact and warning before force-completing anything active.

    207 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Task Orchestrator Item Creator

    jpicklyk/task-orchestrator

    Creates an MCP work item from conversation context, anchoring it under the right container, inferring type and priority and pre-filling the required notes.

    207 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Work Item Dependency Manager

    jpicklyk/task-orchestrator

    Views, creates, deletes and diagnoses BLOCKS, IS_BLOCKED_BY and RELATES_TO links between MCP work items, including why an item cannot start.

    207 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed

Questions about Review Quality

What does Review Quality do?

Review quality framework for the work-to-review transition gate. Review Quality is an agent skill from jpicklyk/task-orchestrator. Review quality framework for the work-to-review transition gate.

When should I use Review Quality?

Review Quality fits situations like: filling review-checklist notes; asked to review completed implementation work.

How do I install Review Quality in Claude Code?

Run `npx skills add jpicklyk/task-orchestrator --skill review-quality -a claude-code`. Or copy the skill folder (.claude/skills/review-quality in jpicklyk/task-orchestrator) into .claude/skills/review-quality in your project. Claude Code loads it when a task matches its description.

How do I install Review Quality in Codex?

Run `npx skills add jpicklyk/task-orchestrator --skill review-quality -a codex`. Or copy the skill folder (.claude/skills/review-quality in jpicklyk/task-orchestrator) into .agents/skills/review-quality in your project. Codex loads it when a task matches its description.

Can I use Review Quality in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jpicklyk/task-orchestrator --skill review-quality -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-quality, .gemini/skills/review-quality, .github/skills/review-quality and .opencode/skills/review-quality in your project.

What does Review Quality need to run?

Going by SKILL.md and its folder, Review Quality needs the command-line tools its instructions call (git).

Does Review Quality access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Review Quality safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Review Quality use?

Review Quality is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Quality use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Review Quality?

Skills that share tags, products or a category with Review Quality: Find Slop (letehaha/moneymatter, 162 stars), Subagent Review (NikiforovAll/claude-code-rules, 141 stars), Ponytail Lazy Developer Mode (DietrichGebert/ponytail, 160k stars) and Ponytail Review (kortix-ai/suna, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Quality?

jpicklyk (a GitHub user) maintains it in jpicklyk/task-orchestrator, which has 207 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 9, 2026.

Source: jpicklyk/task-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.