Agent skill

Review Testing

by posit-dev in posit-dev/skills

Review test code for quality, design, and completeness after implementing a feature or fixing a bug.

MITAuto-check passedTesting & QA

Install Review Testing

skills CLI
$ npx skills add posit-dev/skills --skill review-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install posit-dev/skills review-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/posit-dev/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/posit-dev/review-testing .claude/skills/review-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-testing
GitHub stars
529
Token cost
~2.6k tokens
SKILL.md length
1,301 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Review test code for quality, design, and completeness after implementing a feature or fixing a bug.

  • Works in 7 steps: Assertion Completeness → Test Structure → Fixture and State Management → …
  • The user asks to review my tests
  • SKILL.md covers Review Scope, What Makes a Test Valuable, Review Areas and Response Format, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Review Testing is an agent skill from posit-dev/skills. Review test code for quality, design, and completeness after implementing a feature or fixing a bug. Use when the user asks to "review my tests", "check my test quality", "are these tests good enough", "review testing", or after completing a feature implementation that includes tests. Also use when tests feel brittle, flaky, or superficial. Cross-references production code to find coverage gaps.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test coverage. The repository describes itself as: A collection of Claude Skills from Posit. The licence is MIT.

When your agent uses it

  • The user asks to review my tests
  • Check my test quality
  • Are these tests good enough
  • After completing a feature implementation that includes tests

Example prompts

  • “review my tests”
  • “check my test quality”
  • “are these tests good enough”
  • “/review-testing”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Assertion Completeness
  2. Test Structure
  3. Fixture and State Management
  4. Mocking Boundaries
  5. Test Smells
  6. Naming and Readability
  7. Coverage Gaps

What it can do on your machine

Read from SKILL.md and the folder at commit 0300946. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review Testing loads about 2.6k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,301 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from posit-dev/skills at commit 0300946, republished under its MIT licence (© posit-dev). 1,301 words, ~2,577 tokens.

Download SKILL.mdSave it as .claude/skills/review-testing/SKILL.md (or your agent's skills folder).
name
review-testing
description
Review test code for quality, design, and completeness after implementing a feature or fixing a bug. Use when the user asks to "review my tests", "check my test quality", "are these tests good enough", "review testing", or after completing a feature implementation that includes tests. Also use when tests feel brittle, flaky, or superficial. Cross-references production code to find coverage gaps.
metadata.author
Garrick Aden-Buie (@gadenbuie)
metadata.version
0.1
license
MIT

You are reviewing test code written alongside a feature implementation or bug fix. Ensure the tests are well-designed, thorough, and maintainable — not just that they pass. Tests that merely mirror implementation details create false confidence and become a maintenance burden during refactoring.

Review Scope

Identify what to review:

  1. Find changed test files on the current branch (relative to the base branch).
  2. Find the production code those tests cover — trace imports, function calls, and file naming conventions to map tests to their targets.
  3. Find related existing tests for the same modules or functions that weren't changed — these may need updates or reveal gaps.

Read test files first, before production code. If you can infer the feature's requirements and edge cases from the tests alone, that's a sign the tests are well-written. If you need to read the implementation to understand what the tests are doing, that's a finding worth reporting.

What Makes a Test Valuable

Weigh each test against four qualities:

Regression protection — Does this test actually catch bugs? A test that exercises trivial code or skips complex branches protects against nothing. Check: does the test touch business-critical logic, or only verify the happy path of a simple getter?

Refactoring resilience — Will this test break when someone restructures code without changing behavior? Tests coupled to internal method names, call sequences, or private state punish every cleanup with false failures, eroding trust in the suite.

Fast feedback — Unit tests should run in milliseconds. If a test hits the filesystem, network, or database unnecessarily, that's a design issue. But don't confuse speed with value — an integration test verifying a real database query is better than a fast unit test that mocks everything and verifies nothing.

Maintainability — Can someone unfamiliar with this code read the test and understand what it verifies and why? Tests with sprawling setup, cryptic names, or deeply nested mocking fail this check.

Review Areas

1. Assertion Completeness

The most common weakness in generated tests: asserting only the obvious output and missing the full "blast radius" of a state change.

When a test triggers an action, ask what else changed. If a test adds an item to a cart, does it only check the item count? Or does it also verify the price calculation, subtotal update, and that other items are unaffected?

Flag when:

  • A test asserts a return value but ignores side effects
  • A test checks that an operation succeeded but not that it produced the right result
  • A test verifies the happy path but skips boundary conditions (empty inputs, nulls, maximum values, off-by-one)
  • Error-path tests only check that an error was thrown, not the error message, type, or cleanup behavior
2. Test Structure

Each test should have exactly one Arrange-Act-Assert cycle. If a test acts and asserts multiple times in sequence, it's testing multiple behaviors and should be split.

Flag when:

  • A test has multiple act phases (testing a workflow, not a behavior)
  • The arrange phase is so large the actual behavior being tested is buried
  • Assertions appear inside setup helpers or utility functions (hides what's being verified)
  • Test logic contains conditionals or loops — tests should be straight-line code
3. Fixture and State Management

How test data is created determines whether the suite is maintainable at scale. Inline setup (data created in the test body) is fine for simple tests but painful when constructor signatures change across many tests. Implicit setup (shared beforeEach/setUp blocks) eliminates duplication but obscures what each test actually depends on. Delegated setup (factory functions or builders called explicitly) keeps tests readable while centralizing construction logic.

Flag when:

  • A shared setup block creates objects most tests don't use
  • Tests depend on external files, database state, or globals defined elsewhere ("Mystery Guest" — everything a test needs should be visible in its body or one function call away)
  • Tests assume resources exist without creating them (files, database records, environment variables)
  • Tests mutate shared state without cleanup, causing order-dependent failures
4. Mocking Boundaries

Mocks are essential for isolation, but overuse turns tests into mirrors of the implementation.

The key principle: mock at architectural boundaries, not at every function call. Use stubs (canned data) for query-type dependencies and mocks (interaction verification) only for commands with side effects like sending emails or writing to external systems.

Respect the codebase's existing mocking convention. Flag inconsistency within the project, not deviation from a universal rule.

Flag when:

  • A test mocks types it doesn't own (third-party libraries, framework internals) — when the library updates, the mock passes against outdated assumptions. Prefer: thin adapters, real implementations in integration tests, or official testing utilities (MemoryRouter, test databases, in-memory caches).
  • Mock density is high — 4+ mock configurations suggests too many responsibilities in the production code
  • A test verifies call counts or argument sequences on internal methods (couples the test to how the code works, not what it does)
  • Stubs are being verified (checking a stub was called tests implementation, not behavior)
Show full SKILL.md (485 more words)Show less
5. Test Smells
SmellWhat It Looks LikeWhy It Matters
Assertion RouletteMultiple assertions with no failure messagesCan't tell which assertion broke without debugging
Eager TestOne test exercises several unrelated methodsFailures are ambiguous — which behavior broke?
Lazy TestMultiple tests call the same method with identical inputsRedundant maintenance cost, no coverage gain
Sleepy TestHard-coded sleep()/Sys.sleep()/setTimeout()Flaky in CI, slow everywhere. Use polling or explicit waits.
Rotten GreenAssertions inside try/tryCatch or conditional branchesTest always passes because the assertion is never reached
Sensitive EqualityAsserting against toString()/print() outputBreaks on formatting changes; assert structural properties instead
Print Statementprint()/console.log() instead of assertionsDebugging leftovers that verify nothing
Snapshot AbuseSnapshots as a substitute for behavioral assertionsAny change triggers failure, developers blindly update. Good uses: one snapshot of an HTML component's structure (not every prop combination), error message text, CLI output. Bad: snapshotting entire objects or rendering every variant.
Implementation MirrorExpected values computed using the same logic as production codeTest and production code will always agree — even when both are wrong. Hardcode expected values from a known-good source.
6. Naming and Readability

Test names should describe behavior, not implementation. A well-named test suite reads like a feature specification.

Flag when:

  • Test names reference internal method names (e.g., test_processData_returns_true) — break on rename
  • Test names are generic (test1, test_it_works, test_basic)
  • The name doesn't communicate the scenario or expected outcome

Prefer behavioral names: test_expired_subscription_blocks_access, delivery_with_past_date_is_invalid, empty_cart_shows_zero_total.

7. Coverage Gaps

Cross-reference production code changes against the test suite:

  • Are there branches or conditions no test exercises?
  • Are error paths tested? (Not just "does it throw" but "does it throw the right thing and clean up properly")
  • Are edge cases covered? (Empty collections, null/NA/None/undefined, boundary values, concurrent access if applicable)
  • If production code changed existing behavior, were existing tests updated?

Walk through the implementation and note every decision point — each if, match, switch, error handler, or early return. Check whether the test suite exercises both sides of that decision.

When reviewing R tests using testthat, check if the r-testthat skill is available and invoke it for R-specific conventions and patterns.

Response Format

## Summary
[Overall assessment: How well do these tests protect the codebase?]

## Critical Issues (Blocking)
[Tests that provide false confidence or will cause real problems.]

## Required Changes
[Design problems that weaken the test suite.]

## Strong Suggestions
[Improvements to test quality and maintainability.]

## Noted
[Minor style or convention issues. Mention once, then move on.]

## Verdict
Request Changes | Needs Discussion | Approve

## Next Steps
[Options for proceeding]

Use file:line references for every finding. Quote the specific test code that demonstrates the issue and show what better code looks like.

Next Steps

At the end of the review, offer the user these options:

Discuss and address findings: Use the AskUserQuestion tool to walk through the issues. Group by severity or topic, offer resolution options, and mark the recommended choice.

Fix the issues: Offer to apply fixes directly in priority order — blocking issues first, then required changes, then suggestions. Confirm before continuing after each group.

Add to a pull request: When reviewing in context of a PR, offer to post the review as a PR comment. Include attribution: "Review assisted by the review-testing skill."

If operating as a subagent, skip the next steps and output only the review findings.

© posit-dev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in posit-dev/review-testing of posit-dev/skills.

Open the folder on GitHubat commit 0300946

Compare with similar skills

Review Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Testing this skillposit-dev/skills529—~2.6kAutomated safety check: PassMIT
Requirementsrizsotto/Bear6.5k—~2kAutomated safety check: PassGPL-3.0
Crap Analysisardalis/RiverBooks1342 repos~3.4kAutomated safety check: PassNone
Code Coverages3s-project/s3s311—~789Automated safety check: PassApache-2.0
Project Statusbactopia/bactopia522—~787Automated safety check: PassMIT
Check Coverageldayton/Dippy243—~403Automated safety check: PassMIT

Similar skills

  • Requirements

    rizsotto/Bear

    Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…

    6.5k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed
  • Crap Analysis

    ardalis/RiverBooks

    Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.

    134 GitHub starsUsed in 2 repos~3.4k tokens
    Testing & QAAuto-check passed
  • Code Coverage

    s3s-project/s3s

    Measure and grow the line coverage of the s3s crate. An agent skill from s3s-project/s3s.

    311 GitHub stars~789 tokensUpdated today
    Testing & QAAuto-check passed
  • Project Status

    bactopia/bactopia

    Show a live snapshot of the Bactopia project state — component counts, GroovyDoc coverage, nf-test coverage, and structural issues.

    522 GitHub stars~787 tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Check Coverage

    ldayton/Dippy

    Ensure comprehensive test coverage for a CLI handler. An agent skill from ldayton/Dippy.

    243 GitHub stars~403 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Guarding Destructive Operations

    kajisho5/ffmpeg-skill

    Add and review preconditions on operations that delete, overwrite, rewrite history, or resolve a caller-supplied name to a filesystem path — refusing instead of warning, placing the guard ahead of…

    1.9k GitHub stars~2.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from posit-dev/skills

All 17 skills in this repo
  • R CLI App

    posit-dev/skills

    Build command-line apps in R using the Rapp package. An agent skill from posit-dev/skills.

    529 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • R Lifecycle

    posit-dev/skills

    Guidance for managing R package lifecycle according to tidyverse principles using the lifecycle package.

    529 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • PR Create

    posit-dev/skills

    Creates a pull request from current changes, monitors GitHub CI, and debugs any failures until CI passes.

    529 GitHub stars~4.3k tokensUpdated today
    Auto-check: warnings
  • Shiny Bslib

    posit-dev/skills

    Build modern Shiny dashboards and applications using bslib (Bootstrap 5).

    529 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Shiny Bslib Theming

    posit-dev/skills

    Advanced theming for Shiny apps using bslib and Bootstrap 5.

    529 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Critical Code Reviewer

    posit-dev/skills

    Rigorously review code or pull requests for correctness, security, accessibility, maintainability, tests, and edge cases.

    529 GitHub stars~3.9k tokensUpdated today
    Auto-check passed

Categories

Questions about Review Testing

What does Review Testing do?

Review test code for quality, design, and completeness after implementing a feature or fixing a bug. Review Testing is an agent skill from posit-dev/skills. Review test code for quality, design, and completeness after implementing a feature or fixing a bug.

When should I use Review Testing?

Review Testing fits situations like: the user asks to review my tests; check my test quality; are these tests good enough; after completing a feature implementation that includes tests.

How do I install Review Testing in Claude Code?

Run `npx skills add posit-dev/skills --skill review-testing -a claude-code`. Or copy the skill folder (posit-dev/review-testing in posit-dev/skills) into .claude/skills/review-testing in your project. Claude Code loads it when a task matches its description.

How do I install Review Testing in Codex?

Run `npx skills add posit-dev/skills --skill review-testing -a codex`. Or copy the skill folder (posit-dev/review-testing in posit-dev/skills) into .agents/skills/review-testing in your project. Codex loads it when a task matches its description.

Can I use Review Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add posit-dev/skills --skill review-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-testing, .gemini/skills/review-testing, .github/skills/review-testing and .opencode/skills/review-testing in your project.

What does Review Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Review Testing is instructions for the agent only.

Does Review Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Review Testing use?

Review Testing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Testing use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Review Testing?

Skills that share tags, products or a category with Review Testing: Requirements (rizsotto/Bear, 6.5k stars), Crap Analysis (ardalis/RiverBooks, 134 stars), Code Coverage (s3s-project/s3s, 311 stars) and Project Status (bactopia/bactopia, 522 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Testing?

posit-dev (a GitHub organization) maintains it in posit-dev/skills, which has 529 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 6, 2026.

Source: posit-dev/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.