Agent skill

Test Review

by athola in athola/claude-night-market

Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns.

MITAuto-check passedTesting & QA

Install Test Review

skills CLI
$ npx skills add athola/claude-night-market --skill test-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install athola/claude-night-market test-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/pensive/skills/test-review .claude/skills/test-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-review
GitHub stars
341
Token cost
~1.8k tokens
SKILL.md length
694 words
Files
6
Skills in repo
152
Repo updated
First seen
Licence
MIT

At a glance

Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns.

  • Works in 5 steps: Detect Languages… → Inventory Coverage… → Assess Scenario Quality… → …
  • Auditing test quality
  • SKILL.md covers Quick Start, When To Use, When NOT To Use and Required TodoWrite Items, plus 8 more sections
  • Calls pytest, git and rg

What it does

Test Review is an agent skill from athola/claude-night-market. Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns. Use when auditing test quality or before a major release.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `modules/content-assertion-quality.md`, `modules/coverage-analysis.md` and `modules/framework-detection.md`).

It sits in Testing & QA, covering Test generation, Test coverage and Test-driven development. The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.

When your agent uses it

  • Auditing test quality
  • Before a major release

Example prompts

  • “Use the test-review skill to evaluate test suites for coverage gaps, TDD/BDD compliance, and anti-patterns”
  • “/test-review”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Detect Languages (test-review:languages-detected)
  2. Inventory Coverage (test-review:coverage-inventoried)
  3. Assess Scenario Quality (test-review:scenario-quality)
  4. Plan Remediation (test-review:gap-remediation)
  5. Log Evidence (test-review:evidence-logged)

What it can do on your machine

Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • git
    • rg
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Review loads about 1.8k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 694 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 694 words, ~1,803 tokens.

Download SKILL.mdSave it as .claude/skills/test-review/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
test-review
description
Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns. Use when auditing test quality or before a major release.
alwaysApply
false
category
testing
tags
testing, tdd, bdd, coverage, quality, fixtures
usage_patterns
test-audit, coverage-analysis, quality-improvement, gap-remediation
complexity
intermediate
model_hint
standard
estimated_tokens
200
progressive_loading
true
dependencies
imbue:proof-of-work, imbue:review-core, imbue:structured-output
modules
modules/framework-detection.md, modules/coverage-analysis.md, modules/scenario-quality.md, modules/remediation-planning.md, modules/content-assertion-quality.md

Test Review Workflow

Evaluate and improve test suites with TDD/BDD rigor.

Quick Start

bash
/test-review

Verification: Run pytest -v to verify tests pass.

When To Use

  • Reviewing test suite quality
  • Analyzing coverage gaps
  • Before major releases
  • After test failures
  • Planning test improvements

When NOT To Use

  • Writing new tests - use parseltongue:python-testing
  • Updating existing tests - use sanctum:test-updates

Required TodoWrite Items

  1. test-review:languages-detected
  2. test-review:coverage-inventoried
  3. test-review:scenario-quality
  4. test-review:invariant-preservation
  5. test-review:gap-remediation
  6. test-review:evidence-logged
  7. test-review:findings-verified

Progressive Loading

Load modules as needed based on review depth:

  • Basic review: Core workflow (this file)
  • Framework detection: Load modules/framework-detection.md
  • Coverage analysis: Load modules/coverage-analysis.md
  • Quality assessment: Load modules/scenario-quality.md
  • Remediation planning: Load modules/remediation-planning.md

Workflow

Step 1: Detect Languages (test-review:languages-detected)

Identify testing frameworks and version constraints. → See: modules/framework-detection.md

Quick check:

bash
find . -maxdepth 2 -name "Cargo.toml" -o -name "pyproject.toml" -o -name "package.json" -o -name "go.mod"

Verification: Run the command with --help flag to verify availability.

Step 2: Inventory Coverage (test-review:coverage-inventoried)

Run coverage tools and identify gaps. → See: modules/coverage-analysis.md

Quick check:

bash
git diff --name-only | rg 'tests|spec|feature'

Verification: Run pytest -v to verify tests pass.

Step 3: Assess Scenario Quality (test-review:scenario-quality)

Evaluate test quality using BDD patterns and assertion checks. → See: modules/scenario-quality.md

Focus on:

  • Given/When/Then clarity
  • Assertion specificity
  • Anti-patterns (dead waits, mocking internals, repeated boilerplate)
Step 4: Plan Remediation (test-review:gap-remediation)

Create concrete improvement plan with owners and dates. → See: modules/remediation-planning.md

Step 5: Log Evidence (test-review:evidence-logged)

Record executed commands, outputs, and recommendations. → See: imbue:proof-of-work

Test Quality Checklist (Condensed)

  • Clear test structure (Arrange-Act-Assert)
  • Critical paths covered (auth, validation, errors)
  • Specific assertions with context
  • No flaky tests (dead waits, order dependencies)
  • Reusable fixtures/factories
  • Invariant-encoding tests intact (see below)
Invariant-Encoding Tests

Tests encode design invariants as well as verifying behavior. A test that asserts "module A never imports from module B" encodes a layer boundary. A test that asserts "this function is pure" encodes a concurrency model. These tests are load-bearing in ways that coverage metrics cannot capture.

During review, check:

  1. Were invariant-encoding tests removed or weakened? A test that enforced an architectural boundary, data structure constraint, or API contract should not be deleted without naming the invariant being abandoned and escalating to human judgment.

  2. Were test expectations changed to match a broken implementation? If an assertion value changed, ask: did the requirement change, or did the agent change the test to make its code pass? The latter is the single most dangerous form of test tampering.

  3. Are new invariants encoded as tests? When a design decision is made (choice of data structure, module boundary, error strategy), there should be at least one test whose failure would signal that the invariant was violated.

Red flag patterns:

PatternRisk
@pytest.mark.skip added to a passing testInvariant being silently dropped
Assertion changed from specific to broadConstraint being relaxed
Test renamed to describe new behaviorOld invariant erased from history
Test deleted "because it tested old code"Invariant removed without replacement
Show full SKILL.md (239 more words)Show less

When invariant erosion is detected:

Do NOT approve. Flag as a BLOCKING quality issue and present the three options to the human:

  1. Preserve: Revert the test change, fix the implementation to satisfy the invariant
  2. Layer: Keep the invariant test, add the new behavior alongside it (accepting inelegance)
  3. Revise: The invariant is genuinely wrong; remove the old test AND write a new test encoding the replacement invariant

This is a judgment call that models get wrong far too often. Default to option 1 (preserve) when no human is available.

Output Format

markdown
## Summary
[Brief assessment]

## Framework Detection
- Languages: [list] | Frameworks: [list] | Versions: [constraints]

## Coverage Analysis
- Overall: X% | Critical: X% | Gaps: [list]

## Quality Issues
[Q1] [Issue] - Location - Anchor: `verbatim source text at file:line` - Fix

## Remediation Plan
1. [Action] - Owner - Date

## Recommendation
Approve / Approve with actions / Block

Verification: Run the command with --help flag to verify availability.

Integration Notes

  • Use imbue:proof-of-work for reproducible evidence capture
  • Reference imbue:diff-analysis for risk assessment
  • Format output using imbue:structured-output patterns

Verify Findings Are Grounded (test-review:findings-verified)

Write findings to .review/findings.json, run the citation verifier (Skill(imbue:review-core) Step 5), and drop or label UNVERIFIED any the verifier rejects.

Exit Criteria

  • Frameworks detected and documented
  • Coverage analyzed and gaps identified
  • Scenario quality assessed
  • Remediation plan created with owners and dates
  • Evidence logged with citations
  • Every reported finding carries a Location + verbatim Anchor confirmed by citation_verifier.py (exit 0), or unverified findings were dropped or labeled UNVERIFIED

Troubleshooting

Common Issues

Tests not discovered Ensure test files match pattern test_*.py or *_test.py. Run pytest --collect-only to verify.

Import errors Check that the module being tested is in PYTHONPATH or install with pip install -e .

Async tests failing Install pytest-asyncio and decorate test functions with @pytest.mark.asyncio

© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in plugins/pensive/skills/test-review of athola/claude-night-market.

  • SKILL.md
  • modules/content-assertion-quality.md
  • modules/coverage-analysis.md
  • modules/framework-detection.md
  • modules/remediation-planning.md
  • modules/scenario-quality.md

Open the folder on GitHubat commit 9f3eb00

Compare with similar skills

Test Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Review this skillathola/claude-night-market341—~1.8kAutomated safety check: PassMIT
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Automated Test Planningtestdouble/han281—~6.7kAutomated safety check: PassMIT
Manual Test Planningtestdouble/han281—~2.9kAutomated safety check: PassMIT
Test Experteinverne/dotfiles121—~2.3kAutomated safety check: PassGPL-3.0
TDD Guidealirezarezvani/claude-skills28k—~3.4kAutomated safety check: PassMIT

Similar skills

  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Produce a standalone test plan by analyzing code for test coverage gaps and edge cases.

    281 GitHub stars~6.7k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Manual Test Planning

    testdouble/han

    Produce a plain-language manual test plan from the context supplied to it — an executive summary, a high-level list of named tests, and a detail section per test with the steps a person follows by…

    281 GitHub stars~2.9k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Test Expert

    einverne/dotfiles

    Testing methodologies, test-driven development (TDD), unit and integration testing, and testing best practices across multiple frameworks.

    121 GitHub stars~2.3k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Prd V07 Test Planning

    mattgierhart/PRD-driven-context-engineering

    Define test cases BEFORE implementation, ensuring every API, business rule, and user journey has verifiable acceptance criteria during PRD v0.7 Build Execution.

    180 GitHub stars~3.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes

More from athola/claude-night-market

All 152 skills in this repo
  • Night Market Diagnostics Toolkit

    athola/claude-night-market

    Run and interpret repo diagnostic scripts (ratchets, validators, token stats).

    341 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Skills Eval

    athola/claude-night-market

    Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.

    341 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Agent Teams

    athola/claude-night-market

    Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.

    341 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Delegation Core

    athola/claude-night-market

    Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).

    341 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Elegant Code

    athola/claude-night-market

    Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.

    341 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Skill Library Mission

    athola/claude-night-market

    Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.

    341 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Test Review

What does Test Review do?

Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns. Test Review is an agent skill from athola/claude-night-market. Evaluates test suites for coverage gaps, TDD/BDD compliance, and anti-patterns.

When should I use Test Review?

Test Review fits situations like: auditing test quality; before a major release.

How do I install Test Review in Claude Code?

Run `npx skills add athola/claude-night-market --skill test-review -a claude-code`. Or copy the skill folder (plugins/pensive/skills/test-review in athola/claude-night-market) into .claude/skills/test-review in your project. Claude Code loads it when a task matches its description.

How do I install Test Review in Codex?

Run `npx skills add athola/claude-night-market --skill test-review -a codex`. Or copy the skill folder (plugins/pensive/skills/test-review in athola/claude-night-market) into .agents/skills/test-review in your project. Codex loads it when a task matches its description.

Can I use Test Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill test-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-review, .gemini/skills/test-review, .github/skills/test-review and .opencode/skills/test-review in your project.

What does Test Review need to run?

Going by SKILL.md and its folder, Test Review needs the command-line tools its instructions call (pytest, git, rg and pip). Our summary lists: Python 3.

Does Test Review access the network?

SKILL.md contains no URLs. Its commands use git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Review use?

Test Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Review use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Review?

Skills that share tags, products or a category with Test Review: Designing Tests (CloudAI-X/opencode-workflow, 275 stars), Automated Test Planning (testdouble/han, 281 stars), Manual Test Planning (testdouble/han, 281 stars) and Test Expert (einverne/dotfiles, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Review?

athola (a GitHub user) maintains it in athola/claude-night-market, which has 341 GitHub stars. The repository holds 152 skills in this directory. The repository was last updated on October 9, 2026.

Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.