---
name: test-audit
description: Audit existing tests for low value, implementation coupling, duplication, and test-only production seams; apply the value bar when reviewing test changes.
license: MIT
---

# Test Audit

Adapted from [OpenClaw's test-audit skill](https://github.com/openclaw/openclaw/blob/80930af448ebabc84174146b56bc106d37fab3b4/.agents/skills/test-audit/SKILL.md).
See [NOTICE](NOTICE) and [LICENSE](LICENSE).

Optimize for confidence, not deletion count. Audit mode discovers a few high-confidence
candidates and reports evidence before editing. Authoring mode checks proposed tests at
write time. For an explicitly requested whole-subsystem pruning campaign, read
[CAMPAIGN.md](CAMPAIGN.md).

Read root and scoped `AGENTS.md` files and the repository's [testing skill](../testing/SKILL.md)
and [test selection](../testing/references/selection.md). Those policies govern whether new
automation is justified; this skill does not authorize replacement tests, production changes,
commits, publication, or deployment. Keep reports and investigation evidence outside the
product repository, or in task/PR artifacts.

## Authoring gate

Before adding or expanding a test, establish why automation is necessary under the testing
policy, then answer:

1. What observable behavior, invariant, or independent contract does it protect?
2. What credible regression makes it fail?
3. Why does existing proof not catch that failure? Each contract has one primary owner at
   the strongest boundary. Another layer needs a distinct transport, lifecycle, or adapter
   risk that the owner cannot exercise. Prefer an existing table or fixture to duplication.
4. Does it require a production export, flag, wrapper, or injection hook with no production
   caller? Exercise the real boundary instead.

A missing answer means the test is not ready. Check every junk pattern below; retain a match
only when it independently guards a named contract. Behavior-preserving refactoring should
not break a behavior test. A regression must demonstrably fail on the pre-fix owner for the
intended reason and pass after the repair. Keep one primary regression at its owner boundary.

## Junk patterns

- Assertion-free coverage probes, self-comparisons, and identity copiers.
- Copied fixtures, inventories, manifests, export lists, or schema definitions.
- Exact source, import, or string greps tied to incidental implementation.
- Private predicates or call shapes already exercised at a real boundary.
- Duplicate invocations of a contract, including adapter-local replays of shared helpers.
- Tests that preserve test-only exports, globals, wrappers, or otherwise dead production code.
- Expected values generated by the helper or renderer under test.
- Mocks that implement the asserted behavior, or one identical mock for different APIs.
- Fixtures that supply the receipt, admission, or callback ordering the owner must produce;
  persistence assertions against a store the production path never writes.
- Capability tests that restate flags rather than exercise promised delivery or acknowledgement.
- Negative controls that pass because of an unrelated guard or unreachable rejection path.
- Names that promise more than the input and assertions exercise.

## Discovery and retention

Keep discovery read-only. For broad scope, use independent owner-boundary lanes across core,
capabilities/durability, storage, platform adapters, examples, tooling, and cross-cutting
patterns. Use an available `orchestrate` skill when applicable; keep workers
read-only with disjoint scopes and one primary integrator.

Before judging a candidate, read the complete test and production owner, entry point,
callers, callees, sibling implementations, overlapping tests, CI routing, and relevant
history. Inspect dependency source or types when a test claims dependency-backed behavior.
Judge what the assertions can detect, not their names, length, citations, or test count.

Keep independent public API, protocol, config, migration, storage, security, platform,
default, prompt-byte, generated cross-language, package, release, and architecture contracts
when their failure would escape remaining checks and repeatable workflow verification is
insufficient. Preserve narrow compile-time inference checks and adapter proof with distinct
risks. Observable call ordering and credible regressions can also justify retention.

Source inspection can be the cheapest independent guard when it protects a user-facing key,
byte, path, or authority contract and survives identifier-only refactoring. Static or slow is
not a deletion reason. A retained baseline failure is a possible product bug: reproduce it
and report or repair its owner within scope rather than deleting the test.

## Candidate evidence

Record each field before proposing an edit; a missing field means deletion is not ready:

- Exact test name and location.
- Failure the test can actually detect.
- Non-test callers of the production or support seam.
- Stronger remaining owner-boundary proof, or why no independent contract needs proof.
- Relevant history and the reason the test or seam exists; label unknown history.
- Production or test-support deletion unlocked, if any.
- Risk and the focused Vite+ validation command.

Separate verified candidates, retained false positives, and unresolved hypotheses.

## Authorized edits and validation

When cleanup is in scope, choose one coherent owner-boundary batch. Remove obsolete test-only
seams and orphaned support instead of preserving aliases. Consolidate duplicate package or
dependency assertions at their canonical owner. Preserve public and persisted contracts;
uncertainty is not evidence for deletion. Do not write replacement tests that restate the
same implementation or increase cleanup scope to inflate deletion counts.

Do not edit source or tests while their suites run in the same checkout. Follow repository
command authority: use `vp run -F <workspace> test` for workspace suites or
`vp -C <package-directory> test <file-filter>` for focused proof, consulting command help
and runner configuration before filtering. Root `vp test` filename filters can also
match tests inside `.worktrees`; scope the runner to its package when worktrees exist.
For removed source greps, exercise the executable or dry-run that owns the contract.
Run applicable formatting and `git diff --check`, then the required `vp run ready` handoff
gate. Preserve blocked proof explicitly. Report production/tooling separately from tests
and test support using the final diff. For PRs, use [open-pull-request](../open-pull-request/SKILL.md).

## Handoff

Report actionable findings and their remaining proof, production simplifications available,
retained false positives and reasons, focused/full proof actually run, blocked proof, and
named follow-ups. If edits occurred, include production versus test LOC and PR/merge state.
For read-only audits, say explicitly that tests and production source were unchanged.
