Agent skill

Test Audit

by mblode in mblode/agent-skills

Prunes low-value tests across a whole suite to a measured target while holding coverage.

MITAuto-check passedTesting & QA

Install Test Audit

skills CLI
$ npx skills add mblode/agent-skills --skill test-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mblode/agent-skills test-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-audit .claude/skills/test-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-audit
GitHub stars
143
Token cost
~2.2k tokens
SKILL.md length
1,264 words
Files
4 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Prunes low-value tests across a whole suite to a measured target while holding coverage.

  • Works in 4 steps: Run the owner and sibling tests for each… → Where a deleted test grepped source or… → Re-measure coverage with the same… → …
  • Asked to remove useless tests
  • SKILL.md covers Pick a mode, Set a target before starting, Junk patterns and Retention bar, plus 9 more sections
  • Calls git

What it does

Test Audit is an agent skill from mblode/agent-skills. Prunes low-value tests across a whole suite to a measured target while holding coverage. Use when asked to "remove useless tests", "prune the test suite", "we have too many tests", "agents keep writing pointless tests", or "cut 20% of tests without losing coverage".

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `evals/evals.json`, `references/campaign.md` and `references/coverage.md`). Compatibility notes: Needs the repository's test runner and a coverage tool it already supports (Vitest, Jest, c8, pytest-cov, go test -cover, or equivalent).

It sits in Testing & QA, covering Test generation. The repository describes itself as: Nobody ships AI slop on purpose. These skills make sure you don’t. The licence is MIT.

When your agent uses it

  • Asked to remove useless tests
  • Prune the test suite
  • We have too many tests
  • Agents keep writing pointless tests

Example prompts

  • “remove useless tests”
  • “prune the test suite”
  • “we have too many tests”
  • “/test-audit”

Requirements

  • Compatibility (from SKILL.md): Needs the repository's test runner and a coverage tool it already supports (Vitest, Jest, c8, pytest-cov, go test -cover, or equivalent).

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run the owner and sibling tests for each batch.
  2. Where a deleted test grepped source or asserted a plan, run the script or dry-run that owns the real contract.
  3. Re-measure coverage with the same command as the baseline and compare against the budget, per file.
  4. Report git diff --numstat with production and tooling lines separate from test lines.

What it can do on your machine

Read from SKILL.md and the folder at commit cef4cfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Needs the repository's test runner and a coverage tool it already supports (Vitest, Jest, c8, pytest-cov, go test -cover, or equivalent).

    From compatibility in the SKILL.md frontmatter.

Context cost

Test Audit loads about 2.2k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 1,264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mblode/agent-skills at commit cef4cfa, republished under its MIT licence (© mblode). 1,264 words, ~2,202 tokens.

Download SKILL.mdSave it as .claude/skills/test-audit/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
test-audit
description
Prunes low-value tests across a whole suite to a measured target while holding coverage. Use when asked to "remove useless tests", "prune the test suite", "we have too many tests", "agents keep writing pointless tests", or "cut 20% of tests without losing coverage".
compatibility
Needs the repository's test runner and a coverage tool it already supports (Vitest, Jest, c8, pytest-cov, go test -cover, or equivalent).

Test Audit

Tests earn their place by catching a credible regression that nothing else catches. Agents write a test for every small change, and most of those re-assert the source, replay a stronger test through a mock, or pin a private call shape. Coverage barely moves when they go; maintenance cost and suite time do.

  • IS: sweeping a directory for low-value tests, and campaign-pruning a whole subsystem or repo to a measured target with coverage held, plus deleting the test-only production seams those tests kept alive.
  • IS NOT: gating a test before it lands or reviewing a feature diff (tidy), making CI faster by splitting or sharding (ci-speedup), or writing tests for untested code.

Pick a mode

ModeWhenDone means
AuditA focused sweep of named files or a directoryA handful of high-confidence candidates with full evidence, deleted in one coherent batch with validation green
CampaignA subsystem's or repo's whole test surface, or any request with a target ("remove 20%")The target is met or the ledger shows why it cannot be, coverage held within tolerance, and the preservation review passed. Read references/campaign.md first

Set a target before starting

An open-ended "clean up the tests" stops after a few obvious deletions, because every remaining candidate looks defensible in isolation. Audit and Campaign need a number to work against. When the user gave one, use it. When they did not, propose one and proceed on it rather than waiting:

Remove the least useful 20% of test declarations in <scope>, keeping line and branch coverage within 2 percentage points of the baseline and every retained test green.

Measure the baseline first (references/coverage.md), then keep deleting in order of lowest value until the target is met or the coverage budget is spent. Stopping short is a valid result only with a ledger showing that the remaining candidates each guard a named contract. The target is a floor on effort, not permission to cut uncertain tests: the retention bar below still wins over the number.

Junk patterns

What an audit or campaign hunts for:

  • assertion-free coverage probes, and tests whose only assertion is "does not throw" on a path that cannot throw
  • self-comparisons, identity copies, and expected values produced by the helper or renderer under test
  • copied fixtures, inventories, manifests, export lists, or snapshots of data the test also defines
  • exact source, import, or string greps that fail on a rename and survive a real behavior change
  • private predicate or call-shape tests (toHaveBeenCalledWith on an internal) duplicated at a real boundary
  • duplicate invocations of one contract, and per-provider replays of a shared helper
  • tests whose only purpose is keeping a test-only export, global, or wrapper alive, and production code whose only callers are tests
  • mocks that implement the asserted behavior, or one mock standing in for several different APIs
  • fixtures that supply the ordering, receipt, or callback the code under test should produce, or persistence asserted against a store the path never writes
  • capability tests that restate declared flags instead of exercising what the flag promises
  • negative controls that pass for an unrelated reason, such as a rejection from a different guard
  • names that promise more than the assertions check: judge the test by its assertions, not its title

Retention bar

Keep a test when it independently enforces a public API, protocol, config, migration, storage, security, platform, default value, generated cross-language, package, release, or architecture contract. Also keep:

  • call ordering when the order is observable behavior
  • a regression with a credible failure mode
  • a source inspection when it is the cheapest independent guard: it fails when the user-facing key, byte, or path changes and survives an identifier-only rename
  • a retained test that fails on the baseline: treat it as a possible product bug, reproduce it, and fix the owner rather than deleting it

Static or slow is not a reason to delete. A test that resembles implementation may still be the only proof of a contract; show otherwise before removing it.

Evidence per candidate

Before judging, read the complete test, the production owner, its callers and callees, sibling implementations, overlapping tests, and the history that explains why the test exists. Record for each deletion:

  • test name and location, and what failure it can actually detect
  • non-test callers of the seam it covers
  • the stronger proof that remains (name the keeper test), or why no contract exists
  • the production or test-support code the deletion unlocks
  • coverage delta from removing it, when measured

A missing field means the candidate is not ready. Prefer a few high-confidence deletions over a long speculative list; do not convert uncertain candidates into deletions to reach the number.

Show full SKILL.md (494 more words)Show less

Edit shape

Delete in coherent owner-boundary batches. With each batch, delete the test-only exports, globals, wrappers, reset hooks, and dead production paths it unlocks instead of keeping aliases. Move retained regressions to their canonical owner and fold duplicated setup into shared fixtures. Aim for net-negative production lines as well as test lines. Do not add replacement tests that restate the same implementation.

Validation

  1. Run the owner and sibling tests for each batch.
  2. Where a deleted test grepped source or asserted a plan, run the script or dry-run that owns the real contract.
  3. Re-measure coverage with the same command as the baseline and compare against the budget, per file.
  4. Report git diff --numstat with production and tooling lines separate from test lines.

Permissions

Running tests and coverage locally, and deleting tests and test-only seams in the working tree, are the task: do them without asking. Commit, push, or open a PR only when asked. A test that meets the retention bar stays; deleting one anyway is the user's call, as is changing coverage thresholds or CI gates.

Handoff

  • target, baseline, and final numbers: declarations or lines, and line and branch coverage
  • removed categories and the root cause behind them (what kept producing these tests)
  • production simplifications, counted separately from test lines
  • retained false positives and why each stays
  • validation actually run, with results
  • a proposed gate for the repo's AGENTS.md, drawn from the patterns this audit actually found, so agents stop writing them

Gotchas

  • Line coverage survives deletions that remove the only assertion on a path, because another test still executes the lines without checking them. Where a deletion looks free on coverage, confirm a keeper asserts the behavior, or run a mutation against the owner (references/coverage.md).
  • Aggregate coverage hides a subsystem that lost its only proof while an unrelated area gained. Compare per file for every production file the deleted tests touched.
  • Ranking candidates by unique line coverage alone deletes the wrong tests: a test with zero unique lines can still hold the only assertion on them. Unique coverage makes a test a candidate; a named keeper makes it a deletion.
  • Editing tests while Vitest or Jest runs in watch mode on the same checkout re-runs half-edited files, and the failures read as regressions from the deletion. Stop the watcher before a batch.

Reference files

FileRead when
references/campaign.mdCampaign mode, or any target-driven prune larger than one directory
references/coverage.mdMeasuring the baseline, per-file and per-test coverage, the budget check, and mutation spot checks

Maintenance only: evals/evals.json holds scenarios and routing prompts for changing this skill; it never loads during a task.

  • tidy: gates a test the moment a diff adds or changes it, and reviews the rest of the diff alongside it. This skill only sweeps or campaign-prunes the existing suite; "audit these tests in my PR" routes to tidy, not here.

Credit

Adapted from OpenClaw's test-audit skill (openclaw/openclaw, MIT). Provenance lives in the repository's maintenance notes.

© mblode, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/test-audit of mblode/agent-skills.

  • SKILL.md
  • evals/evals.json
  • references/campaign.md
  • references/coverage.md

Open the folder on GitHubat commit cef4cfa

Compare with similar skills

Test Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Audit this skillmblode/agent-skills143—~2.2kAutomated safety check: PassMIT
Emcaklofas/kicad-happy1.4k1 repos~2.8kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Generate Test Cases342164796/generate-test-cases1191 repos~2.9kAutomated safety check: PassNone
Verify Cc Safety Netkenryu42/cc-safety-net1.6k—~2kAutomated safety check: PassMIT
Wioworkersio/skills190—~5.8kAutomated safety check: PassMIT

Similar skills

  • Emc

    aklofas/kicad-happy

    EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…

    1.4k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check passed
  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Generate Test Cases

    342164796/generate-test-cases

    自主学习型测试文档生成器。从需求文档(Markdown)生成测试用例 XMind 文件,支持持久化记忆和持续学习。当用户提到"生成测试用例"、"根据需求生成测试"时触发。

    119 GitHub starsUsed in 1 repo~2.9k tokens
    Testing & QAAuto-check passed
  • Verify Cc Safety Net

    kenryu42/cc-safety-net

    Launch and drive the real cc-safety-net CLI — the hook decision path, explain, status/doctor, logs, and the local policy GUI — against an isolated home, capturing evidence.

    1.6k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    190 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • File Server

    microsoft/WindowsProtocolTestSuites

    Official

    ALWAYS LOAD THIS SKILL when working with FileServer, SMB, SMB2, SMB3, CIFS, file sharing, MS-SMB2, MS-FSCC, MS-FSA, MS-DFSC, MS-FSRVP, MS-RSVD, MS-SQOS, or any file server protocol test…

    567 GitHub stars~4.1k tokensUpdated 23 days ago
    Testing & QAAuto-check passed

More from mblode/agent-skills

All 28 skills in this repo
  • Agent Ready

    mblode/agent-skills

    Implements agent-readiness on public sites and docs from Mintlify Agent Score, AFDocs, Is Agentic, Is It Agent Ready, or url-discovery-bench reports, or from server logs of agents 404ing on guessed…

    143 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Agent Skills Creator

    mblode/agent-skills

    Creates and improves portable Agent Skills with a validator, routing scenarios, and evidence-based keep, cut, merge, or retire decisions.

    143 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Chat History

    mblode/agent-skills

    Recovers decisions, previous fixes, research, and what followed a prompt from past AI conversations, with source evidence.

    143 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • CI Speedup

    mblode/agent-skills

    Cuts the wait from push to green by measuring a pipeline's critical path from run timestamps, then splitting, sharding, trimming setup and sharing test module state, with a before/after ledger.

    143 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • PR Babysitter

    mblode/agent-skills

    Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.

    143 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • App Verification

    mblode/agent-skills

    Builds and maintains a repo's own verification harness (verify CLI, doctor, worktree isolation, feature map, seed data) and a reproduce-first bug handoff.

    143 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed

Categories

Questions about Test Audit

What does Test Audit do?

Prunes low-value tests across a whole suite to a measured target while holding coverage. Test Audit is an agent skill from mblode/agent-skills. Prunes low-value tests across a whole suite to a measured target while holding coverage.

When should I use Test Audit?

Test Audit fits situations like: asked to remove useless tests; prune the test suite; we have too many tests; agents keep writing pointless tests.

How do I install Test Audit in Claude Code?

Run `npx skills add mblode/agent-skills --skill test-audit -a claude-code`. Or copy the skill folder (skills/test-audit in mblode/agent-skills) into .claude/skills/test-audit in your project. Claude Code loads it when a task matches its description.

How do I install Test Audit in Codex?

Run `npx skills add mblode/agent-skills --skill test-audit -a codex`. Or copy the skill folder (skills/test-audit in mblode/agent-skills) into .agents/skills/test-audit in your project. Codex loads it when a task matches its description.

Can I use Test Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mblode/agent-skills --skill test-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-audit, .gemini/skills/test-audit, .github/skills/test-audit and .opencode/skills/test-audit in your project.

What does Test Audit need to run?

Going by SKILL.md and its folder, Test Audit needs the command-line tools its instructions call (git). Compatibility (from SKILL.md): Needs the repository's test runner and a coverage tool it already supports (Vitest, Jest, c8, pytest-cov, go test -cover, or equivalent)..

Does Test Audit access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Test Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Audit use?

Test Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Audit use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Test Audit?

Skills that share tags, products or a category with Test Audit: Emc (aklofas/kicad-happy, 1.4k stars), Swig Test (swig/swig, 6.3k stars), Generate Test Cases (342164796/generate-test-cases, 119 stars) and Verify Cc Safety Net (kenryu42/cc-safety-net, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Audit?

mblode (a GitHub user) maintains it in mblode/agent-skills, which has 143 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.

Source: mblode/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.