Agent skill

Audit Tests

by dzhng in dzhng/skills

Audit whether tests earn their maintenance cost and which suite owns each contract.

MITAuto-check passedTesting & QA

Install Audit Tests

skills CLI
$ npx skills add dzhng/skills --skill audit-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dzhng/skills audit-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/engineering/audit-tests .claude/skills/audit-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-tests
GitHub stars
1k
Token cost
~2.4k tokens
SKILL.md length
1,266 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Audit whether tests earn their maintenance cost and which suite owns each contract.

  • Works in 7 steps: Bound the audit. Read the repository and… → Trace the claim. Read each candidate in… → Record a disposition before editing. Use… → …
  • Pruning redundant tests
  • SKILL.md covers Workflow, Evidence record, False confidence and What earns retention, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audit Tests is an agent skill from dzhng/skills. Audit whether tests earn their maintenance cost and which suite owns each contract. Use when pruning redundant tests, investigating implementation coupling or test-only production hooks, reviewing the value of proposed coverage, or auditing an entire subsystem. Use write-tests to implement the resulting test changes.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation. The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.

When your agent uses it

  • Pruning redundant tests
  • Investigating implementation coupling
  • Test-only production hooks
  • Reviewing the value of proposed coverage

Example prompts

  • “/audit-tests”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Bound the audit. Read the repository and scoped instructions, then pick
  2. Trace the claim. Read each candidate in full, including parameter rows,
  3. Record a disposition before editing. Use the evidence record below.
  4. Choose the surviving owner. Look across the records for whole redundant
  5. Change one coherent boundary. When implementation is in scope, carry
  6. Prove preservation. Compare the removed assertions with the survivors,
  7. Close the batch. Run review, then reconcile the

What it can do on your machine

Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Tests loads about 2.4k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,266 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,266 words, ~2,356 tokens.

Download SKILL.mdSave it as .claude/skills/audit-tests/SKILL.md (or your agent's skills folder).
name
audit-tests
description
Audit whether tests earn their maintenance cost and which suite owns each contract. Use when pruning redundant tests, investigating implementation coupling or test-only production hooks, reviewing the value of proposed coverage, or auditing an entire subsystem. Use write-tests to implement the resulting test changes.

Audit Tests

Audit independent proof: what bug would escape if this test disappeared? The result is a map of contracts to their surviving tests, with evidence for what to keep, repair, combine, or remove. Fewer tests is not the objective; less maintenance for the same or better confidence is.

write-tests owns how to write and falsify a test. This skill owns whether that test adds proof and where that proof belongs. review owns closing out the resulting changes.

Workflow

  1. Bound the audit. Read the repository and scoped instructions, then pick a production responsibility. Find its tests, fixtures, shared suites, and CI selection rules. Record the starting revision, local changes, and test results before editing. Distinguish existing failures from checks that could not run. For a focused audit, choose a few substantiated candidates; for a subsystem sweep, use the completeness rules below.
  2. Trace the claim. Read each candidate in full, including parameter rows, beside the production entry point and the code it reaches. Inspect callers, sibling implementations, overlapping coverage, and the history explaining why the test exists. Inspect dependency code or types when the claim rests on dependency behavior. Write down what the assertions actually detect, even when the test's name promises something else.
  3. Record a disposition before editing. Use the evidence record below. Share findings as they become clear; the workflow must not hold back a material discovery. An audit-only request ends with this evidence and a proposed batch, not with deletions.
  4. Choose the surviving owner. Look across the records for whole redundant layers, not just individual weak assertions. Name the suite that will own each contract and every unique case it must absorb. Prefer a real path with controlled external dependencies over replaying internal collaborators. Another layer earns coverage only for a failure the chosen owner cannot expose, such as routing or lifecycle behavior.
  5. Change one coherent boundary. When implementation is in scope, carry unique cases into the surviving suite using write-tests before removing their old home. Extend existing cases and fixtures instead of cloning setup. Remove newly unneeded test hooks and support code in the same batch; verify production consumers before deleting an export, wrapper, reset function, global, or injection parameter. Update test discovery and CI registration when moving suites.
  6. Prove preservation. Compare the removed assertions with the survivors, looking for lost contracts and assertions that cannot fail for the intended reason. Falsify repaired or transferred regression coverage using write-tests. Isolate temporary production changes, restore the original bytes without disturbing other work, and confirm the targeted proof passes. Report baseline failures separately; investigate any new failures. Run owner and sibling suites, then required repository checks. When replacing static checks, exercise the actual command or dry-run whose contract they claimed. Keep files stable while a runner reads them.
  7. Close the batch. Run review, then reconcile the evidence record with what shipped. Report unresolved failures and unrun checks explicitly. A clean smaller suite does not prove removed coverage was redundant; the contract map and preservation evidence do.

Evidence record

For each candidate, record its location and name, actual detectable failure, reason for existing, production consumers of any supporting hook, and decision:

  • Keep: identify the independently protected contract. A file move alone does not change this decision.
  • Repair: preserve the contract but replace an assertion that is weak, misleading, or attached to the wrong boundary.
  • Combine: identify the destination suite and the unique cases it must prove before the original goes away.
  • Remove: identify stronger surviving proof, or demonstrate that there is no contract to protect. Suspicion of duplication is insufficient.

Include the supporting source/history evidence, cleanup unlocked, risk, and validation command. Missing evidence stays an open question, not a deletion.

False confidence

Use these as investigation prompts, not automatic removal rules:

  • Circular evidence. Expected output comes from the same helper being tested; a fixture copies an inventory from production; a test compares a value with itself. Find an expectation independent of the implementation.
  • An exercise without a claim. Executing lines, copying identities, or listing exports may raise coverage without detecting a credible defect. Determine whether execution itself enforces anything meaningful.
  • The harness does the job. A mock implements the promised behavior, a fixture supplies the acknowledgement or ordering production must generate, or a persistence assertion reads a store the real path never writes. Using one identical fake for different dependency APIs can hide their differences.
  • The wrong cause goes red. A negative case hits a different guard or never reaches the intended branch. Verify the failing condition, not merely that something rejected the input. Check that every parameter row reaches its claim.
  • Declarations posing as behavior. Source greps, import lists, capability flags, and internal call shapes can certify the declaration while the user path is broken. Trace the behavior the declaration promises.
  • One proof repeated. Private helper tests, per-provider replays of shared logic, and multiple layers exercising the same regression can all protect one contract. Keep distinct transport and lifecycle risks, not copies of the same scenario under different filenames.
  • Tests keeping code alive. A helper or production escape hatch exists only because tests call it. Move proof to the real entry point and remove the abandoned mechanism; do not replace it with another test-only abstraction.
Show full SKILL.md (393 more words)Show less

What earns retention

Behavior includes compatibility and operational contracts: API and SDK shape, wire bytes, prompts, defaults, configuration, migrations, storage, security, platform behavior, generated cross-language agreement, package contents, release rules, and architectural boundaries. A static assertion can be their cheapest independent proof. Observable ordering also belongs here.

Ask whether the test survives an internal rename while detecting a broken external promise. Exact bytes may be essential to a protocol; exact private identifiers usually are not. Neither slowness nor resemblance to implementation is sufficient evidence for deletion. Do not add refactor-sensitive coverage when the real contract can be exercised directly.

An existing failure may be a product defect. Reproduce and investigate it; never remove a valid claim to make the audit green. Repair in-scope defects separately, with a failing control and passing fix on the same harness. Record unrelated defects as follow-ups and keep their coverage.

Whole-subsystem sweeps

Scale the accounting before scaling the edits:

  • Inventory every owned test, including cases in shared suites and integration, QA, or live harnesses. Assign each to one production responsibility, even when file prefixes disagree. Record baseline outcomes and test/support size.
  • Give every declaration a disposition and evidence. A parameter table can share one record only if all rows deserve the same decision. Then make a second pass across the records to choose surviving suites and retired layers.
  • Delegate read-only discovery by responsibility when agents are available. Keep edits to shared fixtures under one owner. Each batch must introduce no unexplained failures; tracked baseline failures remain visible for repair.
  • Use an independent reviewer to compare removed coverage with its new owners. Resolve every reported gap with restored proof or source evidence rejecting it. Demonstrate that each restored contract catches a deliberate defect.
  • If the base advances, inspect new cases in any file being removed and carry their contracts forward. Follow the repository's integration policy, then rerun the subsystem and required integration checks on the resulting revision. Refresh discovery before starting the next batch.

Done

Hand off the contract map, decisions and reasons, production simplifications, valuable tests deliberately retained, preservation gaps and their resolution, and verification actually run. For broad sweeps, include before/after size for production, tests, and support separately, plus remaining batches and integration state. Size describes the outcome; it never sets the deletion target. Claim a complete subsystem audit only when the inventory has no unexamined entries.

© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/engineering/audit-tests of dzhng/skills.

Open the folder on GitHubat commit d513228

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Audit Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Tests this skilldzhng/skills1k—~2.4kAutomated safety check: PassMIT
Emcaklofas/kicad-happy1.4k1 repos~2.8kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Generate Test Cases342164796/generate-test-cases1191 repos~2.9kAutomated safety check: PassNone
Verify Cc Safety Netkenryu42/cc-safety-net1.6k—~2kAutomated safety check: PassMIT
Wioworkersio/skills190—~5.8kAutomated safety check: PassMIT

Similar skills

  • Emc

    aklofas/kicad-happy

    EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…

    1.4k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check passed
  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Generate Test Cases

    342164796/generate-test-cases

    自主学习型测试文档生成器。从需求文档(Markdown)生成测试用例 XMind 文件,支持持久化记忆和持续学习。当用户提到"生成测试用例"、"根据需求生成测试"时触发。

    119 GitHub starsUsed in 1 repo~2.9k tokens
    Testing & QAAuto-check passed
  • Verify Cc Safety Net

    kenryu42/cc-safety-net

    Launch and drive the real cc-safety-net CLI — the hook decision path, explain, status/doctor, logs, and the local policy GUI — against an isolated home, capturing evidence.

    1.6k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    190 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • File Server

    microsoft/WindowsProtocolTestSuites

    Official

    ALWAYS LOAD THIS SKILL when working with FileServer, SMB, SMB2, SMB3, CIFS, file sharing, MS-SMB2, MS-FSCC, MS-FSA, MS-DFSC, MS-FSRVP, MS-RSVD, MS-SQOS, or any file server protocol test…

    567 GitHub stars~4.1k tokensUpdated 23 days ago
    Testing & QAAuto-check passed

More from dzhng/skills

All 27 skills in this repo
  • Compare screenshots against the intended design, distinguishing approved references from historical baselines.

    1k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Claude

    dzhng/skills

    Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.

    1k GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Refactor Clean

    dzhng/skills

    Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.

    1k GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed
  • Write Skills

    dzhng/skills

    Create or revise agent skills. An agent skill from dzhng/skills.

    1k GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Codex

    dzhng/skills

    Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.

    1k GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check: warnings
  • Run [implement-spec](../implement-spec/SKILL.md) with Codex doing the implementation passes while you orchestrate, integrate, and review.

    1k GitHub starsUsed in 1 repo~543 tokens
    Auto-check passed

Categories

Questions about Audit Tests

What does Audit Tests do?

Audit whether tests earn their maintenance cost and which suite owns each contract. Audit Tests is an agent skill from dzhng/skills. Audit whether tests earn their maintenance cost and which suite owns each contract.

When should I use Audit Tests?

Audit Tests fits situations like: pruning redundant tests; investigating implementation coupling; test-only production hooks; reviewing the value of proposed coverage.

How do I install Audit Tests in Claude Code?

Run `npx skills add dzhng/skills --skill audit-tests -a claude-code`. Or copy the skill folder (skills/engineering/audit-tests in dzhng/skills) into .claude/skills/audit-tests in your project. Claude Code loads it when a task matches its description.

How do I install Audit Tests in Codex?

Run `npx skills add dzhng/skills --skill audit-tests -a codex`. Or copy the skill folder (skills/engineering/audit-tests in dzhng/skills) into .agents/skills/audit-tests in your project. Codex loads it when a task matches its description.

Can I use Audit Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill audit-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-tests, .gemini/skills/audit-tests, .github/skills/audit-tests and .opencode/skills/audit-tests in your project.

What does Audit Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Tests is instructions for the agent only.

Does Audit Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Tests use?

Audit Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Tests use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Tests?

Skills that share tags, products or a category with Audit Tests: Emc (aklofas/kicad-happy, 1.4k stars), Swig Test (swig/swig, 6.3k stars), Generate Test Cases (342164796/generate-test-cases, 119 stars) and Verify Cc Safety Net (kenryu42/cc-safety-net, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Tests?

dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,016 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.

Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.