Agent skill

Test Audit

by fcakyon in fcakyon/claude-codex-settings

This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or…

MITAuto-check passedDevelopment

Install Test Audit

skills CLI
$ npx skills add fcakyon/claude-codex-settings --skill test-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/claude-codex-settings test-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/claude-codex-settings.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/test-audit/skills/test-audit .claude/skills/test-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-audit
GitHub stars
1.2k
Token cost
~1.9k tokens
SKILL.md length
1,054 words
Files
3 (incl. references)
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or…

  • Works in 4 steps: What observable behavior, invariant, or… → What credible regression makes it fail? → Why does existing coverage not already… → …
  • Asks to write a test
  • SKILL.md covers Authoring gate, Junk patterns, Value bar and Discovery, plus 6 more sections
  • Calls git

What it does

Test Audit is an agent skill from fcakyon/claude-codex-settings. This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or sweeping tests. Gates new tests and audits low-value, implementation-coupled, or duplicative tests and the test-only production seams they demand.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/campaign.md`).

It sits in Development. The repository describes itself as: Battle-tested Claude Code, OpenAI Codex, Cursor configs, plugins, hooks and agents with Kimi, MiniMax and GLM API support. The licence is MIT.

When your agent uses it

  • Asks to write a test
  • Review these tests
  • Prune low-value tests
  • Find duplicate tests

Example prompts

  • “write a test”
  • “review these tests”
  • “audit tests”
  • “/test-audit”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Each contract has
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no

What it can do on your machine

Read from SKILL.md and the folder at commit a035b4e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Audit loads about 1.9k tokens when it runs, and up to ~3.3k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,054 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/claude-codex-settings at commit a035b4e, republished under its MIT licence (© fcakyon). 1,054 words, ~1,924 tokens.

Download SKILL.mdSave it as .claude/skills/test-audit/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
test-audit
description
This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or sweeping tests. Gates new tests and audits low-value, implementation-coupled, or duplicative tests and the test-only production seams they demand.

Test Audit

Three modes, one value bar. Authoring mode gates every new or changed test at write time. Audit mode runs focused sweeps of tests that re-assert source, duplicate stronger proof, couple behavior to implementation, or keep test-only production seams alive. Continue broad audits as separate coherent follow-up PRs; optimize for confidence, not deletion count. Campaign mode prunes one whole subsystem's test surface (every test file a plugin or core area owns); before starting one, read references/campaign.md.

Authoring gate

Before adding any test, answer four questions; a missing answer means do not add it yet:

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Each contract has one primary test owner at the strongest boundary; another layer needs its own distinct risk, such as a transport or lifecycle failure the owner cannot reach. Prefer extending a table-driven case or shared fixture over a near-duplicate test; consolidate duplicated setup in the same change.
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no production caller needs? If yes, move the test to the real boundary instead.

Then check the test against every junk pattern; a match fails the gate unless the retention bar names the contract it independently guards. A test that would break under behavior-preserving refactoring is asserting implementation, not behavior; rewrite it at the owning boundary before landing it.

Bug regression tests must fail on the pre-fix code for the intended reason and pass after the owner-boundary repair. A regression test that never demonstrably failed proves the mock, not the fix. One regression at the owner boundary covers the bug; do not replay the same scenario at every layer it crosses.

Junk patterns

The shared checklist for both modes: the authoring gate rejects a new test that matches one, and audits hunt for existing tests that do.

  • assertion-free coverage probes;
  • self-comparisons and identity copiers;
  • copied fixtures, inventories, manifests, or export lists;
  • exact source, import, or string greps;
  • private predicate or call-shape tests duplicated at real boundaries;
  • duplicate invocations of the same contract;
  • provider-local replays of shared helpers;
  • tests whose only purpose is preserving test-only exports, globals, or wrappers;
  • dead production code whose only callers are tests;
  • expected values produced by the helper or renderer under test;
  • mocks that implement the asserted behavior, or one identical mock standing in for different APIs;
  • fixtures that supply the receipt, admission, or callback ordering the owner should produce, or persistence asserted against a store the path never writes;
  • capability tests that restate declared flags instead of exercising the delivery or acknowledgement the flag promises;
  • negative controls that pass for an unrelated reason, such as a denial from a different guard or a rejection the production path never reaches;
  • names or fixtures that promise more than the input exercises, such as a "retires the window" test asserting the window was not cleared.

Value bar

Tests justify their maintenance cost by protecting behavior, a credible regression, or an independently meaningful contract. In an audit, an existing test that must change for behavior-preserving source reorganization is suspect, not automatically deletable; the authoring gate still rejects new ones.

Before judging a candidate, read the complete test and production owner, its entry point, callers, callees, sibling implementations, overlapping tests, CI routing, and relevant history. Read root and scoped AGENTS.md files first. When the test claims dependency-backed behavior, inspect the dependency source or types directly.

Discovery

Keep discovery read-only and report evidence before editing. For broad scope, run parallel discovery lanes when available:

  • core source and packages;
  • plugins or extensions;
  • UI, apps, scripts, and tooling;
  • a cross-cutting pattern sweep.

Outside campaign mode, prefer a few high-confidence candidates over a large speculative inventory. Hunt for the junk patterns.

Show full SKILL.md (423 more words)Show less

Retention bar

Keep a test when it independently enforces a public API, plugin SDK, protocol, config, migration, storage, security, platform, default, prompt-byte, generated cross-language, package, release, or architecture contract. Also keep:

  • call ordering when order is observable behavior;
  • regressions with a credible failure mode;
  • source inspection when it is the cheapest independent guard: it fails when the contract changes (the user-facing key, byte, or path) and survives an identifier-only refactor;
  • a retained test that fails on the baseline: treat it as a possible product bug, reproduce it, and repair the owner rather than deleting it.

Static or slow is not a deletion reason. A test that resembles implementation may still be the independent contract; prove otherwise before removing it.

Candidate evidence

Record every field below before editing. A missing field means the candidate is not ready for deletion:

  • exact test name and location;
  • what failure it can actually detect;
  • non-test callers of the covered production or support seam;
  • stronger remaining owner-boundary proof, or why no proof is needed;
  • relevant history and the reason the test or seam exists;
  • production or test-support deletion unlocked;
  • risk and the focused validation command.

Edit shape

Choose one coherent owner-boundary batch. Delete obsolete test-only exports, globals, wrappers, and dead production paths instead of preserving aliases. Move retained regressions to their canonical owners. Consolidate repeated package or dependency assertions into one generic contract.

Prefer net-negative production LOC. Do not add replacement tests that restate the same implementation, and do not convert uncertain candidates into cleanup to increase deletion counts.

Validation

Never edit source or tests while a test watcher is running in the checkout. Follow the repository testing guide when one exists.

  1. Run the smallest owner and sibling tests with the project's test runner.
  2. For removed source greps or plan assertions, run the executable script or dry-run that owns the real contract.
  3. Run targeted formatting, then git diff --check.
  4. Run the changed-file or CI gate required by repository policy.
  5. Inspect git diff --numstat; report production/tooling separately from tests and test support.
  6. After final audit edits, run a code review of the full diff.

Landing and continuation

Commit, push, open a PR, or land only when authorized. Land one coherent PR at a time; after landing, refresh from current main and rerun read-only discovery for the next high-confidence batch.

Handoff

Report:

  • root cause and removed low-value categories;
  • production owner simplifications;
  • retained false positives and why they remain valuable;
  • focused and full proof actually run;
  • production versus test LOC;
  • PR and merge state;
  • named follow-ups.

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/test-audit/skills/test-audit of fcakyon/claude-codex-settings.

  • SKILL.md
  • LICENSE
  • references/campaign.md

Open the folder on GitHubat commit a035b4e

Compare with similar skills

Test Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Audit this skillfcakyon/claude-codex-settings1.2k—~1.9kAutomated safety check: PassMIT
Finishing a Development Branchobra/superpowers296k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k24 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Greplooponyx-dot-app/onyx32k4 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    296k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 24 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed
  • Guidelines

    akash-network/node

    Behavioral guidelines to reduce common LLM coding mistakes. An agent skill from akash-network/node.

    1.1k GitHub starsUsed in 22 repos~577 tokens
    DevelopmentAuto-check passed

More from fcakyon/claude-codex-settings

All 43 skills in this repo
  • Slides

    fcakyon/claude-codex-settings

    Create and edit presentation slide decks (.pptx) with PptxGenJS, bundled layout helpers, and render/validation utilities.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Openobserve API

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "query OpenObserve", "create OpenObserve dashboard", "edit OpenObserve panel", "fetch OpenObserve logs", "run OpenObserve search", "list OpenObserve…

    1.2k GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Supabase JS

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "use supabase-js", "query Supabase database", "supabase auth", "supabase storage", "supabase realtime", "supabase edge functions", or works with the…

    1.2k GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed
  • Yolo Training

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but…

    1.2k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Dokploy Deploy

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "deploy with Dokploy", "use Dokploy Cloud", "manage self-hosted Dokploy", "deploy Docker Compose on Dokploy", "manage Dokploy databases", "configure…

    1.2k GitHub stars~891 tokensUpdated yesterday
    Auto-check passed
  • Openship Config

    fcakyon/claude-codex-settings

    This skill should be used when the user asks to "create openship.json", "configure an OpenShip deployment", "make a repo deployable on OpenShip", or fix "openship config validate" errors.

    1.2k GitHub stars~803 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Test Audit

What does Test Audit do?

This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or…. Test Audit is an agent skill from fcakyon/claude-codex-settings. This skill should be used when the user asks to "write a test", "review these tests", "audit tests", "prune low-value tests", or "find duplicate tests", and whenever writing, changing, reviewing, or sweeping tests.

When should I use Test Audit?

Test Audit fits situations like: asks to write a test; review these tests; prune low-value tests; find duplicate tests.

How do I install Test Audit in Claude Code?

Run `npx skills add fcakyon/claude-codex-settings --skill test-audit -a claude-code`. Or copy the skill folder (plugins/test-audit/skills/test-audit in fcakyon/claude-codex-settings) into .claude/skills/test-audit in your project. Claude Code loads it when a task matches its description.

How do I install Test Audit in Codex?

Run `npx skills add fcakyon/claude-codex-settings --skill test-audit -a codex`. Or copy the skill folder (plugins/test-audit/skills/test-audit in fcakyon/claude-codex-settings) into .agents/skills/test-audit in your project. Codex loads it when a task matches its description.

Can I use Test Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/claude-codex-settings --skill test-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-audit, .gemini/skills/test-audit, .github/skills/test-audit and .opencode/skills/test-audit in your project.

What does Test Audit need to run?

Going by SKILL.md and its folder, Test Audit needs the command-line tools its instructions call (git).

Does Test Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Audit use?

Test Audit is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Audit use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Test Audit?

Skills that share tags, products or a category with Test Audit: Finishing a Development Branch (obra/superpowers, 296k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars) and Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Audit?

fcakyon (a GitHub user) maintains it in fcakyon/claude-codex-settings, which has 1,163 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on October 6, 2026.

Source: fcakyon/claude-codex-settings on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.