Agent skill

Test Audit

by get-bb in get-bb/bb

Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive.

MITAuto-check passed

Install Test Audit

skills CLI
$ npx skills add get-bb/bb --skill test-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install get-bb/bb test-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/get-bb/bb.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.bb/skills/test-audit .claude/skills/test-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-audit
GitHub stars
4.2k
Token cost
~3.1k tokens
SKILL.md length
1,580 words
Files
3
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive.

  • Works in 4 steps: What observable behavior, invariant, or… → What credible regression makes it fail? → Why does existing coverage not already… → …
  • SKILL.md covers Authoring gate, Junk patterns, Value bar and Discovery, plus 6 more sections
  • Calls pnpm, git and vitest

What it does

Test Audit is an agent skill from get-bb/bb. Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive. Use when writing, changing, reviewing, or sweeping tests.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `CAMPAIGN.md`).

The repository describes itself as: The agent IDE that builds itself. The licence is MIT.

Example prompts

  • “/test-audit”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Each contract has
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no

What it can do on your machine

Read from SKILL.md and the folder at commit 68a1e8b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • git
    • vitest
    • turbo
    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Audit loads about 3.1k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,580 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from get-bb/bb at commit 68a1e8b, republished under its MIT licence (© get-bb). 1,580 words, ~3,061 tokens.

Download SKILL.mdSave it as .claude/skills/test-audit/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
test-audit
description
Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive. Use when writing, changing, reviewing, or sweeping tests.

Test Audit

Three modes, one value bar. Authoring mode gates every new or changed test at write time. Audit mode runs focused sweeps of tests that re-assert source, duplicate stronger proof, couple behavior to implementation, or keep test-only production seams alive. Continue broad audits as separate coherent follow-up PRs; optimize for confidence, not deletion count. Campaign mode prunes one whole subsystem's test surface (every test file a plugin, app, or package owns); before starting one, read CAMPAIGN.md.

Authoring gate

Before adding any test, answer four questions; a missing answer means do not add it yet:

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Each contract has one primary test owner at the strongest boundary; another layer needs its own distinct risk, such as a transport or lifecycle failure the owner cannot reach. Prefer extending a table-driven case or shared fixture over a near-duplicate test; consolidate duplicated setup in the same change.
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no production caller needs? If yes, move the test to the real boundary instead.

Then check the test against every junk pattern; a match fails the gate unless the retention bar names the contract it independently guards. A test that would break under behavior-preserving refactoring is asserting implementation, not behavior; rewrite it at the owning boundary before landing it.

Bug regression tests must fail on the pre-fix code for the intended reason and pass after the owner-boundary repair. A regression test that never demonstrably failed proves the mock, not the fix. One regression at the owner boundary covers the bug; do not replay the same scenario at every layer it crosses.

bb test rules: never mock the database; use createConnection(":memory:") and migrate(db) from @bb/db. Vitest projects come from sharedWorkerProjects in vitest.shared.ts: Node test files share workers (isolate: false), and files that use vi.mock, vi.stubGlobal, fake timers, process.env writes, or global assignments are isolated automatically. Restore any other global state a test changes. Tests follow the no-comments rule.

Junk patterns

The shared checklist for both modes: the authoring gate rejects a new test that matches one, and audits hunt for existing tests that do.

  • assertion-free coverage probes;
  • self-comparisons and identity copiers;
  • copied fixtures, inventories, manifests, or export lists;
  • exact source, import, or string greps;
  • private predicate or call-shape tests duplicated at real boundaries;
  • duplicate invocations of the same contract;
  • provider-local replays of shared helpers;
  • tests whose only purpose is preserving test-only exports, globals, or wrappers;
  • dead production code whose only callers are tests;
  • expected values produced by the helper or renderer under test;
  • mocks that implement the asserted behavior, or one identical mock standing in for different APIs;
  • fakes that skip the real code they front, such as a hand-rolled bb stub or fake provider that never loads the real plugin entry, bridge, or extension;
  • fixtures that supply the receipt, admission, or callback ordering the owner should produce, or persistence asserted against a store the path never writes;
  • capability tests that restate declared flags instead of exercising the delivery or acknowledgement the flag promises;
  • negative controls that pass for an unrelated reason, such as a denial from a different guard or a rejection the production path never reaches;
  • names or fixtures that promise more than the input exercises, such as a "retires the window" test asserting the window was not cleared.

Value bar

Tests justify their maintenance cost by protecting behavior, a credible regression, or an independently meaningful contract. In an audit, an existing test that must change for behavior-preserving source reorganization is suspect, not automatically deletable; the authoring gate still rejects new ones.

Before judging a candidate, read the complete test and production owner, its entry point, callers, callees, sibling implementations, overlapping tests, CI routing (turbo.json task inputs and .github/workflows/ci.yml shards), and relevant history. Read AGENTS.md first. When the test claims dependency-backed behavior, inspect the dependency source or types directly.

Discovery

Keep discovery read-only and report evidence before editing. For broad scope, run parallel discovery lanes when available:

  • app UI: apps/app (the largest suite), apps/web, apps/mobile, apps/desktop, packages/thread-view, packages/client-core;
  • server and data: apps/server, packages/db, packages/domain;
  • host and runtime: apps/host-daemon, packages/agent-runtime, packages/provider-bridge-*, packages/host-*;
  • plugins and SDK: plugins/ (built-in and provider-* plugins), packages/plugin-sdk, packages/plugin-build;
  • CLI, tooling, and end to end: apps/cli, packages/sdk, packages/scripts, packages/bb-app, tests/integration;
  • a cross-cutting pattern sweep.

Outside campaign mode, prefer a few high-confidence candidates over a large speculative inventory. Hunt for the junk patterns.

Retention bar

Keep a test when it independently enforces a public API, Plugin SDK, protocol, config, migration, storage, security, platform, default, prompt-byte, generated cross-language, package, release, or architecture contract. In bb that includes server/daemon wire shapes behind HOST_DAEMON_PROTOCOL_VERSION (packages/host-daemon-contract/src/protocol.ts) and Plugin Guide surfaces and anchors (plugins/plugin-api-docs/src/surfaces.ts, anatomy-manifest.json). Also keep:

  • call ordering when order is observable behavior;
  • regressions with a credible failure mode;
  • source inspection when it is the cheapest independent guard: it fails when the contract changes (the user-facing key, byte, or path) and survives an identifier-only refactor;
  • a retained test that fails on the baseline: treat it as a possible product bug, reproduce it, and repair the owner rather than deleting it.

Static or slow is not a deletion reason. A test that resembles implementation may still be the independent contract; prove otherwise before removing it.

An unprefixed export of a published @get-bb/plugin-sdk subpath is public API, not a test-only seam, even with zero in-repo consumers. Removing one needs an entry under "Scheduled removals (next major)" in docs/api_to_audit.md.

Provider event shapes and timeline rows are owned by the recorded parity conformance (packages/provider-parity, recordings in packages/provider-bridge-protocol/recordings/) and the provider corpus gates (apps/server/test/provider-corpus/). A plugin test that replays a recorded case is a candidate; one covering a case recordings cannot capture is not.

Show full SKILL.md (619 more words)Show less

Candidate evidence

Record every field below before editing. A missing field means the candidate is not ready for deletion:

  • exact test name and location;
  • what failure it can actually detect;
  • non-test callers of the covered production or support seam;
  • stronger remaining owner-boundary proof, or why no proof is needed;
  • relevant history and the reason the test or seam exists;
  • production or test-support deletion unlocked;
  • risk and the focused validation command.

Edit shape

Choose one coherent owner-boundary batch. Delete obsolete test-only exports, globals, wrappers, and dead production paths instead of preserving aliases. Move retained regressions to their canonical owners. Consolidate repeated package or dependency assertions into one generic contract.

Prefer net-negative production LOC. Do not add replacement tests that restate the same implementation, and do not convert uncertain candidates into cleanup to increase deletion counts. Generated modules (packages/templates/src/generated/, packages/plugin-build/src/generated/, packages/plugin-sdk/bundled-types/) are gitignored; never commit them or add a --check mode to guard them.

Validation

Never edit source or tests while Vitest is running in the checkout. Never run git checkout <rev> -- . for a baseline in a checkout with uncommitted edits; it overwrites them. Commit a WIP first, then compare through git worktree add --detach <dir> <rev> or git show <rev>:<path>.

  1. Run the smallest owner and sibling tests through Turbo. Arguments after -- go to vitest run; paths are relative to the package directory. Plugin packages are named bb-plugin-<id>. Pipe slow output to a file: pnpm exec turbo run test --filter=@bb/server -- test/<file>.test.ts -t "<name>" > /tmp/test.log 2>&1.
  2. turbo run test skips some suites; when a touched test or its owner belongs to one, run it too (list them with rg -n '^\s*"(@bb/[a-z-]+#)?(test|smoke)[a-z:_-]*"\s*:' turbo.json). Suites that need credentials must still pass setup and discovery; see docs/debugging-and-qa.md for corpus and parity inputs.
    • packages/agent-runtime/src/integration*.test.ts: pnpm exec turbo run test:integration --filter=@bb/agent-runtime;
    • tests/integration/real/: pnpm exec turbo run test:integration --filter=@bb/integration-tests;
    • apps/server/test/provider-corpus/ (skips without BB_PROVIDER_CORPUS_DIR): pnpm exec turbo run test:provider-corpus --filter=@bb/server;
    • provider bridges: pnpm parity --old <before-worktree> --new .;
    • packaging: pnpm exec turbo run smoke:tarball --filter=bb-app, and the @bb/desktop smoke:* tasks.
  3. For removed source greps or plan assertions, run the executable script or dry-run that owns the real contract.
  4. Run pnpm exec oxfmt <changed-paths>, then pnpm exec turbo run lint typecheck --filter=<pkg>, then git diff --check.
  5. Rule out environment-only failures before calling a baseline failure a product bug: apps/server/test/internal/internal-skill-trees.test.ts fails under umask 0002, and apps/host-daemon/src/command-discovery.test.ts and plugins/provider-acp/src/native-roots/native-roots.test.ts fail when /tmp/.agents or ~/.agents exists. Turbo drops TMPDIR, so after a Turbo run has built dependencies, rerun those with TMPDIR=/var/tmp/<dir> pnpm --filter <pkg> exec vitest run --config vitest.config.ts <file>.
  6. For every removed or merged test, prove no coverage was lost: write the bug it caught as a mutation of the production owner, confirm it fails the old test in a before-worktree and at least one keeper in the after-worktree, then restore the source byte for byte. A past cleanup lost 19 real assertions despite the author's own mutation checks, so have an independent agent run this. Use --concurrency=1 when the machine is loaded.
  7. Inspect git diff --numstat; report production/tooling separately from tests and test support.
  8. After final audit edits, review the diff against docs/CODE_REVIEW.md (for example with /code-review) and run the deslop skill.

Landing and continuation

Commit, push, open a PR, or land only when authorized. Fill .github/PULL_REQUEST_TEMPLATE.md and end the body with > AGENT GENERATED. Use the close-out skill to land when it is available. Land one coherent PR at a time; after landing, refresh from current main and rerun read-only discovery for the next high-confidence batch.

Handoff

Report:

  • root cause and removed low-value categories;
  • production owner simplifications;
  • retained false positives and why they remain valuable;
  • focused and full proof actually run, including off-pipeline suites;
  • production versus test LOC;
  • PR and merge state;
  • named follow-ups.

© get-bb, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .bb/skills/test-audit of get-bb/bb.

  • SKILL.md
  • CAMPAIGN.md
  • LICENSE

Open the folder on GitHubat commit 68a1e8b

Compare with similar skills

Test Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Audit this skillget-bb/bb4.2k—~3.1kAutomated safety check: PassMIT
Orch Change Featureaffaan-m/ECC274k1 repos~420Automated safety check: PassMIT
PR Feedback Quality Gatenexu-io/open-design100k—~572Automated safety check: PassApache-2.0
Change Verification Gatefengshao1227/ccg-workflow5.9k—~511Automated safety check: NotesMIT
Gateplugin87/ux-ui-agent-skills1.5k—~532Automated safety check: PassMIT
Make Changesremix-run/remix33k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Orchestrate altering an existing, working feature to new desired behavior — update its tests to the new spec, change the implementation to match, review, and gated commit.

    274k GitHub starsUsed in 1 repo~420 tokens
    Auto-check passed
  • PR Feedback Quality Gate

    nexu-io/open-design

    Safely track pull request feedback, resolve review comments or merge conflicts, validate fixes, and use a read-only cross-review before committing or pushing follow-up changes.

    100k GitHub stars~572 tokensUpdated today
    DevelopmentAuto-check passed
  • Change Verification Gate

    fengshao1227/ccg-workflow

    Analyzes a code diff for documentation sync, test coverage and impact scope, warning when docs or tests lag behind a design-level change or a large edit.

    5.9k GitHub stars~511 tokensUpdated 22 days ago
    DevelopmentAuto-check: notes
  • Gate

    plugin87/ux-ui-agent-skills

    Run the one-command quality gate and report the real N/N result.

    1.5k GitHub stars~532 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Make Changes

    remix-run/remix

    Create or update Remix repo change files under packages//.changes.

    33k GitHub stars~2.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Brain Ingest Gate

    garrytan/gbrain

    Pre-write quality gate for content entering the brain. An agent skill from garrytan/gbrain.

    31k GitHub stars~3.9k tokensUpdated today
    Testing & QAAuto-check passed

More from get-bb/bb

All 33 skills in this repo
  • Bb CLI

    get-bb/bb

    Inspect or manage BB state with the bb CLI; use for BB commands and configuration.

    4.2k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Skill Creator

    get-bb/bb

    Create or improve BB skills, including their triggers, instructions, and supporting resources.

    4.2k GitHub stars~702 tokensUpdated today
    Auto-check passed
  • Prepare and submit a BB plugin to the Community marketplace when publication or a marketplace PR is requested.

    4.2k GitHub stars~873 tokensUpdated today
    Auto-check passed
  • Workflows

    get-bb/bb

    Author or run durable BB workflows when the user requests workflow execution or multi-agent orchestration.

    4.2k GitHub stars~890 tokensUpdated today
    Auto-check passed
  • Verify Bb

    get-bb/bb

    Verify BB user journeys in an isolated source dev app using dev-browser@next and the matching source CLI.

    4.2k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Create or edit BB color themes and inspect them in the Theme Preview panel.

    4.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Test Audit

What does Test Audit do?

Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive. Test Audit is an agent skill from get-bb/bb. Gate new or changed tests, and audit existing tests for low value, implementation coupling, duplication, and the test-only production seams they keep alive.

How do I install Test Audit in Claude Code?

Run `npx skills add get-bb/bb --skill test-audit -a claude-code`. Or copy the skill folder (.bb/skills/test-audit in get-bb/bb) into .claude/skills/test-audit in your project. Claude Code loads it when a task matches its description.

How do I install Test Audit in Codex?

Run `npx skills add get-bb/bb --skill test-audit -a codex`. Or copy the skill folder (.bb/skills/test-audit in get-bb/bb) into .agents/skills/test-audit in your project. Codex loads it when a task matches its description.

Can I use Test Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add get-bb/bb --skill test-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-audit, .gemini/skills/test-audit, .github/skills/test-audit and .opencode/skills/test-audit in your project.

What does Test Audit need to run?

Going by SKILL.md and its folder, Test Audit needs the command-line tools its instructions call (pnpm, git, vitest, turbo and rg).

Does Test Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Audit use?

Test Audit is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Audit use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Audit?

Skills that share tags, products or a category with Test Audit: Orch Change Feature (affaan-m/ECC, 274k stars), PR Feedback Quality Gate (nexu-io/open-design, 100k stars), Change Verification Gate (fengshao1227/ccg-workflow, 5.9k stars) and Gate (plugin87/ux-ui-agent-skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Audit?

get-bb (a GitHub organization) maintains it in get-bb/bb, which has 4,155 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 7, 2026.

Source: get-bb/bb on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.