Agent skill

Test Audit

by simstudioai in simstudioai/sim

Invoke whenever writing, changing, reviewing, or sweeping tests.

Apache-2.0Auto-check passedTesting & QA

Install Test Audit

skills CLI
$ npx skills add simstudioai/sim --skill test-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install simstudioai/sim test-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/simstudioai/sim.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-audit .claude/skills/test-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-audit
GitHub stars
30k
Token cost
~2k tokens
SKILL.md length
1,097 words
Files
1
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Invoke whenever writing, changing, reviewing, or sweeping tests.

  • Works in 4 steps: What observable behavior, invariant, or… → What credible regression makes it fail? → Why does existing coverage not already… → …
  • Testing & QA work in your project
  • SKILL.md covers The three rules, Authoring gate, Junk patterns and Retention bar, plus 4 more sections
  • Calls bun, git and rg

What it does

Test Audit is an agent skill from simstudioai/sim. Invoke whenever writing, changing, reviewing, or sweeping tests. Authoring gate for new tests, plus an audit workflow for low-value, implementation-coupled, or duplicative tests and the test-only production seams they demand.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. The repository describes itself as: Sim is the collaborative workspace to build, deploy, and monitor AI agents and workflows. Used by 100,000+ builders. The licence is Apache-2.0.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “/test-audit”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Type-check, next build,
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no production

What it can do on your machine

Read from SKILL.md and the folder at commit b1b084d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • git
    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Audit loads about 2k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from simstudioai/sim at commit b1b084d, republished under its Apache-2.0 licence (© simstudioai). 1,097 words, ~2,037 tokens.

Download SKILL.mdSave it as .claude/skills/test-audit/SKILL.md (or your agent's skills folder).
name
test-audit
description
Invoke whenever writing, changing, reviewing, or sweeping tests. Authoring gate for new tests, plus an audit workflow for low-value, implementation-coupled, or duplicative tests and the test-only production seams they demand.
argument-hint
[author | audit <path> | campaign <subsystem>]

Test Audit

Three modes, one value bar. Authoring gates every new or changed test at write time. Audit runs a focused sweep of existing tests that re-assert source, duplicate stronger proof, couple to implementation, or keep test-only production seams alive. Campaign prunes one subsystem's whole test surface in parallel lanes. Optimize for confidence, not deletion count — but a test that cannot name the bug it catches is cost, not coverage.

Read .claude/rules/sim-testing.md first: it defines the test layers, file naming, and the mechanics (global mocks, @sim/testing, performance rules).

The three rules

CLAUDE.md → Testing sets them: no unit tests after the code; prefer E2E at the real boundary, ending in a verifiable artifact; list failure modes before testing in isolation. A mode you cannot list is not a test you should write. This skill enforces them.

Authoring gate

Before adding any test, answer all four. A missing answer means do not add it.

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch that failure? Type-check, next build, bun run check:audits, and the integration/E2E suites are coverage too. Each contract has one primary owner at the strongest boundary; another layer needs its own distinct risk. Prefer extending an existing table-driven case over a near-duplicate test.
  4. Does it need a production seam (export, flag, wrapper, injection hook) that no production caller needs? If yes, test at the real boundary instead.

Then check it against every junk pattern below. A match fails the gate unless the retention bar names the contract it independently guards. A test that would break under behavior-preserving refactoring asserts implementation, not behavior.

Regression tests must fail on the pre-fix code for the intended reason. Revert each guard of the fix separately and watch the test named for that guard go red, then restore. A regression test that never demonstrably failed proves the mock, not the fix. One regression at the owner boundary covers the bug; do not replay it at every layer it crosses.

Junk patterns

  • assertion-free or toBeDefined()-only tests; "renders without crashing";
  • restating declarations: block/tool/trigger/provider config (subBlock ids, params, outputs, URL templates, header maps), constant tables, registries, enums, export lists — type-check and check:audits own these;
  • route/handler tests that mock every collaborator and assert toHaveBeenCalledWith on the mocks, or re-assert a mock's canned return;
  • mocks that implement the asserted behavior, or one mock standing in for different APIs;
  • Zod contract tests proving a schema accepts a valid object or rejects an obviously invalid one;
  • React tests of text, class names, aria presence, snapshots, "calls onClick";
  • hook tests asserting query keys or fetch URLs; store tests of trivial setters;
  • tests of test infrastructure (mocks, factories, builders testing themselves);
  • source-text or import greps (readFileSync(src) + toContain);
  • expected values produced by the helper under test;
  • duplicate invocations of the same contract, or provider-local replays of a shared helper;
  • fixtures that supply the ordering, receipt, or callback the owner should produce;
  • negative controls that pass for an unrelated reason (a different guard short-circuits first);
  • names or fixtures that promise more than the input exercises;
  • dead production code or exports whose only callers are tests.
  • a hand-rolled vi.mock factory for a module vitest.setup.ts or @sim/testing already mocks, or a local copy of a @sim/testing helper (bun run check:test-patterns fails on these).

Retention bar

Keep a test when it independently enforces one of:

  • security — authn/authz denial, tenant/workspace isolation, SSRF/URL validation, secret redaction, encryption, signature verification, path traversal, injection, rate limits;
  • money and data integrity — billing/usage math, metering, quotas, idempotency, migrations, persistence semantics, concurrency/locking/leases, outbox, retries;
  • executor semantics — DAG traversal, loops/parallels, conditions/routers, reference resolution, streaming, pause/resume, run-from-block, cancellation;
  • real algorithms with edge cases — chunkers, parsers, query builders, cron, diff/merge, pagination, encoding, date math, ranking;
  • cross-process wire contracts — realtime protocol, desktop bridge/IPC, CLI/SDK wire, provider webhooks, MCP — that type-check cannot see;
  • a regression with a credible repeat, shown red on the pre-fix code.

Also keep call ordering when order is observable, and a source inspection when it is the cheapest independent guard of a user-facing byte, key, or path. A retained test that fails on the baseline is a possible product bug: reproduce it and fix the owner rather than deleting it. Static or slow is not a deletion reason.

Show full SKILL.md (388 more words)Show less

Audit mode

Keep discovery read-only and report evidence before editing. Before judging a candidate, read the complete test and its production owner, callers, sibling implementations, overlapping tests, CI routing (CI discovers *.integration.ts by glob; .github/workflows/*.yml names a few scripts and files by path), and relevant history (git log --format='%h %s' -5 -- <file>).

Record for every deletion candidate: the test and location; the failure it can actually detect; non-test callers of the seam it covers; the stronger remaining proof (or why none is needed); the production or test-support code its deletion unlocks; and the focused validation command.

Edit shape. One coherent owner-boundary batch per PR. When pruning inside a file, also delete now-unused imports, mocks, fixtures, and helpers. Delete test-only exports and dead production paths instead of preserving aliases (rg -n '<name>' --glob '!**/*.test.*' must show no other reference, including string and dynamic-import references; never delete route files, registry entries, or generated files). Prefer net-negative production LOC. Do not add replacement tests that restate the same implementation.

Campaign mode

For a whole subsystem or the whole repo:

  1. Partition test files into lanes of ~150–350 files by owning directory, and list the protected set (every *.integration.ts, *.live.test.ts, __integration__/**, apps/desktop/e2e/**, and every path named in .github/workflows/*.yml).
  2. Give each lane its own git worktree and branch (git worktree add -b <branch> <path> <base>, then bun install --frozen-lockfile inside it). Lanes never share a checkout, never symlink node_modules, and never use git stash — the stash is shared across worktrees.
  3. Each lane commits once and writes a report: counts, categories removed with examples, notable keeps and why, production seams removed with grep evidence, and the commands it ran.
  4. Merge lane branches, then sweep orphaned shared test support (packages/testing/**, fixtures, helpers) that no remaining test imports.

Validation

Never edit source or tests while Vitest is running in the same checkout.

  1. Run the touched and sibling test files (.claude/rules/sim-testing.md → Running).
  2. If production code changed: bun run type-check in that workspace.
  3. bun run check:audits from the repo root (some audits list test files by path).
  4. bun run lint, then git diff --check.
  5. Report git diff --shortstat with production and test changes counted separately.

Handoff

Report the categories removed, production simplifications, retained false positives and why they stay, the validation actually run, production vs test LOC, and named follow-ups.

© simstudioai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/test-audit of simstudioai/sim.

Open the folder on GitHubat commit b1b084d

Compare with similar skills

Test Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Audit this skillsimstudioai/sim30k—~2kAutomated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k32 repos~2.1kAutomated safety check: PassApache-2.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 32 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    381 GitHub starsUsed in 9 repos~2.9k tokens
    Testing & QAAuto-check passed

More from simstudioai/sim

All 40 skills in this repo
  • Sim Helm

    simstudioai/sim

    Install, upgrade, and operate the Sim Helm chart on Kubernetes.

    30k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Add Column Type

    simstudioai/sim

    Add a new table column type to Sim — registry entry, icon, storage shape, coercion, and the behavioral hooks the grid and API read.

    30k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Add Enrichment

    simstudioai/sim

    Add a code-defined table enrichment (registry entry) under apps/sim/enrichments/ backed by an ordered provider cascade, ensuring every provider tool it calls has hosted-key support.

    30k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Add Hosted Key

    simstudioai/sim

    Add hosted API key support to a tool so Sim provides the key (metered and billed to the workspace) when a user has not brought their own.

    30k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Add Managed CLI

    simstudioai/sim

    Add or upgrade a curated, immutable managed CLI for Sim Function sandboxes, including client-safe catalog metadata, a pinned server-only installation recipe, checksum and executable verification…

    30k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Add Selector

    simstudioai/sim

    Add or update a Sim dynamic selector using the shared manifest, server attachment, and selectors.execute path.

    30k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Test Audit

What does Test Audit do?

Invoke whenever writing, changing, reviewing, or sweeping tests. Test Audit is an agent skill from simstudioai/sim. Invoke whenever writing, changing, reviewing, or sweeping tests.

When should I use Test Audit?

Test Audit fits situations like: testing & QA work in your project.

How do I install Test Audit in Claude Code?

Run `npx skills add simstudioai/sim --skill test-audit -a claude-code`. Or copy the skill folder (.agents/skills/test-audit in simstudioai/sim) into .claude/skills/test-audit in your project. Claude Code loads it when a task matches its description.

How do I install Test Audit in Codex?

Run `npx skills add simstudioai/sim --skill test-audit -a codex`. Or copy the skill folder (.agents/skills/test-audit in simstudioai/sim) into .agents/skills/test-audit in your project. Codex loads it when a task matches its description.

Can I use Test Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add simstudioai/sim --skill test-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-audit, .gemini/skills/test-audit, .github/skills/test-audit and .opencode/skills/test-audit in your project.

What does Test Audit need to run?

Going by SKILL.md and its folder, Test Audit needs the command-line tools its instructions call (bun, git and rg).

Does Test Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Audit use?

Test Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Audit use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Audit?

Skills that share tags, products or a category with Test Audit: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Audit?

simstudioai (a GitHub organization) maintains it in simstudioai/sim, which has 29,792 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on October 9, 2026.

Source: simstudioai/sim on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.