Agent skill

Auditing Tests

by purefunctor in purefunctor/purescript-iris

Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams.

MITAuto-check passedTesting & QA

Install Auditing Tests

skills CLI
$ npx skills add purefunctor/purescript-iris --skill auditing-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install purefunctor/purescript-iris auditing-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/purefunctor/purescript-iris.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/auditing-tests .claude/skills/auditing-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
auditing-tests
GitHub stars
117
Token cost
~2.7k tokens
SKILL.md length
1,392 words
Files
3
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams.

  • Works in 4 steps: What observable behavior, invariant, or… → What credible regression makes it fail… → Why does existing coverage not already… → …
  • Reviewing tests
  • SKILL.md covers Authoring gate, Choose the Iris owner, Junk patterns and Value and retention bar, plus 4 more sections
  • Calls just, cargo and git

What it does

Auditing Tests is an agent skill from purefunctor/purescript-iris. Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams. Use when writing or reviewing tests, sweeping a compiler subsystem, or pruning fixtures and test support in purescript-iris.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `CAMPAIGN.md`).

It sits in Testing & QA. The repository describes itself as: A compiler for the PureScript programming language. The licence is MIT.

When your agent uses it

  • Reviewing tests
  • Sweeping a compiler subsystem
  • Pruning fixtures and test support in purescript-iris

Example prompts

  • “Use the auditing-tests skill to gate new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only…”
  • “/auditing-tests”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail or changes its observed report?
  3. Why does existing coverage not already catch that failure? Each contract has
  4. Does it need an export, flag, wrapper, or injection hook no production caller

What it can do on your machine

Read from SKILL.md and the folder at commit e5bd7b2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • just
    • cargo
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Auditing Tests loads about 2.7k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,392 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from purefunctor/purescript-iris at commit e5bd7b2, republished under its MIT licence (© purefunctor). 1,392 words, ~2,739 tokens.

Download SKILL.mdSave it as .claude/skills/auditing-tests/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
auditing-tests
description
Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams. Use when writing or reviewing tests, sweeping a compiler subsystem, or pruning fixtures and test support in purescript-iris.
license
MIT

Auditing Iris Tests

Three modes, one value bar. Authoring gates new or changed tests at write time. Audit investigates a few high-confidence candidates for consolidation or deletion. Campaign reviews one subsystem's whole test surface; read CAMPAIGN.md before starting one. Optimize for confidence, not deletion counts or coverage percentages.

Adapted from OpenClaw's test-audit skill. The upstream license is preserved in LICENSE.

Read root and scoped AGENTS.md files before working. This skill adds a test value gate, not a replacement for Iris's ownership rules or verification policy. Use the existing workflow-integration-tests, workflow-regression-tests, and running-compatibility-checks skills for their respective workflows.

Authoring gate

Before adding or changing a test, answer four questions:

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes it fail or changes its observed report?
  3. Why does existing coverage not already catch that failure? Each contract has one primary owner at the strongest useful boundary. Another layer needs a distinct risk, such as a local algorithm invariant versus its integration into compilation, or a process lifecycle failure compiler APIs cannot reach. Prefer extending a fixture or table over duplicating the same scenario.
  4. Does it need an export, flag, wrapper, or injection hook no production caller needs? If so, test through the real owning boundary instead.

A missing answer means the test is not ready. Check it against the junk patterns below; a match fails the gate unless the retention bar identifies an independent contract. Behavior-preserving internal refactoring should not break a behavioral test. Explicit representation or output contracts may legitimately constrain refactoring; name those contracts instead of treating all snapshots as junk.

For a bug regression, demonstrate the undesirable behavior before the repair and the intended behavior afterward at the owning boundary. Follow workflow-regression-tests for Iris's failing-fixture/fix history. A deliberately accepted baseline snapshot may pass while recording a bug: the decisive evidence is the reviewed before/after report, not merely the runner's exit status. Runtime regressions must exercise fresh generated code and fail for the intended reason on the buggy compiler. Do not replay one bug at every layer it crosses.

Choose the Iris owner

OwnerContract and suitable proof
Subsystem unit tests beside Rust codeSmall algorithms and data structures using constructed local data, such as functional-dependency closure or pattern-matrix operations
tests-integrationSource-file behavior through compiler APIs and the existing fixture harness; do not build another compiler pipeline in a unit test
compiler fixturesChecked types/kinds, diagnostics, semantic recovery, functional conversion, generated JavaScript, and optional runtime verification
lowering, resolving, lsp fixturesThe corresponding lowered/source-link, name/import/export, or editor-analysis contract
tests-e2eReal CLI, Spago, filesystem, shell/process, watch/run, and development-environment behavior through temporary workspaces
tests-compatibilityReal package-set compatibility and benchmarks; extract a focused compiler regression into tests-integration
tests-supportRegistry preparation, resolution, digest verification, extraction, and cache-publication invariants, not compiler semantics

Compiler fixtures enter through Main.purs. Their checking, diagnostics, semantic, and functional reports protect different observations; sharing one input does not make those reports redundant. Keep generated goldens limited to reachable fixture-owned modules. Use verify.mjs only when execution is the contract; it must test fresh output, never tracked goldens. Use real registry modules and the existing replacements.json mechanism for deliberate substitutes.

Junk patterns

  • Assertion-free coverage probes, self-comparisons, and identity copiers.
  • Expected values computed by the helper, renderer, or compiler under test.
  • Copied inventories, manifests, export lists, or exact source/import greps without an independent contract.
  • Private predicate or call-shape tests duplicated by stronger boundary proof.
  • Multiple invocations of the same contract without distinct failure modes.
  • Source-compilation unit tests duplicating the integration fixture pipeline.
  • Handwritten library stand-ins where real prepared registry modules belong.
  • Mocks that implement the behavior being asserted or stand in identically for APIs with different contracts.
  • Fixtures that supply the ordering, persistence, or evidence the owner should produce, or assert a store the exercised path never writes.
  • Capability tests that restate declared flags rather than exercising delivery.
  • Tests that exist only to preserve test-only exports, globals, or wrappers; production code whose only callers are tests.
  • Negative controls that pass because of an unrelated parser, resolver, checker, or environment failure before reaching the intended contract.
  • Names promising more than the inputs and assertions exercise.
  • Snapshot acceptance without checking types, diagnostics, locations, recovery, generated code, or runtime results against the intended semantics.

Value and retention bar

Keep tests that independently protect public APIs, language semantics, compiler representations with an intentional contract, LSP payloads, configuration, storage, security, platforms, defaults, generated code, packages, releases, or architecture. Also retain observable call ordering, credible regressions, and source inspection when it is the cheapest independent guard and survives an identifier-only refactor.

Static, slow, snapshot-based, or implementation-adjacent is not a deletion reason. A test that resembles implementation may still be its independent contract. Prove redundancy before removing it. Passing snapshots do not prove semantic correctness. A retained baseline failure may be a product defect; reproduce it instead of deleting the evidence or accepting it away.

Show full SKILL.md (580 more words)Show less

Read-only discovery and candidate evidence

Keep discovery read-only and report evidence before editing. For each candidate, read the complete test or fixture, production owner, entry points, callers, callees, sibling implementations, overlapping coverage, relevant history, and CI routing. Inspect dependency source or types when a claim depends on them. Consult .github/workflows/checks.yml, .github/workflows/platform-tests.yml, and .buildkite/ as relevant; do not infer CI coverage from crate membership.

Record every field before deleting or consolidating a candidate:

  • Exact test declaration or fixture/report and its location.
  • Failure it can actually detect, not merely its name or intended purpose.
  • Non-test callers of its production or support seam.
  • Stronger remaining owner-boundary proof, or why no contract needs proof.
  • Relevant history and why the test or seam exists.
  • Production or test-support deletion unlocked.
  • Risk and focused validation command.

A missing field means the candidate is not ready. Prefer a few well-supported candidates to a speculative inventory. For broad discovery, split lanes by production owner and give workers disjoint scope when delegation is available and useful; a small audit needs no delegation ceremony.

Edit shape

Choose one coherent owner-boundary batch. Move retained contracts into their canonical owners before removing weaker proof. Delete obsolete test-only exports, globals, wrappers, and dead paths rather than preserving aliases. Consolidate repeated assertions at a shared owner when their risks are identical.

Prefer simpler production code, not a target LOC reduction. Do not add replacement tests that restate implementation, weaken assertions, or turn uncertain candidates into cleanup to inflate deletion counts. Leave unrelated bugs alone; report them as follow-ups unless their repair is authorized.

Validation

Do not edit source, fixtures, or expectations during a test run in the checkout. Choose checks by affected behavior and preserve root verification gates:

  1. Check changed Rust crates with cargo check -p <crate-name> --tests and run their unit tests with cargo nextest run -p <crate-name>.
  2. Iterate on fixtures with just t <category> <filters>. Before pushing a change affecting integration tests, run every affected category without filters and confirm it passes with no pending snapshots.
  3. Never edit .snap files or JavaScript goldens by hand. Review diffs with just t <category> <filters> --diff, accept intended snapshots with just t <category> <filters> --accept, and regenerate JavaScript through just t compiler <filters> --update-output. Inspect every expectation.
  4. For CLI/environment changes, run just e2e-prepare and the relevant just e2e tests; preserve the shared harness's cross-platform assertions.
  5. For compatibility impact, follow running-compatibility-checks and use just compatibility <base-ref> when justified. Do not substitute package corpus results for focused regression proof or confuse them with benchmarks.
  6. For removed source greps or plan assertions, execute the script or operation owning the real contract. Run just format for Rust changes and git diff --check. Before a PR push, run just format and just licenses and fold resulting changes into their relevant commits as repository policy requires.
  7. Review the final diff for lost contracts and accidental snapshot acceptance. Use git diff --numstat to report production/tooling separately from tests, fixtures, goldens, and support. State which checks actually ran and any limits.

Landing and handoff

Commit, push, open a PR, or merge only as authorized. Follow the root commit and PR conventions; do not import another project's review bots or landing scripts. Keep one coherent audit batch reviewable. After landing, refresh the baseline before starting another batch.

Report removed low-value categories, owner simplifications, retained false positives and their contracts, proof actually run, production versus test/support LOC, delivery state, and named follow-ups. Campaigns also use the handoff in CAMPAIGN.md.

© purefunctor, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .agents/skills/auditing-tests of purefunctor/purescript-iris.

  • SKILL.md
  • CAMPAIGN.md
  • LICENSE

Open the folder on GitHubat commit e5bd7b2

Compare with similar skills

Auditing Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Auditing Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Auditing Tests this skillpurefunctor/purescript-iris117—~2.7kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k31 repos~2.1kAutomated safety check: PassApache-2.0
TDDfossasia/eventyay-interpretation1.6k28 repos~1.1kAutomated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 31 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    fossasia/eventyay-interpretation

    Test-driven development. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 28 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    380 GitHub starsUsed in 8 repos~2.9k tokens
    Testing & QAAuto-check passed

More from purefunctor/purescript-iris

  • Running Compatibility Checks

    purefunctor/purescript-iris

    Runs Iris package-set compatibility comparisons with release-built verifiers.

    117 GitHub stars~605 tokensUpdated today
    Auto-check passed
  • Cutting Releases

    purefunctor/purescript-iris

    Cuts Iris GitHub releases through the version-bump PR, merge commit, tag-driven build workflow, attestations, installer tests, and generated release notes.

    117 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Watch

    purefunctor/purescript-iris

    Ask a running iris watch about a PureScript project with iris watch query for signatures, module exports, definitions, references, instances, dependent modules, name search, diagnostics, and…

    117 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Workflow Integration Tests

    purefunctor/purescript-iris

    Workflow for adding and updating Iris integration-test fixtures for unified compiler, lowering, resolving, and LSP behavior.

    117 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Workflow Regression Tests

    purefunctor/purescript-iris

    Workflow for producing auditable Git or jj history for a known compiler bug fix.

    117 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Writing Code Commentary

    purefunctor/purescript-iris

    Writes and reviews Iris compiler comments, algorithm traces, and documentation examples.

    117 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Auditing Tests

What does Auditing Tests do?

Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams. Auditing Tests is an agent skill from purefunctor/purescript-iris. Gates new or changed Iris tests and audits low-value, implementation-coupled, or duplicative coverage and test-only production seams.

When should I use Auditing Tests?

Auditing Tests fits situations like: reviewing tests; sweeping a compiler subsystem; pruning fixtures and test support in purescript-iris.

How do I install Auditing Tests in Claude Code?

Run `npx skills add purefunctor/purescript-iris --skill auditing-tests -a claude-code`. Or copy the skill folder (.agents/skills/auditing-tests in purefunctor/purescript-iris) into .claude/skills/auditing-tests in your project. Claude Code loads it when a task matches its description.

How do I install Auditing Tests in Codex?

Run `npx skills add purefunctor/purescript-iris --skill auditing-tests -a codex`. Or copy the skill folder (.agents/skills/auditing-tests in purefunctor/purescript-iris) into .agents/skills/auditing-tests in your project. Codex loads it when a task matches its description.

Can I use Auditing Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add purefunctor/purescript-iris --skill auditing-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/auditing-tests, .gemini/skills/auditing-tests, .github/skills/auditing-tests and .opencode/skills/auditing-tests in your project.

What does Auditing Tests need to run?

Going by SKILL.md and its folder, Auditing Tests needs the command-line tools its instructions call (just, cargo and git).

Does Auditing Tests access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Auditing Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Auditing Tests use?

Auditing Tests is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Auditing Tests use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Auditing Tests?

Skills that share tags, products or a category with Auditing Tests: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (fossasia/eventyay-interpretation, 1.6k stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Auditing Tests?

purefunctor (a GitHub user) maintains it in purefunctor/purescript-iris, which has 117 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: purefunctor/purescript-iris on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.