Agent skill

Test Suite Curation

by petrkindlmann in petrkindlmann/qa-skills

Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…

MITAuto-check passedTesting & QA

Install Test Suite Curation

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-suite-curation .claude/skills/test-suite-curation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-suite-curation
GitHub stars
168
Token cost
~6.3k tokens
SKILL.md length
3,348 words
Files
8 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…

  • Works in 8 steps: Coverage Fingerprinting (per-test) → The Coverage-Equality Trap (the… → Near-Duplicate Clustering → …
  • : audit the test suite
  • SKILL.md covers Quick Route, Discovery Questions, Core Principles and 1. Coverage Fingerprinting…, plus 12 more sections
  • Calls pytest, sqlite3 and git

What it does

Test Suite Curation is an agent skill from petrkindlmann/qa-skills. Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune decision rules (redundant/obsolete/low-value/keep), smoke/core/extended tiering by risk and defect-detection history, and a defensible "what we deleted and why" record. Deletion is destructive — quarantine and human sign-off are mandatory. Use when: "audit the test suite," "prune redundant tests," "find duplicate tests,"…

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/audit-record.md`, `references/ci-history-mining.md` and `references/clustering.md`).

It sits in Testing & QA, covering Test generation, Failing and flaky tests and Test coverage. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • : audit the test suite
  • Prune redundant tests
  • Find duplicate tests
  • Which tests can we delete

Example prompts

  • “what we deleted and why”
  • “audit the test suite,”
  • “prune redundant tests,”
  • “/test-suite-curation”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Coverage Fingerprinting (per-test)
  2. The Coverage-Equality Trap (the load-bearing rule)
  3. Near-Duplicate Clustering
  4. Mining CI History
  5. Prune Decision Rules
  6. Tiering (smoke / core / extended)
  7. Destructive Safety (the grace period)
  8. The Audit Record

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • sqlite3
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Suite Curation loads about 6.3k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 255 tokens; SKILL.md has 3,348 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~255
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 3,348 words, ~6,288 tokens.

Download SKILL.mdSave it as .claude/skills/test-suite-curation/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
test-suite-curation
description
Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune decision rules (redundant/obsolete/low-value/keep), smoke/core/extended tiering by risk and defect-detection history, and a defensible "what we deleted and why" record. Deletion is destructive — quarantine and human sign-off are mandatory. Use when: "audit the test suite," "prune redundant tests," "find duplicate tests," "which tests can we delete," "restructure into smoke/core/extended," "is this test pulling its weight," "shrink the regression suite." Not for: Judging whether an individual test is WELL-WRITTEN (smells, assertions) — that is ai-qa-review. Healing one flaky test at runtime — that is test-reliability. Bulk selector regeneration after a UI refactor — that is selector-drift-recovery. Related: ai-qa-review, coverage-analysis, test-reliability, risk-based-testing, qa-project-context.
license
MIT
metadata.author
kindlmann
metadata.version
1.0
metadata.category
process
<objective>
Test A and Test B cover the exact same lines, so a tired engineer deletes B — and three weeks later a production defect slips through because B was the only test that asserted the rounding was correct. Coverage equality is not redundancy. This skill audits an entire regression suite as a corpus (the redundancy analysis no human does by hand), prunes it on evidence rather than vibes, and treats every deletion as a destructive change that requires a quarantine grace period, human sign-off, and a record you can defend to an auditor.
</objective>

Quick Route

You want to...Go to
Find which lines/branches each test uniquely coversCoverage Fingerprinting
Decide if "same coverage" means "delete one"The Coverage-Equality Trap
Surface copy-pasted near-duplicate testsNear-Duplicate Clustering
Find never-failing and always-flaky testsMining CI History
Decide redundant vs obsolete vs low-value vs keepPrune Decision Rules
Split a flat suite into smoke/core/extendedTiering
Safely delete the tests you flaggedDestructive Safety
Produce the "what we deleted and why" recordThe Audit Record

Discovery Questions

First, check .agents/qa-project-context.md in the project root and skip anything it already answers. Then clarify:

  • Which language/runner? pytest+coverage.py, Jest/Vitest+Istanbul, Go, JUnit — the per-test context mechanism differs per stack (and so does the mutation tool).
  • Is there CI test-result history, and how far back? No JUnit/Datadog/Trunk history means you cannot mine never-failing or flaky signals — you only have coverage and clustering.
  • What is the suite size and current wall-clock time? This sizes the tiering target (e.g. smoke under 5 min) and whether per-test coverage is feasible in one run or must be sharded.
  • What is the business risk map / critical paths? Tiering and the "keep" disposition both depend on it. If absent, run risk-based-testing first.
  • Who signs off on deletions, and is there a CODEOWNERS file? Deletion is destructive; you need a named approver before this skill removes anything.
  • What is the acceptable observation window? How long the team will run quarantined tests as skipped before permanent removal (default: 2 sprints / 2 releases).

Core Principles

  1. Coverage equality is not redundancy. Two tests hitting the same lines can assert completely different things — different oracles, inputs, edge cases. Line/branch coverage tells you what code ran, never what was checked. The only evidence that one test subsumes another is that the survivor catches the same faults, which you prove with mutation testing, not a coverage diff.

  2. The agent's edge is whole-corpus analysis, not deletion authority. An agent can fingerprint 4,000 tests, cluster near-duplicates, and cross-reference CI history in minutes — work no human does by hand. That is the entire value. But the agent proposes; a human approves. Never let the corpus-scale analysis become corpus-scale auto-deletion.

  3. Deletion is destructive and must be reversible in practice. "It's in git history" is not a recovery plan. Quarantine first (skip/xfail, or move to a deprecated suite), observe for a defined window, watch for escaped defects, then delete with sign-off. The grace period is the safety net, not the commit log.

  4. Every disposition is differentiated. Redundant, obsolete, and low-value are three different states with three different actions. Collapsing them into one "delete" bucket is how you lose real coverage. A test that never failed is not the same as a test that cannot fail.

  5. Evidence over intuition, recorded per test. Each removal carries its own row: category, the test that supersedes it, the coverage delta, who approved, and how to restore. If you cannot fill the row, you cannot delete the test.


1. Coverage Fingerprinting (per-test)

The wrong answer is a single combined --cov report that tells you per-file percentages. That cannot tell you which test covered which line, so it cannot tell you which tests overlap. You need per-test (dynamic) contexts.

pytest / coverage.py — record which test hit each line with dynamic contexts, and turn on branch coverage:

bash
pytest --cov=src --cov-branch --cov-context=test

--cov-context=test makes coverage.py call switch_context() around each test, tagging every measured line with the test that ran it. Add --cov-branch so a test that takes the if and one that takes the else are not treated as covering "the same line." The result is written into the .coverage SQLite database, in the context and line_bits/arc tables.

Then read the contexts table out of the .coverage SQLite DB to build a per-test fingerprint: for each test, the exact set of (file, line) and (file, branch-arc) pairs it covered. From those sets you compute:

  • Uniquely covered lines — lines/branches that only one test covers. Lose that test and you lose that coverage outright. These tests are pulling their weight; protect them.
  • Subsumption — test A's covered set is a superset of test B's. A candidate for redundancy (but see §2 — it is not proof).

See references/coverage-fingerprinting.md for the SQL to pull contexts from .coverage, the Python that builds per-test line/branch sets and computes uniquely-covered and subsumption relations, and the JS/Vitest+Istanbul --coverage equivalent (coverage-final.json with per-test reporters).

Do not rank tests by line count per test file, and do not delete tests merely for having a low overall coverage percentage — a one-line test can be the only thing guarding a critical branch.


2. The Coverage-Equality Trap (the load-bearing rule)

This is the single most important rule in the skill. When per-test data shows Test A and Test B cover exactly the same lines, the naive conclusion is "redundant, delete one." That is wrong, and here is the gotcha:

Coverage equality does not prove the tests have the same assertions, the same inputs, or the same oracle. Two tests can cover identical lines while one asserts the HTTP status and the other asserts the response body, or while they pass different edge-case inputs. One covers same lines but asserts different values; the other covers same lines but checks different state. Coverage measures execution, not verification.

To find out whether B is actually redundant — whether A truly subsumes B's fault-detection — run mutation testing:

  • mutmut (3.6.0+) or cosmic-ray for Python, StrykerJS (9.x) for JS/TS, PIT for Java/JVM, cargo-mutants for Rust.
  • Mutation testing injects faults (mutants) into the covered code. A test "kills" a mutant if it fails on the mutated code. If A kills every mutant that B kills, A genuinely subsumes B's fault detection and B is a defensible delete candidate. If B kills a mutant A misses, B catches a defect class A does not — keep B, even though coverage was identical.

Decision: identical coverage → flag as a candidate → confirm with mutation testing → only then propose deletion. When assertions differ and mutation results differ, retain both, do not delete. See references/mutation-confirmation.md for the mutmut/Stryker config that scopes mutation runs to the suspect lines and the kill-set comparison.


3. Near-Duplicate Clustering

Goal: surface copy-pasted tests without flagging every test in the same file. Grouping tests by filename is not clustering — it tells you nothing about similarity. Two defensible signals, combined:

  1. AST (abstract syntax tree) similarity. Parse each test into an AST, normalize away identifier names and literals, then compare structure. Use a token/tree similarity metric (Jaccard over normalized token shingles, cosine over AST n-grams, or tree edit distance). AST-based comparison ignores formatting and variable-name noise that defeats exact string matching or raw diff. Never use raw line numbers as a similarity signal.

  2. Coverage-profile signature. From §1, each test already has a covered-line/branch set — its execution profile. Tests with near-identical coverage signatures and near-identical ASTs are strong near-duplicate candidates; either signal alone is weak.

Cluster with a tunable similarity threshold (e.g. agglomerative/hierarchical clustering, cut at a configurable cutoff — start ~0.85, tune to your false-positive tolerance). Output clusters, never deletions.

Every cluster is routed to human review. The agent does not delete a whole cluster automatically — copy-paste tests frequently diverge in one assertion that matters. Present each cluster with its members, the pairwise similarity, and the coverage-profile overlap, and let a human confirm which (if any) collapse.

See references/clustering.md for the AST normalization, the shingle/Jaccard and tree-edit similarity functions, and the agglomerative clustering with the tunable threshold.


4. Mining CI History

Parse your test-result history — JUnit XML archives, or a platform that already stores it: Datadog Test Optimization, Trunk Flaky Tests, BuildPulse, CircleCI test insights. For each test compute pass rate / fail rate and the flip rate (how often consecutive runs transition pass↔fail). Two findings, two very different meanings:

Never-failing tests (zero failures / 100% pass over the window). The naive move is to delete any test that has never once failed. Wrong — never-failing does not mean delete or useless. A test most often never fails because it guards low-churn, low-risk, stable code — exactly the code nobody touches, so the test never trips. That is low defect-detection signal in this window, not zero value. Disposition: investigate, do not delete — check churn and risk of the code under test. If it covers a critical path that simply has not regressed, it stays.

Always-flaky tests. Flakiness is not decided by a single run. The real definition is different results on the same SHA / same commit — the same code produced a pass and a fail. Detect that by grouping runs by commit SHA and finding tests with both outcomes on one SHA (Trunk and BuildPulse do this natively). The naive move is "delete flaky tests to clean up CI." Wrong: a flaky test may still be your only coverage of a real path. Disposition: quarantine and fix the flake, never delete to clean up CI. Quarantine de-noises CI immediately; the root cause still gets fixed. See test-reliability for runtime quarantine and self-healing of an individual flaky test.

See references/ci-history-mining.md for the JUnit-XML aggregation script, the same-SHA flake query, and the Datadog/Trunk API pulls.


5. Prune Decision Rules

Stop deleting everything that "looks redundant." The three failure-categories are distinct states, and each gets a different disposition:

CategoryDefinition (the test is...)Disposition
Redundantsubsumed by another test — covers the same lines AND the survivor kills the same mutants (§2)merge or delete — but only after the mutation-kill check confirms subsumption; quarantine first
Obsoletetesting a feature that was removed / dead code / a path that no longer existsdelete — the code it tested is gone; verify the target truly no longer exists, then remove
Low-valuenever failed AND trivial (a getter, a no-op, no meaningful assertion)quarantine / route to review — low value is not zero value; confirm before removal
Keepcovers something uniquely, catches defects (positive defect-detection history), or guards a high-risk pathkeep — protected regardless of coverage overlap

The discipline: a different action per category. redundant => merge/delete-after-mutation-check, obsolete => delete, low-value => quarantine/review, keep => keep. Applying a single blanket rule to every flagged test is the anti-pattern that loses coverage. Note that "redundant" and "low-value" both route through quarantine, not straight to rm.


6. Tiering (smoke / core / extended)

Restructuring a flat suite into tiers is not sorting by speed. Smoke is not the first N tests in file order, it is not a random sample, and you must not tier solely on how fast each test runs. Tier on two evidence inputs:

  1. Business risk / critical path — from the risk map (risk-based-testing). Tests guarding revenue, auth, data integrity, and the top user journeys are smoke/core regardless of speed.
  2. Defect-detection history — tests that have actually caught real bugs (mine CI history + linked bug tickets for tests that failed on a commit that fixed a defect). A test with a track record of catching regressions earns a high tier on merit.
TierGoalSelection evidence
Smokefast, runs on every push, a few minutes, critical paths onlyhighest-risk paths + proven defect-catchers; fast enough to gate every commit
Coreper-PR / merge gateall critical + high-risk coverage, broader than smoke
Extendedfull / nightly / slow / pre-releaseeverything else — exhaustive, long-running, edge cases

Execution time as a tiering input is allowed — but only as a secondary tiebreaker, not the primary axis. Runtime and duration break ties; they never set the tier. Among equally-risky tests, prefer the faster ones for smoke. Encode tiers as markers/tags — @pytest.mark.smoke / @pytest.mark.extended, Jest/Vitest tag, or a JUnit category — so pytest -m smoke selects a tier without moving files. See references/tiering.md for the marker scheme and the defect-detection-history query.


Show full SKILL.md (1,348 more words)Show less

7. Destructive Safety (the grace period)

Asked to delete 600 tests, the naive agent opens one PR that rms 600 files. Never remove the whole cohort in a single PR, and do not rm the test files. Deletion is destructive; gate it:

  1. Quarantine before delete — never delete directly. Mark the flagged tests skipped (@pytest.mark.skip(reason="curation-2026-Q2, see audit row"), xfail, Jest .skip), or move them to a quarantine//deprecated/ suite that still lives in the repo. They stop running but stay visible and instantly restorable.
  2. Human sign-off is required. A named approver — via CODEOWNERS on the test directories, or an explicit reviewer on the PR — must approve. No "redundant tests need no review." Redundancy was a hypothesis until §2 confirmed it.
  3. Define an observation window. Run the suite with the cohort quarantined for a grace period — default 2 sprints / 2 releases — and monitor for escaped defects: watch production incidents and bug tickets for anything the quarantined tests would have caught. An escaped defect in the window means a test was not redundant; restore it.
  4. Only then delete, in small batches, each linked to its audit row (§8).
  5. Keep it restorable. Record the quarantine location and the commit so any test can be restored/reverted in one step. Git history is the floor, not the recovery plan.

See references/destructive-safety.md for the quarantine markers, a CODEOWNERS snippet for test paths, and the escaped-defect watch checklist.


8. The Audit Record

The deliverable that makes every removal defensible to a future engineer or an auditor is not a count of deleted tests. It is a per-test record — one row per deleted test — with a real justification each. A bare "we removed 600 redundant tests" is unauditable.

Each row carries:

  • Test id / path — what was removed.
  • Category & reason — redundant / obsolete / low-value, with the specific justification ("subsumed; survivor kills identical mutant set").
  • Superseded by — the surviving (superseding) test id that covers/replaces it (for redundant): it is covered by the test that subsumes it, or replaced by the test that supersedes it.
  • Coverage delta — lines/branches no longer covered after removal (before/after); ideally zero net loss, proven by §1.
  • Approval — approver name, the PR/commit SHA, and the date signed off.
  • Restore path — quarantine location and the one-line command to recover/revert it.

Store it as a committed CSV/Markdown table (e.g. docs/test-curation-log.md) so it is versioned alongside the deletions. See references/audit-record.md for the full column schema and a worked example row.


Anti-Patterns

1. "Same coverage, delete one"

The single most damaging mistake. Identical line coverage proves the lines ran, not that the assertions match. Confirm subsumption with mutation testing (§2) before proposing any deletion.

2. Two green signals read as a green light

A test that has never failed in 18 months AND is fully covered (same lines) by another is the classic "surely safe to delete" case. It is not sufficient on its own — still verify, do not delete automatically. A test that never failed may mean stable code not worthless code — it often means low risk and low change, not that the test is redundant. A different oracle could still be unique to this test (a different assertion, input, or edge case it alone checks). Ask the mutation question: does the OTHER test assert/catch/kill what this one does? Only after that, quarantine and observe for a grace period, then delete. Both signals are necessary-not-sufficient (§9 below / §2 / §7).

3. One merged coverage run with no per-test context

--cov=src with no context gives per-file percentages and cannot identify per-test overlap. Use --cov-context=test and read per-test contexts from the .coverage DB.

4. Grouping near-duplicates by filename

Filename proximity is not similarity. Cluster on AST similarity + coverage-profile signature with a tunable threshold (§3), and route clusters to humans — never auto-delete a whole cluster.

5. Deleting never-failing or flaky tests

Never-failing → investigate (likely low-churn code), do not delete. Flaky → quarantine and fix, do not delete to clean up CI. Flakiness is same-SHA divergence, not a single failed run.

6. One disposition for all flagged tests

Redundant, obsolete, and low-value need different actions (§5). Collapsing them into one "delete" bucket loses real coverage.

7. Tiering by speed (or "first N") alone

Smoke is risk + defect-detection history, not the fastest or first N tests. Runtime is a tiebreaker, not the axis (§6).

8. Big-bang deletion PR

600 tests rm'd in one PR with no quarantine, no sign-off, no observation window. Always quarantine → sign-off → observe → delete in batches (§7).

9. "It's in git history, just delete it"

Git history is not a recovery plan — nobody watches for the defect the deleted test would have caught. The quarantine grace period plus escaped-defect monitoring is the actual safety net.


Verification

Prove the audit actually holds before anyone deletes anything, smallest check first:

  • Per-test contexts were captured, not a flat report. sqlite3 .coverage "SELECT COUNT(*) FROM context WHERE context != '';" returns a count roughly equal to your test count (not 0/1). Zero means --cov-context=test did not run.
  • No deletion rests on coverage equality alone. For every redundant row, the audit log's mutation_check column is filled (yes + run id). grep -c 'redundant' docs/test-curation-log.md equals the number of rows whose mutation_check is non-blank.
  • Quarantine, not removal, hit the repo first. git log --diff-filter=D --name-only -- tests/ | grep test_ shows no test files deleted in the quarantine PR; the flagged tests are skipped/xfail or moved under tests/quarantine/, still collectible (pytest --collect-only -m skip or equivalent lists them).
  • Tiers select correctly. pytest -m smoke --collect-only returns only the risk/defect-catcher cohort and runs under the smoke budget; pytest -m "smoke or core or extended" --collect-only accounts for every test (no test is untagged).
  • The audit record is complete. No redundant row has a blank superseded_by, coverage_delta, or mutation_check; coverage_delta is 0 lines / 0 branches for every removal claiming zero net loss.

Done When

  • A per-test coverage fingerprint exists for the suite (built with --cov-context=test/--cov-branch or the Istanbul per-test equivalent), and uniquely-covered lines per test are computed.
  • Every "redundant" candidate has a mutation-testing result attached proving the survivor kills the same mutants — no deletion proposed on coverage equality alone.
  • Near-duplicate clusters were produced from AST + coverage-profile similarity at a stated threshold and routed to human review; no cluster was auto-deleted.
  • CI history was mined for never-failing (disposition: investigate) and same-SHA flaky tests (disposition: quarantine/fix) — neither category deleted on those signals.
  • Each flagged test has exactly one of {redundant, obsolete, low-value, keep} with the matching disposition applied.
  • Tiers (smoke/core/extended) are assigned via markers/tags using risk + defect-detection history as inputs; pytest -m smoke (or equivalent) selects a tier.
  • No test was deleted without: a quarantine period served, a named approver / CODEOWNERS sign-off recorded, and the observation window completed with no escaped defect.
  • A committed per-test audit record exists with one row per deletion containing category, superseded-by, coverage delta, approver+SHA+date, and restore path.

  • ai-qa-review — Judges whether an individual test is well-written (smells, weak assertions, testability). This skill decides whether a test should exist at all; ai-qa-review decides whether an existing one is good. Run ai-qa-review on the survivors after curation.
  • coverage-analysis — Owns coverage thresholds, gap analysis, and the mutation-testing setup at the project level. This skill consumes per-test coverage for redundancy decisions; go there for ratchets and CI gating.
  • test-reliability — Runtime self-healing and quarantine of a single flaky test as it fails. This skill finds the flaky cohort across CI history; test-reliability fixes one at a time.
  • risk-based-testing — Produces the risk matrix this skill's "keep" disposition and tiering depend on. Run it first if no risk map exists.
  • qa-project-context — Universal dependency: stack, runner, risk map, and ownership that drive every decision here.

Reference Files (in references/)

  • coverage-fingerprinting.md — .coverage SQLite context queries, per-test line/branch set construction, uniquely-covered + subsumption computation, and the JS/Istanbul per-test equivalent.
  • mutation-confirmation.md — mutmut/cosmic-ray/StrykerJS config scoped to suspect lines and the kill-set comparison that confirms subsumption.
  • clustering.md — AST normalization, Jaccard/cosine/tree-edit similarity, and agglomerative clustering with a tunable threshold.
  • ci-history-mining.md — JUnit-XML aggregation, same-SHA flake detection query, and Datadog/Trunk API pulls.
  • tiering.md — Marker/tag scheme for smoke/core/extended and the defect-detection-history query.
  • destructive-safety.md — Quarantine markers, CODEOWNERS for test paths, and the escaped-defect watch checklist.
  • audit-record.md — Full column schema for the deletion log and a worked example row.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/test-suite-curation of petrkindlmann/qa-skills.

  • SKILL.md
  • references/audit-record.md
  • references/ci-history-mining.md
  • references/clustering.md
  • references/coverage-fingerprinting.md
  • references/destructive-safety.md
  • references/mutation-confirmation.md
  • references/tiering.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

Test Suite Curation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Suite Curation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Suite Curation this skillpetrkindlmann/qa-skills168—~6.3kAutomated safety check: PassMIT
Caliber Testingcaliber-ai-org/ai-setup1.3k—~3.2kAutomated safety check: PassMIT
Memstack Development Test Writercwinvestments/memstack423—~3.7kAutomated safety check: PassProprietary
Frappe Testing UnitImpertio-Studio/Frappe_Claude_Skill_Package188—~3kAutomated safety check: PassMIT
Ut Checkintel/torch-xpu-ops115—~1.6kAutomated safety check: PassApache-2.0
Test Experteinverne/dotfiles121—~2.3kAutomated safety check: PassGPL-3.0

Similar skills

  • Caliber Testing

    caliber-ai-org/ai-setup

    Writes Vitest tests following project patterns: tests/ directories, vi.mock() for module mocking with vi.hoisted() for test-time factories, global LLM mock from src/test/setup.ts, environment…

    1.3k GitHub stars~3.2k tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • Memstack Development Test Writer

    cwinvestments/memstack

    A skill your agent uses when the user says 'write tests', 'add tests', 'test coverage', 'unit tests', 'integration tests', 'component tests', 'mocking', 'edge cases', or needs to generate tests with…

    423 GitHub stars~3.7k tokensUpdated 12 days ago
    Testing & QAAuto-check passed
  • Frappe Testing Unit

    Impertio-Studio/Frappe_Claude_Skill_Package

    A skill your agent uses when writing unit tests, integration tests, creating test fixtures, or running tests with bench run-tests.

    188 GitHub stars~3k tokensUpdated 22 days ago
    Testing & QAAuto-check passed
  • Ut Check

    intel/torch-xpu-ops

    Official

    Analyze UT (unit test) results for a torch-xpu-ops PR. An agent skill from intel/torch-xpu-ops.

    115 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Expert

    einverne/dotfiles

    Testing methodologies, test-driven development (TDD), unit and integration testing, and testing best practices across multiple frameworks.

    121 GitHub stars~2.3k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Scenario Testing

    aiskillstore/marketplace

    This skill should be used when writing tests, validating features, or needing to verify code works.

    430 GitHub starsUsed in 1 repo~830 tokens
    Testing & QAAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    168 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    168 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    168 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    168 GitHub stars~2.7k tokensUpdated 4 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    168 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    168 GitHub stars~4.6k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Test Suite Curation

What does Test Suite Curation do?

Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…. Test Suite Curation is an agent skill from petrkindlmann/qa-skills. Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune decision rules (redundant/obsolete/low-value/keep), smoke/core/extended tiering by risk and defect-detection history, and a defensible "what we deleted and why" record.

When should I use Test Suite Curation?

Test Suite Curation fits situations like: : audit the test suite; prune redundant tests; find duplicate tests; which tests can we delete.

How do I install Test Suite Curation in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a claude-code`. Or copy the skill folder (skills/test-suite-curation in petrkindlmann/qa-skills) into .claude/skills/test-suite-curation in your project. Claude Code loads it when a task matches its description.

How do I install Test Suite Curation in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a codex`. Or copy the skill folder (skills/test-suite-curation in petrkindlmann/qa-skills) into .agents/skills/test-suite-curation in your project. Codex loads it when a task matches its description.

Can I use Test Suite Curation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-suite-curation, .gemini/skills/test-suite-curation, .github/skills/test-suite-curation and .opencode/skills/test-suite-curation in your project.

What does Test Suite Curation need to run?

Going by SKILL.md and its folder, Test Suite Curation needs the command-line tools its instructions call (pytest, sqlite3 and git). Our summary lists: Python 3.

Does Test Suite Curation access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Suite Curation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Suite Curation use?

Test Suite Curation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Suite Curation use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Test Suite Curation?

Skills that share tags, products or a category with Test Suite Curation: Caliber Testing (caliber-ai-org/ai-setup, 1.3k stars), Memstack Development Test Writer (cwinvestments/memstack, 423 stars), Frappe Testing Unit (Impertio-Studio/Frappe_Claude_Skill_Package, 188 stars) and Ut Check (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Suite Curation?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 168 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.