Caliber Testing
caliber-ai-org/ai-setup
Writes Vitest tests following project patterns: tests/ directories, vi.mock() for module mocking with vi.hoisted() for test-time factories, global LLM mock from src/test/setup.ts, environment…
Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-suite-curation .claude/skills/test-suite-curation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .claude/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/test-suite-curation .agents/skills/test-suite-curation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .agents/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/test-suite-curation .cursor/skills/test-suite-curation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .cursor/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/test-suite-curation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/test-suite-curation .gemini/skills/test-suite-curation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .gemini/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills test-suite-curationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/test-suite-curation .github/skills/test-suite-curation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .github/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills test-suite-curation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/test-suite-curation .opencode/skills/test-suite-curation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-suite-curation" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-suite-curation into .opencode/skills/test-suite-curation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-suite-curation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-suite-curationAudit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…
Test Suite Curation is an agent skill from petrkindlmann/qa-skills. Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune decision rules (redundant/obsolete/low-value/keep), smoke/core/extended tiering by risk and defect-detection history, and a defensible "what we deleted and why" record. Deletion is destructive — quarantine and human sign-off are mandatory. Use when: "audit the test suite," "prune redundant tests," "find duplicate tests,"…
Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/audit-record.md`, `references/ci-history-mining.md` and `references/clustering.md`).
It sits in Testing & QA, covering Test generation, Failing and flaky tests and Test coverage. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pytestsqlite3gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Suite Curation loads about 6.3k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 255 tokens; SKILL.md has 3,348 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 3,348 words, ~6,288 tokens.
.claude/skills/test-suite-curation/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.<objective>
Test A and Test B cover the exact same lines, so a tired engineer deletes B — and three weeks later a production defect slips through because B was the only test that asserted the rounding was correct. Coverage equality is not redundancy. This skill audits an entire regression suite as a corpus (the redundancy analysis no human does by hand), prunes it on evidence rather than vibes, and treats every deletion as a destructive change that requires a quarantine grace period, human sign-off, and a record you can defend to an auditor.
</objective>
| You want to... | Go to |
|---|---|
| Find which lines/branches each test uniquely covers | Coverage Fingerprinting |
| Decide if "same coverage" means "delete one" | The Coverage-Equality Trap |
| Surface copy-pasted near-duplicate tests | Near-Duplicate Clustering |
| Find never-failing and always-flaky tests | Mining CI History |
| Decide redundant vs obsolete vs low-value vs keep | Prune Decision Rules |
| Split a flat suite into smoke/core/extended | Tiering |
| Safely delete the tests you flagged | Destructive Safety |
| Produce the "what we deleted and why" record | The Audit Record |
First, check .agents/qa-project-context.md in the project root and skip anything it already answers. Then clarify:
risk-based-testing first.Coverage equality is not redundancy. Two tests hitting the same lines can assert completely different things — different oracles, inputs, edge cases. Line/branch coverage tells you what code ran, never what was checked. The only evidence that one test subsumes another is that the survivor catches the same faults, which you prove with mutation testing, not a coverage diff.
The agent's edge is whole-corpus analysis, not deletion authority. An agent can fingerprint 4,000 tests, cluster near-duplicates, and cross-reference CI history in minutes — work no human does by hand. That is the entire value. But the agent proposes; a human approves. Never let the corpus-scale analysis become corpus-scale auto-deletion.
Deletion is destructive and must be reversible in practice. "It's in git history" is not a recovery plan. Quarantine first (skip/xfail, or move to a deprecated suite), observe for a defined window, watch for escaped defects, then delete with sign-off. The grace period is the safety net, not the commit log.
Every disposition is differentiated. Redundant, obsolete, and low-value are three different states with three different actions. Collapsing them into one "delete" bucket is how you lose real coverage. A test that never failed is not the same as a test that cannot fail.
Evidence over intuition, recorded per test. Each removal carries its own row: category, the test that supersedes it, the coverage delta, who approved, and how to restore. If you cannot fill the row, you cannot delete the test.
The wrong answer is a single combined --cov report that tells you per-file percentages. That cannot tell you which test covered which line, so it cannot tell you which tests overlap. You need per-test (dynamic) contexts.
pytest / coverage.py — record which test hit each line with dynamic contexts, and turn on branch coverage:
pytest --cov=src --cov-branch --cov-context=test--cov-context=test makes coverage.py call switch_context() around each test, tagging every measured line with the test that ran it. Add --cov-branch so a test that takes the if and one that takes the else are not treated as covering "the same line." The result is written into the .coverage SQLite database, in the context and line_bits/arc tables.
Then read the contexts table out of the .coverage SQLite DB to build a per-test fingerprint: for each test, the exact set of (file, line) and (file, branch-arc) pairs it covered. From those sets you compute:
See references/coverage-fingerprinting.md for the SQL to pull contexts from .coverage, the Python that builds per-test line/branch sets and computes uniquely-covered and subsumption relations, and the JS/Vitest+Istanbul --coverage equivalent (coverage-final.json with per-test reporters).
Do not rank tests by line count per test file, and do not delete tests merely for having a low overall coverage percentage — a one-line test can be the only thing guarding a critical branch.
This is the single most important rule in the skill. When per-test data shows Test A and Test B cover exactly the same lines, the naive conclusion is "redundant, delete one." That is wrong, and here is the gotcha:
Coverage equality does not prove the tests have the same assertions, the same inputs, or the same oracle. Two tests can cover identical lines while one asserts the HTTP status and the other asserts the response body, or while they pass different edge-case inputs. One covers same lines but asserts different values; the other covers same lines but checks different state. Coverage measures execution, not verification.
To find out whether B is actually redundant — whether A truly subsumes B's fault-detection — run mutation testing:
Decision: identical coverage → flag as a candidate → confirm with mutation testing → only then propose deletion. When assertions differ and mutation results differ, retain both, do not delete. See references/mutation-confirmation.md for the mutmut/Stryker config that scopes mutation runs to the suspect lines and the kill-set comparison.
Goal: surface copy-pasted tests without flagging every test in the same file. Grouping tests by filename is not clustering — it tells you nothing about similarity. Two defensible signals, combined:
AST (abstract syntax tree) similarity. Parse each test into an AST, normalize away identifier names and literals, then compare structure. Use a token/tree similarity metric (Jaccard over normalized token shingles, cosine over AST n-grams, or tree edit distance). AST-based comparison ignores formatting and variable-name noise that defeats exact string matching or raw diff. Never use raw line numbers as a similarity signal.
Coverage-profile signature. From §1, each test already has a covered-line/branch set — its execution profile. Tests with near-identical coverage signatures and near-identical ASTs are strong near-duplicate candidates; either signal alone is weak.
Cluster with a tunable similarity threshold (e.g. agglomerative/hierarchical clustering, cut at a configurable cutoff — start ~0.85, tune to your false-positive tolerance). Output clusters, never deletions.
Every cluster is routed to human review. The agent does not delete a whole cluster automatically — copy-paste tests frequently diverge in one assertion that matters. Present each cluster with its members, the pairwise similarity, and the coverage-profile overlap, and let a human confirm which (if any) collapse.
See references/clustering.md for the AST normalization, the shingle/Jaccard and tree-edit similarity functions, and the agglomerative clustering with the tunable threshold.
Parse your test-result history — JUnit XML archives, or a platform that already stores it: Datadog Test Optimization, Trunk Flaky Tests, BuildPulse, CircleCI test insights. For each test compute pass rate / fail rate and the flip rate (how often consecutive runs transition pass↔fail). Two findings, two very different meanings:
Never-failing tests (zero failures / 100% pass over the window). The naive move is to delete any test that has never once failed. Wrong — never-failing does not mean delete or useless. A test most often never fails because it guards low-churn, low-risk, stable code — exactly the code nobody touches, so the test never trips. That is low defect-detection signal in this window, not zero value. Disposition: investigate, do not delete — check churn and risk of the code under test. If it covers a critical path that simply has not regressed, it stays.
Always-flaky tests. Flakiness is not decided by a single run. The real definition is different results on the same SHA / same commit — the same code produced a pass and a fail. Detect that by grouping runs by commit SHA and finding tests with both outcomes on one SHA (Trunk and BuildPulse do this natively). The naive move is "delete flaky tests to clean up CI." Wrong: a flaky test may still be your only coverage of a real path. Disposition: quarantine and fix the flake, never delete to clean up CI. Quarantine de-noises CI immediately; the root cause still gets fixed. See test-reliability for runtime quarantine and self-healing of an individual flaky test.
See references/ci-history-mining.md for the JUnit-XML aggregation script, the same-SHA flake query, and the Datadog/Trunk API pulls.
Stop deleting everything that "looks redundant." The three failure-categories are distinct states, and each gets a different disposition:
| Category | Definition (the test is...) | Disposition |
|---|---|---|
| Redundant | subsumed by another test — covers the same lines AND the survivor kills the same mutants (§2) | merge or delete — but only after the mutation-kill check confirms subsumption; quarantine first |
| Obsolete | testing a feature that was removed / dead code / a path that no longer exists | delete — the code it tested is gone; verify the target truly no longer exists, then remove |
| Low-value | never failed AND trivial (a getter, a no-op, no meaningful assertion) | quarantine / route to review — low value is not zero value; confirm before removal |
| Keep | covers something uniquely, catches defects (positive defect-detection history), or guards a high-risk path | keep — protected regardless of coverage overlap |
The discipline: a different action per category. redundant => merge/delete-after-mutation-check, obsolete => delete, low-value => quarantine/review, keep => keep. Applying a single blanket rule to every flagged test is the anti-pattern that loses coverage. Note that "redundant" and "low-value" both route through quarantine, not straight to rm.
Restructuring a flat suite into tiers is not sorting by speed. Smoke is not the first N tests in file order, it is not a random sample, and you must not tier solely on how fast each test runs. Tier on two evidence inputs:
risk-based-testing). Tests guarding revenue, auth, data integrity, and the top user journeys are smoke/core regardless of speed.| Tier | Goal | Selection evidence |
|---|---|---|
| Smoke | fast, runs on every push, a few minutes, critical paths only | highest-risk paths + proven defect-catchers; fast enough to gate every commit |
| Core | per-PR / merge gate | all critical + high-risk coverage, broader than smoke |
| Extended | full / nightly / slow / pre-release | everything else — exhaustive, long-running, edge cases |
Execution time as a tiering input is allowed — but only as a secondary tiebreaker, not the primary axis. Runtime and duration break ties; they never set the tier. Among equally-risky tests, prefer the faster ones for smoke. Encode tiers as markers/tags — @pytest.mark.smoke / @pytest.mark.extended, Jest/Vitest tag, or a JUnit category — so pytest -m smoke selects a tier without moving files. See references/tiering.md for the marker scheme and the defect-detection-history query.
Asked to delete 600 tests, the naive agent opens one PR that rms 600 files. Never remove the whole cohort in a single PR, and do not rm the test files. Deletion is destructive; gate it:
@pytest.mark.skip(reason="curation-2026-Q2, see audit row"), xfail, Jest .skip), or move them to a quarantine//deprecated/ suite that still lives in the repo. They stop running but stay visible and instantly restorable.See references/destructive-safety.md for the quarantine markers, a CODEOWNERS snippet for test paths, and the escaped-defect watch checklist.
The deliverable that makes every removal defensible to a future engineer or an auditor is not a count of deleted tests. It is a per-test record — one row per deleted test — with a real justification each. A bare "we removed 600 redundant tests" is unauditable.
Each row carries:
Store it as a committed CSV/Markdown table (e.g. docs/test-curation-log.md) so it is versioned alongside the deletions. See references/audit-record.md for the full column schema and a worked example row.
The single most damaging mistake. Identical line coverage proves the lines ran, not that the assertions match. Confirm subsumption with mutation testing (§2) before proposing any deletion.
A test that has never failed in 18 months AND is fully covered (same lines) by another is the classic "surely safe to delete" case. It is not sufficient on its own — still verify, do not delete automatically. A test that never failed may mean stable code not worthless code — it often means low risk and low change, not that the test is redundant. A different oracle could still be unique to this test (a different assertion, input, or edge case it alone checks). Ask the mutation question: does the OTHER test assert/catch/kill what this one does? Only after that, quarantine and observe for a grace period, then delete. Both signals are necessary-not-sufficient (§9 below / §2 / §7).
--cov=src with no context gives per-file percentages and cannot identify per-test overlap. Use --cov-context=test and read per-test contexts from the .coverage DB.
Filename proximity is not similarity. Cluster on AST similarity + coverage-profile signature with a tunable threshold (§3), and route clusters to humans — never auto-delete a whole cluster.
Never-failing → investigate (likely low-churn code), do not delete. Flaky → quarantine and fix, do not delete to clean up CI. Flakiness is same-SHA divergence, not a single failed run.
Redundant, obsolete, and low-value need different actions (§5). Collapsing them into one "delete" bucket loses real coverage.
Smoke is risk + defect-detection history, not the fastest or first N tests. Runtime is a tiebreaker, not the axis (§6).
600 tests rm'd in one PR with no quarantine, no sign-off, no observation window. Always quarantine → sign-off → observe → delete in batches (§7).
Git history is not a recovery plan — nobody watches for the defect the deleted test would have caught. The quarantine grace period plus escaped-defect monitoring is the actual safety net.
Prove the audit actually holds before anyone deletes anything, smallest check first:
sqlite3 .coverage "SELECT COUNT(*) FROM context WHERE context != '';" returns a count roughly equal to your test count (not 0/1). Zero means --cov-context=test did not run.redundant row, the audit log's mutation_check column is filled (yes + run id). grep -c 'redundant' docs/test-curation-log.md equals the number of rows whose mutation_check is non-blank.git log --diff-filter=D --name-only -- tests/ | grep test_ shows no test files deleted in the quarantine PR; the flagged tests are skipped/xfail or moved under tests/quarantine/, still collectible (pytest --collect-only -m skip or equivalent lists them).pytest -m smoke --collect-only returns only the risk/defect-catcher cohort and runs under the smoke budget; pytest -m "smoke or core or extended" --collect-only accounts for every test (no test is untagged).redundant row has a blank superseded_by, coverage_delta, or mutation_check; coverage_delta is 0 lines / 0 branches for every removal claiming zero net loss.--cov-context=test/--cov-branch or the Istanbul per-test equivalent), and uniquely-covered lines per test are computed.pytest -m smoke (or equivalent) selects a tier.references/).coverage SQLite context queries, per-test line/branch set construction, uniquely-covered + subsumption computation, and the JS/Istanbul per-test equivalent.© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in skills/test-suite-curation of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
Test Suite Curation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Suite Curation this skillpetrkindlmann/qa-skills | 168 | — | ~6.3k | Automated safety check: Pass | MIT | |
| Caliber Testingcaliber-ai-org/ai-setup | 1.3k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Memstack Development Test Writercwinvestments/memstack | 423 | — | ~3.7k | Automated safety check: Pass | Proprietary | |
| Frappe Testing UnitImpertio-Studio/Frappe_Claude_Skill_Package | 188 | — | ~3k | Automated safety check: Pass | MIT | |
| Ut Checkintel/torch-xpu-ops | 115 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Test Experteinverne/dotfiles | 121 | — | ~2.3k | Automated safety check: Pass | GPL-3.0 |
caliber-ai-org/ai-setup
Writes Vitest tests following project patterns: tests/ directories, vi.mock() for module mocking with vi.hoisted() for test-time factories, global LLM mock from src/test/setup.ts, environment…
cwinvestments/memstack
A skill your agent uses when the user says 'write tests', 'add tests', 'test coverage', 'unit tests', 'integration tests', 'component tests', 'mocking', 'edge cases', or needs to generate tests with…
Impertio-Studio/Frappe_Claude_Skill_Package
A skill your agent uses when writing unit tests, integration tests, creating test fixtures, or running tests with bench run-tests.
intel/torch-xpu-ops
Analyze UT (unit test) results for a torch-xpu-ops PR. An agent skill from intel/torch-xpu-ops.
einverne/dotfiles
Testing methodologies, test-driven development (TDD), unit and integration testing, and testing best practices across multiple frameworks.
aiskillstore/marketplace
This skill should be used when writing tests, validating features, or needing to verify code works.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Categories
Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune…. Test Suite Curation is an agent skill from petrkindlmann/qa-skills. Audit a whole regression suite and prune/restructure it with evidence: per-test coverage fingerprinting, AST near-duplicate clustering, CI-history mining for never-failing and flaky tests, prune decision rules (redundant/obsolete/low-value/keep), smoke/core/extended tiering by risk and defect-detection history, and a defensible "what we deleted and why" record.
Test Suite Curation fits situations like: : audit the test suite; prune redundant tests; find duplicate tests; which tests can we delete.
Run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a claude-code`. Or copy the skill folder (skills/test-suite-curation in petrkindlmann/qa-skills) into .claude/skills/test-suite-curation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a codex`. Or copy the skill folder (skills/test-suite-curation in petrkindlmann/qa-skills) into .agents/skills/test-suite-curation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill test-suite-curation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-suite-curation, .gemini/skills/test-suite-curation, .github/skills/test-suite-curation and .opencode/skills/test-suite-curation in your project.
Going by SKILL.md and its folder, Test Suite Curation needs the command-line tools its instructions call (pytest, sqlite3 and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Suite Curation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Test Suite Curation: Caliber Testing (caliber-ai-org/ai-setup, 1.3k stars), Memstack Development Test Writer (cwinvestments/memstack, 423 stars), Frappe Testing Unit (Impertio-Studio/Frappe_Claude_Skill_Package, 188 stars) and Ut Check (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 168 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.