Agent skill

Test Reliability

by petrkindlmann in petrkindlmann/qa-skills

Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.

MITAuto-check passedTesting & QA

Install Test Reliability

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills test-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-reliability .claude/skills/test-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-reliability
GitHub stars
165
Token cost
~5.6k tokens
SKILL.md length
2,275 words
Files
2 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.

  • Works in 8 steps: Silent Selector Replacement → "Just Retry It" as a Fix → Disabling Flaky Tests Permanently → …
  • Self-healing locator
  • SKILL.md covers Quick Route, Discovery Questions, Core Principles and Locator Resilience, plus 6 more sections
  • Calls npx

What it does

Test Reliability is an agent skill from petrkindlmann/qa-skills. Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair. Goes beyond simple locator fallbacks to cover action-level healing, data healing, and observable repair workflows. Use when: "flaky test," "test stability," "self-healing locator," "broken locator recovery," "unreliable test," "quarantine flaky test." Not for: bulk regenerating selectors after a planned UI refactor — use…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/flaky-test-runbook.md`).

It sits in Testing & QA, covering Failing and flaky tests, Issue triage and QA and bug reports. It works with Playwright. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • Self-healing locator
  • Broken locator recovery
  • Unreliable test
  • That one rewrites many tests offline)

Example prompts

  • “flaky test,”
  • “test stability,”
  • “self-healing locator,”
  • “/test-reliability”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Silent Selector Replacement
  2. "Just Retry It" as a Fix
  3. Disabling Flaky Tests Permanently
  4. Treating All Flakiness the Same
  5. waitForTimeout as a Stability Fix
  6. Healing Without Observability
  7. Over-Engineering Healing Before Writing Stable Tests
  8. No Quarantine Expiry

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Reliability loads about 5.6k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 202 tokens; SKILL.md has 2,275 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~202
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,275 words, ~5,554 tokens.

Download SKILL.mdSave it as .claude/skills/test-reliability/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-reliability
description
Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair. Goes beyond simple locator fallbacks to cover action-level healing, data healing, and observable repair workflows. Use when: "flaky test," "test stability," "self-healing locator," "broken locator recovery," "unreliable test," "quarantine flaky test." Not for: bulk regenerating selectors after a planned UI refactor — use selector-drift-recovery (this skill heals ONE test at runtime; that one rewrites many tests offline). Not for: classifying/clustering CI failures into bug reports — use ai-bug-triage. Related: playwright-automation, selector-drift-recovery, ci-cd-integration, qa-metrics, ai-bug-triage.
license
MIT
metadata.author
kindlmann
metadata.version
2.0
metadata.category
ai-qa
<objective>
A retried test that still flakes will eventually fail 3-of-3 during your most critical release, and a silently auto-repaired test may now verify a different element entirely. This skill builds suites teams can trust: resilient locators, classified flakes, environment-aware healing, data healing, and observable repair with confidence scoring — every automated fix produces evidence a human can review.
</objective>

Quick Route

SituationGo to
Test flakes in CI, root cause unknownFlake Classification → decision tree
Locator broke, want resilient replacementLocator Resilience
Action fails, suspect slow backend not UI bugEnvironment-Aware Healing
401/404/429 mid-test from stale dataData Healing
Want auto-repair with review gateObservable Repair Workflow
Isolate a flaky test without blocking CIQuarantine Management
Step-by-step triage of one flaky testreferences/flaky-test-runbook.md

Discovery Questions

Check .agents/qa-project-context.md first — it carries known flaky areas, selector strategy, and CI environment details. Skip any question it already answers.

  • What is your current flaky test rate? Check CI failure stats over the last 30 days. Below 2% is healthy; 2-5% needs attention; above 5% is eroding team trust.
  • Where is the pain concentrated? Locator breakage? Timing? Test data? Environment? If unknown, instrument first (see Flake Classification).
  • What is your current selector strategy? data-testid everywhere, mixed CSS and role-based, or no strategy (whatever works)? This sets the stability-score baseline.
  • How do you handle flaky tests today? Retry and hope, skip and forget, or something structured? Decides how much process you need to add.
  • What CI environment runs the tests? Same machine every time or different runners? Consistent or variable resources? Drives the environment-vs-test diagnosis.
  • What is your test data strategy? Shared database, per-test fixtures, factory seeding, or external services? Decides whether data healing applies.

Core Principles

  1. Prevention over cure. Writing a resilient test costs 1x. Investigating a flaky one costs 10x. Losing team trust in the suite costs 100x.

  2. Healing must be observable and reviewable. Every automated repair produces evidence: what broke, what was tried, what worked, the confidence score. Silent fixes erode trust as fast as silent failures.

  3. Classify before fixing. The fix for a timing issue is completely different from the fix for a data dependency. Wrong diagnosis wastes effort and can make things worse.

  4. Flaky tests are bugs. Not annoyances to tolerate. A flaky test either has a test bug (fix the test), reveals an app bug (fix the app), or exposes an environment issue (fix the environment).

  5. Track reliability as a metric, not a feeling. Measure flaky rate, mean time to heal, quarantine age, and selector stability. What gets measured gets fixed.

  6. Self-healing is a spectrum. Start with resilient locators (Level 1), add fallback strategies (Level 2), then environment-aware healing (Level 3), then confidence-scored auto-repair (Level 4). Do not jump to Level 4 before mastering Level 1.

Locator Resilience

Multi-Attribute Selectors (Beyond Fallback Chains)

A single locator strategy is a single point of failure. Multi-attribute selectors combine multiple signals for one element lookup — resilience without fallback-chain complexity.

The key insight: instead of "try A, then B, then C," use "find element matching A AND B AND C with tolerance for one signal missing."

typescript
// Multi-attribute locator: tries combinations from most specific to least
const submitBtn = await multiAttributeLocator(page, {
  testId: 'checkout-submit',              // most stable signal
  role: 'button',                          // semantic signal
  name: /place order/i,                    // accessible name
  nearText: 'Order Summary',              // visual context
});
// Internally: tries testId+role+name first, then testId alone, then role+name,
// then text, then nearText+role. Returns first visible match.
// Unlike fallback chains, it combines signals for higher confidence.
DOM Similarity / Neighbor Context Matching

When a locator fails, the element may still exist with changed attributes. Use surrounding DOM context to find it:

  1. Parent + tag + type: Find the container (by testId), then locate by tag and type within it.
  2. Preceding label: Find sibling text (label), then locate the adjacent input/button.
  3. Nearby text context: Find visible text near the target, then locate the element type in the same parent.

These are repair candidates scored by the confidence system below — not runtime fallbacks.

Selector Stability Scoring

Rate every selector on a 0-5 scale to prioritize refactoring.

ScoreStrategySurvives
5getByTestId('submit-order')CSS, text, and structural changes
4getByRole('button', { name: 'Submit' })CSS and structural changes
3getByLabel('Email')CSS changes; breaks on label rewording
2getByText('Submit Order')Breaks on any copy change
1locator('.btn-primary.submit')Breaks on CSS or structural change
0locator('//div[3]/button[1]')Breaks on any DOM change

Target: Average score of 3.5+ across the suite. Audit monthly. Prioritize fixing score-0 and score-1 selectors. Emit one score per locator to selector-stability.md (or a CI step) and report the suite average so the 3.5 target is verifiable, not asserted.

Flake Classification Framework

Every flaky test has a root cause category. Classifying correctly determines the fix.

Categories
CategorySignalRoot CauseFix Direction
TimingTimeout errors, passes on retry, worse in CIRace condition, animation, async operationWait for condition, not time
Data dependencyFails with other tests, passes aloneShared state, missing cleanupIsolate per-test, fixture cleanup
EnvironmentFails on specific runner, correlates with loadResource contention, network latencyMock externals, increase resources
Order dependencyFails with --shard or fullyParallelDepends on another test's side effectSelf-contained setup
Time sensitivityFails at specific times (midnight, month-end)Uses real clock, date boundaryMock clock, relative comparisons
Visual renderingScreenshot diff flickers, subpixel differencesFont rendering, antialiasing, animation frameIncrease threshold, mask dynamic regions
External serviceCorrelates with third-party statusReal HTTP calls in testsMock external APIs
Classification Decision Tree
Test is flaky
│
├── Does it pass when run alone?
│   ├── YES → ORDER DEPENDENCY or DATA DEPENDENCY
│   │   ├── Does another test create/modify data it needs? → ORDER DEPENDENCY
│   │   └── Does it share a database/file/cache? → DATA DEPENDENCY
│   │
│   └── NO → Not order/data dependent. Continue below.
│
├── Does it fail more often in CI than locally?
│   ├── YES → TIMING or ENVIRONMENT
│   │   ├── Timeout errors? → TIMING (CI is slower)
│   │   ├── Connection errors? → ENVIRONMENT (network latency / service)
│   │   └── Resource errors (OOM, disk)? → ENVIRONMENT (resource contention)
│   │
│   └── NO → Same rate locally and CI. Continue below.
│
├── Does it fail at specific times?
│   ├── YES → TIME SENSITIVITY
│   │   ├── Near midnight? → Date boundary issue
│   │   ├── Near month/year end? → Calendar calculation
│   │   └── Specific hour? → Timezone issue
│   │
│   └── NO → Continue below.
│
├── Does it involve screenshots or visual comparison?
│   ├── YES → VISUAL RENDERING
│   │
│   └── NO → Continue below.
│
├── Does it call external HTTP APIs?
│   ├── YES → EXTERNAL SERVICE
│   │
│   └── NO → TIMING (most likely — default classification)
│       └── Investigate: what async operation is not being awaited?

The fix per category lives in references/flaky-test-runbook.md (Step 4) with full code patterns.

Environment-Aware Healing

Not all test failures are test problems. Some are environment problems. Environment-aware healing distinguishes the two and adapts.

Slow Backend vs True UI Failure

When an action fails, check backend health before blaming the test:

  1. Action fails → Hit /api/health.
  2. Backend unhealthy (5xx or timeout) → Retry with exponential backoff. Diagnose as backend_down. This is not a UI bug.
  3. Backend healthy (2xx) → This is a real UI/test failure. Do not retry.

Return a structured diagnosis: { success: boolean; diagnosis: 'backend_down' | 'ui_failure' | 'backend_slow_recovered' }. This feeds flake classification — backend issues are environment issues, not test bugs.

Resource Contention Detection

Before declaring a test failure in CI, check for resource contention:

  • Browser health: Load about:blank. If it takes > 2s (baseline < 500ms), the runner is overloaded.
  • API health: Hit /api/health. If it takes > 5s (baseline < 1s), the backend is under pressure.
  • Diagnosis: If either check fails, classify as an environment issue and annotate the test result. Do not count resource-contention failures toward flaky test rates.

Data Healing

Test data expires, gets cleaned up, or becomes invalid. Data healing detects and regenerates stale test data.

Common Data Failure Patterns
PatternSignalFix
Expired auth token401 response during testRegenerate token in fixture
Deleted test record404 when accessing seeded dataRe-seed before test
Uniqueness violation409 or constraint errorGenerate unique identifiers per run
Stale cacheWrong data returnedClear cache in setup
Exceeded quota429 or rate limit errorReset quotas or use dedicated test account
Self-Healing Test Data Fixture Pattern

Build fixtures that verify data exists and regenerate if stale.

typescript
// Pattern: verify → heal → use → cleanup
testUser: async ({ request }, use, testInfo) => {
  // 1. Try to find existing test user by deterministic email
  // 2. Verify auth token is still valid (GET /api/me)
  // 3. If token expired → refresh it (POST /refresh-token), mark as healed
  // 4. If user missing → create new one, mark as healed
  // 5. If healed → annotate testInfo for observability
  // 6. use(user) → run the test
  // 7. Cleanup: delete test user (guaranteed by fixture, even on failure)
}

Key patterns:

  • Use testInfo.testId in email/identifiers for per-test uniqueness.
  • Annotate testInfo.annotations when healing occurs, for observability.
  • Always clean up in the fixture's post-use block, not in afterEach — fixtures guarantee cleanup on failure.

Observable Repair Workflow

Core guardrail: Healing must be observable and reviewable. Every repair follows this flow:

Failure Detected
  │
  ▼
Candidate Repair Generated
  │
  ▼
Confidence Score Computed (0.0 - 1.0)
  │
  ▼
Evidence Diff Produced (what changed, what was tried)
  │
  ▼
Approval Policy Applied
  │ ├── Score >= 0.9   → Auto-apply, log for batch review
  │ ├── Score 0.7-0.89 → Apply in quarantine, flag for individual review
  │ ├── Score 0.5-0.69 → Do NOT apply, open PR with evidence for review
  │ └── Score < 0.5    → Discard, manual investigation required
  │
  ▼
Intent Fidelity Check (does repaired test still test the same thing?)
  │
  ▼
Rollback if intent fidelity drops
Confidence Scoring

Score each repair candidate on six dimensions (weighted sum, 0.0-1.0):

DimensionWeightScoring
Match specificity0.30testId=1.0, role=0.9, text=0.7, context=0.5, CSS=0.3
Element visible0.151.0 if visible, 0.0 if not
Same parent container0.151.0 if same container, 0.0 if different
Same element type0.151.0 if same tag+role, 0.0 if different
Text similarity0.150.0-1.0 (Levenshtein ratio of accessible name)
Attribute overlap0.100.0-1.0 (Jaccard of shared attributes)

Score thresholds (half-open bands — 0.9 belongs to the auto-apply tier only):

  • >= 0.9 — Auto-apply, log for batch review.
  • 0.7-0.89 — Apply in quarantine, flag for individual review.
  • 0.5-0.69 — Do not apply; open a PR with evidence for review.
  • < 0.5 — Discard, manual investigation required.

references/flaky-test-runbook.md (Score Interpretation) uses these exact four bands — keep all three locations identical if you edit them.

Repair Evidence

Every repair produces an evidence record containing: test file, test name, failure type, original locator, candidate replacements (each with confidence, evidence string, and intent-preserved flag), which candidate was selected, timestamp, approval path, rollback trigger, and a screencast.webm recording of the repair run (Playwright 1.59+ page.screencast with showActions annotations — the "agentic video receipt"). Without the screencast, a Level-3/4 healed test asks reviewers to trust a JSON record; with it, the diff and the runtime are both inspectable.

Show full SKILL.md (902 more words)Show less
Buy vs Build

Hand Levels 3-4 to a vendor when you don't want to own a confidence-scored healer; build in-house when you need on-prem/air-gapped deployment, an explicit AI-prompt audit trail, or quarantine logic your tracker can't express. Building stops being worth it once a hosted tool already fingerprints + clusters + auto-quarantines for you.

  • For Selenium / Selenide / Robot Framework suites, Healenium remains the dedicated self-healing layer (now on AWS Marketplace).
  • For agent-driven self-healing in Playwright, Playwright MCP is the canonical path — it gives the agent live browser control to discover replacement locators when CLI + skills isn't enough.
  • Hosted platforms shipping this workflow: Trunk Flaky Tests (auto-quarantine + AI failure clustering; its 2026 Quarantined Tests API returns the current quarantine list programmatically — same hygiene workflow this skill builds by hand), CloudBees Smart Tests (formerly Launchable; agents searching old docs may find the old name), and Datadog Test Optimization (renamed from "Datadog Test Visibility" in December 2024). All three offer fingerprinting + clustering + auto-quarantine that maps onto Levels 3-4.
Intent Fidelity Checking

After applying a repair, verify the test still exercises the same user intent:

  • Element type changed (e.g. button → a) — intent NOT preserved, rollback.
  • ARIA role changed (e.g. button → link) — intent NOT preserved, rollback.
  • Form action changed (targets a different endpoint) — intent NOT preserved, rollback.
  • Same tag, role, and form action — intent preserved, keep repair.

A repair that changes WHAT the test verifies (not just HOW it finds elements) must be rolled back.

Quarantine Management

Quarantine isolates flaky tests so they run but do not block CI. See references/flaky-test-runbook.md (Step 6) for the full config and CI wiring; the essentials:

typescript
// Tag the flaky test
test('intermittent WebSocket reconnect', {
  tag: ['@quarantine'],
  annotation: {
    type: 'quarantine',
    description: 'Flaky since 2026-03-15. Race condition in WebSocket handler. Ticket: BUG-1234.',
  },
}, async ({ page }) => { /* ... */ });
typescript
// playwright.config.ts — separate projects
projects: [
  { name: 'stable', testMatch: /.*\.spec\.ts/, grep: /^(?!.*@quarantine)/ },  // exclude quarantine
  { name: 'quarantine', grep: /@quarantine/, retries: 3 },
],

In CI, run --project=stable as a blocking step and --project=quarantine with continue-on-error: true so the quarantine project never blocks the pipeline.

Quarantine Lifecycle
1. DETECT    — Test identified as flaky (CI reporter or manual triage)
2. TAG       — Add @quarantine annotation with ticket link and date
3. ISOLATE   — Quarantine project runs separately, does not block
4. DIAGNOSE  — Follow the flaky test runbook (references/flaky-test-runbook.md)
5. FIX       — Apply the fix pattern for the classified category
6. VERIFY    — Run 50x with --repeat-each, zero failures required
7. RELEASE   — Remove @quarantine tag, add annotation documenting the fix
Quarantine Hygiene Rules
  • Maximum quarantine age: 14 days. After 14 days, fix it or delete it. Permanent quarantine is permanent rot.
  • Every quarantine entry has a ticket link. No anonymous quarantines.
  • Weekly review. Check the quarantine list every sprint. Aging quarantines get escalated.
  • Track quarantine size. More than 5% of tests in quarantine signals a systemic problem requiring process change, not just test fixes.

Anti-Patterns

1. Silent Selector Replacement

Replacing a broken selector with no logging, review, or confidence scoring. The repaired test may now verify a different element entirely. Every repair must produce evidence.

2. "Just Retry It" as a Fix

Retries are a detection mechanism, not a fix. A test that needs retry 2-of-3 will eventually fail 3-of-3 during your most critical release.

3. Disabling Flaky Tests Permanently

test.skip('flaky, will fix later') — "later" never comes. Either quarantine with tracking or delete entirely. Skipped tests with no ticket are dead code.

4. Treating All Flakiness the Same

Timing issues and data dependencies need completely different fixes. Adding waitForTimeout(5000) to a data-dependency problem makes the test slower and still flaky.

5. waitForTimeout as a Stability Fix
typescript
// NEVER the right fix
await page.waitForTimeout(5000);

// Wait for the actual condition
await expect(page.getByRole('table')).toBeVisible();
await page.waitForResponse(resp => resp.url().includes('/api/data') && resp.status() === 200);
6. Healing Without Observability

Auto-repair that produces no logs, evidence, or confidence scores. You cannot improve what you cannot measure, and you cannot trust what you cannot review.

7. Over-Engineering Healing Before Writing Stable Tests

Building a complex self-healing framework before adopting basic resilient-locator patterns. Start with multi-attribute selectors and proper waits. Add healing infrastructure only when data shows where breakage occurs.

8. No Quarantine Expiry

Tests sit in quarantine for months. Quarantine is a temporary state, not a permanent home. Enforce a 14-day maximum.

Failure Modes

SymptomLikely causeFix or check
Backend health check itself flaps → false backend_down diagnosisHealth endpoint is itself flaky/slowTrack the health endpoint's own p99 separately; don't gate diagnosis on a single probe
Artifact storage balloons after enabling repair videopage.screencast recording on every run, not just repairsRecord only on the repair path, not the happy path
Auto-repair accuracy drops below 80%Confidence threshold too low, or intent-fidelity check skippedRaise the auto-apply floor; never skip the intent check
Quarantine project blocks the pipelineMissing continue-on-error on the quarantine CI stepAdd it; the quarantine project must never block merges

Verification

  • Reproduce the flake: npx playwright test <spec> --repeat-each=20 --workers=4 --trace=on — it must fail at least once before you trust any fix.
  • After fixing, prove stability: npx playwright test <spec> --repeat-each=50 --workers=4 — require 50/50 passes, in CI conditions too.
  • Confirm quarantine routing: npx playwright test --project=stable excludes @quarantine tests and --project=quarantine runs only them.
  • Confirm selector audit emits numbers: the stability report lists a score per locator and a suite average.

Done When

  • Every flaky test is identified and categorized by root cause (timing, data dependency, environment, etc.).
  • Each flaky test is quarantined or fixed — no test silently retried without a documented plan and ticket reference.
  • selector-stability.md (or the CI report) lists a stability score per locator and reports a suite average >= 3.5.
  • The flaky-test-rate metric (% of tests passing on retry) is published to the CI dashboard and visible to the team.
  • Every quarantine entry has a ticket reference and an expiry date <= 14 days out.
  • selector-drift-recovery — go there to bulk-regenerate many selectors offline after a planned UI refactor; this skill heals one test at runtime.
  • playwright-automation — full Playwright setup, Page Object Model, fixtures, and CI integration that the patterns here build on.
  • ci-cd-integration — pipeline configuration, parallel execution, and the quarantine job wiring referenced above.
  • qa-metrics — track flaky rate, mean-time-to-heal, quarantine size, and selector stability over time.
  • ai-bug-triage — when flake investigation reveals a real app bug, hand the failure to the triage pipeline to classify and report it.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/test-reliability of petrkindlmann/qa-skills.

  • SKILL.md
  • references/flaky-test-runbook.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

Test Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Reliability this skillpetrkindlmann/qa-skills165—~5.6kAutomated safety check: PassMIT
Diagnose Playwright Failure as Product Bugappsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0
Playwright Regression Testingfugazi/test-automation-skills-agents247—~1.4kAutomated safety check: PassMIT
Claude Code QAPramodDutta/qaskills232—~2.3kAutomated safety check: PassMIT
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence
Fix Failing Playwright Specappsmithorg/appsmith41k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

    41k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Regression Testing

    fugazi/test-automation-skills-agents

    Govern Playwright TypeScript regression suites across many tests.

    247 GitHub stars~1.4k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Claude Code QA

    PramodDutta/qaskills

    The complete QA skill for Claude Code — turn Claude into an expert QA engineer that picks the right test type, writes reliable Playwright, Cypress, and pytest tests, eliminates flaky tests, enforces…

    232 GitHub stars~2.3k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • Fix Failing Playwright Spec

    appsmithorg/appsmith

    Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions.

    41k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

    2.8k GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    165 GitHub stars~4.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    165 GitHub stars~4.5k tokensUpdated 3 mo ago
    Auto-check passed
  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    165 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    165 GitHub stars~2.7k tokensUpdated 3 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    165 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    165 GitHub stars~4.6k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Categories

Questions about Test Reliability

What does Test Reliability do?

Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair. Test Reliability is an agent skill from petrkindlmann/qa-skills. Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.

When should I use Test Reliability?

Test Reliability fits situations like: self-healing locator; broken locator recovery; unreliable test; that one rewrites many tests offline).

How do I install Test Reliability in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a claude-code`. Or copy the skill folder (skills/test-reliability in petrkindlmann/qa-skills) into .claude/skills/test-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Test Reliability in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a codex`. Or copy the skill folder (skills/test-reliability in petrkindlmann/qa-skills) into .agents/skills/test-reliability in your project. Codex loads it when a task matches its description.

Can I use Test Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-reliability, .gemini/skills/test-reliability, .github/skills/test-reliability and .opencode/skills/test-reliability in your project.

What does Test Reliability need to run?

Going by SKILL.md and its folder, Test Reliability needs the command-line tools its instructions call (npx).

Does Test Reliability access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Reliability use?

Test Reliability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Reliability use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Test Reliability?

Skills that share tags, products or a category with Test Reliability: Diagnose Playwright Failure as Product Bug (appsmithorg/appsmith, 41k stars), Playwright Regression Testing (fugazi/test-automation-skills-agents, 247 stars), Claude Code QA (PramodDutta/qaskills, 232 stars) and Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Reliability?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 165 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.