Diagnose Playwright Failure as Product Bug
appsmithorg/appsmith
Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.
Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills test-reliability --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-reliability .claude/skills/test-reliability && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .claude/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliabilityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills test-reliability --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/test-reliability .agents/skills/test-reliability && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .agents/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills test-reliability --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/test-reliability .cursor/skills/test-reliability && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .cursor/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/test-reliability--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills test-reliability --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/test-reliability .gemini/skills/test-reliability && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .gemini/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills test-reliabilityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/test-reliability .github/skills/test-reliability && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .github/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill test-reliability -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills test-reliability --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/test-reliability .opencode/skills/test-reliability && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-reliability" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/test-reliability into .opencode/skills/test-reliability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-reliability", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-reliabilityRuntime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.
Test Reliability is an agent skill from petrkindlmann/qa-skills. Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair. Goes beyond simple locator fallbacks to cover action-level healing, data healing, and observable repair workflows. Use when: "flaky test," "test stability," "self-healing locator," "broken locator recovery," "unreliable test," "quarantine flaky test." Not for: bulk regenerating selectors after a planned UI refactor — use…
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/flaky-test-runbook.md`).
It sits in Testing & QA, covering Failing and flaky tests, Issue triage and QA and bug reports. It works with Playwright. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Reliability loads about 5.6k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 202 tokens; SKILL.md has 2,275 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,275 words, ~5,554 tokens.
.claude/skills/test-reliability/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.<objective>
A retried test that still flakes will eventually fail 3-of-3 during your most critical release, and a silently auto-repaired test may now verify a different element entirely. This skill builds suites teams can trust: resilient locators, classified flakes, environment-aware healing, data healing, and observable repair with confidence scoring — every automated fix produces evidence a human can review.
</objective>
| Situation | Go to |
|---|---|
| Test flakes in CI, root cause unknown | Flake Classification → decision tree |
| Locator broke, want resilient replacement | Locator Resilience |
| Action fails, suspect slow backend not UI bug | Environment-Aware Healing |
| 401/404/429 mid-test from stale data | Data Healing |
| Want auto-repair with review gate | Observable Repair Workflow |
| Isolate a flaky test without blocking CI | Quarantine Management |
| Step-by-step triage of one flaky test | references/flaky-test-runbook.md |
Check .agents/qa-project-context.md first — it carries known flaky areas, selector strategy, and CI environment details. Skip any question it already answers.
Prevention over cure. Writing a resilient test costs 1x. Investigating a flaky one costs 10x. Losing team trust in the suite costs 100x.
Healing must be observable and reviewable. Every automated repair produces evidence: what broke, what was tried, what worked, the confidence score. Silent fixes erode trust as fast as silent failures.
Classify before fixing. The fix for a timing issue is completely different from the fix for a data dependency. Wrong diagnosis wastes effort and can make things worse.
Flaky tests are bugs. Not annoyances to tolerate. A flaky test either has a test bug (fix the test), reveals an app bug (fix the app), or exposes an environment issue (fix the environment).
Track reliability as a metric, not a feeling. Measure flaky rate, mean time to heal, quarantine age, and selector stability. What gets measured gets fixed.
Self-healing is a spectrum. Start with resilient locators (Level 1), add fallback strategies (Level 2), then environment-aware healing (Level 3), then confidence-scored auto-repair (Level 4). Do not jump to Level 4 before mastering Level 1.
A single locator strategy is a single point of failure. Multi-attribute selectors combine multiple signals for one element lookup — resilience without fallback-chain complexity.
The key insight: instead of "try A, then B, then C," use "find element matching A AND B AND C with tolerance for one signal missing."
// Multi-attribute locator: tries combinations from most specific to least
const submitBtn = await multiAttributeLocator(page, {
testId: 'checkout-submit', // most stable signal
role: 'button', // semantic signal
name: /place order/i, // accessible name
nearText: 'Order Summary', // visual context
});
// Internally: tries testId+role+name first, then testId alone, then role+name,
// then text, then nearText+role. Returns first visible match.
// Unlike fallback chains, it combines signals for higher confidence.When a locator fails, the element may still exist with changed attributes. Use surrounding DOM context to find it:
These are repair candidates scored by the confidence system below — not runtime fallbacks.
Rate every selector on a 0-5 scale to prioritize refactoring.
| Score | Strategy | Survives |
|---|---|---|
| 5 | getByTestId('submit-order') | CSS, text, and structural changes |
| 4 | getByRole('button', { name: 'Submit' }) | CSS and structural changes |
| 3 | getByLabel('Email') | CSS changes; breaks on label rewording |
| 2 | getByText('Submit Order') | Breaks on any copy change |
| 1 | locator('.btn-primary.submit') | Breaks on CSS or structural change |
| 0 | locator('//div[3]/button[1]') | Breaks on any DOM change |
Target: Average score of 3.5+ across the suite. Audit monthly. Prioritize fixing score-0 and score-1 selectors. Emit one score per locator to selector-stability.md (or a CI step) and report the suite average so the 3.5 target is verifiable, not asserted.
Every flaky test has a root cause category. Classifying correctly determines the fix.
| Category | Signal | Root Cause | Fix Direction |
|---|---|---|---|
| Timing | Timeout errors, passes on retry, worse in CI | Race condition, animation, async operation | Wait for condition, not time |
| Data dependency | Fails with other tests, passes alone | Shared state, missing cleanup | Isolate per-test, fixture cleanup |
| Environment | Fails on specific runner, correlates with load | Resource contention, network latency | Mock externals, increase resources |
| Order dependency | Fails with --shard or fullyParallel | Depends on another test's side effect | Self-contained setup |
| Time sensitivity | Fails at specific times (midnight, month-end) | Uses real clock, date boundary | Mock clock, relative comparisons |
| Visual rendering | Screenshot diff flickers, subpixel differences | Font rendering, antialiasing, animation frame | Increase threshold, mask dynamic regions |
| External service | Correlates with third-party status | Real HTTP calls in tests | Mock external APIs |
Test is flaky
│
├── Does it pass when run alone?
│ ├── YES → ORDER DEPENDENCY or DATA DEPENDENCY
│ │ ├── Does another test create/modify data it needs? → ORDER DEPENDENCY
│ │ └── Does it share a database/file/cache? → DATA DEPENDENCY
│ │
│ └── NO → Not order/data dependent. Continue below.
│
├── Does it fail more often in CI than locally?
│ ├── YES → TIMING or ENVIRONMENT
│ │ ├── Timeout errors? → TIMING (CI is slower)
│ │ ├── Connection errors? → ENVIRONMENT (network latency / service)
│ │ └── Resource errors (OOM, disk)? → ENVIRONMENT (resource contention)
│ │
│ └── NO → Same rate locally and CI. Continue below.
│
├── Does it fail at specific times?
│ ├── YES → TIME SENSITIVITY
│ │ ├── Near midnight? → Date boundary issue
│ │ ├── Near month/year end? → Calendar calculation
│ │ └── Specific hour? → Timezone issue
│ │
│ └── NO → Continue below.
│
├── Does it involve screenshots or visual comparison?
│ ├── YES → VISUAL RENDERING
│ │
│ └── NO → Continue below.
│
├── Does it call external HTTP APIs?
│ ├── YES → EXTERNAL SERVICE
│ │
│ └── NO → TIMING (most likely — default classification)
│ └── Investigate: what async operation is not being awaited?The fix per category lives in references/flaky-test-runbook.md (Step 4) with full code patterns.
Not all test failures are test problems. Some are environment problems. Environment-aware healing distinguishes the two and adapts.
When an action fails, check backend health before blaming the test:
/api/health.backend_down. This is not a UI bug.Return a structured diagnosis: { success: boolean; diagnosis: 'backend_down' | 'ui_failure' | 'backend_slow_recovered' }. This feeds flake classification — backend issues are environment issues, not test bugs.
Before declaring a test failure in CI, check for resource contention:
about:blank. If it takes > 2s (baseline < 500ms), the runner is overloaded./api/health. If it takes > 5s (baseline < 1s), the backend is under pressure.Test data expires, gets cleaned up, or becomes invalid. Data healing detects and regenerates stale test data.
| Pattern | Signal | Fix |
|---|---|---|
| Expired auth token | 401 response during test | Regenerate token in fixture |
| Deleted test record | 404 when accessing seeded data | Re-seed before test |
| Uniqueness violation | 409 or constraint error | Generate unique identifiers per run |
| Stale cache | Wrong data returned | Clear cache in setup |
| Exceeded quota | 429 or rate limit error | Reset quotas or use dedicated test account |
Build fixtures that verify data exists and regenerate if stale.
// Pattern: verify → heal → use → cleanup
testUser: async ({ request }, use, testInfo) => {
// 1. Try to find existing test user by deterministic email
// 2. Verify auth token is still valid (GET /api/me)
// 3. If token expired → refresh it (POST /refresh-token), mark as healed
// 4. If user missing → create new one, mark as healed
// 5. If healed → annotate testInfo for observability
// 6. use(user) → run the test
// 7. Cleanup: delete test user (guaranteed by fixture, even on failure)
}Key patterns:
testInfo.testId in email/identifiers for per-test uniqueness.testInfo.annotations when healing occurs, for observability.afterEach — fixtures guarantee cleanup on failure.Core guardrail: Healing must be observable and reviewable. Every repair follows this flow:
Failure Detected
│
▼
Candidate Repair Generated
│
▼
Confidence Score Computed (0.0 - 1.0)
│
▼
Evidence Diff Produced (what changed, what was tried)
│
▼
Approval Policy Applied
│ ├── Score >= 0.9 → Auto-apply, log for batch review
│ ├── Score 0.7-0.89 → Apply in quarantine, flag for individual review
│ ├── Score 0.5-0.69 → Do NOT apply, open PR with evidence for review
│ └── Score < 0.5 → Discard, manual investigation required
│
▼
Intent Fidelity Check (does repaired test still test the same thing?)
│
▼
Rollback if intent fidelity dropsScore each repair candidate on six dimensions (weighted sum, 0.0-1.0):
| Dimension | Weight | Scoring |
|---|---|---|
| Match specificity | 0.30 | testId=1.0, role=0.9, text=0.7, context=0.5, CSS=0.3 |
| Element visible | 0.15 | 1.0 if visible, 0.0 if not |
| Same parent container | 0.15 | 1.0 if same container, 0.0 if different |
| Same element type | 0.15 | 1.0 if same tag+role, 0.0 if different |
| Text similarity | 0.15 | 0.0-1.0 (Levenshtein ratio of accessible name) |
| Attribute overlap | 0.10 | 0.0-1.0 (Jaccard of shared attributes) |
Score thresholds (half-open bands — 0.9 belongs to the auto-apply tier only):
references/flaky-test-runbook.md (Score Interpretation) uses these exact four bands — keep all three locations identical if you edit them.
Every repair produces an evidence record containing: test file, test name, failure type, original locator, candidate replacements (each with confidence, evidence string, and intent-preserved flag), which candidate was selected, timestamp, approval path, rollback trigger, and a screencast.webm recording of the repair run (Playwright 1.59+ page.screencast with showActions annotations — the "agentic video receipt"). Without the screencast, a Level-3/4 healed test asks reviewers to trust a JSON record; with it, the diff and the runtime are both inspectable.
Hand Levels 3-4 to a vendor when you don't want to own a confidence-scored healer; build in-house when you need on-prem/air-gapped deployment, an explicit AI-prompt audit trail, or quarantine logic your tracker can't express. Building stops being worth it once a hosted tool already fingerprints + clusters + auto-quarantines for you.
After applying a repair, verify the test still exercises the same user intent:
button → a) — intent NOT preserved, rollback.button → link) — intent NOT preserved, rollback.A repair that changes WHAT the test verifies (not just HOW it finds elements) must be rolled back.
Quarantine isolates flaky tests so they run but do not block CI. See references/flaky-test-runbook.md (Step 6) for the full config and CI wiring; the essentials:
// Tag the flaky test
test('intermittent WebSocket reconnect', {
tag: ['@quarantine'],
annotation: {
type: 'quarantine',
description: 'Flaky since 2026-03-15. Race condition in WebSocket handler. Ticket: BUG-1234.',
},
}, async ({ page }) => { /* ... */ });// playwright.config.ts — separate projects
projects: [
{ name: 'stable', testMatch: /.*\.spec\.ts/, grep: /^(?!.*@quarantine)/ }, // exclude quarantine
{ name: 'quarantine', grep: /@quarantine/, retries: 3 },
],In CI, run --project=stable as a blocking step and --project=quarantine with continue-on-error: true so the quarantine project never blocks the pipeline.
1. DETECT — Test identified as flaky (CI reporter or manual triage)
2. TAG — Add @quarantine annotation with ticket link and date
3. ISOLATE — Quarantine project runs separately, does not block
4. DIAGNOSE — Follow the flaky test runbook (references/flaky-test-runbook.md)
5. FIX — Apply the fix pattern for the classified category
6. VERIFY — Run 50x with --repeat-each, zero failures required
7. RELEASE — Remove @quarantine tag, add annotation documenting the fixReplacing a broken selector with no logging, review, or confidence scoring. The repaired test may now verify a different element entirely. Every repair must produce evidence.
Retries are a detection mechanism, not a fix. A test that needs retry 2-of-3 will eventually fail 3-of-3 during your most critical release.
test.skip('flaky, will fix later') — "later" never comes. Either quarantine with tracking or delete entirely. Skipped tests with no ticket are dead code.
Timing issues and data dependencies need completely different fixes. Adding waitForTimeout(5000) to a data-dependency problem makes the test slower and still flaky.
// NEVER the right fix
await page.waitForTimeout(5000);
// Wait for the actual condition
await expect(page.getByRole('table')).toBeVisible();
await page.waitForResponse(resp => resp.url().includes('/api/data') && resp.status() === 200);Auto-repair that produces no logs, evidence, or confidence scores. You cannot improve what you cannot measure, and you cannot trust what you cannot review.
Building a complex self-healing framework before adopting basic resilient-locator patterns. Start with multi-attribute selectors and proper waits. Add healing infrastructure only when data shows where breakage occurs.
Tests sit in quarantine for months. Quarantine is a temporary state, not a permanent home. Enforce a 14-day maximum.
| Symptom | Likely cause | Fix or check |
|---|---|---|
Backend health check itself flaps → false backend_down diagnosis | Health endpoint is itself flaky/slow | Track the health endpoint's own p99 separately; don't gate diagnosis on a single probe |
| Artifact storage balloons after enabling repair video | page.screencast recording on every run, not just repairs | Record only on the repair path, not the happy path |
| Auto-repair accuracy drops below 80% | Confidence threshold too low, or intent-fidelity check skipped | Raise the auto-apply floor; never skip the intent check |
| Quarantine project blocks the pipeline | Missing continue-on-error on the quarantine CI step | Add it; the quarantine project must never block merges |
npx playwright test <spec> --repeat-each=20 --workers=4 --trace=on — it must fail at least once before you trust any fix.npx playwright test <spec> --repeat-each=50 --workers=4 — require 50/50 passes, in CI conditions too.npx playwright test --project=stable excludes @quarantine tests and --project=quarantine runs only them.selector-stability.md (or the CI report) lists a stability score per locator and reports a suite average >= 3.5.© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/test-reliability of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
Test Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Reliability this skillpetrkindlmann/qa-skills | 165 | — | ~5.6k | Automated safety check: Pass | MIT | |
| Diagnose Playwright Failure as Product Bugappsmithorg/appsmith | 41k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Playwright Regression Testingfugazi/test-automation-skills-agents | 247 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Claude Code QAPramodDutta/qaskills | 232 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Cucumber and Playwright E2E Testslanggenius/dify | 158k | — | ~682 | Automated safety check: Pass | Custom licence | |
| Fix Failing Playwright Specappsmithorg/appsmith | 41k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
appsmithorg/appsmith
Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.
fugazi/test-automation-skills-agents
Govern Playwright TypeScript regression suites across many tests.
PramodDutta/qaskills
The complete QA skill for Claude Code — turn Claude into an expert QA engineer that picks the right test type, writes reliable Playwright, Cypress, and pytest tests, eliminates flaky tests, enforces…
langgenius/dify
Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.
appsmithorg/appsmith
Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions.
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Works with
Categories
Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair. Test Reliability is an agent skill from petrkindlmann/qa-skills. Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.
Test Reliability fits situations like: self-healing locator; broken locator recovery; unreliable test; that one rewrites many tests offline).
Run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a claude-code`. Or copy the skill folder (skills/test-reliability in petrkindlmann/qa-skills) into .claude/skills/test-reliability in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a codex`. Or copy the skill folder (skills/test-reliability in petrkindlmann/qa-skills) into .agents/skills/test-reliability in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill test-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-reliability, .gemini/skills/test-reliability, .github/skills/test-reliability and .opencode/skills/test-reliability in your project.
Going by SKILL.md and its folder, Test Reliability needs the command-line tools its instructions call (npx).
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Reliability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Test Reliability: Diagnose Playwright Failure as Product Bug (appsmithorg/appsmith, 41k stars), Playwright Regression Testing (fugazi/test-automation-skills-agents, 247 stars), Claude Code QA (PramodDutta/qaskills, 232 stars) and Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 165 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.