Agent skill

Diagnose Playwright Failure as Product Bug

by appsmithorg in appsmithorg/appsmith

Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

Apache-2.0Auto-check passedTesting & QA

Install Diagnose Playwright Failure as Product Bug

skills CLI
$ npx skills add appsmithorg/appsmith --skill diagnose-pw-failure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install appsmithorg/appsmith diagnose-pw-failure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/appsmithorg/appsmith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/diagnose-pw-failure .claude/skills/diagnose-pw-failure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnose-pw-failure
GitHub stars
41k
Token cost
~1.5k tokens
SKILL.md length
489 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

  • Works in 4 steps: Collect Evidence → Classify the Bug → Verify It's Not a Spec Bug → …
  • A Playwright assertion keeps failing after several spec fixes
  • SKILL.md covers When to Use, Step 1 — Collect Evidence, Step 2 — Classify the Bug and Step 3 — Verify It's Not a…, plus 2 more sections
  • Calls npx

What it does

Use it when a Playwright spec has been fixed two or three times and the same assertion keeps failing, when the selectors and waits look right but the app renders the wrong thing, or when the server returns unexpected 500 or 403 responses. It also applies once the write-and-verify or spec-fixing skills have used up their retries, and it steps aside when the error is plainly a test-code problem such as an import or selector syntax error.

Step one gathers evidence: the exact failing assertion with expected and received values, failure screenshots from the Playwright results folder, trace files that can be opened with the Playwright trace viewer, captured API responses, and optionally the server code behind the endpoint plus recent commits to that path. Step two classifies the bug as UI rendering, data regression, server error or auth and permission, each with a severity.

The output is a structured bug report with expected versus actual behavior, reproduction steps and a likely root cause, which gives developers something to act on instead of another round of spec edits.

When your agent uses it

  • A Playwright assertion keeps failing after several spec fixes
  • The test seems correct but the app shows the wrong content
  • A valid request in a test gets a 500 or 403 from the server

Example prompts

  • “This Playwright test still fails after three fixes, so is it a product bug?”
  • “Diagnose why the checkout spec gets a 403 and write up the bug report.”
  • “The test is correct but the page renders the wrong table, so collect the evidence.”

Requirements

  • A Playwright test run with its results folder
  • Access to the application's server code for deeper diagnosis

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Collect Evidence
  2. Classify the Bug
  3. Verify It's Not a Spec Bug
  4. Produce Bug Report

What it can do on your machine

Read from SKILL.md and the folder at commit fb3eb95. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diagnose Playwright Failure as Product Bug loads about 1.5k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from appsmithorg/appsmith at commit fb3eb95, republished under its Apache-2.0 licence (© appsmithorg). 489 words, ~1,482 tokens.

Download SKILL.mdSave it as .claude/skills/diagnose-pw-failure/SKILL.md (or your agent's skills folder).
name
diagnose-pw-failure
description
Diagnose a Playwright test failure as a product bug by analyzing error output, screenshots, traces, and server logs. Produces a structured bug report with expected vs actual behavior, reproduction steps, and likely root cause. Use when a Playwright test keeps failing after spec fixes, "is this a product bug", "diagnose this test failure", or "the test is correct but the app is broken".

Diagnose Playwright Failure as Product Bug

When to Use

Use this skill when:

  • A spec has been fixed 2-3 times and the same assertion keeps failing with the same received value
  • The spec's selectors and waits are correct but the app renders wrong content
  • The server returns unexpected status codes (500, 403) on valid requests
  • The write-and-verify-pw-test or fix-pw-spec skill has exhausted its retry budget

Do not use if the error is clearly a test code issue (import error, wrong selector syntax, TypeScript compilation error). Use fix-pw-spec instead.

Step 1 — Collect Evidence

Gather all available evidence:

a) Playwright error output

Read the terminal output from the failed test run. Note:

  • The exact assertion that failed
  • Expected vs received values
  • Which test step failed (line number in the spec)
b) Screenshots

Check app/client/playwright/results/ for failure screenshots:

bash
ls -la app/client/playwright/results/

If screenshots exist, read them (the Read tool supports images). They show exactly what the browser rendered at failure time.

c) Trace (if available)

Traces are at app/client/playwright/results/<test-path>/trace.zip. Note their location for the bug report — they can be viewed with npx playwright show-trace <path>.

d) Network responses

If the spec uses waitForResponse, check if the API response was captured in the error output. Common patterns:

  • API returned 200 but with wrong data → backend logic bug
  • API returned 500 → server crash
  • API returned 403 → permission/auth regression
  • API never responded (timeout) → endpoint broken or renamed
e) Server-side code (optional, for deeper diagnosis)

If the failure involves an API call, trace the server code:

  1. Identify the API endpoint from the spec (e.g., API.actionsExecute → /api/v1/actions/execute)
  2. Find the controller: grep app/server/ for @PostMapping("/api/v1/actions/execute") or similar
  3. Read the service method to understand what could go wrong
  4. Check if recent commits changed this code path
Show full SKILL.md (202 more words)Show less

Step 2 — Classify the Bug

CategorySignalsSeverity
UI renderingWrong text, missing element, broken layout (screenshot shows it)Medium
Data regressionAPI returns correct status but wrong payloadHigh
Server error500 response, stack trace in logsCritical
Auth/permission401/403 on previously working endpointHigh
Feature flagFeature works with flag on but not off (or vice versa)Medium
Deployment issueECONNREFUSED, DNS failure, unhealthy containersBlocker
Race conditionIntermittent — passes sometimes, fails othersMedium (flaky)

Step 3 — Verify It's Not a Spec Bug

Before declaring "product bug", do a sanity check:

  1. Manual verification: Does the spec's assertion make sense? Re-read the test name and expected behavior.
  2. Check the deployment manually: Navigate to PLAYWRIGHT_BASE_URL in your analysis and verify the page actually shows what the test expects.
  3. Check if the feature exists on this deployment: The deployment might be on an older version that doesn't have the feature yet.
  4. Check feature flags: If the feature is behind a flag, verify the flag is enabled on the deployment (check PW_FLAG_OVERRIDES or query /api/v1/users/features).

Step 4 — Produce Bug Report

Output a structured diagnosis:

## Playwright Failure Diagnosis: PRODUCT BUG

**Spec**: playwright/tests/sanity/widgets/table-filter.spec.ts
**Test**: "filters table by country"
**Deployment**: https://my-dp.appsmith.com
**Category**: Data regression

### Expected behavior
Filtering the table by Country "starts with Ba" should show rows including "Bangladesh".

### Actual behavior
Filter returns 0 rows. The table is empty after applying the filter.

### Evidence
- **Assertion**: `expect(table.cell(2, 0)).toContainText("Bangladesh")` — timed out, cell doesn't exist
- **Screenshot**: playwright/results/sanity-widgets-table-filter/test-failed-1.png
  - Shows table with "No data" message after filter is applied
- **API response**: GET /api/v1/actions/execute returned 200 with `{ data: [] }`
- **Spec verified correct**: Selector targets the right table, filter UI interaction works (filter chip appears)

### Likely root cause
The execute API returns empty results for the MySQL "starts with" filter. Possible causes:
- Query generation bug in the server's filter-to-SQL translation
- Datasource connection issue (DATASOURCE_HOST may be unreachable from this deployment)

### Reproduction steps
1. Open the app in deployed mode
2. Click "Add Filter" on the data_table widget
3. Set Column: Country, Condition: starts with, Value: Ba
4. Observe: table shows "No data" instead of filtered results

### Suggested investigation
- Check server logs for the execute query
- Verify DATASOURCE_HOST is reachable from the deployment
- Test the same filter on dev.appsmith.com to compare

When Diagnosis Is Inconclusive

If you can't determine whether it's a spec bug or product bug:

## Playwright Failure Diagnosis: INCONCLUSIVE

**Spec**: <path>
**Test**: <name>

### What we know
- [facts from error output]

### What we don't know
- [ambiguities]

### Recommended next steps
1. Run the test with `--debug` flag for step-by-step execution
2. Check server logs on the deployment
3. Try reproducing manually in the browser

© appsmithorg, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/diagnose-pw-failure of appsmithorg/appsmith.

Open the folder on GitHubat commit fb3eb95

Compare with similar skills

Diagnose Playwright Failure as Product Bug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diagnose Playwright Failure as Product Bug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diagnose Playwright Failure as Product Bug this skillappsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0
Debugging Opik E2E Testscomet-ml/opik22k—~1.8kAutomated safety check: PassApache-2.0
Playwright Regression Testingfugazi/test-automation-skills-agents247—~1.4kAutomated safety check: PassMIT
Debug E2E Testbitovi/ai-enablement-prompts121—~677Automated safety check: PassMIT
Claude Code QAPramodDutta/qaskills233—~2.3kAutomated safety check: PassMIT
AI Bug Triagepetrkindlmann/qa-skills168—~5.2kAutomated safety check: PassMIT

Similar skills

  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Regression Testing

    fugazi/test-automation-skills-agents

    Govern Playwright TypeScript regression suites across many tests.

    247 GitHub stars~1.4k tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • Debug E2E Test

    bitovi/ai-enablement-prompts

    Debug and fix failing Playwright E2E tests. An agent skill from bitovi/ai-enablement-prompts.

    121 GitHub stars~677 tokensUpdated 29 days ago
    Testing & QAAuto-check passed
  • Claude Code QA

    PramodDutta/qaskills

    The complete QA skill for Claude Code — turn Claude into an expert QA engineer that picks the right test type, writes reliable Playwright, Cypress, and pytest tests, eliminates flaky tests, enforces…

    233 GitHub stars~2.3k tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • AI Bug Triage

    petrkindlmann/qa-skills

    Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

    168 GitHub stars~5.2k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Test Reliability

    petrkindlmann/qa-skills

    Runtime per-test healing with evidence: multi-attribute selector healing, environment-aware diagnosis, flake classification, quarantine management, and confidence-scored auto-repair.

    168 GitHub stars~5.6k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed

More from appsmithorg/appsmith

  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check: notes
  • Fix Failing Playwright Spec

    appsmithorg/appsmith

    Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions.

    41k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Diagnose Playwright Failure as Product Bug

What does Diagnose Playwright Failure as Product Bug do?

Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report. Use it when a Playwright spec has been fixed two or three times and the same assertion keeps failing, when the selectors and waits look right but the app renders the wrong thing, or when the server returns unexpected 500 or 403 responses. It also applies once the write-and-verify or spec-fixing skills have used up their retries, and it steps aside when the error is plainly a test-code problem such as an import or selector syntax error.

When should I use Diagnose Playwright Failure as Product Bug?

Diagnose Playwright Failure as Product Bug fits situations like: A Playwright assertion keeps failing after several spec fixes; the test seems correct but the app shows the wrong content; A valid request in a test gets a 500 or 403 from the server.

How do I install Diagnose Playwright Failure as Product Bug in Claude Code?

Run `npx skills add appsmithorg/appsmith --skill diagnose-pw-failure -a claude-code`. Or copy the skill folder (.cursor/skills/diagnose-pw-failure in appsmithorg/appsmith) into .claude/skills/diagnose-pw-failure in your project. Claude Code loads it when a task matches its description.

How do I install Diagnose Playwright Failure as Product Bug in Codex?

Run `npx skills add appsmithorg/appsmith --skill diagnose-pw-failure -a codex`. Or copy the skill folder (.cursor/skills/diagnose-pw-failure in appsmithorg/appsmith) into .agents/skills/diagnose-pw-failure in your project. Codex loads it when a task matches its description.

Can I use Diagnose Playwright Failure as Product Bug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add appsmithorg/appsmith --skill diagnose-pw-failure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnose-pw-failure, .gemini/skills/diagnose-pw-failure, .github/skills/diagnose-pw-failure and .opencode/skills/diagnose-pw-failure in your project.

What does Diagnose Playwright Failure as Product Bug need to run?

Going by SKILL.md and its folder, Diagnose Playwright Failure as Product Bug needs the command-line tools its instructions call (npx). Our summary lists: A Playwright test run with its results folder; Access to the application's server code for deeper diagnosis.

Does Diagnose Playwright Failure as Product Bug access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Diagnose Playwright Failure as Product Bug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Diagnose Playwright Failure as Product Bug use?

Diagnose Playwright Failure as Product Bug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diagnose Playwright Failure as Product Bug use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diagnose Playwright Failure as Product Bug?

Skills that share tags, products or a category with Diagnose Playwright Failure as Product Bug: Debugging Opik E2E Tests (comet-ml/opik, 22k stars), Playwright Regression Testing (fugazi/test-automation-skills-agents, 247 stars), Debug E2E Test (bitovi/ai-enablement-prompts, 121 stars) and Claude Code QA (PramodDutta/qaskills, 233 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diagnose Playwright Failure as Product Bug?

appsmithorg (a GitHub organization) maintains it in appsmithorg/appsmith, which has 41,042 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 8, 2026.

Source: appsmithorg/appsmith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.