Agent skill

Debug Playwright

by quay in quay/quay

Debug Playwright E2E test failures from GitHub Actions CI runs.

Apache-2.0Auto-check passedTesting & QA

Install Debug Playwright

skills CLI
$ npx skills add quay/quay --skill debug-playwright -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install quay/quay debug-playwright --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/quay/quay.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debug-playwright .claude/skills/debug-playwright && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-playwright
GitHub stars
2.8k
Token cost
~1.2k tokens
SKILL.md length
516 words
Files
2 (incl. scripts)
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Debug Playwright E2E test failures from GitHub Actions CI runs.

  • Works in 5 steps: Fetch and Categorize → Report Overview → Diagnose Each Real Failure → …
  • Tasks that involve Browser testing
  • SKILL.md covers Step 1: Fetch and Categorize, Step 2: Report Overview, Step 3: Diagnose Each Real… and Step 4: Classify and Explain, plus 2 more sections
  • Runs Shell scripts from its folder; calls bash and jq

What it does

Debug Playwright is an agent skill from quay/quay. Debug Playwright E2E test failures from GitHub Actions CI runs. Downloads artifacts, categorizes failures (flaky/real/infra), correlates with backend logs and Jaeger traces, and offers fixes.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/playwright-debug.sh`).

It sits in Testing & QA, covering Browser testing, Failing and flaky tests and End-to-end testing. It works with Playwright and GitHub Actions. The repository describes itself as: Build, Store, and Distribute your Applications and Containers. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Browser testing
  • Tasks that involve Failing and flaky tests
  • Tasks that involve End-to-end testing

Example prompts

  • “/debug-playwright”

Requirements

  • Python 3
  • A Bash shell
  • Pre-approved tools (allowed-tools): Bash(bash scripts/playwright-debug.sh *), Bash(gh run view *), Bash(gh run download *), Bash(gh pr view *), Read, Grep, Edit, AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Fetch and Categorize
  2. Report Overview
  3. Diagnose Each Real Failure
  4. Classify and Explain
  5. Offer Fixes

What it can do on your machine

Read from SKILL.md and the folder at commit 8759f55. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(bash scripts/playwright-debug.sh *)
    • Bash(gh run view *)
    • Bash(gh run download *)
    • Bash(gh pr view *)
    • Read
    • Grep
    • Edit
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Playwright loads about 1.2k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 516 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from quay/quay at commit 8759f55, republished under its Apache-2.0 licence (© quay). 516 words, ~1,174 tokens.

Download SKILL.mdSave it as .claude/skills/debug-playwright/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
debug-playwright
description
Debug Playwright E2E test failures from GitHub Actions CI runs. Downloads artifacts, categorizes failures (flaky/real/infra), correlates with backend logs and Jaeger traces, and offers fixes.
allowed-tools
Bash(bash scripts/playwright-debug.sh *), Bash(gh run view *), Bash(gh run download *), Bash(gh pr view *), Read, Grep, Edit, AskUserQuestion
argument-hint
PR_NUMBER | RUN_URL

Debug Playwright CI Failures

Debug Playwright test failures for $ARGUMENTS.

Step 1: Fetch and Categorize

bash
bash scripts/playwright-debug.sh $ARGUMENTS

Capture the full JSON output, then extract artifacts_dir for use in later commands:

bash
PW_JSON=$(bash scripts/playwright-debug.sh $ARGUMENTS)
ARTIFACTS_DIR=$(echo "$PW_JSON" | jq -r '.artifacts_dir')

Key fields:

  • artifacts_dir — temp directory with downloaded artifacts
  • failed — tests that failed on all attempts (real failures), includes error_message, error_stack, last_step, and attachments
  • flaky — tests that failed then passed on retry
  • interrupted — tests where a worker crashed
  • stats — overall run statistics
  • surge_url — link to the HTML report
  • has_container_logs / has_jaeger_traces — what extra data is available
  • global_setup_failure — if true, no tests ran at all (check errors field)

If exit code is 2, the run is still in progress — tell the user to wait or use /dev:poll.

Step 2: Report Overview

Summarize what happened conversationally:

  • Total tests, pass/fail/flaky counts
  • Link to the Surge HTML report (if available)
  • List flaky tests briefly (name + file) — note them but don't deep-dive unless asked
  • Note any interrupted tests (worker crashes)

If global_setup_failure is true, report the setup errors and stop.

If there are no real failures, report "all failures were flaky" with the list and stop.

Step 3: Diagnose Each Real Failure

For each entry in failed, perform root cause analysis:

3a: Read the test source

Read the failing spec file at the reported line number. The file path from the JSON is relative to web/playwright/e2e/ — resolve it against the quay/quay repo root (e.g., auth/signin.spec.ts → web/playwright/e2e/auth/signin.spec.ts).

Understand what the test does — what page it navigates to, what selectors it uses, what API calls it makes.

Check last_step from the JSON — this tells you which Playwright action timed out (e.g., locator.click("button.submit")).

3b: Correlate with container logs

If has_container_logs is true, search for backend errors within a ~30-second window around the test's startTime:

bash
grep -n "Traceback\|Internal Server Error\|FATAL" \
  "$ARTIFACTS_DIR/quay-container-logs/quay-quay.log" | head -30

Look for Python tracebacks, 500 responses, or gunicorn crashes that coincide with the test failure.

Show full SKILL.md (219 more words)Show less
3c: Correlate with Jaeger traces

If has_jaeger_traces is true, infer the API endpoint from the test source (e.g., a test navigating to /organization/myorg/teams likely hits GET /api/v1/organization/.*/team). Search for matching spans:

bash
jq '[.data[].spans[] | select(.operationName | test("ENDPOINT_PATTERN")) |
  {operationName, duration, tags: [.tags[] |
  select(.key == "http.status_code" or .key == "error")]}]' \
  "$ARTIFACTS_DIR/jaeger-traces/traces.json"

Look for slow spans, error spans, and 4xx/5xx status codes.

3d: Determine auth phase

Check the test's tags for auth:OIDC or auth:LDAP. Tests without auth-specific tags run in the DB auth phase (the first phase).

Step 4: Classify and Explain

For each failure, classify the root cause and explain conversationally:

  • Selector change — element not found but backend responded fine
  • Backend error — 500/traceback in container logs or error spans in traces
  • Timing/race — intermittent, slow spans in traces, or missing waitFor
  • Auth/config — failure only in one auth phase, related to auth swap
  • Test isolation — leftover state from prior tests causing interference
  • Infra — browser crash, connection refused, worker timeout

For each one, state what the test was trying to do, what went wrong, what the backend logs/traces show, and what a fix would look like.

Step 5: Offer Fixes

Ask: "Want me to apply fixes for any of these?"

If yes, edit the spec files under web/playwright/e2e/. Show what you're changing and why. Only edit backend code if the user explicitly asks.

Do NOT auto-commit — let the user review the changes.

Cleanup

When diagnosis is complete, remove the temp artifacts directory:

bash
rm -rf "$ARTIFACTS_DIR"

© quay, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .agents/skills/debug-playwright of quay/quay.

  • SKILL.md
  • scripts/playwright-debug.sh

Open the folder on GitHubat commit 8759f55

Compare with similar skills

Debug Playwright next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Playwright compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Playwright this skillquay/quay2.8k—~1.2kAutomated safety check: PassApache-2.0
Web Testing with Playwright and Vitestwithkynam/vibecode-pro-max-kit1.1k—~892Automated safety check: PassApache-2.0
Playwright E2Eforcedotcom/salesforcedx-vscode1k—~3kAutomated safety check: PassBSD-3-Clause
E2E TestingaAAaqwq/AGI-Super-Team1056 repos~2kAutomated safety check: PassMIT
Playwright Testmizchi/skills356—~4.7kAutomated safety check: NotesNone
E2E Testingxu-xiang/everything-claude-code-zh2k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Web Testing with Playwright and Vitest

    withkynam/vibecode-pro-max-kit

    Covers web testing from unit to E2E, load, visual, accessibility and security checks, with Playwright, Vitest and k6 guides plus a Playwright setup script.

    1.1k GitHub stars~892 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Playwright E2E

    forcedotcom/salesforcedx-vscode

    writing, running, and debugging Playwright tests; creating and recreating scratch orgs (Dreamhouse, minimal, non-tracking); working with their output from github actions

    1k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Testing

    aAAaqwq/AGI-Super-Team

    Playwright E2E testing patterns, Page Object Model, configuration, CI/CD integration, artifact management, and flaky test strategies.

    105 GitHub starsUsed in 6 repos~2k tokens
    Testing & QAAuto-check passed
  • Playwright Test

    mizchi/skills

    Best practices and reference for Playwright Test (E2E). An agent skill from mizchi/skills.

    356 GitHub stars~4.7k tokensUpdated 5 days ago
    Testing & QAAuto-check: notes
  • E2E Testing

    xu-xiang/everything-claude-code-zh

    Playwright E2E 测试模式、页面对象模型(POM)、配置、CI/CD 集成、产物管理以及不稳定测试(flaky test)策略。

    2k GitHub stars~1.9k tokensUpdated 7 mo ago
    Testing & QAAuto-check passed
  • E2E Testing

    xu-xiang/everything-claude-code-zh

    Playwright E2E 测试模式、页面对象模型 (Page Object Model)、配置、CI/CD 集成、产物管理以及不稳定测试 (Flaky Test) 策略。

    2k GitHub stars~2k tokensUpdated 7 mo ago
    Testing & QAAuto-check passed

More from quay/quay

  • Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

    2.8k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Diagnose any Quay Prow job failure end to end: prowjob.json - top-level build log - JUnit - resolved failing step - Playwright results.json when the failing step is Playwright, continuing through…

    2.8k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Triage a flaky Playwright test end to end, from a Sippy signal to a written fix proposal: Sippy numbers and failing run URLs, Prow artifacts (or the access gap), the spec, a local reproduction, and…

    2.8k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Pilot Update

    quay/quay

    Post a biweekly Agentic SDLC pilot update comment to PROJQUAY-11352.

    2.8k GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Debug Playwright

What does Debug Playwright do?

Debug Playwright E2E test failures from GitHub Actions CI runs. Debug Playwright is an agent skill from quay/quay. Debug Playwright E2E test failures from GitHub Actions CI runs.

When should I use Debug Playwright?

Debug Playwright fits situations like: tasks that involve Browser testing; tasks that involve Failing and flaky tests; tasks that involve End-to-end testing.

How do I install Debug Playwright in Claude Code?

Run `npx skills add quay/quay --skill debug-playwright -a claude-code`. Or copy the skill folder (.agents/skills/debug-playwright in quay/quay) into .claude/skills/debug-playwright in your project. Claude Code loads it when a task matches its description.

How do I install Debug Playwright in Codex?

Run `npx skills add quay/quay --skill debug-playwright -a codex`. Or copy the skill folder (.agents/skills/debug-playwright in quay/quay) into .agents/skills/debug-playwright in your project. Codex loads it when a task matches its description.

Can I use Debug Playwright in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add quay/quay --skill debug-playwright -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-playwright, .gemini/skills/debug-playwright, .github/skills/debug-playwright and .opencode/skills/debug-playwright in your project.

What does Debug Playwright need to run?

Going by SKILL.md and its folder, Debug Playwright needs a shell for the scripts in its folder and the command-line tools its instructions call (bash and jq). Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Bash(bash scripts/playwright-debug.sh *), Bash(gh run view *), Bash(gh run download *), Bash(gh pr view *), Read, Grep, Edit, AskUserQuestion.

Does Debug Playwright access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Playwright safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Debug Playwright use?

Debug Playwright is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Playwright use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Playwright?

Skills that share tags, products or a category with Debug Playwright: Web Testing with Playwright and Vitest (withkynam/vibecode-pro-max-kit, 1.1k stars), Playwright E2E (forcedotcom/salesforcedx-vscode, 1k stars), E2E Testing (aAAaqwq/AGI-Super-Team, 105 stars) and Playwright Test (mizchi/skills, 356 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Playwright?

quay (a GitHub organization) maintains it in quay/quay, which has 2,827 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: quay/quay on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.