Official agent skill

Web Application Testing

by anthropics in anthropics/skills

Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

OfficialApache-2.0Auto-check passedTesting & QA

Install Web Application Testing

skills CLI
$ npx skills add anthropics/skills --skill webapp-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anthropics/skills webapp-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anthropics/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/webapp-testing .claude/skills/webapp-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
webapp-testing
GitHub stars
180k
Used in
51 other repos
Token cost
~966 tokens
SKILL.md length
255 words
Files
6 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

  • Works in 3 steps: Inspect rendered DOM → Identify selectors from inspection results → Execute actions using discovered selectors
  • Verifying that a frontend change works in a real browser
  • SKILL.md covers Decision Tree: Choosing Your…, Example: Using with_server.py, Reconnaissance-Then-Action… and Common Pitfall, plus 2 more sections
  • Runs Python scripts from its folder; calls python and npm

What it does

This toolkit has the agent write native Python Playwright scripts to exercise a web app running on your machine. A decision tree picks the approach: static HTML can be read directly for selectors, while dynamic apps need a running server, inspection of the rendered page and then actions on the selectors found there.

A helper, scripts/with_server.py, starts one or more servers (for example a backend and a frontend), waits for their ports and then runs the automation, so the script itself contains only Playwright logic. The agent is told to wait for network idle before inspecting dynamic pages, use descriptive selectors, close the browser when done, and call the helpers with --help as black boxes instead of reading their source. Example scripts show console logging, element discovery and static HTML automation.

When your agent uses it

  • Verifying that a frontend change works in a real browser
  • Debugging UI behavior with screenshots and console logs
  • Testing an app that needs both a backend and a frontend server running

Example prompts

  • “Start the dev server and check that the login form shows an error for a wrong password.”
  • “Take full-page screenshots of the dashboard at mobile and desktop widths.”
  • “Click through checkout on localhost and tell me which step throws a console error.”

Requirements

  • Python with Playwright
  • The app's start commands, such as npm run dev

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Inspect rendered DOM
  2. Identify selectors from inspection results
  3. Execute actions using discovered selectors

What it can do on your machine

Read from SKILL.md and the folder at commit 683bc88. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Application Testing loads about 966 tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 255 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~966

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from anthropics/skills at commit 683bc88, republished under its Apache-2.0 licence (© anthropics). 255 words, ~966 tokens.

Download SKILL.mdSave it as .claude/skills/webapp-testing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
webapp-testing
description
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
license
Complete terms in LICENSE.txt

Web Application Testing

To test local web applications, write native Python Playwright scripts.

Helper Scripts Available:

  • scripts/with_server.py - Manages server lifecycle (supports multiple servers)

Always run scripts with --help first to see usage. DO NOT read the source until you try running the script first and find that a customized solution is abslutely necessary. These scripts can be very large and thus pollute your context window. They exist to be called directly as black-box scripts rather than ingested into your context window.

Decision Tree: Choosing Your Approach

User task → Is it static HTML?
    ├─ Yes → Read HTML file directly to identify selectors
    │         ├─ Success → Write Playwright script using selectors
    │         └─ Fails/Incomplete → Treat as dynamic (below)
    │
    └─ No (dynamic webapp) → Is the server already running?
        ├─ No → Run: python scripts/with_server.py --help
        │        Then use the helper + write simplified Playwright script
        │
        └─ Yes → Reconnaissance-then-action:
            1. Navigate and wait for networkidle
            2. Take screenshot or inspect DOM
            3. Identify selectors from rendered state
            4. Execute actions with discovered selectors

Example: Using with_server.py

To start a server, run --help first, then use the helper:

Single server:

bash
python scripts/with_server.py --server "npm run dev" --port 5173 -- python your_automation.py

Multiple servers (e.g., backend + frontend):

bash
python scripts/with_server.py \
  --server "cd backend && python server.py" --port 3000 \
  --server "cd frontend && npm run dev" --port 5173 \
  -- python your_automation.py

To create an automation script, include only Playwright logic (servers are managed automatically):

python
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True) # Always launch chromium in headless mode
    page = browser.new_page()
    page.goto('http://localhost:5173') # Server already running and ready
    page.wait_for_load_state('networkidle') # CRITICAL: Wait for JS to execute
    # ... your automation logic
    browser.close()

Reconnaissance-Then-Action Pattern

  1. Inspect rendered DOM:

    python
    page.screenshot(path='/tmp/inspect.png', full_page=True)
    content = page.content()
    page.locator('button').all()
  2. Identify selectors from inspection results

  3. Execute actions using discovered selectors

Common Pitfall

❌ Don't inspect the DOM before waiting for networkidle on dynamic apps ✅ Do wait for page.wait_for_load_state('networkidle') before inspection

Best Practices

  • Use bundled scripts as black boxes - To accomplish a task, consider whether one of the scripts available in scripts/ can help. These scripts handle common, complex workflows reliably without cluttering the context window. Use --help to see usage, then invoke directly.
  • Use sync_playwright() for synchronous scripts
  • Always close the browser when done
  • Use descriptive selectors: text=, role=, CSS selectors, or IDs
  • Add appropriate waits: page.wait_for_selector() or page.wait_for_timeout()

Reference Files

  • examples/ - Examples showing common patterns:
    • element_discovery.py - Discovering buttons, links, and inputs on a page
    • static_html_automation.py - Using file:// URLs for local HTML
    • console_logging.py - Capturing console logs during automation

© anthropics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts) in skills/webapp-testing of anthropics/skills.

  • SKILL.md
  • LICENSE.txt
  • examples/console_logging.py
  • examples/element_discovery.py
  • examples/static_html_automation.py
  • scripts/with_server.py

Open the folder on GitHubat commit 683bc88

Used in at least 33 other repositories

We found 87 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 51 other GitHub owners. This page covers the copy in anthropics/skills, which our catalogue first saw on October 7, 2026.

…and 37 more copies not listed here.

Compare with similar skills

Web Application Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Application Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Application Testing this skillanthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
RStudio Selenium to Playwright Migrationrstudio/rstudio5.1k—~3.6kAutomated safety check: PassCustom licence
Python Testingmacalbert/envilder138—~3.1kAutomated safety check: PassMIT
Webapp TestingxenitV1/Antigravity-Workflows130—~947Automated safety check: NotesMIT
Playwrightsecondsky/claude-skills227—~3.7kAutomated safety check: NotesMIT
Acarshub Tool Additionssdr-enthusiasts/docker-acarshub117—~707Automated safety check: PassGPL-3.0

Similar skills

  • Converts RStudio Python Selenium electron tests into TypeScript Playwright tests, checking each against a live RStudio before counting it as migrated.

    5.1k GitHub stars~3.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Python Testing

    macalbert/envilder

    Mandatory testing conventions including AAA pattern, test naming, assertions, and mocks.

    138 GitHub stars~3.1k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Webapp Testing

    xenitV1/Antigravity-Workflows

    Web application testing principles. An agent skill from xenitV1/Antigravity-Workflows.

    130 GitHub stars~947 tokensUpdated 8 mo ago
    Testing & QAAuto-check: notes
  • Playwright

    secondsky/claude-skills

    Browser automation and E2E testing with Playwright. An agent skill from secondsky/claude-skills.

    227 GitHub stars~3.7k tokensUpdated 9 days ago
    Testing & QAAuto-check: notes
  • Acarshub Tool Additions

    sdr-enthusiasts/docker-acarshub

    Use ONLY when working in the docker-acarshub repository AND a task may require adding a system tool, npm package, or other dependency.

    117 GitHub stars~707 tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • E2E Testing

    ericrisco/rsc-harness

    A skill your agent uses when writing or stabilizing Playwright tests that drive a real browser through multi-step journeys — durable locators, web-first assertions, storageState auth, trace/retries…

    156 GitHub stars~3.2k tokensUpdated today
    Testing & QAAuto-check passed

More from anthropics/skills

All 16 skills in this repo
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Auto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Auto-check passed
  • Web Artifacts Builder

    anthropics/skills

    Official

    Builds multi-component claude.ai HTML artifacts as a small React, TypeScript and Tailwind project, then bundles it into one shareable HTML file.

    180k GitHub starsUsed in 40 repos~769 tokens
    Auto-check passed
  • Official

    Creates original generative art in two steps: a written algorithmic philosophy, then a p5.js sketch with seeded randomness and an interactive viewer for exploring parameters.

    180k GitHub starsUsed in 37 repos~4.9k tokens
    Auto-check passed
  • Canvas Design

    anthropics/skills

    Official

    Creates original posters and static art as PNG or PDF by first writing a short design philosophy, then expressing it visually on a canvas.

    180k GitHub starsUsed in 51 repos~3k tokens
    Auto-check passed
  • DOCX Creation and Editing

    anthropics/skills

    Official

    Creates, edits and reviews Word documents: new files with docx-js, edits through the underlying XML, plus tracked changes, comments and conversions.

    180k GitHub starsUsed in 6 repos~1.7k tokens
    Auto-check passed

Questions about Web Application Testing

What does Web Application Testing do?

Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs. This toolkit has the agent write native Python Playwright scripts to exercise a web app running on your machine. A decision tree picks the approach: static HTML can be read directly for selectors, while dynamic apps need a running server, inspection of the rendered page and then actions on the selectors found there.

When should I use Web Application Testing?

Web Application Testing fits situations like: verifying that a frontend change works in a real browser; debugging UI behavior with screenshots and console logs; testing an app that needs both a backend and a frontend server running.

How do I install Web Application Testing in Claude Code?

Run `npx skills add anthropics/skills --skill webapp-testing -a claude-code`. Or copy the skill folder (skills/webapp-testing in anthropics/skills) into .claude/skills/webapp-testing in your project. Claude Code loads it when a task matches its description.

How do I install Web Application Testing in Codex?

Run `npx skills add anthropics/skills --skill webapp-testing -a codex`. Or copy the skill folder (skills/webapp-testing in anthropics/skills) into .agents/skills/webapp-testing in your project. Codex loads it when a task matches its description.

Can I use Web Application Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/skills --skill webapp-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/webapp-testing, .gemini/skills/webapp-testing, .github/skills/webapp-testing and .opencode/skills/webapp-testing in your project.

What does Web Application Testing need to run?

Going by SKILL.md and its folder, Web Application Testing needs Python for the scripts in its folder and the command-line tools its instructions call (python and npm). Our summary lists: Python with Playwright; The app's start commands, such as npm run dev.

Does Web Application Testing access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Web Application Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Web Application Testing use?

Web Application Testing is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Application Testing use?

About 966 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Application Testing?

Skills that share tags, products or a category with Web Application Testing: RStudio Selenium to Playwright Migration (rstudio/rstudio, 5.1k stars), Python Testing (macalbert/envilder, 138 stars), Webapp Testing (xenitV1/Antigravity-Workflows, 130 stars) and Playwright (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Application Testing?

anthropics (a GitHub organization, an official publisher) maintains it in anthropics/skills, which has 179,935 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 5, 2026.

Source: anthropics/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.