Agent skill

Webapp Testing

by bobmatnyc in bobmatnyc/claude-mpm

Automated webapp testing with Playwright. An agent skill from bobmatnyc/claude-mpm.

Apache-2.0Auto-check passedTesting & QA

Install Webapp Testing

skills CLI
$ npx skills add bobmatnyc/claude-mpm --skill webapp-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bobmatnyc/claude-mpm webapp-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bobmatnyc/claude-mpm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/claude_mpm/skills/bundled/testing/webapp-testing .claude/skills/webapp-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
webapp-testing
GitHub stars
155
Token cost
~1.7k tokens
SKILL.md length
537 words
Files
11 (incl. scripts)
Skills in repo
63
Repo updated
First seen
Licence
Apache-2.0

At a glance

Automated webapp testing with Playwright. An agent skill from bobmatnyc/claude-mpm.

  • Works in 4 steps: Verify Server State → Start Server (If Needed) → Write Test with Reconnaissance → …
  • Tasks that involve Browser testing
  • SKILL.md covers Overview, When to Use This Skill, The Iron Law and Quick Start, plus 6 more sections
  • Runs Python scripts from its folder; calls python, curl and npm

What it does

Webapp Testing is an agent skill from bobmatnyc/claude-mpm. Automated webapp testing with Playwright. Server management, UI testing, visual debugging, and reconnaissance-first approach.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts (for example `decision-tree.md`, `examples/console_logging.py` and `examples/element_discovery.py`).

It sits in Testing & QA, covering Browser testing. It works with Playwright. The repository describes itself as: Claude Multi-Agent Project Manager — multi-channel orchestration, GitHub-first SDK mode, and plugin system for Claude. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Browser testing

Example prompts

  • “/webapp-testing”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Verify Server State
  2. Start Server (If Needed)
  3. Write Test with Reconnaissance
  4. Verify Results

What it can do on your machine

Read from SKILL.md and the folder at commit 25203d3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • curl
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Webapp Testing loads about 1.7k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 537 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bobmatnyc/claude-mpm at commit 25203d3, republished under its Apache-2.0 licence (© bobmatnyc). 537 words, ~1,689 tokens.

Download SKILL.mdSave it as .claude/skills/webapp-testing/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
webapp-testing
description
Automated webapp testing with Playwright. Server management, UI testing, visual debugging, and reconnaissance-first approach.
version
2.0.0
category
testing
license
Complete terms in LICENSE.txt
progressive_disclosure.references
playwright-patterns.md, server-management.md, reconnaissance-pattern.md, decision-tree.md, troubleshooting.md
effort
medium

Webapp Testing

Overview

Core Principle: Reconnaissance Before Action

Automated webapp testing using Playwright with a focus on verifying system state (server status, page load, element presence) before taking any action. This ensures reliable, debuggable tests that fail for clear reasons.

Key capabilities:

  • Automated browser testing with Playwright
  • Server lifecycle management
  • Visual reconnaissance (screenshots, DOM inspection)
  • Network monitoring and debugging

When to Use This Skill

  • Web application testing - UI behavior, forms, navigation, integration testing
  • Frontend debugging - Screenshots, DOM inspection, console monitoring
  • Regression testing - Ensure changes don't break existing functionality
  • Server verification - Check servers are running and responding

Not suitable for: Unit testing (use Jest/pytest), load testing, or API-only testing.

The Iron Law

RECONNAISSANCE BEFORE ACTION

Never execute test actions without first:

  1. Verify server state - lsof -i :PORT and curl checks
  2. Wait for page ready - page.wait_for_load_state('networkidle')
  3. Visual confirmation - Screenshot before actions
  4. Read complete output - Examine full results before claiming success

Why: Tests fail mysteriously when servers aren't ready, selectors break when DOM is still building, and 5 seconds of reconnaissance saves 30 minutes of debugging.

Quick Start

Step 1: Verify Server State
bash
lsof -i :3000 -sTCP:LISTEN  # Check server listening
curl -f http://localhost:3000/health  # Test response
Step 2: Start Server (If Needed)
bash
# Single server
python scripts/with_server.py --server "npm run dev" --port 5173 -- python test.py

# Multiple servers (backend + frontend)
python scripts/with_server.py \
  --server "cd backend && python server.py" --port 3000 \
  --server "cd frontend && npm run dev" --port 5173 \
  -- python test.py
Step 3: Write Test with Reconnaissance
python
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()

    # 1. Navigate and wait
    page.goto('http://localhost:5173')
    page.wait_for_load_state('networkidle')  # CRITICAL

    # 2. Reconnaissance
    page.screenshot(path='/tmp/before.png', full_page=True)
    buttons = page.locator('button').all()
    print(f"Found {len(buttons)} buttons")

    # 3. Execute
    page.click('button.submit')

    # 4. Verify
    page.wait_for_selector('.success-message')
    page.screenshot(path='/tmp/after.png', full_page=True)

    browser.close()
Step 4: Verify Results

Review console output, check for errors, verify state changes, examine screenshots.

Key Patterns

Server Management - Check → Start → Wait → Test → Cleanup

  • Use with_server.py for automatic lifecycle management
  • Check status with lsof, test with curl
  • Automatic cleanup on exit

Reconnaissance - Inspect → Understand → Act → Verify

  • Screenshot current state
  • Inspect DOM for elements
  • Act on discovered selectors
  • Verify results visually

Wait Strategy - Load → Idle → Element → Action

  • Always wait for networkidle on dynamic apps
  • Wait for specific elements before interaction
  • Playwright auto-waits but explicit waits prevent race conditions

Selector Priority - data-testid > role > text > CSS > XPath

  • [data-testid="submit"] - most stable
  • role=button[name="Submit"] - semantic
  • text=Submit - readable
  • button.submit - acceptable
  • XPath - last resort
Show full SKILL.md (235 more words)Show less

Common Pitfalls

❌ Testing without server verification - Always check lsof and curl first ❌ Ignoring timeout errors - TimeoutError means something is wrong, investigate ❌ Not waiting for networkidle - Dynamic apps need full page load ❌ Poor selector strategies - Use data-testid for stability ❌ Missing network verification - Check API responses complete ❌ Incomplete cleanup - Close browsers, stop servers properly

Reference Documentation

playwright-patterns.md - Complete Playwright reference Selectors, waits, interactions, assertions, test organization, network interception, screenshots, debugging

server-management.md - Server lifecycle and operations with_server.py usage, manual management, port management, process control, environment config, health checks

reconnaissance-pattern.md - Philosophy and practice Why reconnaissance first, complete process, server checks, network diagnostics, DOM inspection, log analysis

decision-tree.md - Flowcharts for every scenario New test decisions, server state paths, test failure diagnosis, debugging flows, selector/wait strategies

troubleshooting.md - Solutions to common problems Timeout issues, selector problems, server crashes, network errors, environment config, debugging workflow

Examples and Scripts

Examples (examples/ directory):

  • element_discovery.py - Discovering page elements
  • static_html_automation.py - Testing local HTML files
  • console_logging.py - Capturing console output

Scripts (scripts/ directory):

  • with_server.py - Server lifecycle management (run with --help first)

Integration with Other Skills

Mandatory: verification-before-completion Recommended: systematic-debugging, test-driven-development Related: playwright-testing, selenium-automation

Bottom Line

  1. Reconnaissance always comes first - Verify before acting
  2. Never skip server checks - 5 seconds saves 30 minutes
  3. Wait for networkidle - Dynamic apps need time
  4. Read complete output - Verify before claiming success
  5. Screenshot everything - Visual evidence is invaluable

The reconnaissance-then-action pattern is not optional - it's the foundation of reliable webapp testing.

© bobmatnyc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts) in src/claude_mpm/skills/bundled/testing/webapp-testing of bobmatnyc/claude-mpm.

  • SKILL.md
  • LICENSE.txt
  • decision-tree.md
  • examples/console_logging.py
  • examples/element_discovery.py
  • examples/static_html_automation.py
  • playwright-patterns.md
  • reconnaissance-pattern.md
  • scripts/with_server.py
  • server-management.md
  • troubleshooting.md

Open the folder on GitHubat commit 25203d3

Compare with similar skills

Webapp Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Webapp Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Webapp Testing this skillbobmatnyc/claude-mpm155—~1.7kAutomated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Playwright CLIsanity-io/sanity6.4k18 repos~1.9kAutomated safety check: PassMIT
playwright-cli Browser Automationgithub/gh-aw5.3k23 repos~2.8kAutomated safety check: PassMIT
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Playwright CLI

    sanity-io/sanity

    Official

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    6.4k GitHub starsUsed in 18 repos~1.9k tokens
    Testing & QAAuto-check passed
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.3k GitHub starsUsed in 23 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Testing

    langflow-ai/langflow

    Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow.

    156k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed

More from bobmatnyc/claude-mpm

All 63 skills in this repo
  • Build MCP Server

    bobmatnyc/claude-mpm

    Create high-quality MCP servers that enable LLMs to effectively interact with external services.

    155 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Env Manager

    bobmatnyc/claude-mpm

    Environment variable validation, synchronization, and management across local development, CI/CD, and deployment platforms

    155 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Session Analyzer

    bobmatnyc/claude-mpm

    Debug and teach agentic coding: a deterministic-first session timeline + cost report, with optional narrative polish and a standalone JSX visualiser.

    155 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Software Patterns

    bobmatnyc/claude-mpm

    Decision framework for architectural patterns including DI, SOA, Repository, Domain Events, Circuit Breaker, and Anti-Corruption Layer.

    155 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Verification Before Completion

    bobmatnyc/claude-mpm

    Run verification commands and confirm output before claiming success

    155 GitHub starsUsed in 2 repos~1k tokens
    Auto-check passed
  • Dependency Audit

    bobmatnyc/claude-mpm

    Dependency audit and cleanup workflow for maintaining healthy project dependencies.

    155 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Webapp Testing

What does Webapp Testing do?

Automated webapp testing with Playwright. An agent skill from bobmatnyc/claude-mpm. Webapp Testing is an agent skill from bobmatnyc/claude-mpm. Automated webapp testing with Playwright.

When should I use Webapp Testing?

Webapp Testing fits situations like: tasks that involve Browser testing.

How do I install Webapp Testing in Claude Code?

Run `npx skills add bobmatnyc/claude-mpm --skill webapp-testing -a claude-code`. Or copy the skill folder (src/claude_mpm/skills/bundled/testing/webapp-testing in bobmatnyc/claude-mpm) into .claude/skills/webapp-testing in your project. Claude Code loads it when a task matches its description.

How do I install Webapp Testing in Codex?

Run `npx skills add bobmatnyc/claude-mpm --skill webapp-testing -a codex`. Or copy the skill folder (src/claude_mpm/skills/bundled/testing/webapp-testing in bobmatnyc/claude-mpm) into .agents/skills/webapp-testing in your project. Codex loads it when a task matches its description.

Can I use Webapp Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bobmatnyc/claude-mpm --skill webapp-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/webapp-testing, .gemini/skills/webapp-testing, .github/skills/webapp-testing and .opencode/skills/webapp-testing in your project.

What does Webapp Testing need to run?

Going by SKILL.md and its folder, Webapp Testing needs Python for the scripts in its folder and the command-line tools its instructions call (python, curl and npm). Our summary lists: Python 3.

Does Webapp Testing access the network?

SKILL.md contains no URLs. Its commands use curl and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Webapp Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Webapp Testing use?

Webapp Testing is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Webapp Testing use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Webapp Testing?

Skills that share tags, products or a category with Webapp Testing: Web Application Testing (anthropics/skills, 180k stars), Playwright CLI (sanity-io/sanity, 6.4k stars), playwright-cli Browser Automation (github/gh-aw, 5.3k stars) and Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Webapp Testing?

bobmatnyc (a GitHub user) maintains it in bobmatnyc/claude-mpm, which has 155 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on August 31, 2026.

Source: bobmatnyc/claude-mpm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.