Agent skill

QA Testing

by espennilsen in espennilsen/pi

Perform QA testing on running web applications using cmuxbrowser.

MITAuto-check passedTesting & QA

Install QA Testing

skills CLI
$ npx skills add espennilsen/pi --skill qa-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install espennilsen/pi qa-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa-testing .claude/skills/qa-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-testing
GitHub stars
122
Token cost
~2.9k tokens
SKILL.md length
693 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Perform QA testing on running web applications using cmuxbrowser.

  • Works in 8 steps: Setup → Smoke Test → Functional Testing → …
  • — use this skill when: - User says test this
  • SKILL.md covers Prerequisites, Workflow, Report Template and Grading Rubric Reference, plus 1 more section
  • Calls docker, npx and npm; reaches cdnjs.cloudflare.com

What it does

QA Testing is an agent skill from espennilsen/pi. Perform QA testing on running web applications using cmuxbrowser. Covers functional testing, accessibility auditing, edge case testing, and structured reporting with evidence. Triggers — use this skill when: - User says "test this", "QA this", "QA this PR", "check the app" - User says "run QA", "verify the build", "test the UI" - User asks to "check if it works", "validate the feature" - User says "acceptance testing", "smoke test", "regression test" - User asks to "find bugs", "test for bugs", "break it" - User…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports and Accessibility. It works with Playwright. The licence is MIT.

When your agent uses it

  • — use this skill when: - User says test this
  • Check the app - User says run QA
  • Verify the build
  • Test the UI - User asks to check if it works

Example prompts

  • “test this”
  • “QA this”
  • “QA this PR”
  • “/qa-testing”

Requirements

  • Node.js
  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Smoke Test
  3. Functional Testing
  4. Edge Case Testing
  5. Accessibility Audit
  6. Performance (Optional)
  7. Report
  8. Cleanup

What it can do on your machine

Read from SKILL.md and the folder at commit 79d019b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • npx
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cdnjs.cloudflare.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA Testing loads about 2.9k tokens when it runs. Until then it costs about 180 tokens; SKILL.md has 693 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~180
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from espennilsen/pi at commit 79d019b, republished under its MIT licence (© espennilsen). 693 words, ~2,873 tokens.

Download SKILL.mdSave it as .claude/skills/qa-testing/SKILL.md (or your agent's skills folder).
name
qa-testing
description
Perform QA testing on running web applications using cmux_browser. Covers functional testing, accessibility auditing, edge case testing, and structured reporting with evidence. **Triggers — use this skill when:** - User says "test this", "QA this", "QA this PR", "check the app" - User says "run QA", "verify the build", "test the UI" - User asks to "check if it works", "validate the feature" - User says "acceptance testing", "smoke test", "regression test" - User asks to "find bugs", "test for bugs", "break it" - User wants to "check accessibility", "run axe", "a11y audit" **Covers:** Any web application accessible via HTTP. Uses cmux_browser (Playwright under the hood) for all browser interactions.

QA Testing — Web Application Testing with cmux_browser

Test running web applications by interacting as a real user via cmux_browser, then report findings with structured evidence.

Prerequisites

  • The application must be running and accessible via HTTP
  • cmux_browser tool must be available
  • If the app isn't running, start it first (see Setup below)

Workflow

Step 1: Setup

If the app is already running, skip to Step 2.

If the app needs to be started:

# Start in a background cmux pane
cmux_split({ direction: "down", command: "cd /path/to/project && npm run dev\n" })

# Wait for it to be ready (check the pane output)
cmux_read({ surface: "surface:N" })

# Or poll the URL
cmux_browser({ action: "open", url: "http://localhost:3000" })
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })

If Docker is needed:

bash
cd /path/to/project && docker compose up -d
# Wait for healthy
timeout 60 bash -c 'until curl -sf http://localhost:3000 > /dev/null; do sleep 2; done'
Step 2: Smoke Test

Verify the app loads at all before deeper testing.

# Open the app
cmux_browser({ action: "open", url: "http://localhost:3000" })

# Wait for it to load
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })

# Screenshot the initial state
cmux_browser({ action: "screenshot" })

# Check for JavaScript errors
cmux_browser({ action: "errors" })

# Check console output
cmux_browser({ action: "console" })

# Get a DOM snapshot (accessibility tree)
cmux_browser({ action: "snapshot" })

If the smoke test fails (page doesn't load, crash errors), stop and report immediately.

Step 3: Functional Testing

For each acceptance criterion, follow this pattern:

# 1. Navigate to the relevant state
cmux_browser({ action: "navigate", url: "/some-page" })
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })

# 2. Interact as a real user
cmux_browser({ action: "click", selector: "button.submit" })
cmux_browser({ action: "fill", selector: "input[name='email']", value: "test@example.com" })
cmux_browser({ action: "press", key: "Enter" })

# 3. Verify outcomes
cmux_browser({ action: "wait", waitCondition: "text", text: "Success" })
cmux_browser({ action: "snapshot", interactive: true })  # check element states
cmux_browser({ action: "get", selector: ".result", subaction: "textContent" })
cmux_browser({ action: "is", selector: ".modal", subaction: "visible" })

# 4. Screenshot as evidence
cmux_browser({ action: "screenshot" })

# 5. Check for errors after the interaction
cmux_browser({ action: "errors" })

Tips for reliable element selection:

  • Use snapshot to see the accessibility tree and find the right selectors
  • Use find with role for semantic targeting: cmux_browser({ action: "find", subaction: "role", name: "Submit" })
  • Use identify to see interactive elements on the page
  • Prefer data-testid, aria-label, or semantic selectors over fragile CSS paths
Step 4: Edge Case Testing

Be skeptical. Agents often ship code that works for the happy path but breaks on edge cases. Actively try to break things:

TestHowWhat to look for
Empty stateClear all data, visit pages with no contentCrashes, blank screens, missing "no data" messages
Empty inputsSubmit forms with empty required fieldsMissing validation, silent failures, crashes
Long textPaste 500+ character strings into inputsOverflow, layout breaking, truncation without indication
Special charactersInput <script>alert(1)</script>, emoji 🎉, Unicode ñXSS, encoding errors, display issues
Rapid clicksDouble-click submit buttons, rapidly toggle switchesDuplicate submissions, race conditions, broken state
Back buttonNavigate forward through a flow, then press backLost state, stale data, errors
RefreshF5 / reload mid-flowLost state, errors, unexpected redirects
Network errorsDisconnect WiFi / block API calls (if possible)Missing error handling, infinite spinners, blank screens
# Example: test empty form submission
cmux_browser({ action: "click", selector: "button[type='submit']" })
cmux_browser({ action: "screenshot" })
cmux_browser({ action: "errors" })

# Example: test long text
cmux_browser({ action: "fill", selector: "input[name='title']", value: "A".repeat(500) })
cmux_browser({ action: "screenshot" })

# Example: test special characters
cmux_browser({ action: "fill", selector: "input[name='name']", value: "<script>alert('xss')</script>" })
cmux_browser({ action: "click", selector: "button[type='submit']" })
cmux_browser({ action: "screenshot" })
cmux_browser({ action: "errors" })
Step 5: Accessibility Audit

Inject axe-core and run a full accessibility audit:

# Inject axe-core library
cmux_browser({ action: "eval", value: `
  await new Promise((resolve, reject) => {
    const script = document.createElement('script');
    script.src = 'https://cdnjs.cloudflare.com/ajax/libs/axe-core/4.10.2/axe.min.js';
    script.onload = resolve;
    script.onerror = reject;
    document.head.appendChild(script);
  });
  const results = await axe.run();
  return JSON.stringify({
    violations: results.violations.map(v => ({
      id: v.id,
      impact: v.impact,
      description: v.description,
      helpUrl: v.helpUrl,
      nodes: v.nodes.length
    })),
    passes: results.passes.length,
    incomplete: results.incomplete.length,
    inapplicable: results.inapplicable.length
  }, null, 2);
` })

CDN fallback: If the CDN is unreachable (air-gapped CI, network restriction, outage), the script injection will fail. In that case, fall back to:

  • npx axe-cli http://localhost:3000 --save axe-report.json (CLI-based audit)
  • or skip the accessibility audit and note "A11y audit skipped — axe-core CDN unavailable" in the report

Security note: For production use, add Subresource Integrity (SRI) verification to the script tag: script.integrity = "sha384-..."; script.crossOrigin = "anonymous"; (hash available on cdnjs.com).

Interpreting axe-core results:

  • critical impact — Must fix. Screen readers can't use the page.
  • serious impact — Should fix. Major barrier for some users.
  • moderate impact — Nice to fix. Some users affected.
  • minor impact — Optional. Best practice improvements.

Scoring accessibility:

ViolationsScore
0 violations10
1-3 minor8-9
1-3 serious6-7
4-10 mixed4-5
10+ or any critical2-3
Show full SKILL.md (249 more words)Show less
Step 6: Performance (Optional)

Run Lighthouse for performance auditing when requested:

bash
npx lighthouse http://localhost:3000 \
  --output=json \
  --output-path=./lighthouse-report.json \
  --chrome-flags="--headless --no-sandbox" \
  --only-categories=performance,accessibility,best-practices \
  --quiet

Then read and summarize the results:

read lighthouse-report.json
Step 7: Report

Produce a structured QA report. Always include:

  1. Verdict — PASS (all dimensions ≥ 6, no critical failures) or FAIL
  2. Scores — 1-10 for each dimension with brief justification
  3. Acceptance criteria checklist — each item explicitly PASS or FAIL
  4. Bugs — with severity, reproduction steps, expected vs actual, screenshot
  5. Edge cases tested — what you tried and what happened
  6. Accessibility results — axe-core violation count and details
  7. Screenshots — numbered, referenced in bug descriptions
Step 8: Cleanup

If you started a dev server or Docker container in Step 1, clean up:

# Stop a dev server running in a cmux pane
cmux_close({ surface: "surface:N" })

# Or stop Docker containers
cd /path/to/project && docker compose down

This prevents orphaned processes and keeps the environment clean for the next test run.

Report Template

markdown
# QA Report: [Feature/PR Name]

**Date:** YYYY-MM-DD
**App URL:** http://localhost:XXXX
**Tested by:** QA Agent

## Verdict: PASS ✅ / FAIL ❌

## Scores

| Dimension | Score | Notes |
|-----------|-------|-------|
| Functionality | X/10 | Brief justification |
| Completeness | X/10 | Brief justification |
| UX | X/10 | Brief justification |
| Robustness | X/10 | Brief justification |
| Accessibility | X/10 | N violations (N critical, N serious) |

**Average:** X.X/10

## Acceptance Criteria

- [x] Criterion 1 — PASS
- [x] Criterion 2 — PASS
- [ ] Criterion 3 — FAIL: [specific issue]

## Bugs Found

### 🔴 Bug 1: [Title] (Critical)
- **Steps:** 1. Navigate to /page 2. Click button 3. ...
- **Expected:** Form submits and shows success
- **Actual:** Page crashes with TypeError in console
- **Screenshot:** #3

### 🟡 Bug 2: [Title] (Major)
...

### 🔵 Bug 3: [Title] (Minor)
...

## Edge Cases Tested

| Test | Result | Notes |
|------|--------|-------|
| Empty form submission | ✅ PASS | Shows validation errors |
| Long text (500 chars) | ⚠️ WARN | Text overflows container |
| Special characters | ✅ PASS | Properly escaped |
| Back button | ❌ FAIL | State lost, shows blank page |
| Rapid double-click | ✅ PASS | Button disabled after first click |

## Accessibility (axe-core)

- **Violations:** N
- **Passes:** N
- **Details:**
  - [serious] button-name: 2 buttons missing accessible names
  - [moderate] color-contrast: 3 elements with insufficient contrast

## Screenshots

1. Initial load — [description]
2. After form submission — [description]
3. Bug #1 evidence — [description]

Grading Rubric Reference

Dimension10 (Exceptional)8-9 (Very Good)6-7 (Good)4-5 (Acceptable)2-3 (Poor)
FunctionalityAll criteria pass, flows smooth—Most pass, minor issuesSome criteria failCore flows broken
CompletenessEverything built and working—Minor features missingSignificant gapsMostly stubs
UXPolished, delightful—Good, minor rough edgesFunctional but clunkyConfusing/broken
RobustnessHandles everything gracefully—Handles common casesSome edge cases crashFragile
Accessibility0 violations1-3 minor violations1-3 serious violations4-10 mixed violations10+ or any critical

Verdict Rules

  • PASS = All dimensions ≥ 6 AND no critical functionality failures
  • FAIL = Any dimension < 6 OR acceptance criteria not met

When reporting FAIL, always provide specific, actionable feedback the builder can use to fix the issues. Reference exact elements, URLs, and steps.

© espennilsen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qa-testing of espennilsen/pi.

Open the folder on GitHubat commit 79d019b

Compare with similar skills

QA Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA Testing this skillespennilsen/pi122—~2.9kAutomated safety check: PassMIT
Agentic Browser Testingpetrkindlmann/qa-skills165—~4.5kAutomated safety check: PassMIT
Ds Test Componentbaloise/design-system114—~2.1kAutomated safety check: PassApache-2.0
Scoutqa Testgithub/awesome-copilot40k1 repos~3.5kAutomated safety check: PassMIT
Playwright VerifierCommunity-Access/accessibility-agents422—~1.1kAutomated safety check: PassMIT
E2E Testingericrisco/rsc-harness167—~3.2kAutomated safety check: PassMIT

Similar skills

  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    165 GitHub stars~4.5k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Ds Test Component

    baloise/design-system

    Auto-generate all test files for DS components including visual, a11y, component, page object, and unit tests.

    114 GitHub stars~2.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Scoutqa Test

    github/awesome-copilot

    Official

    This skill should be used when the user asks to "test this website", "run exploratory testing", "check for accessibility issues", "verify the login flow works", "find bugs on this page", or requests…

    40k GitHub starsUsed in 1 repo~3.5k tokens
    Testing & QAAuto-check passed
  • Playwright Verifier

    Community-Access/accessibility-agents

    Internal helper: re-run targeted scans to confirm a fix works at runtime.

    422 GitHub stars~1.1k tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • E2E Testing

    ericrisco/rsc-harness

    A skill your agent uses when writing or stabilizing Playwright tests that drive a real browser through multi-step journeys — durable locators, web-first assertions, storageState auth, trace/retries…

    167 GitHub stars~3.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Scanner

    Community-Access/accessibility-agents

    Internal helper: behavioral scans via Playwright for keyboard and state.

    422 GitHub stars~911 tokensUpdated 15 days ago
    Testing & QAAuto-check passed

More from espennilsen/pi

All 36 skills in this repo
  • GitHub

    espennilsen/pi

    Interact with GitHub repos, PRs, issues, CI, and notifications via the pi-github extension commands and gh CLI.

    122 GitHub stars~1k tokensUpdated 16 days ago
    Auto-check passed
  • Skill Creator

    espennilsen/pi

    Create, review, and improve skills for Pi agents. An agent skill from espennilsen/pi.

    122 GitHub stars~2.1k tokensUpdated 16 days ago
    Auto-check passed
  • Dry Code Review

    espennilsen/pi

    Perform a comprehensive DRY (Don't Repeat Yourself) code review on a codebase.

    122 GitHub stars~1.7k tokensUpdated 16 days ago
    Auto-check passed
  • Extract Design System

    espennilsen/pi

    Reverse-engineer a design system from a live website (public URL or localhost).

    122 GitHub stars~2k tokensUpdated 16 days ago
    Auto-check passed
  • Herdr Operations

    espennilsen/pi

    A skill your agent uses when inspecting or operating Herdr sessions, workspaces, tabs, panes, agents, terminal output, agent messaging, or waits.

    122 GitHub starsUsed in 1 repo~525 tokens
    Auto-check passed
  • PDF Reader

    espennilsen/pi

    Read and extract content from PDF files — text, tables, metadata, and images.

    122 GitHub stars~1.6k tokensUpdated 16 days ago
    Auto-check passed

Works with

Questions about QA Testing

What does QA Testing do?

Perform QA testing on running web applications using cmuxbrowser. QA Testing is an agent skill from espennilsen/pi. Perform QA testing on running web applications using cmuxbrowser.

When should I use QA Testing?

QA Testing fits situations like: — use this skill when: - User says test this; check the app - User says run QA; verify the build; test the UI - User asks to check if it works.

How do I install QA Testing in Claude Code?

Run `npx skills add espennilsen/pi --skill qa-testing -a claude-code`. Or copy the skill folder (skills/qa-testing in espennilsen/pi) into .claude/skills/qa-testing in your project. Claude Code loads it when a task matches its description.

How do I install QA Testing in Codex?

Run `npx skills add espennilsen/pi --skill qa-testing -a codex`. Or copy the skill folder (skills/qa-testing in espennilsen/pi) into .agents/skills/qa-testing in your project. Codex loads it when a task matches its description.

Can I use QA Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add espennilsen/pi --skill qa-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-testing, .gemini/skills/qa-testing, .github/skills/qa-testing and .opencode/skills/qa-testing in your project.

What does QA Testing need to run?

Going by SKILL.md and its folder, QA Testing needs the command-line tools its instructions call (docker, npx and npm). Our summary lists: Node.js; Docker.

Does QA Testing access the network?

SKILL.md names 1 domain. In commands or code: cdnjs.cloudflare.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is QA Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QA Testing use?

QA Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA Testing use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA Testing?

Skills that share tags, products or a category with QA Testing: Agentic Browser Testing (petrkindlmann/qa-skills, 165 stars), Ds Test Component (baloise/design-system, 114 stars), Scoutqa Test (github/awesome-copilot, 40k stars) and Playwright Verifier (Community-Access/accessibility-agents, 422 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA Testing?

espennilsen (a GitHub user) maintains it in espennilsen/pi, which has 122 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 21, 2026.

Source: espennilsen/pi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.