Agentic Browser Testing
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
Perform QA testing on running web applications using cmuxbrowser.
$ npx skills add espennilsen/pi --skill qa-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install espennilsen/pi qa-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa-testing .claude/skills/qa-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .claude/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/espennilsen/pi/tree/main/skills/qa-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add espennilsen/pi --skill qa-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install espennilsen/pi qa-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/qa-testing .agents/skills/qa-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .agents/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add espennilsen/pi --skill qa-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install espennilsen/pi qa-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/qa-testing .cursor/skills/qa-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .cursor/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/espennilsen/pi.git --path skills/qa-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add espennilsen/pi --skill qa-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install espennilsen/pi qa-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/qa-testing .gemini/skills/qa-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .gemini/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install espennilsen/pi qa-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add espennilsen/pi --skill qa-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/qa-testing .github/skills/qa-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .github/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add espennilsen/pi --skill qa-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install espennilsen/pi qa-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/qa-testing .opencode/skills/qa-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "qa-testing" agent skill from https://github.com/espennilsen/pi/tree/main/skills/qa-testing into .opencode/skills/qa-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
qa-testingPerform QA testing on running web applications using cmuxbrowser.
QA Testing is an agent skill from espennilsen/pi. Perform QA testing on running web applications using cmuxbrowser. Covers functional testing, accessibility auditing, edge case testing, and structured reporting with evidence. Triggers — use this skill when: - User says "test this", "QA this", "QA this PR", "check the app" - User says "run QA", "verify the build", "test the UI" - User asks to "check if it works", "validate the feature" - User says "acceptance testing", "smoke test", "regression test" - User asks to "find bugs", "test for bugs", "break it" - User…
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering QA and bug reports and Accessibility. It works with Playwright. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 79d019b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockernpxnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
cdnjs.cloudflare.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
QA Testing loads about 2.9k tokens when it runs. Until then it costs about 180 tokens; SKILL.md has 693 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from espennilsen/pi at commit 79d019b, republished under its MIT licence (© espennilsen). 693 words, ~2,873 tokens.
.claude/skills/qa-testing/SKILL.md (or your agent's skills folder).Test running web applications by interacting as a real user via cmux_browser, then report findings with structured evidence.
If the app is already running, skip to Step 2.
If the app needs to be started:
# Start in a background cmux pane
cmux_split({ direction: "down", command: "cd /path/to/project && npm run dev\n" })
# Wait for it to be ready (check the pane output)
cmux_read({ surface: "surface:N" })
# Or poll the URL
cmux_browser({ action: "open", url: "http://localhost:3000" })
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })If Docker is needed:
cd /path/to/project && docker compose up -d
# Wait for healthy
timeout 60 bash -c 'until curl -sf http://localhost:3000 > /dev/null; do sleep 2; done'Verify the app loads at all before deeper testing.
# Open the app
cmux_browser({ action: "open", url: "http://localhost:3000" })
# Wait for it to load
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })
# Screenshot the initial state
cmux_browser({ action: "screenshot" })
# Check for JavaScript errors
cmux_browser({ action: "errors" })
# Check console output
cmux_browser({ action: "console" })
# Get a DOM snapshot (accessibility tree)
cmux_browser({ action: "snapshot" })If the smoke test fails (page doesn't load, crash errors), stop and report immediately.
For each acceptance criterion, follow this pattern:
# 1. Navigate to the relevant state
cmux_browser({ action: "navigate", url: "/some-page" })
cmux_browser({ action: "wait", waitCondition: "load-state", loadState: "networkidle" })
# 2. Interact as a real user
cmux_browser({ action: "click", selector: "button.submit" })
cmux_browser({ action: "fill", selector: "input[name='email']", value: "test@example.com" })
cmux_browser({ action: "press", key: "Enter" })
# 3. Verify outcomes
cmux_browser({ action: "wait", waitCondition: "text", text: "Success" })
cmux_browser({ action: "snapshot", interactive: true }) # check element states
cmux_browser({ action: "get", selector: ".result", subaction: "textContent" })
cmux_browser({ action: "is", selector: ".modal", subaction: "visible" })
# 4. Screenshot as evidence
cmux_browser({ action: "screenshot" })
# 5. Check for errors after the interaction
cmux_browser({ action: "errors" })Tips for reliable element selection:
snapshot to see the accessibility tree and find the right selectorsfind with role for semantic targeting: cmux_browser({ action: "find", subaction: "role", name: "Submit" })identify to see interactive elements on the pagedata-testid, aria-label, or semantic selectors over fragile CSS pathsBe skeptical. Agents often ship code that works for the happy path but breaks on edge cases. Actively try to break things:
| Test | How | What to look for |
|---|---|---|
| Empty state | Clear all data, visit pages with no content | Crashes, blank screens, missing "no data" messages |
| Empty inputs | Submit forms with empty required fields | Missing validation, silent failures, crashes |
| Long text | Paste 500+ character strings into inputs | Overflow, layout breaking, truncation without indication |
| Special characters | Input <script>alert(1)</script>, emoji 🎉, Unicode ñ | XSS, encoding errors, display issues |
| Rapid clicks | Double-click submit buttons, rapidly toggle switches | Duplicate submissions, race conditions, broken state |
| Back button | Navigate forward through a flow, then press back | Lost state, stale data, errors |
| Refresh | F5 / reload mid-flow | Lost state, errors, unexpected redirects |
| Network errors | Disconnect WiFi / block API calls (if possible) | Missing error handling, infinite spinners, blank screens |
# Example: test empty form submission
cmux_browser({ action: "click", selector: "button[type='submit']" })
cmux_browser({ action: "screenshot" })
cmux_browser({ action: "errors" })
# Example: test long text
cmux_browser({ action: "fill", selector: "input[name='title']", value: "A".repeat(500) })
cmux_browser({ action: "screenshot" })
# Example: test special characters
cmux_browser({ action: "fill", selector: "input[name='name']", value: "<script>alert('xss')</script>" })
cmux_browser({ action: "click", selector: "button[type='submit']" })
cmux_browser({ action: "screenshot" })
cmux_browser({ action: "errors" })Inject axe-core and run a full accessibility audit:
# Inject axe-core library
cmux_browser({ action: "eval", value: `
await new Promise((resolve, reject) => {
const script = document.createElement('script');
script.src = 'https://cdnjs.cloudflare.com/ajax/libs/axe-core/4.10.2/axe.min.js';
script.onload = resolve;
script.onerror = reject;
document.head.appendChild(script);
});
const results = await axe.run();
return JSON.stringify({
violations: results.violations.map(v => ({
id: v.id,
impact: v.impact,
description: v.description,
helpUrl: v.helpUrl,
nodes: v.nodes.length
})),
passes: results.passes.length,
incomplete: results.incomplete.length,
inapplicable: results.inapplicable.length
}, null, 2);
` })CDN fallback: If the CDN is unreachable (air-gapped CI, network restriction, outage), the script injection will fail. In that case, fall back to:
npx axe-cli http://localhost:3000 --save axe-report.json (CLI-based audit)Security note: For production use, add Subresource Integrity (SRI) verification to the script tag: script.integrity = "sha384-..."; script.crossOrigin = "anonymous"; (hash available on cdnjs.com).
Interpreting axe-core results:
Scoring accessibility:
| Violations | Score |
|---|---|
| 0 violations | 10 |
| 1-3 minor | 8-9 |
| 1-3 serious | 6-7 |
| 4-10 mixed | 4-5 |
| 10+ or any critical | 2-3 |
Run Lighthouse for performance auditing when requested:
npx lighthouse http://localhost:3000 \
--output=json \
--output-path=./lighthouse-report.json \
--chrome-flags="--headless --no-sandbox" \
--only-categories=performance,accessibility,best-practices \
--quietThen read and summarize the results:
read lighthouse-report.jsonProduce a structured QA report. Always include:
If you started a dev server or Docker container in Step 1, clean up:
# Stop a dev server running in a cmux pane
cmux_close({ surface: "surface:N" })
# Or stop Docker containers
cd /path/to/project && docker compose downThis prevents orphaned processes and keeps the environment clean for the next test run.
# QA Report: [Feature/PR Name]
**Date:** YYYY-MM-DD
**App URL:** http://localhost:XXXX
**Tested by:** QA Agent
## Verdict: PASS ✅ / FAIL ❌
## Scores
| Dimension | Score | Notes |
|-----------|-------|-------|
| Functionality | X/10 | Brief justification |
| Completeness | X/10 | Brief justification |
| UX | X/10 | Brief justification |
| Robustness | X/10 | Brief justification |
| Accessibility | X/10 | N violations (N critical, N serious) |
**Average:** X.X/10
## Acceptance Criteria
- [x] Criterion 1 — PASS
- [x] Criterion 2 — PASS
- [ ] Criterion 3 — FAIL: [specific issue]
## Bugs Found
### 🔴 Bug 1: [Title] (Critical)
- **Steps:** 1. Navigate to /page 2. Click button 3. ...
- **Expected:** Form submits and shows success
- **Actual:** Page crashes with TypeError in console
- **Screenshot:** #3
### 🟡 Bug 2: [Title] (Major)
...
### 🔵 Bug 3: [Title] (Minor)
...
## Edge Cases Tested
| Test | Result | Notes |
|------|--------|-------|
| Empty form submission | ✅ PASS | Shows validation errors |
| Long text (500 chars) | ⚠️ WARN | Text overflows container |
| Special characters | ✅ PASS | Properly escaped |
| Back button | ❌ FAIL | State lost, shows blank page |
| Rapid double-click | ✅ PASS | Button disabled after first click |
## Accessibility (axe-core)
- **Violations:** N
- **Passes:** N
- **Details:**
- [serious] button-name: 2 buttons missing accessible names
- [moderate] color-contrast: 3 elements with insufficient contrast
## Screenshots
1. Initial load — [description]
2. After form submission — [description]
3. Bug #1 evidence — [description]| Dimension | 10 (Exceptional) | 8-9 (Very Good) | 6-7 (Good) | 4-5 (Acceptable) | 2-3 (Poor) |
|---|---|---|---|---|---|
| Functionality | All criteria pass, flows smooth | — | Most pass, minor issues | Some criteria fail | Core flows broken |
| Completeness | Everything built and working | — | Minor features missing | Significant gaps | Mostly stubs |
| UX | Polished, delightful | — | Good, minor rough edges | Functional but clunky | Confusing/broken |
| Robustness | Handles everything gracefully | — | Handles common cases | Some edge cases crash | Fragile |
| Accessibility | 0 violations | 1-3 minor violations | 1-3 serious violations | 4-10 mixed violations | 10+ or any critical |
When reporting FAIL, always provide specific, actionable feedback the builder can use to fix the issues. Reference exact elements, URLs, and steps.
© espennilsen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/qa-testing of espennilsen/pi.
Open the folder on GitHubat commit 79d019b
QA Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| QA Testing this skillespennilsen/pi | 122 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Agentic Browser Testingpetrkindlmann/qa-skills | 165 | — | ~4.5k | Automated safety check: Pass | MIT | |
| Ds Test Componentbaloise/design-system | 114 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Scoutqa Testgithub/awesome-copilot | 40k | 1 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Playwright VerifierCommunity-Access/accessibility-agents | 422 | — | ~1.1k | Automated safety check: Pass | MIT | |
| E2E Testingericrisco/rsc-harness | 167 | — | ~3.2k | Automated safety check: Pass | MIT |
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
baloise/design-system
Auto-generate all test files for DS components including visual, a11y, component, page object, and unit tests.
github/awesome-copilot
This skill should be used when the user asks to "test this website", "run exploratory testing", "check for accessibility issues", "verify the login flow works", "find bugs on this page", or requests…
Community-Access/accessibility-agents
Internal helper: re-run targeted scans to confirm a fix works at runtime.
ericrisco/rsc-harness
A skill your agent uses when writing or stabilizing Playwright tests that drive a real browser through multi-step journeys — durable locators, web-first assertions, storageState auth, trace/retries…
Community-Access/accessibility-agents
Internal helper: behavioral scans via Playwright for keyboard and state.
espennilsen/pi
Interact with GitHub repos, PRs, issues, CI, and notifications via the pi-github extension commands and gh CLI.
espennilsen/pi
Create, review, and improve skills for Pi agents. An agent skill from espennilsen/pi.
espennilsen/pi
Perform a comprehensive DRY (Don't Repeat Yourself) code review on a codebase.
espennilsen/pi
Reverse-engineer a design system from a live website (public URL or localhost).
espennilsen/pi
A skill your agent uses when inspecting or operating Herdr sessions, workspaces, tabs, panes, agents, terminal output, agent messaging, or waits.
espennilsen/pi
Read and extract content from PDF files — text, tables, metadata, and images.
Works with
Categories
Perform QA testing on running web applications using cmuxbrowser. QA Testing is an agent skill from espennilsen/pi. Perform QA testing on running web applications using cmuxbrowser.
QA Testing fits situations like: — use this skill when: - User says test this; check the app - User says run QA; verify the build; test the UI - User asks to check if it works.
Run `npx skills add espennilsen/pi --skill qa-testing -a claude-code`. Or copy the skill folder (skills/qa-testing in espennilsen/pi) into .claude/skills/qa-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add espennilsen/pi --skill qa-testing -a codex`. Or copy the skill folder (skills/qa-testing in espennilsen/pi) into .agents/skills/qa-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add espennilsen/pi --skill qa-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-testing, .gemini/skills/qa-testing, .github/skills/qa-testing and .opencode/skills/qa-testing in your project.
Going by SKILL.md and its folder, QA Testing needs the command-line tools its instructions call (docker, npx and npm). Our summary lists: Node.js; Docker.
SKILL.md names 1 domain. In commands or code: cdnjs.cloudflare.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
QA Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with QA Testing: Agentic Browser Testing (petrkindlmann/qa-skills, 165 stars), Ds Test Component (baloise/design-system, 114 stars), Scoutqa Test (github/awesome-copilot, 40k stars) and Playwright Verifier (Community-Access/accessibility-agents, 422 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
espennilsen (a GitHub user) maintains it in espennilsen/pi, which has 122 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 21, 2026.
Source: espennilsen/pi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.