Category
Best testing and QA skills for Claude Code, Codex and other agents.
- skills
- 4,904
- official
- 371
Testing & QA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking… | vercel-labs/ | 44k | 24 repos | ~864 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 2 | Explores a web app with the agent-browser CLI to find bugs and UX problems, then writes a report with screenshots, repro videos and step-by-step reproduction for each issue. | vercel-labs/ | 44k | 8 repos | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way. | openinterpreter/ | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers. | obra/ | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 5 | Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs. | anthropics/ | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 6 | 6.Grilling Grill the user relentlessly about a plan, decision, or idea. | bestofjs/ | 3.1k | 30 repos | ~464 | Automated safety check: Pass | MIT | 3 days ago |
| 7 | Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data. | MemTensor/ | 12k | 3 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 8 | Automates browser interactions for web testing, form filling, screenshots, and data extraction. | sanity-io/ | 6.4k | 18 repos | ~1.9k | Automated safety check: Pass | MIT | today |
| 9 | Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port. | vercel-labs/ | 44k | 5 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. | HKUDS/ | 16k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | 4 mo ago |
| 11 | iOS Simulator control from inside Orca, with the live device view in Orca's emulator pane. Use when driving a booted Apple Simulator on macOS: taps… | stablyai/ | 87k | 1 repo | ~584 | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests. | github/ | 5.3k | 23 repos | ~2.8k | Automated safety check: Pass | MIT | today |
| 13 | Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations. | PowerShell/ | 56k | — | ~5.1k | Automated safety check: Pass | MIT | yesterday |
| 14 | Jest patterns for React Native style tests: TDD discipline, mock factory functions, module and GraphQL hook mocking, custom render helpers and anti-patterns to avoid. | ChrisWiles/ | 6.1k | 7 repos | ~1.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 15 | Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes. | appsmithorg/ | 41k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 16 | Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation. | fossasia/ | 1.6k | 31 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 17 | Enforces red-green-refactor for Rust work, with idiomatic test patterns, a naming convention and a pre-commit gate of cargo fmt, clippy and test. | rtk-ai/ | 83k | — | ~753 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 18 | Captures screenshots of the running Mailspring dev app for docs, PRs or visual checks by launching it with a debugging port, driving the UI and clipping to an element. | Foundry376/ | 18k | — | ~1.4k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 19 | Core usage guide for the agent-browser CLI: the snapshot-and-ref workflow for navigating, clicking, filling forms, extracting data and running parallel sessions. | vercel-labs/ | 44k | 4 repos | ~9.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | 20.Clawteam Dev A skill your agent uses when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually… | HKUDS/ | 5.5k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 21 | Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report. | bytedance/ | 83k | — | ~2.5k | Automated safety check: Notes | MIT | today |
| 22 | Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version. | colbymchenry/ | 73k | — | ~950 | Automated safety check: Pass | MIT | today |
| 23 | Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions. | langgenius/ | 158k | — | ~682 | Automated safety check: Pass | Unknown | today |
| 24 | 24.E2E Testing Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow. | langflow-ai/ | 156k | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 25 | 25.Effect V4 Write, review, or upgrade Effect v4 code in the Composio CLI, cli-keyring, and json-schema-to-effect-schema packages, all pinned exactly to effect@4.0.0-rc.117 — Context.Service and explicit layers… | ComposioHQ/ | 30k | — | ~1k | Automated safety check: Pass | MIT | today |
| 26 | 26.Diagnose Disciplined diagnosis loop for hard bugs and performance regressions. | ywwynm/ | 144 | 18 repos | ~1.8k | Automated safety check: Pass | GPL-3.0 | 23 days ago |
| 27 | Sets the test-writing workflow for the repository: risk-first scenario lists, behavior-focused Vitest tests, a full run before each commit and a coverage target. | iOfficeAI/ | 33k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | 28 days ago |
| 28 | Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely. | steipete/ | 22k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 29 | Connects an agent to a real Chrome instance through the Chrome DevTools MCP server, so it can inspect the DOM, read console errors and profile performance directly. | addyosmani/ | 102k | 4 repos | ~3.5k | Automated safety check: Warn | MIT | 4 days ago |
| 30 | Triages and lands a batch of open Dependabot PRs in the Onyx repo, where main is gated exclusively by GitHub's merge queue: approves and enqueues green PRs, closes superseded duplicates, fixes… | onyx-dot-app/ | 32k | 1 repo | ~2.2k | Automated safety check: Pass | MIT | today |
| 31 | Analyze an Android pull request, branch, commit, or patch for user-visible changes and produce reproducible before/after screenshots from isolated builds. | permissionlesstech/ | 7.7k | — | ~2.6k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 32 | Drives a real browser through the omowright library, either the user's own signed-in browser or a separate browser the code launches, for forms, QA, screenshots and scraping. | code-yeongyu/ | 70k | — | ~2.2k | Automated safety check: Pass | Unknown | today |
| 33 | 33.Expert Panel Score, evaluate, and iteratively improve any content or strategy using an auto-assembled panel of domain experts. | ericosiu/ | 3.6k | 2 repos | ~2.1k | Automated safety check: Pass | MIT | 15 days ago |
| 34 | Explains how to run agent integration tests against remote executors, using Docker for Linux or Wine for Windows, and how to opt tests in or skip them. | openinterpreter/ | 69k | 2 repos | ~842 | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Starts a throwaway MongoDB 7.0 replica set and runs the Airbyte spec, check, discover and read commands against source-mongodb-v2 images for local end-to-end testing. | airbytehq/ | 22k | — | ~1.9k | Automated safety check: Pass | Unknown | yesterday |
| 36 | Records a project's quality bar in CONSTRAINTS.md and watches diffs for signs an agent quietly weakened it, such as suppressions, skipped tests or lowered thresholds. | addyosmani/ | 102k | 2 repos | ~5.2k | Automated safety check: Pass | MIT | 4 days ago |
| 37 | Automates a Chrome or Chromium browser through the agent-browser CLI: navigate, fill forms, click, screenshot and extract data using element refs. | sipeed/ | 30k | — | ~1.1k | Automated safety check: Pass | MIT | 13 days ago |
| 38 | Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report. | appsmithorg/ | 41k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 39 | 39.Go Pedantry This skill should be used when the user is writing Go code and needs guidance on Go-specific pedantry: error wrapping with fmt.Errorf and %w, interface design (accept interfaces return structs)… | chromedp/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 2 days ago |
| 40 | Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI. | lobehub/ | 83k | — | ~9.7k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail… | google/ | 22k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | 42.Papercuts Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that… | every-app/ | 23k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 43 | 43.Use Yaak A skill your agent uses when the user mentions Yaak, a Yaak workspace, or the yaak command, or asks to call, hit, or smoke test HTTP/REST endpoints, save or organize API requests for reuse or manual… | mountain-loop/ | 19k | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 44 | 44.TDD Test-driven development. An agent skill from fossasia/eventyay-interpretation. | fossasia/ | 1.6k | 28 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 45 | Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or… | kedro-org/ | 11k | — | ~4k | Automated safety check: Pass | Unknown | yesterday |
| 46 | Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots. | nanocoai/ | 31k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 47 | Drives the Expo showcase app in an iOS simulator with agent-device to sweep every UI Kitten component in all theme and mapping combinations, reporting regressions with evidence. | akveo/ | 11k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 48 | Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication. | github/ | 11k | — | ~1.8k | Automated safety check: Pass | MIT | today |
Questions, answered from the data.
What is the best testing and QA skill?
Agent Browser CLI (official) from vercel-labs/agent-browser ranks first of the 4,904 testing and QA skills listed here, with the highest score: its repository has 44k GitHub stars, 24 other GitHub owners carry a copy, its SKILL.md loads about 864 tokens and it passes the automated safety check with no findings. Next come Dogfood Exploratory QA and PR Babysitter.
Which testing and QA skills are official?
371 of the 4,904 testing and QA skills are official, published by the vendor's own GitHub organization: Agent Browser CLI, Dogfood Exploratory QA, Web Application Testing, Playwright CLI, Electron App Automation and 366 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Topics in Testing & QA
Products these skills work with
Roles that use these skills
Other categories
- Development15,057
- Frontend & Design5,859
- Backend & APIs6,424
- DevOps & Cloud5,486
- Databases2,270
- Data & Analytics3,530
- AI & LLM Engineering4,913
- Agent Workflows8,317
- Documents & Office4,276
- Writing & Content3,844
- Marketing & SEO3,982
- Sales & Support1,788
- Research & Science6,175
- Security3,019
- Productivity & Automation4,118
- Business, Finance & HR4,048
- Legal & Compliance1,640
- Education1,144
- Media & Creative4,289
- Mobile2,786
- Product & Project Management2,437
- Knowledge Management1,458
- Game Development1,703