Topic · Testing & QA
Best end-to-end testing skills for Claude Code, Codex and other agents.
- skills
- 783
- official
- 58
End-to-end testing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs. | anthropics/ | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port. | vercel-labs/ | 44k | 5 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. | HKUDS/ | 16k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | 4 mo ago |
| 4 | Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests. | github/ | 5.3k | 23 repos | ~2.8k | Automated safety check: Pass | MIT | today |
| 5 | Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes. | appsmithorg/ | 41k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | today |
| 6 | Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions. | langgenius/ | 158k | — | ~682 | Automated safety check: Pass | Unknown | today |
| 7 | Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow. | langflow-ai/ | 156k | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 8 | Write, review, or upgrade Effect v4 code in the Composio CLI, cli-keyring, and json-schema-to-effect-schema packages, all pinned exactly to effect@4.0.0-rc.117 — Context.Service and explicit layers… | ComposioHQ/ | 30k | — | ~1k | Automated safety check: Pass | MIT | today |
| 9 | Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely. | steipete/ | 22k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 10 | Starts a throwaway MongoDB 7.0 replica set and runs the Airbyte spec, check, discover and read commands against source-mongodb-v2 images for local end-to-end testing. | airbytehq/ | 22k | — | ~1.9k | Automated safety check: Pass | Unknown | today |
| 11 | Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI. | lobehub/ | 83k | — | ~9.7k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication. | github/ | 11k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 13 | Stands up a throwaway local SQL Server 2022 backend, applies SQL fixtures and runs Airbyte spec, check, discover and read against source-mssql images. | airbytehq/ | 22k | — | ~4.4k | Automated safety check: Pass | Unknown | today |
| 14 | Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic. | QwenLM/ | 28k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Run Wix Engine (mobile-apps-engine) iOS E2E tests locally to validate RNN changes. | wix/ | 13k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 16 | Guides writing and changing Playwright end-to-end tests for Handsontable using page objects, data-testid hooks and deterministic waits. | handsontable/ | 22k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 17 | 17.TDD Workflow A skill your agent uses when writing new features, fixing bugs, or refactoring code. | hellangleZ/ | 112 | 11 repos | ~2.4k | Automated safety check: Pass | No licence | 8 mo ago |
| 18 | Runs the json_repair docs demo against a local Flask API and static server, so changes to docs/app.py or the docs UI are checked end to end before publishing. | mangiucugna/ | 5.1k | — | ~549 | Automated safety check: Pass | MIT | 6 days ago |
| 19 | Stands up a throwaway local MySQL 8.0 backend, applies SQL fixtures and sweeps the Airbyte spec, check, discover and read commands against a source-mysql image. | airbytehq/ | 22k | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 20 | Master end-to-end testing with Playwright and Cypress to build reliable test suites that catch bugs, improve confidence, and enable fast deployment. | try-works/ | 117 | 13 repos | ~990 | Automated safety check: Pass | Unknown | today |
| 21 | 21.E2E Agentic end-to-end tests with e2e, the e2e runner. An agent skill from kortix-ai/suna. | kortix-ai/ | 20k | — | ~2.3k | Automated safety check: Pass | Unknown | today |
| 22 | Explains how to run go-redis tests: the Docker Compose stack, make targets, focusing a single Ginkgo spec, the e2e suite and the version environment variables. | redis/ | 22k | — | ~786 | Automated safety check: Pass | BSD-2-Clause | yesterday |
| 23 | Reproduces iPolloWork enterprise TLS behavior in a Daytona Windows sandbox by installing a fake corporate CA and proving the app and its runtimes use the OS trust store. | Devin-AXIS/ | 6.7k | 1 repo | ~3.2k | Automated safety check: Pass | Unknown | yesterday |
| 24 | 24.E2E Testing This skill should be used when the user asks to "add an E2E test", "run live E2E", "run mock-LLM tests", "debug Playwright CI", "test the Docker image", or changes tests/e2e, Playwright configs, E2E… | OpenHands/ | 90k | — | ~308 | Automated safety check: Pass | MIT | today |
| 25 | 25.E2E Testing A skill your agent uses when an InsForge maintainer has finished an OSS repo change and is ready to open, update, or submit the InsForge PR. | InsForge/ | 13k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | 26.E2E Run end-to-end OTA verification for examples/v0.85.0 with agent-device. | gronxb/ | 1.8k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 27 | 27.Senior QA Comprehensive QA and testing skill for quality assurance, test automation, and testing strategies for ReactJS, NextJS, NodeJS applications. | nicepkg/ | 192 | 3 repos | ~1.1k | Automated safety check: Notes | No licence | 7 mo ago |
| 28 | Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools. | kurotu/ | 373 | 3 repos | ~615 | Automated safety check: Pass | MIT | 7 days ago |
| 29 | Converts RStudio Python Selenium electron tests into TypeScript Playwright tests, checking each against a live RStudio before counting it as migrated. | rstudio/ | 5.1k | — | ~3.6k | Automated safety check: Pass | Unknown | today |
| 30 | Runs Cherry Studio's critical-path regression suite as deterministic Playwright E2E tests through a GitHub workflow on macOS and Windows runners. | CherryHQ/ | 52k | — | ~1.2k | Automated safety check: Pass | AGPL-3.0 | today |
| 31 | 31.E2E Selects and runs the appropriate AgentsMesh end-to-end suite for Web, Desktop, MCP, or iOS, including worktree-specific environment setup and browser-level verification. | AgentsMesh/ | 2.4k | — | ~629 | Automated safety check: Notes | Unknown | 14 days ago |
| 32 | Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests. | comet-ml/ | 22k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | A skill your agent uses when UI changes are complete and e2e tests need updating. | payloadcms/ | 45k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 34 | 34.Supergoal Plan and autonomously build a software task end-to-end. An agent skill from robzilla1738/supergoal. | robzilla1738/ | 679 | — | ~10k | Automated safety check: Pass | MIT | 1 mo ago |
| 35 | Starts a local PostgreSQL 16 container, loads SQL fixtures and runs the Airbyte spec, check, discover and read commands against a chosen source-postgres image. | airbytehq/ | 22k | — | ~2.5k | Automated safety check: Pass | Unknown | today |
| 36 | Write and maintain Playwright end-to-end tests for the Onyx application. | onyx-dot-app/ | 32k | 1 repo | ~2.8k | Automated safety check: Notes | Unknown | today |
| 37 | Add, complete, or audit a locale across a Flutter application's ARB catalogs and a localized Docusaurus manual, including generated localization code, locale selectors, native platform declarations… | matthiasn/ | 1.2k | — | ~1.2k | Automated safety check: Pass | GPL-3.0 | today |
| 38 | Writes and reviews Geb browser automation specs with Spock, using Page Objects, at checkers, Modules and explicit waits, for Groovy and Gradle projects. | apache/ | 1.2k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 39 | MUST USE for any hotel or accommodation intent in any language, including hotel search, hotel recommendations, nearby accommodation, hostels, guesthouses, resorts, where-to-stay questions, room… | tourmind-com/ | 1.6k | — | ~13k | Automated safety check: Pass | MIT | 13 days ago |
| 40 | Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence. | XiaomiMiMo/ | 14k | — | ~3.9k | Automated safety check: Pass | MIT | 4 days ago |
| 41 | A skill your agent uses for test selection and quality across backend, frontend, and e2e; load its backend reference for Vitest projects, mocking, DB fixtures, and performance. | archestra-ai/ | 4.3k | — | ~3.2k | Automated safety check: Pass | Unknown | today |
| 42 | Generate and execute PR-aware OTA E2E scenarios for examples/v0.85.0 by diffing the checked-out branch against its PR base branch or default branch, inferring the affected runtime, rollout, and… | gronxb/ | 1.8k | — | ~1k | Automated safety check: Pass | Unknown | today |
| 43 | A skill your agent uses when the user wants to actually PLACE, manage, or redeem bets on Polymarket (not just read odds — that's the blockrunpredexon data tools). | BlockRunAI/ | 6.6k | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 44 | 44.E2E Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh. | stackia/ | 2.2k | — | ~517 | Automated safety check: Pass | GPL-2.0 | 5 days ago |
| 45 | Runs the Airbyte protocol sweep for a source-snowflake image against the shared Snowflake integration account, in an isolated per-run schema. | airbytehq/ | 22k | — | ~918 | Automated safety check: Pass | Unknown | today |
| 46 | Run and debug Agmente iOS end-to-end tests against a real local Codex CLI app-server instance. | rebornix/ | 545 | — | ~625 | Automated safety check: Pass | MIT | 4 mo ago |
| 47 | Continuously maintain automatic NemoClaw main E2E results through coordinated repairs. | NVIDIA/ | 23k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 48 | Build a capability end-to-end — backend vertical slice (Contracts→handler→validator→endpoint) AND the React page wired to it. | fullstackhero/ | 6.8k | — | ~783 | Automated safety check: Pass | MIT | 7 days ago |
Questions, answered from the data.
What is the best end-to-end testing skill?
Web Application Testing (official) from anthropics/skills ranks first of the 783 end-to-end testing skills listed here, with the highest score: its repository has 180k GitHub stars, 51 other GitHub owners carry a copy, its SKILL.md loads about 966 tokens and it passes the automated safety check with no findings. Next come Electron App Automation and OpenHarness End-to-End Evals.
Which end-to-end testing skills are official?
58 of the 783 end-to-end testing skills are official, published by the vendor's own GitHub organization: Web Application Testing, Electron App Automation, playwright-cli Browser Automation, MongoDB Source Connector E2E Harness, Java SDK E2E Test with Replay Snapshot and 53 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.