Agent skill

Full Stack E2E Review

by BetterWright in BetterWright/betterwright

Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX.

MITAuto-check passedTesting & QA

Install Full Stack E2E Review

skills CLI
$ npx skills add BetterWright/betterwright --skill full-stack-e2e-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BetterWright/betterwright full-stack-e2e-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BetterWright/betterwright.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/full-stack-e2e-review .claude/skills/full-stack-e2e-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
full-stack-e2e-review
GitHub stars
329
Token cost
~2.7k tokens
SKILL.md length
1,489 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX.

  • Works in 8 steps: Establish the real system and review… → Model realistic fixtures → Boot and prove the stack → …
  • Requests such as review this branch end to end
  • SKILL.md covers 1. Establish the real system…, Browser harness, 2. Model realistic fixtures and 3. Boot and prove the stack, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Full Stack E2E Review is an agent skill from BetterWright/betterwright. Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX. Use for requests such as review this branch end to end, test everything, ensure no bugs, perform a full visual pass, or validate a multi-user workflow before merge or deploy.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing. The repository describes itself as: A persistent, policy-guarded Playwright browser for AI agents — network policy, encrypted credential vault, proof screenshots, and CAPTCHA solving. The licence is MIT.

When your agent uses it

  • Requests such as review this branch end to end
  • Test everything
  • Perform a full visual pass
  • Validate a multi-user workflow before merge

Example prompts

  • “/full-stack-e2e-review”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Establish the real system and review boundary
  2. Model realistic fixtures
  3. Boot and prove the stack
  4. Review code for missing invariants
  5. Execute the proof ladder
  6. Perform a true visual and interaction pass
  7. Close the loop on every failure
  8. Completion gate and report

What it can do on your machine

Read from SKILL.md and the folder at commit ec55a32. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Full Stack E2E Review loads about 2.7k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BetterWright/betterwright at commit ec55a32, republished under its MIT licence (© BetterWright). 1,489 words, ~2,725 tokens.

Download SKILL.mdSave it as .claude/skills/full-stack-e2e-review/SKILL.md (or your agent's skills folder).
name
full-stack-e2e-review
description
Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX. Use for requests such as review this branch end to end, test everything, ensure no bugs, perform a full visual pass, or validate a multi-user workflow before merge or deploy.

Full-stack end-to-end review

Treat the task as a closed-loop product validation, not as "run the existing test suite." Build evidence across code, live services, APIs, and ordinary browser interactions. Do not claim full coverage while a requested layer remains untested.

Hosts load only this skill's name and description until the user asks for an end-to-end review. When that happens, follow the ladder below. Load BetterWright's browser skill or MCP tool guidance before driving the browser — this playbook does not repeat that protocol.

1. Establish the real system and review boundary

  1. Read repository instructions and relevant skills first.
  2. Inspect the branch diff, recent commits, migrations, environment files, and existing tests.
  3. If a prior agent, branch, PR, or environment booted the stack, resolve its exact refs and inspect its environment metadata, setup scripts, transcript/start logs, and diff before reuse. Record both setup and feature refs; avoid copying setup files into the reviewed working tree unless necessary.
  4. Map the feature through every affected layer:
    • data model and migrations
    • authorization and ownership rules
    • backend services, routes, jobs, and third-party gates
    • CLI or SDK surfaces
    • frontend state and actions
    • billing, quotas, lifecycle, and recovery paths
  5. Write a short review matrix before testing. Its rows are user journeys or invariants; its columns are code inspection, focused test, live API, desktop UI, and mobile UI. Mark every cell passed, failed, not applicable, or blocked.

Use independent subagents in parallel for bounded exploration, setup reconstruction, or visual review. Keep blockers and the final integration in the parent. Give subagents exact files, URLs, actors, fixtures, expected results, prohibited side effects, and required evidence.

Browser harness

This skill defines what to prove, not one vendor's browser syntax.

When BetterWright is available through MCP, CLI, or a native tool:

  1. Verify the integration with browser_doctor or betterwright doctor, then prove navigation and screenshot output before the review. Confirm its network policy permits the local test hosts; loopback must remain allowed for localhost applications.
  2. Use the host agent's normal shell, repository, file, and test-runner tools for environment setup, database fixtures, raw API checks, code edits, and durable regression tests. BetterWright's browser sandbox has no host filesystem, process, module-loader, or unrestricted network-interception access and is not a replacement for Playwright Test or the project's test runner.
  3. Use one named BetterWright session per independent browser lane. Sessions share a cookie jar, so use separate BetterWright profiles for different user identities. For parallel role testing over MCP, configure separate server aliases with distinct BETTERWRIGHT_PROFILE values; otherwise test roles sequentially with explicit sign-out/sign-in and verify identity after every switch.
  4. Capture debug images while diagnosing and a visually inspected proof image for each important completed flow. Keep command output, API matrices, test results, and repository diffs as host-agent evidence.

With another browser harness, preserve the same semantics: fresh observed locators, action verification, isolated identities, mobile and desktop viewports, ordinary user controls, and proof screenshots.

2. Model realistic fixtures

Create deterministic local or test-only fixtures that expose boundary failures. Use the application's real persistence layer where practical; do not mock away the behavior being reviewed.

Cover the smallest meaningful matrix:

  • actors: owner, admin, member, outsider, and empty/new user
  • scopes: personal and shared/tenant scope
  • resources: owned by the actor, owned by another member, archived or failed, and empty state
  • account states: active, not started/free, past due or suspended when relevant
  • collaboration: pending invitation and at least two members
  • limits: below limit, at limit, and disallowed action

Use fake credentials and test-only auth helpers. For OAuth-only applications, seed test users and mint local sessions through the application's supported session mechanism without bypassing authorization middleware. Never reuse or expose production secrets. Make fixture scripts idempotent so they can be rerun after fixes.

3. Boot and prove the stack

  1. Inspect install and database-preparation scripts before running them. Confirm all database URLs and external services are disposable local/test resources; never reset, migrate, or seed shared or production infrastructure without explicit approval.
  2. Start long-lived services through a persistent, inspectable mechanism. Keep startup commands separate from health checks.
  3. Verify each required port or health endpoint and inspect recent logs. A running process is not proof that the service is healthy.
  4. Record the exact URLs, ports, test actors, and known unavailable integrations.
  5. Distinguish harness problems from product failures. Repair the harness only when the change is isolated, reversible, and within the requested scope; otherwise mark it blocked.

4. Review code for missing invariants

Do not rely only on tests that already exist. Trace each journey and actively search for holes such as:

  • scope identifiers accepted inconsistently across endpoints
  • owner-only fields, URLs, tokens, IPs, or secrets leaking after partial redaction
  • UI actions enabled even though the API will reject them
  • list, detail, metrics, snapshots, search, or billing endpoints using different tenancy fences
  • actor identity confused with billing subject or resource owner
  • permission checks missing on mutation paths
  • "personal" behavior differing between UI, CLI, API, and stored identifiers
  • stale, loading, empty, partial-error, retry, and past-due states
  • migration, rollback, idempotency, concurrency, and post-credit or post-plan behavior

For sensitive fields, assert absence as well as presence. For mutations, verify both the response and persisted state.

Show full SKILL.md (631 more words)Show less

5. Execute the proof ladder

Run from narrow and fast to broad and realistic:

  1. Focused unit or integration tests around changed invariants.
  2. Typecheck, lint, or compile every affected package.
  3. Live HTTP/API checks against the running stack.
  4. CLI or SDK checks if the feature is exposed there.
  5. Real browser interaction and visual review.
  6. Broader regression suites after any fix.

For live API checks, build an explicit actor-by-scenario matrix. Verify status code and response body for positive and negative cases, including:

  • personal versus shared scope
  • owner versus non-owner resource access
  • each role's read and mutation rights
  • outsider isolation
  • invitation, billing, quota, and lifecycle gates
  • sensitive-field redaction
  • consistency across list, detail, metrics, snapshots, and related endpoints

Do not use a destructive billing, invitation, deletion, or external side effect merely to prove a button works. Prefer local fakes, dry runs, seeded states, or stop before submission and mark the final step blocked when approval is required.

6. Perform a true visual and interaction pass

Use ordinary browser controls. Click through; do not infer UI behavior from source or screenshots alone.

For each actor and state:

  1. Start from a clean authenticated entry point.
  2. Walk navigation, switchers, tabs, menus, dropdowns, inline editors, dialogs, and back/refresh paths.
  3. Check enabled actions and disabled reasons against the permission matrix.
  4. Confirm submitted or selected state persists after navigation or reload when it should.
  5. Inspect loading flashes, error toasts, stale data, empty states, truncation, wrapping, contrast, hierarchy, alignment, copy consistency, and focus behavior.
  6. Test a normal desktop viewport and a narrow mobile viewport. Check drawers and overlays, not only static layout.
  7. Capture each distinct important state with a descriptive filename. Tie every visual finding to its screenshot.

A dedicated visual subagent should receive a numbered walkthrough with exact fixtures and expected states. Its return must include what worked, bugs by severity, role/action confirmation, interaction failures, and screenshot paths.

7. Close the loop on every failure

For each failure:

  1. Reproduce it with the smallest reliable case.
  2. Classify it as product bug, test bug, fixture problem, environment problem, or expected limitation.
  3. Trace the failure across layers. If edits are authorized, make the smallest coherent root-cause fix and add or strengthen a regression test; otherwise report the proposed fix and evidence without modifying product code.
  4. Rerun the targeted check.
  5. Rerun affected broader tests, typechecks, and builds.
  6. Recheck the corresponding live API and UI path.
  7. Review git status and diff; remove incidental artifacts such as generated files or lockfile churn that are unrelated to the fix.

Never commit, push, open or update a PR, upgrade global toolchains, or mutate shared services unless explicitly requested. Keep dependency-install and generated-file churn out of the reviewed diff.

Do not stop at "tests pass" if the user asked for an end-to-end or visual review. Do not silently weaken an assertion to make a test green.

8. Completion gate and report

The review is complete only when every requested journey is passed or explicitly marked [blocked] with the missing dependency or approval.

Return:

  1. Scope tested: branch/commit, services, actors, states, and viewports.
  2. What worked: concise confirmed behavior.
  3. Bugs found and fixed: severity, root cause, affected layers, regression coverage.
  4. Remaining findings: severity, reproduction steps, and evidence.
  5. Permission and state matrix: what each actor can see and do.
  6. Verification evidence: exact commands, test counts, builds/typechecks, API cases, and screenshot paths.
  7. Deploy blockers: migrations, infra, third-party dependencies, data safety, or untested production-only paths.
  8. Honest limitations: distinguish a fully exercised flow from code-inspected or locally simulated coverage.

Prefer concrete statements such as "member received 403 and no secret fields were returned" over "permissions look correct." Attach focused visual proof when browser work was performed.

Adapted from full-stack-e2e-review (MIT).

© BetterWright, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/full-stack-e2e-review of BetterWright/betterwright.

Open the folder on GitHubat commit ec55a32

Compare with similar skills

Full Stack E2E Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Full Stack E2E Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Full Stack E2E Review this skillBetterWright/betterwright329—~2.7kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Uloop Replay Inputkurotu/VRCQuestTools3733 repos~615Automated safety check: PassMIT
Ui4 Convert Testspayloadcms/payload45k—~3.5kAutomated safety check: PassMIT
E2Estackia/rtp2httpd2.2k—~517Automated safety check: PassGPL-2.0

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Uloop Replay Input

    kurotu/VRCQuestTools

    Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.

    373 GitHub starsUsed in 3 repos~615 tokens
    Testing & QAAuto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E

    stackia/rtp2httpd

    Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.

    2.2k GitHub stars~517 tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from BetterWright/betterwright

All 8 skills in this repo
  • Browser

    BetterWright/betterwright

    Drive a persistent, policy-guarded real web browser via the betterwright CLI.

    329 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • 1password

    BetterWright/betterwright

    Read this when the browser profile has the 1Password extension and a login, signup, or payment form should be filled with it.

    329 GitHub stars~618 tokensUpdated 8 days ago
    Auto-check passed
  • Credential Manager

    BetterWright/betterwright

    Read this before any login, signup, password change, or checkout so the right credential source is used in the right order without ever exposing a secret.

    329 GitHub stars~1.1k tokensUpdated 8 days ago
    Auto-check passed
  • Browser Console

    BetterWright/betterwright

    Diagnose browser console errors and uncaught JavaScript exceptions without dumping routine logs.

    329 GitHub stars~404 tokensUpdated 8 days ago
    Auto-check passed
  • Bitwarden

    BetterWright/betterwright

    Read this when the browser profile has the Bitwarden extension and a login, signup, or payment form should be filled with it.

    329 GitHub stars~408 tokensUpdated 8 days ago
    Auto-check passed
  • GitHub

    BetterWright/betterwright

    GitHub navigation, review, and account-context guidance for working repos, issues, and pull requests in the browser.

    329 GitHub stars~359 tokensUpdated 8 days ago
    Auto-check passed

Categories

Questions about Full Stack E2E Review

What does Full Stack E2E Review do?

Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX. Full Stack E2E Review is an agent skill from BetterWright/betterwright. Run a rigorous end-to-end product or feature review across architecture, data, APIs, permissions, billing, tests, and real browser UX.

When should I use Full Stack E2E Review?

Full Stack E2E Review fits situations like: requests such as review this branch end to end; test everything; perform a full visual pass; validate a multi-user workflow before merge.

How do I install Full Stack E2E Review in Claude Code?

Run `npx skills add BetterWright/betterwright --skill full-stack-e2e-review -a claude-code`. Or copy the skill folder (skills/full-stack-e2e-review in BetterWright/betterwright) into .claude/skills/full-stack-e2e-review in your project. Claude Code loads it when a task matches its description.

How do I install Full Stack E2E Review in Codex?

Run `npx skills add BetterWright/betterwright --skill full-stack-e2e-review -a codex`. Or copy the skill folder (skills/full-stack-e2e-review in BetterWright/betterwright) into .agents/skills/full-stack-e2e-review in your project. Codex loads it when a task matches its description.

Can I use Full Stack E2E Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BetterWright/betterwright --skill full-stack-e2e-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/full-stack-e2e-review, .gemini/skills/full-stack-e2e-review, .github/skills/full-stack-e2e-review and .opencode/skills/full-stack-e2e-review in your project.

What does Full Stack E2E Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Full Stack E2E Review is instructions for the agent only.

Does Full Stack E2E Review access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Full Stack E2E Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Full Stack E2E Review use?

Full Stack E2E Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Full Stack E2E Review use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Full Stack E2E Review?

Skills that share tags, products or a category with Full Stack E2E Review: Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Uloop Replay Input (kurotu/VRCQuestTools, 373 stars) and Ui4 Convert Tests (payloadcms/payload, 45k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Full Stack E2E Review?

BetterWright (a GitHub organization) maintains it in BetterWright/betterwright, which has 329 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 29, 2026.

Source: BetterWright/betterwright on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.