Web Application Testing
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
End-to-end validation of coordinator and agent template changes
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add bradygaster/squad --skill e2e-template-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bradygaster/squad e2e-template-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .claude/skills/e2e-template-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .claude/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bradygaster/squad --skill e2e-template-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bradygaster/squad e2e-template-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .agents/skills/e2e-template-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .agents/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradygaster/squad --skill e2e-template-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bradygaster/squad e2e-template-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .cursor/skills/e2e-template-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .cursor/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bradygaster/squad.git --path .squad-templates/skills/e2e-template-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bradygaster/squad --skill e2e-template-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bradygaster/squad e2e-template-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .gemini/skills/e2e-template-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .gemini/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bradygaster/squad e2e-template-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bradygaster/squad --skill e2e-template-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .github/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .github/skills/e2e-template-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .github/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradygaster/squad --skill e2e-template-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bradygaster/squad e2e-template-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bradygaster/squad.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.squad-templates/skills/e2e-template-testing .opencode/skills/e2e-template-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "e2e-template-testing" agent skill from https://github.com/bradygaster/squad/tree/dev/.squad-templates/skills/e2e-template-testing into .opencode/skills/e2e-template-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-template-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
e2e-template-testingEnd-to-end validation of coordinator and agent template changes
E2E Template Testing is an agent skill from bradygaster/squad. End-to-end validation of coordinator and agent template changes
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering End-to-end testing. The repository describes itself as: Squad: AI agent teams for any project. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1fd7e03. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitnpmghFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, npm and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
E2E Template Testing loads about 5.9k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 2,288 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
character (`⟨U+FEFF⟩`) at the start of the comment.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from bradygaster/squad at commit 1fd7e03, republished under its MIT licence (© bradygaster). 2,288 words, ~5,886 tokens.
.claude/skills/e2e-template-testing/SKILL.md (or your agent's skills folder).Squad's coordinator prompt (squad.agent.md) and agent charters (e.g.
scribe-charter.md) are shipped as templates in .squad-templates/. Changes to
these files affect how every squad session behaves — but unit tests can't catch
prompt-level regressions because the prompts are interpreted by an LLM at
runtime.
This skill describes how to validate template changes end-to-end by running real squad sessions against a locally-built CLI that includes your modified templates.
.squad-templates/squad.agent.md (coordinator prompt).squad-templates/scribe-charter.md or other agent charters.squad-templates/notes-protocol.md or helper scriptscopilot or ghcs) installedIf PR_NUMBER and REPO are both set, the absolute first thing you do — before
fast-fail checks, before building, before creating any repos — is post the initial
tracking comment with all steps marked as :hourglass_flowing_sand: Pending.
This gives reviewers immediate visibility that a run is in progress and what to expect.
$runStart = Get-Date
$body = @"
## E2E Progress - PR $env:PR_NUMBER
| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build :cd: link :cd: ``squad version``) | :hourglass_flowing_sand: Pending | --:-- | -- |
| 2. Create test repo(s) | :hourglass_flowing_sand: Pending | --:-- | -- |
| 3. ``squad init`` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- |
| 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- |
| 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- |
| 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- |
| Symbol | Meaning |
|---|---|
| :hourglass_flowing_sand: | Not started |
| :arrows_counterclockwise: | Running |
| :white_check_mark: | Passed |
| :x: | Failed |
| :warning: | Passed with caveats |
*Run started: $($runStart.ToString('HH:mm')) — all steps pending*
"@
$tmpFile = [System.IO.Path]::GetTempFileName()
$utf8NoBom = New-Object System.Text.UTF8Encoding $false
[System.IO.File]::WriteAllText($tmpFile, $body, $utf8NoBom)
$response = gh api "repos/$env:REPO/issues/$env:PR_NUMBER/comments" --method POST --field "body=@$tmpFile" | ConvertFrom-Json
$env:COMMENT_ID = $response.id
Remove-Item $tmpFile -Force
Write-Host "Progress comment posted — ID: $($response.id)"After posting, immediately update the comment to mark Step 1 as :arrows_counterclockwise: Running (do NOT wait — this is a two-step sequence: post all-pending, then immediately update to Step 1 running). Then proceed to Step 1.
See Progress Reporting for the full comment lifecycle and update patterns.
⚠️ STOP before continuing: If posting the comment fails (network error, auth error), abort the run and report the failure. Do not proceed silently without a tracking comment.
cd /path/to/squad # your feature branch
npm install
npm run build -w packages/squad-sdk && npm run build -w packages/squad-cli
# Link so `squad` command uses your local build (workspace flag — no cd required)
npm link -w packages/squad-cliVerify: squad version output includes the -preview suffix (e.g., x.y.z-preview),
confirming the local dev build is active. If the output shows a plain semver without
-preview, the globally-installed npm package is still in use — re-check the link step.
See CONTRIBUTING.md — Making the squad Command Use Your Local Build
for the full guidance on local dev versioning.
mkdir /tmp/sq-test-1 && cd /tmp/sq-test-1
git init
echo "# Test Project" > README.md
echo '{"name":"test-project","version":"1.0.0"}' > package.json
mkdir src
echo "export function hello() { return 'world' }" > src/index.ts
git add -- README.md package.json src/index.ts
git commit -m "init: test project"Keep the project small — you only need enough for the coordinator to recognize a codebase and hire a team.
squad init
# If testing a specific feature (e.g. state backends):
# squad init --state-backend git-notesVerify the init produced the expected files:
ls -la .squad/
cat .squad/team.md # should have ## Members with 3+ agents
cat .squad/config.json # should reflect any CLI flags you passedUse the Copilot CLI's -p flag with --allow-all-tools for non-interactive sessions.
--allow-all-tools is required for automated/non-interactive runs — without it,
tool calls (including file writes) prompt for confirmation and block.
# PowerShell (Windows)
copilot --agent squad --allow-all-tools -p "Lead, decide what testing framework to use. Write your decision." `
2>&1 | Tee-Object evidence/session-task.log# Bash (macOS/Linux)
copilot --agent squad --allow-all-tools -p "Lead, decide what testing framework to use. Write your decision." \
2>&1 | tee evidence/session-task.logAlternatively, set the COPILOT_ALLOW_ALL=1 environment variable instead of the flag.
For multi-turn workflows, run sequential sessions:
# Session A: give the team a task
copilot --agent squad --allow-all-tools -p "prompt A" 2>&1 | Tee-Object evidence/session-A.log
# Session B: verify state persisted
copilot --agent squad --allow-all-tools -p "What decisions has the team made?" 2>&1 | Tee-Object evidence/session-B.logCheck that your template change had the expected effect. Common checks:
# State location (for state-backend changes)
git notes --ref=squad list # git-notes backend
git ls-tree -r squad-state # orphan backend
ls .squad/agents/*/history.md # worktree backend
# Coordinator behavior (grep session log)
grep "STATE_BACKEND" evidence/session-task.log
grep "spawn" evidence/session-task.log
# File tree diff
git diff --stat HEAD~1 # what changed on working branch
git log --all --oneline # commits across all branchesCreate an evidence/verdict.md in each test repo:
## Test: [scenario name]
**Backend:** worktree | git-notes | orphan | two-layer
**Branch:** [your feature branch]
**Result:** PASS | PARTIAL | FAIL
**Duration:** Xm Ys
### What was verified
- [ ] Coordinator identified feature correctly (from session log)
- [ ] Agent was spawned via `task` tool (not simulated)
- [ ] team.md has ## Members with 3+ agents
- [ ] State landed in correct location
- [ ] No unexpected side effects
### Evidence files
- session-task.log — full session output
- git-log.txt — `git log --all --oneline`
### Notes
[anything unusual or noteworthy]Record the wall-clock time from the start of Step 1 (fast-fail checks) to the end of Step 6 (verdict posted). This is the full E2E run duration for this scenario.
Use this section only when you are running E2E validation for an open PR. If
PR_NUMBER and REPO are both set, post and maintain a live tracking comment
in the PR thread. If either value is missing (for example, a local-only run),
skip progress reporting silently.
The initial comment must be posted as Step 0 — the absolute first action before anything else. See the Step 0 block in the Workflow section for the exact code.
The subsections below describe how to update the comment at each step boundary. For reference, here is the initial all-pending comment body posted in Step 0:
gh pr comment "$PR_NUMBER" --repo "$REPO" --body "## E2E Progress\n\n| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build · link · \\`squad version\\`) | ⏳ Pending | --:-- | -- |
| 2. Create test repo(s) | ⏳ Pending | --:-- | -- |
| 3. \\`squad init\\` + file verification | ⏳ Pending | --:-- | -- |
| 4. Run sessions | ⏳ Pending | --:-- | -- |
| 5. Verify outcomes | ⏳ Pending | --:-- | -- |
| 6. Record verdicts + post final comment | ⏳ Pending | --:-- | -- |
\n| Symbol | Meaning |
|---|---|
| ⏳ | Not started |
| 🔄 | Running |
| ✅ | Passed |
| ❌ | Failed |
| ⚠️ | Passed with caveats |"COMMENT_ID=$(gh api "repos/$REPO/issues/$PR_NUMBER/comments" --jq '.[-1].id')🔄 Running and every later step remains ⏳ Pending.🔄 Running, record $startTime = Get-Date and store the
HH:MM start time in that row's Started column.gh api --method PATCH "repos/$REPO/issues/comments/$COMMENT_ID" --field body="..."✅, ❌, or ⚠️, compute
$duration = (Get-Date) - $startTime and format it as
"{0}m {1}s" -f [int]$duration.TotalMinutes, $duration.Seconds.✅, ❌, or ⚠️, keep its original
HH:MM value in Started, and write the formatted duration in Duration.🔄 Running and set its Started value.⏳ Pending with --:-- for Started and -- for
Duration.❌ with its original start time and computed duration, and Step 6
becomes 🔄 Running while you prepare the final verdict.| Symbol | Meaning |
|---|---|
| ⏳ | Not started |
| 🔄 | Running |
| ✅ | Passed |
| ❌ | Failed |
| ⚠️ | Passed with caveats |
Keep these six rows in this exact order every time you update the comment:
squad version)squad init + file verificationOn Windows PowerShell 5.1, use the --field body=@file pattern to post comment
bodies. Write the content to a temp file using UTF-8 without BOM, then pass
--field "body=@$tmpFile" to gh api. This is more reliable than piping JSON
through --input - on PS 5.1, which can silently corrupt multi-byte characters
even with [Console]::OutputEncoding = UTF8.
Key rules:
New-Object System.Text.UTF8Encoding $false (the $false disables the BOM).
[System.Text.Encoding]::UTF8 writes a BOM which GitHub renders as a stray
character (``) at the start of the comment.--field "body=@$tmpFile", NOT --input - or --input filename, for
comment body updates. The @ prefix tells gh to read the field value from
the file rather than treating the path as a literal string.$step1StartTime = Get-Date
$step1Started = $step1StartTime.ToString('HH:mm')
$step1Duration = (Get-Date) - $step1StartTime
$step1DurationText = "{0}m {1}s" -f [int]$step1Duration.TotalMinutes, $step1Duration.Seconds
$step2StartTime = Get-Date
$step2Started = $step2StartTime.ToString('HH:mm')
$body = @"
## E2E Progress
| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build · link · `squad version`) | :white_check_mark: Passed | $step1Started | $step1DurationText |
| 2. Create test repo(s) | :arrows_counterclockwise: Running | $step2Started | -- |
| 3. `squad init` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- |
| 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- |
| 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- |
| 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- |
| Symbol | Meaning |
|---|---|
| :hourglass_flowing_sand: | Not started |
| :arrows_counterclockwise: | Running |
| :white_check_mark: | Passed |
| :x: | Failed |
| :warning: | Passed with caveats |
"@
$tmpFile = "$env:TEMP\e2e-comment-body.md"
$utf8NoBom = New-Object System.Text.UTF8Encoding $false
[System.IO.File]::WriteAllText($tmpFile, $body, $utf8NoBom)
gh api --method PATCH "repos/$env:REPO/issues/comments/$env:COMMENT_ID" --field "body=@$tmpFile"
Remove-Item $tmpFile -ForceDo NOT batch all scenario results to the end. This is the most common cause of lost verdicts. After each scenario completes, immediately PATCH the tracking comment with that scenario's result before moving to the next one.
The pattern for each scenario:
# After scenario N completes — PATCH immediately, before starting scenario N+1
$scenarioNDuration = "{0}m {1}s" -f [int]((Get-Date) - $scenarioNStartTime).TotalMinutes, ((Get-Date) - $scenarioNStartTime).Seconds
# ...rebuild the full comment body with this scenario updated to PASS/FAIL/PARTIAL...
$tmpFile = [System.IO.Path]::GetTempFileName()
$utf8NoBom = New-Object System.Text.UTF8Encoding $false
[System.IO.File]::WriteAllText($tmpFile, $body, $utf8NoBom)
gh api --method PATCH "repos/$env:REPO/issues/comments/$env:COMMENT_ID" --field "body=@$tmpFile"
Remove-Item $tmpFile -Force
Write-Host "Scenario N verdict posted"This guarantees that even if the AI model connection drops mid-run, the last successfully PATCHed state is always visible in the PR.
⚠️ Critical: Background agents lose their AI model connection after ~15 minutes of continuous execution. This is a platform limit, not a bug in your code. The verdict stage appears to "hang" because the connection drops right at the end when the agent has been running too long.
Per-agent scenario budget:
| Scenario type | Estimated time | Budget |
|---|---|---|
| Static checks only (file existence, grep, size) | 1-3 min | 4 per agent |
squad init + file verification (no copilot session) | 3-5 min | 3 per agent |
squad init + one copilot --agent squad session | 8-15 min | 1 per agent |
| Build + link + one copilot session | 12-20 min | 1 per agent |
Rule: Limit yourself to 1 scenario that includes a copilot --agent squad session
per agent run. For a plan with multiple copilot-session scenarios, run them in
separate agents — not in sequence within a single agent.
If your scenario plan has N copilot-session scenarios, request N separate test agents to run them in parallel (one scenario each). Static scenarios may be batched up to 4 per agent.
If you are running a scenario with a copilot --agent squad session:
When you reach Step 6, replace the tracking comment body entirely with the final structured verdict table. Do not post a separate final comment. The tracking comment is the final verdict comment.
Include a summary row at the bottom of the final table showing the total elapsed time for the full run:
| **Total** | — | HH:MM | Xm Ys |If the connection drops before Step 6: The last progressive verdict PATCH already shows the partial state. The next agent run should read the existing comment, pick up where it left off, and add remaining scenario rows rather than starting fresh.
Use this matrix when planning validation for a template change. Not every change needs every row — pick the scenarios relevant to your modification.
| # | Scenario | What to verify | Duration |
|---|---|---|---|
| 1 | Basic init + task | Templates applied, agent spawned, work produced | — |
| 2 | Cross-branch persistence | State survives git checkout (if state-backend) | — |
| 3 | Scribe behavior | Scribe commits to correct target | — |
| 4 | PR cleanliness | Feature branch PR has no leaked state files | — |
| 5 | Migration path | Existing squad picks up new template behavior | — |
| 6 | Edge case: empty repo | Init works in repo with single commit | — |
| 7 | Edge case: monorepo | Init works in subdirectory of monorepo | — |
Note: Keep — during planning, then replace it with the actual elapsed time when
recording the verdict for each scenario.
sq-test-notes-crossbranch, not test1.Tee-Object replaces tee. Paths use \.These checks must pass before running any scenario. If any fail, stop immediately and report the failure — do not attempt workarounds or mark scenarios as SKIPPED.
npm run build, remove any
stale published copy of @bradygaster/squad-sdk that may have been installed
into packages/squad-cli/node_modules/. This local copy shadows the workspace
symlink in the root node_modules/ and causes TypeScript to see the published
version instead of the local source. Run from the repo root:$stale = "packages\squad-cli\node_modules\@bradygaster\squad-sdk"
if (Test-Path $stale) { Remove-Item -Recurse -Force $stale; Write-Host "Cleaned stale SDK" }node_modules\@bradygaster\squad-sdk workspace symlink remains
intact and npm will use it automatically.npm run build from the repo root. A build
failure blocks all scenarios; report BUILD_FAILED and stop.tsc: not found or similar missing-binary errors, run
npm install first to reconcile node_modules with the lock file, then
retry. This can happen after git checkout HEAD -- package-lock.json
restores the lock file without reinstalling.cd packages/squad-cli && npm link must exitLINK_FAILED and stop.squad version must run. After linking, squad version must output a
version string. If not, report CLI_NOT_FOUND and stop.Do not mark scenarios as SKIPPED due to build or environment errors — that obscures real failures from reviewers. SKIPPED is only acceptable when the user explicitly requests it.
When posting evidence to PR comments, issues, or any shared document:
C:\Users\username\... or /home/username/...).~ notation for home-relative paths: ~/AppData/Local/Temp/...
or ~/tmp/sq-test-1.<repo-root> as the prefix. If evidence files live
inside the repository (e.g. .e2e/, tmp/, or any subdirectory of the repo),
write them as <repo-root>\.e2e\pr-1035\evidence — not as ~\..\..\... backward
navigation. <repo-root> is a clear, portable placeholder for the repository root
that does not expose the machine's directory layout.~.Example — ❌ wrong: C:\Users\johndoe\AppData\Local\Temp\sq-e2e-pr1035\evidence
Example — ✅ right: ~/AppData/Local/Temp/sq-e2e-pr1035/evidence
Example — ❌ wrong: ~\..\..\repos\squad\.e2e\pr-1035\evidence
Example — ✅ right: <repo-root>\.e2e\pr-1035\evidence
This applies to all evidence tables, verdict files, and PR comments.
~-relative paths
before sharing. See PII Protection above.copilot --agent squad sessions in one agent. Each
session takes 5-15 minutes; combined with build time, you'll hit the ~15-minute
connection budget. One copilot session per agent — split into parallel agents
if your plan has more.--allow-all-tools in non-interactive modeThe Copilot CLI requires explicit permission to run tools automatically. In
interactive mode the user approves each tool call; in non-interactive mode (-p)
those prompts cannot be displayed and writes fail silently or with a "Permission
denied and could not request permission from user" error.
Fix: always include --allow-all-tools (or --yolo / --allow-all) in Step 4
commands, or export COPILOT_ALLOW_ALL=1 before running E2E sessions.
This also applies when copilot --agent squad is launched as a subprocess from
inside a Copilot CLI background agent (e.g. a test engineer running via the task tool) —
the flag is still needed.
--allow-all-paths for repos outside the CWDBy default the CLI restricts file access to the current directory tree. If the
coordinator needs to read files from a parent repo while running in a disposable
test repo, add --allow-all-paths:
copilot --agent squad --allow-all-tools --allow-all-paths -p "..."Or use the combined shorthand: --allow-all / --yolo.
high — Validated through 12 real E2E test sessions during state-backend
development (PR #1004). --allow-all-tools requirement confirmed in PR #1035.
© bradygaster, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. 1 hidden character (zero-width or bidirectional) removed. Raw file
Just SKILL.md in .squad-templates/skills/e2e-template-testing of bradygaster/squad.
Open the folder on GitHubat commit 1fd7e03
E2E Template Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| E2E Template Testing this skillbradygaster/squad | 3.3k | — | ~5.9k | Automated safety check: Warn | MIT | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph | 112 | 11 repos | ~2.4k | Automated safety check: Pass | None | |
| Uloop Replay Inputkurotu/VRCQuestTools | 373 | 3 repos | ~615 | Automated safety check: Pass | MIT | |
| Ui4 Convert Testspayloadcms/payload | 45k | — | ~3.5k | Automated safety check: Pass | MIT | |
| E2Estackia/rtp2httpd | 2.2k | — | ~517 | Automated safety check: Pass | GPL-2.0 |
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
hellangleZ/burn-in-cceverywhere-ralph
A skill your agent uses when writing new features, fixing bugs, or refactoring code.
kurotu/VRCQuestTools
Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.
payloadcms/payload
A skill your agent uses when UI changes are complete and e2e tests need updating.
stackia/rtp2httpd
Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.
MotherofallVPNs/MoaV
Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.
bradygaster/squad
Review and validate claims using counter-hypothesis testing.
bradygaster/squad
How to review PRs for architectural quality — module boundaries, dependency direction, export surface, pattern consistency
bradygaster/squad
Preserve content when moving entries between tracked Squad state files
bradygaster/squad
Defensive CI/CD patterns: semver validation, token checks, retry logic, and draft detection
bradygaster/squad
Checklist and patterns for wiring new CLI commands into cli-entry.ts
bradygaster/squad
Enables squad agents on different machines to share work via git-based task queuing
Categories
End-to-end validation of coordinator and agent template changes. E2E Template Testing is an agent skill from bradygaster/squad.
E2E Template Testing fits situations like: tasks that involve End-to-end testing.
Run `npx skills add bradygaster/squad --skill e2e-template-testing -a claude-code`. Or copy the skill folder (.squad-templates/skills/e2e-template-testing in bradygaster/squad) into .claude/skills/e2e-template-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bradygaster/squad --skill e2e-template-testing -a codex`. Or copy the skill folder (.squad-templates/skills/e2e-template-testing in bradygaster/squad) into .agents/skills/e2e-template-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bradygaster/squad --skill e2e-template-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-template-testing, .gemini/skills/e2e-template-testing, .github/skills/e2e-template-testing and .opencode/skills/e2e-template-testing in your project.
Going by SKILL.md and its folder, E2E Template Testing needs the command-line tools its instructions call (git, npm and gh). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use git, npm and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): contains zero-width characters. Read the flagged lines before installing; the check is not a guarantee either way.
E2E Template Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with E2E Template Testing: Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Uloop Replay Input (kurotu/VRCQuestTools, 373 stars) and Ui4 Convert Tests (payloadcms/payload, 45k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bradygaster (a GitHub user) maintains it in bradygaster/squad, which has 3,257 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.
Source: bradygaster/squad on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.