Web Application Testing
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install asgeirtj/system_prompts_leaks verify --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .claude/skills/verify && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .claude/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verifyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install asgeirtj/system_prompts_leaks verify --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .agents/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .agents/skills/verify && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .agents/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install asgeirtj/system_prompts_leaks verify --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .cursor/skills/verify && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .cursor/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/asgeirtj/system_prompts_leaks.git --path Anthropic/claude-code/skills/verify--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install asgeirtj/system_prompts_leaks verify --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .gemini/skills/verify && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .gemini/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install asgeirtj/system_prompts_leaks verifyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .github/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .github/skills/verify && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .github/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install asgeirtj/system_prompts_leaks verify --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .opencode/skills/verify && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "verify" agent skill from https://github.com/asgeirtj/system_prompts_leaks/tree/main/Anthropic/claude-code/skills/verify into .opencode/skills/verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "verify", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
verifyVerify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.
Verify is an agent skill from asgeirtj/system_prompts_leaks. Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck. Run before committing nontrivial changes; bootstraps this repo's project verify skill if none exists yet. Don't invoke it on a diff that only touches tests, docs, or other code with no runtime surface to drive (a change to product source always has one) — there's nothing to observe.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/cli.md` and `examples/server.md`).
It sits in Testing & QA. The repository describes itself as: Documented system prompts from Anthropic - Claude Fable 5.1, Opus 5.5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-6-Astra, Codex. Google - Gemini 3.8 Flash, 3.1 Pro… The licence is CC0-1.0.
Read from SKILL.md and the folder at commit 60d44cc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitghFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Verify loads about 3k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 1,449 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from asgeirtj/system_prompts_leaks at commit 60d44cc, republished under its CC0-1.0 licence (© asgeirtj). 1,449 words, ~3,045 tokens.
.claude/skills/verify/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Verification is runtime observation. You build the app, run it, drive it to where the changed code executes, and capture what you see. That capture is your evidence. Nothing else is.
Don't run tests. Don't typecheck. Running them here proves you can run CI — not that the change works. Not as a warm-up, not "just to be sure," not as a regression sweep after. The time goes to running the app instead.
Don't import-and-call. import { foo } from './src/...' then
console.log(foo(x)) is a unit test you wrote. The function did what
the function does — you knew that from reading it. The app never ran.
Whatever calls foo in the real codebase ends at a CLI, a socket, or
a window. Go there.
The scope is what you're verifying — usually a diff, sometimes just "does X work." In a git repo, establish the full range (a branch may be many commits, or the change may still be uncommitted):
git log --oneline @{u}.. # count commits (if upstream set)
git diff @{u}.. --stat # full range, not HEAD~1
git diff origin/HEAD... --stat # no upstream: committed vs base
git diff HEAD --stat # uncommitted: working tree vs HEAD
gh pr diff # if in a PR contextState the commit count. Large diff truncating? Redirect to a file then Read it. Repo but no diff from any of these → say so, stop. No repo → the scope is whatever the user named; ask if they didn't.
The diff is ground truth. Any description is a claim about it. Read both. If they disagree, that's a finding.
The surface is where a user — human or programmatic — meets the change. That's where you observe.
| Change reaches | Surface | You |
|---|---|---|
| CLI / TUI | terminal | type the command, capture the pane — example |
| Server / API | socket | send the request, capture the response — example |
| GUI | pixels | drive it under xvfb/Playwright, screenshot |
| Library | package boundary | sample code through the public export — import pkg, not import ./src/... |
| Prompt / agent config | the agent | run the agent, capture its behavior |
| CI workflow | Actions | dispatch it, read the run |
Internal function? Not a surface. Something in the repo calls it and that caller ends at one of the rows above. Follow it there. A bash security gate's surface isn't the function's return value — it's the CLI prompting or auto-allowing when you type the command.
No runtime surface at all — docs-only, type declarations with no emit, build config that produces no behavioral diff — report SKIP — no runtime surface: (reason). Don't run tests to fill the space.
Tests in the diff are the author's evidence, not a surface. CI runs them. You'd be re-running CI. Tests-only PR → SKIP, one line. Mixed src+tests → verify the src, ignore the test files. Reading a test to learn what to check is fine — it's a spec. But then go run the app. Checking that assertions match source is code review.
Check .claude/skills/ first — even if you already know how to
build and run. A matching verifier-* skill is the repo's
evidence-capture protocol: it wraps the session so a reviewer can
replay what you saw (recording, screenshots). Drive the surface
without it and you get a verdict with no replay.
Skills live at the repo root and in the package/app dirs the
diff touches — in a monorepo the unlock for apps/desktop/ is
usually apps/desktop/.claude/skills/, not the root. Probe both:
ls .claude/skills/ # repo root
ls <touched-dir>/.claude/skills/ # each dir level the diff namesverifier-* matching your surface (CLI verifier for a CLI
change, etc.) → invoke it with the Skill tool and follow its
setup. Mismatched surface → skip that one, try the next. Stale
verifier (fails on mechanics unrelated to the change) → ask the
user whether to patch it; don't FAIL the change for verifier rot.run-* but no matching verifier → use its build/launch
primitives as your handle./run-skill-generator prompt. Got through → persist what you
learned: create .claude/skills/verify/SKILL.md at the level you
probed above — repo root for a single-package repo; the touched
package/app dir (apps/desktop/.claude/skills/verify/SKILL.md) in
a monorepo where verification is per-package — capturing the
build/launch/drive recipe that worked, so the next session skips
this cold start. Keep it short: the commands that worked, the
flows worth driving, any gotchas. A project verify skill already
exists → edit it only when it steered you wrong: a documented
command failed or turned out wrong, or a needed step it doesn't
cover. Routine learnings don't warrant an edit, and never rewrite
or reorganize existing content for style.Smallest path that makes the changed code execute:
Read your plan back before running. If every step is build / typecheck / run test file — you've planned a CI rerun, not a verification. Find a step that reaches the surface or report BLOCKED.
The verdict is table stakes. Your observations are the signal. A PASS with three sharp "hey, I noticed…" lines is worth more than a bare PASS. You're the only reviewer who actually ran the thing — anything that made you pause, work around, or go "huh" is information the author doesn't have. Don't filter for "is this a bug." Filter for "would I mention this if they were sitting next to me."
End-to-end, through the real interface. Pieces passing in isolation doesn't mean the flow works — seams are where bugs hide. If users click buttons, test by clicking buttons, not by curling the API underneath.
Destructive path? If the change touches code that deletes, publishes, sends, or writes outside the workspace and there's no dry-run or safe target, don't drive it live. Verify what you can around it and say which path you didn't exercise and why.
The claim checked out — that's the first half. Confirming is step one, not the job. The description is what the author intended; your value is what they didn't.
You know exactly what changed. Probe around it, at the same surface you just drove:
These aren't a checklist — pick the ones the change points at. Stop
when you've covered the obvious adjacents or hit something worth a
⚠️. A probe that finds nothing is still a step: "🔍 passed --from ''
→ clean error: --from requires a value, exit 2." That the author
didn't test it is exactly why it's worth knowing it holds.
Still not a test run. You're at the surface, typing what a user would type wrong.
Stdout, response bodies, screenshots, pane dumps. Captured output is evidence; your memory isn't. Something unexpected? Don't route around it — capture, note, decide if it's the change or the environment. Unrelated breakage is a finding, not noise.
Shared process state (tmux, ports, lockfiles) — isolate. tmux -L name, bind :0, mktemp -d. You share a namespace with your host.
Inline, final message:
## Verification: <one-line what changed>
**Verdict:** PASS | FAIL | BLOCKED | SKIP
**Claim:** <what it's supposed to do — your read of the diff and/or
the stated claim; note any mismatch>
**Method:** <how you got a handle — which verifier/run-skill, or
cold start; what you launched>
### Steps
Each step is one thing you did to the **running app** and what it
showed. Build/install/checkout are setup, not steps. Test runs and
typecheck don't belong here — they're CI's output.
1. ✅/❌/⚠️/🔍 <what you did to the running app> → <what you observed>
<evidence: the app's own output — pane capture, response body,
screenshot>
🔍 marks a probe — a step off the claim's happy path, trying to
break it. At least one. A Steps list that's all ✅ and no 🔍 is a
happy-path replay: still PASS, but you stopped at the first half.
**Screenshot / sample:** <the one frame a reviewer looks at to see
the feature — an image for GUI/TUI, code block for library/API;
omit for build/types-only>
### Findings
<Things you noticed. Not just bugs — friction, surprises, anything
a first-time user would trip on. "Took three tries to find the right
flag." "Error message on typo was unhelpful." "Default seems odd for
the common case." "Works, but slower than I expected." Lower the bar:
if it made you pause, it goes here. But the pause has to be yours,
from running the app — not from reading the PR page. A red CI check,
a review comment, someone else's bot: visible to anyone already, and
you relaying it isn't an observation. Claim/diff mismatch, pre-existing
breakage, and env notes also belong.
Each probe gets a line here even when it held — "🔍 empty `--from`
→ clean error" tells the author what *was* covered, which they
can't see from a bare PASS.
Lead with ⚠️ for lines worth interrupting the reviewer for; plain
bullets are context. Empty is fine if nothing stuck out — but nothing
sticking out is itself rare.>Evidence has to reach the reader. A file path is only evidence
if the person reading the report can open it. If the SendUserFile
tool is in your toolset, you're on a remote surface where they
can't — send the screenshots and recordings with it and let the
report name what you sent. Without it, reference the path and keep
the evidence that matters inline — pane captures and response
bodies travel in the report; a bare path only works when the reader
shares your filesystem.
Verdicts:
/run-skill-generator prompt.No partial pass. "3 of 4 passed" is FAIL until 4 passes or is explained away.
When in doubt, FAIL. False PASS ships broken code; false FAIL costs one more human look. Ambiguous output is FAIL with the raw capture attached — don't interpret.
© asgeirtj, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in Anthropic/claude-code/skills/verify of asgeirtj/system_prompts_leaks.
Open the folder on GitHubat commit 60d44cc
Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Verify this skillasgeirtj/system_prompts_leaks | 69k | — | ~3k | Automated safety check: Pass | CC0-1.0 | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| Diagnosing Bugsfossasia/eventyay-interpretation | 1.6k | 32 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| TDDpietheinstrengholt/rssmonster | 564 | 30 repos | ~906 | Automated safety check: Pass | MIT | |
| TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph | 112 | 11 repos | ~2.4k | Automated safety check: Pass | None | |
| TDDsanity-io/sanity | 6.4k | 20 repos | ~1k | Automated safety check: Pass | MIT |
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
fossasia/eventyay-interpretation
Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.
pietheinstrengholt/rssmonster
Test-driven development. An agent skill from pietheinstrengholt/rssmonster.
hellangleZ/burn-in-cceverywhere-ralph
A skill your agent uses when writing new features, fixing bugs, or refactoring code.
sanity-io/sanity
Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.
Ibrahim-3d/orchestrator-supaconductor
A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…
asgeirtj/system_prompts_leaks
Shows one digest of coding-agent sessions across your connected machines and lets you open, read, steer, approve, stop and close them, over Herdr, tmux or MSP.
asgeirtj/system_prompts_leaks
Diagnoses a Muse Code installation's own failures from binary and session evidence, instead of treating the report as an ordinary repository bug.
asgeirtj/system_prompts_leaks
A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx) or Word templates (.dotx).
asgeirtj/system_prompts_leaks
Runs a goal as a project in which the agent coordinates separate agent threads, judging when to split the work, and interviews you first when nothing can be verified.
asgeirtj/system_prompts_leaks
Creates and validates a new native Muse plugin package in the current workspace, limited to five capability families, and leaves installation to you.
asgeirtj/system_prompts_leaks
A skill your agent uses when the user's prompt requires (1) researching a topic across multiple sources, comparing options or alternatives, analyzing trends or history, understanding markets or…
Categories
Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck. Verify is an agent skill from asgeirtj/system_prompts_leaks. Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.
Verify fits situations like: testing & QA work in your project.
Run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a claude-code`. Or copy the skill folder (Anthropic/claude-code/skills/verify in asgeirtj/system_prompts_leaks) into .claude/skills/verify in your project. Claude Code loads it when a task matches its description.
Run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a codex`. Or copy the skill folder (Anthropic/claude-code/skills/verify in asgeirtj/system_prompts_leaks) into .agents/skills/verify in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify, .gemini/skills/verify, .github/skills/verify and .opencode/skills/verify in your project.
Going by SKILL.md and its folder, Verify needs the command-line tools its instructions call (git and gh).
SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Verify is published under the CC0-1.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Verify: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
asgeirtj (a GitHub user) maintains it in asgeirtj/system_prompts_leaks, which has 69,280 GitHub stars. The repository holds 128 skills in this directory. The repository was last updated on October 10, 2026.
Source: asgeirtj/system_prompts_leaks on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.