Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.

CC0-1.0Auto-check passedTesting & QA

Install Verify

skills CLI
$ npx skills add asgeirtj/system_prompts_leaks --skill verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install asgeirtj/system_prompts_leaks verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/asgeirtj/system_prompts_leaks.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Anthropic/claude-code/skills/verify .claude/skills/verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify
GitHub stars
69k
Token cost
~3k tokens
SKILL.md length
1,449 words
Files
3
Skills in repo
128
Repo updated
First seen
Licence
CC0-1.0

At a glance

Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.

  • Testing & QA work in your project
  • SKILL.md covers Find the change, Surface, Get a handle and Drive it, plus 3 more sections
  • Calls git and gh

What it does

Verify is an agent skill from asgeirtj/system_prompts_leaks. Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck. Run before committing nontrivial changes; bootstraps this repo's project verify skill if none exists yet. Don't invoke it on a diff that only touches tests, docs, or other code with no runtime surface to drive (a change to product source always has one) — there's nothing to observe.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/cli.md` and `examples/server.md`).

It sits in Testing & QA. The repository describes itself as: Documented system prompts from Anthropic - Claude Fable 5.1, Opus 5.5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-6-Astra, Codex. Google - Gemini 3.8 Flash, 3.1 Pro… The licence is CC0-1.0.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “s project verify skill if none exists yet. Don”
  • “/verify”

What it can do on your machine

Read from SKILL.md and the folder at commit 60d44cc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify loads about 3k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 1,449 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from asgeirtj/system_prompts_leaks at commit 60d44cc, republished under its CC0-1.0 licence (© asgeirtj). 1,449 words, ~3,045 tokens.

Download SKILL.mdSave it as .claude/skills/verify/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
verify
description
Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck. Run before committing nontrivial changes; bootstraps this repo's project verify skill if none exists yet. Don't invoke it on a diff that only touches tests, docs, or other code with no runtime surface to drive (a change to product source always has one) — there's nothing to observe.
disable-model-invocation
true

Verification is runtime observation. You build the app, run it, drive it to where the changed code executes, and capture what you see. That capture is your evidence. Nothing else is.

Don't run tests. Don't typecheck. Running them here proves you can run CI — not that the change works. Not as a warm-up, not "just to be sure," not as a regression sweep after. The time goes to running the app instead.

Don't import-and-call. import { foo } from './src/...' then console.log(foo(x)) is a unit test you wrote. The function did what the function does — you knew that from reading it. The app never ran. Whatever calls foo in the real codebase ends at a CLI, a socket, or a window. Go there.

Find the change

The scope is what you're verifying — usually a diff, sometimes just "does X work." In a git repo, establish the full range (a branch may be many commits, or the change may still be uncommitted):

bash
git log --oneline @{u}..              # count commits (if upstream set)
git diff @{u}.. --stat                # full range, not HEAD~1
git diff origin/HEAD... --stat        # no upstream: committed vs base
git diff HEAD --stat                  # uncommitted: working tree vs HEAD
gh pr diff                            # if in a PR context

State the commit count. Large diff truncating? Redirect to a file then Read it. Repo but no diff from any of these → say so, stop. No repo → the scope is whatever the user named; ask if they didn't.

The diff is ground truth. Any description is a claim about it. Read both. If they disagree, that's a finding.

Surface

The surface is where a user — human or programmatic — meets the change. That's where you observe.

Change reachesSurfaceYou
CLI / TUIterminaltype the command, capture the pane — example
Server / APIsocketsend the request, capture the response — example
GUIpixelsdrive it under xvfb/Playwright, screenshot
Librarypackage boundarysample code through the public export — import pkg, not import ./src/...
Prompt / agent configthe agentrun the agent, capture its behavior
CI workflowActionsdispatch it, read the run

Internal function? Not a surface. Something in the repo calls it and that caller ends at one of the rows above. Follow it there. A bash security gate's surface isn't the function's return value — it's the CLI prompting or auto-allowing when you type the command.

No runtime surface at all — docs-only, type declarations with no emit, build config that produces no behavioral diff — report SKIP — no runtime surface: (reason). Don't run tests to fill the space.

Tests in the diff are the author's evidence, not a surface. CI runs them. You'd be re-running CI. Tests-only PR → SKIP, one line. Mixed src+tests → verify the src, ignore the test files. Reading a test to learn what to check is fine — it's a spec. But then go run the app. Checking that assertions match source is code review.

Get a handle

Check .claude/skills/ first — even if you already know how to build and run. A matching verifier-* skill is the repo's evidence-capture protocol: it wraps the session so a reviewer can replay what you saw (recording, screenshots). Drive the surface without it and you get a verdict with no replay.

Skills live at the repo root and in the package/app dirs the diff touches — in a monorepo the unlock for apps/desktop/ is usually apps/desktop/.claude/skills/, not the root. Probe both:

bash
ls .claude/skills/                    # repo root
ls <touched-dir>/.claude/skills/      # each dir level the diff names
  • verifier-* matching your surface (CLI verifier for a CLI change, etc.) → invoke it with the Skill tool and follow its setup. Mismatched surface → skip that one, try the next. Stale verifier (fails on mechanics unrelated to the change) → ask the user whether to patch it; don't FAIL the change for verifier rot.
  • run-* but no matching verifier → use its build/launch primitives as your handle.
  • Neither → cold start from README/package.json/Makefile. Timebox ~15min. Stuck → BLOCKED with exactly where, plus a filled-in /run-skill-generator prompt. Got through → persist what you learned: create .claude/skills/verify/SKILL.md at the level you probed above — repo root for a single-package repo; the touched package/app dir (apps/desktop/.claude/skills/verify/SKILL.md) in a monorepo where verification is per-package — capturing the build/launch/drive recipe that worked, so the next session skips this cold start. Keep it short: the commands that worked, the flows worth driving, any gotchas. A project verify skill already exists → edit it only when it steered you wrong: a documented command failed or turned out wrong, or a needed step it doesn't cover. Routine learnings don't warrant an edit, and never rewrite or reorganize existing content for style.

Drive it

Smallest path that makes the changed code execute:

  • Changed a flag? Run with it.
  • Changed a handler? Hit that route.
  • Changed error handling? Trigger the error.
  • Changed an internal function? Find the CLI command / request / render that reaches it. Run that.

Read your plan back before running. If every step is build / typecheck / run test file — you've planned a CI rerun, not a verification. Find a step that reaches the surface or report BLOCKED.

The verdict is table stakes. Your observations are the signal. A PASS with three sharp "hey, I noticed…" lines is worth more than a bare PASS. You're the only reviewer who actually ran the thing — anything that made you pause, work around, or go "huh" is information the author doesn't have. Don't filter for "is this a bug." Filter for "would I mention this if they were sitting next to me."

End-to-end, through the real interface. Pieces passing in isolation doesn't mean the flow works — seams are where bugs hide. If users click buttons, test by clicking buttons, not by curling the API underneath.

Destructive path? If the change touches code that deletes, publishes, sends, or writes outside the workspace and there's no dry-run or safe target, don't drive it live. Verify what you can around it and say which path you didn't exercise and why.

Show full SKILL.md (529 more words)Show less

Push on it

The claim checked out — that's the first half. Confirming is step one, not the job. The description is what the author intended; your value is what they didn't.

You know exactly what changed. Probe around it, at the same surface you just drove:

  • New flag / option → empty value, passed twice, combined with a conflicting flag, typo'd (does the error name it?)
  • New handler / route → wrong method, malformed body, missing required field, oversized payload
  • Changed error path → the adjacent errors it didn't touch — did the refactor catch them too, or only the one in the diff?
  • Interactive / TUI → Ctrl-C mid-op, resize the pane, paste garbage, rapid-fire the key, Esc at the wrong moment
  • State / persistence → do it twice, do it with stale state underneath, do it in two sessions at once
  • Wander → what's adjacent? What looked off while you were confirming? Go back to it.

These aren't a checklist — pick the ones the change points at. Stop when you've covered the obvious adjacents or hit something worth a ⚠️. A probe that finds nothing is still a step: "🔍 passed --from '' → clean error: --from requires a value, exit 2." That the author didn't test it is exactly why it's worth knowing it holds.

Still not a test run. You're at the surface, typing what a user would type wrong.

Capture

Stdout, response bodies, screenshots, pane dumps. Captured output is evidence; your memory isn't. Something unexpected? Don't route around it — capture, note, decide if it's the change or the environment. Unrelated breakage is a finding, not noise.

Shared process state (tmux, ports, lockfiles) — isolate. tmux -L name, bind :0, mktemp -d. You share a namespace with your host.

Report

Inline, final message:

## Verification: <one-line what changed>

**Verdict:** PASS | FAIL | BLOCKED | SKIP

**Claim:** <what it's supposed to do — your read of the diff and/or
the stated claim; note any mismatch>

**Method:** <how you got a handle — which verifier/run-skill, or
cold start; what you launched>

### Steps

Each step is one thing you did to the **running app** and what it
showed. Build/install/checkout are setup, not steps. Test runs and
typecheck don't belong here — they're CI's output.

1. ✅/❌/⚠️/🔍 <what you did to the running app> → <what you observed>
   <evidence: the app's own output — pane capture, response body,
   screenshot>

🔍 marks a probe — a step off the claim's happy path, trying to
break it. At least one. A Steps list that's all ✅ and no 🔍 is a
happy-path replay: still PASS, but you stopped at the first half.

**Screenshot / sample:** <the one frame a reviewer looks at to see
the feature — an image for GUI/TUI, code block for library/API;
omit for build/types-only>

### Findings
<Things you noticed. Not just bugs — friction, surprises, anything
a first-time user would trip on. "Took three tries to find the right
flag." "Error message on typo was unhelpful." "Default seems odd for
the common case." "Works, but slower than I expected." Lower the bar:
if it made you pause, it goes here. But the pause has to be yours,
from running the app — not from reading the PR page. A red CI check,
a review comment, someone else's bot: visible to anyone already, and
you relaying it isn't an observation. Claim/diff mismatch, pre-existing
breakage, and env notes also belong.

Each probe gets a line here even when it held — "🔍 empty `--from`
→ clean error" tells the author what *was* covered, which they
can't see from a bare PASS.

Lead with ⚠️ for lines worth interrupting the reviewer for; plain
bullets are context. Empty is fine if nothing stuck out — but nothing
sticking out is itself rare.>

Evidence has to reach the reader. A file path is only evidence if the person reading the report can open it. If the SendUserFile tool is in your toolset, you're on a remote surface where they can't — send the screenshots and recordings with it and let the report name what you sent. Without it, reference the path and keep the evidence that matters inline — pane captures and response bodies travel in the report; a bare path only works when the reader shares your filesystem.

Verdicts:

  • PASS — you ran the app, the change did what it should at its surface. Not: tests pass, builds clean, code looks right.
  • FAIL — you ran it and it doesn't. Or it breaks something else. Or claim and diff disagree materially.
  • BLOCKED — couldn't reach a state where the change is observable. Build broke, env missing a dep, handle wouldn't come up. Not a verdict on the change. Never report an approach blocked or impossible until you've enumerated the skills along the touched subtree — environment-specific unlocks (headless runners, login helpers, VM harnesses) usually live there. Say exactly where it stopped + /run-skill-generator prompt.
  • SKIP — no runtime surface exists. Docs-only, types-only, tests-only. Nothing went wrong; there's just nothing here to run. One line why.

No partial pass. "3 of 4 passed" is FAIL until 4 passes or is explained away.

When in doubt, FAIL. False PASS ships broken code; false FAIL costs one more human look. Ambiguous output is FAIL with the raw capture attached — don't interpret.

© asgeirtj, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in Anthropic/claude-code/skills/verify of asgeirtj/system_prompts_leaks.

  • SKILL.md
  • examples/cli.md
  • examples/server.md

Open the folder on GitHubat commit 60d44cc

Compare with similar skills

Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify this skillasgeirtj/system_prompts_leaks69k—~3kAutomated safety check: PassCC0-1.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k32 repos~2.1kAutomated safety check: PassApache-2.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 32 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    381 GitHub starsUsed in 9 repos~2.9k tokens
    Testing & QAAuto-check passed

More from asgeirtj/system_prompts_leaks

All 128 skills in this repo
  • Fleet Manager for Agent Sessions

    asgeirtj/system_prompts_leaks

    Shows one digest of coding-agent sessions across your connected machines and lets you open, read, steer, approve, stop and close them, over Herdr, tmux or MSP.

    69k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Muse Code Product Doctor

    asgeirtj/system_prompts_leaks

    Diagnoses a Muse Code installation's own failures from binary and session evidence, instead of treating the report as an ordinary repository bug.

    69k GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • DOCX

    asgeirtj/system_prompts_leaks

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx) or Word templates (.dotx).

    69k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Agents Project Coordinator

    asgeirtj/system_prompts_leaks

    Runs a goal as a project in which the agent coordinates separate agent threads, judging when to split the work, and interviews you first when nothing can be verified.

    69k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Muse Plugin Creator

    asgeirtj/system_prompts_leaks

    Creates and validates a new native Muse plugin package in the current workspace, limited to five capability families, and leaves installation to you.

    69k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deep Research

    asgeirtj/system_prompts_leaks

    A skill your agent uses when the user's prompt requires (1) researching a topic across multiple sources, comparing options or alternatives, analyzing trends or history, understanding markets or…

    69k GitHub stars~3.3k tokensUpdated today
    Auto-check passed

Categories

Questions about Verify

What does Verify do?

Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck. Verify is an agent skill from asgeirtj/system_prompts_leaks. Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.

When should I use Verify?

Verify fits situations like: testing & QA work in your project.

How do I install Verify in Claude Code?

Run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a claude-code`. Or copy the skill folder (Anthropic/claude-code/skills/verify in asgeirtj/system_prompts_leaks) into .claude/skills/verify in your project. Claude Code loads it when a task matches its description.

How do I install Verify in Codex?

Run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a codex`. Or copy the skill folder (Anthropic/claude-code/skills/verify in asgeirtj/system_prompts_leaks) into .agents/skills/verify in your project. Codex loads it when a task matches its description.

Can I use Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgeirtj/system_prompts_leaks --skill verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify, .gemini/skills/verify, .github/skills/verify and .opencode/skills/verify in your project.

What does Verify need to run?

Going by SKILL.md and its folder, Verify needs the command-line tools its instructions call (git and gh).

Does Verify access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Verify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verify use?

Verify is published under the CC0-1.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify?

Skills that share tags, products or a category with Verify: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify?

asgeirtj (a GitHub user) maintains it in asgeirtj/system_prompts_leaks, which has 69,280 GitHub stars. The repository holds 128 skills in this directory. The repository was last updated on October 10, 2026.

Source: asgeirtj/system_prompts_leaks on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.