Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

MITAuto-check passedProduct & Project Management

Install Proof Verify

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill proof-verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config proof-verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/development/proof-verify .claude/skills/proof-verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
proof-verify
GitHub stars
154
Token cost
~2.6k tokens
SKILL.md length
873 words
Files
4 (incl. scripts, references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

  • Works in 4 steps: Create Plan → Build → Verify (Independent Agent) → …
  • - verify against plan
  • SKILL.md covers When to Use, The Pattern, Stage Ledger - only when proof… and Phase 1: Create Plan, plus 7 more sections
  • Runs Python scripts from its folder

What it does

Proof Verify is an agent skill from AnastasiyaW/codex-claude-code-config. Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work). For multi-stage work, seal accepted inputs with commit/tree, contract, input/output digests, and a fresh verdict so downstream stages do not reopen them. Use when - "verify against plan", "proof check", "independent review", "check the implementation", or confirming a feature built from a plan meets spec. Do NOT use for quick one-off checks…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/kb-aware-verification.md`, `references/proven-stage-contracts.md` and `scripts/validate_stage_ledger.py`).

It sits in Product & Project Management, covering User stories. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • - verify against plan
  • Independent review
  • Check the implementation
  • Confirming a feature built from a plan meets spec

Example prompts

  • “verify against plan”
  • “proof check”
  • “independent review”
  • “/proof-verify”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Create Plan
  2. Build
  3. Verify (Independent Agent)
  4. Fix Loop

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Proof Verify loads about 2.6k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 873 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 873 words, ~2,645 tokens.

Download SKILL.mdSave it as .claude/skills/proof-verify/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
proof-verify
description
Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work). For multi-stage work, seal accepted inputs with commit/tree, contract, input/output digests, and a fresh verdict so downstream stages do not reopen them. Use when - "verify against plan", "proof check", "independent review", "check the implementation", or confirming a feature built from a plan meets spec. Do NOT use for quick one-off checks with no plan, or for letting the builder self-verify.

Proof Verify

Plan-based verification: freeze acceptance criteria BEFORE building, verify AFTER with independent agents.

When to Use

  • After completing a feature/fix that was built from a plan
  • When you need independent confirmation that work meets spec
  • When the builder should NOT verify their own work
  • Trigger phrases: "verify against plan", "check the implementation", "proof check", "independent review"

The Pattern

PHASE 1: PLAN (before any code)
  Create .proof/PLAN.md with numbered acceptance criteria
  Each AC: testable, specific, has a verification command or check
  Plan is FROZEN - no changes during build

PHASE 2: BUILD (normal work)
  Implement against the plan
  Mark progress in .proof/PROGRESS.md
  Builder does NOT self-verify

PHASE 3: VERIFY (after build, independent agent)
  Fresh agent reads PLAN.md (never saw the build process)
  Walks through each AC, runs verification commands
  Writes .proof/VERDICT.md with PASS/FAIL per criterion
  If any FAIL → .proof/PROBLEMS.md with specific fixes

PHASE 4: FIX (if needed)
  Builder reads PROBLEMS.md, makes minimal fixes
  Back to PHASE 3 (re-verify)
  Loop until all PASS

Stage Ledger - only when proof feeds another stage

For a release, integration, migration, hardware, signer, or other multi-stage task, a green check is not enough. The next stage needs a stable input, not a summary that the previous stage once looked good.

  1. Freeze the stage contract and its scope before building it.
  2. Run the focused proof and obtain the fresh verdict as usual.
  3. Write the stage to .proof/stage-ledger.json as VERIFIED or SEALED. A sealed stage records its source commit/tree, contract digest, named input and output digests, fresh verdict digest, and invalidation keys.
  4. A downstream stage names the sealed parent and exact output digest it consumes. It must not consume a merely VERIFIED or BLOCKED stage.
  5. If an external dependency is missing, record BLOCKED with the exact missing prerequisite. Do not invalidate the sealed upstream code.
  6. If code, contract, or an input digest changes, add a SUPERSEDED successor and re-run proof for that successor; do not overwrite the old receipt.

Validate the ledger deterministically:

Resolve <proof-verify-skill-dir> to the directory containing the loaded proof-verify SKILL.md; do not assume that a project checkout has a skills/development/ copy:

text
python <proof-verify-skill-dir>/scripts/validate_stage_ledger.py \
  .proof/stage-ledger.json

The ledger is not a signing ceremony for every edit. Use it only when an accepted result crosses a real project boundary. Full format, status semantics, and examples: references/proven-stage-contracts.md.

Phase 1: Create Plan

Create .proof/PLAN.md in the project root:

markdown
# Verification Plan

**Created:** YYYY-MM-DD HH:MM
**Task:** [one-line description]
**Builder:** [session ID or "current"]
**Status:** FROZEN

## Acceptance Criteria

### AC1: [short name]
**Description:** [what must be true]
**Verify:** [exact command or check to run]
**Expected:** [what success looks like]

### AC2: [short name]
**Description:** [what must be true]
**Verify:** [exact command or check to run]
**Expected:** [what success looks like]

### AC3: [short name]
...

## Out of Scope
- [explicitly what this plan does NOT cover]

## Constraints
- [time, resource, or technical constraints]

Rules for good ACs:

  • Testable - there is a command or check that produces PASS/FAIL
  • Specific - "function returns correct value" not "code works"
  • Independent - each AC can be verified without the others
  • Sufficient - use one criterion when one observable contract is all that changed; split criteria only when their behavior, owner, or verification command is meaningfully independent
  • Frozen - once written, do not modify during build

Phase 2: Build

Normal implementation. The only additions:

  1. Create .proof/PROGRESS.md as you work:
markdown
# Build Progress

### AC1: [name]
- [x] Implemented in `src/foo.py:42`
- Files changed: `src/foo.py`, `tests/test_foo.py`

### AC2: [name]
- [x] Implemented in `src/bar.py:18`
- Files changed: `src/bar.py`
- Note: chose approach B because [reason]
  1. After build is complete, write .proof/EVIDENCE.md:
markdown
# Evidence

### AC1: [name]
**Command:** `pytest tests/test_foo.py -v`
**Output:**
\```
tests/test_foo.py::test_returns_correct PASSED
tests/test_foo.py::test_handles_edge PASSED
\```
**Result:** PASS

### AC2: [name]
**Command:** `grep -c "TODO" src/bar.py`
**Output:** `0`
**Result:** PASS

Builder collects evidence but does NOT write the verdict. That is the verifier's job.

Phase 3: Verify (Independent Agent)

This is the critical phase. The verifier MUST be:

  • A fresh agent (new session or subagent) that never saw the build
  • Given ONLY: PLAN.md + access to the codebase
  • NOT given: PROGRESS.md, EVIDENCE.md, or any build context
Verifier prompt template
You are an independent verifier. Your job is to check whether
the implementation meets the acceptance criteria in .proof/PLAN.md.

Rules:
1. Read .proof/PLAN.md first. This is your ONLY specification.
2. For each AC, run the verification command yourself.
3. Do NOT read .proof/PROGRESS.md or .proof/EVIDENCE.md
   (those are the builder's claims - you verify independently).
4. Write your verdict to .proof/VERDICT.md in this format:

# Verification Verdict

**Verifier:** [your session ID]
**Date:** YYYY-MM-DD HH:MM
**Plan hash:** [first 8 chars of md5 of PLAN.md]

## Results

### AC1: [name]
**Status:** PASS | FAIL
**Evidence:** [what you saw when you ran the check]
**Notes:** [any observations]

### AC2: [name]
...

## Summary
- Total: N criteria
- Passed: X
- Failed: Y
- **Overall:** PASS | FAIL

5. If any AC fails, also create .proof/PROBLEMS.md:

# Problems

### AC2: [name]
**Expected:** [from PLAN.md]
**Actual:** [what you found]
**Suggested fix:** [smallest change that would fix it]
**Affected files:** [list]

6. Do NOT fix anything. You are read-only. Report only.
How to spawn the verifier

Option A: Subagent (same session)

Agent({
  description: "Independent verification against plan",
  prompt: "[verifier prompt above]",
  mode: "plan"  // read-only first
})

Option B: Fresh session (stronger isolation) Write handoff with instruction: "Start by reading .proof/PLAN.md and running verification."

Option C: Multiple verifiers (highest confidence) Spawn 2-3 verifiers independently. If they disagree on any AC, that AC needs investigation.

Phase 4: Fix Loop

If VERDICT.md shows any FAIL:

  1. Builder reads PROBLEMS.md
  2. Makes minimal fixes (not refactoring, not "while I'm here")
  3. Updates EVIDENCE.md with new evidence for failed ACs
  4. Verifier runs again (Phase 3)
  5. Loop until all PASS

If repeated failures stop distinguishing causal hypotheses, re-triage the affected owner and evidence. Do not weaken or rewrite an acceptance criterion merely to turn the current implementation green.

Show full SKILL.md (324 more words)Show less

File Structure

.proof/
  PLAN.md        # frozen acceptance criteria (Phase 1)
  PROGRESS.md    # builder's notes (Phase 2)
  EVIDENCE.md    # builder's evidence (Phase 2)
  VERDICT.md     # verifier's verdict (Phase 3)
  PROBLEMS.md    # verifier's findings (Phase 3, if failures)
  stage-ledger.json # only for multi-stage work; accepted inputs and blockers

Gotchas

  • Builder reads VERDICT, not the reverse. Verifier never sees builder's evidence. This prevents confirmation bias.
  • "PASS with concerns" is FAIL. Either it passes or it doesn't. No soft passes.
  • Plan hash in verdict. If someone edited PLAN.md mid-build, the hash won't match. Catch.
  • A later blocker is not retroactive failure. A missing VM, signer, or account blocks its own stage. It does not turn a sealed source or artifact into a failed one.
  • Do not reuse stale proof. A changed contract, source tree, or recorded input requires a successor stage and a fresh verdict.
  • Duration is not a verdict. Give a long-running check a bounded timeout appropriate to its environment. Prefer a smaller check only when it proves the same contract; do not discard a valid runtime boundary merely because it takes longer than a fixed threshold.
  • Don't verify style. ACs should be functional ("function returns X"), not stylistic ("code is clean"). Style is for code review, not proof loop.

Troubleshooting

SymptomCauseFix
Verifier passes everythingACs too vagueRewrite with specific commands
Repeated failures no longer distinguish causal hypothesesCurrent owner, evidence, or remedy is no longer discriminatingRe-triage the affected owner and evidence; preserve the frozen acceptance contract unless an explicitly authorized successor contract is required
Verifier disagrees with builder's evidenceDifferent env or stale stateBoth run from clean state
Builder keeps editing PLAN.mdNot frozenHash check catches this
A new audit says an old stage is "missing"It mixed an unavailable next prerequisite with already-proven scopeCheck .proof/stage-ledger.json; record the external BLOCKED stage separately
Downstream proof cannot identify its inputThe prior result was a chat claim, not a sealed receiptSeal the prior stage with commit/tree, digests, and fresh verdict before proceeding

Sources

  • Proof Loop (Principle 02) - the theoretical foundation
  • references/proven-stage-contracts.md - immutable stage promotion, provenance, and local ledger contract
  • OpenClaw-RL - spec freeze → build → fresh verify
  • Agent-R - failed-then-fixed trajectories
  • oh-my-claudecode Ralph - PRD-driven persistence (practical inspiration)

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/development/proof-verify of AnastasiyaW/codex-claude-code-config.

  • SKILL.md
  • references/kb-aware-verification.md
  • references/proven-stage-contracts.md
  • scripts/validate_stage_ledger.py

Open the folder on GitHubat commit 67709af

Compare with similar skills

Proof Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Proof Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Proof Verify this skillAnastasiyaW/codex-claude-code-config154—~2.6kAutomated safety check: PassMIT
User Story Writerdeanpeters/Product-Manager-Skills7.2k2 repos~2.9kAutomated safety check: PassCustom licence
Ralph Tui Create Beadssubsy/ralph-tui2.5k1 repos~2.6kAutomated safety check: PassMIT
Agile Product Owneralirezarezvani/claude-skills28k3 repos~3.2kAutomated safety check: PassMIT
Ralph Tui Create Beads Rustsubsy/ralph-tui2.5k1 repos~2.8kAutomated safety check: PassMIT
To Specbestofjs/bestofjs3.1k21 repos~757Automated safety check: PassMIT

Similar skills

  • User Story Writer

    deanpeters/Product-Manager-Skills

    Writes user stories in Mike Cohn's format with Gherkin acceptance criteria, turning user needs into development-ready work with testable conditions.

    7.2k GitHub starsUsed in 2 repos~2.9k tokens
    Product & Project ManagementAuto-check passed
  • Ralph Tui Create Beads

    subsy/ralph-tui

    Convert PRDs to beads for ralph-tui execution. An agent skill from subsy/ralph-tui.

    2.5k GitHub starsUsed in 1 repo~2.6k tokens
    Product & Project ManagementAuto-check passed
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Product & Project ManagementAuto-check passed
  • Convert PRDs to beads for ralph-tui execution using beads-rust (br CLI).

    2.5k GitHub starsUsed in 1 repo~2.8k tokens
    Product & Project ManagementAuto-check passed
  • To Spec

    bestofjs/bestofjs

    Turn the current conversation into a spec and publish it to the project issue tracker — no interview, just synthesis of what you've already discussed.

    3.1k GitHub starsUsed in 21 repos~757 tokens
    Product & Project ManagementAuto-check passed
  • Ralph Tui Create JSON

    subsy/ralph-tui

    Convert PRDs to prd.json format for ralph-tui execution. An agent skill from subsy/ralph-tui.

    2.5k GitHub starsUsed in 1 repo~2.6k tokens
    Product & Project ManagementAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Desktop Sessions Discovery

    AnastasiyaW/codex-claude-code-config

    Discover, search, and selectively restore Claude desktop app sessions hidden across multiple accountIds.

    154 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Questions about Proof Verify

What does Proof Verify do?

Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work). Proof Verify is an agent skill from AnastasiyaW/codex-claude-code-config. Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

When should I use Proof Verify?

Proof Verify fits situations like: - verify against plan; independent review; check the implementation; confirming a feature built from a plan meets spec.

How do I install Proof Verify in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill proof-verify -a claude-code`. Or copy the skill folder (skills/development/proof-verify in AnastasiyaW/codex-claude-code-config) into .claude/skills/proof-verify in your project. Claude Code loads it when a task matches its description.

How do I install Proof Verify in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill proof-verify -a codex`. Or copy the skill folder (skills/development/proof-verify in AnastasiyaW/codex-claude-code-config) into .agents/skills/proof-verify in your project. Codex loads it when a task matches its description.

Can I use Proof Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill proof-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/proof-verify, .gemini/skills/proof-verify, .github/skills/proof-verify and .opencode/skills/proof-verify in your project.

What does Proof Verify need to run?

Going by SKILL.md and its folder, Proof Verify needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Proof Verify access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Proof Verify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Proof Verify use?

Proof Verify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Proof Verify use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Proof Verify?

Skills that share tags, products or a category with Proof Verify: User Story Writer (deanpeters/Product-Manager-Skills, 7.2k stars), Ralph Tui Create Beads (subsy/ralph-tui, 2.5k stars), Agile Product Owner (alirezarezvani/claude-skills, 28k stars) and Ralph Tui Create Beads Rust (subsy/ralph-tui, 2.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Proof Verify?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.