Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict.

MITAuto-check passed

Install Verify This

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill verify-this -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config verify-this --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/development/verify-this .claude/skills/verify-this && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-this
GitHub stars
154
Token cost
~1.2k tokens
SKILL.md length
559 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict.

  • Works in 6 steps: Restate the claim as a condition,… → Select the smallest local surface that… → Capture a baseline from the parent… → …
  • Asked to verify
  • SKILL.md covers Workflow, Evidence Contract, Verdict Rules and Boundaries, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Verify This is an agent skill from AnastasiyaW/codex-claude-code-config. Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict. Use when asked to verify, prove, compare before and after, show evidence, or confirm that a fix works. Do not use for vague claims such as cleaner code, a full plan-based release verification, or a known bug that needs a red-to-green reproducer.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Asked to verify
  • Compare before and after
  • Confirm that a fix works
  • Vague claims such as cleaner code

Example prompts

  • “/verify-this”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Restate the claim as a condition, metric, and threshold. If the claim cannot
  2. Select the smallest local surface that can disprove it.
  3. Capture a baseline from the parent commit, merge base, current failing
  4. Capture treatment with the same command, data, warmup, environment, and
  5. Compare raw artifacts: test output, timings, screenshots, HTTP responses,
  6. Return exactly one verdict: VERIFIED, NOT VERIFIED, or INCONCLUSIVE.

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify This loads about 1.2k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 559 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 559 words, ~1,175 tokens.

Download SKILL.mdSave it as .claude/skills/verify-this/SKILL.md (or your agent's skills folder).
name
verify-this
description
Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict. Use when asked to verify, prove, compare before and after, show evidence, or confirm that a fix works. Do not use for vague claims such as cleaner code, a full plan-based release verification, or a known bug that needs a red-to-green reproducer.

Verify This

Verification is a falsifiable comparison, not a recap of what the agent believes it changed. Turn one claim into a measurable check and preserve enough evidence for another agent to repeat it.

Workflow

  1. Restate the claim as a condition, metric, and threshold. If the claim cannot be measured, ask for a measurable form or classify it as INCONCLUSIVE.
  2. Select the smallest local surface that can disprove it.
  3. Capture a baseline from the parent commit, merge base, current failing reproducer, or unchanged fixture.
  4. Capture treatment with the same command, data, warmup, environment, and measurement method.
  5. Compare raw artifacts: test output, timings, screenshots, HTTP responses, traces, profiles, or heap snapshots. Do not compare summaries alone.
  6. Return exactly one verdict: VERIFIED, NOT VERIFIED, or INCONCLUSIVE.

Evidence Contract

Record:

  • claim and threshold;
  • revision or baseline identity for both runs;
  • frozen contract and stage identity when this result will feed later work;
  • exact commands and input fixture;
  • environment differences and skipped checks;
  • artifact paths or hashes, with sensitive payloads kept outside public Git;
  • the verdict and one short explanation of confounders.

For durable project work, use the repository's existing proof artifact location, for example .agent/tasks/<task-id>/verification/<claim-slug>/. Temporary or sensitive evidence may stay outside the repository; retain only safe metadata and hashes in the project. Never put credentials, private prompts, customer data, or heap contents in a public checkout.

Verdict Rules

  • VERIFIED: baseline and treatment move in the predicted direction, meet the stated threshold, and have no material confound.
  • NOT VERIFIED: behavior is unchanged, moves the wrong way, or misses the threshold.
  • INCONCLUSIVE: there is no valid baseline, the signal is too noisy, the command failed, or the environments are not comparable.

Use this output shape:

text
VERIFIED | NOT VERIFIED | INCONCLUSIVE
Claim: <falsifiable claim>
Evidence:
<artifact or metric>: baseline=<...>, treatment=<...>, delta=<...>, threshold=<...>
Reasoning:
<one tight paragraph naming evidence and confounders>
Show full SKILL.md (272 more words)Show less

Boundaries

  • Use proof-verify when the work has frozen multi-criterion acceptance criteria and needs a fresh-context verifier.
  • When the claim becomes an input to another stage, seal it in .proof/stage-ledger.json with the exact commit/tree and input/output digests. A changed contract, source, or input invalidates that claim; an unavailable later environment does not.
  • Use bug-reproducer when a concrete defect needs a minimal red-to-green test and separate approval gates.
  • Use testing-strategy to choose test levels before running this comparison.
  • A single green test is not enough for a performance, release, UI, or memory claim unless it is the stated evidence surface.

Gotchas

  • A different fixture, warm cache, compiler, or machine can invalidate a baseline comparison; report it instead of smoothing it away.
  • A missing baseline is not a passing baseline. Use INCONCLUSIVE.
  • A test can pass while user-visible behavior remains wrong; use the real CLI, browser, API, or artifact boundary when that is the claim.
  • Do not turn a failed comparison green with retries, wider tolerances, or a changed workload unless the claim itself was explicitly re-scoped.

Troubleshooting

SymptomLikely causeAction
No comparable baselineParent state or repro is unavailableReport INCONCLUSIVE; capture a new baseline before changing the claim
Results vary between runsWarmup, shared state, timing noise, or nondeterminismFix isolation and repeat with a fixed workload; record variance
Treatment passes but claim is still doubtfulWrong evidence surfaceMove to the real boundary or add one focused integration/UI/CLI check
Evidence contains sensitive dataRaw artifact is not suitable for GitKeep it private and record only safe metadata or a hash

Source

Adapted from Cursor Team Kit's MIT-licensed verify-this workflow: https://github.com/cursor/plugins/tree/main/cursor-team-kit/skills/verify-this

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/development/verify-this of AnastasiyaW/codex-claude-code-config.

Open the folder on GitHubat commit 67709af

Compare with similar skills

Verify This next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify This compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify This this skillAnastasiyaW/codex-claude-code-config154—~1.2kAutomated safety check: PassMIT
Principle Prove It Workscursor/plugins11k9 repos~343Automated safety check: PassNone
Cement And Concrete Researchbrycewang-stanford/Awesome-Journal-Skills1.2k—~2.1kAutomated safety check: PassMIT
Direct Provingfrenzymath/Danus476—~1.3kAutomated safety check: PassApache-2.0
Lean4 Theorem Provingbenchflow-ai/skillsbench1.8k—~2.2kAutomated safety check: PassApache-2.0
Create Proveaeonfun/aeon770—~1.2kAutomated safety check: PassMIT

Similar skills

  • Official

    Apply after completing a task, before declaring done. An agent skill from cursor/plugins.

    11k GitHub starsUsed in 9 repos~343 tokens
    Auto-check passed
  • Cement And Concrete Research

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when targeting Cement and Concrete Research or deciding whether a cementitious-materials manuscript fits this venue.

    1.2k GitHub stars~2.1k tokensUpdated 13 days ago
    Research & ScienceAuto-check passed
  • Direct Proving

    frenzymath/Danus

    Screen a decomposition plan by first trying to prove all of its subgoals directly, then identifying the key stuck points if the plan does not fully go through.

    476 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Lean4 Theorem Proving

    benchflow-ai/skillsbench

    A skill your agent uses when working with Lean 4 (.lean files), writing mathematical proofs, seeing "failed to synthesize instance" errors, managing sorry/axiom elimination, or searching mathlib for…

    1.8k GitHub stars~2.2k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Create Prove

    aeonfun/aeon

    Run a changed Aeon skill for real and attach SHA-bound behavioral evidence to its PR

    770 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Prove Feature

    searlsco/prove_it

    Create a temporary real project and prove a proveit feature works (or doesn't) end-to-end.

    198 GitHub stars~5.7k tokensUpdated 18 days ago
    Testing & QAAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Verify This

What does Verify This do?

Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict. Verify This is an agent skill from AnastasiyaW/codex-claude-code-config. Prove a concrete behavior, performance, UI, CLI, API, or memory claim with fresh baseline-versus-treatment evidence and one explicit verdict.

When should I use Verify This?

Verify This fits situations like: asked to verify; compare before and after; confirm that a fix works; vague claims such as cleaner code.

How do I install Verify This in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill verify-this -a claude-code`. Or copy the skill folder (skills/development/verify-this in AnastasiyaW/codex-claude-code-config) into .claude/skills/verify-this in your project. Claude Code loads it when a task matches its description.

How do I install Verify This in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill verify-this -a codex`. Or copy the skill folder (skills/development/verify-this in AnastasiyaW/codex-claude-code-config) into .agents/skills/verify-this in your project. Codex loads it when a task matches its description.

Can I use Verify This in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill verify-this -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-this, .gemini/skills/verify-this, .github/skills/verify-this and .opencode/skills/verify-this in your project.

What does Verify This need to run?

SKILL.md names no scripts, command-line tools or credentials: Verify This is instructions for the agent only.

Does Verify This access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Verify This safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verify This use?

Verify This is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify This use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify This?

Skills that share tags, products or a category with Verify This: Principle Prove It Works (cursor/plugins, 11k stars), Cement And Concrete Research (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars), Direct Proving (frenzymath/Danus, 476 stars) and Lean4 Theorem Proving (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify This?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.