Agent skill

Verifying Claims

by oaustegard in oaustegard/claude-skills

Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree.

MITAuto-check passedTesting & QA

Install Verifying Claims

skills CLI
$ npx skills add oaustegard/claude-skills --skill verifying-claims -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills verifying-claims --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/verifying-claims .claude/skills/verifying-claims && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verifying-claims
GitHub stars
150
Token cost
~1.3k tokens
SKILL.md length
685 words
Files
5 (incl. scripts, references)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree.

  • Works in 6 steps: Identify the document(s) to check and… → Gather consistent input: run… → Extract the claims the prose makes —… → …
  • The user wants to verify a README
  • SKILL.md covers What changed (v0.1 → v0.2), Division of labor — read this…, Procedure and Verdicts, plus 4 more sections
  • Runs Python scripts from its folder

What it does

Verifying Claims is an agent skill from oaustegard/claude-skills. Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documentation drift, doc-code sync, "is this still accurate", stale docs, or keeping docs/tests/code consistent; before publishing or merging a docs change; or as a periodic doc-accuracy sweep. The agent reads the prose's meaning directly —…

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `CHANGELOG.md`, `README.md` and `references/drift-report-example.md`).

It sits in Testing & QA, covering Technical documentation, Test generation and Test-driven development. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • The user wants to verify a README
  • Docstring still matches the code
  • Whenever they mention documentation drift
  • Is this still accurate

Example prompts

  • “is this still accurate”
  • “/verifying-claims”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify the document(s) to check and the code + tests they describe.
  2. Gather consistent input: run `scripts/gather_context.py --doc DOC --src
  3. Extract the claims the prose makes — every checkable assertion about the
  4. Judge each claim against the API surface and the tests
  5. Report drift, ranked by severity, each finding citing the prose claim and
  6. Optionally fix: rewrite the prose to match reality, and/or flag claims

What it can do on your machine

Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verifying Claims loads about 1.3k tokens when it runs, and up to ~1.8k if it reads all its reference files. Until then it costs about 174 tokens; SKILL.md has 685 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~174
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 685 words, ~1,296 tokens.

Download SKILL.mdSave it as .claude/skills/verifying-claims/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
verifying-claims
description
Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documentation drift, doc-code sync, "is this still accurate", stale docs, or keeping docs/tests/code consistent; before publishing or merging a docs change; or as a periodic doc-accuracy sweep. The agent reads the prose's meaning directly — there is no claim-comment DSL to maintain. Pairs with TDD — the test suite is the deterministic behavioral gate, this skill is the semantic prose-vs-reality review.
metadata.version
0.2.0

verifying-claims

Check that what a document says about code is true, by reading the document, the code, and the tests together and reporting where they disagree.

What changed (v0.1 → v0.2)

v0.1 was a comment-DSL: you hand-wrote <!-- claim: ... --> next to prose and a script checked the comment against the code. That had a fatal gap — the comment and the prose were two artifacts stapled together, and only the comment was checked, while humans read the prose. The prose could lie with a green run.

v0.2 drops the DSL. The reviewer is the agent: it reads the prose's meaning directly and compares it to what the code does and what the tests assert. No shadow copy, because the thing being checked is the thing the human reads. (Existing tools already own the alternatives — Gherkin binds executable scenarios, Lean's Verso transcludes facts into prose, TDD couples code to tests. This fills the remaining slot: free-prose documentation, judged.)

Division of labor — read this first

This skill does NOT gate merges and is NOT a test framework.

  • The test suite (TDD/CI) owns the behavioral contract: deterministic, cheap, auditable, gated. A green check is something you can hold CI to.
  • This skill owns the prose layer: does the documentation match reality? That needs semantic judgment across artifacts, which is non-deterministic and fallible — so it runs as a triggered review (before docs ship, on request, as a sweep), not as a per-commit gate. "The agent said the docs match" is not a guarantee you gate a merge on; it's a review you act on.

Tests are the anchor. The docs are correct when they agree with what the tests assert about the code. So write/keep good tests first; this skill keeps the prose pinned to them.

Procedure

  1. Identify the document(s) to check and the code + tests they describe.
  2. Gather consistent input: run scripts/gather_context.py --doc DOC --src SRC --tests TESTS. It ast-parses source (no imports, no execution) and bundles the document text, the public API surface, and the test inventory.
  3. Extract the claims the prose makes — every checkable assertion about the code (signatures, behavior, return shapes, defaults, guarantees, examples). Do this by reading; there are no claim markers.
  4. Judge each claim against the API surface and the tests:
    • Does the code actually do what the prose says?
    • Is the claim backed by a test, or merely asserted?
    • Does it reference something that no longer exists?
  5. Report drift, ranked by severity, each finding citing the prose claim and the contradicting reality (file/function). Use the verdicts below.
  6. Optionally fix: rewrite the prose to match reality, and/or flag claims that need a test (an UNSUPPORTED claim is a missing test, not just a doc bug).
Show full SKILL.md (236 more words)Show less

Verdicts

  • PASS — the prose claim matches the code and is exercised by a test.
  • FAIL — the code contradicts the claim (the doc is wrong, or the code regressed and the doc caught it).
  • UNSUPPORTED — the claim matches the current code but no test backs it, so nothing protects it from future drift. Surface as a missing test.
  • STALE — the claim refers to something removed or renamed.

Invoking

  • "Check the README against the code before I publish it."
  • "Does docs/api.md still match pkg/?"
  • "Sweep the docs for drift after this refactor."

Run it at moments that matter — pre-publish, post-refactor, on a docs PR — not on every commit. The deterministic gate is the test suite; this is the layer tests can't reach.

Honest limits

  • Non-deterministic and fallible: a review can miss drift or misjudge. Treat output as a careful review, not a proof.
  • Cost/latency: reading three artifacts and reasoning is expensive next to a test run. Don't wire it where a cheap deterministic check belongs.
  • It checks prose against code+tests; it does not verify the tests themselves are correct. Garbage tests → confident-but-wrong PASS. TDD discipline upstream still matters.

When NOT to use

  • As a CI merge gate (use the test suite).
  • To verify behavior (write a test).
  • On prose with no factual claims about code (nothing to check).

Files

  • scripts/gather_context.py — deterministic input bundler (doc + API surface + test inventory), ast-only, no imports.
  • references/drift-report-example.md — what a review report looks like.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in verifying-claims of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • references/drift-report-example.md
  • scripts/gather_context.py

Open the folder on GitHubat commit 90b0f1b

Compare with similar skills

Verifying Claims next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verifying Claims compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verifying Claims this skilloaustegard/claude-skills150—~1.3kAutomated safety check: PassMIT
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0
Devnpc-live/clawfirm156—~642Automated safety check: PassNone
Claude Superpowersmohitagw15856/pm-claude-skills1.4k—~2.7kAutomated safety check: PassMIT
Test Designjjjkkkjjj/Matft147—~2kAutomated safety check: PassBSD-3-Clause
Write Test389ds/389-ds-base294—~1.8kAutomated safety check: PassCustom licence

Similar skills

  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Dev

    npc-live/clawfirm

    Software development workflow dispatcher. An agent skill from npc-live/clawfirm.

    156 GitHub stars~642 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Claude Superpowers

    mohitagw15856/pm-claude-skills

    Activate a 4-stage coding discipline framework that forces Claude to plan before coding, isolate changes on a branch, write tests first, and self-review output twice before presenting it.

    1.4k GitHub stars~2.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Design

    jjjkkkjjj/Matft

    Procedure for designing and writing Matft's XCTest cases with high coverage — boundary values, dtypes, memory layouts, NaN/inf, empty arrays, broadcasting, platform differences, performance and…

    147 GitHub stars~2k tokensUpdated 14 days ago
    Testing & QAAuto-check passed
  • Write Test

    389ds/389-ds-base

    Add or extend a pytest integration test for 389 Directory Server under dirsrvtests/.

    294 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • TDD Guide

    alirezarezvani/claude-code-skill-factory

    Comprehensive Test Driven Development guide for engineering subagents with multi-framework support, coverage analysis, and intelligent test generation

    882 GitHub starsUsed in 1 repo~2.7k tokens
    Testing & QAAuto-check passed

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Deciding With Confidence

    oaustegard/claude-skills

    Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

    150 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Verifying Claims

What does Verifying Claims do?

Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Verifying Claims is an agent skill from oaustegard/claude-skills. Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree.

When should I use Verifying Claims?

Verifying Claims fits situations like: the user wants to verify a README; docstring still matches the code; whenever they mention documentation drift; is this still accurate.

How do I install Verifying Claims in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill verifying-claims -a claude-code`. Or copy the skill folder (verifying-claims in oaustegard/claude-skills) into .claude/skills/verifying-claims in your project. Claude Code loads it when a task matches its description.

How do I install Verifying Claims in Codex?

Run `npx skills add oaustegard/claude-skills --skill verifying-claims -a codex`. Or copy the skill folder (verifying-claims in oaustegard/claude-skills) into .agents/skills/verifying-claims in your project. Codex loads it when a task matches its description.

Can I use Verifying Claims in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill verifying-claims -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verifying-claims, .gemini/skills/verifying-claims, .github/skills/verifying-claims and .opencode/skills/verifying-claims in your project.

What does Verifying Claims need to run?

Going by SKILL.md and its folder, Verifying Claims needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Verifying Claims access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verifying Claims safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Verifying Claims use?

Verifying Claims is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verifying Claims use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 454 tokens, read only when the agent opens those files.

What are the alternatives to Verifying Claims?

Skills that share tags, products or a category with Verifying Claims: MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars), Dev (npc-live/clawfirm, 156 stars), Claude Superpowers (mohitagw15856/pm-claude-skills, 1.4k stars) and Test Design (jjjkkkjjj/Matft, 147 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verifying Claims?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.