Agent skill

Verify Math

by flonat in flonat/flonat-research

Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report.

MITAuto-check: notes

Install Verify Math

skills CLI
$ npx skills add flonat/flonat-research --skill verify-math -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research verify-math --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/verify-math .claude/skills/verify-math && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-math
GitHub stars
145
Token cost
~1.9k tokens
SKILL.md length
825 words
Files
2 (incl. scripts)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report.

  • Works in 5 steps: Decompose the result into atomic claims → Classify each claim → rung → Order: falsify cheap → prove expensive → …
  • Paper-wide mathematical argument needs the appropriate combination of verification methods
  • SKILL.md covers When to Use, When NOT to Use, The verification spectrum (the… and Procedure, plus 4 more sections
  • Runs Python scripts from its folder; calls uv

What it does

Verify Math is an agent skill from flonat/flonat-research. Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report. Use when a theorem, proposition, conjecture, or paper-wide mathematical argument needs the appropriate combination of verification methods. For one known method, use $numerical-check, $symbolic-check, or $lean-check directly.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/extract_block.py`).

The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Paper-wide mathematical argument needs the appropriate combination of verification methods

Example prompts

  • “/verify-math”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Grep, Bash, AskUserQuestion, Skill, Task, Write, Edit

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Decompose the result into atomic claims
  2. Classify each claim → rung
  3. Order: falsify cheap → prove expensive
  4. Dispatch
  5. Aggregate → one verdict

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Bash
    • AskUserQuestion
    • Skill
    • Task
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Math loads about 1.9k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 825 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Bash, AskUserQuestion, Skill, Task, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 825 words, ~1,934 tokens.

Download SKILL.mdSave it as .claude/skills/verify-math/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
verify-math
description
Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report. Use when a theorem, proposition, conjecture, or paper-wide mathematical argument needs the appropriate combination of verification methods. For one known method, use $numerical-check, $symbolic-check, or $lean-check directly.
allowed-tools
Read, Grep, Bash, AskUserQuestion, Skill, Task, Write, Edit
agent-dependencies
domain-reviewer

Verify Math: Route a Math Claim Through the Verification Spectrum

The front door for verifying self-authored mathematics. Classify each claim, dispatch it to the strongest applicable rung(s), and merge the sub-verdicts into one report. This skill does not verify anything itself — it routes and aggregates; the rungs do the work.

When to Use

  • You have a Proposition / Theorem / Conjecture (or a whole paper's worth) and want it verified with the right method(s), possibly combined.
  • verify-math, "verify this theorem", "check all the math in the paper", "is this result correct".
  • The operational front end of mark-unverified: run this before asserting a self-authored result.

When NOT to Use

  • You already know the single method → call it directly (numerical-check, symbolic-check, lean-check, or domain-reviewer).
  • Non-mathematical claims (citations, prose) → proofread, bib-validate, domain-reviewer.

The verification spectrum (the rungs it routes to)

RungMethodCan it…Tool
R0adversarial deductive readcatch conceptual/assumption gaps (no proof)domain-reviewer (agent)
R1numerical falsificationfalsify definitively; support (never prove)numerical-check
R2symbolic / CASprove or falsify an algebra stepsymbolic-check
R3formal Lean proofprove (strongest)lean-check

Procedure

1. Decompose the result into atomic claims

A theorem is usually several obligations. List each separately: the algebra steps, the distributional/parameter-space claims, the load-bearing lemma, the conceptual assumptions. Verify each with the rung that fits — a single "verdict" on a compound theorem hides which part is shaky.

When the claim lives in LaTeX, extract a self-contained theorem/proof block before dispatch:

bash
uv run python <skill-dir>/scripts/extract_block.py paper/sections/model.tex "prop:concavity"

The helper includes an immediately following proof and any displayed equations referenced by label. Inspect the output and add missing definitions or assumptions before giving it to a reviewer or computational rung.

2. Classify each claim → rung
Claim shapeRungRoute to
algebra / derivative / limit / closed-form identityR2symbolic-check
monotonicity / threshold / comparative-static / inequality over a parameter space (distributional, probabilistic)R1numerical-check
a critical, faithfully-formalizable lemmaR3lean-check
conceptual correctness, assumption completeness, code-theory alignmentR0domain-reviewer (Task/agent)
3. Order: falsify cheap → prove expensive
  • Falsify first with the cheapest applicable rung (usually R1 numerical, or R2 for algebra). A counterexample stops the pipeline — no point proving a false claim.
  • Then prove the survivors with a proving rung (R2 for algebra, R3 for the key lemma) if the claim is important enough to warrant a positive guarantee.
  • R0 domain-reviewer runs across the whole result for the conceptual/assumption layer the computational rungs can't see.
  • Present the routing plan (claim → rung) and, when it involves R3 (expensive) or ambiguity, confirm with the user before dispatching.
4. Dispatch
  • Invoke each rung skill via the skill-routing mechanism (numerical-check, symbolic-check, lean-check); dispatch domain-reviewer via the Agent tool (fresh context — it reviews math it didn't write, per agents-vs-skills).
  • Each rung writes its own reviews/<scope>/verify-<method>/… report (shared shape).

For a hard proof strategy, existence/uniqueness question, or counterexample search that resists the standard rungs, use an adversarial reasoning pass as lead generation, never as a verdict. Make the prompt self-contained: state the exact claim, define every symbol and parameter domain, describe failed approaches, request multiple concrete strategies, and demand explicit algebra or numerical counterexamples. Route every proposed proof step or counterexample back through R1, R2, R3, or the fresh-context R0 reviewer before relying on it.

Show full SKILL.md (303 more words)Show less
5. Aggregate → one verdict
ConditionAggregate verdict
Any rung returns FALSIFIEDFALSIFIED (name which claim + counterexample)
A proving rung (R2/R3) VERIFIES the core claim, nothing falsifiedVERIFIED (note the guarantee level: CAS vs Lean)
Only R1 supports (no counterexample), no proving rungINCONCLUSIVE (supported: no counterexample in N)
Rungs couldn't decide / not formalizableINCONCLUSIVE
Setup/toolchain failureERROR

Numerical alone is never VERIFIED. Only a proving rung (R2/R3) upgrades a claim from "unfalsified" to "verified".

Anti-Patterns

  • Don't collapse a compound theorem into one verdict — decompose; report per-obligation so the shaky step is visible.
  • Don't report VERIFIED off R1 (numerical) alone — that's INCONCLUSIVE(supported); numerical can't prove.
  • Don't prove before falsifying — a cheap counterexample saves an expensive Lean effort.
  • Don't self-review the conceptual layer — dispatch domain-reviewer as a fresh-context agent (the session that wrote the math is blind to its own gaps).
  • Don't skip R0 — the computational rungs verify the math as stated; they can't catch a wrong assumption or a statement that doesn't mean what you think.
  • Don't treat a persuasive exploratory proof or counterexample as verification — it remains a lead until a verification rung checks it.

Output — Aggregate Verification Report

Write to reviews/<scope>/verify-math/<YYYY-MM-DD-HHMM>.md:

result:   <the theorem/proposition being verified>
claims:   (one row per atomic obligation)
  - <claim 1> | rung R? | verdict | evidence/report link
  - <claim 2> | rung R? | verdict | ...
verdict:  VERIFIED (Lean|CAS) | FALSIFIED (<which claim>) | INCONCLUSIVE(supported|undecided) | ERROR
guarantee: <the weakest link — the result is only as verified as its least-verified obligation>
next:     <what to escalate — e.g. "obligation 3 unfalsified by R1; Lean-prove it (R3)">

Verification (did this skill work?)

  • Each atomic claim has a rung + a sub-verdict + a linked sub-report.
  • The aggregate verdict follows the table (FALSIFIED dominates; VERIFIED requires a proving rung).
  • The guarantee line names the weakest link honestly.

Worked example — 2026-07-04 (median-collapse paper), how it would route

  • Threshold ρ* = (μ_med/μ_max)² (closed-form identity) → R2 symbolic-check → VERIFIED.
  • Prop 3.1 d/dρ Φ(μ/√ρ) < 0 (derivative sign) → R2 → VERIFIED.
  • "Q₀(∞;ρ) monotone in ρ for all competence" (distributional monotonicity) → R1 numerical-check → FALSIFIED (bimodal counterexample) ⇒ aggregate FALSIFIED for that claim; conjecture removed.
  • Median collapse / assumptions (conceptual) → R0 domain-reviewer → substantive check.
  • Aggregate: the paper's symmetric theorem VERIFIED (CAS); the general conjecture FALSIFIED (R1) — exactly the split the paper now reports.

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/verify-math of flonat/flonat-research.

  • SKILL.md
  • scripts/extract_block.py

Open the folder on GitHubat commit da27600

Compare with similar skills

Verify Math next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Math compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Math this skillflonat/flonat-research145—~1.9kAutomated safety check: NotesMIT
Mathparcadei/Continuous-Claude-v33.9k3 repos~1.6kAutomated safety check: NotesMIT
Hermes Agent Skill AuthoringNousResearch/hermes-agent252k—~3.6kAutomated safety check: PassMIT
Configuring Oauth2 Authorization Flowmukul975/Anthropic-Cybersecurity-Skills34k—~1.7kAutomated safety check: PassApache-2.0
Authoring Skillsvercel/next.js143k—~1kAutomated safety check: PassMIT
Abp Authorizationabpframework/abp14k—~1.3kAutomated safety check: PassLGPL-3.0

Similar skills

  • Math

    parcadei/Continuous-Claude-v3

    Unified math capabilities - computation, solving, and explanation.

    3.9k GitHub starsUsed in 3 repos~1.6k tokens
    Research & ScienceAuto-check: notes
  • Hermes Agent Skill Authoring

    NousResearch/hermes-agent

    Author in-repo SKILL.md files: frontmatter and structure. An agent skill from NousResearch/hermes-agent.

    252k GitHub stars~3.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Configuring Oauth2 Authorization Flow

    mukul975/Anthropic-Cybersecurity-Skills

    Configures secure OAuth 2.0 authorization flows, including Authorization Code with PKCE, Client Credentials, and Device Authorization Grant, covering flow selection, PKCE implementation, token…

    34k GitHub stars~1.7k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Authoring Skills

    vercel/next.js

    Official

    How to create and maintain agent skills in .agents/skills/. An agent skill from vercel/next.js.

    143k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Abp Authorization

    abpframework/abp

    ABP permission system - PermissionDefinitionProvider, [Authorize] attribute, CheckPolicyAsync, IsGrantedAsync, ICurrentUser, IPermissionManager, multi-tenancy side.

    14k GitHub stars~1.3k tokensUpdated today
    Backend & APIsAuto-check passed
  • Math Computation

    tradecatlabs/vibe-coding-cn

    Runs reproducible math computations and counterexample searches with SymPy, NumPy and mpmath, logging evidence without presenting results as proofs.

    17k GitHub stars~881 tokensUpdated today
    Research & ScienceAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 9 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 9 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 9 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 9 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check: notes

Questions about Verify Math

What does Verify Math do?

Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report. Verify Math is an agent skill from flonat/flonat-research. Verify a self-authored mathematical result end to end by routing claims across adversarial review, numerical falsification, symbolic or CAS checks, and Lean, then aggregating one report.

When should I use Verify Math?

Verify Math fits situations like: paper-wide mathematical argument needs the appropriate combination of verification methods.

How do I install Verify Math in Claude Code?

Run `npx skills add flonat/flonat-research --skill verify-math -a claude-code`. Or copy the skill folder (skills/verify-math in flonat/flonat-research) into .claude/skills/verify-math in your project. Claude Code loads it when a task matches its description.

How do I install Verify Math in Codex?

Run `npx skills add flonat/flonat-research --skill verify-math -a codex`. Or copy the skill folder (skills/verify-math in flonat/flonat-research) into .agents/skills/verify-math in your project. Codex loads it when a task matches its description.

Can I use Verify Math in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill verify-math -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-math, .gemini/skills/verify-math, .github/skills/verify-math and .opencode/skills/verify-math in your project.

What does Verify Math need to run?

Going by SKILL.md and its folder, Verify Math needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Bash, AskUserQuestion, Skill, Task, Write, Edit.

Does Verify Math access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Verify Math safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Verify Math use?

Verify Math is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Math use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify Math?

Skills that share tags, products or a category with Verify Math: Math (parcadei/Continuous-Claude-v3, 3.9k stars), Hermes Agent Skill Authoring (NousResearch/hermes-agent, 252k stars), Configuring Oauth2 Authorization Flow (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Authoring Skills (vercel/next.js, 143k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Math?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.