Agent skill

Numerical Check

by flonat in flonat/flonat-research

Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations.

MITAuto-check: notesTesting & QA

Install Numerical Check

skills CLI
$ npx skills add flonat/flonat-research --skill numerical-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research numerical-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/numerical-check .claude/skills/numerical-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
numerical-check
GitHub stars
146
Token cost
~2.2k tokens
SKILL.md length
928 words
Files
1
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations.

  • Works in 7 steps: Formalize the claim as a predicate over… → Sample the domain to approximate the… → Evaluate on an INTERIOR grid, smoothly,… → …
  • Checking monotonicity
  • SKILL.md covers When to Use, When NOT to Use, Position in the verification… and Procedure, plus 5 more sections
  • Calls uv

What it does

Numerical Check is an agent skill from flonat/flonat-research. Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations. Use when checking monotonicity, thresholds, inequalities, comparative statics, or limits computationally. For algebraic proof or Lean formalization, use $symbolic-check or $lean-check.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Load testing. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Checking monotonicity
  • Comparative statics
  • Limits computationally

Example prompts

  • “/numerical-check”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, AskUserQuestion

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Formalize the claim as a predicate over a domain
  2. Sample the domain to approximate the TRUE object — not finite-n atoms
  3. Evaluate on an INTERIOR grid, smoothly, with a noise-aware tolerance
  4. Sweep, count, capture the worst — seeded
  5. Diagnose a surprise BEFORE trusting it
  6. Characterize the violators — mechanism, not just a rate
  7. Emit the verification report + wire numbers if feeding a paper

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Numerical Check loads about 2.2k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 928 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 928 words, ~2,186 tokens.

Download SKILL.mdSave it as .claude/skills/numerical-check/SKILL.md (or your agent's skills folder).
name
numerical-check
description
Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations. Use when checking monotonicity, thresholds, inequalities, comparative statics, or limits computationally. For algebraic proof or Lean formalization, use $symbolic-check or $lean-check.
allowed-tools
Read, Write, Edit, Bash, AskUserQuestion

Numerical Check: Falsify a Self-Authored Math Claim by Sweep

Empirically stress-test a mathematical claim you wrote but have not proven. The goal is falsification: throw many random instances at the claim and try to break it. A single genuine counterexample kills the claim; a large clean sweep is evidence, never proof.

When to Use

  • You wrote a Proposition / Theorem / Conjecture (monotonicity, threshold, comparative-static, inequality, closed-form, limit) and want to know if it's actually true before claiming it.
  • numerical-check, "stress-test my conjecture", "find a counterexample to X", "is Q(ρ) really monotone", "does the threshold hold for all …".
  • The write-time empirical arm of the mark-unverified rule (self-authored math must be checked before assertion).

When NOT to Use

SituationUse instead
Verify an algebra / derivative / limit / closed-form identitysymbolic-check (R2)
Machine-prove a lemma (want a proof, not a stress-test)lean-check (R3)
Re-verify a computed empirical result in another languagecross-language-check
Conceptual / assumption-completeness reviewdomain-reviewer (agent)

Position in the verification spectrum

R1 — numerical falsification. Can FALSIFY definitively (a confirmed counterexample refutes the claim) but can never VERIFY (no counterexample ≠ proof). The strongest positive result is INCONCLUSIVE (supported): no counterexample in N draws. Pair with lean-check (R3) to prove the claim once it survives.

Procedure

1. Formalize the claim as a predicate over a domain

Restate the claim as P(x) that must hold for all x in a domain D. Make the failure condition explicit and quantitative.

  • "Q(ρ) is monotone decreasing in ρ" → P(instance) := max_i (Q(ρ_{i+1}) − Q(ρ_i)) ≤ tol over a ρ-grid.
  • "threshold ρ* separates help/hurt" → P := (Q<p_max) iff (ρ>ρ*).
  • Write down the domain D precisely (which parameters, which ranges, which side-conditions — e.g. "mean competence > ½, dispersed").
2. Sample the domain to approximate the TRUE object — not finite-n atoms

This is the step that fools people. If the claim is about a continuous or large-n limit object, a tiny discrete instance is NOT that object — it carries finite-n artifacts (ties, atoms, degenerate medians, staircase discontinuities) that manufacture fake violations.

  • If the claim is a large-n / continuous-distribution statement, represent each random instance with dense sampling (hundreds of points), so the computed quantity approximates the limit.
  • Generate instances from varied shapes (uniform, skewed, bimodal, heavy-tailed) so the sweep is adversarial, not cherry-picked.
3. Evaluate on an INTERIOR grid, smoothly, with a noise-aware tolerance
  • Grid the parameter on the interior (e.g. ρ ∈ [0.02, 0.98]); endpoints breed boundary/degeneracy artifacts.
  • Prefer a smooth evaluation (e.g. root-find the crossing) over indicator-quadrature, which quantization-jitters and creates false steps.
  • Set the violation tolerance an order of magnitude above the numerical noise floor (measure the floor on a case you believe holds). Too tight → false positives; too loose → misses real breaks.
4. Sweep, count, capture the worst — seeded
  • Run over thousands of random instances (uv run --no-project --with numpy --with scipy python; never bare python3). Seed the RNG.
  • Count genuine violations; capture the worst counterexample (the instance + violation magnitude) for reporting and for the figure.
5. Diagnose a surprise BEFORE trusting it

If you find violations (or a suspiciously high/low rate), do not report the raw number yet. Check it is not an artifact:

  • Re-plot the worst case on a fine grid — is the "violation" a real interior feature, or a jump at an endpoint / a finite-n tie?
  • Re-run with denser sampling and odd vs even n — does the rate persist as you approach the continuous limit? An artifact shrinks; a real effect stays.
  • Only once it survives these is it real. (Document the surprise + the diagnosis — design-before-results.)
Show full SKILL.md (355 more words)Show less
6. Characterize the violators — mechanism, not just a rate

A bare "X% violate" is weak; find when it breaks. Break the sweep down by instance feature (shape, skew, competence-gap, bimodality) and report the driver: "non-monotonicity is a bimodal phenomenon — 18% of bimodal vs ~0% unimodal." This turns a number into a result.

7. Emit the verification report + wire numbers if feeding a paper

Write the report (shape below). If the result feeds a LaTeX paper, emit every number via a generated macro file (results-numbers.tex, no-hardcoded-results) and keep the seeded script in experiments/.

Script skeleton (adapt; keep it seeded + uv-run)

python
# uv run --no-project --with numpy --with scipy --with matplotlib python <script>.py
import numpy as np; from scipy.stats import norm; from scipy.optimize import brentq
RNG = np.random.default_rng(0)
def quantity(instance, t):        # smooth eval of the claimed object at parameter t
    ...                           # prefer root-find over indicator-quadrature
def sweep(n_trials, nsamp):       # dense instances, interior grid, noise-aware tol
    grid = np.linspace(0.02, 0.98, 90); viol = 0; worst = (-1, None); by = {}
    for _ in range(n_trials):
        inst = draw_instance(RNG, nsamp)          # varied shapes, dense
        if not in_domain(inst): continue
        Q = np.array([quantity(inst, t) for t in grid])
        up = np.diff(Q).max()                     # violation statistic
        if up > worst[0]: worst = (up, inst)
        if up > 1e-4: viol += 1; bump(by, feature(inst))
    return viol, worst, by                        # characterize by feature

Anti-Patterns

  • Don't test a continuous/large-n claim with tiny discrete instances — finite-n ties/atoms fabricate violations. (2026-07-04: a small-n sweep read 29%; dense sampling read the real 6%.)
  • Don't evaluate at exact parameter endpoints — degeneracies live there.
  • Don't set the tolerance near the numerical-noise floor (false positives) or absurdly loose (misses real breaks).
  • Don't trust the first surprising number — diagnose artifact-vs-real first.
  • Don't report a bare violation rate — characterize the driver.
  • Don't claim VERIFIED — numerical can only FALSIFY or SUPPORT. Say "no counterexample in N draws".
  • Don't use bare python3 — it's blocked by the uv-only allowlist; use uv run.
  • Don't confirm "clean" from a compile log or a grep alone — check the actual object (e.g. grep the rendered PDF for ??, not the build log).

Output — Verification Report (shared *-check shape)

Write to reviews/<scope>/verify-numerical/<YYYY-MM-DD-HHMM>.md:

claim:    <the exact statement tested, with its domain>
method:   R1 numerical falsification (N=<trials>, nsamp=<density>, grid=<interior>, tol=<t>)
verdict:  FALSIFIED | INCONCLUSIVE (supported: no counterexample in N) | INCONCLUSIVE | ERROR
evidence: <worst counterexample instance + magnitude>  OR  <"no counterexample; sweep params">
mechanism:<what feature drives violations, if any>
reproduce: uv run --no-project --with ... python experiments/<script>.py   (seed=<s>)

Verification (did this skill work?)

  • The script runs seeded via uv run and prints a violation count + worst case.
  • The verdict is one of the four values; VERIFIED is never emitted.
  • If a paper consumes the result, numbers come from a generated macro file (nothing hand-typed).

Worked example — 2026-07-04 (median-collapse paper)

Claim: "Q₀(∞;ρ) is monotone decreasing in ρ for all competence distributions (mean>½, dispersed)" → single-threshold conjecture.

  • Naive small-n sweep: 29% violations → suspicious; worst case was a balanced even-n council.
  • Diagnosed: finite-even-n ties + endpoint staircase — an artifact.
  • Corrected (dense sampling, interior grid, smooth root-find eval, tol=10⁻⁴): 6% real violations, worst step +0.013.
  • Characterized: non-monotonicity is a bimodal phenomenon (18% of bimodal vs 3/3363 unimodal).
  • Verdict: FALSIFIED — the conjecture was removed and replaced with the characterized finding.

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/numerical-check of flonat/flonat-research.

Open the folder on GitHubat commit da27600

Compare with similar skills

Numerical Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Numerical Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Numerical Check this skillflonat/flonat-research146—~2.2kAutomated safety check: NotesMIT
Writing Livekit Scenarioslivekit-examples/agent-starter-python2641 repos~2.5kAutomated safety check: PassMIT
Go Testingcxuu/golang-skills1721 repos~1.3kAutomated safety check: PassApache-2.0
Goalcraftgrp06/goalcraft102—~3.8kAutomated safety check: PassMIT
Thinking Partnermattnowdev/thinking-partner206—~4.4kAutomated safety check: PassMIT
Visionkunchenguid/vision331—~2.9kAutomated safety check: PassMIT

Similar skills

  • Writing Livekit Scenarios

    livekit-examples/agent-starter-python

    Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

    264 GitHub starsUsed in 1 repo~2.5k tokens
    Testing & QAAuto-check passed
  • Go Testing

    cxuu/golang-skills

    A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.

    172 GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Goalcraft

    grp06/goalcraft

    Turn a rough draft, vague ambition, or messy task brief into a powerful Codex /goal objective for persistent, evidence-checked work.

    102 GitHub stars~3.8k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Thinking Partner

    mattnowdev/thinking-partner

    A deterministic thinking partner that challenges assumptions and applies mental models to sharpen decisions, solve problems, and think more clearly.

    206 GitHub stars~4.4k tokensUpdated 6 mo ago
    Testing & QAAuto-check passed
  • Vision

    kunchenguid/vision

    Draft and stress-test a VISION.md for a repository, then iterate with the author on an interactive review board until approved.

    331 GitHub stars~2.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Volt Load Testing

    owenHochwald/volt

    Safely exercise and evaluate HTTP APIs with the Volt CLI, including authenticated requests, JSON bodies, staged load, machine-readable results, performance baselines, and before/after comparisons.

    141 GitHub stars~1.2k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    146 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    146 GitHub stars~4.4k tokensUpdated 10 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    146 GitHub stars~1.2k tokensUpdated 10 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    146 GitHub stars~488 tokensUpdated 10 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    146 GitHub stars~1.6k tokensUpdated 10 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    146 GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check: notes

Categories

Questions about Numerical Check

What does Numerical Check do?

Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations. Numerical Check is an agent skill from flonat/flonat-research. Numerically stress-test a self-authored mathematical claim over its parameter space to seek counterexamples or characterize violations.

When should I use Numerical Check?

Numerical Check fits situations like: checking monotonicity; comparative statics; limits computationally.

How do I install Numerical Check in Claude Code?

Run `npx skills add flonat/flonat-research --skill numerical-check -a claude-code`. Or copy the skill folder (skills/numerical-check in flonat/flonat-research) into .claude/skills/numerical-check in your project. Claude Code loads it when a task matches its description.

How do I install Numerical Check in Codex?

Run `npx skills add flonat/flonat-research --skill numerical-check -a codex`. Or copy the skill folder (skills/numerical-check in flonat/flonat-research) into .agents/skills/numerical-check in your project. Codex loads it when a task matches its description.

Can I use Numerical Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill numerical-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/numerical-check, .gemini/skills/numerical-check, .github/skills/numerical-check and .opencode/skills/numerical-check in your project.

What does Numerical Check need to run?

Going by SKILL.md and its folder, Numerical Check needs the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, AskUserQuestion.

Does Numerical Check access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Numerical Check safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Numerical Check use?

Numerical Check is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Numerical Check use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Numerical Check?

Skills that share tags, products or a category with Numerical Check: Writing Livekit Scenarios (livekit-examples/agent-starter-python, 264 stars), Go Testing (cxuu/golang-skills, 172 stars), Goalcraft (grp06/goalcraft, 102 stars) and Thinking Partner (mattnowdev/thinking-partner, 206 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Numerical Check?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 146 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.