Agent skill

Math Auditor

by CliMA in CliMA/EnsembleKalmanProcesses.jl

Run an adversarial mathematical-accuracy review of a Julia package's src/ and test/ directories, producing a dated markdown report plus concise, self-contained fix-prompt markdowns suitable for…

Apache-2.0Auto-check passedDevelopment

Install Math Auditor

skills CLI
$ npx skills add CliMA/EnsembleKalmanProcesses.jl --skill math-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CliMA/EnsembleKalmanProcesses.jl math-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CliMA/EnsembleKalmanProcesses.jl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/math-auditor .claude/skills/math-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
math-auditor
GitHub stars
127
Token cost
~2.6k tokens
SKILL.md length
1,174 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run an adversarial mathematical-accuracy review of a Julia package's src/ and test/ directories, producing a dated markdown report plus concise, self-contained fix-prompt markdowns suitable for…

  • Works in 6 steps: Partition → Fan out reviewers (parallel agents) → Verify → …
  • The user asks for an adversarial review
  • SKILL.md covers What "adversarial" means here, Workflow, Calibration and Improving this skill
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Math Auditor is an agent skill from CliMA/EnsembleKalmanProcesses.jl. Run an adversarial mathematical-accuracy review of a Julia package's src/ and test/ directories, producing a dated markdown report plus concise, self-contained fix-prompt markdowns suitable for handing to a smaller model (e.g. Sonnet) in a later session. Use this skill whenever the user asks for an adversarial review, a math audit, a full code review focused on mathematical/numerical correctness, a check of algorithmic consistency with the literature, or says things like "review the math in src/", "is the algebra…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Markdown, Statistics and Code review. The repository describes itself as: Derivative-free parameter calibration and uncertainty quantification for expensive models using ensemble Kalman methods. The licence is Apache-2.0.

When your agent uses it

  • The user asks for an adversarial review
  • A full code review focused on mathematical/numerical correctness
  • A check of algorithmic consistency with the literature
  • Says things like review the math in src/

Example prompts

  • “review the math in src/”
  • “is the algebra right?”
  • “audit the update equations”
  • “/math-auditor”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Partition
  2. Fan out reviewers (parallel agents)
  3. Verify
  4. Write the report
  5. Write fix prompts
  6. Report back

What it can do on your machine

Read from SKILL.md and the folder at commit d10e521. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Math Auditor loads about 2.6k tokens when it runs. Until then it costs about 199 tokens; SKILL.md has 1,174 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~199
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from CliMA/EnsembleKalmanProcesses.jl at commit d10e521, republished under its Apache-2.0 licence (© CliMA). 1,174 words, ~2,570 tokens.

Download SKILL.mdSave it as .claude/skills/math-auditor/SKILL.md (or your agent's skills folder).
name
math-auditor
description
Run an adversarial mathematical-accuracy review of a Julia package's src/ and test/ directories, producing a dated markdown report plus concise, self-contained fix-prompt markdowns suitable for handing to a smaller model (e.g. Sonnet) in a later session. Use this skill whenever the user asks for an adversarial review, a math audit, a full code review focused on mathematical/numerical correctness, a check of algorithmic consistency with the literature, or says things like "review the math in src/", "is the algebra right?", "audit the update equations", "check the statistics/linear algebra for bugs", or "construct a code review as markdown". Trigger even when the user does not say "audit" — any request for a correctness-focused sweep of a scientific Julia codebase qualifies.

Math Audit

Adversarial review of a scientific Julia package for mathematical accuracy and consistency — not software architecture (flag architecture only when it causes mathematical wrongness, e.g. mutation aliasing, accidental type demotion, or inconsistent conventions between modules).

The output is written for the package's own developers: findings must cite exact file:line, state the correct mathematics, and give a concrete failure scenario. A finding that can't survive an attempt at refutation doesn't ship.

What "adversarial" means here

Each reviewer's job is to break the code, not describe it. Concretely, hunt for:

  • Wrong equations: update formulas, gradients, covariances, likelihoods that differ from the cited papers or from the docstring's own LaTeX. Derive the correct expression independently and diff it against the code.
  • Convention drift: rows-vs-columns for ensemble members, N-1 vs N normalization, factor-of-2 / sign errors, Cholesky L vs U, covariance vs precision, whether noise is added in obs-space or transformed space — especially inconsistencies between modules that must agree.
  • Statistical validity: is added noise sampled with the right covariance and scaling (e.g. Δt scaling in stochastic dynamics)? Are means/covariances computed over the right dimension? Deterministic vs stochastic variants actually equivalent in expectation?
  • Numerical soundness: unguarded inv/\ on possibly-singular matrices, loss of symmetry/PSD-ness, subtraction-based variance formulas, missing regularization, sqrt of negative-by-roundoff eigenvalues.
  • Edge cases the math must survive: ensemble size 1–2, dimension 1 (scalar vs matrix degeneracy), zero variance, NaN/failed ensemble members, empty minibatches.
  • Test-math consistency: do the tests actually pin the mathematics (analytic solutions, invariants, convergence rates), or just check shapes and "it runs"? A wrong equation whose test only checks size() is a double finding: the bug and the missing test.

Workflow

1. Partition

List src/*.jl and test/**/*.jl with line counts. Group into 4–8 review units of roughly comparable size, pairing each source module with the tests that exercise it. Group modules that share mathematical conventions together (e.g. all Kalman-update variants in one or two units) so the reviewer can catch cross-module inconsistencies.

2. Fan out reviewers (parallel agents)

Spawn one agent per unit, in a single message so they run concurrently. Each agent prompt must include:

  • the exact file list for its unit,
  • the "What adversarial means here" hunting list above (copy it in — agents don't see this skill),
  • instructions to check code against docstrings/comments and against the standard form of the algorithm from the literature,
  • a required output format: a JSON-like list of findings, each with file, line, severity (critical / major / minor / hygiene), claim (one sentence), evidence (the code vs the correct math), failure_scenario (concrete inputs → wrong output), verified (numerical / inspection), and suggested_fix (optional, a few lines).

Tell agents explicitly:

  • "Prefer few, well-evidenced findings over many speculative ones — but do report genuine minor inconsistencies. If the module's math is correct, say so and note the strongest invariants the tests pin."
  • "When a finding concerns a fixed point, a statistical scaling, or a crash, verify it numerically in a scratch script if cheap — a small linear-Gaussian fixed-point run, a quick Monte Carlo of the statistic, or reproducing the error — and tag it verified: numerical. Numerically verified findings are worth far more than inspection-only ones." (In one audit, the agents that ran code delivered the critical finding pre-verified to 1e-16; the only refuted claim of the run came from an inspection-only unit.)
  • "Before claiming anything is 'silent' or 'has no warning/guard', grep for @warn, @error, and throw at the constructors and call sites of the code path, not just the function you are reading — guards often live at construction time."

As each agent's report arrives, save its raw findings verbatim to a scratchpad file (one per unit). A full audit plus verification is long enough that context summarization mid-run can silently lose findings; the scratchpad files are the durable record the report is assembled from.

Show full SKILL.md (552 more words)Show less
3. Verify

For each critical/major finding, attempt refutation before it enters the report: re-read the cited lines yourself, re-derive the math, and check whether a test or an upstream transformation already accounts for it (common false positives: a transpose hidden in a helper, normalization done at construction time, a convention documented elsewhere, a @warn at the constructor that the reviewer never read — re-run the warn/throw grep yourself for any "silent" claim). Prioritise findings tagged verified: inspection; numerically verified ones usually need only a sanity re-read. Spawn skeptic agents for findings you can't settle from the main context. Demote or drop findings that don't survive; mark surviving ones CONFIRMED vs PLAUSIBLE (couldn't fully verify).

Then build a small conventions matrix from the unit reports before writing anything: rows = modules, columns = the conventions the units commented on (covariance normalization N vs N−1, Δt/noise-scale placement, order of regularization vs localization, RNG threading, rows-vs-columns). Any mismatched cell between modules that must agree is a finding candidate in itself — in practice the worst bugs are a scale factor applied to one block or module but not its sibling, and they only become visible side by side.

4. Write the report

Create full-code-review/<YYYY-MM-DD>/ (date from date +%F, never from memory). Write review.md:

markdown
# Adversarial Mathematical Review — <Package> (<date>)
## Scope and method          <!-- files covered, units, verification policy -->
## Summary table             <!-- ID | severity | verdict | file:line | one-line claim -->
## Critical findings         <!-- full detail: evidence, math, failure scenario, fix sketch -->
## Major findings
## Minor findings & hygiene  <!-- terser -->
## Cross-module consistency notes
## Test-coverage gaps        <!-- where math is unpinned by tests -->
## What was checked and found sound   <!-- credit where due; prevents re-auditing -->

Findings get stable IDs used everywhere, including fix prompts: C1… critical, M1… major, m1… minor, h1… hygiene; grouped-minor fix prompts get G1….

5. Write fix prompts

For each finding with an actionable fix (usually critical + major, plus grouped minors), write full-code-review/<date>/fix-prompts/<ID>-<slug>.md. These are consumed by a smaller model in a fresh session with no context, so each must be self-contained:

markdown
# Fix <ID>: <one-line title>
**File**: `src/Foo.jl`, function `bar!`, around line NNN.
**Problem**: <2–4 sentences: what the code does vs what the math requires.
Include the incorrect snippet verbatim.>
**Required change**: <exact edit, or precise description with the correct formula>
**Do not**: <guardrails — e.g. "do not change the API", "do not touch other methods">
**Verify**: <the test to run or add, with the invariant it should pin>

Keep each under ~40 lines. One finding per file; group only truly mechanical repeats (e.g. the same typo pattern in five docstrings) into one prompt. Quote enough of the offending snippet that the fixer can locate it by function name + snippet — line numbers drift between the audit and the fix session, so present them as hints, not anchors. Also write fix-prompts/README.md listing prompts in recommended application order (independent fixes first, same-file prompts sequenced, conflicting ones flagged) and naming which fixes intentionally change numerical results — so the fixer checks a failing loose regression test against the analytic reference in the prompt before "fixing" the test.

6. Report back

Final message: lead with the headline (how many confirmed critical/major findings and the single worst one), then the report path, then a compact summary table. Do not paste the whole report into the chat.

Calibration

  • Severity: critical = produces mathematically wrong results in mainstream use; major = wrong in common configurations or silently degrades statistical properties; minor = wrong in edge cases, misleading docs math, dead/misnamed math; hygiene = style-level (only if math-adjacent).
  • A docstring–code mismatch is a real finding even when the code is right — users implement against docstrings.
  • A statistically unjustified combination that the package explicitly warns about (e.g. a constructor @warn "... experimental ...") caps at minor unless the warning itself is wrong — the trap isn't silent.
  • Don't pad. If a unit is sound, the report says so in one paragraph; an audit that cries wolf gets ignored next time.

Improving this skill

After delivering the report, offer: "Would you like to improve the math-auditor skill itself using skill-creator? You can share suggestions, or I can analyse this run — finding quality, false-positive rate, fix-prompt usability — to refine the skill for next time."

© CliMA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/math-auditor of CliMA/EnsembleKalmanProcesses.jl.

Open the folder on GitHubat commit d10e521

Compare with similar skills

Math Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Math Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Math Auditor this skillCliMA/EnsembleKalmanProcesses.jl127—~2.6kAutomated safety check: PassApache-2.0
Quarto Authoringquarto-dev/quarto-r1601 repos~1.9kAutomated safety check: PassMIT
Markdown Mermaid WritingK-Dense-AI/scientific-agent-skills48k1 repos~4.2kAutomated safety check: NotesApache-2.0
Format Markdown Tablemaslennikov-ig/claude-code-orchestrator-kit259—~1kAutomated safety check: PassCustom licence
Cc Structured Abstractfranklee16/academic-research-skills2231 repos~921Automated safety check: PassNone
WooCommerce Markdown Guidelineswoocommerce/woocommerce11k1 repos~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Quarto Authoring

    quarto-dev/quarto-r

    A skill your agent uses when the user is explicitly working with Quarto, .qmd files, quarto.yml, Quarto projects, or Quarto features such as callouts, cross-references, citations, Mermaid diagrams…

    160 GitHub starsUsed in 1 repo~1.9k tokens
    DevelopmentAuto-check passed
  • Markdown Mermaid Writing

    K-Dense-AI/scientific-agent-skills

    Writes scientific Markdown documentation and Mermaid diagrams for workflows, relationships, timelines, and schemas.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    DevelopmentAuto-check: notes
  • Format Markdown Table

    maslennikov-ig/claude-code-orchestrator-kit

    Generate well-formatted markdown tables from data with proper alignment and spacing.

    259 GitHub stars~1k tokensUpdated 7 mo ago
    Documents & OfficeAuto-check passed
  • Cc Structured Abstract

    franklee16/academic-research-skills

    A skill your agent uses when writing the Cancer Cell (Cell Press) front matter — the Summary, eTOC blurb, Highlights, and graphical abstract.

    223 GitHub starsUsed in 1 repo~921 tokens
    Data & AnalyticsAuto-check passed
  • WooCommerce Markdown Guidelines

    woocommerce/woocommerce

    Rules for writing and editing markdown in the WooCommerce repository, with the project's markdownlint settings for headings, lists and code blocks.

    11k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Docs Conventions

    flet-dev/flet

    A skill your agent uses when writing or reviewing Flet documentation, including Python docstrings (Google style, reST roles, admonitions), Markdown docs (cross-references, images, code examples)…

    17k GitHub stars~1.6k tokensUpdated today
    DevelopmentAuto-check passed

More from CliMA/EnsembleKalmanProcesses.jl

  • Slurm Pipeline Manager

    CliMA/EnsembleKalmanProcesses.jl

    Scaffold and maintain a SLURM/HPC job-dependency tree for an EnsembleKalmanProcesses.jl (EKP) calibration pipeline.

    127 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Base Show

    CliMA/EnsembleKalmanProcesses.jl

    Add concise Base.show and Base.summary methods to Julia types whose default REPL representation is unhelpful or overwhelming.

    127 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Docstrings

    CliMA/EnsembleKalmanProcesses.jl

    Add or normalise Julia docstrings on public symbols (exported types, functions, and constants) so the package's public API is fully self-documenting and the Documenter.jl docs build passes its…

    127 GitHub stars~5.5k tokensUpdated today
    Auto-check passed
  • Error Message Manager

    CliMA/EnsembleKalmanProcesses.jl

    Rewrite vague, delayed, or low-context Julia error messages into structured, actionable diagnostics.

    127 GitHub stars~8.2k tokensUpdated today
    Auto-check passed

Questions about Math Auditor

What does Math Auditor do?

Run an adversarial mathematical-accuracy review of a Julia package's src/ and test/ directories, producing a dated markdown report plus concise, self-contained fix-prompt markdowns suitable for…. jl.g.

When should I use Math Auditor?

Math Auditor fits situations like: the user asks for an adversarial review; A full code review focused on mathematical/numerical correctness; A check of algorithmic consistency with the literature; says things like review the math in src/.

How do I install Math Auditor in Claude Code?

Run `npx skills add CliMA/EnsembleKalmanProcesses.jl --skill math-auditor -a claude-code`. Or copy the skill folder (.claude/skills/math-auditor in CliMA/EnsembleKalmanProcesses.jl) into .claude/skills/math-auditor in your project. Claude Code loads it when a task matches its description.

How do I install Math Auditor in Codex?

Run `npx skills add CliMA/EnsembleKalmanProcesses.jl --skill math-auditor -a codex`. Or copy the skill folder (.claude/skills/math-auditor in CliMA/EnsembleKalmanProcesses.jl) into .agents/skills/math-auditor in your project. Codex loads it when a task matches its description.

Can I use Math Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CliMA/EnsembleKalmanProcesses.jl --skill math-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/math-auditor, .gemini/skills/math-auditor, .github/skills/math-auditor and .opencode/skills/math-auditor in your project.

What does Math Auditor need to run?

SKILL.md names no scripts, command-line tools or credentials: Math Auditor is instructions for the agent only.

Does Math Auditor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Math Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Math Auditor use?

Math Auditor is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Math Auditor use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Math Auditor?

Skills that share tags, products or a category with Math Auditor: Quarto Authoring (quarto-dev/quarto-r, 160 stars), Markdown Mermaid Writing (K-Dense-AI/scientific-agent-skills, 48k stars), Format Markdown Table (maslennikov-ig/claude-code-orchestrator-kit, 259 stars) and Cc Structured Abstract (franklee16/academic-research-skills, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Math Auditor?

CliMA (a GitHub organization) maintains it in CliMA/EnsembleKalmanProcesses.jl, which has 127 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: CliMA/EnsembleKalmanProcesses.jl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.