Agent skill

Statement Extract And Prove

by OneWave-AI in OneWave-AI/claude-skills

Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus…

MITAuto-check passedDocuments & Office

Install Statement Extract And Prove

skills CLI
$ npx skills add OneWave-AI/claude-skills --skill statement-extract-and-prove -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OneWave-AI/claude-skills statement-extract-and-prove --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/statement-extract-and-prove .claude/skills/statement-extract-and-prove && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
statement-extract-and-prove
GitHub stars
336
Token cost
~2.2k tokens
SKILL.md length
1,138 words
Files
7 (incl. scripts, references)
Skills in repo
69
Repo updated
First seen
Licence
MIT

At a glance

Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus…

  • Works in 7 steps: Check for a text layer → Extract → Prove → …
  • The user wants to convert a bank statement PDF to Excel
  • SKILL.md covers Setup, Workflow, Rules and Files
  • Runs Python scripts from its folder; calls python, pip and pdftoppm

What it does

Statement Extract And Prove is an agent skill from OneWave-AI/claude-skills. Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus transactions equals closing balance, every running balance chains, page totals, row counts and summary totals all tie, signs and decimal separators are consistent. Handles digital PDFs with word-position column bands and detects scanned PDFs that need OCR first. Use it whenever the user wants to convert a bank statement PDF to…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/formats.md`, `references/ocr.md` and `scripts/extract_statement.py`).

It sits in Documents & Office, covering Excel spreadsheets, PDF and Forms and invoices. It works with Microsoft Excel and QuickBooks. The repository describes itself as: 200+ production-ready Claude Code skills for sales, marketing, design, engineering, and AI agent architecture. Built and maintained by OneWave AI. The licence is MIT.

When your agent uses it

  • The user wants to convert a bank statement PDF to Excel
  • Extract transactions from a statement
  • Turn a PDF table into a spreadsheet
  • Import statements into QuickBooks

Example prompts

  • “Use the statement-extract-and-prove skill to convert bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV…”
  • “/statement-extract-and-prove”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Check for a text layer
  2. Extract
  3. Prove
  4. If the proof fails, localize before touching anything
  5. Re-extract the broken page with adjusted bands
  6. Visual reading, last resort only
  7. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit fc5b785. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip
    • pdftoppm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statement Extract And Prove loads about 2.2k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 229 tokens; SKILL.md has 1,138 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~229
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from OneWave-AI/claude-skills at commit fc5b785, republished under its MIT licence (© OneWave-AI). 1,138 words, ~2,242 tokens.

Download SKILL.mdSave it as .claude/skills/statement-extract-and-prove/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
statement-extract-and-prove
description
Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus transactions equals closing balance, every running balance chains, page totals, row counts and summary totals all tie, signs and decimal separators are consistent. Handles digital PDFs with word-position column bands and detects scanned PDFs that need OCR first. Use it whenever the user wants to convert a bank statement PDF to Excel or CSV, extract transactions from a statement, turn a PDF table into a spreadsheet, import statements into QuickBooks or Xero, or says the converted statement does not balance, rows are missing, or signs look wrong, even if they do not ask for a proof. Never reads amounts off the page image when a text layer exists, and never fixes a failed tie-out by guessing.

Statement Extract and Prove

A model reading a statement PDF by eye drops rows on long tables, flips signs, reads 1.234,56 as 1.234, and invents numbers where a scan is faint. Roughly one statement in seven fails to tie on the first pass. The fix is deterministic extraction from the PDF text layer plus arithmetic that the statement itself supplies: the bank already printed the opening balance, the closing balance, often a running balance on every row, and a summary box. If the extracted rows reproduce all of those, the extraction is proven. If not, the arithmetic points at the exact row.

Scope: this skill gets statement and table rows out of PDFs and proves them. For categorizing and reconciling against the books, hand the CSV to bookkeeping-close (its scripts/reconcile.py reads this CSV directly). For pulling header fields out of receipts and invoices (vendor, total, due date), use financial-parser; that skill does field extraction, while this one handles row tables that must tie out.

Setup

pip install pdfplumber openpyxl        # required; openpyxl only for --xlsx
# optional, scanned PDFs only: ocrmypdf + tesseract (see references/ocr.md)

Workflow

1. Check for a text layer

Run the extractor. It checks the text layer first and exits with code 3 and NO TEXT LAYER when the PDF is an image. In that case, OCR it as described in references/ocr.md and continue with the OCR'd file. Do not start reading the page images yourself: OCR plus the proof below catches misreads, and eyeballing catches none.

2. Extract
python scripts/extract_statement.py statement.pdf --out out/stmt --xlsx \
    [--account-type bank|card] [--decimal auto|dot|comma] [--date-order auto|mdy|dmy]
  • --account-type card for credit cards: charges raise the balance owed, so they come out positive and payments negative. For bank accounts, deposits are positive. This matches what bookkeeping-close expects.
  • --decimal defaults to auto: it counts .dd against ,dd endings in the amount columns and stops with exit 2 if the evidence is mixed. Pass the flag explicitly whenever you know the locale, because an explicit flag cannot be fooled by a statement with few amounts.
  • Dates printed without a year get it from the statement period; those rows are flagged year_inferred.
  • Read the summary line it prints: row count, layout, decimal, date order, period, low-confidence rows, unassigned lines. Any WARNING needs a look before you go on.

Outputs: stmt.csv (one row per transaction, with page and y for traceability, raw printed text, and flags), stmt.meta.json (summary box values, forward/closing markers, the column bands used on each page, and every table line that did not become a row), and optionally stmt.xlsx.

3. Prove
python scripts/prove.py out/stmt.csv [--opening X --closing Y --stated-count N]

Hard checks: opening + sum = closing; every printed running balance equals the previous one plus the amounts in between; page continuity (brought forward = carried forward, start + page rows = page end); credits and debits against the summary box and totals rows; row counts against any count the statement prints; duplicates across page breaks; and a number-format check. The number-format check is needed because the other invariants do not depend on scale: a statement where every amount was read 100x too large still ties out perfectly. Exit 0 prints PROVEN; anything else is NOT PROVEN, with proof_report.md, proof.json and low_confidence.csv written.

If the opening balance is not printed and the tool had to derive it from the first row, the verdict says so. Read the opening balance off the statement and pass --opening, because a derived opening proves nothing about the first row.

4. If the proof fails, localize before touching anything

The report tells you where the fault is. Work from it, in this order:

  1. Per-page table: the page whose start + rows does not equal its end is the page with the fault.
  2. First broken running balance: the diagnosis names a culprit row or a span between two y-positions. It identifies: wrong sign (the gap is twice one row's amount), power of ten (a decimal or thousands separator misread), extra row (removing it closes the gap, usually a duplicate or a subtotal read as a transaction), missing row (the gap equals a printed row that is not in the CSV; an unassigned line carrying that amount is named when one exists), no amount (usually wrong column bands), and misread balance (two consecutive breaks of opposite size).
  3. Tie-out explained: when one break accounts for the whole gap, the report prints LOCALIZED. Fix that one thing, not the rest of the statement.
Show full SKILL.md (443 more words)Show less
5. Re-extract the broken page with adjusted bands

Most failures are column bands: a page printed with a different template, a header the detector missed, or amounts drifting left of their header. Look at the real word positions:

python scripts/extract_statement.py statement.pdf --out /tmp/x --dump-words 2

Choose x-ranges that separate the columns, then re-run the whole document with overrides for that page only:

python scripts/extract_statement.py statement.pdf --out out/stmt \
    --bands "date=40-95,description=95-285,amount=285-480,balance=480-570" --band-pages 2
python scripts/prove.py out/stmt.csv

Other levers: --decimal when the number-format check fails, --date-order when dates land outside the period, --account-type when every sign is inverted, --invert-amount for a signed column printed from the other party's view, --allow-integers for price lists without cents. See references/formats.md for layouts and their traps.

6. Visual reading, last resort only

If a page still fails after band and flag adjustments (a damaged scan, handwriting, a stamp over a figure), render just that page (pdftoppm -r 200 -f N -l N -png statement.pdf page) and read only the rows in the localized span. Every value obtained this way goes into the deliverable marked source=visual in the flags column, the proof is re-run, and the handoff says which rows were read by eye. Never type in a number so that the statement ties; if the visual reading does not close the gap, report the gap.

7. Deliver

Hand over the CSV or XLSX, proof_report.md, and the low-confidence queue. State the verdict in one line: "PROVEN: 44 rows, opening 4,210.33 + 25,374.85 = closing 29,585.18, 20 running balances chained, totals and counts match." For NOT PROVEN, give the gap, the localized row or span, and what was tried. For bookkeeping, point to bookkeeping-close and pass the proven balances as --statement-begin / --statement-end.

Rules

  • Extract from the text layer. Only read the page image for a localized span after deterministic options are exhausted, and mark those rows.
  • Report, do not repair. A plug, a dropped "extra" row, or a retyped amount must be confirmed against the page, never made up to close the gap.
  • A tie-out with a derived opening balance, or with number-format conflicts, is not a proof.
  • Keep page and y in every deliverable, so any figure can be traced back to its spot on the PDF.
  • Multi-statement PDFs (a year of statements in one file): split by statement period and prove each statement on its own, because a combined tie-out hides offsetting errors.

Files

  • scripts/extract_statement.py: text-layer check, header and band detection, row assembly, number and date parsing, CSV/XLSX/meta output, --dump-words
  • scripts/prove.py: invariants, diagnosis, report, low-confidence queue
  • references/formats.md: statement layouts, sign conventions, locale number and date formats, invoice and price-list notes
  • references/ocr.md: when to OCR, how, confidence, and re-checking OCR output
  • tests/make_fixtures.py, tests/run_tests.py: synthetic statements (US 3-page running balance, German decimal comma, UK paid-out/paid-in with CR/DR, misaligned page, scanned) and corruption tests

© OneWave-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in statement-extract-and-prove of OneWave-AI/claude-skills.

  • SKILL.md
  • references/formats.md
  • references/ocr.md
  • scripts/extract_statement.py
  • scripts/prove.py
  • tests/make_fixtures.py
  • tests/run_tests.py

Open the folder on GitHubat commit fc5b785

Compare with similar skills

Statement Extract And Prove next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statement Extract And Prove compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statement Extract And Prove this skillOneWave-AI/claude-skills336—~2.2kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Research Integrity Auditxuzhougeng/wisp-science1k—~2.6kAutomated safety check: PassAGPL-3.0
File ReadingWide-Moat/open-computer-use1261 repos~3.1kAutomated safety check: PassProprietary
Compdf Documents To PDFComPDFKit/compdf-skills109—~850Automated safety check: PassNone
Multi Source Data Integration ExtractionDrchronx/ai-agent-research-starter-kit139—~671Automated safety check: PassCustom licence

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • File Reading

    Wide-Moat/open-computer-use

    A skill your agent uses when a file has been uploaded but its content is NOT in your context — only its path at /mnt/user-data/uploads/ is listed in an uploadedfiles block.

    126 GitHub starsUsed in 1 repo~3.1k tokens
    Documents & OfficeAuto-check passed
  • Compdf Documents To PDF

    ComPDFKit/compdf-skills

    Convert Word, Excel, PPT, HTML, TXT, CSV, RTF, PNG, and JPG files into PDF with ComPDF.

    109 GitHub stars~850 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Multi Source Data Integration Extraction

    Drchronx/ai-agent-research-starter-kit

    Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

    139 GitHub stars~671 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Light File Reading

    Light0305/Light-skills

    Light 多格式文件深度理解常驻技能:强大地读 Word / PDF / PPTX / Excel / CSV / 图片 / 视频 / 代码 / 压缩包,不只提取文字,而是理解结构 / 图表 / 数据 / 格式要求 / 隐含意图,产结构化"理解笔记"五面 (结构逻辑·关键内容·格式约束·视觉风格·可复用)并映射到下游技能动作(这个文件→接下来能做什么)。

    640 GitHub stars~4.1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed

More from OneWave-AI/claude-skills

All 69 skills in this repo
  • CRM Data Cleanup

    OneWave-AI/claude-skills

    Finds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM.

    336 GitHub stars~2.5k tokensUpdated 7 days ago
    Auto-check passed
  • Design Export Repair

    OneWave-AI/claude-skills

    Repairs broken decks and PDFs exported from Claude Design or similar AI deck generators: clipped text, wrong fonts and corrupted .pptx package structure.

    336 GitHub stars~2.6k tokensUpdated 7 days ago
    Auto-check passed
  • Bi Measure Builder

    OneWave-AI/claude-skills

    Writes, explains, debugs, and optimizes BI calculations - Power BI / Fabric DAX measures and calculated columns, Tableau calculated fields (FIXED/INCLUDE/EXCLUDE LOD expressions, table…

    336 GitHub stars~2.2k tokensUpdated 7 days ago
    Auto-check passed
  • Bookkeeping Close

    OneWave-AI/claude-skills

    Categorizes transactions, reconciles bank and card statements to the ledger, works a month-end checklist and prepares a close package, without ever forcing a balance.

    336 GitHub stars~2k tokensUpdated 7 days ago
    Auto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    336 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • Sec Filing Puller

    OneWave-AI/claude-skills

    Pulls financial statement numbers for US public companies straight from SEC EDGAR's free official XBRL APIs (companyfacts, companyconcept, frames, submissions) into a cited table.

    336 GitHub stars~2.1k tokensUpdated 7 days ago
    Auto-check passed

Questions about Statement Extract And Prove

What does Statement Extract And Prove do?

Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus…. Statement Extract And Prove is an agent skill from OneWave-AI/claude-skills. Converts bank, credit card and brokerage statement PDFs, invoices and vendor price lists into clean CSV or Excel, then proves the extraction is complete and correct - opening balance plus transactions equals closing balance, every running balance chains, page totals, row counts and summary totals all tie, signs and decimal separators are consistent.

When should I use Statement Extract And Prove?

Statement Extract And Prove fits situations like: the user wants to convert a bank statement PDF to Excel; extract transactions from a statement; turn a PDF table into a spreadsheet; import statements into QuickBooks.

How do I install Statement Extract And Prove in Claude Code?

Run `npx skills add OneWave-AI/claude-skills --skill statement-extract-and-prove -a claude-code`. Or copy the skill folder (statement-extract-and-prove in OneWave-AI/claude-skills) into .claude/skills/statement-extract-and-prove in your project. Claude Code loads it when a task matches its description.

How do I install Statement Extract And Prove in Codex?

Run `npx skills add OneWave-AI/claude-skills --skill statement-extract-and-prove -a codex`. Or copy the skill folder (statement-extract-and-prove in OneWave-AI/claude-skills) into .agents/skills/statement-extract-and-prove in your project. Codex loads it when a task matches its description.

Can I use Statement Extract And Prove in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OneWave-AI/claude-skills --skill statement-extract-and-prove -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/statement-extract-and-prove, .gemini/skills/statement-extract-and-prove, .github/skills/statement-extract-and-prove and .opencode/skills/statement-extract-and-prove in your project.

What does Statement Extract And Prove need to run?

Going by SKILL.md and its folder, Statement Extract And Prove needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip and pdftoppm). Our summary lists: Python 3.

Does Statement Extract And Prove access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Statement Extract And Prove safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Statement Extract And Prove use?

Statement Extract And Prove is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Statement Extract And Prove use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Statement Extract And Prove?

Skills that share tags, products or a category with Statement Extract And Prove: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Research Integrity Audit (xuzhougeng/wisp-science, 1k stars), File Reading (Wide-Moat/open-computer-use, 126 stars) and Compdf Documents To PDF (ComPDFKit/compdf-skills, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statement Extract And Prove?

OneWave-AI (a GitHub organization) maintains it in OneWave-AI/claude-skills, which has 336 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 2, 2026.

Source: OneWave-AI/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.