Agent skill

Extract Document Data

by sickn33 in sickn33/agentic-awesome-skills

Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating.

Apache-2.0Auto-check passedDocuments & Office

Install Extract Document Data

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill extract-document-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills extract-document-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/extract-document-data .claude/skills/extract-document-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extract-document-data
GitHub stars
47k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
390 words
Files
1
Skills in repo
1,394
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating.

  • Works in 4 steps: Get the document. URL or local file path… → Choose the extraction mode → Interpret the response. → …
  • Parsing invoices
  • SKILL.md covers When to use, Instructions, Output format and Limitations and Safety, plus 1 more section
  • Needs STIPPLE_API_KEY

What it does

Extract Document Data is an agent skill from sickn33/agentic-awesome-skills. Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating. Use for parsing invoices, payslips, statements, contracts.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Forms and invoices and Data cleaning. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is Apache-2.0.

When your agent uses it

  • Parsing invoices
  • Tasks that involve Forms and invoices
  • Tasks that involve Data cleaning

Example prompts

  • “/extract-document-data”

Requirements

  • A credential in STIPPLE_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Get the document. URL or local file path (PDF, PNG, JPEG, DOCX).
  2. Choose the extraction mode
  3. Interpret the response.
  4. Report honestly. This is extraction, not verification — values are what the document shows, not proof it's genuine

What it can do on your machine

Read from SKILL.md and the folder at commit 1e53ce2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • stipple.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • STIPPLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extract Document Data loads about 1.1k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 390 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 1e53ce2, republished under its Apache-2.0 licence (© sickn33). 390 words, ~1,111 tokens.

Download SKILL.mdSave it as .claude/skills/extract-document-data/SKILL.md (or your agent's skills folder).
name
extract-document-data
description
Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating. Use for parsing invoices, payslips, statements, contracts.
category
document-verification
risk
critical
source
community
source_repo
Sketchjar/stipple-agent-skills
source_type
community
date_added
2026-08-31
author
Sketchjar
tags
document-verification, fact-checking, stipple, authenticity
tools
claude, cursor, gemini, codex
license
Apache-2.0

Extract Document Data

Extract structured JSON from documents with per-value grounding: every extracted value cites where it came from (page number, confidence), and values that aren't clearly present are reported in not_found rather than hallucinated. Uses the Stipple API (free anonymous tier).

When to use

  • Parsing payslips, invoices, bank statements, receipts, or contracts
  • Converting unstructured documents to JSON for downstream systems
  • Any extraction where hallucinated values are worse than missing values (lending, accounting, compliance)

Instructions

  1. Get the document. URL or local file path (PDF, PNG, JPEG, DOCX).

  2. Choose the extraction mode:

    • Ad-hoc fields — tell the API exactly which fields you want:
      bash
      curl -X POST https://www.stipple.sh/v1/extract \
        -F "file=@payslip.pdf" \
        -F 'fields=[{"name":"employer_name"},{"name":"net_pay"},{"name":"pay_date"}]' \
        -H "Authorization: Bearer $STIPPLE_API_KEY"
    • Template — use a built-in schema: payslip, tax_invoice, bank_statement, receipt, contract
    • Schema-free — omit fields and let the model extract what it finds
  3. Interpret the response.

    json
    {
      "mode": "schema_free",
      "document_type": "payslip",
      "pages_read": 1,
      "fields": {
        "employer_name": {"value": "Acme Cleaning Pty Ltd", "confidence": 0.95, "page": 1},
        "net_pay": {"value": "2845.10", "confidence": 0.97, "page": 1}
      },
      "not_found": ["ytd_tax"]
    }
    • Every value carries confidence (the model's self-report) and page (grounding)
    • not_found[] lists requested fields the model couldn't find — absences are reported, never guessed
    • pages_read shows how many pages were processed (page limits apply per document)
  4. Report honestly. This is extraction, not verification — values are what the document shows, not proof it's genuine:

    • "Employer: Acme Cleaning Pty Ltd (confidence 0.95, page 1)"
    • "ytd_tax: not found in document" — never "ytd_tax: 0" or a guess
    • For "is this document genuine?", pair with the verify-document skill first
Show full SKILL.md (165 more words)Show less

Output format

Payslip fields (grounded, not guessed):

  Employer          Acme Cleaning Pty Ltd  (confidence 0.95, page 1)
  Employee          J. Citizen             (confidence 0.98, page 1)
  Net pay           2,845.10               (confidence 0.97, page 1)
  Superannuation    268.20                 (confidence 0.93, page 1)

not_found: ytd_tax
(absences are reported, never hallucinated)

Limitations and Safety

  • Invoices, statements, payslips, and contracts often contain sensitive personal, financial, or commercial data. Obtain explicit approval before uploading them to a hosted third party, minimize the submitted content, and confirm current retention, residency, access, and deletion terms.
  • Confidence and page grounding do not prove that an extracted value is correct or that the source document is authentic. Reconcile consequential values against the original document and authoritative systems before payment, lending, accounting, compliance, or legal action.
  • Keep the original file and extraction response so a human reviewer can reproduce and correct disputed fields.

Notes

  • Costs 1 credit per page read by the model (minimum 1); free weekly allowance applies
  • Templates: payslip, tax_invoice, bank_statement, receipt, contract — pass as the template form field
  • Tables are extracted with structure preserved; multi-page documents are processed page by page
  • Pairs with verify-document (run first, for authenticity) — an extracted value from a tampered document is still wrong
  • Free key at https://www.stipple.sh for metering beyond the anonymous allowance

© sickn33, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/extract-document-data of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 1e53ce2

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Extract Document Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extract Document Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extract Document Data this skillsickn33/agentic-awesome-skills47k1 repos~1.1kAutomated safety check: PassApache-2.0
Paginated Reportdata-goblin/power-bi-agentic-development1k—~3.5kAutomated safety check: PassGPL-3.0
Data Cleanupsgharlow/claude-code-recipes388—~566Automated safety check: PassCustom licence
Browser Ops SkillOpenLoaf/OpenLoaf107—~1.5kAutomated safety check: PassAGPL-3.0
Sn Da Excel WorkflowMichaelYang-lyx/AIDABench1111 repos~2.5kAutomated safety check: PassNone
Numeric Format NormalizationMichaelYang-lyx/AIDABench1111 repos~505Automated safety check: PassNone

Similar skills

  • Paginated Report

    data-goblin/power-bi-agentic-development

    Author, validate, publish, and test Power BI paginated reports in the RDL format.

    1k GitHub stars~3.5k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Data Cleanup

    sgharlow/claude-code-recipes

    Clean and standardize messy tabular data (CSV, spreadsheet paste, system exports) into an analysis-ready dataset — consistent dates and names, typed columns, duplicates identified, missing values…

    388 GitHub stars~566 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Browser Ops Skill

    OpenLoaf/OpenLoaf

    Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or…

    107 GitHub stars~1.5k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed
  • Numeric Format Normalization

    MichaelYang-lyx/AIDABench

    对 Excel 数据进行数值格式标准化与清洗,支持大规模数据的 Parquet 转换流程,并完成关键指标的合计核对与结果文件导出。

    111 GitHub starsUsed in 1 repo~505 tokens
    Documents & OfficeAuto-check passed
  • Functional

    citypaul/.dotfiles

    Functional programming patterns with immutable data. An agent skill from citypaul/.dotfiles.

    739 GitHub stars~3.6k tokensUpdated 4 days ago
    Testing & QAAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,394 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Extract Document Data

What does Extract Document Data do?

Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating. Extract Document Data is an agent skill from sickn33/agentic-awesome-skills. Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating.

When should I use Extract Document Data?

Extract Document Data fits situations like: parsing invoices; tasks that involve Forms and invoices; tasks that involve Data cleaning.

How do I install Extract Document Data in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill extract-document-data -a claude-code`. Or copy the skill folder (skills/extract-document-data in sickn33/agentic-awesome-skills) into .claude/skills/extract-document-data in your project. Claude Code loads it when a task matches its description.

How do I install Extract Document Data in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill extract-document-data -a codex`. Or copy the skill folder (skills/extract-document-data in sickn33/agentic-awesome-skills) into .agents/skills/extract-document-data in your project. Codex loads it when a task matches its description.

Can I use Extract Document Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill extract-document-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract-document-data, .gemini/skills/extract-document-data, .github/skills/extract-document-data and .opencode/skills/extract-document-data in your project.

What does Extract Document Data need to run?

Going by SKILL.md and its folder, Extract Document Data needs credentials named STIPPLE_API_KEY. Our summary lists: A credential in STIPPLE_API_KEY.

Does Extract Document Data access the network?

SKILL.md names 1 domain. As links in the text: stipple.sh. This is read from the text; nothing was executed.

Is Extract Document Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extract Document Data use?

Extract Document Data is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extract Document Data use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extract Document Data?

Skills that share tags, products or a category with Extract Document Data: Paginated Report (data-goblin/power-bi-agentic-development, 1k stars), Data Cleanup (sgharlow/claude-code-recipes, 388 stars), Browser Ops Skill (OpenLoaf/OpenLoaf, 107 stars) and Sn Da Excel Workflow (MichaelYang-lyx/AIDABench, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extract Document Data?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,304 GitHub stars. The repository holds 1,394 skills in this directory. The repository was last updated on October 6, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.