Agent skill

Extracting Structured Data

by GAIK-project in GAIK-project/gaik-toolkit

Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…

MITAuto-check passedDocuments & Office

Install Extracting Structured Data

skills CLI
$ npx skills add GAIK-project/gaik-toolkit --skill extracting-structured-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GAIK-project/gaik-toolkit extracting-structured-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/extracting-structured-data .claude/skills/extracting-structured-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extracting-structured-data
GitHub stars
100
Token cost
~3.2k tokens
SKILL.md length
1,587 words
Files
2 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…

  • Pulling fields from invoices
  • SKILL.md covers Choose providers separately…, Fix the schema, not the prompt, The binding limit is property… and Three failures that look like…, plus 5 more sections
  • Calls pip; needs AITTA_API_KEY and AITTA_API_TOKEN
  • Scanned documents into JSON

What it does

Extracting Structured Data is an agent skill from GAIK-project/gaik-toolkit. Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable evidence. Use when pulling fields from invoices, forms, contracts, reports, or scanned documents into JSON or a Pydantic model; when an extraction request fails with a 400, returns a truncated list, or silently drops its deepest fields; when deciding whether to parse a document first or feed the PDF straight to the…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/measuring-extraction.md`).

It sits in Documents & Office, covering Schema markup, PDF and Database schema design. It works with Pydantic and OpenAI. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.

When your agent uses it

  • Pulling fields from invoices
  • Scanned documents into JSON
  • A Pydantic model
  • An extraction request fails with a 400

Example prompts

  • “Use the extracting-structured-data skill to extract structured data — fields, tables, line items — out of documents into a validated schema using…”
  • “/extracting-structured-data”

Requirements

  • Python 3
  • A credential in AITTA_API_KEY
  • A credential in AITTA_API_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AITTA_API_KEY
    • AITTA_API_TOKEN
    • AITTA_TOKEN
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extracting Structured Data loads about 3.2k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 195 tokens; SKILL.md has 1,587 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~195
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 1,587 words, ~3,246 tokens.

Download SKILL.mdSave it as .claude/skills/extracting-structured-data/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
extracting-structured-data
description
Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable evidence. Use when pulling fields from invoices, forms, contracts, reports, or scanned documents into JSON or a Pydantic model; when an extraction request fails with a 400, returns a truncated list, or silently drops its deepest fields; when deciding whether to parse a document first or feed the PDF straight to the model; when adding page numbers, quotes, or confidence scores as evidence; or when extraction accuracy needs to be measured rather than guessed. Covers schema design, property limits, prompt-versus-schema tuning, and citation verification.

Extracting structured data with gaik

bash
pip install "gaik[extract]"
pip install "gaik[vision-extract]"   # PDF/image extraction
# Add llm-google for shared Google/Vertex configs, or llm-litellm for LiteLLM.

Two entry points, and the choice is about where the document is in its lifecycle:

  • VisionExtractor — PDF or image straight to structured data in one call. It renders the pages, so it sees layout, stamps, and handwriting.
  • DataExtractor — takes document text you already parsed. Pair it with the parsing-documents skill.
python
from gaik.software_components.vision_extractor import VisionExtractor
from gaik.software_components.llm import get_llm_config

config = get_llm_config("azure")   # or another supported vision/structured-output model
result = VisionExtractor(api_config=config).extract(
    file_paths=["invoice.pdf"],
    user_requirements="Supplier name, invoice number, and every line item with quantity and unit price.",
)
result.data          # dict
result.usage         # None on the shared structured-output path; do not invent token cost

Arguments are keyword-only. With no extraction_model, gaik generates the schema from the prose (an extra LLM call); pass schema_dir=... and it caches the generated schema to disk and reuses it, which makes runs both cheaper and reproducible.

For text extraction, DataExtractor(config=get_llm_config(provider)) accepts the shared configuration from gaik.software_components.llm. Choose azure, openai, google, vertex, anthropic, anthropic_foundry, aitta, openai_compatible, or optional litellm; the model must support structured output. Native Google/Vertex and Anthropic need the llm-google and llm-anthropic extras respectively. LiteLLM requires gaik[llm-litellm] and a provider-prefixed model identifier. Aitta uses AITTA_API_KEY (aliases AITTA_API_TOKEN, AITTA_TOKEN) and AITTA_MODEL, defaulting to google/gemma-4-31b-it. Other compatible servers require OPENAI_API_KEY, OPENAI_BASE_URL, and OPENAI_MODEL, or equivalent explicit config overrides.

VisionExtractor(api_config=config) uses the same interface, renders PDF pages as PNG images, and generates schemas with that provider. Its model must support both images and structured output. SchemaGenerator(config=config) supports the shared provider path; use a hand-built Pydantic schema and ExtractionRequirements when you already know the schema. LLMJudge(config=config) also accepts the shared config; visual judging requires images, while judge_text_pair(...) only needs a chat model. Keep credentials in runtime configuration or environment variables, not generated source code.

Choose providers separately for audio, vision and extraction

The modules accept stage configs so transcription can stay on a supported audio service while text extraction uses Aitta or another chat provider:

bash
pip install "gaik[audio-to-structured-data,documents-to-structured-data]"
python
from gaik.software_modules.audio_to_structured_data import AudioToStructuredData
from gaik.software_modules.documents_to_structured_data import DocumentsToStructuredData

audio_pipeline = AudioToStructuredData(
    transcription_config=get_llm_config("azure"),
    extraction_config=get_llm_config("aitta"),
)
document_pipeline = DocumentsToStructuredData(
    parser_config=get_llm_config("openai", model="gpt-6-luna"),
    extraction_config=get_llm_config("aitta"),
)

Omitted stage configs use api_config; without a shared config, missing stages load the legacy OpenAI/Azure default. Supplying every stage config avoids loading unrelated default credentials. Audio transcription remains OpenAI/Azure-only: a chat provider selection, including LiteLLM, does not make toolkit transcription support that provider.

Fix the schema, not the prompt

This is the highest-leverage thing to know, because the instinct is the opposite.

Five deliberately different extraction prompts — stripped-down, "read the characters first", a role frame, "walk the pages before answering" — produced identical F1 on every document and every repeat. One of them had been written specifically to counter a known corpus error and did not move a single value. The verbose variants cost about 30% more per document for exactly the same answer.

The reason is that a well-specified schema has already pinned the task: a description on each field, a closed object, and structured output constraining the response shape regardless of what the prose asks for. Prompt wording only has room to matter when the schema is vague — and then fixing the schema is still the better move.

So when extraction is wrong: add or sharpen field descriptions, tighten types, close the object. Do not spend the afternoon rewording the instruction.

The binding limit is property count, not schema size

Providers reject large schemas, but the thing being counted is the number of properties, not characters. Measured on one provider: 200 properties passed and 250 failed; a 300,000-character schema passed while carrying 100 properties.

$ref reuse does not help. The natural assumption is that a definition appearing once in $defs and referenced eight times counts once — it does not. Tested directly: 8 × 50 properties inline (18,263 characters) and the same thing via $ref (2,663 characters) were rejected identically. The schema is expanded before counting, so the refactor that shrinks it sevenfold on disk changes nothing.

The numbers are provider-specific and they move. The method is what to carry: generate synthetic schemas at increasing property counts and bisect until one is rejected. That takes about ten minutes and replaces a guess with a number.

Three failures that look like degrees of one problem

They are not, and the fix for two of them does nothing for the third:

SymptomCauseFix
400, request refused outrightschema too largesplit the schema
Output arrives but the deepest fields are missinga nesting layer droppedsplit the schema
MAX_TOKENS, truncated or no outputthe output list ran longwindow the document

Splitting the schema cannot help the third: however you divide the fields, one call is still asked to return the whole list. Staged and one-shot pipelines fail identically at a thousand rows. That one needs the document chunked into windows and results concatenated.

Only put fields in the schema that the document actually contains

Two traps, both of which produce numbers that look like model errors:

A field with no ground truth does not belong in the schema. Adding a unit field with expected value null because the source system does not record it means a correctly extracted "KG" scores as a mistake. Leave it out.

A field the document does not carry measures enrichment, not extraction. If the truth holds an internal part number and the page shows the supplier's code, reaching one from the other needs a lookup table. Scoring it grades the wrong system.

Both are caught the same cheap way: check every expected value against the document text when building the evaluation set.

Show full SKILL.md (743 more words)Show less

Evidence that is worth something

Ask for evidence in the same call as the value. A separate "now justify your previous answer" call sees its own earlier output and confirms it — that measures the model's agreement with itself. In the same call, the quote constrains the value instead of decorating it afterwards.

Self-reported confidence usually ranks nothing. In one run, 960 citations carried three distinct confidence values. Raising the threshold to 0.99 dropped the accepted set's precision below what accepting everything gave, while sending three quarters of the work back to a human. Before building a threshold on a confidence score, run len(set(scores)). Three means there is no ordering to threshold on — no new run required.

A quote is valuable because it is text you can run a rule over, not because it points somewhere. Once you hold the quoted string, a deterministic rule can screen it with no model, no ground truth, and no annotation — for example, flagging a number written so that two readings differ by a factor of a thousand. The specific rule is corpus-specific; the method is not. Look at which quotes separate the wrong values from the right ones, and write that down as a rule.

Unverified evidence is not evidence. The page number and the quote are both tokens the model generated, and either can be wrong while reading perfectly. What makes them evidence is opening the document and checking the quoted characters are on the named page. When checking, look at the named page first and on its own terms — scoring every page and taking the best match reports correct citations as wrong whenever a common string appears in several places.

gaik ships the wrapping form of this: VisionExtractor(include_verification=True) wraps each scalar in {value, confidence_score, confidence_reason}, and results land in result.verification. Wrapping buys paths that cannot drift onto the wrong field, because the path is a fact about the response tree rather than a string the model wrote. It costs roughly four times the property count — and property count is the limit that decides whether the request is accepted at all. A separate citation table costs a flat +5 properties regardless of schema size. Choose against your distance from the limit.

Parse first only when you need boxes

Feeding the model the native PDF, plain extracted text, or model-produced HTML gave the same F1 in measured comparison, while parsing cost 22–25× the wall time. Parsing earns its price when the deliverable is a reviewer clicking a highlighted region, because a parser that returns bounding boxes lets a quote be matched to a box afterwards. See the parsing-documents skill for the full argument.

Measuring whether any of this worked

Comparing two configurations, or answering "did that change help", has its own set of ways to get a confidently wrong number — including the fact that stability can be measured with no ground truth at all, which unblocks measurement on the day rather than after the labelling project. Read references/measuring-extraction.md before reporting any figure or comparing two runs.

Gotchas

  • generate_schema() returns only the model, but DataExtractor.extract() also needs requirements. Use generate_schema_with_usage(), which returns both on SchemaGenerationResult (.schema, .requirements), plus usage when the client exposes it. Shared parsed-output clients may not provide token counts.
  • DataExtractor.extract(documents=...) takes document text, not file paths. VisionExtractor.extract(file_paths=...) takes paths. Passing paths to DataExtractor extracts from the literal filename string and returns confident nonsense.
  • Without a shared api_config, VisionExtractor and MultimodalParser retain legacy use_azure=True and vertex_ai=True defaults. Supplying a direct provider key without choosing the direct hosting flag produces an auth error that reads like a bad key. An explicit shared provider config supersedes these routing flags.
  • Every layer that calls a model must classify its own errors. A rate-limit response arriving through a layer that does not recognise it gets re-raised as permanent, the runner does not retry, and the affected documents score zero — which is indistinguishable from model instability in the aggregate. A run whose variance is 0 but whose score is 0 is a document that produced nothing, not an unstable model. Read the failure column before calling anything unstable.
  • A failed document must score zero and stay in the sample. Dropping it lets a pipeline improve its own average by crashing on the hardest inputs.
  • Stability alone rewards silence: a branch that leaves deep fields empty is perfectly repeatable, because empty is always the same empty. Read a stability metric next to a completeness metric or it will point the wrong way.

© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in implementation_layer/no-code-assets/agent-plugin/skills/extracting-structured-data of GAIK-project/gaik-toolkit.

  • SKILL.md
  • references/measuring-extraction.md

Open the folder on GitHubat commit e516ece

Compare with similar skills

Extracting Structured Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extracting Structured Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extracting Structured Data this skillGAIK-project/gaik-toolkit100—~3.2kAutomated safety check: PassMIT
Glmocr SDKzai-org/GLM-skills476—~2.7kAutomated safety check: NotesApache-2.0
Azure AI Document Intelligence TSmicrosoft/skills3.1k6 repos~2.4kAutomated safety check: PassMIT
PDF Text Replaceinstavm/coderunner893—~536Automated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Glmocr SDK

    zai-org/GLM-skills

    Trigger when: (1) User wants to extract text, tables, formulas, or structured data from images/PDFs/scanned documents, (2) User mentions "OCR", "文字识别", "文档解析", (3) User has a document (screenshot…

    476 GitHub stars~2.7k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check: notes
  • Official

    Extract text, tables, and structured data from documents using Azure Document Intelligence (@azure-rest/ai-document-intelligence).

    3.1k GitHub starsUsed in 6 repos~2.4k tokens
    Documents & OfficeAuto-check passed
  • PDF Text Replace

    instavm/coderunner

    Replace text in fillable PDF forms by updating form field values.

    893 GitHub stars~536 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Kimi PDF

    thvroyal/kimi-skills

    Professional PDF solution. An agent skill from thvroyal/kimi-skills.

    238 GitHub stars~1.9k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

More from GAIK-project/gaik-toolkit

All 15 skills in this repo
  • Brief To Slides

    GAIK-project/gaik-toolkit

    Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…

    100 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated today
    Auto-check passed
  • Parsing Documents

    GAIK-project/gaik-toolkit

    Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

    100 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Searching Documents

    GAIK-project/gaik-toolkit

    Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

    100 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Construction Diary Creation

    GAIK-project/gaik-toolkit

    Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.

    100 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Gaik Add Examples

    GAIK-project/gaik-toolkit

    Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/.

    100 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes

Works with

Questions about Extracting Structured Data

What does Extracting Structured Data do?

Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…. Extracting Structured Data is an agent skill from GAIK-project/gaik-toolkit. Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable evidence.

When should I use Extracting Structured Data?

Extracting Structured Data fits situations like: pulling fields from invoices; scanned documents into JSON; A Pydantic model; an extraction request fails with a 400.

How do I install Extracting Structured Data in Claude Code?

Run `npx skills add GAIK-project/gaik-toolkit --skill extracting-structured-data -a claude-code`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/extracting-structured-data in GAIK-project/gaik-toolkit) into .claude/skills/extracting-structured-data in your project. Claude Code loads it when a task matches its description.

How do I install Extracting Structured Data in Codex?

Run `npx skills add GAIK-project/gaik-toolkit --skill extracting-structured-data -a codex`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/extracting-structured-data in GAIK-project/gaik-toolkit) into .agents/skills/extracting-structured-data in your project. Codex loads it when a task matches its description.

Can I use Extracting Structured Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill extracting-structured-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-structured-data, .gemini/skills/extracting-structured-data, .github/skills/extracting-structured-data and .opencode/skills/extracting-structured-data in your project.

What does Extracting Structured Data need to run?

Going by SKILL.md and its folder, Extracting Structured Data needs the command-line tools its instructions call (pip) and credentials named AITTA_API_KEY, AITTA_API_TOKEN, AITTA_TOKEN and OPENAI_API_KEY. Our summary lists: Python 3; A credential in AITTA_API_KEY; A credential in AITTA_API_TOKEN.

Does Extracting Structured Data access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Extracting Structured Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extracting Structured Data use?

Extracting Structured Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extracting Structured Data use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Extracting Structured Data?

Skills that share tags, products or a category with Extracting Structured Data: Glmocr SDK (zai-org/GLM-skills, 476 stars), Azure AI Document Intelligence TS (microsoft/skills, 3.1k stars), PDF Text Replace (instavm/coderunner, 893 stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extracting Structured Data?

GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.