Agent skill

Parsing Documents

by GAIK-project in GAIK-project/gaik-toolkit

Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

MITAuto-check passedDocuments & Office

Install Parsing Documents

skills CLI
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .claude/skills/parsing-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
parsing-documents
GitHub stars
100
Token cost
~2.3k tokens
SKILL.md length
1,045 words
Files
2 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

  • Pulling tables out of a document
  • SKILL.md covers Choose by what must survive,…, Why the cheap path scores zero…, Parsing quality is usually a… and Use the API correctly, plus 2 more sections
  • Calls pip
  • Running OCR on scans

What it does

Parsing Documents is an agent skill from GAIK-project/gaik-toolkit. Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. Use when reading a PDF or DOCX into text, pulling tables out of a document, running OCR on scans, feeding documents into a RAG pipeline or an LLM, deciding between PyMuPDF, Docling, and vision-LLM parsing, or when a parse appeared to succeed but the tables, columns, or whole pages came out wrong or empty. Also use when…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/parser-selection.md`).

It sits in Documents & Office, covering Document parsing, Word documents and Computer vision. It works with Microsoft Word. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.

When your agent uses it

  • Pulling tables out of a document
  • Running OCR on scans
  • Feeding documents into a RAG pipeline
  • Deciding between PyMuPDF

Example prompts

  • “Use the parsing-documents skill to convert PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the…”
  • “/parsing-documents”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Parsing Documents loads about 2.3k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 1,045 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~175
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 1,045 words, ~2,308 tokens.

Download SKILL.mdSave it as .claude/skills/parsing-documents/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
parsing-documents
description
Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. Use when reading a PDF or DOCX into text, pulling tables out of a document, running OCR on scans, feeding documents into a RAG pipeline or an LLM, deciding between PyMuPDF, Docling, and vision-LLM parsing, or when a parse appeared to succeed but the tables, columns, or whole pages came out wrong or empty. Also use when document parsing is costing more time or money than expected. Covers parser selection, per-page verification, and escalation from cheap local parsing to vision models.

Parsing documents with gaik

bash
pip install "gaik[parser]"              # PyMuPDF, python-docx, Docling
pip install "gaik[multimodal-parser]"   # multi-provider vision parsing
# Add llm-google for shared Google/Vertex configs, or llm-litellm for LiteLLM.

Choose by what must survive, not by file type

A parse that returns fluent, plausible text can still have destroyed the one thing the task needed. Decide first what has to still be true afterwards, then pick:

What must surviveParserCost
Plain prose, simple single-column layoutPyMuPDFParserfree, local, milliseconds
A Word document's textDocxParserfree, local
Text on scans / no text layerDoclingParser (OCR)slow on CPU, free
Table structure — merged cells, multi-row headersMultimodalParser or VisionParserAPI calls
Images explained in place, for RAG chunksVisionPlusParserDocling + API call
Docling quality without the local installDoclingApiClientParsera Docling service (api_base=, password=)

Escalate only when a check fails — start at the cheapest row that could plausibly work, verify (below), and move down one row if it did not. When the text is headed for a search index, the searching-documents skill picks up from here.

Why the cheap path scores zero on tables

Measured on a public benchmark's table split (40 documents, GriTS and TEDS scored against ground-truth HTML table trees):

ParserOutput formatTable structure score
Vision-LLM parsing (MultimodalParser)HTML <table>0.90 – 0.96
Docling, serialized as HTMLHTML <table>0.89
Docling, as shipped (use_markdown=True)markdown pipe table0.00
PyMuPDFplain text0.00

The zeros are not "much worse" — they are structurally unable to score, and that is the transferable point. The output format decides what can survive. A markdown pipe table has no way to express a merged cell or a two-row header, so a document containing one comes back looking clean and quietly wrong. Plain text loses column boundaries entirely.

So the rule is not "always use vision". It is: if the tables carry merged cells or stacked headers, the parser must emit HTML — and among the paths that do, vision-LLM parsing led the specialized parser, with the gap widest on the messiest layouts.

Treat those numbers as a dated snapshot on one corpus, not a constant. What generalizes is the format argument; re-measure the ranking on documents that look like yours.

Parsing quality is usually a traceability decision, not an accuracy one

The expensive counterexample, measured on an extraction task: feeding the model the native PDF, plain extracted text, or model-produced HTML gave the same F1. Parsing cost 22–25× the wall time and was 93% of total spend, and bought no accuracy at all.

What it did buy was bounding boxes. When the parser returns coordinates for each element, a citation can be matched to a box afterwards, and a human reviewer clicks a highlighted region on the page. The model cannot invent that, and both halves stay independently checkable.

Decide on that basis:

  • The deliverable is a reviewer clicking a highlighted box → pay for the rich parse.
  • Page-level evidence is enough, or nothing is reviewed by hand → do not. Feed the model the document and skip the parsing bill.

Use the API correctly

The class method and the module-level function of the same name do not return the same type, which is the easiest mistake to make here:

python
from gaik.software_components.parsers import PyMuPDFParser, parse_pdf

text = PyMuPDFParser().parse_pdf("doc.pdf")     # -> str
result = parse_pdf("doc.pdf")                   # -> dict
text = result["text_content"]

Every parse_document returns a dict, and the key differs by parser:

python
from gaik.software_components.parsers import (
    DocxParser, DoclingParser, VisionParser, VisionPlusParser, MultimodalParser,
)
from gaik.software_components.llm import get_llm_config

DocxParser().parse_docx("doc.docx")                             # -> str
DoclingParser().parse_document("scan.pdf")["text_content"]      # OCR
config = get_llm_config("azure")                               # explicit provider
VisionPlusParser(vision_config=config).parse_document(
    "doc.pdf"
)["parsed_markdown"]                                            # note: different key
VisionParser(openai_config=config).convert_image("page.jpg")    # -> str
MultimodalParser(api_config=config).parse("doc.pdf")             # -> ParseResult

VisionPlusParser and DoclingApiClientParser take required keyword arguments — vision_config=, and api_base= plus password= — and a bare constructor call fails.

MultimodalParser takes keyword arguments only. Use api_config=get_llm_config(...) for the shared provider interface; the legacy model_provider flags read credentials from the environment. It has no config parameter. VisionParser uses openai_config and VisionPlusParser uses vision_config for the same shared dictionary. Choose native openai, azure, google, vertex, anthropic, anthropic_foundry, aitta, openai_compatible, or optional litellm through get_llm_config. The configured model must accept images. Shared multimodal parsing renders PDF pages as PNG images, so native PDF uploads are not required.

ParseResult is a plain dataclass with raw_markdown, clean_markdown, html (populated only when create_html=True) and usage; it has no save() method, so write the files yourself. DoclingParser has no parse() method.

Set merge_table=True when a table runs across a page break — it instructs the model to stitch the halves back together, which no local parser can do.

For which environment variables each provider needs, read references/parser-selection.md.

Show full SKILL.md (369 more words)Show less

Verify before building on the output

Parsers fail quietly far more often than they raise, so check the output rather than the exception. Three checks catch nearly everything:

1. Emptiness, per page — never per document. In one measured corpus 18 of 66 pages had no text layer, spread across half the documents. A document-level if not text check passes such a document as normal and those pages simply never reach the model: no error, no warning, no missing file. Loop the pages and assert each one produced characters; report which page numbers came back empty.

2. Table structure, if tables matter. Search the output for <table> (or pipe rows). If the source has a merged cell and the output has no <table>, the structure is already gone — escalate rather than patch the text.

3. A known token round-trip. Pick a handful of values you can see in the document — a total, an invoice number, a date — and assert they appear in the parsed text. This catches column-collapse and page-drop, which both otherwise read as fine prose.

Gotchas

  • parse_document returns text_content on PyMuPDFParser and DoclingParser, but parsed_markdown on VisionPlusParser and DoclingApiClientParser. Same method name, different key — read the dict, don't assume.
  • Docling on CPU runs roughly 20–30 s/page. A GPU makes it faster but does not change its accuracy, so never reach for Docling to improve a quality result you measured on CPU — the number will be identical.
  • DoclingParser requires the parser extra, not parser-cpu.
  • VisionParser accepts all shared provider configs, including native Google/Vertex, Anthropic, Aitta and optional LiteLLM. The selected model must support image inputs. MultimodalParser(api_config=...) uses the same interface. LiteLLM requires gaik[llm-litellm] and a provider-prefixed model identifier; capabilities still depend on that backend/model.
  • Audio transcription components require OpenAI/Azure audio configs. Aitta, native Google/Anthropic, and an arbitrary compatible chat endpoint do not acquire audio support through the shared chat client. Use separate audio and text stage configs.
  • On Windows, write parsed output with encoding="utf-8" explicitly. Path.write_text() defaults to the platform codepage, which raises on characters a document parser routinely produces — and a crashed write downstream looks exactly like a bad parse.
  • A parser returning fluent text is not evidence it read the whole document. Only the per-page check is.

© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents of GAIK-project/gaik-toolkit.

  • SKILL.md
  • references/parser-selection.md

Open the folder on GitHubat commit e516ece

Compare with similar skills

Parsing Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Parsing Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Parsing Documents this skillGAIK-project/gaik-toolkit100—~2.3kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Document ConversionHarryoung/efka104—~505Automated safety check: PassApache-2.0
Doc To Markdowndaymade/claude-code-skills1.4k—~2.5kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Document Conversion

    Harryoung/efka

    Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.

    104 GitHub stars~505 tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Doc To Markdown

    daymade/claude-code-skills

    Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.

    1.4k GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.

    276k GitHub starsUsed in 4 repos~1.5k tokens
    Documents & OfficeAuto-check passed

More from GAIK-project/gaik-toolkit

All 15 skills in this repo
  • Brief To Slides

    GAIK-project/gaik-toolkit

    Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…

    100 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated today
    Auto-check passed
  • Extracting Structured Data

    GAIK-project/gaik-toolkit

    Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…

    100 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Searching Documents

    GAIK-project/gaik-toolkit

    Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

    100 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Construction Diary Creation

    GAIK-project/gaik-toolkit

    Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.

    100 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Gaik Add Examples

    GAIK-project/gaik-toolkit

    Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/.

    100 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes

Works with

Questions about Parsing Documents

What does Parsing Documents do?

Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. Parsing Documents is an agent skill from GAIK-project/gaik-toolkit. Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

When should I use Parsing Documents?

Parsing Documents fits situations like: pulling tables out of a document; running OCR on scans; feeding documents into a RAG pipeline; deciding between PyMuPDF.

How do I install Parsing Documents in Claude Code?

Run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a claude-code`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents in GAIK-project/gaik-toolkit) into .claude/skills/parsing-documents in your project. Claude Code loads it when a task matches its description.

How do I install Parsing Documents in Codex?

Run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a codex`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents in GAIK-project/gaik-toolkit) into .agents/skills/parsing-documents in your project. Codex loads it when a task matches its description.

Can I use Parsing Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/parsing-documents, .gemini/skills/parsing-documents, .github/skills/parsing-documents and .opencode/skills/parsing-documents in your project.

What does Parsing Documents need to run?

Going by SKILL.md and its folder, Parsing Documents needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Parsing Documents access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Parsing Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Parsing Documents use?

Parsing Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Parsing Documents use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Parsing Documents?

Skills that share tags, products or a category with Parsing Documents: Markitdown (ImCa0/just-laws, 782 stars), Markitdown (jimmc414/Kosmos, 595 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Document Conversion (Harryoung/efka, 104 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Parsing Documents?

GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.