Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .claude/skills/parsing-documents && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .claude/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .agents/skills/parsing-documents && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .agents/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .cursor/skills/parsing-documents && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .cursor/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GAIK-project/gaik-toolkit.git --path implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .gemini/skills/parsing-documents && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .gemini/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GAIK-project/gaik-toolkit parsing-documentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .github/skills/parsing-documents && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .github/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GAIK-project/gaik-toolkit parsing-documents --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents .opencode/skills/parsing-documents && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "parsing-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents into .opencode/skills/parsing-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "parsing-documents", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
parsing-documentsConverts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.
Parsing Documents is an agent skill from GAIK-project/gaik-toolkit. Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. Use when reading a PDF or DOCX into text, pulling tables out of a document, running OCR on scans, feeding documents into a RAG pipeline or an LLM, deciding between PyMuPDF, Docling, and vision-LLM parsing, or when a parse appeared to succeed but the tables, columns, or whole pages came out wrong or empty. Also use when…
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/parser-selection.md`).
It sits in Documents & Office, covering Document parsing, Word documents and Computer vision. It works with Microsoft Word. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.
Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Parsing Documents loads about 2.3k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 1,045 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 1,045 words, ~2,308 tokens.
.claude/skills/parsing-documents/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.pip install "gaik[parser]" # PyMuPDF, python-docx, Docling
pip install "gaik[multimodal-parser]" # multi-provider vision parsing
# Add llm-google for shared Google/Vertex configs, or llm-litellm for LiteLLM.A parse that returns fluent, plausible text can still have destroyed the one thing the task needed. Decide first what has to still be true afterwards, then pick:
| What must survive | Parser | Cost |
|---|---|---|
| Plain prose, simple single-column layout | PyMuPDFParser | free, local, milliseconds |
| A Word document's text | DocxParser | free, local |
| Text on scans / no text layer | DoclingParser (OCR) | slow on CPU, free |
| Table structure — merged cells, multi-row headers | MultimodalParser or VisionParser | API calls |
| Images explained in place, for RAG chunks | VisionPlusParser | Docling + API call |
| Docling quality without the local install | DoclingApiClientParser | a Docling service (api_base=, password=) |
Escalate only when a check fails — start at the cheapest row that could plausibly work,
verify (below), and move down one row if it did not. When the text is headed for a search
index, the searching-documents skill picks up from here.
Measured on a public benchmark's table split (40 documents, GriTS and TEDS scored against ground-truth HTML table trees):
| Parser | Output format | Table structure score |
|---|---|---|
Vision-LLM parsing (MultimodalParser) | HTML <table> | 0.90 – 0.96 |
| Docling, serialized as HTML | HTML <table> | 0.89 |
Docling, as shipped (use_markdown=True) | markdown pipe table | 0.00 |
| PyMuPDF | plain text | 0.00 |
The zeros are not "much worse" — they are structurally unable to score, and that is the transferable point. The output format decides what can survive. A markdown pipe table has no way to express a merged cell or a two-row header, so a document containing one comes back looking clean and quietly wrong. Plain text loses column boundaries entirely.
So the rule is not "always use vision". It is: if the tables carry merged cells or stacked headers, the parser must emit HTML — and among the paths that do, vision-LLM parsing led the specialized parser, with the gap widest on the messiest layouts.
Treat those numbers as a dated snapshot on one corpus, not a constant. What generalizes is the format argument; re-measure the ranking on documents that look like yours.
The expensive counterexample, measured on an extraction task: feeding the model the native PDF, plain extracted text, or model-produced HTML gave the same F1. Parsing cost 22–25× the wall time and was 93% of total spend, and bought no accuracy at all.
What it did buy was bounding boxes. When the parser returns coordinates for each element, a citation can be matched to a box afterwards, and a human reviewer clicks a highlighted region on the page. The model cannot invent that, and both halves stay independently checkable.
Decide on that basis:
The class method and the module-level function of the same name do not return the same type, which is the easiest mistake to make here:
from gaik.software_components.parsers import PyMuPDFParser, parse_pdf
text = PyMuPDFParser().parse_pdf("doc.pdf") # -> str
result = parse_pdf("doc.pdf") # -> dict
text = result["text_content"]Every parse_document returns a dict, and the key differs by parser:
from gaik.software_components.parsers import (
DocxParser, DoclingParser, VisionParser, VisionPlusParser, MultimodalParser,
)
from gaik.software_components.llm import get_llm_config
DocxParser().parse_docx("doc.docx") # -> str
DoclingParser().parse_document("scan.pdf")["text_content"] # OCR
config = get_llm_config("azure") # explicit provider
VisionPlusParser(vision_config=config).parse_document(
"doc.pdf"
)["parsed_markdown"] # note: different key
VisionParser(openai_config=config).convert_image("page.jpg") # -> str
MultimodalParser(api_config=config).parse("doc.pdf") # -> ParseResultVisionPlusParser and DoclingApiClientParser take required keyword arguments —
vision_config=, and api_base= plus password= — and a bare constructor call fails.
MultimodalParser takes keyword arguments only. Use api_config=get_llm_config(...)
for the shared provider interface; the legacy model_provider flags read credentials
from the environment. It has no config parameter. VisionParser uses openai_config
and VisionPlusParser uses vision_config for the same shared dictionary. Choose native
openai, azure, google, vertex, anthropic, anthropic_foundry, aitta,
openai_compatible, or optional litellm through get_llm_config. The configured model
must accept images. Shared multimodal parsing renders PDF pages as PNG images, so native
PDF uploads are not required.
ParseResult is a plain dataclass with
raw_markdown, clean_markdown, html (populated only when create_html=True) and
usage; it has no save() method, so write the files yourself. DoclingParser has no
parse() method.
Set merge_table=True when a table runs across a page break — it instructs the model to
stitch the halves back together, which no local parser can do.
For which environment variables each provider needs, read
references/parser-selection.md.
Parsers fail quietly far more often than they raise, so check the output rather than the exception. Three checks catch nearly everything:
1. Emptiness, per page — never per document. In one measured corpus 18 of 66 pages had
no text layer, spread across half the documents. A document-level if not text check
passes such a document as normal and those pages simply never reach the model: no error, no
warning, no missing file. Loop the pages and assert each one produced characters; report
which page numbers came back empty.
2. Table structure, if tables matter. Search the output for <table> (or pipe rows).
If the source has a merged cell and the output has no <table>, the structure is already
gone — escalate rather than patch the text.
3. A known token round-trip. Pick a handful of values you can see in the document — a total, an invoice number, a date — and assert they appear in the parsed text. This catches column-collapse and page-drop, which both otherwise read as fine prose.
parse_document returns text_content on PyMuPDFParser and DoclingParser, but
parsed_markdown on VisionPlusParser and DoclingApiClientParser. Same method name,
different key — read the dict, don't assume.DoclingParser requires the parser extra, not parser-cpu.VisionParser accepts all shared provider configs, including native Google/Vertex,
Anthropic, Aitta and optional LiteLLM. The selected model must support image inputs.
MultimodalParser(api_config=...) uses the same interface. LiteLLM requires
gaik[llm-litellm] and a provider-prefixed model identifier; capabilities still depend
on that backend/model.encoding="utf-8" explicitly. Path.write_text()
defaults to the platform codepage, which raises on characters a document parser routinely
produces — and a crashed write downstream looks exactly like a bad parse.© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents of GAIK-project/gaik-toolkit.
Open the folder on GitHubat commit e516ece
Parsing Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Parsing Documents this skillGAIK-project/gaik-toolkit | 100 | — | ~2.3k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 782 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 595 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Document ConversionHarryoung/efka | 104 | — | ~505 | Automated safety check: Pass | Apache-2.0 | |
| Doc To Markdowndaymade/claude-code-skills | 1.4k | — | ~2.5k | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
Harryoung/efka
Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.
daymade/claude-code-skills
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
affaan-m/ECC
Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.
GAIK-project/gaik-toolkit
Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…
GAIK-project/gaik-toolkit
GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.
GAIK-project/gaik-toolkit
Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…
GAIK-project/gaik-toolkit
Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…
GAIK-project/gaik-toolkit
Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.
GAIK-project/gaik-toolkit
Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/.
Works with
Categories
Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. Parsing Documents is an agent skill from GAIK-project/gaik-toolkit. Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.
Parsing Documents fits situations like: pulling tables out of a document; running OCR on scans; feeding documents into a RAG pipeline; deciding between PyMuPDF.
Run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a claude-code`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents in GAIK-project/gaik-toolkit) into .claude/skills/parsing-documents in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a codex`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/parsing-documents in GAIK-project/gaik-toolkit) into .agents/skills/parsing-documents in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill parsing-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/parsing-documents, .gemini/skills/parsing-documents, .github/skills/parsing-documents and .opencode/skills/parsing-documents in your project.
Going by SKILL.md and its folder, Parsing Documents needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Parsing Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Parsing Documents: Markitdown (ImCa0/just-laws, 782 stars), Markitdown (jimmc414/Kosmos, 595 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Document Conversion (Harryoung/efka, 104 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.