Agent skill

Mineru

by Nebutra in Nebutra/MinerU-Skill

An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

MITAuto-check passedDocuments & Office

Install Mineru

skills CLI
$ npx skills add Nebutra/MinerU-Skill --skill mineru -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Nebutra/MinerU-Skill mineru --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mineru
GitHub stars
123
Token cost
~1.4k tokens
SKILL.md length
301 words
Files
63 (incl. scripts, references, assets)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

  • Converting PDF/Word/PPT/Excel/image to Markdown
  • SKILL.md covers Zero-config quick start (no…, Power mode (token) — large…, Supported modalities and Common options, plus 6 more sections
  • Calls python3, uv and pip; reaches mineru.net; needs MINERU_TOKEN
  • Extracting text

What it does

Mineru is an agent skill from Nebutra/MinerU-Skill. An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. Works with NO token via the lightweight Agent API and auto-upgrades to the Standard API (token) for large files, batches, and DOCX/HTML/LaTeX export. Use when: (1) Converting PDF/Word/PPT/Excel/image to Markdown, (2) Extracting text, tables, formulas, or running OCR on scanned docs, (3) Batch-parsing a folder in parallel, (4) Piping parsed Markdown straight back to an…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 67 other files, including scripts, reference files and assets (for example `.claude-plugin/plugin.json`, `.github/dependabot.yml` and `.github/workflows/ci.yml`).

It sits in Documents & Office, covering PDF, LaTeX and Word documents. It works with LaTeX, Microsoft Word, Obsidian and Microsoft Excel. The repository describes itself as: AI-Native document parser: PDF, Office & images → clean Markdown with LaTeX, tables & OCR. Zero-dependency CLI & skill for Claude Code, Cursor & AI agents. The licence is MIT.

When your agent uses it

  • Converting PDF/Word/PPT/Excel/image to Markdown
  • Extracting text
  • Running OCR on scanned docs
  • Batch-parsing a folder in parallel

Example prompts

  • “/mineru”

Requirements

  • Python 3
  • A credential in MINERU_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 02433d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • mineru.net

    Also links to:

    • peps.python.org
    • docs.astral.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINERU_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mineru loads about 1.4k tokens when it runs, and up to ~9.5k if it reads all its reference files. Until then it costs about 136 tokens; SKILL.md has 301 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Nebutra/MinerU-Skill at commit 02433d8, republished under its MIT licence (© Nebutra). 301 words, ~1,394 tokens.

Download SKILL.mdSave it as .claude/skills/mineru/SKILL.md (or your agent's skills folder). This skill also uses 62 other files; get the full folder from GitHub.
name
mineru
description
An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. Works with NO token via the lightweight Agent API and auto-upgrades to the Standard API (token) for large files, batches, and DOCX/HTML/LaTeX export. Use when: (1) Converting PDF/Word/PPT/Excel/image to Markdown, (2) Extracting text, tables, formulas, or running OCR on scanned docs, (3) Batch-parsing a folder in parallel, (4) Piping parsed Markdown straight back to an agent or into Obsidian.
homepage
https://mineru.net
metadata.author
Nebutra
metadata.version
3.3.1
metadata.argument-hint
<pdf-file-or-url>

MinerU PDF Parser

Parse PDF, Office (Word/PPT/Excel), and image files into clean Markdown — with LaTeX formulas, tables, images, and OCR. One zero-dependency script, two backends, automatic routing.

Zero-config quick start (no token, no install)

bash
# Parse a local file or URL — the Agent API needs no login
python3 scripts/mineru.py paper.pdf

# Pipe the Markdown straight back to an agent
python3 scripts/mineru.py paper.pdf --stdout

# Machine-readable status for tool pipelines
python3 scripts/mineru.py paper.pdf --json

No pip install, no API key. The free Agent API handles files ≤ 10 MB / ≤ 20 pages.

Run with uv (zero-install, managed Python)

scripts/mineru.py carries PEP 723 inline metadata, so uv runs it directly — no venv, no pip install, with a uv-managed interpreter:

bash
uv run scripts/mineru.py paper.pdf --stdout       # zero-install run
uv run --no-project --with pytest pytest -q       # dev suite via uv

Power mode (token) — large files, batches, extra formats

bash
export MINERU_TOKEN="..."          # https://mineru.net/apiManage/token

# Parallel batch a directory, resume on re-run
python3 scripts/mineru.py ./pdfs/ --output ./out/ --workers 8 --resume

# Export DOCX/HTML/LaTeX alongside Markdown (auto-routes to the Standard API)
python3 scripts/mineru.py report.pdf --format docx --format latex

When a token is set, the tool auto-routes: small single files still use the free Agent API; anything large (> 10 MB / > 20 pages), batched, or needing extra export formats uses the Standard API (≤ 200 MB / ≤ 200 pages). If the Agent API hits a size/page limit, it auto-escalates to the Standard API.

Supported modalities

ModalityExtensionsOCR
PDF.pdf--ocr
Image.png .jpg .jpeg .jp2 .webp .gif .bmpbuilt-in
Word.doc .docx—
Slides.ppt .pptx—
Sheet.xls .xlsx—
HTML.html (Standard API, MinerU-HTML model)—

Common options

INPUT...          One or more files, a directory, or a URL
--output, -o      Output directory (default: ./output)
--api             auto | agent | standard   (default: auto)
--model           pipeline | vlm | MinerU-HTML  (default: vlm)
--format          docx | html | latex  (repeatable; forces Standard API)
--lang            OCR/document language (default: ch)
--ocr             Enable OCR for scanned documents
--pages           Page range, e.g. "1-10" or "2,4-6"
--workers, -w     Concurrent submit/upload/download slots (default: 8)
--resume          Skip inputs already parsed
--stdout          Print Markdown to stdout
--json            Print machine-readable status to stdout
--to SINK         Deliver into a content tool (repeatable); --list-sinks to enumerate
--obsidian PATH   Shortcut for --to obsidian with this vault
--engine          cloud | local | auto  (local/auto parse born-digital PDFs offline)
--split           Split oversized PDFs past the page caps, parse parts, merge (needs pypdf)
--chunk           Emit heading-aware RAG chunks (.chunks.json + --json)
--doctor          Environment self-check and exit

MCP server

Expose MinerU over MCP (zero-dependency stdio JSON-RPC) so an MCP host can call it:

bash
python3 scripts/mineru_mcp.py

Tools: mineru_parse, mineru_parse_to (parse + deliver to sinks), mineru_list_sinks.

Deliver into your tools (--to)

Parse once and push the Markdown into content tools via each one's official path:

bash
python3 scripts/mineru.py paper.pdf --to obsidian --to notion --to feishu

Targets: obsidian logseq siyuan notion linear yuque coda slack feishu confluence onenote ticktick dingtalk airtable wecom (all zero-dependency), plus roam and wps via optional extras. Each reads its config from env vars (run --list-sinks). Per-target auth, fidelity, and image notes: references/integrations.md.

Output

output/
└── document-name/
    ├── document-name.md    # clean Markdown
    └── images/             # extracted figures (Standard API)

Performance (real, measured)

End-to-end latency for the official demo PDF via the free Agent API: cold ≈ 14 s · warm ≈ 13 s (submit → poll → download). Batches scale with --workers. Numbers come from the no-mock live benchmark in tests/test_live.py.

Testing

bash
python3 -m pytest                      # fast unit suite (offline)
MINERU_LIVE=1 python3 -m pytest -m live -s   # real API + benchmark (no mocks)

API Reference

See references/api_reference.md. Official docs: https://mineru.net/apiManage/docs · Token: https://mineru.net/apiManage/token

© Nebutra, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 62 other files (scripts, references, assets) in the repository root of Nebutra/MinerU-Skill.

  • SKILL.md
  • .claude-plugin/plugin.json
  • .github/dependabot.yml
  • .github/workflows/ci.yml
  • .github/workflows/publish-skill.yml
  • .github/workflows/release.yml
  • .gitignore
  • LICENSE
  • README.md
  • README_CN.md
  • assets/mineru-skill.jpg
  • pyproject.toml
  • references/api_reference.md
  • references/comparison.md
  • references/integrations.md
  • requirements.txt
  • … and 47 more

Open the folder on GitHubat commit 02433d8

Compare with similar skills

Mineru next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mineru compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mineru this skillNebutra/MinerU-Skill123—~1.4kAutomated safety check: PassMIT
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Lexoid CLIoidlabs-com/Lexoid109—~2kAutomated safety check: NotesApache-2.0
Document ConverterBlackBeltTechnology/pi-agent-dashboard315—~999Automated safety check: PassMIT
Markdown Exporterbowenliang123/markdown-exporter2721 repos~5.3kAutomated safety check: PassApache-2.0
Lexoid Pythonoidlabs-com/Lexoid109—~3.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Lexoid CLI

    oidlabs-com/Lexoid

    Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI.

    109 GitHub stars~2k tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    315 GitHub stars~999 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    272 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed
  • Lexoid Python

    oidlabs-com/Lexoid

    Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library.

    109 GitHub stars~3.5k tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes
  • Docling

    zhuzhaoyun/Molio

    PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown.

    433 GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed

More from Nebutra/MinerU-Skill

  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    123 GitHub stars~504 tokensUpdated 16 days ago
    Auto-check passed

Questions about Mineru

What does Mineru do?

An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. Mineru is an agent skill from Nebutra/MinerU-Skill. An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

When should I use Mineru?

Mineru fits situations like: converting PDF/Word/PPT/Excel/image to Markdown; extracting text; running OCR on scanned docs; batch-parsing a folder in parallel.

How do I install Mineru in Claude Code?

Run `npx skills add Nebutra/MinerU-Skill --skill mineru -a claude-code`. Or copy the skill folder (the Nebutra/MinerU-Skill repository) into .claude/skills/mineru in your project. Claude Code loads it when a task matches its description.

How do I install Mineru in Codex?

Run `npx skills add Nebutra/MinerU-Skill --skill mineru -a codex`. Or copy the skill folder (the Nebutra/MinerU-Skill repository) into .agents/skills/mineru in your project. Codex loads it when a task matches its description.

Can I use Mineru in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Nebutra/MinerU-Skill --skill mineru -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mineru, .gemini/skills/mineru, .github/skills/mineru and .opencode/skills/mineru in your project.

What does Mineru need to run?

Going by SKILL.md and its folder, Mineru needs the command-line tools its instructions call (python3, uv and pip) and credentials named MINERU_TOKEN. Our summary lists: Python 3; A credential in MINERU_TOKEN.

Does Mineru access the network?

SKILL.md names 3 domains. In commands or code: mineru.net; the agent is likely to contact it when it follows the instructions. As links in the text: peps.python.org and docs.astral.sh. This is read from the text; nothing was executed.

Is Mineru safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mineru use?

Mineru is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mineru use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.2k tokens, read only when the agent opens those files.

What are the alternatives to Mineru?

Skills that share tags, products or a category with Mineru: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Lexoid CLI (oidlabs-com/Lexoid, 109 stars), Document Converter (BlackBeltTechnology/pi-agent-dashboard, 315 stars) and Markdown Exporter (bowenliang123/markdown-exporter, 272 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mineru?

Nebutra (a GitHub organization) maintains it in Nebutra/MinerU-Skill, which has 123 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 24, 2026.

Source: Nebutra/MinerU-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.