Agent skill

Mineru Document Explorer

by opendatalab in opendatalab/MinerU-Document-Explorer

MinerU Document Explorer — Agent-native knowledge engine. An agent skill from opendatalab/MinerU-Document-Explorer.

MITAuto-check: notesDocuments & Office

Install Mineru Document Explorer

skills CLI
$ npx skills add opendatalab/MinerU-Document-Explorer --skill mineru-document-explorer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opendatalab/MinerU-Document-Explorer mineru-document-explorer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opendatalab/MinerU-Document-Explorer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mineru-document-explorer .claude/skills/mineru-document-explorer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mineru-document-explorer
GitHub stars
638
Token cost
~6.2k tokens
SKILL.md length
2,008 words
Files
2 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

MinerU Document Explorer — Agent-native knowledge engine. An agent skill from opendatalab/MinerU-Document-Explorer.

  • Works in 6 steps: Check qmd is installed → Check Python for binary document support → Check and install Python packages → …
  • Users ask to search their documents
  • SKILL.md covers Quick Reference, Agent Principles, Key Concepts and Playbook 0: First-Run Setup &…, plus 11 more sections
  • Calls pip, python3 and npm; reaches mineru.net and api.openai.com; needs MINERU_API_KEY and OPENAI_API_KEY

What it does

Mineru Document Explorer is an agent skill from opendatalab/MinerU-Document-Explorer. MinerU Document Explorer — Agent-native knowledge engine. Use when users ask to search their documents, look up information in PDFs/DOCX/PPTX/Markdown, navigate inside large documents, extract tables/figures, or build wiki knowledge bases. Provides three tool groups: information retrieval (query, get, multiget, status), document deep reading (doctoc, docread, docgrep, docquery, docelements, doclinks), and knowledge ingestion (wikiingest, docwrite, wikilint, wikilog, wikiindex).

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/mcp-setup.md`). Compatibility notes: Requires qmd CLI or MCP server. Install via npm (npm install -g mineru-document-explorer) or from source…

It sits in Documents & Office, covering LLM wikis, Word documents and PowerPoint presentations. It works with Microsoft PowerPoint and Microsoft Word. The repository describes itself as: Agent-native knowledge engine with MCP tools for document indexing, wiki organization, fast retrieval and deep reading across PDF/DOCX/PPTX/Markdown. The licence is MIT.

When your agent uses it

  • Users ask to search their documents
  • Look up information in PDFs/DOCX/PPTX/Markdown
  • Navigate inside large documents
  • Extract tables/figures

Example prompts

  • “/mineru-document-explorer”

Requirements

  • Python 3
  • Node.js
  • A credential in MINERU_API_KEY
  • A credential in OPENAI_API_KEY
  • Compatibility (from SKILL.md): Requires qmd CLI or MCP server. Install via npm (npm install -g mineru-document-explorer) or from source (https://github.com/opendatalab/MinerU-Document-Explorer). PDF/DOCX/PPTX support requires Python 3.10+ with pymupdf, python-docx, python-pptx.
  • Pre-approved tools (allowed-tools): Bash(qmd:*), mcp__qmd__*, mcp__mineru-document-explorer__*

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Check qmd is installed
  2. Check Python for binary document support
  3. Check and install Python packages
  4. Ask about advanced PDF processing (optional)
  5. Index documents and verify
  6. Configure MCP server (for AI agent integration)

What it can do on your machine

Read from SKILL.md and the folder at commit a7e9c6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(qmd:*)
    • mcp__qmd__*
    • mcp__mineru-document-explorer__*

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python3
    • npm
    • bun
    • git
    • brew
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • mineru.net
    • api.openai.com

    Also links to:

    • python.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINERU_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires qmd CLI or MCP server. Install via npm (npm install -g mineru-document-explorer) or from source (https://github.com/opendatalab/MinerU-Document-Explorer). PDF/DOCX/PPTX support requires Python 3.10+ with pymupdf, python-docx, python-pptx.

    From compatibility in the SKILL.md frontmatter.

Context cost

Mineru Document Explorer loads about 6.2k tokens when it runs, and up to ~7.2k if it reads all its reference files. Until then it costs about 130 tokens; SKILL.md has 2,008 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:143
    - **Ubuntu/Debian**: `sudo apt install python3 python3-pip`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from opendatalab/MinerU-Document-Explorer at commit a7e9c6c, republished under its MIT licence (© opendatalab). 2,008 words, ~6,228 tokens.

Download SKILL.mdSave it as .claude/skills/mineru-document-explorer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
mineru-document-explorer
description
MinerU Document Explorer — Agent-native knowledge engine. Use when users ask to search their documents, look up information in PDFs/DOCX/PPTX/Markdown, navigate inside large documents, extract tables/figures, or build wiki knowledge bases. Provides three tool groups: information retrieval (query, get, multi_get, status), document deep reading (doc_toc, doc_read, doc_grep, doc_query, doc_elements, doc_links), and knowledge ingestion (wiki_ingest, doc_write, wiki_lint, wiki_log, wiki_index).
allowed-tools
Bash(qmd:*), mcp__qmd__*, mcp__mineru-document-explorer__*
compatibility
Requires qmd CLI or MCP server. Install via npm (npm install -g mineru-document-explorer) or from source (https://github.com/opendatalab/MinerU-Document-Explorer). PDF/DOCX/PPTX support requires Python 3.10+ with pymupdf, python-docx, python-pptx.
license
MIT
metadata.author
MinerU
metadata.version
3.3.0

MinerU Document Explorer

Agent-native knowledge engine — hybrid search and deep reading over Markdown, PDF, DOCX, PPTX. Designed for AI agents to organize knowledge and retrieve information autonomously.

Quick Reference

I want to...ToolExample
Search across all docsquery{ "query": "authentication flow" }
Get a specific fileget{ "file": "#abc123" } or { "file": "docs/readme.md" }
Get multiple filesmulti_get{ "pattern": "docs/*.md" }
See document structuredoc_toc{ "file": "paper.pdf" }
Read specific sectionsdoc_read{ "file": "paper.pdf", "addresses": ["page:3"] }
Find keyword in a docdoc_grep{ "file": "report.md", "pattern": "revenue" }
Semantic search in docdoc_query{ "file": "paper.pdf", "query": "methodology" }
Extract tables/figuresdoc_elements{ "file": "report.pdf", "element_types": ["table"] }
Write a wiki pagedoc_write{ "collection": "wiki", "path": "topic.md", "content": "..." }
Check wiki healthwiki_lint{}

Agent Principles

Follow these rules to use the tools effectively:

  1. Collection-relative paths only. All file paths are prefixed by collection name: mydocs/readme.md, papers/survey.pdf. Never use absolute filesystem paths like /Users/.../file.md. You can also use qmd://mydocs/readme.md.

  2. Navigate before reading large documents. For PDFs, DOCX, PPTX, or Markdown files >100 lines, always use doc_toc → doc_read instead of get. The get tool dumps the entire document — wasteful for large files.

  3. Addresses bridge navigation and reading. doc_toc, doc_grep, and doc_query return address strings like line:45-120. Pass these directly to doc_read. Never call doc_read without addresses from one of these.

  4. Use simple query first. Start with { "query": "your terms" }. Only switch to advanced searches mode when simple mode misses. The system auto-expands into BM25 + semantic + reranking.

  5. Always pass source when writing wiki pages. This enables provenance tracking and staleness detection via wiki_lint.

  6. Prefer MCP over CLI. The MCP server keeps models loaded in memory (~3GB). CLI reloads on every invocation (~5-15s overhead). If MCP is not available, qmd search (BM25 only) is instant and needs no model loading.

Key Concepts

Collections and File Paths

Documents live in collections — named groups with a filesystem path and glob mask. Collections have two types:

  • raw (default) — immutable source documents, read-only for agents
  • wiki — LLM-maintained pages, agents create/update via doc_write

File paths in all results are collection-relative: mydocs/readme.md, papers/survey.pdf. Use these exact paths when calling tools.

Document IDs (docid)

Every document has a short hash ID like #abc123 shown in search results. Use docids anywhere a file path is accepted: get("#abc123"), doc_toc("#abc123"). The # prefix is optional.

Addresses

Addresses identify locations within a document. They are the bridge between navigation tools (doc_toc, doc_grep, doc_query) and the reading tool (doc_read).

FormatMeaningUsed by
line:N or line:N-MLine or line rangeMarkdown
page:NPDF pagePDF
slide:NPPTX slidePPTX
section:NDOCX sectionDOCX
Three Tool Groups (15 tools)
GroupPurposeTools
RetrievalFind and fetch documentsquery, get, multi_get, status
Deep ReadingNavigate within a documentdoc_toc, doc_read, doc_grep, doc_query, doc_elements, doc_links
Knowledge IngestionBuild wiki knowledge basewiki_ingest, doc_write, wiki_lint, wiki_log, wiki_index

Playbook 0: First-Run Setup & Configuration

Use when a user first connects MinerU Document Explorer, gives you the project link, or when PDF/DOCX/PPTX operations fail. Walk the user through setup interactively — check each prerequisite and guide them step by step.

Step 1 — Check qmd is installed
bash
which qmd && qmd status

If not installed:

bash
# Option A: npm (recommended)
npm install -g mineru-document-explorer

# Option B: from source
git clone https://github.com/opendatalab/MinerU-Document-Explorer.git
cd MinerU-Document-Explorer && bun install && bun link
Step 2 — Check Python for binary document support

PDF, DOCX, and PPTX processing requires Python 3.10+:

bash
python3 --version

If Python is missing, guide the user to install it for their platform:

  • macOS: brew install python@3.12
  • Ubuntu/Debian: sudo apt install python3 python3-pip
  • Windows: Download from https://python.org
Step 3 — Check and install Python packages

Three packages are required for binary document processing:

bash
python3 -c "import pymupdf; import docx; import pptx; print('All dependencies OK')"

If any import fails, install the missing packages:

bash
pip install pymupdf python-docx python-pptx
PackageFormatWhat it does
pymupdfPDFText extraction, bookmarks, page-level reading
python-docxDOCXSection extraction, table extraction
python-pptxPPTXSlide text, table extraction
Step 4 — Ask about advanced PDF processing (optional)

Ask the user: "Do you need high-quality PDF extraction for scanned documents or complex layouts? MinerU Cloud provides significantly better results than basic PyMuPDF."

If yes, guide them to set up MinerU Cloud:

  1. Get an API key from https://mineru.net
  2. Configure it (pick one method):
bash
# Method A: Environment variable
export MINERU_API_KEY="your-key-here"

# Method B: Config file (~/.config/qmd/doc-reading.json)
mkdir -p ~/.config/qmd
cat > ~/.config/qmd/doc-reading.json << 'EOF'
{
  "docReading": {
    "providers": {
      "fullText": { "pdf": ["mineru_cloud", "pymupdf"] }
    },
    "credentials": {
      "mineru": { "api_key": "YOUR_API_KEY_HERE" }
    }
  }
}
EOF

When MINERU_API_KEY is set, MinerU Cloud is automatically used as the primary PDF provider with PyMuPDF as fallback — no config file needed.

Additional Python package for MinerU Cloud:

bash
pip install mineru-open-sdk
Step 5 — Index documents and verify
bash
# Index a folder (adjust path to user's documents)
qmd collection add ~/Documents --name mydocs --mask '**/*.{md,pdf,docx,pptx}'

# Verify indexing worked
qmd status

# Test search (instant, no model downloads)
qmd search "test"
Step 6 — Configure MCP server (for AI agent integration)

Ask the user which AI client they use and provide the matching config:

Claude Code (~/.claude/settings.json):

json
{ "mcpServers": { "qmd": { "command": "qmd", "args": ["mcp"] } } }

Cursor (.cursor/mcp.json) — HTTP mode recommended:

bash
qmd mcp --http --daemon   # start the server first
json
{ "mcpServers": { "qmd": { "url": "http://localhost:8181/mcp" } } }

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):

json
{ "mcpServers": { "qmd": { "command": "qmd", "args": ["mcp"] } } }
Configuration Reference

Config file locations (later overrides earlier):

  1. ~/.config/qmd/doc-reading.json — global settings
  2. ./qmd.config.json — project-level overrides
  3. Environment variables — highest priority

Full config example (~/.config/qmd/doc-reading.json):

json
{
  "docReading": {
    "providers": {
      "fullText": { "pdf": ["mineru_cloud", "pymupdf"] },
      "toc":      { "pdf": ["native_bookmarks"] },
      "elements": { "docx": ["python_docx_local"], "pptx": ["python_pptx_local"] }
    },
    "credentials": {
      "mineru": {
        "api_key": "your-mineru-api-key",
        "api_url": "https://mineru.net/api/v4"
      },
      "openai": {
        "api_key": "your-openai-api-key",
        "base_url": "https://api.openai.com/v1"
      }
    }
  }
}

Environment variables:

VariablePurpose
MINERU_API_KEYMinerU Cloud PDF (auto-enables mineru_cloud provider)
OPENAI_API_KEYGPT PageIndex (LLM-inferred TOC for PDFs)
OPENAI_BASE_URLCustom OpenAI-compatible endpoint

Provider options:

CapabilityProviderRequires
PDF full textpymupdf (default)pip install pymupdf
PDF full textmineru_cloudpip install mineru-open-sdk + API key
PDF full textmineru_localpip install mineru-vl-utils[transformers] + model
PDF TOCnative_bookmarks (default)pip install pymupdf
PDF TOCgpt_pageindexpip install tiktoken openai pyyaml + API key
DOCX tablespython_docx_local (default)pip install python-docx
PPTX tablespython_pptx_local (default)pip install python-pptx

Playbook 1: Search & Answer a Question

Use when the user asks a question and you need to find information.

Step 1 — Search:

json
query({ "query": "how does authentication work" })

Results include docid, file, score, snippet. Use these to decide what to read.

Step 2 — Read the top result:

If the document is short (snippet suggests it's a small file):

json
get({ "file": "#abc123" })

If the document is large or structured (PDF, long Markdown):

json
doc_toc({ "file": "#abc123" })
doc_read({ "file": "#abc123", "addresses": ["line:11-20", "line:31-49"] })

Step 3 — Synthesize and answer using the content you read.

Tips:

  • Add intent to disambiguate: { "query": "performance", "intent": "web page load times" }
  • Use collections to narrow scope: { "query": "...", "collections": ["papers"] }
  • The get response header includes Total lines: — if >100, switch to doc_toc + doc_read

Playbook 2: Deep-Read a Large Document

Use when the user points to a specific large document (PDF, DOCX, PPTX, or long Markdown) and wants to understand it.

Step 1 — Get the table of contents:

json
doc_toc({ "file": "papers/survey.pdf" })

Returns a nested tree of sections with addresses. This is your map.

Step 2 — Read relevant sections:

Pick addresses from the TOC and read them:

json
doc_read({ "file": "papers/survey.pdf", "addresses": ["line:11-20", "line:45-60"] })

Step 3 — Search within the document (if you need to find something specific):

For keywords:

json
doc_grep({ "file": "papers/survey.pdf", "pattern": "attention mechanism" })

For concepts (semantic, requires embeddings):

json
doc_query({ "file": "papers/survey.pdf", "query": "what evaluation metrics were used" })

Both return addresses — pass them to doc_read.

Step 4 — Extract structured elements (tables, figures):

json
doc_elements({ "file": "report.pdf", "element_types": ["table"], "query": "revenue" })

Example flow:

doc_toc("papers/survey.pdf")
  → sees section "3. Methodology" at line:45-80
doc_read("papers/survey.pdf", ["line:45-80"])
  → reads methodology section
doc_grep("papers/survey.pdf", "dataset")
  → finds mentions at line:62, line:78
doc_read("papers/survey.pdf", ["line:60-65", "line:76-80"])
  → reads specific paragraphs around dataset mentions

Playbook 3: Build Wiki from Sources

Use when the user wants to build a persistent knowledge base from their documents. Requires a wiki-type collection.

Step 1 — Ingest a source document:

json
wiki_ingest({ "source": "mydocs/distributed-systems.md", "wiki_collection": "mywiki" })

Returns: source content, TOC, related existing wiki pages, and suggestions for what pages to create. Incremental — skips unchanged sources unless force: true.

Step 2 — Deep-read key sections (for large sources):

json
doc_toc({ "file": "mydocs/distributed-systems.md" })
doc_read({ "file": "mydocs/distributed-systems.md", "addresses": ["line:11-30"] })

Step 3 — Write wiki pages:

json
doc_write({
  "collection": "mywiki",
  "path": "concepts/cap-theorem.md",
  "content": "# CAP Theorem\n\n**Source:** [[sources/distributed-systems]]\n\n## Overview\n\nThe CAP theorem states that...\n\n## Connections\n- Related to [[concepts/consistency-models]]\n- See also [[concepts/consensus-algorithms]]",
  "title": "CAP Theorem",
  "source": "mydocs/distributed-systems.md"
})

Use [[wikilinks]] to create cross-references. Always pass source for provenance tracking.

Step 4 — Health-check:

json
wiki_lint({ "collection": "mywiki", "stale_days": 30 })

Detects orphan pages, broken links, missing pages, stale content.

Wiki page template:

markdown
# Page Title

**Source:** [[sources/paper-name]]

## Key Points
- ...

## Connections
- Related to [[concepts/topic-a]]
- Extends [[concepts/topic-b]]

Playbook 4: Maintain Wiki Health

Use periodically to keep the wiki knowledge base healthy.

json
wiki_lint({ "collection": "mywiki" })

Act on the results:

  • Orphan pages → add [[wikilinks]] from related pages
  • Broken links → fix the link target or create the missing page
  • Missing pages → create them with doc_write
  • Stale pages → re-read the source with doc_read, update with doc_write

View activity history:

json
wiki_log({ "since": "2025-01-01", "limit": 20 })

Generate or update the wiki index:

json
wiki_index({ "collection": "mywiki", "write": true })

Tool Reference

Show full SKILL.md (825 more words)Show less
Retrieval Tools

query — Search the knowledge base (primary search tool)

ParamTypeDefaultDescription
querystring—Simple search (mutually exclusive with searches)
searchesarray—Advanced: `[{type: "lex"
intentstring—Disambiguation context (steers ranking, not searched)
collectionsstring[]allFilter to specific collections
limitnumber10Max results
minScorenumber0Min relevance 0-1

Simple mode auto-expands into BM25 + semantic + reranking. For advanced mode, first sub-query gets 2x weight.

Sub-query typeMethodBest for
lexBM25 keywordsExact terms, names, "quoted phrases", -negation
vecVector semanticNatural language questions
hydeHypothetical answerWrite 50-100 words resembling the answer

get — Retrieve a single document

ParamTypeDefaultDescription
filestring—Path, docid (#abc123), or path:line
fromLinenumber—Start line (1-indexed)
maxLinesnumber—Max lines to return
lineNumbersbooleanfalseAdd line numbers

Response header includes Total lines: — if >100, prefer doc_toc + doc_read. On "not found" errors, check "Did you mean?" suggestions.

multi_get — Batch retrieve

ParamTypeDefaultDescription
patternstring—Glob, comma-separated paths, or comma-separated globs
maxLinesnumber—Max lines per file
maxBytesnumber10240Skip files larger than this
lineNumbersbooleanfalseAdd line numbers

Pattern examples: journals/2025-05*.md, readme.md, config.md, docs/api*.md, docs/config*.md, #abc123, #def456.

status — Index health (no parameters)

Returns document counts, embedding status, collection list. When connected via MCP, this info is already in the system prompt.

Deep Reading Tools

doc_toc — Table of contents

ParamTypeDescription
filestringFile path or docid

Returns a nested tree of sections with address fields. Start here for any large document.

doc_read — Read at addresses

ParamTypeDefaultDescription
filestring—File path or docid
addressesstring[]—Addresses from doc_toc / doc_grep / doc_query
max_tokensnumber2000Max tokens per section

doc_grep — Keyword search within a document

ParamTypeDefaultDescription
filestring—File path or docid
patternstring—Regex or keyword (e.g. "revenue|profit")
flagsstring"gi"Regex flags

Returns matches with address fields for doc_read.

doc_query — Semantic search within a document

ParamTypeDefaultDescription
filestring—File path or docid
querystring—Natural language query
top_knumber5Max ranked chunks to return

Returns ranked chunks with address fields. Requires embeddings.

doc_elements — Extract tables, figures, equations

ParamTypeDescription
filestringFile path or docid
addressesstring[]Optional: restrict extraction scope
querystringOptional: filter by relevance
element_typesstring[]Filter: "table", "figure", "equation"

doc_links — Forward/backward link graph

ParamTypeDefaultDescription
filestring—File path or docid
directionstring"both""forward", "backward", or "both"
link_typestring"all""wikilink", "markdown", "url", or "all"
Knowledge Ingestion Tools

wiki_ingest — Prepare source for wiki processing

ParamTypeDefaultDescription
sourcestring—Source file path or docid
wiki_collectionstringautoTarget wiki collection
forcebooleanfalseForce re-ingest even if unchanged

Returns source content, TOC, related pages, suggestions. Large docs (>50k chars) are truncated — use doc_read for details.

doc_write — Write a document

ParamTypeDescription
collectionstringTarget collection name
pathstringRelative path (e.g. "concepts/topic.md")
contentstringFull markdown content
titlestringOptional: document title
sourcestringOptional: source path for provenance

Writes to disk and immediately re-indexes. Wiki collections auto-log.

wiki_lint — Health check

ParamTypeDefaultDescription
collectionstring—Optional: limit to collection
stale_daysnumber30Days threshold for staleness

wiki_log — Activity timeline

ParamTypeDefaultDescription
sincestring—ISO date filter (e.g. "2025-01-01")
operationstring—Filter: "ingest", "update", "lint", "query", "index"
limitnumber20Max entries
formatstring"markdown""markdown" or "json"

wiki_index — Generate index page

ParamTypeDefaultDescription
collectionstring—Wiki collection to index
writebooleanfalseWrite index.md to disk

Decision Tree

START
  │
  ├─ "What's indexed?" → status
  │
  ├─ "Find documents about X" → query
  │     Next: get (small docs) or doc_toc → doc_read (large docs)
  │
  ├─ "Get this specific file" → get (path or #docid)
  │     ⚠ For large docs: use doc_toc + doc_read instead
  │
  ├─ "Get several files" → multi_get (glob or comma-list)
  │
  ├─ "Read section of a large doc" → doc_toc → doc_read
  │
  ├─ "Find keyword in one doc" → doc_grep → doc_read
  │
  ├─ "Conceptual search in one doc" → doc_query → doc_read
  │
  ├─ "Extract tables/figures" → doc_elements
  │
  ├─ "What links to this page?" → doc_links
  │
  ├─ "Build wiki from source" → wiki_ingest → doc_read → doc_write
  │
  └─ "Check wiki health" → wiki_lint

Troubleshooting

ProblemCauseFix
"Document not found"Wrong path or missing collection prefixCheck "Did you mean?" suggestions in error; use status to see collections
"No results found"Query too specific or wrong collectionTry simpler keywords; omit collections to search all; check status
"No vector embeddings" warningEmbeddings not generatedTell the user to run qmd embed (one-time, downloads ~2GB models)
get returns too much textDocument is largeUse doc_toc → doc_read for targeted sections
doc_read returns emptyNo addresses provided or wrong formatGet addresses from doc_toc, doc_grep, or doc_query first
Slow first query (~5-15s)LLM models loadingNormal for MCP startup; subsequent queries are fast. CLI always reloads.
PDF/DOCX/PPTX not workingMissing Python dependenciesFollow Playbook 0 to check and install: python3 -c "import pymupdf; import docx; import pptx", then pip install pymupdf python-docx python-pptx
Wiki page has broken linksTarget page doesn't existCreate the missing page with doc_write, or fix the [[wikilink]]
Stale wiki pagesSource document updated after wiki page writtenRun wiki_lint to detect; re-read source with doc_read and update
multi_get returns no filesPattern doesn't match any indexed filesCheck exact collection names via status; try broader glob

CLI Reference (when MCP is not available)

bash
qmd status                                # Index health
qmd query "question"                      # Hybrid search (recommended)
qmd search "keywords"                     # BM25 only (fast, no LLM)
qmd get "#abc123"                         # Get by docid
qmd get "docs/readme.md:100" -l 50        # Line slice
qmd multi-get "journals/2026-*.md" -l 40  # Glob batch
qmd multi-get "a.md, b.md, c.md"          # Comma-separated
qmd doc-toc "paper.pdf"                   # Document TOC
qmd doc-read "paper.pdf" "line:45-120"    # Read section
qmd doc-grep "report.md" "revenue"        # Search in document
qmd mcp                                   # MCP server (stdio)
qmd mcp --http --daemon                   # MCP server (HTTP, background)

Setup

For first-time users, use Playbook 0 above — it walks through the full setup interactively, including dependency checks and configuration.

bash
# Install
npm install -g mineru-document-explorer

# Python dependencies for PDF/DOCX/PPTX (required for binary formats)
pip install pymupdf python-docx python-pptx

# Optional: MinerU Cloud for high-quality PDF (scanned docs, complex layouts)
pip install mineru-open-sdk
export MINERU_API_KEY="your-key"  # get from https://mineru.net

# Index documents
qmd collection add ~/notes --name notes
qmd collection add ~/papers --name papers --mask '**/*.{md,pdf,docx,pptx}'

# Verify
qmd status
qmd search "test query"  # instant, no model download

# Optional: enable semantic search (downloads ~2GB models on first run)
qmd embed

MCP Configuration

Claude Code (~/.claude/settings.json):

json
{ "mcpServers": { "qmd": { "command": "qmd", "args": ["mcp"] } } }

Cursor (.cursor/mcp.json) — stdio:

json
{ "mcpServers": { "qmd": { "command": "qmd", "args": ["mcp"] } } }

Cursor (.cursor/mcp.json) — HTTP (recommended):

json
{ "mcpServers": { "qmd": { "url": "http://localhost:8181/mcp" } } }

Start daemon first: qmd mcp --http --daemon

Skill Installation

bash
qmd skill install              # install to current project
qmd skill install --global     # install globally

© opendatalab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/mineru-document-explorer of opendatalab/MinerU-Document-Explorer.

  • SKILL.md
  • references/mcp-setup.md

Open the folder on GitHubat commit a7e9c6c

Compare with similar skills

Mineru Document Explorer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mineru Document Explorer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mineru Document Explorer this skillopendatalab/MinerU-Document-Explorer638—~6.2kAutomated safety check: NotesMIT
Exam IngestZeKaiNie/universal-examprep-skill303—~5.6kAutomated safety check: PassMIT
Knowledge Ingestevolution-foundation/evo-nexus545—~922Automated safety check: PassCustom licence
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
GenOffice Document CLIgenspark-ai/genoffice8.9k—~19kAutomated safety check: PassApache-2.0
Docsagentdocsagent/docsagent625—~834Automated safety check: PassNone

Similar skills

  • Exam Ingest

    ZeKaiNie/universal-examprep-skill

    从学生上传的课件/大纲/老师勾的重点/真题,一键初始化并验证备考工作区:解析 PDF、DOCX、PPTX、 XLSX、常见独立图片与 txt/md,建立分章节 LLM Wiki、标准题库、结构化接管队列与进度状态;仅在 Python 确实无法运行时 明确降级为手动写盘。当工作区尚未建立、资料发生变化、或建库 readiness 被阻断时使用。

    303 GitHub stars~5.6k tokensUpdated 10 days ago
    Documents & OfficeAuto-check passed
  • Knowledge Ingest

    evolution-foundation/evo-nexus

    Upload a file (PDF, DOCX, PPTX, XLSX, HTML, EPUB, image) or URL to the Knowledge base.

    545 GitHub stars~922 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Docsagent

    docsagent/docsagent

    Search and manage private, local document collections (PDF, PPTX, DOCX) offline.

    625 GitHub stars~834 tokensUpdated 10 days ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed

Questions about Mineru Document Explorer

What does Mineru Document Explorer do?

MinerU Document Explorer — Agent-native knowledge engine. An agent skill from opendatalab/MinerU-Document-Explorer. Mineru Document Explorer is an agent skill from opendatalab/MinerU-Document-Explorer. MinerU Document Explorer — Agent-native knowledge engine.

When should I use Mineru Document Explorer?

Mineru Document Explorer fits situations like: users ask to search their documents; look up information in PDFs/DOCX/PPTX/Markdown; navigate inside large documents; extract tables/figures.

How do I install Mineru Document Explorer in Claude Code?

Run `npx skills add opendatalab/MinerU-Document-Explorer --skill mineru-document-explorer -a claude-code`. Or copy the skill folder (skills/mineru-document-explorer in opendatalab/MinerU-Document-Explorer) into .claude/skills/mineru-document-explorer in your project. Claude Code loads it when a task matches its description.

How do I install Mineru Document Explorer in Codex?

Run `npx skills add opendatalab/MinerU-Document-Explorer --skill mineru-document-explorer -a codex`. Or copy the skill folder (skills/mineru-document-explorer in opendatalab/MinerU-Document-Explorer) into .agents/skills/mineru-document-explorer in your project. Codex loads it when a task matches its description.

Can I use Mineru Document Explorer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opendatalab/MinerU-Document-Explorer --skill mineru-document-explorer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mineru-document-explorer, .gemini/skills/mineru-document-explorer, .github/skills/mineru-document-explorer and .opencode/skills/mineru-document-explorer in your project.

What does Mineru Document Explorer need to run?

Going by SKILL.md and its folder, Mineru Document Explorer needs the command-line tools its instructions call (pip, python3, npm, bun, git and brew) and credentials named MINERU_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; Node.js; A credential in MINERU_API_KEY; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Bash(qmd:*), mcp__qmd__*, mcp__mineru-document-explorer__*. Compatibility (from SKILL.md): Requires qmd CLI or MCP server. Install via npm (npm install -g mineru-document-explorer) or from source (https://github.com/opendatalab/MinerU-Document-Explorer). PDF/DOCX/PPTX support requires Python 3.10+ with pymupdf, python-docx, python-pptx. .

Does Mineru Document Explorer access the network?

SKILL.md names 3 domains. In commands or code: mineru.net and api.openai.com; the agent is likely to contact these when it follows the instructions. As links in the text: python.org. This is read from the text; nothing was executed.

Is Mineru Document Explorer safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Mineru Document Explorer use?

Mineru Document Explorer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mineru Document Explorer use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 954 tokens, read only when the agent opens those files.

What are the alternatives to Mineru Document Explorer?

Skills that share tags, products or a category with Mineru Document Explorer: Exam Ingest (ZeKaiNie/universal-examprep-skill, 303 stars), Knowledge Ingest (evolution-foundation/evo-nexus, 545 stars), Markitdown (ImCa0/just-laws, 781 stars) and GenOffice Document CLI (genspark-ai/genoffice, 8.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mineru Document Explorer?

opendatalab (a GitHub organization) maintains it in opendatalab/MinerU-Document-Explorer, which has 638 GitHub stars. The repository was last updated on April 26, 2026.

Source: opendatalab/MinerU-Document-Explorer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.