Mineru
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.
Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI.
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install oidlabs-com/Lexoid lexoid-cli --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/lexoid-cli .claude/skills/lexoid-cli && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .claude/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cliType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install oidlabs-com/Lexoid lexoid-cli --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/lexoid-cli .agents/skills/lexoid-cli && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .agents/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install oidlabs-com/Lexoid lexoid-cli --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/lexoid-cli .cursor/skills/lexoid-cli && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .cursor/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/oidlabs-com/Lexoid.git --path skills/lexoid-cli--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install oidlabs-com/Lexoid lexoid-cli --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/lexoid-cli .gemini/skills/lexoid-cli && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .gemini/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install oidlabs-com/Lexoid lexoid-cliInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/lexoid-cli .github/skills/lexoid-cli && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .github/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install oidlabs-com/Lexoid lexoid-cli --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oidlabs-com/Lexoid.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/lexoid-cli .opencode/skills/lexoid-cli && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "lexoid-cli" agent skill from https://github.com/oidlabs-com/Lexoid/tree/main/skills/lexoid-cli into .opencode/skills/lexoid-cli/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lexoid-cli", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
lexoid-cliParse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI.
Lexoid CLI is an agent skill from oidlabs-com/Lexoid. Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI. Use when the user wants to extract markdown / JSON / LaTeX from a file or URL without writing Python, run schema-based structured extraction, or batch-parse from shell scripts. Triggers include "parse this PDF", "convert document to markdown", "extract JSON from PDF", "convert PDF to LaTeX", "use lexoid CLI", or any pipe/shell-style document-processing request.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering Document parsing, LaTeX and PDF. It works with LaTeX, Python, Microsoft Word and Microsoft Excel. The repository describes itself as: The open-source universal adapter for LLMs. Turn messy real-world data into clean, agent-ready context. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b45d174. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ollamapythonjqpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GOOGLE_API_KEYOPENAI_API_KEYANTHROPIC_API_KEYMISTRAL_API_KEYHUGGINGFACEHUB_API_TOKENTOGETHER_API_KEYOPENROUTER_API_KEYFIREWORKS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Lexoid CLI loads about 2k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 769 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
### Loading API keys from `.env`API keys are commonly stored in a `.env` file at the project root rather than exported in the shell. The `lexoid` CLI do**Always load `.env` in a subshell so the keys never leak into the surrounding environment** — the parentheses scope the# Load .env for this command only (if it exists); env is reset to its prior state on exit( set -a; [ -f .env ] && . ./.env; set +a; lexoid parse --input document.pdf --parser-type LLM_PARSE --model gpt-4o )Do **not** run a bare `set -a; . ./.env; set +a` in the parent shell — that persists secrets into the session.regardless of routing). LLM-based: load .env first.`schema` commands are LLM-based — load `.env` first (see "Loading API keys from .env") unless keys are already exportedLaTeX conversion is LLM-based — load `.env` first (see "Loading API keys from .env") unless keys are already exported.the env var. The CLI does not auto-load `.env`; load it via the subshell wrapper (see "Loading API keys from .env") andAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from oidlabs-com/Lexoid at commit b45d174, republished under its Apache-2.0 licence (© oidlabs-com). 769 words, ~1,969 tokens.
.claude/skills/lexoid-cli/SKILL.md (or your agent's skills folder).Lexoid ships a lexoid console script (also runnable as python -m lexoid) for document parsing without writing Python. There are three sub-commands: parse, schema, and latex.
If the user wants to embed parsing into a Python application or library, use the lexoid-python skill instead.
Before invoking, confirm:
lexoid is installed (lexoid --help or python -m lexoid --help). If not, run pip install lexoid.GOOGLE_API_KEY (Gemini, default), OPENAI_API_KEY, ANTHROPIC_API_KEY, MISTRAL_API_KEY, HUGGINGFACEHUB_API_TOKEN, TOGETHER_API_KEY, OPENROUTER_API_KEY, FIREWORKS_API_KEY.ollama serve running at OLLAMA_BASE_URL (default http://localhost:11434) and the target model pulled (ollama pull <model>).lowriter) must be installed..envAPI keys are commonly stored in a .env file at the project root rather than exported in the shell. The lexoid CLI does not auto-load .env, so for any LLM-based command (LLM_PARSE, schema, latex, or AUTO when it routes to an LLM) you must load it yourself.
Always load .env in a subshell so the keys never leak into the surrounding environment — the parentheses scope the exports to that one command, restoring the environment to its previous state automatically afterward. Guard the source with [ -f .env ] so the command still runs (using already-exported keys) when no .env is present:
# Load .env for this command only (if it exists); env is reset to its prior state on exit
( set -a; [ -f .env ] && . ./.env; set +a; lexoid parse --input document.pdf --parser-type LLM_PARSE --model gpt-4o )Do not run a bare set -a; . ./.env; set +a in the parent shell — that persists secrets into the session.
The examples in the rest of this skill are written as plain lexoid … for readability. Apply the wrapper above to any LLM-based command (LLM_PARSE, schema, latex, or AUTO when it routes to an LLM). If your keys are already exported in the environment, run the commands as-is.
lexoid parse — Markdown / JSON outputConvert a document to markdown (default) or JSON (with segments, token usage, parser info).
# Default: AUTO routing, markdown to stdout
lexoid parse --input document.pdf
# Save to file
lexoid parse --input document.pdf --output output.md
# Full result as JSON (includes per-page segments, token usage, parsers_used)
lexoid parse --input document.pdf --format json --output result.json
# Explicit LLM parsing (forces an LLM regardless of routing). LLM-based: load .env first.
lexoid parse --input document.pdf --parser-type LLM_PARSE --model gpt-4o
lexoid parse --input document.pdf --parser-type LLM_PARSE --model claude-3-5-sonnet-20241022
lexoid parse --input scanned.pdf --parser-type LLM_PARSE --model mistral-ocr-latest
# Force static parsing (no LLM, no API key needed for PDFs)
lexoid parse --input document.pdf --parser-type STATIC_PARSE --framework pdfplumber
# Parse a URL
lexoid parse --input https://example.com --output page.md
# Tune chunking / parallelism
lexoid parse --input big.pdf --pages-per-split 8 --max-processes 8Key flags: --parser-type (AUTO/LLM_PARSE/STATIC_PARSE), --model, --framework (pdfplumber/paddleocr), --api (override provider), --format (markdown/json), --pages-per-split, --max-processes, --verbose.
lexoid schema — Structured extractionExtract data conforming to a JSON schema. Schema can be a file path or inline JSON.
All schema commands are LLM-based — load .env first (see "Loading API keys from .env") unless keys are already exported.
# Inline schema
lexoid schema \
--input invoice.pdf \
--schema '{"type":"object","properties":{"invoice_number":{"type":"string"},"total":{"type":"number"}}}' \
--output invoice.json
# Schema from file, explicit provider
lexoid schema --input invoice.pdf --schema schema.json --api openai --model gpt-4o
# Example-guided extraction (improves accuracy)
lexoid schema --input invoice.pdf --schema schema.json \
--example-schema example.json
# Treat the whole doc as one instance (vs. one per page)
lexoid schema --input contract.pdf --schema schema.json --fill-single-schemaDefaults: model gpt-4o-mini. The provider is auto-detected from the model name unless --api is given.
lexoid latex — LaTeX conversionConvert a document to a self-contained LaTeX source.
LaTeX conversion is LLM-based — load .env first (see "Loading API keys from .env") unless keys are already exported.
lexoid latex --input paper.pdf --output paper.tex
lexoid latex --input paper.pdf --model gpt-4oWhen no --output is given, only the parsed content is written to stdout; status messages, token usage, and parser info go to stderr. This means standard piping works:
lexoid parse --input report.pdf | grep -i "revenue"
lexoid parse --input report.pdf --format json | jq '.token_usage'AUTO: with no --parser-type, the CLI uses AUTO, which inspects the document and routes to the best parser (often an LLM for scans, charts, or complex tables). This is the right choice unless the user asks otherwise.STATIC_PARSE is opt-in, not the default: choose it when the user explicitly wants no API calls / no cost, or you know the input is a clean native-text PDF. It returns empty output on scanned/image-only pages, so it is not a safe first guess for unknown documents.LLM_PARSE for quality: force it with --parser-type LLM_PARSE --model <model> for scans, chart/figure-heavy pages, or messy tables where layout fidelity matters.--parser-type STATIC_PARSE --framework paddleocr (no API key) or an LLM model with vision.for f in inputs/*.pdf; do lexoid parse -i "$f" -o "out/${f%.pdf}.md"; done).--verbose to surface loguru logs to stderr..env; load it via the subshell wrapper (see "Loading API keys from .env") and retry.libreoffice.--model and --api mismatch → use --api only to override an auto-inferred provider (e.g., to send a model through OpenRouter).ollama serve and ollama pull <model> first; the CLI does not start the server.docs/cli.rst and docs/api.rst in this repo.lexoid-python skill.© oidlabs-com, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/lexoid-cli of oidlabs-com/Lexoid.
Open the folder on GitHubat commit b45d174
Lexoid CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Lexoid CLI this skilloidlabs-com/Lexoid | 109 | — | ~2k | Automated safety check: Notes | Apache-2.0 | |
| MineruNebutra/MinerU-Skill | 122 | — | ~1.4k | Automated safety check: Pass | MIT | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| Doclingzhuzhaoyun/Molio | 431 | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Document ConverterBlackBeltTechnology/pi-agent-dashboard | 316 | — | ~999 | Automated safety check: Pass | MIT |
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
zhuzhaoyun/Molio
PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown.
BlackBeltTechnology/pi-agent-dashboard
Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…
bowenliang123/markdown-exporter
Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.
oidlabs-com/Lexoid
Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library.
Categories
Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI. Lexoid CLI is an agent skill from oidlabs-com/Lexoid. Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI.
Lexoid CLI fits situations like: the user wants to extract markdown / JSON / LaTeX from a file; URL without writing Python; run schema-based structured extraction; batch-parse from shell scripts.
Run `npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a claude-code`. Or copy the skill folder (skills/lexoid-cli in oidlabs-com/Lexoid) into .claude/skills/lexoid-cli in your project. Claude Code loads it when a task matches its description.
Run `npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a codex`. Or copy the skill folder (skills/lexoid-cli in oidlabs-com/Lexoid) into .agents/skills/lexoid-cli in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oidlabs-com/Lexoid --skill lexoid-cli -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lexoid-cli, .gemini/skills/lexoid-cli, .github/skills/lexoid-cli and .opencode/skills/lexoid-cli in your project.
Going by SKILL.md and its folder, Lexoid CLI needs the command-line tools its instructions call (ollama, python, jq and pip) and credentials named GOOGLE_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY and MISTRAL_API_KEY. Our summary lists: Python 3; A credential in GOOGLE_API_KEY; A credential in OPENAI_API_KEY.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Lexoid CLI is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Lexoid CLI: Mineru (Nebutra/MinerU-Skill, 122 stars), Mineru (Nebutra/MinerU-Skill, 122 stars), Markdown Converter (Team-Commonly/commonly, 1.4k stars) and Docling (zhuzhaoyun/Molio, 431 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
oidlabs-com (a GitHub organization) maintains it in oidlabs-com/Lexoid, which has 109 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.
Source: oidlabs-com/Lexoid on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.