Paper Research on arXiv
XiaomiMiMo/MiMo-Code
Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.
Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature-pdf-ocr-library .claude/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .claude/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-libraryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/literature-pdf-ocr-library .agents/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .agents/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/literature-pdf-ocr-library .cursor/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .cursor/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LigphiDonk/Oh-my--paper.git --path skills/literature-pdf-ocr-library--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/literature-pdf-ocr-library .gemini/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .gemini/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-libraryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/literature-pdf-ocr-library .github/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .github/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/literature-pdf-ocr-library .opencode/skills/literature-pdf-ocr-library && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "literature-pdf-ocr-library" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/literature-pdf-ocr-library into .opencode/skills/literature-pdf-ocr-library/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "literature-pdf-ocr-library", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
literature-pdf-ocr-librarySearches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.
This skill builds a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages: it narrows the topic, searches official or stable sources such as arXiv or Hugging Face, downloads only legally accessible PDFs, and runs OCR or layout parsing before writing a clean Markdown library with machine-readable metadata.
Each corpus gets its own named folder - under a fixed pipeline path in Oh My Paper projects, or under a literature folder in standalone projects - never a flat directory. A bundled script downloads papers by arXiv ID or search query, another converts PDFs or page images to Markdown through a PaddleOCR layout-parsing API with a local pdfminer fallback, a third builds a JSON and JSONL index of the library, and a fourth runs the full search-to-ingest pipeline in one pass.
OCR output for each paper is written inside that paper's own folder rather than a shared top-level OCR directory, and every paper's OCR path gets recorded in a literature bank file so later work can read the actual extracted content rather than guessing from the PDF filename.
Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
PADDLEOCR_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Literature PDF OCR Library Builder loads about 1.1k tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 186 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 186 words, ~1,114 tokens.
.claude/skills/literature-pdf-ocr-library/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Use this skill to build a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages. The default workflow is: narrow the topic, search official or stable APIs, download only legally accessible PDFs, run OCR or layout parsing, then emit a clean Markdown library with machine-readable metadata.
In Oh My Paper projects, the corpus always lives under .pipeline/literature/<corpus-name>/.
In standalone projects, use research/literature/<corpus-name>/.
Never dump papers into the root or a flat directory without a corpus name.
.pipeline/
literature/
<corpus-name>/ ← one folder per topic/session, e.g. "humanoid-locomotion"
search_results.json ← raw search/ID-lookup results
library_index.json ← consolidated index for the whole corpus
library_index.jsonl
papers/
<arxiv-id>-<title-slug>/ ← one folder per paper
metadata.json
paper.pdf
ocr/ ← OCR output lives here, next to the PDF
paper/
doc_0.md ← main OCR markdown (PaddleOCR: multiple pages)
manifest.json
doc_0.md ← pdfminer fallback: single flat fileRules:
--out-dir always points to .pipeline/literature/<corpus-name>/ — never to .pipeline/literature/ directly.papers/<slug>/ocr/), not in a top-level ocr/ directory.ocr/ path in literature_bank.md so agents can read the actual content.# Download by arXiv IDs (recommended when IDs are known from web search)
python .claude/skills/literature-pdf-ocr-library/scripts/search_and_download_papers.py \
--arxiv-ids 2502.13817 2501.14459 \
--out-dir .pipeline/literature/my-corpus \
--download-pdfs
# Download by query
python .claude/skills/literature-pdf-ocr-library/scripts/search_and_download_papers.py \
--query "humanoid locomotion reinforcement learning" \
--out-dir .pipeline/literature/my-corpus \
--limit 20 --sources arxiv semanticscholar openalex hf_daily \
--download-pdfs
# OCR: PaddleOCR API (best quality)
export PADDLEOCR_TOKEN="<token>" # ask user, never hardcode
python .claude/skills/literature-pdf-ocr-library/scripts/paddleocr_layout_to_markdown.py \
.pipeline/literature/my-corpus/papers/*/paper.pdf \
--output-dir .pipeline/literature/my-corpus/papers \
--skip-existing
# OCR: pdfminer fallback (text-only, no layout — confirm with user first)
python .claude/skills/literature-pdf-ocr-library/scripts/paddleocr_layout_to_markdown.py \
.pipeline/literature/my-corpus/papers/*/paper.pdf \
--output-dir .pipeline/literature/my-corpus/papers \
--fallback-pdfminer
# Build index
python .claude/skills/literature-pdf-ocr-library/scripts/build_library_index.py \
--library-root .pipeline/literature/my-corpusscripts/search_and_download_papers.py for traceable search and PDF download (supports --query and --arxiv-ids).scripts/paddleocr_layout_to_markdown.py for single-file or batch OCR conversion (supports --fallback-pdfminer).scripts/build_library_index.py to generate library_index.json and library_index.jsonl.scripts/ingest_literature_library.py when the user wants the full ingestion workflow in one go.© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in skills/literature-pdf-ocr-library of LigphiDonk/Oh-my--paper.
Open the folder on GitHubat commit 6baece9
Literature PDF OCR Library Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Literature PDF OCR Library Builder this skillLigphiDonk/Oh-my--paper | 739 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Paper Research on arXivXiaomiMiMo/MiMo-Code | 14k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Arxiv MCP Serverblazickjp/arxiv-mcp-server | 3.2k | — | ~353 | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Paper Publisherhuggingface/skills | 11k | 4 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Daily arXiv Paper Briefjuliye2025/evil-read-arxiv | 1.7k | — | ~4.3k | Automated safety check: Pass | None | |
| Arxiv Paper Writeryunshenwuchuxun/latex-paper-skills | 267 | — | ~3.2k | Automated safety check: Pass | MIT |
XiaomiMiMo/MiMo-Code
Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.
blazickjp/arxiv-mcp-server
A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…
huggingface/skills
Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.
juliye2025/evil-read-arxiv
Searches arXiv for recent papers matching your research interests, scores them and writes a daily recommendation note into an Obsidian vault.
yunshenwuchuxun/latex-paper-skills
Writes ML/AI review and survey papers for arXiv using the IEEEtran LaTeX template with verified BibTeX citations.
juliye2025/evil-read-arxiv
Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them.
LigphiDonk/Oh-my--paper
Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.
LigphiDonk/Oh-my--paper
Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.
LigphiDonk/Oh-my--paper
Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.
LigphiDonk/Oh-my--paper
Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.
LigphiDonk/Oh-my--paper
Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.
LigphiDonk/Oh-my--paper
Create academic presentation slide decks and optionally demo videos from research papers.
Works with
Categories
Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library. This skill builds a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages: it narrows the topic, searches official or stable sources such as arXiv or Hugging Face, downloads only legally accessible PDFs, and runs OCR or layout parsing before writing a clean Markdown library with machine-readable metadata.
Literature PDF OCR Library Builder fits situations like: building a traceable corpus of papers on a research topic; batch-converting a folder of PDFs into Markdown for a knowledge base; fetching arXiv paper leads by ID for a literature survey.
Run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a claude-code`. Or copy the skill folder (skills/literature-pdf-ocr-library in LigphiDonk/Oh-my--paper) into .claude/skills/literature-pdf-ocr-library in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a codex`. Or copy the skill folder (skills/literature-pdf-ocr-library in LigphiDonk/Oh-my--paper) into .agents/skills/literature-pdf-ocr-library in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/literature-pdf-ocr-library, .gemini/skills/literature-pdf-ocr-library, .github/skills/literature-pdf-ocr-library and .opencode/skills/literature-pdf-ocr-library in your project.
Going by SKILL.md and its folder, Literature PDF OCR Library Builder needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named PADDLEOCR_TOKEN. Our summary lists: PaddleOCR layout-parsing API (or local pdfminer fallback); Python.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Literature PDF OCR Library Builder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 360 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Literature PDF OCR Library Builder: Paper Research on arXiv (XiaomiMiMo/MiMo-Code, 14k stars), Arxiv MCP Server (blazickjp/arxiv-mcp-server, 3.2k stars), Hugging Face Paper Publisher (huggingface/skills, 11k stars) and Daily arXiv Paper Brief (juliye2025/evil-read-arxiv, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 739 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.
Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.