Agent skill

Literature PDF OCR Library Builder

by LigphiDonk in LigphiDonk/Oh-my--paper

Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

MITAuto-check passedResearch & Science

Install Literature PDF OCR Library Builder

skills CLI
$ npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LigphiDonk/Oh-my--paper literature-pdf-ocr-library --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature-pdf-ocr-library .claude/skills/literature-pdf-ocr-library && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
literature-pdf-ocr-library
GitHub stars
739
Token cost
~1.1k tokens
SKILL.md length
186 words
Files
9 (incl. scripts, references)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

  • Building a traceable corpus of papers on a research topic
  • SKILL.md covers Overview, Canonical Directory Layout, Commands and Resources
  • Runs Python scripts from its folder; calls python; needs PADDLEOCR_TOKEN
  • Batch-converting a folder of PDFs into Markdown for a knowledge base

What it does

This skill builds a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages: it narrows the topic, searches official or stable sources such as arXiv or Hugging Face, downloads only legally accessible PDFs, and runs OCR or layout parsing before writing a clean Markdown library with machine-readable metadata.

Each corpus gets its own named folder - under a fixed pipeline path in Oh My Paper projects, or under a literature folder in standalone projects - never a flat directory. A bundled script downloads papers by arXiv ID or search query, another converts PDFs or page images to Markdown through a PaddleOCR layout-parsing API with a local pdfminer fallback, a third builds a JSON and JSONL index of the library, and a fourth runs the full search-to-ingest pipeline in one pass.

OCR output for each paper is written inside that paper's own folder rather than a shared top-level OCR directory, and every paper's OCR path gets recorded in a literature bank file so later work can read the actual extracted content rather than guessing from the PDF filename.

When your agent uses it

  • Building a traceable corpus of papers on a research topic
  • Batch-converting a folder of PDFs into Markdown for a knowledge base
  • Fetching arXiv paper leads by ID for a literature survey

Example prompts

  • “Build a literature corpus on humanoid locomotion from arXiv.”
  • “OCR these downloaded PDFs into Markdown with layout parsing.”
  • “Fetch these three arXiv IDs and add them to the paper library.”

Requirements

  • PaddleOCR layout-parsing API (or local pdfminer fallback)
  • Python

What it can do on your machine

Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • PADDLEOCR_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Literature PDF OCR Library Builder loads about 1.1k tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 186 words, ~1,114 tokens.

Download SKILL.mdSave it as .claude/skills/literature-pdf-ocr-library/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
literature-pdf-ocr-library
description
Search traceable academic papers, download legally accessible PDFs from arXiv and open-access sources, convert PDFs or page images to Markdown with a PaddleOCR layout-parsing API (or local pdfminer fallback), and organize the results into an AI-readable literature library. Use when Claude Code needs to build a paper corpus, batch OCR PDFs to Markdown, ingest real literature into a knowledge base, fetch arXiv or Hugging Face paper leads, or turn a directory of papers into structured Markdown plus metadata.
triggers
/literature-library, /paper-library, build paper corpus, build literature library, ingest papers, batch ocr papers, download arxiv papers, search and download…

Literature PDF OCR Library

Overview

Use this skill to build a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages. The default workflow is: narrow the topic, search official or stable APIs, download only legally accessible PDFs, run OCR or layout parsing, then emit a clean Markdown library with machine-readable metadata.


Canonical Directory Layout

In Oh My Paper projects, the corpus always lives under .pipeline/literature/<corpus-name>/.
In standalone projects, use research/literature/<corpus-name>/.
Never dump papers into the root or a flat directory without a corpus name.

text
.pipeline/
  literature/
    <corpus-name>/              ← one folder per topic/session, e.g. "humanoid-locomotion"
      search_results.json       ← raw search/ID-lookup results
      library_index.json        ← consolidated index for the whole corpus
      library_index.jsonl
      papers/
        <arxiv-id>-<title-slug>/   ← one folder per paper
          metadata.json
          paper.pdf
          ocr/                  ← OCR output lives here, next to the PDF
            paper/
              doc_0.md          ← main OCR markdown (PaddleOCR: multiple pages)
              manifest.json
            doc_0.md            ← pdfminer fallback: single flat file

Rules:

  • --out-dir always points to .pipeline/literature/<corpus-name>/ — never to .pipeline/literature/ directly.
  • OCR output lives inside the paper's own folder (papers/<slug>/ocr/), not in a top-level ocr/ directory.
  • After OCR, record each paper's ocr/ path in literature_bank.md so agents can read the actual content.

Commands

bash
# Download by arXiv IDs (recommended when IDs are known from web search)
python .claude/skills/literature-pdf-ocr-library/scripts/search_and_download_papers.py \
  --arxiv-ids 2502.13817 2501.14459 \
  --out-dir .pipeline/literature/my-corpus \
  --download-pdfs

# Download by query
python .claude/skills/literature-pdf-ocr-library/scripts/search_and_download_papers.py \
  --query "humanoid locomotion reinforcement learning" \
  --out-dir .pipeline/literature/my-corpus \
  --limit 20 --sources arxiv semanticscholar openalex hf_daily \
  --download-pdfs

# OCR: PaddleOCR API (best quality)
export PADDLEOCR_TOKEN="<token>"  # ask user, never hardcode
python .claude/skills/literature-pdf-ocr-library/scripts/paddleocr_layout_to_markdown.py \
  .pipeline/literature/my-corpus/papers/*/paper.pdf \
  --output-dir .pipeline/literature/my-corpus/papers \
  --skip-existing

# OCR: pdfminer fallback (text-only, no layout — confirm with user first)
python .claude/skills/literature-pdf-ocr-library/scripts/paddleocr_layout_to_markdown.py \
  .pipeline/literature/my-corpus/papers/*/paper.pdf \
  --output-dir .pipeline/literature/my-corpus/papers \
  --fallback-pdfminer

# Build index
python .claude/skills/literature-pdf-ocr-library/scripts/build_library_index.py \
  --library-root .pipeline/literature/my-corpus

Resources

  • Read source-strategy.md when you need source-specific behavior, file layout conventions, or legal constraints.
  • Use scripts/search_and_download_papers.py for traceable search and PDF download (supports --query and --arxiv-ids).
  • Use scripts/paddleocr_layout_to_markdown.py for single-file or batch OCR conversion (supports --fallback-pdfminer).
  • Use scripts/build_library_index.py to generate library_index.json and library_index.jsonl.
  • Use scripts/ingest_literature_library.py when the user wants the full ingestion workflow in one go.

© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/literature-pdf-ocr-library of LigphiDonk/Oh-my--paper.

  • SKILL.md
  • agents/openai.yaml
  • references/source-strategy.md
  • scripts/__pycache__/literature_lib.cpython-313.pyc
  • scripts/build_library_index.py
  • scripts/ingest_literature_library.py
  • scripts/literature_lib.py
  • scripts/paddleocr_layout_to_markdown.py
  • scripts/search_and_download_papers.py

Open the folder on GitHubat commit 6baece9

Compare with similar skills

Literature PDF OCR Library Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Literature PDF OCR Library Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Literature PDF OCR Library Builder this skillLigphiDonk/Oh-my--paper739—~1.1kAutomated safety check: PassMIT
Paper Research on arXivXiaomiMiMo/MiMo-Code14k—~1.5kAutomated safety check: PassMIT
Arxiv MCP Serverblazickjp/arxiv-mcp-server3.2k—~353Automated safety check: PassApache-2.0
Hugging Face Paper Publisherhuggingface/skills11k4 repos~4.2kAutomated safety check: PassApache-2.0
Daily arXiv Paper Briefjuliye2025/evil-read-arxiv1.7k—~4.3kAutomated safety check: PassNone
Arxiv Paper Writeryunshenwuchuxun/latex-paper-skills267—~3.2kAutomated safety check: PassMIT

Similar skills

  • Paper Research on arXiv

    XiaomiMiMo/MiMo-Code

    Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.

    14k GitHub stars~1.5k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Arxiv MCP Server

    blazickjp/arxiv-mcp-server

    A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…

    3.2k GitHub stars~353 tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Official

    Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.

    11k GitHub starsUsed in 4 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • Daily arXiv Paper Brief

    juliye2025/evil-read-arxiv

    Searches arXiv for recent papers matching your research interests, scores them and writes a daily recommendation note into an Obsidian vault.

    1.7k GitHub stars~4.3k tokensUpdated 26 days ago
    Research & ScienceAuto-check passed
  • Arxiv Paper Writer

    yunshenwuchuxun/latex-paper-skills

    Writes ML/AI review and survey papers for arXiv using the IEEEtran LaTeX template with verified BibTeX citations.

    267 GitHub stars~3.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Paper Figure Extractor

    juliye2025/evil-read-arxiv

    Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them.

    1.7k GitHub stars~298 tokensUpdated 26 days ago
    Research & ScienceAuto-check passed

More from LigphiDonk/Oh-my--paper

All 27 skills in this repo
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    739 GitHub starsUsed in 12 repos~3.7k tokens
    Auto-check passed
  • Inno Code Survey

    LigphiDonk/Oh-my--paper

    Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.

    739 GitHub stars~3.6k tokensUpdated 5 mo ago
    Auto-check passed
  • Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

    739 GitHub stars~3k tokensUpdated 5 mo ago
    Auto-check passed
  • Citation Verification Guide

    LigphiDonk/Oh-my--paper

    Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.

    739 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    739 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Making Academic Presentations

    LigphiDonk/Oh-my--paper

    Create academic presentation slide decks and optionally demo videos from research papers.

    739 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check passed

Questions about Literature PDF OCR Library Builder

What does Literature PDF OCR Library Builder do?

Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library. This skill builds a real, traceable literature corpus instead of fabricating references or scraping arbitrary publisher pages: it narrows the topic, searches official or stable sources such as arXiv or Hugging Face, downloads only legally accessible PDFs, and runs OCR or layout parsing before writing a clean Markdown library with machine-readable metadata.

When should I use Literature PDF OCR Library Builder?

Literature PDF OCR Library Builder fits situations like: building a traceable corpus of papers on a research topic; batch-converting a folder of PDFs into Markdown for a knowledge base; fetching arXiv paper leads by ID for a literature survey.

How do I install Literature PDF OCR Library Builder in Claude Code?

Run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a claude-code`. Or copy the skill folder (skills/literature-pdf-ocr-library in LigphiDonk/Oh-my--paper) into .claude/skills/literature-pdf-ocr-library in your project. Claude Code loads it when a task matches its description.

How do I install Literature PDF OCR Library Builder in Codex?

Run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a codex`. Or copy the skill folder (skills/literature-pdf-ocr-library in LigphiDonk/Oh-my--paper) into .agents/skills/literature-pdf-ocr-library in your project. Codex loads it when a task matches its description.

Can I use Literature PDF OCR Library Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill literature-pdf-ocr-library -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/literature-pdf-ocr-library, .gemini/skills/literature-pdf-ocr-library, .github/skills/literature-pdf-ocr-library and .opencode/skills/literature-pdf-ocr-library in your project.

What does Literature PDF OCR Library Builder need to run?

Going by SKILL.md and its folder, Literature PDF OCR Library Builder needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named PADDLEOCR_TOKEN. Our summary lists: PaddleOCR layout-parsing API (or local pdfminer fallback); Python.

Does Literature PDF OCR Library Builder access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Literature PDF OCR Library Builder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Literature PDF OCR Library Builder use?

Literature PDF OCR Library Builder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Literature PDF OCR Library Builder use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 360 tokens, read only when the agent opens those files.

What are the alternatives to Literature PDF OCR Library Builder?

Skills that share tags, products or a category with Literature PDF OCR Library Builder: Paper Research on arXiv (XiaomiMiMo/MiMo-Code, 14k stars), Arxiv MCP Server (blazickjp/arxiv-mcp-server, 3.2k stars), Hugging Face Paper Publisher (huggingface/skills, 11k stars) and Daily arXiv Paper Brief (juliye2025/evil-read-arxiv, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Literature PDF OCR Library Builder?

LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 739 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.

Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.