Literature PDF OCR Library Builder
LigphiDonk/Oh-my--paper
Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.
Mine open access full-text repositories for research data extraction
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .claude/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .claude/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .agents/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .agents/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .cursor/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .cursor/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wentorai/research-plugins.git --path skills/literature/fulltext/open-access-mining-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .gemini/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .gemini/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wentorai/research-plugins open-access-mining-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .github/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .github/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .opencode/skills/open-access-mining-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "open-access-mining-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/literature/fulltext/open-access-mining-guide into .opencode/skills/open-access-mining-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "open-access-mining-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
open-access-mining-guideMine open access full-text repositories for research data extraction
Open Access Mining Guide is an agent skill from wentorai/research-plugins. Mine open access full-text repositories for research data extraction
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Document parsing. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
eutils.ncbi.nlm.nih.govFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Open Access Mining Guide loads about 2.6k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 177 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 177 words, ~2,597 tokens.
.claude/skills/open-access-mining-guide/SKILL.md (or your agent's skills folder).A skill for systematically mining open access full-text repositories to extract structured research data at scale. Covers legal frameworks for text and data mining (TDM), major open access repositories and their APIs, full-text retrieval and parsing, section-level extraction, entity recognition in scientific text, and building reproducible mining pipelines.
Text and data mining of published literature operates within a specific legal framework that varies by jurisdiction. Understanding these rules is essential before starting any mining project.
Legal landscape for TDM:
EU Directive 2019/790 (DSM Directive):
- Article 3: TDM exception for research organizations
- Lawful access required (institutional subscription counts)
- Must be for scientific research purposes
- No opt-out possible for publishers
- Applies to EU/EEA research institutions
- Article 4: General TDM exception
- Available to anyone with lawful access
- Publishers CAN opt out (via robots.txt or metadata)
UK: TDM exception for non-commercial research (CDPA s.29A)
US: No specific TDM law; relies on fair use doctrine
- Transformative use generally favored by courts
- Google Books case (2015) supports large-scale text analysis
- But: database protection via Terms of Service
Practical guidelines:
- Mine open access content (CC-BY, CC-BY-SA) freely
- Mine subscription content under institutional license
- Check publisher TDM policies (Elsevier, Springer, Wiley
all have TDM APIs for licensed content)
- Never redistribute full text; share derived data only
- Credit the data source in publicationsRepository overview for full-text mining:
PubMed Central (PMC):
- Coverage: 8M+ full-text articles (biomedical/life sciences)
- Access: Free, OA subset freely downloadable
- Formats: XML (JATS), PDF
- API: E-utilities (Entrez), bulk FTP download
- License: varies by article (check individual licenses)
- Best for: biomedical systematic reviews, meta-analyses
- Bulk download: ftp.ncbi.nlm.nih.gov/pub/pmc/
Europe PMC:
- Coverage: PMC content + European-funded research
- Access: Free, REST API
- Formats: XML, JSON
- API: europepmc.org/RestfulWebService
- Annotations: sentence-level annotations, concepts, data links
- Best for: European research, annotated text mining
CORE (core.ac.uk):
- Coverage: 200M+ metadata records, 36M+ full texts
- Access: Free API (registration required)
- Formats: JSON, full text as extracted plain text
- Sources: aggregates from 10,000+ repositories worldwide
- Best for: cross-disciplinary mining, thesis/dissertation text
arXiv:
- Coverage: 2M+ preprints (physics, math, CS, etc.)
- Access: Free bulk download, API
- Formats: LaTeX source, PDF
- Bulk: Kaggle dataset, S3 requester-pays bucket
- Best for: STEM preprint analysis, citation studies
Unpaywall / OpenAlex:
- Coverage: tracks OA status of 200M+ works
- Access: Free API, database dump
- Use: Find OA versions of any DOI
- Best for: Locating freely available versions of papers
OpenAlex:
- Coverage: 250M+ works, all disciplines
- Access: Free API, no key required
- Features: Concepts, citation counts, author profiles, institution data
- Best for: Cross-disciplinary metadata and OA discoveryimport requests
import xml.etree.ElementTree as ET
import time
def fetch_pmc_fulltext(pmc_id):
"""
Fetch full-text XML from PubMed Central via E-utilities.
Args:
pmc_id: PMC identifier (e.g., "PMC7096724")
Returns:
Parsed article as structured dictionary
"""
base_url = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi"
params = {
"db": "pmc",
"id": pmc_id.replace("PMC", ""),
"rettype": "xml",
}
response = requests.get(base_url, params=params, timeout=30)
response.raise_for_status()
root = ET.fromstring(response.content)
article = parse_jats_xml(root)
return article
def parse_jats_xml(root):
"""
Parse JATS XML (Journal Article Tag Suite) into structured data.
JATS is the standard XML format for PMC articles.
"""
article = {}
# Title
title_elem = root.find(".//article-title")
article["title"] = "".join(title_elem.itertext()) if title_elem is not None else ""
# Abstract
abstract_elem = root.find(".//abstract")
if abstract_elem is not None:
article["abstract"] = "".join(abstract_elem.itertext()).strip()
# Body sections
body = root.find(".//body")
if body is not None:
article["sections"] = extract_sections(body)
# References
ref_list = root.find(".//ref-list")
if ref_list is not None:
article["references"] = extract_references(ref_list)
return article
def extract_sections(body_element):
"""
Extract sections with their titles and text content.
Preserves the hierarchical structure of the paper.
"""
sections = []
for sec in body_element.findall(".//sec"):
title_elem = sec.find("title")
title = title_elem.text if title_elem is not None else "Untitled"
paragraphs = []
for p in sec.findall("p"):
text = "".join(p.itertext()).strip()
if text:
paragraphs.append(text)
sections.append({
"title": title,
"text": "\n".join(paragraphs),
"id": sec.get("id", ""),
})
return sectionsdef batch_mine_pmc(pmc_ids, output_dir, delay=0.4):
"""
Mine multiple PMC articles with rate limiting.
NCBI E-utilities rate limit:
- Without API key: 3 requests/second
- With API key: 10 requests/second
- Register for API key at ncbi.nlm.nih.gov/account/
"""
import json
import os
results = []
errors = []
for i, pmc_id in enumerate(pmc_ids):
try:
article = fetch_pmc_fulltext(pmc_id)
results.append(article)
# Save individual article
output_path = os.path.join(output_dir, f"{pmc_id}.json")
with open(output_path, "w") as f:
json.dump(article, f, indent=2)
if (i + 1) % 100 == 0:
print(f"Processed {i + 1}/{len(pmc_ids)} articles")
except Exception as e:
errors.append({"pmc_id": pmc_id, "error": str(e)})
# Rate limiting
time.sleep(delay)
print(f"Successfully mined {len(results)} articles, "
f"{len(errors)} errors")
return results, errorsTargeted extraction by paper section:
Introduction:
- Research questions and hypotheses
- Knowledge gaps identified
- Theoretical framework references
Methods:
- Study design (RCT, cohort, case-control, etc.)
- Sample size and population characteristics
- Measurement instruments and their validity
- Statistical analysis methods
- Software and versions used
Results:
- Effect sizes with confidence intervals
- P-values and test statistics
- Participant flow (enrollment, dropout, analysis)
- Tables and figures (structured data)
Discussion:
- Key findings summarized
- Comparison with prior work
- Limitations acknowledged
- Future directions proposed
- Clinical/practical implicationsdef extract_scientific_entities(text):
"""
Extract scientific named entities from full text.
For biomedical text, use specialized NER models:
- SciSpaCy: biomedical NER (diseases, chemicals, genes)
- BioBERT: contextual biomedical NER
- PubTator: NCBI's annotation service
"""
import scispacy
import spacy
nlp = spacy.load("en_ner_bionlp13cg_md")
doc = nlp(text)
entities = []
for ent in doc.ents:
entities.append({
"text": ent.text,
"label": ent.label_,
"start": ent.start_char,
"end": ent.end_char,
})
return entitiesRecommended pipeline structure:
1. Query definition:
- Define search terms, date ranges, inclusion criteria
- Document in a protocol file (version-controlled)
2. Article retrieval:
- Search API for matching articles
- Download full text (XML/PDF)
- Store raw data with metadata
3. Text extraction:
- Parse XML or extract text from PDF
- Section segmentation
- Table and figure extraction (if needed)
4. Information extraction:
- NER for entities of interest
- Relation extraction
- Numeric data extraction (effect sizes, p-values)
5. Quality control:
- Sample-based manual validation (10-20% of results)
- Inter-annotator agreement on validation sample
- Error analysis and pipeline refinement
6. Data export:
- Structured output (CSV, JSON, database)
- Provenance tracking (which article, which section)
- Ready for downstream analysis
Best practices:
- Version control the entire pipeline code
- Log all API queries and responses
- Set random seeds for any sampling steps
- Share the pipeline code in supplementary materials
- Use DOIs or PMCIDs as stable article identifiers
- Cache downloaded articles to avoid re-fetchingOpen access full-text mining enables research at a scale impossible with manual reading. A single researcher can systematically extract data from thousands of papers, enabling comprehensive evidence synthesis, trend analysis, and hypothesis generation. The key requirements are respecting legal and ethical boundaries, building robust parsing pipelines, and rigorously validating extracted data against manual review.
© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/literature/fulltext/open-access-mining-guide of wentorai/research-plugins.
Open the folder on GitHubat commit bf44b3c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.
Open Access Mining Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Open Access Mining Guide this skillwentorai/research-plugins | 298 | 1 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Literature PDF OCR Library BuilderLigphiDonk/Oh-my--paper | 738 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Paper Figure Extractorjuliye2025/evil-read-arxiv | 1.7k | — | ~298 | Automated safety check: Pass | None | |
| Bilingual Paper ReaderYuan1z0825/nature-skills | 47k | — | ~961 | Automated safety check: Pass | Apache-2.0 | |
| Paper Image ExtractorLigphiDonk/Oh-my--paper | 738 | 1 repos | ~810 | Automated safety check: Pass | MIT | |
| ARA Research CompilerOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT |
LigphiDonk/Oh-my--paper
Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.
juliye2025/evil-read-arxiv
Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them.
Yuan1z0825/nature-skills
Creates a source-grounded Chinese-English reader for a research paper, with aligned text, figures, tables and equations, or answers questions about a given passage.
LigphiDonk/Oh-my--paper
Extracts figures from a research paper, preferring the arXiv source package for original-quality images and falling back to PDF extraction.
Orchestra-Research/AI-Research-SKILLs
Turns papers, repositories, logs or notes into an Agent-Native Research Artifact with claims, concepts, configs, an exploration graph and grounded evidence.
Mathews-Tom/armory
Converts research papers into executable skill packages via document conversion, critical analysis, and co-evolutionary refinement.
wentorai/research-plugins
Craft structured research abstracts that maximize clarity and journal acceptance
wentorai/research-plugins
Manage academic citations across BibTeX, APA, MLA, and Chicago formats
wentorai/research-plugins
Summarize academic papers with structured extraction of key elements
wentorai/research-plugins
Evidence-based study techniques for academic learning and retention
wentorai/research-plugins
Adjust writing tone and register for academic audiences and venues
wentorai/research-plugins
Academic translation, post-editing, and Chinglish correction guide
Categories
Mine open access full-text repositories for research data extraction. Open Access Mining Guide is an agent skill from wentorai/research-plugins.
Open Access Mining Guide fits situations like: tasks that involve Document parsing.
Run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a claude-code`. Or copy the skill folder (skills/literature/fulltext/open-access-mining-guide in wentorai/research-plugins) into .claude/skills/open-access-mining-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a codex`. Or copy the skill folder (skills/literature/fulltext/open-access-mining-guide in wentorai/research-plugins) into .agents/skills/open-access-mining-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-access-mining-guide, .gemini/skills/open-access-mining-guide, .github/skills/open-access-mining-guide and .opencode/skills/open-access-mining-guide in your project.
SKILL.md names no scripts, command-line tools or credentials: Open Access Mining Guide is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: eutils.ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Open Access Mining Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Open Access Mining Guide: Literature PDF OCR Library Builder (LigphiDonk/Oh-my--paper, 738 stars), Paper Figure Extractor (juliye2025/evil-read-arxiv, 1.7k stars), Bilingual Paper Reader (Yuan1z0825/nature-skills, 47k stars) and Paper Image Extractor (LigphiDonk/Oh-my--paper, 738 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.
Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.