Agent skill

Open Access Mining Guide

by wentorai in wentorai/research-plugins

Mine open access full-text repositories for research data extraction

MITAuto-check passedResearch & Science

Install Open Access Mining Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill open-access-mining-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins open-access-mining-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature/fulltext/open-access-mining-guide .claude/skills/open-access-mining-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
open-access-mining-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
177 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Mine open access full-text repositories for research data extraction

  • Tasks that involve Document parsing
  • SKILL.md covers Legal Framework for Text and…, Major Open Access Repositories, Full-Text Retrieval and Parsing and Information Extraction from…, plus 1 more section
  • Reaches eutils.ncbi.nlm.nih.gov

What it does

Open Access Mining Guide is an agent skill from wentorai/research-plugins. Mine open access full-text repositories for research data extraction

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Document parsing. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Document parsing

Example prompts

  • “/open-access-mining-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • eutils.ncbi.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Open Access Mining Guide loads about 2.6k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 177 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 177 words, ~2,597 tokens.

Download SKILL.mdSave it as .claude/skills/open-access-mining-guide/SKILL.md (or your agent's skills folder).
name
open-access-mining-guide
description
Mine open access full-text repositories for research data extraction

Open Access Mining Guide

A skill for systematically mining open access full-text repositories to extract structured research data at scale. Covers legal frameworks for text and data mining (TDM), major open access repositories and their APIs, full-text retrieval and parsing, section-level extraction, entity recognition in scientific text, and building reproducible mining pipelines.

Rights and Regulations

Text and data mining of published literature operates within a specific legal framework that varies by jurisdiction. Understanding these rules is essential before starting any mining project.

Legal landscape for TDM:

EU Directive 2019/790 (DSM Directive):
  - Article 3: TDM exception for research organizations
    - Lawful access required (institutional subscription counts)
    - Must be for scientific research purposes
    - No opt-out possible for publishers
    - Applies to EU/EEA research institutions
  - Article 4: General TDM exception
    - Available to anyone with lawful access
    - Publishers CAN opt out (via robots.txt or metadata)

UK: TDM exception for non-commercial research (CDPA s.29A)

US: No specific TDM law; relies on fair use doctrine
  - Transformative use generally favored by courts
  - Google Books case (2015) supports large-scale text analysis
  - But: database protection via Terms of Service

Practical guidelines:
  - Mine open access content (CC-BY, CC-BY-SA) freely
  - Mine subscription content under institutional license
  - Check publisher TDM policies (Elsevier, Springer, Wiley
    all have TDM APIs for licensed content)
  - Never redistribute full text; share derived data only
  - Credit the data source in publications

Major Open Access Repositories

Repository Comparison
Repository overview for full-text mining:

PubMed Central (PMC):
  - Coverage: 8M+ full-text articles (biomedical/life sciences)
  - Access: Free, OA subset freely downloadable
  - Formats: XML (JATS), PDF
  - API: E-utilities (Entrez), bulk FTP download
  - License: varies by article (check individual licenses)
  - Best for: biomedical systematic reviews, meta-analyses
  - Bulk download: ftp.ncbi.nlm.nih.gov/pub/pmc/

Europe PMC:
  - Coverage: PMC content + European-funded research
  - Access: Free, REST API
  - Formats: XML, JSON
  - API: europepmc.org/RestfulWebService
  - Annotations: sentence-level annotations, concepts, data links
  - Best for: European research, annotated text mining

CORE (core.ac.uk):
  - Coverage: 200M+ metadata records, 36M+ full texts
  - Access: Free API (registration required)
  - Formats: JSON, full text as extracted plain text
  - Sources: aggregates from 10,000+ repositories worldwide
  - Best for: cross-disciplinary mining, thesis/dissertation text

arXiv:
  - Coverage: 2M+ preprints (physics, math, CS, etc.)
  - Access: Free bulk download, API
  - Formats: LaTeX source, PDF
  - Bulk: Kaggle dataset, S3 requester-pays bucket
  - Best for: STEM preprint analysis, citation studies

Unpaywall / OpenAlex:
  - Coverage: tracks OA status of 200M+ works
  - Access: Free API, database dump
  - Use: Find OA versions of any DOI
  - Best for: Locating freely available versions of papers

OpenAlex:
  - Coverage: 250M+ works, all disciplines
  - Access: Free API, no key required
  - Features: Concepts, citation counts, author profiles, institution data
  - Best for: Cross-disciplinary metadata and OA discovery

Full-Text Retrieval and Parsing

Retrieving from PubMed Central
python
import requests
import xml.etree.ElementTree as ET
import time

def fetch_pmc_fulltext(pmc_id):
    """
    Fetch full-text XML from PubMed Central via E-utilities.

    Args:
        pmc_id: PMC identifier (e.g., "PMC7096724")

    Returns:
        Parsed article as structured dictionary
    """
    base_url = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi"
    params = {
        "db": "pmc",
        "id": pmc_id.replace("PMC", ""),
        "rettype": "xml",
    }

    response = requests.get(base_url, params=params, timeout=30)
    response.raise_for_status()

    root = ET.fromstring(response.content)
    article = parse_jats_xml(root)

    return article


def parse_jats_xml(root):
    """
    Parse JATS XML (Journal Article Tag Suite) into structured data.
    JATS is the standard XML format for PMC articles.
    """
    article = {}

    # Title
    title_elem = root.find(".//article-title")
    article["title"] = "".join(title_elem.itertext()) if title_elem is not None else ""

    # Abstract
    abstract_elem = root.find(".//abstract")
    if abstract_elem is not None:
        article["abstract"] = "".join(abstract_elem.itertext()).strip()

    # Body sections
    body = root.find(".//body")
    if body is not None:
        article["sections"] = extract_sections(body)

    # References
    ref_list = root.find(".//ref-list")
    if ref_list is not None:
        article["references"] = extract_references(ref_list)

    return article


def extract_sections(body_element):
    """
    Extract sections with their titles and text content.
    Preserves the hierarchical structure of the paper.
    """
    sections = []
    for sec in body_element.findall(".//sec"):
        title_elem = sec.find("title")
        title = title_elem.text if title_elem is not None else "Untitled"
        paragraphs = []
        for p in sec.findall("p"):
            text = "".join(p.itertext()).strip()
            if text:
                paragraphs.append(text)

        sections.append({
            "title": title,
            "text": "\n".join(paragraphs),
            "id": sec.get("id", ""),
        })

    return sections
Batch Processing Pipeline
python
def batch_mine_pmc(pmc_ids, output_dir, delay=0.4):
    """
    Mine multiple PMC articles with rate limiting.

    NCBI E-utilities rate limit:
    - Without API key: 3 requests/second
    - With API key: 10 requests/second
    - Register for API key at ncbi.nlm.nih.gov/account/
    """
    import json
    import os

    results = []
    errors = []

    for i, pmc_id in enumerate(pmc_ids):
        try:
            article = fetch_pmc_fulltext(pmc_id)
            results.append(article)

            # Save individual article
            output_path = os.path.join(output_dir, f"{pmc_id}.json")
            with open(output_path, "w") as f:
                json.dump(article, f, indent=2)

            if (i + 1) % 100 == 0:
                print(f"Processed {i + 1}/{len(pmc_ids)} articles")

        except Exception as e:
            errors.append({"pmc_id": pmc_id, "error": str(e)})

        # Rate limiting
        time.sleep(delay)

    print(f"Successfully mined {len(results)} articles, "
          f"{len(errors)} errors")
    return results, errors

Information Extraction from Full Text

Section-Level Extraction
Targeted extraction by paper section:

Introduction:
  - Research questions and hypotheses
  - Knowledge gaps identified
  - Theoretical framework references

Methods:
  - Study design (RCT, cohort, case-control, etc.)
  - Sample size and population characteristics
  - Measurement instruments and their validity
  - Statistical analysis methods
  - Software and versions used

Results:
  - Effect sizes with confidence intervals
  - P-values and test statistics
  - Participant flow (enrollment, dropout, analysis)
  - Tables and figures (structured data)

Discussion:
  - Key findings summarized
  - Comparison with prior work
  - Limitations acknowledged
  - Future directions proposed
  - Clinical/practical implications
Named Entity Recognition for Science
python
def extract_scientific_entities(text):
    """
    Extract scientific named entities from full text.

    For biomedical text, use specialized NER models:
    - SciSpaCy: biomedical NER (diseases, chemicals, genes)
    - BioBERT: contextual biomedical NER
    - PubTator: NCBI's annotation service
    """
    import scispacy
    import spacy

    nlp = spacy.load("en_ner_bionlp13cg_md")
    doc = nlp(text)

    entities = []
    for ent in doc.ents:
        entities.append({
            "text": ent.text,
            "label": ent.label_,
            "start": ent.start_char,
            "end": ent.end_char,
        })

    return entities

Building Reproducible Pipelines

Pipeline Architecture
Recommended pipeline structure:

1. Query definition:
   - Define search terms, date ranges, inclusion criteria
   - Document in a protocol file (version-controlled)

2. Article retrieval:
   - Search API for matching articles
   - Download full text (XML/PDF)
   - Store raw data with metadata

3. Text extraction:
   - Parse XML or extract text from PDF
   - Section segmentation
   - Table and figure extraction (if needed)

4. Information extraction:
   - NER for entities of interest
   - Relation extraction
   - Numeric data extraction (effect sizes, p-values)

5. Quality control:
   - Sample-based manual validation (10-20% of results)
   - Inter-annotator agreement on validation sample
   - Error analysis and pipeline refinement

6. Data export:
   - Structured output (CSV, JSON, database)
   - Provenance tracking (which article, which section)
   - Ready for downstream analysis

Best practices:
  - Version control the entire pipeline code
  - Log all API queries and responses
  - Set random seeds for any sampling steps
  - Share the pipeline code in supplementary materials
  - Use DOIs or PMCIDs as stable article identifiers
  - Cache downloaded articles to avoid re-fetching

Open access full-text mining enables research at a scale impossible with manual reading. A single researcher can systematically extract data from thousands of papers, enabling comprehensive evidence synthesis, trend analysis, and hypothesis generation. The key requirements are respecting legal and ethical boundaries, building robust parsing pipelines, and rigorously validating extracted data against manual review.

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/literature/fulltext/open-access-mining-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Open Access Mining Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Open Access Mining Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Open Access Mining Guide this skillwentorai/research-plugins2981 repos~2.6kAutomated safety check: PassMIT
Literature PDF OCR Library BuilderLigphiDonk/Oh-my--paper738—~1.1kAutomated safety check: PassMIT
Paper Figure Extractorjuliye2025/evil-read-arxiv1.7k—~298Automated safety check: PassNone
Bilingual Paper ReaderYuan1z0825/nature-skills47k—~961Automated safety check: PassApache-2.0
Paper Image ExtractorLigphiDonk/Oh-my--paper7381 repos~810Automated safety check: PassMIT
ARA Research CompilerOrchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Literature PDF OCR Library Builder

    LigphiDonk/Oh-my--paper

    Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

    738 GitHub stars~1.1k tokensUpdated 5 mo ago
    Research & ScienceAuto-check passed
  • Paper Figure Extractor

    juliye2025/evil-read-arxiv

    Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them.

    1.7k GitHub stars~298 tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Bilingual Paper Reader

    Yuan1z0825/nature-skills

    Creates a source-grounded Chinese-English reader for a research paper, with aligned text, figures, tables and equations, or answers questions about a given passage.

    47k GitHub stars~961 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Paper Image Extractor

    LigphiDonk/Oh-my--paper

    Extracts figures from a research paper, preferring the arXiv source package for original-quality images and falling back to PDF extraction.

    738 GitHub starsUsed in 1 repo~810 tokens
    Research & ScienceAuto-check passed
  • ARA Research Compiler

    Orchestra-Research/AI-Research-SKILLs

    Turns papers, repositories, logs or notes into an Agent-Native Research Artifact with claims, concepts, configs, an exploration graph and grounded evidence.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Paper To Skill

    Mathews-Tom/armory

    Converts research papers into executable skill packages via document conversion, critical analysis, and co-evolutionary refinement.

    328 GitHub stars~1.8k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Open Access Mining Guide

What does Open Access Mining Guide do?

Mine open access full-text repositories for research data extraction. Open Access Mining Guide is an agent skill from wentorai/research-plugins.

When should I use Open Access Mining Guide?

Open Access Mining Guide fits situations like: tasks that involve Document parsing.

How do I install Open Access Mining Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a claude-code`. Or copy the skill folder (skills/literature/fulltext/open-access-mining-guide in wentorai/research-plugins) into .claude/skills/open-access-mining-guide in your project. Claude Code loads it when a task matches its description.

How do I install Open Access Mining Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a codex`. Or copy the skill folder (skills/literature/fulltext/open-access-mining-guide in wentorai/research-plugins) into .agents/skills/open-access-mining-guide in your project. Codex loads it when a task matches its description.

Can I use Open Access Mining Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill open-access-mining-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-access-mining-guide, .gemini/skills/open-access-mining-guide, .github/skills/open-access-mining-guide and .opencode/skills/open-access-mining-guide in your project.

What does Open Access Mining Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Open Access Mining Guide is instructions for the agent only. Our summary lists: Python 3.

Does Open Access Mining Guide access the network?

SKILL.md names 1 domain. In commands or code: eutils.ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Open Access Mining Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Open Access Mining Guide use?

Open Access Mining Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Open Access Mining Guide use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Open Access Mining Guide?

Skills that share tags, products or a category with Open Access Mining Guide: Literature PDF OCR Library Builder (LigphiDonk/Oh-my--paper, 738 stars), Paper Figure Extractor (juliye2025/evil-read-arxiv, 1.7k stars), Bilingual Paper Reader (Yuan1z0825/nature-skills, 47k stars) and Paper Image Extractor (LigphiDonk/Oh-my--paper, 738 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Open Access Mining Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.