Agent skill

Fulltext Retrieval

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set.

MITAuto-check passedDocuments & Office

Install Fulltext Retrieval

skills CLI
$ npx skills add Aperivue/medsci-skills --skill fulltext-retrieval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills fulltext-retrieval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fulltext-retrieval .claude/skills/fulltext-retrieval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fulltext-retrieval
GitHub stars
331
Token cost
~1.9k tokens
SKILL.md length
730 words
Files
13 (incl. references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set.

  • You need full-text PDFs for a list of DOIs
  • SKILL.md covers Pipeline, Run, Output and Retrieval report (--report), plus 2 more sections
  • Runs Python, Shell and JavaScript scripts from its folder; calls python and pip
  • Such as a meta-analysis screening set

What it does

Fulltext Retrieval is an agent skill from Aperivue/medsci-skills. Use when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set. Batch-downloads open-access copies via Unpaywall, PMC, OpenAlex and Crossref, lists paywalled papers for manual access, and can convert PDFs to Markdown.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `fetch_oa.py`, `fetch_oa_report_challenge/expected/projection.json` and `fetch_oa_report_challenge/extracted_text.json`).

It sits in Documents & Office, covering Academic paper search, PDF and Markdown. It works with arXiv. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • You need full-text PDFs for a list of DOIs
  • Such as a meta-analysis screening set

Example prompts

  • “/fulltext-retrieval”

Requirements

  • Python 3
  • Node.js
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, Shell and JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pypi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fulltext Retrieval loads about 1.9k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 730 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 730 words, ~1,938 tokens.

Download SKILL.mdSave it as .claude/skills/fulltext-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
fulltext-retrieval
description
Use when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set. Batch-downloads open-access copies via Unpaywall, PMC, OpenAlex and Crossref, lists paywalled papers for manual access, and can convert PDFs to Markdown.
metadata.triggers
PDF download, fulltext retrieval, open access PDF, batch download papers, meta-analysis PDF, PDF to markdown, convert PDF

Fulltext Retrieval Skill

Batch download open-access full-text PDFs from a DOI list using legitimate OA APIs only. Paywalled articles fail by design and are listed in manual_needed.txt for institutional access or ILL; never work around a paywall or publisher access control.

Pipeline

DOI → arXiv (10.48550/arXiv.* DOIs) → Unpaywall → PMC (Europe PMC / OA FTP / web) → OpenAlex → Crossref → landing page

Each DOI goes through these sources in order until a valid PDF (≥10 KB, %PDF- header) is found. arXiv DOIs (10.48550/arXiv.2401.01234, version suffixes, old-style hep-th/9901001, or a bare arXiv: id) resolve directly to the arXiv PDF first.

Run

Requires Python 3.10+ (stdlib only) and a contact email, which Unpaywall's Terms of Service require. The script paces its requests (0.3–0.5 s delays) for the APIs' rate limits.

bash
python "${CLAUDE_SKILL_DIR}/fetch_oa.py" dois.txt --output pdfs/ --email your@email.com

# Verbose mode for debugging (per-DOI source trace)
python "${CLAUDE_SKILL_DIR}/fetch_oa.py" dois.txt -o pdfs/ -e your@email.com --verbose

Input formats:

  • Plain text — one DOI per line.
  • TSV / CSV with header, or a Markdown pipe table — must contain a DOI column; optional PMID, Title, and FirstAuthor (surname or full name) columns.

A PMID makes the PMC lookup more reliable (PMID → PMCID conversion). Supply Title where available: a DOI-only worklist can download a PDF but cannot establish title agreement. FirstAuthor is optional additional evidence.

Output

  • PDFs saved as {DOI_safe}.pdf (slashes replaced with underscores).
  • pdfs/retrieval_report.json — structured per-DOI report (below); override with --report PATH.
  • <output>/manual_needed.txt — DOIs that could not be retrieved via OA; when a PMCID was resolved, the line also carries it and the PubMed Central article URL to open in a browser.
  • Summary with arXiv/OA/PMC/fail/skip counts.

Retrieval report (--report)

Every run writes the report (default <output>/retrieval_report.json), schema 2:

json
{
  "schema_version": 2,
  "generated_by": "fetch_oa.py",
  "counts": {"total": 4, "retrieved": 3, "not_retrieved": 1, "title_mismatch": 1,
             "source_identity": {"consistent": 1, "conflict": 1, "unresolved": 1, "unavailable": 1}},
  "items": [
    {"doi": "10.1000/synthetic.example", "pmid": "", "title": "Example title",
     "first_author": "", "status": "oa", "source": "unpaywall",
     "file": "10.1000_synthetic.example.pdf", "size_bytes": 482113, "page_count": 9,
     "file_sha256": "<SHA-256 of the downloaded file>", "title_match": "match",
     "source_identity": {"status": "consistent", "reason": "title_and_identifier_agree",
                         "text_scope": "first_page_front_matter", "title_match": "match",
                         "doi_match": "match", "observed_identifiers": ["10.1000/synthetic.example"],
                         "first_author_match": "unavailable"}}
  ]
}

status (arxiv | oa | pmc | skip | fail), source, and counts.retrieved describe the resolver result, including existing files (skip). They do not count identity-verified papers. No PDF is automatically deleted or rejected. page_count comes from Poppler's pdfinfo (null without it) and is recorded, not judged — a 3-page "article" or a 4-page "book" is worth opening.

source_identity.statusMeaning / action
consistentComplete normalized title and a compatible DOI/arXiv identifier occur in the bounded first-page front matter, with no supplement / preface / table-of-contents heading and no retraction / erratum / correction / corrigendum / expression-of-concern heading there; an optional supplied author must also match. Evidence agrees, but this is not independent source verification or claim validation.
conflictBoth the title and observed identifier differ. Inspect the PDF and requested record.
unresolvedEvidence is incomplete or ambiguous: title-only, DOI-only, missing author, multiple identifiers, a matching title with another DOI/version, or a supplement / preface / table-of-contents file that names the work without being it (supplement_or_front_matter). A retraction notice, erratum, correction, corrigendum or expression of concern whose heading line sits in that area is likewise unresolved (correction_or_retraction_notice). Inspect before using as evidence.
unavailableNo usable extracted text, Poppler unavailable, no output PDF, or the PDF changed during assessment. No current identity assessment was possible.

Evidence is limited to the first page before a recognized abstract/body/reference heading (at most 40 lines / 4,000 characters), so a title cited in the body or references does not count. A title_match of match needs the complete normalized title on up to six consecutive lines; partial overlap is unavailable and low overlap an advisory mismatch. Cover sheets, unusual reading order, short or changed titles and DOI footers outside that area can stay unresolved. PDF metadata and the filename alone are not identity evidence. Explicit arXiv versions must agree; a preprint/published-version DOI difference needs review, not automatic rejection.

Show full SKILL.md (197 more words)Show less

Downstream reports must preserve source_identity and file_sha256, keep unresolved items visible, and check the hash still identifies the file being used. Older reports without identity evidence remain unassessed; do not infer identity from retrieved or title_match=match. Full-text conversion does not resolve an identity warning.

Attach PDFs into Zotero ("Find Available PDF")

For paywalled-but-licensed papers the OA resolvers miss, references/find_available_pdf.js is a user-run snippet for Zotero's Tools → Developer → Run JavaScript (no-code equivalent: right-click → "Find Available PDF"). It triggers Zotero's own addAvailablePDF / addAvailablePDFs, so it reuses the user's OpenURL resolver / institutional proxy config; no credentials, proxy hosts, or institutional identifiers are hard-coded or leave the Zotero client. It is user-initiated and depends on the live Zotero session, so record its results manually — they are not reproducible CI evidence. /lit-sync Phase 2.7 orchestrates both routes and reconciles them in a report.

PDF → Markdown Conversion (Optional)

Convert downloaded PDFs to Markdown when the same papers will be read repeatedly (data extraction from k≥5 studies, a meta-analysis pipeline); for one-pass screening, read the PDF directly.

bash
# Install (one-time)
pip install pymupdf4llm

# Convert all PDFs in a directory (.md files land alongside the .pdf files)
python "${CLAUDE_SKILL_DIR}/pdf_to_md.py" pdfs/ -v

# Custom output directory
python "${CLAUDE_SKILL_DIR}/pdf_to_md.py" pdfs/ -o markdown/

# First 10 pages only (useful for long supplements)
python "${CLAUDE_SKILL_DIR}/pdf_to_md.py" pdfs/ --pages 0-9

# Overwrite existing conversions
python "${CLAUDE_SKILL_DIR}/pdf_to_md.py" pdfs/ --force

Images are skipped, so figures survive only as caption text. Scanned-only PDFs (no text layer) convert poorly. pdf_to_md.py needs pymupdf4llm (AGPL-3.0), an optional dependency; fetch_oa.py stays stdlib-only.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (references) in skills/fulltext-retrieval of Aperivue/medsci-skills.

  • SKILL.md
  • fetch_oa.py
  • fetch_oa_report_challenge/expected/projection.json
  • fetch_oa_report_challenge/extracted_text.json
  • fetch_oa_report_challenge/results.json
  • fetch_oa_report_challenge/run_challenge.py
  • fetch_oa_report_challenge/verify.sh
  • fetch_oa_report_challenge/worklist.tsv
  • pdf_to_md.py
  • references/find_available_pdf.js
  • skill.yml
  • tests/test_pdf_to_md.py
  • tests/test_source_identity.py

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Fulltext Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fulltext Retrieval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fulltext Retrieval this skillAperivue/medsci-skills331—~1.9kAutomated safety check: PassMIT
Paper Interpretationdigoal/blog8.6k—~1.5kAutomated safety check: PassGPL-2.0
Paper LensYSQ-boop/paper-lens101—~1.3kAutomated safety check: PassApache-2.0
Paper Interpreterchujianyun/skills740—~810Automated safety check: PassCustom licence
Omh Paper Learningrlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMIT
Paper2htmlcnfjlhj/ai-collab-playbook453—~2.4kAutomated safety check: PassNone

Similar skills

  • 从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…

    8.6k GitHub stars~1.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Paper Lens

    YSQ-boop/paper-lens

    Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.

    101 GitHub stars~1.3k tokensUpdated 10 days ago
    Documents & OfficeAuto-check passed
  • Paper Interpreter

    chujianyun/skills

    论文解读助手。适用于用户发送 arXiv 论文链接,并希望下载论文、解读论文、生成读书笔记、做论文拆解或输出详细报告时使用。会在工作目录创建论文文件夹、下载 PDF 与 TeX Source(如有)、生成中文 Markdown 报告。默认先交付初稿,不自动复查;如果用户明确同意,再安排后续复查。不适用于只要简短推荐语的情况。

    740 GitHub stars~810 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Omh Paper Learning

    rlaope/oh-my-hermes

    [omh] Paper or paper PDF to understand: explain a supplied paper or paper/PDF at a selected level while preserving full section coverage and source evidence boundaries.

    3.2k GitHub stars~2.2k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Paper2html

    cnfjlhj/ai-collab-playbook

    A skill your agent uses when turning an academic paper PDF/arXiv/OpenReview page/local LaTeX source into a single-file Chinese HTML deep-reading page, especially when the user wants a Cheat-Sheet…

    453 GitHub stars~2.4k tokensUpdated 25 days ago
    Documents & OfficeAuto-check passed
  • Paper Reading

    Edwardxlai/easyread

    论文共读:用 EasyRead(E:\CursorProject\easyread)这个本地工具读论文。用户给一篇 PDF 或 arXiv 编号时,导入文献库、翻译成中文(后台引擎或对话里的 agent 亲自译);之后边读边讨论,agent 读用户在页面上的笔记和提问,把回答、解释、原文核对提示追加到对应段落旁。用于…

    871 GitHub stars~567 tokensUpdated today
    Research & ScienceAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    331 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    331 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    331 GitHub stars~3.9k tokensUpdated 4 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    331 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Fill Protocol

    Aperivue/medsci-skills

    A skill your agent uses when an institutional Word form (.doc/.docx IRB protocol, ethics application, grant template) must be filled without breaking its styles, tables, fonts or page layout.

    331 GitHub stars~1.7k tokensUpdated 4 days ago
    Auto-check passed
  • Find Cohort Gap

    Aperivue/medsci-skills

    A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).

    331 GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Fulltext Retrieval

What does Fulltext Retrieval do?

A skill your agent uses when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set. Fulltext Retrieval is an agent skill from Aperivue/medsci-skills. Use when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set.

When should I use Fulltext Retrieval?

Fulltext Retrieval fits situations like: you need full-text PDFs for a list of DOIs; such as a meta-analysis screening set.

How do I install Fulltext Retrieval in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill fulltext-retrieval -a claude-code`. Or copy the skill folder (skills/fulltext-retrieval in Aperivue/medsci-skills) into .claude/skills/fulltext-retrieval in your project. Claude Code loads it when a task matches its description.

How do I install Fulltext Retrieval in Codex?

Run `npx skills add Aperivue/medsci-skills --skill fulltext-retrieval -a codex`. Or copy the skill folder (skills/fulltext-retrieval in Aperivue/medsci-skills) into .agents/skills/fulltext-retrieval in your project. Codex loads it when a task matches its description.

Can I use Fulltext Retrieval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill fulltext-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fulltext-retrieval, .gemini/skills/fulltext-retrieval, .github/skills/fulltext-retrieval and .opencode/skills/fulltext-retrieval in your project.

What does Fulltext Retrieval need to run?

Going by SKILL.md and its folder, Fulltext Retrieval needs Python, a shell and JavaScript for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3; Node.js; A Bash shell.

Does Fulltext Retrieval access the network?

SKILL.md names 1 domain. As links in the text: pypi.org. This is read from the text; nothing was executed.

Is Fulltext Retrieval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fulltext Retrieval use?

Fulltext Retrieval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fulltext Retrieval use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 765 tokens, read only when the agent opens those files.

What are the alternatives to Fulltext Retrieval?

Skills that share tags, products or a category with Fulltext Retrieval: Paper Interpretation (digoal/blog, 8.6k stars), Paper Lens (YSQ-boop/paper-lens, 101 stars), Paper Interpreter (chujianyun/skills, 740 stars) and Omh Paper Learning (rlaope/oh-my-hermes, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fulltext Retrieval?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 331 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.