Agent skill

Mining Pubmed Literature

by maziyarpanahi in maziyarpanahi/openmed

Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora.

Apache-2.0Auto-check passedResearch & Science

Install Mining Pubmed Literature

skills CLI
$ npx skills add maziyarpanahi/openmed --skill mining-pubmed-literature -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed mining-pubmed-literature --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mining-pubmed-literature .claude/skills/mining-pubmed-literature && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mining-pubmed-literature
GitHub stars
5.5k
Token cost
~1.7k tokens
SKILL.md length
525 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora.

  • Works in 5 steps: Build the query. Combine… → ESearch with usehistory=y to capture… → Batch-fetch with EFetch/ESummary in… → …
  • The user wants citations for a condition
  • SKILL.md covers When to use, Quick start (real E-utilities…, ESummary for structured metadata and Workflow, plus 3 more sections
  • Calls curl; reaches eutils.ncbi.nlm.nih.gov; needs API_KEY

What it does

Mining Pubmed Literature is an agent skill from maziyarpanahi/openmed. Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora. Use when the user wants citations for a condition or drug, abstracts to summarize, MeSH-based searches, or a corpus of literature to run NER over. Trigger keywords: PubMed, PMC, NCBI, E-utilities, ESearch, EFetch, ESummary, MeSH, PMID, literature search, abstracts, evidence. Pairs adjacent to OpenMed: fetched abstracts feed openmed.analyzetext for biomedical NER, and…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Academic paper search, Rate limiting and Literature review. It works with PubMed and NCBI. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants citations for a condition
  • Abstracts to summarize
  • MeSH-based searches
  • A corpus of literature to run NER over

Example prompts

  • “Use the mining-pubmed-literature skill to search and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical…”
  • “/mining-pubmed-literature”

Requirements

  • Python 3
  • A credential in API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Build the query. Combine OpenMed-extracted terms with MeSH tags and field
  2. ESearch with usehistory=y to capture WebEnv + query_key and the count.
  3. Batch-fetch with EFetch/ESummary in pages of ≤ ~200 IDs (or by history),
  4. Parse abstracts/metadata; store PMID, title, journal, date, abstract text.
  5. NER the abstracts with openmed.analyze_text to extract diseases, drugs,

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • eutils.ncbi.nlm.nih.gov

    Also links to:

    • ncbi.nlm.nih.gov
    • support.nlm.nih.gov
    • pubmed.ncbi.nlm.nih.gov
    • meshb.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mining Pubmed Literature loads about 1.7k tokens when it runs. Until then it costs about 175 tokens; SKILL.md has 525 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~175
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 525 words, ~1,747 tokens.

Download SKILL.mdSave it as .claude/skills/mining-pubmed-literature/SKILL.md (or your agent's skills folder).
name
mining-pubmed-literature
description
Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora. Use when the user wants citations for a condition or drug, abstracts to summarize, MeSH-based searches, or a corpus of literature to run NER over. Trigger keywords: PubMed, PMC, NCBI, E-utilities, ESearch, EFetch, ESummary, MeSH, PMID, literature search, abstracts, evidence. Pairs adjacent to OpenMed: fetched abstracts feed openmed.analyze_text for biomedical NER, and OpenMed-extracted diagnoses/drugs/genes become the search terms. E-utilities are public; an optional free API key raises rate limits from 3 to 10 requests/second.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
research-genomics
metadata.pairs
adjacent
metadata.version
1.0

Mining PubMed & PMC literature (NCBI E-utilities)

Search PubMed (citations/abstracts) and PMC (full text) programmatically with NCBI E-utilities — the stable HTTP interface to Entrez. The core pattern is two steps: ESearch returns matching record IDs (PMIDs), then EFetch (or ESummary) downloads the records. The Entrez History server (usehistory=y) lets you chain the two without re-sending thousands of IDs.

E-utilities are public. No key is required, but a free API key raises your limit from 3 to 10 requests/second and is strongly recommended for batch work.

When to use

  • OpenMed extracted a diagnosis, drug, or gene and you want supporting literature.
  • You need abstracts to summarize or to assemble a corpus for biomedical NER.
  • You want MeSH-anchored, reproducible searches (date ranges, article types).

For ClinicalTrials.gov use searching-clinicaltrials; this skill is for the published literature.

Quick start (real E-utilities calls)

Base URL: https://eutils.ncbi.nlm.nih.gov/entrez/eutils/. JSON for ESearch/ ESummary via retmode=json; EFetch returns text or XML (no JSON for PubMed).

python
import requests, time

BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
API_KEY = None   # set to your free NCBI key to get 10 req/s instead of 3

def _params(**kw):
    if API_KEY:
        kw["api_key"] = API_KEY
    return kw

def esearch(term: str, retmax: int = 50) -> dict:
    """Find PMIDs; usehistory=y stores them on the Entrez History server."""
    r = requests.get(f"{BASE}/esearch.fcgi", params=_params(
        db="pubmed", term=term, retmax=retmax,
        usehistory="y", retmode="json"), timeout=30)
    r.raise_for_status()
    res = r.json()["esearchresult"]
    return {"count": int(res["count"]), "ids": res["idlist"],
            "webenv": res["webenv"], "query_key": res["querykey"]}

def efetch_abstracts(webenv: str, query_key: str, retmax: int = 50) -> str:
    """Pull abstracts by reference to the stored result set (no ID list needed)."""
    r = requests.get(f"{BASE}/efetch.fcgi", params=_params(
        db="pubmed", WebEnv=webenv, query_key=query_key,
        retmax=retmax, rettype="abstract", retmode="text"), timeout=60)
    r.raise_for_status()
    return r.text

hits = esearch('("type 2 diabetes"[MeSH]) AND metformin AND 2023:2025[pdat]')
print(hits["count"], "papers")
abstracts = efetch_abstracts(hits["webenv"], hits["query_key"])

Equivalent cURL (search then fetch one PMID's abstract):

bash
curl "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=metformin&retmode=json"
curl "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=38000000&rettype=abstract&retmode=text"

ESummary for structured metadata

When you need titles/authors/journal/date as JSON (not the full abstract), use ESummary — it returns one record per ID:

python
def esummary(ids: list[str]) -> dict:
    r = requests.get(f"{BASE}/esummary.fcgi", params=_params(
        db="pubmed", id=",".join(ids), retmode="json"), timeout=30)
    r.raise_for_status()
    return r.json()["result"]   # keyed by PMID: title, pubdate, source, authors…

For PMC full text, repeat with db=pmc and EFetch rettype=""/retmode=xml (JATS XML). Respect each article's license before redistributing full text.

Workflow

  1. Build the query. Combine OpenMed-extracted terms with MeSH tags and field filters: "<disease>"[MeSH] AND <drug>[tiab] AND 2020:2025[pdat]. Use [tiab] (title/abstract), [au] (author), [pdat] (publication date).
  2. ESearch with usehistory=y to capture WebEnv + query_key and the count.
  3. Batch-fetch with EFetch/ESummary in pages of ≤ ~200 IDs (or by history), sleeping to stay under your rate limit.
  4. Parse abstracts/metadata; store PMID, title, journal, date, abstract text.
  5. NER the abstracts with openmed.analyze_text to extract diseases, drugs, genes, and oncology entities for downstream synthesis.
Show full SKILL.md (236 more words)Show less

Hand-off to / from OpenMed

  • OpenMed facts → query. openmed.analyze_text(note) yields Disease, Pharmaceutical, Genomics, and Oncology entities. Turn the top spans into the ESearch term (optionally grounded: ICD-10 label, RxNorm ingredient, gene symbol) to retrieve targeted evidence.
  • Abstracts → OpenMed. Feed fetched abstracts straight into openmed.analyze_text(abstract, model_name="disease_detection_superclinical") (or a Genomics/Oncology model) to structure the literature into entities for evidence tables or knowledge-graph edges.
  • Queries and abstracts are public literature, not PHI. Still run locally and never embed patient text in a search term.

Edge cases & gotchas

  • Rate limits. 3 req/s without a key, 10 with one — exceed it and NCBI returns HTTP 429. Add api_key, throttle, and retry with backoff. NCBI also requests a tool= and email= parameter identifying your application.
  • EFetch has no JSON for PubMed. Use retmode=text (human-readable) or retmode=xml (PubMedArticle XML) and parse XML for structured fields.
  • History expires. WebEnv/query_key are session-scoped — fetch promptly after searching, or re-run ESearch.
  • Large result sets. Page with retstart/retmax (or history) rather than pulling everything at once; cap total fetches.
  • MeSH lag. Very recent articles may not yet be MeSH-indexed — include [tiab] term variants so you do not miss them.
  • Full-text licensing. PMC full text carries per-article licenses; many are not redistributable. Store PMIDs/abstracts freely; check the license before republishing full text.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/mining-pubmed-literature of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Mining Pubmed Literature next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mining Pubmed Literature compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mining Pubmed Literature this skillmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
PubMed REST API Searchdavila7/claude-code-templates33k14 repos~3.9kAutomated safety check: PassMIT
Pubmed Databasejaechang-hits/SciAgent-Skills3741 repos~4.4kAutomated safety check: PassCC-BY-4.0
Molecular Review Workflowaipoch/medical-research-skills1.9k—~1.8kAutomated safety check: PassMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Academic Search and Citation RouterYuan1z0825/nature-skills47k—~884Automated safety check: PassApache-2.0

Similar skills

  • PubMed REST API Search

    davila7/claude-code-templates

    Searches PubMed directly through its E-utilities REST API, with guidance on Boolean and MeSH query syntax, batch retrieval and citation data.

    33k GitHub starsUsed in 14 repos~3.9k tokens
    Research & ScienceAuto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Research & ScienceAuto-check passed
  • Molecular Review Workflow

    aipoch/medical-research-skills

    Generates academic reviews for molecules in diseases using PubMed research.

    1.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Academic Search and Citation Router

    Yuan1z0825/nature-skills

    Finds papers across literature sources, verifies and converts citations, builds MeSH strategies and audits independent citations of a paper.

    47k GitHub stars~884 tokensUpdated today
    Research & ScienceAuto-check passed
  • Literature Review

    K-Dense-AI/scientific-agent-skills

    Runs systematic, scoping or narrative literature reviews across PubMed, arXiv, bioRxiv and Semantic Scholar, with citation checks and Markdown or PDF output.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check: notes

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Mining Pubmed Literature

What does Mining Pubmed Literature do?

Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora. Mining Pubmed Literature is an agent skill from maziyarpanahi/openmed. Searches and fetches PubMed and PMC via NCBI E-utilities (ESearch then EFetch/ESummary) to gather biomedical evidence and build text corpora.

When should I use Mining Pubmed Literature?

Mining Pubmed Literature fits situations like: the user wants citations for a condition; abstracts to summarize; meSH-based searches; A corpus of literature to run NER over.

How do I install Mining Pubmed Literature in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill mining-pubmed-literature -a claude-code`. Or copy the skill folder (skills/mining-pubmed-literature in maziyarpanahi/openmed) into .claude/skills/mining-pubmed-literature in your project. Claude Code loads it when a task matches its description.

How do I install Mining Pubmed Literature in Codex?

Run `npx skills add maziyarpanahi/openmed --skill mining-pubmed-literature -a codex`. Or copy the skill folder (skills/mining-pubmed-literature in maziyarpanahi/openmed) into .agents/skills/mining-pubmed-literature in your project. Codex loads it when a task matches its description.

Can I use Mining Pubmed Literature in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill mining-pubmed-literature -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mining-pubmed-literature, .gemini/skills/mining-pubmed-literature, .github/skills/mining-pubmed-literature and .opencode/skills/mining-pubmed-literature in your project.

What does Mining Pubmed Literature need to run?

Going by SKILL.md and its folder, Mining Pubmed Literature needs the command-line tools its instructions call (curl) and credentials named API_KEY. Our summary lists: Python 3; A credential in API_KEY.

Does Mining Pubmed Literature access the network?

SKILL.md names 5 domains. In commands or code: eutils.ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. As links in the text: ncbi.nlm.nih.gov, support.nlm.nih.gov, pubmed.ncbi.nlm.nih.gov and meshb.nlm.nih.gov. This is read from the text; nothing was executed.

Is Mining Pubmed Literature safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Mining Pubmed Literature use?

Mining Pubmed Literature is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mining Pubmed Literature use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mining Pubmed Literature?

Skills that share tags, products or a category with Mining Pubmed Literature: PubMed REST API Search (davila7/claude-code-templates, 33k stars), Pubmed Database (jaechang-hits/SciAgent-Skills, 374 stars), Molecular Review Workflow (aipoch/medical-research-skills, 1.9k stars) and Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mining Pubmed Literature?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.