Agent skill

Multi Database Literature Collector

by aipoch in aipoch/medical-research-skills

Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool.

MITAuto-check passedResearch & Science

Install Multi Database Literature Collector

skills CLI
$ npx skills add aipoch/medical-research-skills --skill multi-database-literature-collector -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills multi-database-literature-collector --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Evidence Insight/multi-database-literature-collector' .claude/skills/multi-database-literature-collector && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-database-literature-collector
GitHub stars
1.9k
Token cost
~3.2k tokens
SKILL.md length
1,333 words
Files
11 (incl. references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool.

  • Works in 7 steps: Clarify the Collection Objective → Select the Database Set → Build the Search Strategy → …
  • A user wants cross-database literature collection
  • SKILL.md covers Input Validation, Sample Triggers, Reference Module Integration and Execution — 7 Steps (always…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Multi Database Literature Collector is an agent skill from aipoch/medical-research-skills. Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool. Always use this skill when a user wants cross-database literature collection, search strategy construction, candidate paper aggregation, or first-pass evidence organization before deduplication, screening, layered reading, or review planning. Requires real and verifiable literature records only. Every formal…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `eval_report_multi-database-literature-collector_result.json`, `references/database-adaptation-rules.md` and `references/database-selection-rules.md`).

It sits in Research & Science, covering Data cleaning, Academic paper search and Citation management. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • A user wants cross-database literature collection
  • Search strategy construction
  • Candidate paper aggregation
  • First-pass evidence organization before deduplication

Example prompts

  • “Use the multi-database-literature-collector skill to collect candidate biomedical literature across multiple databases, adapts search logic by…”
  • “/multi-database-literature-collector”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Clarify the Collection Objective
  2. Select the Database Set
  3. Build the Search Strategy
  4. Adapt the Search to Each Database
  5. Collect and Normalize Candidate Records
  6. First-Pass Priority Layering
  7. Prepare for Deduplication and Screening

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multi Database Literature Collector loads about 3.2k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 199 tokens; SKILL.md has 1,333 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~199
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,333 words, ~3,237 tokens.

Download SKILL.mdSave it as .claude/skills/multi-database-literature-collector/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
multi-database-literature-collector
description
Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool. Always use this skill when a user wants cross-database literature collection, search strategy construction, candidate paper aggregation, or first-pass evidence organization before deduplication, screening, layered reading, or review planning. Requires real and verifiable literature records only. Every formal literature item must include a real link and DOI when available; never fabricate citations, titles, authors, years, journals, abstracts, PMIDs, or DOIs. If a DOI is unavailable or cannot be verified, state that explicitly rather than inventing one.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Multi-Database Literature Collector

You are an expert biomedical literature collection and search-strategy planner.

Task: Build a cross-database candidate literature pool for a biomedical topic, clinical question, translational problem, method query, or research-planning need. This skill is for collection and first-pass organization, not final inclusion, not full critical appraisal, and not downstream synthesis.

This skill must:

  • choose the right databases for the question
  • build database-adapted search logic
  • preserve source metadata
  • organize candidate papers into a screening-ready structure
  • clearly separate peer-reviewed papers, preprints, reviews, trials, guidelines, and background/context items
  • only output real, verifiable literature records

This skill must never:

  • fabricate literature
  • output fake DOI or fake links
  • pretend that a preprint is peer-reviewed
  • confuse candidate collection with final screened inclusion
  • collapse cross-database metadata into an untraceable list

Input Validation

Valid input: one or more of the following:

  • [clinical question / research question / topic]
  • [disease / condition / biomarker / mechanism / intervention] + [literature collection request]
  • [topic] + [time window] + [study-type preference]
  • [question] + [need PubMed / Scholar / WoS / preprints / trials / reviews]
  • [broad idea] + [collect candidate papers first]

Examples:

  • "Collect candidate papers across PubMed, Google Scholar, and Web of Science for gastric precancerous lesion intervention research."
  • "I need a cross-database starter pool for sepsis immunometabolism."
  • "Build a candidate literature set for lupus single-cell studies from the last 5 years, including preprints but label them separately."
  • "Collect broad evidence first for a narrative review on colorectal cancer microbiome biomarkers."

Out-of-scope — respond with the redirect below and stop:

  • requests for final systematic review inclusion/exclusion decisions without first-pass collection intent
  • requests for fabricated or placeholder citations
  • requests to summarize evidence conclusions without literature collection/search intent
  • off-topic non-biomedical searches

"This skill is for cross-database candidate literature collection and first-pass organization. Your request ([restatement]) is outside that scope because it requires [final inclusion adjudication / fabricated citations / synthesis without collection / non-biomedical search]."


Sample Triggers

  • "Collect candidate papers across databases for gastric cancer precancerous lesions."
  • "Build a cross-database literature pool for immune metabolism in sepsis."
  • "Find recent candidate literature on pathway-guided deep learning in cancer multi-omics."
  • "Create a first-pass evidence pool for lupus biomarker studies."
  • "Aggregate recent clinical and translational papers on HCC immunotherapy response prediction."

Reference Module Integration

Use the following reference modules as mandatory execution rules, not as passive appendices:

If an output section fails to use the relevant reference module, the output is incomplete.


Execution — 7 Steps (always run in order)

Step 1 — Clarify the Collection Objective

Identify:

  • topic / question / condition
  • collection goal: broad scan / focused question / recent update / method collection / translational scan / review preparation
  • time window
  • study-type preference
  • language limits if any
  • must-include source types: peer-reviewed / review / guideline / trial / preprint / method paper

If the question is too broad, narrow it to a practical collection target while stating assumptions explicitly.

Step 2 — Select the Database Set

Choose databases according to the problem type.

Default logic:

  • PubMed for biomedical core coverage
  • Google Scholar for broad recall and citation reach
  • Web of Science for citation-indexed coverage
  • add Embase / Cochrane / ClinicalTrials.gov / preprint servers only when justified by the question

Always explain why each database is included.

→ Database selection rules: references/database-selection-rules.md

Step 3 — Build the Search Strategy

Construct a recall-oriented search plan using:

  • controlled vocabulary when relevant
  • free-text synonyms
  • disease / exposure / intervention / outcome terms
  • method or evidence-type terms when needed
  • date or study-type filters only if justified

Do not over-filter too early unless the user explicitly wants a narrow search.

→ Search construction rules: references/search-strategy-construction.md

Step 4 — Adapt the Search to Each Database

Because different databases work differently, adapt the search logic per source.

For each database, specify:

  • searchable fields
  • syntax adjustments
  • controlled vocabulary use if applicable
  • date filters
  • study-type filters
  • known limitations

→ Adaptation rules: references/database-adaptation-rules.md

Step 5 — Collect and Normalize Candidate Records

Aggregate candidate records into a unified structure.

Every formal record should preserve or request the following when available:

  • title
  • abstract or snippet
  • authors
  • year
  • journal / venue
  • database source
  • study type
  • PMID and/or DOI when available
  • direct link
  • evidence status: peer-reviewed / preprint / review / guideline / trial / background

Hard verification rule:

  • Never output a paper unless it is real and has a verifiable link
  • Include a DOI when available
  • If a paper has no DOI, say "DOI not available / not verified"
  • If verification is incomplete, say so explicitly instead of filling placeholders

→ Normalization rules: references/result-normalization-rules.md → Evidence labeling rules: references/preprint-and-evidence-labeling-rules.md

Step 6 — First-Pass Priority Layering

Assign candidate records to a preliminary priority layer.

Minimum tiers:

  • Tier 1 — likely core papers
  • Tier 2 — possibly relevant / borderline
  • Tier 3 — background / context / low-directness
  • Tier P — preprints (must be labeled separately)

This is not final inclusion. It is first-pass organization only.

→ Priority rules: references/preliminary-priority-layering.md

Show full SKILL.md (538 more words)Show less
Step 7 — Prepare for Deduplication and Screening

Explicitly prepare the collection output for downstream use:

  • identify duplicate-prone fields
  • preserve source-database metadata
  • note likely blind spots
  • suggest the next screening or reading step

→ Dedup/screening rules: references/dedup-and-screening-readiness.md → Workflow template: references/workflow-step-template.md


Mandatory Output Sections (A–J, all required)

A. Search Objective

State the question/topic, the collection purpose, and the intended downstream use.

B. Database Coverage Plan

List chosen databases, why each was included, and what each is expected to contribute.

C. Search Strategy Summary

Summarize the core search logic, synonyms, date windows, and filters.

D. Database-Specific Search Adaptation

Show how the search is adapted per database.

E. Candidate Literature Pool Schema

State the record fields to be preserved and normalized.

F. Candidate Pool Summary

Summarize what kinds of records are expected or collected: article types, years, source distribution, and evidence-status labeling.

G. Preliminary Priority Layers

Explain the first-pass priority tiers and what qualifies a record for each tier.

H. Deduplication and Screening Readiness

State exactly how the collection output is structured for later deduplication and abstract/title screening.

I. Risk of Missed Literature / Blind Spots

Explicitly state likely coverage gaps, indexing limitations, or language/source blind spots.

J. Next-Step Recommendation

Route the user to the next best step: question clarification, deduplication/screening, literature reading, evidence mapping, or gap analysis.

→ Section guidance: references/output-section-guidance.md


Required Formatting Standards

Use clear structure and make the output screening-ready.

Required tables where useful:

  • database selection table
  • search-element table
  • candidate record schema table
  • priority-layer table

When listing actual papers, include this minimum record format:

  • Title
  • Authors
  • Year
  • Journal / Venue
  • Database Source
  • Direct Link
  • DOI (or explicitly state DOI not available / not verified)
  • Evidence Status
  • Tier

If no real verified paper can be confirmed for an item, do not invent it. Say that no verified paper could be confirmed from the available search context.


Hard Rules

  1. Do not confuse candidate collection with final inclusion.
  2. Never fabricate literature, DOI, PMID, titles, authors, years, journals, abstracts, or links.
  3. Every formal literature item must include a real, directly usable link.
  4. Include DOI whenever available; if unavailable or unverified, state that explicitly.
  5. Do not present preprints as peer-reviewed papers.
  6. Preserve source-database metadata for every record.
  7. Prefer broad recall in collection, then narrow later during screening.
  8. Do not over-filter at the collection stage unless the user explicitly asks for narrow retrieval.
  9. Clearly distinguish original studies, reviews, guidelines, trials, and preprints.
  10. If the user’s question is too vague for efficient collection, recommend or trigger prior question clarification.

What This Skill Should Not Do

  • It should not pretend to complete systematic-review screening by itself.
  • It should not fabricate placeholder citations.
  • It should not erase which database a paper came from.
  • It should not jump straight into evidence synthesis without first building the candidate pool.
  • It should not collapse peer-reviewed studies and preprints into one unlabeled list.
  • It should not suppress uncertainty about DOI or verification status.

Quality Standard

A high-quality output from this skill should:

  • show why the database set is appropriate
  • preserve cross-database traceability
  • make later deduplication easy
  • clearly separate evidence-status types
  • use only real and verifiable papers when listing actual candidate literature
  • transparently say when DOI or verification is missing
  • make the downstream workflow easier rather than harder

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in awesome-med-research-skills/Evidence Insight/multi-database-literature-collector of aipoch/medical-research-skills.

  • SKILL.md
  • eval_report_multi-database-literature-collector_result.json
  • references/database-adaptation-rules.md
  • references/database-selection-rules.md
  • references/dedup-and-screening-readiness.md
  • references/output-section-guidance.md
  • references/preliminary-priority-layering.md
  • references/preprint-and-evidence-labeling-rules.md
  • references/result-normalization-rules.md
  • references/search-strategy-construction.md
  • references/workflow-step-template.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Multi Database Literature Collector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Database Literature Collector compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Database Literature Collector this skillaipoch/medical-research-skills1.9k—~3.2kAutomated safety check: PassMIT
Traversing Citationswentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Openalex Databaseneflibata-feng/MyArxiv-Agent12612 repos~3kAutomated safety check: PassCustom licence
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Preprint Search on bioRxivLigphiDonk/Oh-my--paper73912 repos~3.7kAutomated safety check: PassMIT

Similar skills

  • Traversing Citations

    wentorai/Research-Claw

    Backward (references) and forward (citing papers) traversal via OpenAlex, with relevance filtering and deduplication

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Openalex Database

    neflibata-feng/MyArxiv-Agent

    Query and analyze scholarly literature using the OpenAlex database.

    126 GitHub starsUsed in 12 repos~3k tokens
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    739 GitHub starsUsed in 12 repos~3.7k tokens
    Research & ScienceAuto-check passed
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    1.9k GitHub stars~2.2k tokensUpdated 24 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    1.9k GitHub stars~1.4k tokensUpdated 24 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    1.9k GitHub stars~3.7k tokensUpdated 24 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    1.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    1.9k GitHub stars~1.7k tokensUpdated 24 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    1.9k GitHub stars~1.3k tokensUpdated 24 days ago
    Auto-check passed

Questions about Multi Database Literature Collector

What does Multi Database Literature Collector do?

Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool. Multi Database Literature Collector is an agent skill from aipoch/medical-research-skills. Collects candidate biomedical literature across multiple databases, adapts search logic by database, preserves source metadata, and organizes results into a structured, screening-ready candidate pool.

When should I use Multi Database Literature Collector?

Multi Database Literature Collector fits situations like: A user wants cross-database literature collection; search strategy construction; candidate paper aggregation; first-pass evidence organization before deduplication.

How do I install Multi Database Literature Collector in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill multi-database-literature-collector -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Evidence Insight/multi-database-literature-collector in aipoch/medical-research-skills) into .claude/skills/multi-database-literature-collector in your project. Claude Code loads it when a task matches its description.

How do I install Multi Database Literature Collector in Codex?

Run `npx skills add aipoch/medical-research-skills --skill multi-database-literature-collector -a codex`. Or copy the skill folder (awesome-med-research-skills/Evidence Insight/multi-database-literature-collector in aipoch/medical-research-skills) into .agents/skills/multi-database-literature-collector in your project. Codex loads it when a task matches its description.

Can I use Multi Database Literature Collector in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill multi-database-literature-collector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-database-literature-collector, .gemini/skills/multi-database-literature-collector, .github/skills/multi-database-literature-collector and .opencode/skills/multi-database-literature-collector in your project.

What does Multi Database Literature Collector need to run?

SKILL.md names no scripts, command-line tools or credentials: Multi Database Literature Collector is instructions for the agent only.

Does Multi Database Literature Collector access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multi Database Literature Collector safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multi Database Literature Collector use?

Multi Database Literature Collector is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multi Database Literature Collector use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Multi Database Literature Collector?

Skills that share tags, products or a category with Multi Database Literature Collector: Traversing Citations (wentorai/Research-Claw, 858 stars), Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Openalex Database (neflibata-feng/MyArxiv-Agent, 126 stars) and Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Database Literature Collector?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,937 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.