Agent skill

Keyword Literature Harvester

by Jinze-Lee in Jinze-Lee/codex-skills-workbench

Local workflow that searches PubMed, Europe PMC, Crossref and OpenAlex for topic keywords, downloads accessible PDFs into a PDF-only folder and removes duplicates.

MITAuto-check passedResearch & Science

Install Keyword Literature Harvester

skills CLI
$ npx skills add Jinze-Lee/codex-skills-workbench --skill keyword-literature-download -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jinze-Lee/codex-skills-workbench keyword-literature-download --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jinze-Lee/codex-skills-workbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/keyword-literature-download .claude/skills/keyword-literature-download && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
keyword-literature-download
GitHub stars
109
Token cost
~952 tokens
SKILL.md length
427 words
Files
20 (incl. scripts, references)
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Local workflow that searches PubMed, Europe PMC, Crossref and OpenAlex for topic keywords, downloads accessible PDFs into a PDF-only folder and removes duplicates.

  • Works in 6 steps: Choose any output parent folder where… → Copy references/config_template.json and… → Run… → …
  • Collecting open-access papers on a set of topic keywords into one folder
  • SKILL.md covers What this skill does, Dependency model, Files in this skill and Recommended workflow, plus 2 more sections
  • Runs Python scripts from its folder

What it does

The skill bundles its own search and download scripts, so the target project needs no existing harvest code. Search scripts query PubMed and PMC, Europe PMC, Crossref and OpenAlex, and a candidate table is built first without deduplication. Legally accessible PDFs are then downloaded in parallel into a folder you choose with --pdf-output-dir or the download.pdf_output_dir setting, and that folder stays PDF-only.

HTML and XML helper files are cached separately under download_work/non_pdf_payloads/, and a second pass follows PDF links found in the saved pages. The run ends with a deduplicated PDF-only folder and a manifest. Large jobs can be rerun with the same run name and --skip-search so the APIs are not searched again.

To use it, pick an output parent folder, copy references/config_template.json and edit it for the topic keywords, then run scripts/run_keyword_harvest_no_dedup.py with --output-root, --config and --run-name, optionally tuning --download-workers. scripts/continue_download_and_dedup.py resumes pending downloads and builds the deduplicated set. A prompt template is included for handing the task to another agent.

When your agent uses it

  • Collecting open-access papers on a set of topic keywords into one folder
  • Resuming an interrupted download run and deduplicating the PDFs afterwards
  • Building a candidate table of papers from several scholarly APIs for review

Example prompts

  • “Harvest papers on perovskite solar cell stability into ./papers/perovskite using the config template.”
  • “Resume the pending downloads for the run named crispr-review and deduplicate the PDFs.”
  • “Search PubMed and OpenAlex for gut microbiome and depression and give me the candidate table.”

Requirements

  • Python to run the bundled scripts
  • Network access to PubMed, Europe PMC, Crossref and OpenAlex

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose any output parent folder where the new harvest run folder should be created.
  2. Copy references/config_template.json and edit it for the user's topic keywords.
  3. Run scripts/run_keyword_harvest_no_dedup.py with
  4. If the job is large, rerun with the same --run-name and --skip-search to avoid repeating API search.
  5. Run scripts/continue_download_and_dedup.py --run-root --retry-failed; it reuses the stored PDF folder unless --pdf-output-dir is supplied…
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7aa34a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Keyword Literature Harvester loads about 952 tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 427 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~952
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Jinze-Lee/codex-skills-workbench at commit 7aa34a5, republished under its MIT licence (© Jinze-Lee). 427 words, ~952 tokens.

Download SKILL.mdSave it as .claude/skills/keyword-literature-download/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
keyword-literature-download
description
Use when the user wants a reusable local workflow to search scholarly APIs for topic keywords, build a candidate table, rapidly download accessible PDFs into a PDF-only folder, optionally cache HTML/XML outside that folder for second-pass PDF discovery, and deduplicate the PDF files.

Keyword Research Harvest

Use this skill when the user gives a topic or keyword set and wants a local literature-harvesting workflow, not a one-off manual search.

What this skill does

  • Bundles its own literature_harvest/scripts stack, including search_pubmed.py, search_europepmc.py, search_crossref.py, search_openalex.py, merge_and_deduplicate.py, download_fulltexts.py, and harvest_utils.py.
  • Searches PubMed/PMC, Europe PMC, Crossref, and OpenAlex through the bundled pipeline.
  • Builds a no-dedup candidate table first.
  • Downloads legal PDFs in parallel for faster collection.
  • Lets the user specify the PDF output folder with --pdf-output-dir or download.pdf_output_dir.
  • Keeps the PDF output folder PDF-only; HTML/XML helper files are cached under download_work/non_pdf_payloads/.
  • Runs a second pass to chase PDF links from saved HTML pages.
  • Produces a PDF-only deduplicated file folder and manifest after downloading.

Dependency model

This skill is self-contained. The target project does not need to already contain:

  • literature_harvest/
  • search_pubmed.py
  • download_fulltexts.py

The bundled copies live under:

  • literature_harvest/scripts/

Files in this skill

  • scripts/run_keyword_harvest_no_dedup.py Use this to launch a new broad keyword harvest run into a new run folder.
  • scripts/continue_download_and_dedup.py Use this to resume pending downloads, try HTML-to-PDF second pass, and build a deduplicated download set.
  • literature_harvest/scripts/ Bundled search and download dependencies copied from a working local harvest stack so the skill is portable.
  • references/config_template.json Copy and edit this for the topic-specific query set and filtering terms.
  • references/prompt_template.md Reusable prompt for another AI/agent.
Show full SKILL.md (213 more words)Show less
  1. Choose any output parent folder where the new harvest run folder should be created.
  2. Copy references/config_template.json and edit it for the user's topic keywords.
  3. Run scripts/run_keyword_harvest_no_dedup.py with:
    • --output-root
    • --config
    • --run-name
    • optional --pdf-output-dir <folder> to put generated PDFs in a specific PDF-only folder
    • optional --download-workers <N> to tune parallel downloads
  4. If the job is large, rerun with the same --run-name and --skip-search to avoid repeating API search.
  5. Run scripts/continue_download_and_dedup.py --run-root <run-folder> --retry-failed; it reuses the stored PDF folder unless --pdf-output-dir is supplied again.
  6. Report:
    • candidate count
    • PDF downloaded count
    • true PDF count
    • HTML/XML cache count
    • remaining pending
    • deduplicated keep count

Constraints

  • Do not silently drop failed downloads.
  • Keep metadata even when download fails.
  • Prefer original research when the config asks for it, but do not hard-code one research domain.
  • Do not claim HTML/XML is PDF.
  • Never write HTML, XML, CSV, logs, manifests, or other helper files into the PDF output folder.
  • Verify a file starts with the PDF signature before saving it into the PDF output folder.
  • Treat second-pass cluster-like or support-like annotations as auxiliary only; downloading remains the primary task.

When to read references

  • Read references/config_template.json before preparing a new run config.
  • Read references/prompt_template.md when the user wants to hand this skill to another AI/agent.

© Jinze-Lee, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts, references) in skills/keyword-literature-download of Jinze-Lee/codex-skills-workbench.

  • SKILL.md
  • LICENSE
  • README.md
  • README_CN.md
  • README_EN.md
  • agents/openai.yaml
  • literature_harvest/scripts/config_search_terms.yaml
  • literature_harvest/scripts/download_fulltexts.py
  • literature_harvest/scripts/harvest_utils.py
  • literature_harvest/scripts/merge_and_deduplicate.py
  • literature_harvest/scripts/search_crossref.py
  • literature_harvest/scripts/search_europepmc.py
  • literature_harvest/scripts/search_openalex.py
  • literature_harvest/scripts/search_pubmed.py
  • references/config_template.json
  • references/prompt_template.md
  • references/source-note.md
  • … and 3 more

Open the folder on GitHubat commit 7aa34a5

Compare with similar skills

Keyword Literature Harvester next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Keyword Literature Harvester compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Keyword Literature Harvester this skillJinze-Lee/codex-skills-workbench109—~952Automated safety check: PassMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Academic Search and Citation RouterYuan1z0825/nature-skills47k—~884Automated safety check: PassApache-2.0
Literature ReviewK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: NotesMIT
PubMed REST API Searchdavila7/claude-code-templates33k14 repos~3.9kAutomated safety check: PassMIT
Literature ReviewNorman-bury/research-writing-skill3.4k—~2.2kAutomated safety check: NotesMIT

Similar skills

  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Academic Search and Citation Router

    Yuan1z0825/nature-skills

    Finds papers across literature sources, verifies and converts citations, builds MeSH strategies and audits independent citations of a paper.

    47k GitHub stars~884 tokensUpdated today
    Research & ScienceAuto-check passed
  • Literature Review

    K-Dense-AI/scientific-agent-skills

    Runs systematic, scoping or narrative literature reviews across PubMed, arXiv, bioRxiv and Semantic Scholar, with citation checks and Markdown or PDF output.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check: notes
  • PubMed REST API Search

    davila7/claude-code-templates

    Searches PubMed directly through its E-utilities REST API, with guidance on Boolean and MeSH query syntax, batch retrieval and citation data.

    33k GitHub starsUsed in 14 repos~3.9k tokens
    Research & ScienceAuto-check passed
  • Literature Review

    Norman-bury/research-writing-skill

    A skill your agent uses when writing literature review sections - guides searching, organizing, and synthesizing academic sources

    3.4k GitHub stars~2.2k tokensUpdated 4 mo ago
    Research & ScienceAuto-check: notes
  • Ma Search Bibliography

    htlin222/meta-pipe

    Conduct literature searches for meta-analysis using Python with uv, query PubMed and other databases, deduplicate results, and store round-based bibliographies with notes.

    139 GitHub stars~2.1k tokensUpdated 18 days ago
    Research & ScienceAuto-check: notes

More from Jinze-Lee/codex-skills-workbench

All 17 skills in this repo
  • Master Thesis Studio

    Jinze-Lee/codex-skills-workbench

    Guides the writing of a Chinese master's thesis and generates the Word file from a template, only when you invoke it by name.

    109 GitHub stars~3.5k tokensUpdated 5 mo ago
    Auto-check passed
  • Ggplot2 Richtext Fixes

    Jinze-Lee/codex-skills-workbench

    Fix ggplot2 superscript and richtext rendering issues using the bundled note.

    109 GitHub stars~293 tokensUpdated 5 mo ago
    Auto-check passed
  • Hypervolume Workflow

    Jinze-Lee/codex-skills-workbench

    Follow the bundled ecology hypervolume workflow. Use this skill only when the user explicitly invokes `$hypervolume-workflow`. Do not auto-select this skill…

    109 GitHub stars~290 tokensUpdated 5 mo ago
    Auto-check passed
  • Literature Synthesis Guide

    Jinze-Lee/codex-skills-workbench

    Follow the bundled literature-synthesis guide. An agent skill from Jinze-Lee/codex-skills-workbench.

    109 GitHub stars~297 tokensUpdated 5 mo ago
    Auto-check passed
  • Method Transfer Guide

    Jinze-Lee/codex-skills-workbench

    Follow the bundled method-transfer guide. An agent skill from Jinze-Lee/codex-skills-workbench.

    109 GitHub stars~282 tokensUpdated 5 mo ago
    Auto-check passed
  • Phylogeny Workflow

    Jinze-Lee/codex-skills-workbench

    Follow the bundled phylogeny workflow. An agent skill from Jinze-Lee/codex-skills-workbench.

    109 GitHub stars~299 tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Keyword Literature Harvester

What does Keyword Literature Harvester do?

Local workflow that searches PubMed, Europe PMC, Crossref and OpenAlex for topic keywords, downloads accessible PDFs into a PDF-only folder and removes duplicates. The skill bundles its own search and download scripts, so the target project needs no existing harvest code. Search scripts query PubMed and PMC, Europe PMC, Crossref and OpenAlex, and a candidate table is built first without deduplication.

When should I use Keyword Literature Harvester?

Keyword Literature Harvester fits situations like: collecting open-access papers on a set of topic keywords into one folder; resuming an interrupted download run and deduplicating the PDFs afterwards; building a candidate table of papers from several scholarly APIs for review.

How do I install Keyword Literature Harvester in Claude Code?

Run `npx skills add Jinze-Lee/codex-skills-workbench --skill keyword-literature-download -a claude-code`. Or copy the skill folder (skills/keyword-literature-download in Jinze-Lee/codex-skills-workbench) into .claude/skills/keyword-literature-download in your project. Claude Code loads it when a task matches its description.

How do I install Keyword Literature Harvester in Codex?

Run `npx skills add Jinze-Lee/codex-skills-workbench --skill keyword-literature-download -a codex`. Or copy the skill folder (skills/keyword-literature-download in Jinze-Lee/codex-skills-workbench) into .agents/skills/keyword-literature-download in your project. Codex loads it when a task matches its description.

Can I use Keyword Literature Harvester in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jinze-Lee/codex-skills-workbench --skill keyword-literature-download -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/keyword-literature-download, .gemini/skills/keyword-literature-download, .github/skills/keyword-literature-download and .opencode/skills/keyword-literature-download in your project.

What does Keyword Literature Harvester need to run?

Going by SKILL.md and its folder, Keyword Literature Harvester needs Python for the scripts in its folder. Our summary lists: Python to run the bundled scripts; Network access to PubMed, Europe PMC, Crossref and OpenAlex.

Does Keyword Literature Harvester access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Keyword Literature Harvester safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Keyword Literature Harvester use?

Keyword Literature Harvester is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Keyword Literature Harvester use?

About 952 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Keyword Literature Harvester?

Skills that share tags, products or a category with Keyword Literature Harvester: Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Academic Search and Citation Router (Yuan1z0825/nature-skills, 47k stars), Literature Review (K-Dense-AI/scientific-agent-skills, 48k stars) and PubMed REST API Search (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Keyword Literature Harvester?

Jinze-Lee (a GitHub user) maintains it in Jinze-Lee/codex-skills-workbench, which has 109 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on May 2, 2026.

Source: Jinze-Lee/codex-skills-workbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.