Authoritative Data Harvester
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
Harvest metadata from open repositories using OAI-PMH protocol
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wentorai/research-plugins repository-harvesting-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .claude/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .claude/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wentorai/research-plugins repository-harvesting-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .agents/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .agents/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wentorai/research-plugins repository-harvesting-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .cursor/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .cursor/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wentorai/research-plugins.git --path skills/tools/scraping/repository-harvesting-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wentorai/research-plugins repository-harvesting-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .gemini/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .gemini/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wentorai/research-plugins repository-harvesting-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .github/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .github/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wentorai/research-plugins repository-harvesting-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tools/scraping/repository-harvesting-guide .opencode/skills/repository-harvesting-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "repository-harvesting-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/scraping/repository-harvesting-guide into .opencode/skills/repository-harvesting-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "repository-harvesting-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
repository-harvesting-guideHarvest metadata from open repositories using OAI-PMH protocol
Repository Harvesting Guide is an agent skill from wentorai/research-plugins. Harvest metadata from open repositories using OAI-PMH protocol
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Web scraping and Data cleaning. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
arxiv.orgopenarchives.orgpurl.orgexport.arxiv.orgncbi.nlm.nih.govoai.europeana.euapi.archives-ouvertes.frdblp.orgciteseerx.ist.psu.eduv2.sherpa.ac.ukFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Repository Harvesting Guide loads about 2.4k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 166 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 166 words, ~2,447 tokens.
.claude/skills/repository-harvesting-guide/SKILL.md (or your agent's skills folder).A skill for harvesting metadata from open access repositories using the OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting) protocol. Covers protocol fundamentals, building harvesters in Python, handling resumption tokens for large collections, metadata format parsing (Dublin Core, MARC, METS), selective harvesting by date and set, and integrating harvested data into research workflows.
OAI-PMH is a standardized protocol that allows metadata to be harvested from repository systems. It is the backbone of library interoperability and is supported by virtually every institutional repository, preprint server, and digital library worldwide.
OAI-PMH Architecture:
Data Providers (repositories):
- Expose metadata through a standardized HTTP interface
- Must support Dublin Core as minimum metadata format
- May support additional formats (MARC, MODS, DataCite, etc.)
- Examples: arXiv, PubMed Central, DSpace repositories,
EPrints, institutional repositories
Service Providers (harvesters):
- Send HTTP requests to data providers
- Collect, aggregate, and index metadata
- Build search services, union catalogs, analytics
- Examples: BASE (Bielefeld), CORE, OpenDOAR
Protocol Version: 2.0 (current, since 2002)
Transport: HTTP GET or POST
Response format: XML
Base URL example: https://arxiv.org/oai2OAI-PMH defines exactly six request types (verbs):
1. Identify
Purpose: Describe the repository
URL: baseURL?verb=Identify
Returns: repository name, admin email, earliest datestamp,
granularity, compression support
2. ListMetadataFormats
Purpose: List available metadata formats
URL: baseURL?verb=ListMetadataFormats
Returns: format prefixes (oai_dc, marc21, datacite, etc.)
Optional: identifier parameter to check formats for one record
3. ListSets
Purpose: List available sets (collections/categories)
URL: baseURL?verb=ListSets
Returns: set names and specs for selective harvesting
Example sets: physics:hep-th, cs:AI, math:AG
4. ListIdentifiers
Purpose: List record identifiers (headers only, no metadata)
URL: baseURL?verb=ListIdentifiers&metadataPrefix=oai_dc
Optional: from, until, set parameters
Returns: identifiers, datestamps, set memberships
5. ListRecords
Purpose: Harvest full metadata records
URL: baseURL?verb=ListRecords&metadataPrefix=oai_dc
Optional: from, until, set parameters
Returns: complete metadata records in requested format
6. GetRecord
Purpose: Retrieve a single record by identifier
URL: baseURL?verb=GetRecord&identifier=oai:arxiv:2301.00001
&metadataPrefix=oai_dc
Returns: one complete metadata recordimport requests
import xml.etree.ElementTree as ET
import time
OAI_NS = "http://www.openarchives.org/OAI/2.0/"
DC_NS = "http://purl.org/dc/elements/1.1/"
def harvest_records(base_url, metadata_prefix="oai_dc",
from_date=None, until_date=None,
set_spec=None):
"""
Harvest all records from an OAI-PMH endpoint.
Handles resumption tokens for paginated results.
Args:
base_url: OAI-PMH base URL
metadata_prefix: metadata format (default: oai_dc)
from_date: selective harvest start (YYYY-MM-DD)
until_date: selective harvest end (YYYY-MM-DD)
set_spec: restrict to a specific set
"""
params = {
"verb": "ListRecords",
"metadataPrefix": metadata_prefix,
}
if from_date:
params["from"] = from_date
if until_date:
params["until"] = until_date
if set_spec:
params["set"] = set_spec
all_records = []
request_count = 0
while True:
response = requests.get(base_url, params=params, timeout=30)
response.raise_for_status()
request_count += 1
root = ET.fromstring(response.content)
# Parse records from this page
records = root.findall(
f".//{{{OAI_NS}}}record"
)
for record in records:
parsed = parse_dublin_core(record)
if parsed:
all_records.append(parsed)
# Check for resumption token
token_elem = root.find(
f".//{{{OAI_NS}}}resumptionToken"
)
if token_elem is not None and token_elem.text:
params = {
"verb": "ListRecords",
"resumptionToken": token_elem.text,
}
# Polite delay between requests
time.sleep(2)
else:
break
print(f"Harvested {len(all_records)} records "
f"in {request_count} requests")
return all_records
def parse_dublin_core(record_element):
"""
Parse a Dublin Core metadata record into a dictionary.
"""
header = record_element.find(f"{{{OAI_NS}}}header")
metadata = record_element.find(f"{{{OAI_NS}}}metadata")
if header is None or metadata is None:
return None
# Check if record is deleted
status = header.get("status", "")
if status == "deleted":
return None
identifier = header.findtext(f"{{{OAI_NS}}}identifier", "")
datestamp = header.findtext(f"{{{OAI_NS}}}datestamp", "")
dc = metadata.find(f".//{{{DC_NS}}}../")
result = {
"oai_identifier": identifier,
"datestamp": datestamp,
"title": find_dc_text(metadata, "title"),
"creator": find_dc_all(metadata, "creator"),
"subject": find_dc_all(metadata, "subject"),
"description": find_dc_text(metadata, "description"),
"date": find_dc_text(metadata, "date"),
"type": find_dc_text(metadata, "type"),
"identifier": find_dc_all(metadata, "identifier"),
"language": find_dc_text(metadata, "language"),
"rights": find_dc_text(metadata, "rights"),
}
return result
def find_dc_text(metadata, element_name):
"""Find first Dublin Core element text."""
elem = metadata.find(f".//{{{DC_NS}}}{element_name}")
return elem.text if elem is not None else ""
def find_dc_all(metadata, element_name):
"""Find all values of a Dublin Core element."""
elems = metadata.findall(f".//{{{DC_NS}}}{element_name}")
return [e.text for e in elems if e.text]Incremental harvesting strategy:
First harvest: Get everything
from_date = None (or repository's earliestDatestamp)
until_date = today
Subsequent harvests: Get only new/modified records
from_date = last_harvest_date
until_date = today
Date granularity:
- Day-level: YYYY-MM-DD (most common)
- Second-level: YYYY-MM-DDThh:mm:ssZ (some repositories)
- Check the Identify response for supported granularity
Important: OAI-PMH datestamps reflect the date the METADATA
was last modified, not the publication date. A record edited
yesterday to fix a typo will appear in a harvest with
from=yesterday, even if the paper was published in 2015.Common set structures by repository type:
arXiv:
physics, physics:hep-th, cs, cs:AI, math, math:AG, etc.
DSpace repositories:
com_12345_1 (community), col_12345_2 (collection)
Hierarchical: department -> collection
PubMed Central:
By journal: pmc-journal-name
By funder: pmc-funder-name
Strategy:
1. Call ListSets to see available sets
2. Identify sets relevant to your research topic
3. Harvest only those sets to reduce data volume
4. Store the set membership for each recordQuality problems in harvested metadata:
1. Duplicate records:
- Same paper in multiple repositories
- Same paper in multiple sets within one repository
- Solution: Deduplicate by DOI, then by title similarity
2. Incomplete metadata:
- Missing abstracts (very common)
- Missing author identifiers
- Missing dates or using inconsistent date formats
- Solution: Enrich with Crossref or OpenAlex lookups
3. Encoding issues:
- Non-UTF-8 characters in older repositories
- HTML entities in text fields
- Solution: Normalize encoding, strip HTML tags
4. Inconsistent formats:
- Dates as "2023", "2023-01", "2023-01-15", "January 2023"
- Author names as "Smith, John" vs "John Smith" vs "J. Smith"
- Solution: Parse and normalize to canonical formatsMajor repositories with OAI-PMH support:
arXiv: https://export.arxiv.org/oai2
PubMed Central: https://www.ncbi.nlm.nih.gov/pmc/oai/oai.cgi
Europeana: https://oai.europeana.eu/oai
HAL (France): https://api.archives-ouvertes.fr/oai/hal
DBLP: https://dblp.org/oai
CiteSeerX: https://citeseerx.ist.psu.edu/oai2
To find more endpoints:
- OpenDOAR directory: https://v2.sherpa.ac.uk/opendoar/
- ROAR (Registry of Open Access Repositories)
- BASE (Bielefeld Academic Search Engine) source listOAI-PMH harvesting remains the most reliable method for building comprehensive metadata collections from open repositories. While newer APIs like ResourceSync and Signposting offer richer functionality, OAI-PMH's universal adoption and simplicity make it the practical choice for most academic metadata collection tasks.
© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/tools/scraping/repository-harvesting-guide of wentorai/research-plugins.
Open the folder on GitHubat commit bf44b3c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.
Repository Harvesting Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Repository Harvesting Guide this skillwentorai/research-plugins | 298 | 1 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Authoritative Data Harvesteryushui2022/MathModel-Skill | 454 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Cashclaw Data Scraperertugrulakben/cashclaw | 303 | — | ~3k | Automated safety check: Pass | MIT | |
| Data Cleaningericrisco/rsc-harness | 180 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Glue DiagnosticsKilo-Org/kilo-marketplace | 190 | — | ~2k | Automated safety check: Pass | MIT | |
| Minerals Web Ingestlamm-mit/scienceclaw | 246 | — | ~512 | Automated safety check: Pass | Apache-2.0 |
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
ertugrulakben/cashclaw
Extracts structured data from websites and APIs, delivering clean datasets in multiple formats.
ericrisco/rsc-harness
A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…
Kilo-Org/kilo-marketplace
A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…
lamm-mit/scienceclaw
Ingest and normalize web pages for critical-minerals intelligence, with optional Firecrawl fetching, deduplication manifest, and JSONL export
trpc-group/trpc-agent-go
Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.
wentorai/research-plugins
Craft structured research abstracts that maximize clarity and journal acceptance
wentorai/research-plugins
Manage academic citations across BibTeX, APA, MLA, and Chicago formats
wentorai/research-plugins
Summarize academic papers with structured extraction of key elements
wentorai/research-plugins
Evidence-based study techniques for academic learning and retention
wentorai/research-plugins
Adjust writing tone and register for academic audiences and venues
wentorai/research-plugins
Academic translation, post-editing, and Chinglish correction guide
Categories
Harvest metadata from open repositories using OAI-PMH protocol. Repository Harvesting Guide is an agent skill from wentorai/research-plugins.
Repository Harvesting Guide fits situations like: tasks that involve Web scraping; tasks that involve Data cleaning.
Run `npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a claude-code`. Or copy the skill folder (skills/tools/scraping/repository-harvesting-guide in wentorai/research-plugins) into .claude/skills/repository-harvesting-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a codex`. Or copy the skill folder (skills/tools/scraping/repository-harvesting-guide in wentorai/research-plugins) into .agents/skills/repository-harvesting-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill repository-harvesting-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/repository-harvesting-guide, .gemini/skills/repository-harvesting-guide, .github/skills/repository-harvesting-guide and .opencode/skills/repository-harvesting-guide in your project.
SKILL.md names no scripts, command-line tools or credentials: Repository Harvesting Guide is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 10 domains. In commands or code: arxiv.org, openarchives.org, purl.org, export.arxiv.org, ncbi.nlm.nih.gov, oai.europeana.eu, api.archives-ouvertes.fr, dblp.org, citeseerx.ist.psu.edu and v2.sherpa.ac.uk; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Repository Harvesting Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Repository Harvesting Guide: Authoritative Data Harvester (yushui2022/MathModel-Skill, 454 stars), Cashclaw Data Scraper (ertugrulakben/cashclaw, 303 stars), Data Cleaning (ericrisco/rsc-harness, 180 stars) and Glue Diagnostics (Kilo-Org/kilo-marketplace, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.
Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.