Pullmd
AeternaLabsHQ/pullmd
Read any web page, document, or YouTube video as clean Markdown using PullMD.
Extract structured text, metadata, and references from academic PDFs
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wentorai/research-plugins grobid-pdf-parsing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .claude/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .claude/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wentorai/research-plugins grobid-pdf-parsing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .agents/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .agents/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wentorai/research-plugins grobid-pdf-parsing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .cursor/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .cursor/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wentorai/research-plugins.git --path skills/tools/document/grobid-pdf-parsing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wentorai/research-plugins grobid-pdf-parsing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .gemini/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .gemini/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wentorai/research-plugins grobid-pdf-parsingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .github/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .github/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wentorai/research-plugins grobid-pdf-parsing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tools/document/grobid-pdf-parsing .opencode/skills/grobid-pdf-parsing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "grobid-pdf-parsing" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/tools/document/grobid-pdf-parsing into .opencode/skills/grobid-pdf-parsing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grobid-pdf-parsing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
grobid-pdf-parsingExtract structured text, metadata, and references from academic PDFs
Grobid PDF Parsing is an agent skill from wentorai/research-plugins. Extract structured text, metadata, and references from academic PDFs
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering PDF and REST APIs. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockercurlgitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comtei-c.orgw3.orgAlso links to:
grobid.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Grobid PDF Parsing loads about 2.5k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 329 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 329 words, ~2,501 tokens.
.claude/skills/grobid-pdf-parsing/SKILL.md (or your agent's skills folder).Academic PDFs are the primary format for distributing research, yet extracting structured data from them remains challenging. PDFs encode visual layout, not semantic structure -- headings, paragraphs, equations, tables, and citations are all just positioned text and graphics. GROBID (GeneRation Of BIbliographic Data) is the leading open-source tool for parsing academic PDFs into structured XML/TEI format, extracting metadata, body text, references, and figures with high accuracy.
GROBID is used by major academic platforms including CORE, ResearchGate, and others for large-scale document processing. It combines machine learning models (CRF and deep learning) with heuristic rules to handle the diverse formatting of academic papers across publishers and disciplines.
This guide covers installing and running GROBID, using its REST API for batch processing, extracting specific elements (metadata, references, body sections), and integrating GROBID output into downstream workflows such as knowledge bases, systematic reviews, and literature analysis pipelines.
# Pull the latest GROBID image
docker pull grobid/grobid:0.8.1
# Run GROBID server
docker run --rm --init \
--ulimit core=0 \
-p 8070:8070 \
grobid/grobid:0.8.1
# GROBID is now running at http://localhost:8070
# Web console: http://localhost:8070/consolegit clone https://github.com/kermitt2/grobid.git
cd grobid
./gradlew clean install
./gradlew run# Process a single PDF and get TEI XML
curl -v --form input=@paper.pdf \
http://localhost:8070/api/processFulltextDocument \
-o paper.tei.xml
# With options
curl -v --form input=@paper.pdf \
--form consolidateHeader=1 \
--form consolidateCitations=1 \
--form includeRawCitations=1 \
http://localhost:8070/api/processFulltextDocument \
-o paper.tei.xml| Endpoint | Purpose | Input | Output |
|---|---|---|---|
/api/processFulltextDocument | Full paper parsing | TEI XML | |
/api/processHeaderDocument | Metadata only | TEI XML (header) | |
/api/processReferences | Reference parsing | TEI XML (refs) | |
/api/processCitation | Parse citation string | Text | TEI XML |
/api/processDate | Parse date string | Text | Structured date |
import requests
from pathlib import Path
class GrobidClient:
def __init__(self, base_url='http://localhost:8070'):
self.base_url = base_url
def process_fulltext(self, pdf_path, consolidate_header=True,
consolidate_citations=True):
"""Process a PDF and return TEI XML."""
url = f'{self.base_url}/api/processFulltextDocument'
files = {'input': open(pdf_path, 'rb')}
data = {
'consolidateHeader': '1' if consolidate_header else '0',
'consolidateCitations': '1' if consolidate_citations else '0',
}
response = requests.post(url, files=files, data=data)
response.raise_for_status()
return response.text
def process_header(self, pdf_path):
"""Extract only header metadata from PDF."""
url = f'{self.base_url}/api/processHeaderDocument'
files = {'input': open(pdf_path, 'rb')}
response = requests.post(url, files=files)
response.raise_for_status()
return response.text
def is_alive(self):
"""Check if GROBID server is running."""
try:
resp = requests.get(f'{self.base_url}/api/isalive')
return resp.status_code == 200
except requests.ConnectionError:
return False
# Usage
client = GrobidClient()
if client.is_alive():
tei_xml = client.process_fulltext('paper.pdf')
with open('paper.tei.xml', 'w') as f:
f.write(tei_xml)from lxml import etree
def parse_tei_metadata(tei_xml):
"""Extract title, authors, abstract from TEI XML."""
ns = {'tei': 'http://www.tei-c.org/ns/1.0'}
root = etree.fromstring(tei_xml.encode('utf-8'))
# Title
title_el = root.find('.//tei:titleStmt/tei:title', ns)
title = title_el.text if title_el is not None else ''
# Authors
authors = []
for author in root.findall('.//tei:sourceDesc//tei:author', ns):
forename = author.findtext('.//tei:forename', '', ns)
surname = author.findtext('.//tei:surname', '', ns)
if surname:
authors.append(f'{forename} {surname}'.strip())
# Abstract
abstract_el = root.find('.//tei:profileDesc/tei:abstract', ns)
abstract = ''.join(abstract_el.itertext()).strip() if abstract_el is not None else ''
# DOI
doi_el = root.find('.//tei:idno[@type="DOI"]', ns)
doi = doi_el.text if doi_el is not None else ''
return {
'title': title,
'authors': authors,
'abstract': abstract,
'doi': doi,
}def parse_tei_sections(tei_xml):
"""Extract structured sections from TEI XML body."""
ns = {'tei': 'http://www.tei-c.org/ns/1.0'}
root = etree.fromstring(tei_xml.encode('utf-8'))
sections = []
for div in root.findall('.//tei:body/tei:div', ns):
head = div.findtext('tei:head', '', ns).strip()
paragraphs = []
for p in div.findall('tei:p', ns):
text = ''.join(p.itertext()).strip()
if text:
paragraphs.append(text)
sections.append({
'heading': head,
'n': div.get('n', ''),
'paragraphs': paragraphs,
})
return sectionsdef parse_tei_references(tei_xml):
"""Extract structured references from TEI XML."""
ns = {'tei': 'http://www.tei-c.org/ns/1.0'}
root = etree.fromstring(tei_xml.encode('utf-8'))
refs = []
for bib in root.findall('.//tei:listBibl/tei:biblStruct', ns):
ref = {'id': bib.get('{http://www.w3.org/XML/1998/namespace}id', '')}
# Title
title_el = bib.find('.//tei:title[@level="a"]', ns)
if title_el is None:
title_el = bib.find('.//tei:title', ns)
ref['title'] = title_el.text if title_el is not None else ''
# Authors
ref['authors'] = []
for author in bib.findall('.//tei:author', ns):
name = f"{author.findtext('.//tei:forename', '', ns)} {author.findtext('.//tei:surname', '', ns)}".strip()
if name:
ref['authors'].append(name)
# Year
date_el = bib.find('.//tei:date[@type="published"]', ns)
ref['year'] = date_el.get('when', '') if date_el is not None else ''
# DOI
doi_el = bib.find('.//tei:idno[@type="DOI"]', ns)
ref['doi'] = doi_el.text if doi_el is not None else ''
refs.append(ref)
return refsfrom pathlib import Path
import json
from concurrent.futures import ThreadPoolExecutor
def batch_process(pdf_dir, output_dir, max_workers=4):
"""Process all PDFs in a directory using GROBID."""
client = GrobidClient()
pdf_dir = Path(pdf_dir)
output_dir = Path(output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
pdf_files = list(pdf_dir.glob('*.pdf'))
print(f"Processing {len(pdf_files)} PDFs...")
def process_one(pdf_path):
try:
tei = client.process_fulltext(str(pdf_path))
meta = parse_tei_metadata(tei)
refs = parse_tei_references(tei)
# Save TEI XML
tei_path = output_dir / f'{pdf_path.stem}.tei.xml'
tei_path.write_text(tei)
# Save structured JSON
json_path = output_dir / f'{pdf_path.stem}.json'
json_path.write_text(json.dumps({
'metadata': meta,
'references': refs,
'n_references': len(refs),
}, indent=2))
return pdf_path.name, 'success'
except Exception as e:
return pdf_path.name, f'error: {str(e)}'
with ThreadPoolExecutor(max_workers=max_workers) as executor:
results = list(executor.map(process_one, pdf_files))
for name, status in results:
print(f" {name}: {status}")
batch_process('papers/', 'parsed_output/')consolidateHeader=1 and consolidateCitations=1 cross-reference against Crossref for better metadata.© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/tools/document/grobid-pdf-parsing of wentorai/research-plugins.
Open the folder on GitHubat commit bf44b3c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.
Grobid PDF Parsing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Grobid PDF Parsing this skillwentorai/research-plugins | 298 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| PullmdAeternaLabsHQ/pullmd | 486 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | |
| PDFzai-org/ZCode | 7.7k | — | ~18k | Automated safety check: Notes | Proprietary | |
| Split PDFscunning1975/MixtapeTools | 474 | 2 repos | ~2.9k | Automated safety check: Pass | None | |
| Paper Interpretationdigoal/blog | 8.6k | — | ~1.5k | Automated safety check: Pass | GPL-2.0 | |
| Paper2slidesQuZhan51496/paper2anything | 450 | — | ~3.8k | Automated safety check: Notes | Apache-2.0 |
AeternaLabsHQ/pullmd
Read any web page, document, or YouTube video as clean Markdown using PullMD.
zai-org/ZCode
Professional PDF toolkit covering four production workflows: reports, creative visuals, academic LaTeX, and existing PDF processing.
scunning1975/MixtapeTools
Download, split, and deeply read academic PDFs. An agent skill from scunning1975/MixtapeTools.
digoal/blog
从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…
QuZhan51496/paper2anything
Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.
YSQ-boop/paper-lens
Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.
wentorai/research-plugins
Craft structured research abstracts that maximize clarity and journal acceptance
wentorai/research-plugins
Manage academic citations across BibTeX, APA, MLA, and Chicago formats
wentorai/research-plugins
Summarize academic papers with structured extraction of key elements
wentorai/research-plugins
Evidence-based study techniques for academic learning and retention
wentorai/research-plugins
Adjust writing tone and register for academic audiences and venues
wentorai/research-plugins
Academic translation, post-editing, and Chinglish correction guide
Categories
Extract structured text, metadata, and references from academic PDFs. Grobid PDF Parsing is an agent skill from wentorai/research-plugins.
Grobid PDF Parsing fits situations like: tasks that involve PDF; tasks that involve REST APIs.
Run `npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a claude-code`. Or copy the skill folder (skills/tools/document/grobid-pdf-parsing in wentorai/research-plugins) into .claude/skills/grobid-pdf-parsing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a codex`. Or copy the skill folder (skills/tools/document/grobid-pdf-parsing in wentorai/research-plugins) into .agents/skills/grobid-pdf-parsing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill grobid-pdf-parsing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grobid-pdf-parsing, .gemini/skills/grobid-pdf-parsing, .github/skills/grobid-pdf-parsing and .opencode/skills/grobid-pdf-parsing in your project.
Going by SKILL.md and its folder, Grobid PDF Parsing needs the command-line tools its instructions call (docker, curl and git). Our summary lists: Python 3; Docker.
SKILL.md names 4 domains. In commands or code: github.com, tei-c.org and w3.org; the agent is likely to contact these when it follows the instructions. As links in the text: grobid.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Grobid PDF Parsing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Grobid PDF Parsing: Pullmd (AeternaLabsHQ/pullmd, 486 stars), PDF (zai-org/ZCode, 7.7k stars), Split PDF (scunning1975/MixtapeTools, 474 stars) and Paper Interpretation (digoal/blog, 8.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.
Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.