Agent skill

Paper Parse Guide

by wentorai in wentorai/research-plugins

Deep dual-mode reading of academic papers from PDF or URL sources

MITAuto-check passedResearch & Science

Install Paper Parse Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill paper-parse-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins paper-parse-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/document/paper-parse-guide .claude/skills/paper-parse-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paper-parse-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
601 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Deep dual-mode reading of academic papers from PDF or URL sources

  • Works in 5 steps: Identity Card → Core Argument (1-2 sentences) → Methods Snapshot → …
  • Tasks that involve PDF
  • SKILL.md covers Overview, Paper Acquisition and Parsing, Mode A: Survey Reading and Mode B: Deep Analysis, plus 3 more sections
  • Calls docker and curl; reaches arxiv.org

What it does

Paper Parse Guide is an agent skill from wentorai/research-plugins. Deep dual-mode reading of academic papers from PDF or URL sources

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering PDF. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/paper-parse-guide”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identity Card
  2. Core Argument (1-2 sentences)
  3. Methods Snapshot
  4. Key Findings (3-5 bullets)
  5. Relevance Assessment

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org

    Also links to:

    • github.com
    • pymupdf.readthedocs.io
    • api.openalex.org
    • unpaywall.org
    • ccr.sigcomm.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paper Parse Guide loads about 2k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 601 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 601 words, ~1,979 tokens.

Download SKILL.mdSave it as .claude/skills/paper-parse-guide/SKILL.md (or your agent's skills folder).
name
paper-parse-guide
description
Deep dual-mode reading of academic papers from PDF or URL sources

Paper Parse Guide

Perform structured, dual-mode deep reading of academic papers from PDF files or URLs. Mode A provides a rapid overview suitable for screening during literature reviews. Mode B delivers exhaustive section-by-section analysis for papers central to your research.

Overview

Reading academic papers efficiently is a core research skill, yet the density and conventions of scholarly writing make it time-consuming. A typical researcher reads dozens of papers per week during a literature review phase, requiring different levels of depth for different papers. Some need only a quick scan to determine relevance; others demand line-by-line scrutiny of methods and results.

This skill implements a dual-mode reading system. Mode A (Survey Mode) extracts key metadata, the main argument, methods summary, and key findings in under two minutes of processing time. Mode B (Deep Analysis Mode) performs exhaustive section-by-section analysis including methodology critique, statistical evaluation, figure interpretation, and connection to broader literature.

Both modes begin by parsing the paper's structure from its PDF or HTML source, extracting clean text with section boundaries, figures, tables, equations, and references. The parsing pipeline handles the common challenges of academic PDFs: two-column layouts, footnotes, headers/footers, embedded equations, and supplementary materials.

Paper Acquisition and Parsing

Input Sources
SourceMethodNotes
Local PDFDirect file pathBest quality, no network needed
DOIResolve via CrossRef/UnpaywallAuto-fetches open access version
arXiv IDhttps://arxiv.org/pdf/{id}Always available
URLDirect downloadMay require institutional access
OpenAlex IDOpenAlex API + OA linkIncludes metadata
PDF Parsing Pipeline
python
from pathlib import Path
import fitz  # PyMuPDF

def parse_paper(pdf_path: str) -> dict:
    doc = fitz.open(pdf_path)
    sections = []
    current_section = {"title": "Header", "content": []}

    for page in doc:
        blocks = page.get_text("dict")["blocks"]
        for block in blocks:
            if block["type"] == 0:  # Text block
                for line in block["lines"]:
                    text = " ".join(span["text"] for span in line["spans"])
                    font_size = max(span["size"] for span in line["spans"])
                    is_bold = any("Bold" in span["font"] for span in line["spans"])

                    # Detect section headings
                    if font_size > 11 and is_bold:
                        if current_section["content"]:
                            sections.append(current_section)
                        current_section = {"title": text.strip(), "content": []}
                    else:
                        current_section["content"].append(text)

    sections.append(current_section)
    return {
        "title": extract_title(doc),
        "authors": extract_authors(doc),
        "sections": sections,
        "references": extract_references(doc),
        "page_count": len(doc)
    }
GROBID Integration

For higher-quality structural parsing, use GROBID (GeneRation Of BIbliographic Data):

bash
# Start GROBID server
docker run --rm -p 8070:8070 lfoppiano/grobid:0.8.0

# Parse a paper
curl -X POST "http://localhost:8070/api/processFulltextDocument" \
  -F "input=@paper.pdf" \
  -F "consolidateHeader=1" \
  -F "consolidateCitations=1" \
  -H "Accept: application/xml" \
  -o parsed_paper.xml

GROBID returns TEI XML with structured sections, author affiliations, parsed references, and figure/table captions. It handles two-column layouts, footnotes, and complex formatting better than simple text extraction.

Mode A: Survey Reading

Designed for rapid screening. Produces a structured summary in 5 components:

1. Identity Card
Title:      [Extracted title]
Authors:    [First author et al., year]
Venue:      [Journal/Conference name]
DOI:        [DOI if available]
Pages:      [Page count]
Type:       [Empirical / Theoretical / Review / Methods]
2. Core Argument (1-2 sentences)

Extract from abstract + introduction: What is the main claim?

3. Methods Snapshot
  • Study design (experimental, observational, computational, theoretical)
  • Sample/dataset description
  • Key techniques or models used
4. Key Findings (3-5 bullets)

Extract from results section and abstract.

5. Relevance Assessment
  • Relevance to current research question: High / Medium / Low
  • Methodological quality signal: sample size, controls, statistical rigor
  • Recommended action: Deep read / Cite only / Skip
Show full SKILL.md (233 more words)Show less

Mode B: Deep Analysis

Exhaustive section-by-section reading with critical evaluation.

Introduction Analysis
  • What gap in the literature does this paper address?
  • What is the stated research question or hypothesis?
  • How does the framing position the contribution?
Literature Review Evaluation
  • Which theoretical frameworks are invoked?
  • Are there notable omissions in cited literature?
  • How does the paper position itself relative to competing approaches?
Methodology Critique
  • Is the methodology appropriate for the research question?
  • Sample size and power analysis: reported? adequate?
  • Threats to internal and external validity
  • Reproducibility: are methods described in sufficient detail?
  • Statistical tests: appropriate? assumptions met?
Results Assessment
  • Do the results support the claims?
  • Effect sizes: reported? meaningful?
  • Confidence intervals vs. p-values
  • Figures and tables: do they accurately represent the data?
  • Any signs of p-hacking or selective reporting?
Discussion Evaluation
  • Are limitations adequately acknowledged?
  • Are alternative explanations considered?
  • Do the conclusions follow logically from the results?
  • Are implications overstated?
Reference Network
  • Extract all cited works with metadata
  • Identify seminal references (cited by many papers in this field)
  • Flag self-citations
  • Map citation clusters by topic

Output Formats

Structured Note (Markdown)
markdown
## [Paper Title] ([Year])

**Authors**: [Authors]
**Venue**: [Venue]
**DOI**: [DOI]

### Summary
[2-3 sentence summary]

### Key Contributions
1. [Contribution 1]
2. [Contribution 2]

### Methodology
[Methods description]

### Strengths
- [Strength 1]
- [Strength 2]

### Weaknesses
- [Weakness 1]
- [Weakness 2]

### Relevance to My Research
[How this paper connects to your work]

### Key Quotes
> "[Notable quote]" (p. X)

### References to Follow
- [Ref 1]: [Why relevant]
- [Ref 2]: [Why relevant]
BibTeX Entry

Automatically extract or generate a BibTeX entry for the parsed paper, including all available fields (author, title, journal, year, volume, pages, doi).

Batch Processing

For systematic reviews, process multiple papers in sequence:

python
papers = ["paper1.pdf", "paper2.pdf", "paper3.pdf"]
summaries = []

for pdf in papers:
    parsed = parse_paper(pdf)
    summary = mode_a_survey(parsed)
    summaries.append(summary)

# Generate comparison matrix
comparison = create_comparison_table(summaries,
    columns=["methods", "sample_size", "key_finding", "relevance"])

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/document/paper-parse-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Paper Parse Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paper Parse Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paper Parse Guide this skillwentorai/research-plugins2981 repos~2kAutomated safety check: PassMIT
Studyalaliqing/claude-paper3441 repos~2.5kAutomated safety check: NotesMIT
Summaryalaliqing/claude-paper3441 repos~2kAutomated safety check: NotesMIT
Paper ReadingEdwardxlai/easyread848—~567Automated safety check: PassMIT
Ref Downloaderltczding-gif/ref-downloader139—~5.9kAutomated safety check: PassMIT
Paper Readingsodalone/paper-reading-skill142—~1.3kAutomated safety check: PassNone

Similar skills

  • Study

    alaliqing/claude-paper

    A skill your agent uses when the user wants to read, study, analyze, or deeply understand a research paper (PDF).

    344 GitHub starsUsed in 1 repo~2.5k tokens
    Research & ScienceAuto-check: notes
  • Summary

    alaliqing/claude-paper

    Use this for a quick summary of a research paper's core ideas and key points.

    344 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check: notes
  • Paper Reading

    Edwardxlai/easyread

    论文共读:用 EasyRead(E:\CursorProject\easyread)这个本地工具读论文。用户给一篇 PDF 或 arXiv 编号时,导入文献库、翻译成中文(后台引擎或对话里的 agent 亲自译);之后边读边讨论,agent 读用户在页面上的笔记和提问,把回答、解释、原文核对提示追加到对应段落旁。用于…

    848 GitHub stars~567 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Ref Downloader

    ltczding-gif/ref-downloader

    A skill your agent uses when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B…

    139 GitHub stars~5.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Paper Reading

    sodalone/paper-reading-skill

    系统阅读、解释和批判性分析单篇 arXiv AI / 机器人论文,并生成一份主线清楚、核心内容不缺失、审稿证据可追溯的自包含 Markdown 报告。适用于 arXiv 链接、ID 或其 PDF 的精读、方法/理论/系统讲解、完整 Claim 审查、实验/证明/协议/代码证据核对、相关文献定位和复现或验证判断;默认执行完整精读,仅在用户明确要求速读时缩减覆盖范围。当前流水线不支持无法解析出…

    142 GitHub stars~1.3k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Paper Parse Guide

What does Paper Parse Guide do?

Deep dual-mode reading of academic papers from PDF or URL sources. Paper Parse Guide is an agent skill from wentorai/research-plugins.

When should I use Paper Parse Guide?

Paper Parse Guide fits situations like: tasks that involve PDF.

How do I install Paper Parse Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill paper-parse-guide -a claude-code`. Or copy the skill folder (skills/tools/document/paper-parse-guide in wentorai/research-plugins) into .claude/skills/paper-parse-guide in your project. Claude Code loads it when a task matches its description.

How do I install Paper Parse Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill paper-parse-guide -a codex`. Or copy the skill folder (skills/tools/document/paper-parse-guide in wentorai/research-plugins) into .agents/skills/paper-parse-guide in your project. Codex loads it when a task matches its description.

Can I use Paper Parse Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill paper-parse-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paper-parse-guide, .gemini/skills/paper-parse-guide, .github/skills/paper-parse-guide and .opencode/skills/paper-parse-guide in your project.

What does Paper Parse Guide need to run?

Going by SKILL.md and its folder, Paper Parse Guide needs the command-line tools its instructions call (docker and curl). Our summary lists: Python 3; Docker.

Does Paper Parse Guide access the network?

SKILL.md names 6 domains. In commands or code: arxiv.org; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, pymupdf.readthedocs.io, api.openalex.org, unpaywall.org and ccr.sigcomm.org. This is read from the text; nothing was executed.

Is Paper Parse Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paper Parse Guide use?

Paper Parse Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paper Parse Guide use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Paper Parse Guide?

Skills that share tags, products or a category with Paper Parse Guide: Study (alaliqing/claude-paper, 344 stars), Summary (alaliqing/claude-paper, 344 stars), Paper Reading (Edwardxlai/easyread, 848 stars) and Ref Downloader (ltczding-gif/ref-downloader, 139 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paper Parse Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.