Agent skill

PDF Explore

by HughYau in HughYau/AcademicForge

A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…

Apache-2.0Auto-check passedDocuments & Office

Install PDF Explore

skills CLI
$ npx skills add HughYau/AcademicForge --skill pdf-explore -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HughYau/AcademicForge pdf-explore --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HughYau/AcademicForge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/claude-science/pdf-explore .claude/skills/pdf-explore && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-explore
GitHub stars
2.6k
Token cost
~3.1k tokens
SKILL.md length
1,414 words
Files
3
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…

  • The user has attached a PDF
  • SKILL.md covers Setup (any agent, no API key), Which helper, Recipe — pull the sections you… and Recipe — navigate by outline…, plus 7 more sections
  • Runs Python scripts from its folder; calls pip
  • Other document and the answer needs content from more than one place in it: summarize the methods

What it does

PDF Explore is an agent skill from HughYau/AcademicForge. Use this skill when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section, compare sections, find where a topic is discussed, read a value or label off a figure or chart, or find/list/extract every instance of something across the whole document (datasets, benchmarks, citations, figures, table rows, accession numbers — including appendices). Parses the PDF once with a deterministic Python kernel…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `kernel.py`).

It sits in Documents & Office, covering PDF and Citation management. It works with pypdf and Python. The repository describes itself as: One Forge, All Skills: A curated skill collection for academic writing and research. 点开即用,按需配置的一站式学术研究skills平台。 The licence is Apache-2.0.

When your agent uses it

  • The user has attached a PDF
  • Other document and the answer needs content from more than one place in it: summarize the methods
  • Any other section
  • Compare sections

Example prompts

  • “/pdf-explore”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 01b6d90. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Explore loads about 3.1k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 1,414 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HughYau/AcademicForge at commit 01b6d90, republished under its Apache-2.0 licence (© HughYau). 1,414 words, ~3,080 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-explore/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
pdf-explore
description
Use this skill when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section, compare sections, find where a topic is discussed, read a value or label off a figure or chart, or find/list/extract every instance of something across the whole document (datasets, benchmarks, citations, figures, table rows, accession numbers — including appendices). Parses the PDF once with a deterministic Python kernel: `pdf_pages` (pages as persistent text, or high-res images), `pdf_outline` (embedded TOC), `pdf_scan` (a lexical pre-filter that narrows a long doc to candidate pages), `pdf_grep` (regex sweep for exhaustive pattern extraction). You read the shortlist and do the relevance / summary / extraction judgment yourself. For PDF creation/manipulation, use reportlab/pypdf directly.
license
Apache-2.0

PDF Explore — navigate a PDF too big to embed

A 50-page PDF read in full is ~200K tokens of context. When the answer draws on several sections at once (summarize the methods; compare section 3 and section 5), or when the answer is "every page" (list all the datasets / citations / figures / benchmarks mentioned anywhere in this document), reading the whole thing page-by-page is the expensive way to get it. This skill parses the PDF once into persistent text with a deterministic Python kernel, then lets you narrow — by outline, by lexical scan, by regex — and read only the pages you actually need, reasoning over them yourself. Nothing you read vanishes: it is ordinary text and ordinary files.

Setup (any agent, no API key)

This is a pure skill — kernel.py is deterministic Python and you (the base model) do all the reasoning. There is no host runtime and no LLM API. Load the helpers once per session in a Python cell:

python
exec(open("skills/claude-science/pdf-explore/kernel.py").read())
# adjust the path to wherever this skill is installed

Nothing auto-loads it outside Claude Science. Then call the helpers directly (no import). If a helper is "not defined", you haven't exec'd kernel.py yet — go back and run the line above.

Dependencies: pip install pypdfium2 pillow (pillow does the PNG encoding for mode="image"; it is not pulled in by the pypdfium2 wheel).

Which helper

whenreturns
pdf_pages(path, pages=[...], mode="text")you need several pages/sections at the same time — summaries, comparisons, anything where the answer draws on more than one range[{page, text, n_chars}, ...] — persistent text; write to a file then read it
pdf_outline(path)structured doc (paper, report, book) with an embedded TOC[{page, heading, level}, ...] — the embedded outline, or [] if the PDF has none
pdf_scan(path, query, top_k)narrow a long doc to a handful of candidate pages for a query{hits: [{page, score, matched, text}], n_scanned} — a lexical pre-filter (no LLM); you read the shortlist and judge relevance
pdf_grep(path, pattern)exhaustive regex sweep (DOIs, accession ids, every "Table N", emails)[{page, matches, lines?}, ...] — every match with its page
pdf_pages(mode="image", dpi=200)read a small value, axis label, or legend off a figure[{page, image_path}, ...] — open the PNG with your agent's image tool

These come from kernel.py — load it via exec once per session (see Setup), then call directly. pdf_resolve(path) normalizes a path (a workspace path or a ~/-expanded path); the helpers call it internally, so path can be either form.

Note: the default backend is pypdfium2 (Google PDFium; permissive Apache-2.0/BSD-3-Clause). PyMuPDF is honored as a fallback if already installed, but it is AGPL-3.0-licensed (commercial licenses available from Artifex): if you embed it in a network-accessible service, AGPL's source-sharing terms apply to that service.

Recipe — pull the sections you need as persistent text (synthesis)

For "summarize the methods" / "compare section 3 and section 5" / anything where the answer draws on several page ranges at once, pull all the pages you need in one python call, write them to a file, then read that file:

python
wanted = [5, 21,22,23,24,25, 62,63,64, 124,125,126]  # from pdf_outline
with open("sections.txt", "w") as f:
    for p in pdf_pages("paper.pdf", pages=wanted, mode="text"):
        f.write(f"\n── page {p['page']} ──\n{p['text']}")
import os; print(f"wrote {os.path.getsize('sections.txt'):,} bytes")

Then read sections.txt with your agent's file-read tool (in chunks if it's large) and write the answer from that. It's ordinary text — one parse, and you reason over it directly. Don't print() a full chapter into the cell output: most agents spill large cell output to disk and make you re-read it anyway, so writing + reading a file costs the same two steps without the wasted preview. (For a quick look at ≤5 pages, printing is fine.)

Text is ~800 tokens/page vs ~4,000 tokens/page as vision, and you pay it once. Find the page numbers from pdf_outline (below) or the paper's own table of contents first.

Recipe — navigate by outline (try this first)

python
for e in pdf_outline("report.pdf"):
    print(f"p{e['page']:>3} {'  ' * (e['level'] - 1)}{e['heading']}")
# → then pull the section you want with pdf_pages(pages=[...])

Free and instant when the PDF has an embedded outline (most LaTeX-compiled papers do). pdf_outline reads the embedded TOC only — if the PDF has none it returns []. In that case build the outline yourself: pull the first handful of pages (or a stride sample of a long doc) as text with pdf_pages and pick out the headings by reading them. For a semantic question the outline doesn't obviously answer ("where do they discuss limitations"), fall through to pdf_scan.

Recipe — find the pages relevant to a query

pdf_scan ranks pages by lexical overlap with your query — it is a cheap pre-filter, not a relevance judgment. It narrows a long document to a handful of candidate pages; you then read those pages and decide which actually answer the question.

python
r = pdf_scan("paper.pdf", query="batch-effect correction methods", top_k=8)
for h in r["hits"]:
    print(f"p{h['page']}  score={h['score']:.2f}  matched={h['matched']}")
print(f"[{r['n_scanned']} pages scanned]")

Then read the shortlist's text and make the final call yourself:

python
for h in r["hits"]:
    print(f"\n── page {h['page']} ──\n{h['text'][:2000]}")

Keep the pages that genuinely address the query; discard lexical false positives (a page that merely says "batch" in another sense). Because the ranking is lexical, a synonym the query didn't use won't score — so lean on your own reading, broaden top_k if the shortlist looks thin, and if the term is one you can spell out, cross-check with pdf_grep or pdf_outline.

To skim hit pages as images (layout, tables) instead of text, render them and open the PNGs with your agent's image tool — but a full page is too low-res to read small values off a figure; for that use the next recipe.

python
for p in pdf_pages("paper.pdf", mode="image",
                   pages=[h["page"] for h in r["hits"]], dpi=150):
    print(p["image_path"])   # open each with your image tool
Show full SKILL.md (578 more words)Show less

Recipe — read a figure in detail

A full rendered page is often too low-resolution to read small axis labels, legend text, or values off a dense multi-panel figure. Render the page at high DPI, then crop the figure region before viewing it — the crop is both more legible and cheaper (fewer vision tokens than the whole page).

python
# 1. Find the figure's page (pdf_scan on the caption text, pdf_grep on the
#    figure label, pdf_outline, or you already know it).
# 2. Render that page at dpi=200 — high enough to crop into.
p = pdf_pages("paper.pdf", mode="image", pages=[5], dpi=200)[0]

# 3. Open p["image_path"] with your agent's image tool to locate the
#    figure, then crop it yourself with pillow before a close read:
from PIL import Image
Image.open(p["image_path"]).crop((x0, y0, x1, y1)).save("fig_crop.png")
#    (x0,y0,x1,y1) are pixels in the dpi=200 render.
# 4. Open fig_crop.png with your image tool.

Crop to one panel at a time for multi-panel figures. Always crop from the full-resolution render on disk, not from an already-downsampled view.

Recipe — list / extract every instance of X across the doc

Two shapes, depending on whether X has a pattern.

Pattern-shaped X (DOIs, accession numbers, "Figure N", emails, anything you can write a regex for) → pdf_grep, an exhaustive regex sweep over the parsed text:

python
hits = pdf_grep("paper.pdf", r"10\.\d{4,9}/[-._;()/:A-Za-z0-9]+")  # DOIs
for h in hits:
    print(f"p{h['page']}: {h['matches']}")
dois = sorted({m for h in hits for m in h["matches"]})

Returns [{page, matches, lines?}] — every match with its page, so you can build a page-indexed list. Exhaustive by construction (regex over every page) and free.

Judgment-shaped X (datasets results are actually reported on, key claims, figure captions, table rows) — no regex captures it, so you read and extract. Narrow first if you can (pdf_outline to the results section, or pdf_scan / pdf_grep on a likely term), then pull those pages as text and read them:

python
# e.g. every dataset RESULTS are reported on (not merely cited)
cand = {h["page"] for h in
        pdf_scan("paper.pdf", query="dataset benchmark evaluated", top_k=20)["hits"]}
with open("cand.txt", "w") as f:
    for p in pdf_pages("paper.pdf", pages=sorted(cand)):
        f.write(f"\n── page {p['page']} ──\n{p['text']}")

Read cand.txt and write the list yourself, applying the inclusion criterion ("reported on, not merely cited") as you read — that judgment is now just your own reading. If the document is short, skip the narrowing: pull every page's text into one file and read it straight through — recall-complete, with no filter that could miss anything. De-dupe and normalize as you write the final answer — you have the whole candidate list in context, so collapse "Dataset A" / "DatasetA" / "the A dataset" yourself.

The parse is cached (see Caching), so pulling a few more doubtful pages in a follow-up call is instant and free. Decide all your doubts up front and fetch them in one pdf_pages call rather than dribbling out one-page-at-a-time reads.

When NOT to use this skill

  • A single lookup of 1–4 pages you'll quote immediately: your agent's own PDF/page read tool is fine.
  • Literal keyword / pattern search: use pdf_grep, or filter the extracted text directly — [p for p in pdf_pages(path) if "Harmony" in p["text"]]. pdf_scan earns its keep on multi-word queries where you want a ranked shortlist to read, not on a single exact term.

Mode (scanned PDFs)

All helpers default to mode="auto": try text extraction; if pages average < 80 extractable characters (scanned document, image-only slide export), re-parse with page rendering so you can read the image. You don't need to set this. "text" / "image" force one or the other. Note that pdf_scan and pdf_grep operate on whatever text layer exists — a pure-image scan has none, so for those docs render the pages (mode="image") and read them yourself.

Cost & budget

Text is ~5× fewer tokens per page than a rendered image and it persists, so the winning pattern is: parse once, narrow (outline → scan → grep), and read only the pages you land on. For a very large document you can scan a subset via pages=range(1, n, 3), but stride sampling can miss a narrow relevant span between unrelated neighbors; prefer pdf_outline → read the section you want when the document has structure.

Caching

pdf_pages caches on (abs_path, mtime, mode, dpi) — a second pdf_scan / pdf_grep / pdf_pages with different arguments on the same file skips re-parsing and re-rendering. Pass cache=False to force a fresh parse. Page renders land in ./.cache/pdf-explore/{sha8}-{mtime}/dpi{N}/p{NNN}.png — copy or point your image tool at the ones you want to view.

© HughYau, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/claude-science/pdf-explore of HughYau/AcademicForge.

  • SKILL.md
  • .catalog_stamp
  • kernel.py

Open the folder on GitHubat commit 01b6d90

Compare with similar skills

PDF Explore next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Explore compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Explore this skillHughYau/AcademicForge2.6k—~3.1kAutomated safety check: PassApache-2.0
PDF ExploreJimLiu/science-skills2272 repos~3.6kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
Kimi PDFthvroyal/kimi-skills238—~1.9kAutomated safety check: PassNone
PDFNousResearch/hermes-agent252k—~3.1kAutomated safety check: PassMIT

Similar skills

  • PDF Explore

    JimLiu/science-skills

    A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…

    227 GitHub starsUsed in 2 repos~3.6k tokens
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Kimi PDF

    thvroyal/kimi-skills

    Professional PDF solution. An agent skill from thvroyal/kimi-skills.

    238 GitHub stars~1.9k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    NousResearch/hermes-agent

    PDF files: create, read, merge, fill, OCR, edit text. An agent skill from NousResearch/hermes-agent.

    252k GitHub stars~3.1k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Reader

    espennilsen/pi

    Read and extract content from PDF files — text, tables, metadata, and images.

    122 GitHub stars~1.6k tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed

More from HughYau/AcademicForge

  • Figure Composer

    HughYau/AcademicForge

    Compose one publication-grade multi-panel figure. An agent skill from HughYau/AcademicForge.

    2.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Narrative

    HughYau/AcademicForge

    Judge and reshape the STORY a paper's figures tell. An agent skill from HughYau/AcademicForge.

    2.6k GitHub stars~936 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about PDF Explore

What does PDF Explore do?

A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…. PDF Explore is an agent skill from HughYau/AcademicForge. Use this skill when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section, compare sections, find where a topic is discussed, read a value or label off a figure or chart, or find/list/extract every instance of something across the whole document (datasets, benchmarks, citations, figures, table rows, accession numbers — including appendices).

When should I use PDF Explore?

PDF Explore fits situations like: the user has attached a PDF; other document and the answer needs content from more than one place in it: summarize the methods; any other section; compare sections.

How do I install PDF Explore in Claude Code?

Run `npx skills add HughYau/AcademicForge --skill pdf-explore -a claude-code`. Or copy the skill folder (skills/claude-science/pdf-explore in HughYau/AcademicForge) into .claude/skills/pdf-explore in your project. Claude Code loads it when a task matches its description.

How do I install PDF Explore in Codex?

Run `npx skills add HughYau/AcademicForge --skill pdf-explore -a codex`. Or copy the skill folder (skills/claude-science/pdf-explore in HughYau/AcademicForge) into .agents/skills/pdf-explore in your project. Codex loads it when a task matches its description.

Can I use PDF Explore in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HughYau/AcademicForge --skill pdf-explore -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-explore, .gemini/skills/pdf-explore, .github/skills/pdf-explore and .opencode/skills/pdf-explore in your project.

What does PDF Explore need to run?

Going by SKILL.md and its folder, PDF Explore needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does PDF Explore access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Explore safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Explore use?

PDF Explore is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Explore use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Explore?

Skills that share tags, products or a category with PDF Explore: PDF Explore (JimLiu/science-skills, 227 stars), PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars) and Kimi PDF (thvroyal/kimi-skills, 238 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Explore?

HughYau (a GitHub user) maintains it in HughYau/AcademicForge, which has 2,589 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on August 30, 2026.

Source: HughYau/AcademicForge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.