PDF Explore
HughYau/AcademicForge
A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…
A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…
$ npx skills add JimLiu/science-skills --skill pdf-explore -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install JimLiu/science-skills pdf-explore --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf-explore .claude/skills/pdf-explore && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .claude/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/JimLiu/science-skills/tree/main/skills/pdf-exploreType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add JimLiu/science-skills --skill pdf-explore -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install JimLiu/science-skills pdf-explore --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/pdf-explore .agents/skills/pdf-explore && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .agents/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JimLiu/science-skills --skill pdf-explore -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install JimLiu/science-skills pdf-explore --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/pdf-explore .cursor/skills/pdf-explore && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .cursor/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/JimLiu/science-skills.git --path skills/pdf-explore--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add JimLiu/science-skills --skill pdf-explore -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install JimLiu/science-skills pdf-explore --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/pdf-explore .gemini/skills/pdf-explore && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .gemini/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install JimLiu/science-skills pdf-exploreInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add JimLiu/science-skills --skill pdf-explore -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/pdf-explore .github/skills/pdf-explore && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .github/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JimLiu/science-skills --skill pdf-explore -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install JimLiu/science-skills pdf-explore --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/pdf-explore .opencode/skills/pdf-explore && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-explore" agent skill from https://github.com/JimLiu/science-skills/tree/main/skills/pdf-explore into .opencode/skills/pdf-explore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-explore", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-exploreA skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…
PDF Explore is an agent skill from JimLiu/science-skills. Use this skill when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section, compare sections, find where a topic is discussed, read a value or label off a figure or chart, or find/list/extract every instance of something across the whole document (datasets, benchmarks, citations, figures, table rows, accession numbers — including appendices). Skip it only for a single lookup of 1–4 pages quoted in your…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `kernel.py`).
It sits in Documents & Office, covering PDF and Citation management. It works with pypdf and Python. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit fb309c3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Explore loads about 3.6k tokens when it runs. Until then it costs about 252 tokens; SKILL.md has 1,506 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from JimLiu/science-skills at commit fb309c3, republished under its Apache-2.0 licence (© JimLiu). 1,506 words, ~3,581 tokens.
.claude/skills/pdf-explore/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A 50-page PDF via read_file is ~200K tokens in context, and pages
loaded with read_file(pages=[...]) are dropped from context after one
turn — so multi-section synthesis turns into re-reading the same pages
over and over. And when the answer is "every page" (list all the
datasets / citations / figures / benchmarks mentioned anywhere in this
document), reading the whole thing page-by-page is the expensive way to
get it. This skill parses the PDF once in the Python kernel and runs one
cheap haiku call per page, in parallel, so you load only what
matters — or sweep every page without ever putting the pages in your own
context.
| when | returns | |
|---|---|---|
read_file(pages=[...]) (no skill) | a single lookup of 1–4 pages you will quote in your very next response | pages as vision blocks — dropped from context after one turn |
pdf_pages(path, pages=[...], mode="text") | you need several pages/sections at the same time — summaries, comparisons, anything where the answer draws on more than one range | [{page, text}, ...] — write to a file then read_file; stays in context like any tool output |
pdf_outline(path) | structured doc (paper, report, book) | [{page, heading, level}, ...] — a TOC |
pdf_scan(path, query, top_k) | semantic question, want the K most relevant pages | {hits: [{page, relevance, summary, text}], n_scanned, usage} |
pdf_extract(path, schema) | exhaustive list of X across the whole doc (datasets, citations, figures, table rows, entities) | [{page, data, usage}, ...] — then flatten + dedupe |
pdf_map(path, prompt) | unstructured doc (transcript, slide dump, compilation), or a free-text question of every page | {pages: [{page, text}], n_pages, usage} — every page's answer to prompt |
pdf_pages(mode="image", dpi=200) → host.view_image(img, crop=…) | read a small value, axis label, or legend off a figure | a high-res crop of the figure, auto-attached as vision |
Loading this skill auto-injects these into the Python kernel (from
kernel.py); call directly, no import. Don't check for or install
pypdfium2 first — host-managed installs seed it in the default python
environment, and if it's genuinely missing the first call raises with an
install recipe (on platform-managed hosts the default env rejects
installs — if the recipe's manage_packages call errors, create a domain
env with pypdfium2 and pillow and re-run there; pillow does the PNG
encoding for mode="image" and is not pulled in by the pypdfium2 wheel).
Go straight to the helper call.
Note: the default backend is pypdfium2 (Google PDFium; permissive Apache-2.0/BSD-3-Clause). PyMuPDF is honored as a fallback if already installed, but it is AGPL-3.0-licensed (commercial licenses available from Artifex): if you embed it in a network-accessible service, AGPL's source-sharing terms apply to that service.
For "summarize the methods" / "compare section 3 and section 5" / anything
where the answer draws on several page ranges at once, do not read the
ranges one read_file call at a time — each call's pages are dropped from
context before you finish, and you will loop re-reading them. Pull all
the pages you need in one python call, write them to a file, then
read_file that:
wanted = [5, 21,22,23,24,25, 62,63,64, 124,125,126] # from pdf_outline
with open("sections.txt", "w") as f:
for p in pdf_pages("paper.pdf", pages=wanted, mode="text"):
f.write(f"\n── page {p['page']} ──\n{p['text']}")
import os; print(f"wrote {os.path.getsize('sections.txt'):,} bytes")Then read_file(file_path="sections.txt") — or with offset=/limit= if
it's over 100KB — and write the answer from that. Don't print() the
page text directly: any python output over ~16KB is spilled to disk and
you'll be told to read_file it anyway, so printing a full chapter costs
two tool calls where writing + reading costs the same two without the
wasted preview. (For a quick look at ≤5 pages, printing is fine.)
~800 tokens/page of text vs ~4,000 tokens/page as vision — and you only pay
it once. Find the page numbers from pdf_outline (below) or the paper's
own table of contents first.
for e in pdf_outline("report.pdf"):
print(f"p{e['page']:>3} {' ' * (e['level'] - 1)}{e['heading']}")
# → then read_file(file_path="report.pdf", pages=[the section you want])Free and instant when the PDF has an embedded outline (most
LaTeX-compiled papers do). Falls back to a single batched LLM call on
text-layer ≤150pp docs ($0.001-0.003 total), or per-page LLM heading
extraction on scanned/>150pp ($0.002/page). For a semantic question the
outline doesn't obviously answer ("where do they discuss
limitations"), fall through to pdf_scan.
r = pdf_scan("paper.pdf", query="batch-effect correction methods", top_k=5)
for h in r["hits"]:
print(f"p{h['page']} {h['relevance']:.2f} {h['summary'] or h['text'][:100]}")
print(f"[{r['n_scanned']} pages scanned, "
f"{r['usage']['input_tokens']} in + {r['usage']['output_tokens']} out]")(summary is populated only when strategy="fanout" — the default
"auto" uses single-call comparative ranking on text-mode docs ≤150pp,
which is ~3× cheaper and ranks better but doesn't generate summaries.
Pass strategy="fanout" if you want them.)
Then load only those:
# Either: call read_file from the harness (attaches as vision)
# read_file(file_path="paper.pdf", pages=[h["page"] for h in r["hits"]])
# Or: print the text right here
for h in r["hits"]:
print(f"\n── page {h['page']} ──\n{h['text'][:2000]}")
# Or: render hit pages to cwd (PNGs auto-attach as vision next turn).
# Fine for layout/skimming — but a full page is too low-res to read
# small values off a figure; for that use the next recipe instead.
for p in pdf_pages("paper.pdf", mode="image",
pages=[h["page"] for h in r["hits"]], dpi=150):
import shutil; shutil.copy(p["image_path"], f"./hit_p{p['page']}.png")path can be a workspace path, a ~/-expanded path, or an artifact
version_id (resolved via host.artifact_path).
A full rendered page is too low-resolution to read small axis labels, legend text, or values off a dense multi-panel figure — the attach pipeline downsamples everything to ≤1568px, so the figure region ends up at a few hundred pixels no matter what DPI you render at. Render the page at high DPI, then crop the figure before attaching. The crop is both more legible and cheaper (a figure crop is ~400 vision tokens vs ~1,600 for the full page).
# 1. Find the figure's page (pdf_scan on the caption text, or pdf_extract
# with {"figures": [...]}, or you already know it).
# 2. Render that page at dpi=200 — high enough to crop into. The render
# lands in .cache/ so it is NOT auto-attached.
p = pdf_pages("paper.pdf", mode="image", pages=[5], dpi=200)[0]
# 3. Attach the full page once to locate the figure, OR go straight to a
# crop if the caption/position tells you where it is:
host.view_image(p["image_path"], crop=(x0, y0, x1, y1)) # pixels in the dpi=200 render
# → writes the crop to cwd; it auto-attaches at full resolution next turn.Crop to one panel at a time for multi-panel figures. Always crop from
the .cache/ render, not from a previously attached (downsampled) view.
For documents with no useful section structure (meeting transcripts, slide exports, multi-document compilations), get a 2-sentence summary of every page instead of a ranked subset:
m = pdf_map("transcript.pdf")
for p in m["pages"]:
print(f"p{p['page']}: {p['text']}")Then pick pages and read_file(pages=[...]). 100 pages → ~10K tokens in
context (vs ~400K if you embedded the whole PDF as vision, or ~90K as
extracted text). Nothing is filtered out, so there's no chance the
relevant page was missed. Measured: ~100 output tokens/page.
Pull the same fields from every page in parallel:
rows = pdf_extract("paper.pdf", {
"type": "object",
"properties": {
"figures": {"type": "array", "items": {"type": "object",
"properties": {"label": {"type": "string"},
"caption": {"type": "string"}}}},
},
"required": ["figures"],
})
figs = [(r["page"], f) for r in rows for f in (r["data"] or {}).get("figures", [])]Schemas that work well: {figures:[{label,caption}]} (figure index),
{citations:[str]} (bibliography — follow with one host.llm() call
to dedupe inline-marker noise), {section_headings:[str]} (TOC — this
is what pdf_outline does internally), {gene_symbols:[str]} (entity
lists), {rows:[{col1,col2,...}]} (table rows — whitespace-aligned
tables in the text layer parse cleanly; for rendered/image tables pass
mode="image").
Put the inclusion criterion in the schema's description — e.g.
"datasets on which results are actually reported on this page (in a table or the text), not datasets merely cited or mentioned". The
per-page model sees the full page and applies the criterion for you. If
you leave it out you'll end up re-reading pages afterwards to apply it
yourself, which costs more than the sweep did.
Schemas that don't: anything requiring judgment about "key" vs "all"
({key_claims:[str]} returns ~10/page, unusable). Per-page extraction
is recall-complete but precision-noisy. For ≲300 raw names, print the
sorted unique names with their page lists and dedupe/normalize them
yourself while writing the final answer — you have the whole list in
context, and a host.llm reduce call just hands you back a blob you
then have to parse. Reach for an LLM reduce only above that.
The sweep already read every page. Don't follow it with
read_file(pages=[...]) vision loads to "check for missed items" or to
re-examine specific hits — that re-spends the tokens the sweep just
saved. Call budget for an exhaustive-extraction job: 2 kernel calls
— (1) the sweep, printing unique names + page lists; (2) one batched
check that prints the cached text of every page you have a doubt about,
collected up front:
doubtful = {"DatasetA", "DatasetB", "DatasetC"} # decide ALL of them first
pages = sorted({p for n in doubtful for p in name_pages[n]})
for p in pdf_pages("paper.pdf", pages=pages):
print(f"\n── page {p['page']} ──\n{p['text']}")Then write the answer. The parse is cached so this is instant and free,
the text persists in your context (unlike read_file vision pages,
which vanish after one turn), and it's ~5× fewer tokens per page than
an image. If you're about to make a third kernel call to check one more
item, you're doing it wrong — fold it into call 2 or let it go.
read_file(file_path=..., pages=[...]) is fine — but only if you write
your answer on the very next turn, because the attached pages do not
survive past it.[p for p in pdf_pages(path) if "Harmony" in p["text"]]. pdf_scan
earns its cost on semantic queries ("where are the limitations
discussed") that keywords won't find.All helpers default to mode="auto": try text extraction; if pages
average < 80 extractable characters (scanned document, image-only slide
export), re-parse with page rendering so the LLM sees the image. You
don't need to set this. "text"/"image" force one or the other.
~800 input + 100 output tokens/page in text mode. The sweep helpers in
kernel.py pin PDF_DEFAULT_MODEL = "claude-haiku-4-5" explicitly on
every request — explicit model= wins over the kernel default, so the
sweep is haiku-priced (≈ $0.001–0.003/page depending on page length)
regardless of deployment config. (Ad-hoc host.llm() calls you make
yourself are different: bare calls use the Haiku-class kernel default
via [llm] kernel_default_model — check help(host.llm) if cost
matters there.) Don't
pass a heavier model= for extraction; heavier models cost 10-30×
more and add nothing to recall-complete per-page pulls. For a very
large document you can
scan a subset via pages=range(1, n, 3), but stride sampling can
miss a narrow relevant span between unrelated neighbors; prefer
pdf_outline → read the section you want when the document has
structure.
pdf_pages caches on (abs_path, mtime, mode, dpi) — a second
pdf_scan/pdf_map with a different query/prompt on the same file
skips re-parsing and re-rendering. Page renders land in
./.cache/pdf-explore/{sha8}-{mtime}/dpi{N}/p{NNN}.png (under
.cache/ so they do NOT auto-attach — copy ones you want seen to
./hit_pN.png).
© JimLiu, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/pdf-explore of JimLiu/science-skills.
Open the folder on GitHubat commit fb309c3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in JimLiu/science-skills, which our catalogue first saw on October 7, 2026.
PDF Explore next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Explore this skillJimLiu/science-skills | 227 | 2 repos | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| PDF ExploreHughYau/AcademicForge | 2.6k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| PDF ToolkitXiaomiMiMo/MiMo-Code | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Kimi PDFthvroyal/kimi-skills | 238 | — | ~1.9k | Automated safety check: Pass | None | |
| PDFNousResearch/hermes-agent | 252k | — | ~3.1k | Automated safety check: Pass | MIT |
HughYau/AcademicForge
A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…
XiaomiMiMo/MiMo-Code
Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.
pipeshub-ai/pipeshub-ai
Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.
thvroyal/kimi-skills
Professional PDF solution. An agent skill from thvroyal/kimi-skills.
NousResearch/hermes-agent
PDF files: create, read, merge, fill, OCR, edit text. An agent skill from NousResearch/hermes-agent.
espennilsen/pi
Read and extract content from PDF files — text, tables, metadata, and images.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
JimLiu/science-skills
Set up a compute environment on a remote provider so Claude Science jobs can run there.
JimLiu/science-skills
Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi.
JimLiu/science-skills
Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.
JimLiu/science-skills
Embed proteins with Meta AI's ESM-2 (fair-esm package). An agent skill from JimLiu/science-skills.
JimLiu/science-skills
Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab.
Categories
A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section…. PDF Explore is an agent skill from JimLiu/science-skills. Use this skill when the user has attached a PDF, paper, report, or other document and the answer needs content from more than one place in it: summarize the methods or any other section, compare sections, find where a topic is discussed, read a value or label off a figure or chart, or find/list/extract every instance of something across the whole document (datasets, benchmarks, citations, figures, table rows, accession numbers — including appendices).
PDF Explore fits situations like: the user has attached a PDF; other document and the answer needs content from more than one place in it: summarize the methods; any other section; compare sections.
Run `npx skills add JimLiu/science-skills --skill pdf-explore -a claude-code`. Or copy the skill folder (skills/pdf-explore in JimLiu/science-skills) into .claude/skills/pdf-explore in your project. Claude Code loads it when a task matches its description.
Run `npx skills add JimLiu/science-skills --skill pdf-explore -a codex`. Or copy the skill folder (skills/pdf-explore in JimLiu/science-skills) into .agents/skills/pdf-explore in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/science-skills --skill pdf-explore -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-explore, .gemini/skills/pdf-explore, .github/skills/pdf-explore and .opencode/skills/pdf-explore in your project.
Going by SKILL.md and its folder, PDF Explore needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
PDF Explore is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with PDF Explore: PDF Explore (HughYau/AcademicForge, 2.6k stars), PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars) and Kimi PDF (thvroyal/kimi-skills, 238 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
JimLiu (a GitHub user) maintains it in JimLiu/science-skills, which has 227 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on July 1, 2026.
Source: JimLiu/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.