Split PDF
scunning1975/MixtapeTools
Download, split, and deeply read academic PDFs. An agent skill from scunning1975/MixtapeTools.
Download, split, and deeply read an academic PDF that is not available through Paperpile.
$ npx skills add flonat/flonat-research --skill split-pdf -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install flonat/flonat-research split-pdf --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/split-pdf .claude/skills/split-pdf && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .claude/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/flonat/flonat-research/tree/main/skills/split-pdfType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add flonat/flonat-research --skill split-pdf -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install flonat/flonat-research split-pdf --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/split-pdf .agents/skills/split-pdf && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .agents/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add flonat/flonat-research --skill split-pdf -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install flonat/flonat-research split-pdf --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/split-pdf .cursor/skills/split-pdf && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .cursor/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/flonat/flonat-research.git --path skills/split-pdf--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add flonat/flonat-research --skill split-pdf -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install flonat/flonat-research split-pdf --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/split-pdf .gemini/skills/split-pdf && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .gemini/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install flonat/flonat-research split-pdfInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add flonat/flonat-research --skill split-pdf -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/split-pdf .github/skills/split-pdf && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .github/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add flonat/flonat-research --skill split-pdf -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install flonat/flonat-research split-pdf --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/split-pdf .opencode/skills/split-pdf && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "split-pdf" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/split-pdf into .opencode/skills/split-pdf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "split-pdf", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
split-pdfDownload, split, and deeply read an academic PDF that is not available through Paperpile.
Split PDF is an agent skill from flonat/flonat-research. Download, split, and deeply read an academic PDF that is not available through Paperpile. Use when a long external PDF needs page-wise ingestion. For Paperpile items, use the Paperpile text-extraction route instead.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `methodology.md`).
It sits in Documents & Office, covering PDF. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(uv:*)Bash(uv*)Bash(curl*)Bash(wget*)Bash(mkdir*)Bash(ls*)Bash(rm*)ReadWriteEdit…and 4 more on the same allowed-tools line.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.commccombs.utexas.eduFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Split PDF loads about 3k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,296 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 1,296 words, ~2,979 tokens.
.claude/skills/split-pdf/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.CRITICAL RULE: Never read a full PDF. Never. Only read the 4-page split files, and only 3 splits at a time (~12 pages). Reading a full PDF will either crash the session with an unrecoverable "prompt too long" error — destroying all context — or produce shallow, hallucinated output. There are no exceptions.
The user wants you to read, review, or summarize an academic paper. The input is either:
./articles/smith_2024.pdf)"Gentzkow Shapiro Sinkinson 2014 competition newspapers")Important: You cannot search for a paper you don't know exists. The user MUST provide either a file path or a specific search query. If the user invokes this skill without specifying what paper to read, ask them. Do not guess.
Prefer Paperpile when possible. If the paper is in Paperpile, call paperpile get-pdf-text(citekey=KEY) directly — you get the full structured text without any splitting. Only fall through to this skill's page-split workflow when the PDF is NOT in Paperpile (preprints, referee materials, ad-hoc reading, third-party shared PDFs).
If a local file path is provided:
pdf-extract <file> --preflight) — a non-PASS verdict means page indices from this file are untrustworthy (truncated or silently repaired document).If a search query or paper title is provided:
Determine the download directory:
CLAUDE.md, data/, paper/, etc.): use ./articles/ in the project directory (create if needed).to-sort/downloads/ in the Task Management folder.Then:
scholarly scholarly-search "paper title" --json). Otherwise search the
publisher, DOI landing page, arXiv, or an institutional repository directly.CRITICAL: Always preserve the original PDF. The source PDF must NEVER be deleted, moved, or overwritten at any point in this workflow. The split files are derivatives; the original is the permanent artifact. Do not clean up, do not remove, do not tidy.
First, check for an existing extract. Look for <basename>_text.md in the same folder as the PDF.
If found, ask:
"An extract from a previous deep-read exists (
<basename>_text.md). Use it for this request, or re-read the PDF from scratch?"
<basename>_text.md and use it as the source notes — skip the rest of Steps 2 and 3 entirely.This prevents redundant re-reading of papers you have already processed. The _text.md file is a structured plain-text extraction far cheaper to read than re-processing PDF page images.
Second, check for existing splits. Compute the build directory:
import os
folder_path = os.path.dirname(os.path.abspath(pdf_path))
foldername = os.path.basename(folder_path)
pdf_basename = os.path.splitext(os.path.basename(pdf_path))[0]
build_dir = os.path.join(folder_path, foldername + '_build')
split_dir = os.path.join(build_dir, 'split_' + pdf_basename)If split_dir already exists and contains .pdf files, ask:
"Splits already exist for
<pdf-basename>(N chunks in<foldername>_build/split_<pdf-basename>/). Reuse existing splits, or re-split from scratch?"
split_dir.Otherwise, split. Create <foldername>_build/split_<pdf-basename>/ and run:
from PyPDF2 import PdfReader, PdfWriter
import os
def split_pdf(input_path, output_dir, pages_per_chunk=4):
os.makedirs(output_dir, exist_ok=True)
reader = PdfReader(input_path)
total = len(reader.pages)
prefix = os.path.splitext(os.path.basename(input_path))[0]
for start in range(0, total, pages_per_chunk):
end = min(start + pages_per_chunk, total)
writer = PdfWriter()
for i in range(start, end):
writer.add_page(reader.pages[i])
out_name = f"{prefix}_pp{start+1}-{end}.pdf"
out_path = os.path.join(output_dir, out_name)
with open(out_path, "wb") as f:
writer.write(f)
print(f"Split {total} pages into {-(-total // pages_per_chunk)} chunks in {output_dir}")If PyPDF2 is not installed: uv pip install PyPDF2.
Directory convention:
articles/ # any folder containing a PDF
├── smith_2024.pdf # original — NEVER DELETE
├── smith_2024_text.md # persistent extract — created after deep-read
└── articles_build/ # <foldername>_build/ — shared build folder
└── split_smith_2024/ # split_<pdf-basename>/
├── smith_2024_pp1-4.pdf
├── smith_2024_pp5-8.pdf
├── smith_2024_pp9-12.pdf
├── notes.md # working copy — source for _text.md
└── ...The build directory (<foldername>_build/) keeps split artifacts separate from source material and finished outputs. Multiple PDFs in the same folder share one build directory, each with its own split_<basename>/ subdirectory.
Read exactly 3 split files at a time (~12 pages). After each batch:
notes.md in the split subdirectory)"I have finished reading splits [X-Y] and updated the notes. I have [N] more splits remaining. Would you like me to continue with the next 3?"
Do NOT read ahead. Do NOT read all splits at once. The pause-and-confirm protocol is mandatory.
As you read, collect information along these dimensions and write them into notes.md:
These extract what a researcher needs to build on or replicate the work.
After all batches are complete, write the final notes to <basename>_text.md in the same folder as the source PDF:
articles/smith_2024_text.mdThen notify the user:
"Extract saved to
smith_2024_text.mdalongside the source PDF. Future requests on this paper can reuse it without re-reading."
This file is the persistent, reusable artifact. The notes.md in the build directory is the working copy. Both are kept — never delete either.
If the paper IS in Paperpile (user provides a citekey), skip the page-split workflow entirely:
paperpile get-item KEY --json — title, authors, abstract, affiliationspaperpile get-pdf-text KEY --json — full text from the attached PDFWrite the 8-dimension extraction directly to <basename>_text.md in the working directory. No splits needed.
This skill's page-split workflow is for PDFs NOT in Paperpile.
When split-pdf is invoked by another skill or workflow (any process that continues working after the PDF has been read), the PDF reading MUST run inside a subagent to prevent context bloat in the parent conversation.
Why: Each PDF page rendered by a client's PDF-reading capability produces image data in the conversation context. A 35-page PDF (9 chunks) can add 10-20MB of image data that accumulates permanently. After reading one or two large PDFs on top of prior work, the conversation can hit its request-size limit and become unrecoverable.
Pattern: The parent skill handles splitting (Step 2's Python script) in its own context — this is lightweight. Then it launches an Agent to perform all the reading:
Read PDF split files and produce structured extraction notes.
Split directory: <split_dir>
Files (read in this order, 3 at a time): <file_list>
Notes output: <notes_path> (working copy in split_dir)
Text output: <text_path> (persistent <basename>_text.md)
Process:
1. Read 3 PDF files at a time using the active client's PDF-reading surface
2. After each batch, update notes.md with extracted content
3. Extract along the 8 dimensions (research question, audience, method,
data, statistical methods, findings, contributions, replication feasibility)
4. Write the final structured extraction to <text_path>
Report when done: pages read, figures/tables found, one-sentence content summary.After the agent returns, the parent reads the output files (plain markdown, not PDF images) and continues its workflow.
Standalone invocations (user calls split-pdf directly) use the interactive protocol above with reads in the main conversation and the pause-and-confirm protocol.
paperpile get-pdf-text directly| Step | Action |
|---|---|
| Acquire | Use local PDF in place, or download to ./articles/ (in-project) / to-sort/downloads/ (ad-hoc) |
| Check cache | <basename>_text.md or existing splits — offer to reuse |
| Split | 4-page chunks into <foldername>_build/split_<pdf-basename>/ |
| Read | 3 splits at a time, pause after each batch |
| Notes | Update notes.md with 8-dimension extraction |
| Persist | Save final extract to <basename>_text.md alongside source PDF |
| Confirm | Ask user before continuing to next batch |
The in-place PDF handling, persistent _text.md extraction, split reuse, build directory convention, and agent isolation protocol are adapted from Scott Cunningham's MixtapeTools split-pdf skill (April 2026), which itself incorporated improvements from Ben Bentzin (McCombs School of Business, UT Austin). Structured Paperpile mode is the user's addition.
© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/split-pdf of flonat/flonat-research.
Open the folder on GitHubat commit da27600
Split PDF next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Split PDF this skillflonat/flonat-research | 145 | — | ~3k | Automated safety check: Pass | MIT | |
| Split PDFscunning1975/MixtapeTools | 469 | 2 repos | ~2.9k | Automated safety check: Pass | None | |
| Paper Interpretationdigoal/blog | 8.6k | — | ~1.5k | Automated safety check: Pass | GPL-2.0 | |
| Paper2slidesQuZhan51496/paper2anything | 468 | — | ~3.8k | Automated safety check: Notes | Apache-2.0 | |
| Paper LensYSQ-boop/paper-lens | 101 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Geng Academic Fraud Detectorwooly99/geng-academic-fraud-detector | 276 | — | ~970 | Automated safety check: Pass | None |
scunning1975/MixtapeTools
Download, split, and deeply read academic PDFs. An agent skill from scunning1975/MixtapeTools.
digoal/blog
从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…
QuZhan51496/paper2anything
Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.
YSQ-boop/paper-lens
Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.
wooly99/geng-academic-fraud-detector
学术论文打假检测器,致敬耿同学。分析学术论文 PDF,检测数据造假、图片复用/拼接、Western blot 操纵、统计异常等学术不端行为。当用户提供论文 PDF 要求"查重"、"打假"、"检测造假"、"论文分析"、"学术打假"时使用。
QuZhan51496/paper2anything
Convert academic papers (PDF) into conference posters (HTML/PNG).
flonat/flonat-research
Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.
flonat/flonat-research
Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.
flonat/flonat-research
Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.
flonat/flonat-research
Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.
flonat/flonat-research
Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.
flonat/flonat-research
Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.
Categories
Download, split, and deeply read an academic PDF that is not available through Paperpile. Split PDF is an agent skill from flonat/flonat-research. Download, split, and deeply read an academic PDF that is not available through Paperpile.
Split PDF fits situations like: A long external PDF needs page-wise ingestion; tasks that involve PDF.
Run `npx skills add flonat/flonat-research --skill split-pdf -a claude-code`. Or copy the skill folder (skills/split-pdf in flonat/flonat-research) into .claude/skills/split-pdf in your project. Claude Code loads it when a task matches its description.
Run `npx skills add flonat/flonat-research --skill split-pdf -a codex`. Or copy the skill folder (skills/split-pdf in flonat/flonat-research) into .agents/skills/split-pdf in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill split-pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/split-pdf, .gemini/skills/split-pdf, .github/skills/split-pdf and .opencode/skills/split-pdf in your project.
Going by SKILL.md and its folder, Split PDF needs the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(uv:*), Bash(uv*), Bash(curl*), Bash(wget*), Bash(mkdir*), Bash(ls*), Bash(rm*), Read, Write, Edit, WebSearch, WebFetch, Agent, Bash(paperpile*).
SKILL.md names 2 domains. As links in the text: github.com and mccombs.utexas.edu. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Split PDF is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Split PDF: Split PDF (scunning1975/MixtapeTools, 469 stars), Paper Interpretation (digoal/blog, 8.6k stars), Paper2slides (QuZhan51496/paper2anything, 468 stars) and Paper Lens (YSQ-boop/paper-lens, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.
Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.