Agent skill

Scanned PDF To Epub

by longhaiqwe in longhaiqwe/scanned-pdf-to-epub

A skill your agent uses when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or…

MITAuto-check passedDocuments & Office

Install Scanned PDF To Epub

skills CLI
$ npx skills add longhaiqwe/scanned-pdf-to-epub --skill scanned-pdf-to-epub -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install longhaiqwe/scanned-pdf-to-epub scanned-pdf-to-epub --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scanned-pdf-to-epub
GitHub stars
161
Token cost
~1.9k tokens
SKILL.md length
984 words
Files
10 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or…

  • Works in 6 steps: Inspect the PDF. → Build a sample before full conversion. → OCR text with coordinates. → …
  • Scanned Chinese PDFs
  • SKILL.md covers Overview, Environment Preparation, Workflow and Common Mistakes, plus 1 more section
  • Runs Python scripts from its folder

What it does

Scanned PDF To Epub is an agent skill from longhaiqwe/scanned-pdf-to-epub. Use when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or WeRead/微信读书 compatibility matters.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `README.md`, `agents/openai.yaml` and `references/environment-setup.md`).

It sits in Documents & Office, covering PDF. It works with Linux and macOS. The repository describes itself as: Cross-platform skill for converting scanned Chinese PDFs into WeRead-friendly text EPUBs. The licence is MIT.

When your agent uses it

  • Scanned Chinese PDFs
  • Image-only books need text EPUB conversion on macOS
  • Especially when OCR text
  • Original underlined proper names

Example prompts

  • “/scanned-pdf-to-epub”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Inspect the PDF.
  2. Build a sample before full conversion.
  3. OCR text with coordinates.
  4. Preserve original underlined proper names.
  5. Choose WeRead marking strategy.
  6. Package and validate EPUB.

What it can do on your machine

Read from SKILL.md and the folder at commit 0e95243. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scanned PDF To Epub loads about 1.9k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 984 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from longhaiqwe/scanned-pdf-to-epub at commit 0e95243, republished under its MIT licence (© longhaiqwe). 984 words, ~1,903 tokens.

Download SKILL.mdSave it as .claude/skills/scanned-pdf-to-epub/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
scanned-pdf-to-epub
description
Use when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or WeRead/微信读书 compatibility matters.

Scanned PDF To EPUB

Overview

Convert scanned Chinese PDFs into readable text EPUBs without losing editorial aids such as underlined names. The workflow is cross-platform: choose OCR and rendering tools available on the user's machine, verify a small sample first, then scale.

Environment Preparation

Start every conversion with an environment check; the user should not need to ask separately for tool installation. Follow environment setup to detect the OS and usable runtimes, reuse host-provided tools, and install only missing dependencies for the chosen pipeline in a task-local virtual environment. Reuse a previously verified environment when its executable paths still work.

Explain the selected tools briefly and continue within the user's existing authorization. Do not require confirmation for each routine local dependency. For system-level installation, paid services, or uploading book content, check existing authorization and host permissions; ask only for the concrete missing authorization. Do not silently switch from local to cloud OCR.

Before processing the whole book, run one real body page through rendering, Chinese OCR with coordinates, and a minimal text EPUB package. An installed command or successful import alone does not establish readiness. If setup fails, diagnose the error or choose an available equivalent; report the exact blocker rather than claiming the environment is ready.

Workflow

  1. Inspect the PDF.

    • Use pdfinfo/Poppler when available, or Python libraries such as pypdf, PyMuPDF, or pypdfium2 for page count, metadata, page size, text layer, outline/bookmarks.
    • If extract_text() is empty on early pages, treat it as scan/OCR work.
    • Render a few pages with Poppler, PyMuPDF, or pypdfium2 and visually locate title, table of contents, and first body page.
  2. Build a sample before full conversion.

    • Pick title/front matter plus one complete short section.
    • Confirm PDF page number vs printed page number offset.
    • Make the EPUB title, nav, and NCX point to real section starts.
  3. OCR text with coordinates.

    • Prefer OCR that returns text boxes/character boxes.
    • On Windows, prefer PaddleOCR for local Chinese OCR; use Tesseract only as a fallback for simpler pages. Cloud OCR such as Azure, Baidu, Tencent, or Google is acceptable when privacy/cost is acceptable.
    • On macOS, Vision OCR via a tiny Swift helper is a good local default; PaddleOCR is also fine for portability.
    • On Linux, prefer PaddleOCR or Tesseract, with cloud OCR as the high-accuracy fallback.
    • If OCR returns only line/word boxes, estimate character boxes from the line box and text length; mark underline detection as lower confidence.
    • Sort lines by page coordinates; remove page headers, marginal running titles, and page numbers. In every text path (including the main body), merge same-row OCR fragments left-to-right before computing indentation. Carry character boxes through the merge for proper-name marking; a fragment starting far to the right is not evidence of a new paragraph.
    • Use OCR text for layout/alignment, then校对 against reliable text when available.
    • Reconstruct paragraphs across pages in all prose, including 前记、序跋 and commentary. A PDF page boundary must not create a new paragraph. Keep a paragraph buffer across pages; determine actual breaks from page-local text-region indentation, block style and headings. Do not use punctuation alone, raw x coordinates across facing pages, or one <p> per page.
    • Use logical row merging before paragraph accumulation in both front matter and body; it preserves character boxes. Calibrate row tolerance to OCR coordinates and process columns separately.
    • Read paragraph reconstruction before assembling XHTML. The optional paragraph accumulator supports horizontal prose when supplied with verified region-relative indentation; calibrate its inset-block assumptions to the source edition.
  4. Preserve original underlined proper names.

    • Detect underlines from images as horizontal strokes near the bottom of OCR character boxes.
    • If character boxes are estimated, compare the detected stroke span against the estimated character centers and manually review more tokens.
    • Apply detection only to body lines, not headings; headings create false positives from Hanzi strokes.
    • Convert detected spans into semantic markup before styling, e.g. <span class="proper-name">周威烈王</span>.
  5. Choose WeRead marking strategy.

    • For 微信读书 App and web together, prefer font/weight marking over underline: <strong class="proper-name" style="font-family:'Kaiti SC','STKaiti','KaiTi','楷体','FangSong','STFangsong',serif;font-weight:700;text-decoration:none;border:0;background:transparent;">周威烈王</strong>
    • Do not use border-bottom for WeRead web; its column layout can stretch borders across lines/columns.
    • Do not mix Unicode combining underline marks with web CSS in one WeRead file; web may render combining marks as misplaced strokes.
    • If the user insists on real underlines, produce separate App and web editions. See implementation notes.
  6. Package and validate EPUB.

    • Use EPUB3: mimetype, META-INF/container.xml, OEBPS/content.opf, nav.xhtml, and toc.ncx.
    • XHTML must parse as XML. Zip mimetype first and uncompressed.
    • Validate with unzip -t; parse OPF/nav/NCX/XHTML with XML parser.
    • Audit every page boundary as join, new-paragraph, or section transition; inspect ambiguous cases in the source. Check known continuations in the generated XHTML are inside one <p>, and verify actual new paragraphs and chapter boundaries remain separate. Include front matter and indented commentary in this check; XML/ZIP validity does not validate reading order or paragraph structure.
    • Treat reader rendering as a separate check: compare visible column-boundary text against XHTML on the target reader. If text exists in the EPUB but is missing onscreen, investigate import parsing/rendering; a no-markup comparison edition can test style involvement, but do not claim a confirmed cause or fix without reproducing the result.
    • Report sample scope, remaining OCR risk, and which reader edition to use.
Show full SKILL.md (121 more words)Show less

Common Mistakes

  • Making an image-only EPUB when the user asked for a text EPUB.
  • Splitting prose at PDF page boundaries, or merging all pages into one paragraph without preserving genuine paragraph starts.
  • Removing full-width prose with a fixed margin-x filter; scan alignment can shift between pages.
  • Trusting OCR without a校对 pass on names, dynasties, and section starts.
  • Treating WeRead App and WeRead web as the same renderer.
  • Using border-bottom as an underline fallback in WeRead web.
  • Leaving Unicode combining underline marks in a web-targeted EPUB.
  • Applying underline detection to headings or TOC pages.

When More Detail Is Needed

Read implementation-notes.md for cross-platform OCR choices, Windows PaddleOCR setup, macOS Vision helper shape, underline detection heuristics, EPUB packaging checks, and WeRead compatibility variants.

© longhaiqwe, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in the repository root of longhaiqwe/scanned-pdf-to-epub.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/openai.yaml
  • references/environment-setup.md
  • references/implementation-notes.md
  • scripts/paragraphs.py
  • scripts/test_logical_rows.py
  • scripts/test_paragraphs.py

Open the folder on GitHubat commit 0e95243

Compare with similar skills

Scanned PDF To Epub next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scanned PDF To Epub compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scanned PDF To Epub this skilllonghaiqwe/scanned-pdf-to-epub161—~1.9kAutomated safety check: PassMIT
Rename Pdfsrealspqrk/autorename-pdf122—~541Automated safety check: PassMIT
Seeoil-oil/see-skill160—~632Automated safety check: NotesMIT
Editaplotliushunqi8-hash/editaplot2026128—~5.8kAutomated safety check: PassApache-2.0
Prismer NotionPrismer-AI/PrismerCloud1.6k3 repos~4.6kAutomated safety check: PassMIT
File Documentsiamlukethedev/Herald-OS408—~1.7kAutomated safety check: PassMIT

Similar skills

  • Rename Pdfs

    realspqrk/autorename-pdf

    [Dev] Rename PDF files using AI (Python, requires venv). An agent skill from realspqrk/autorename-pdf.

    122 GitHub stars~541 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • See

    oil-oil/see-skill

    使用外部视觉服务或本地 OCR 分析本地图片、截图、视频及明确提供的媒体 URL,输出 Markdown。用户明确要求 see,或当前主模型缺少完成媒体分析所需的视觉能力时使用。不拦截普通文本、任意 URL 或已有原生视觉可以直接完成的任务;不用于生图、编辑图片或自动操作屏幕。使用外部服务前说明发送的媒体与服务范围,无法识别时如实说明限制。

    160 GitHub stars~632 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Editaplot

    liushunqi8-hash/editaplot2026

    Analyze local scientific CSV, TXT, XLS, or XLSX data; recommend publication-informed charts and Chinese scientific palettes; freeze a reproducible plan; and automate editable figures through a…

    128 GitHub stars~5.8k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed
  • Prismer Notion

    Prismer-AI/PrismerCloud

    Notion API + ntn CLI: pages, databases, markdown, Workers. An agent skill from Prismer-AI/PrismerCloud.

    1.6k GitHub starsUsed in 3 repos~4.6k tokens
    Documents & OfficeAuto-check passed
  • File Documents

    iamlukethedev/Herald-OS

    Find documents such as invoices among badly named files, read them, name them clearly and file them where the person keeps them, with a plan they approve and an undo

    408 GitHub stars~1.7k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Omk PDF Gen

    KaimingWan/oh-my-kiro

    Generate professional PDF documents with correct CJK (Chinese/Japanese/Korean) rendering.

    107 GitHub stars~1.1k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

Works with

Questions about Scanned PDF To Epub

What does Scanned PDF To Epub do?

A skill your agent uses when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or…. Scanned PDF To Epub is an agent skill from longhaiqwe/scanned-pdf-to-epub. Use when scanned Chinese PDFs or image-only books need text EPUB conversion on macOS, Windows, or Linux, especially when OCR text, original underlined proper names, 专名号, or WeRead/微信读书 compatibility matters.

When should I use Scanned PDF To Epub?

Scanned PDF To Epub fits situations like: scanned Chinese PDFs; image-only books need text EPUB conversion on macOS; especially when OCR text; original underlined proper names.

How do I install Scanned PDF To Epub in Claude Code?

Run `npx skills add longhaiqwe/scanned-pdf-to-epub --skill scanned-pdf-to-epub -a claude-code`. Or copy the skill folder (the longhaiqwe/scanned-pdf-to-epub repository) into .claude/skills/scanned-pdf-to-epub in your project. Claude Code loads it when a task matches its description.

How do I install Scanned PDF To Epub in Codex?

Run `npx skills add longhaiqwe/scanned-pdf-to-epub --skill scanned-pdf-to-epub -a codex`. Or copy the skill folder (the longhaiqwe/scanned-pdf-to-epub repository) into .agents/skills/scanned-pdf-to-epub in your project. Codex loads it when a task matches its description.

Can I use Scanned PDF To Epub in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add longhaiqwe/scanned-pdf-to-epub --skill scanned-pdf-to-epub -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scanned-pdf-to-epub, .gemini/skills/scanned-pdf-to-epub, .github/skills/scanned-pdf-to-epub and .opencode/skills/scanned-pdf-to-epub in your project.

What does Scanned PDF To Epub need to run?

Going by SKILL.md and its folder, Scanned PDF To Epub needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Scanned PDF To Epub access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scanned PDF To Epub safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Scanned PDF To Epub use?

Scanned PDF To Epub is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scanned PDF To Epub use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.

What are the alternatives to Scanned PDF To Epub?

Skills that share tags, products or a category with Scanned PDF To Epub: Rename Pdfs (realspqrk/autorename-pdf, 122 stars), See (oil-oil/see-skill, 160 stars), Editaplot (liushunqi8-hash/editaplot2026, 128 stars) and Prismer Notion (Prismer-AI/PrismerCloud, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scanned PDF To Epub?

longhaiqwe (a GitHub user) maintains it in longhaiqwe/scanned-pdf-to-epub, which has 161 GitHub stars. The repository was last updated on September 15, 2026.

Source: longhaiqwe/scanned-pdf-to-epub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.