Agent skill

Multi Format Book Extractor

by Abilityai in Abilityai/cornelius

Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files.

MITAuto-check: notesDocuments & Office

Install Multi Format Book Extractor

skills CLI
$ npx skills add Abilityai/cornelius --skill multi-format-book-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Abilityai/cornelius multi-format-book-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/multi-format-book-extractor .claude/skills/multi-format-book-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-format-book-extractor
GitHub stars
109
Token cost
~872 tokens
SKILL.md length
266 words
Files
1
Skills in repo
56
Repo updated
First seen
Licence
MIT

At a glance

Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files.

  • Extracting non-EPUB books
  • SKILL.md covers Dependencies, Single Book, Scanned / Image-Only PDFs… and Batch Extraction (Parallel), plus 2 more sections
  • Calls uv and brew; needs GEMINI_API_KEY and GOOGLE_API_KEY
  • Tasks that involve PDF

What it does

Multi Format Book Extractor is an agent skill from Abilityai/cornelius. Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files. Use when extracting non-EPUB books. PDF uses pymupdf with TOC-based chapter splitting and auto-OCR fallback (Gemini) for scanned/image-only PDFs; MOBI/AZW3 routes through calibre then epub-chapter-extractor.

Its SKILL.md is about 870 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: AI-powered second brain template for Claude Code + Obsidian. The licence is MIT.

When your agent uses it

  • Extracting non-EPUB books
  • Tasks that involve PDF

Example prompts

  • “/multi-format-book-extractor”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • A credential in GOOGLE_API_KEY
  • Pre-approved tools (allowed-tools): Bash

What it can do on your machine

Read from SKILL.md and the folder at commit fd5e9a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • calibre-ebook.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multi Format Book Extractor loads about 872 tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 266 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~872

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:17
    _API_KEY` (read from `cornelius-internal/.env` if unset)
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Abilityai/cornelius at commit fd5e9a4, republished under its MIT licence (© Abilityai). 266 words, ~872 tokens.

Download SKILL.mdSave it as .claude/skills/multi-format-book-extractor/SKILL.md (or your agent's skills folder).
name
multi-format-book-extractor
description
Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files. Use when extracting non-EPUB books. PDF uses pymupdf with TOC-based chapter splitting and auto-OCR fallback (Gemini) for scanned/image-only PDFs; MOBI/AZW3 routes through calibre then epub-chapter-extractor.
allowed-tools
Bash
user-invocable
true

Multi-Format Book Extractor

Extract chapters from PDF, MOBI, and AZW3 files into individual numbered markdown files. Mirrors the epub-chapter-extractor pattern for non-EPUB formats.

Dependencies

  • PDF: pymupdf — auto-installed via --with pymupdf
  • Scanned PDF OCR fallback: requests (--with requests) + a Gemini key in $GEMINI_API_KEY / $GOOGLE_API_KEY (read from cornelius-internal/.env if unset)
  • MOBI / AZW3: calibre must be installed
    bash
    brew install --cask calibre

Single Book

bash
cd $PROJECT_ROOT/.claude/skills/multi-format-book-extractor && \
uv run --with pymupdf --with requests python extract_book.py "/path/to/book.pdf" [output_dir]

If output_dir is omitted, creates a folder named after the file in the same directory.

Scanned / Image-Only PDFs (auto-OCR)

Some PDFs are pure page scans with no text layer — pymupdf extracts nothing and would silently write empty files. extract_book.py now samples the text layer first (_has_text_layer) and, when a PDF looks image-only, automatically routes it to Gemini 2.5 Flash OCR (ocr_pdf.py).

  • Renders each page to PNG (~180 DPI), batches pages per API call, transcribes verbatim to Markdown (preserves verse line breaks, Sanskrit/Tibetan diacritics, footnotes), and writes NN_pages-S-E.md chunk files.
  • Resumable: re-running skips chunk files already written (≥200 chars), so an interrupted long book picks up where it left off.
  • Cost is ~$0.002/page on Flash (a 300-page book ≈ $0.60).

Run OCR directly (e.g. to tune params or force it):

bash
cd $PROJECT_ROOT/.claude/skills/multi-format-book-extractor && \
uv run --with pymupdf --with requests python ocr_pdf.py "/path/to/scan.pdf" [output_dir] \
  [--model gemini-2.5-flash] [--dpi 180] [--pages-per-call 5] [--pages-per-file 20] [--concurrency 6]

Batch Extraction (Parallel)

Extract all PDFs, MOBIs, and AZW3s from a directory in parallel:

bash
SKILL="$PROJECT_ROOT/.claude/skills/multi-format-book-extractor"
UVX="uv"
DIR="/path/to/books"

for book in "$DIR"/*.{pdf,mobi,azw3}; do
  [ -f "$book" ] || continue
  stem="${book%.*}"
  [ -d "$stem" ] && echo "Skipping: $(basename "$book") (already extracted)" && continue
  echo "Starting: $(basename "$book")"
  (cd "$SKILL" && $UVX run --with pymupdf --with requests python extract_book.py "$book") &
done
wait
echo "All done"

Output Format

Each book gets a subfolder named after the file (without extension):

book-name/
├── 01_introduction.md
├── 02_chapter-one.md
└── ...
  • PDF with TOC: one file per chapter (uses level-1 TOC; falls through to level-2 if < 5 top-level entries, e.g. books structured as Parts)
  • PDF without TOC: 20-page chunks (01_pages-1-20.md, etc.)
  • Scanned / image-only PDF: auto-routed to Gemini OCR → 20-page chunks (01_pages-1-20.md) with <!-- page N --> markers
  • MOBI / AZW3: same output quality as epub-chapter-extractor (converts via calibre first)

After Extraction

bash
open /path/to/books/

© Abilityai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/multi-format-book-extractor of Abilityai/cornelius.

Open the folder on GitHubat commit fd5e9a4

Compare with similar skills

Multi Format Book Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Format Book Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Format Book Extractor this skillAbilityai/cornelius109—~872Automated safety check: NotesMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.8k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.8k GitHub stars~19k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from Abilityai/cornelius

All 56 skills in this repo
  • Nano Banana Image Generator

    Abilityai/cornelius

    Generate images using Google's Nano Banana (Gemini 2.5 Flash Image).

    109 GitHub stars~1.2k tokensUpdated 15 days ago
    Auto-check: notes
  • Changelog Protocol

    Abilityai/cornelius

    Protocol for creating dated changelog files after significant agent sessions.

    109 GitHub stars~555 tokensUpdated 15 days ago
    Auto-check passed
  • Create Article

    Abilityai/cornelius

    Create long-form articles from knowledge base insights. An agent skill from Abilityai/cornelius.

    109 GitHub stars~2.1k tokensUpdated 15 days ago
    Auto-check: notes
  • Epistemic Classification

    Abilityai/cornelius

    Framework for distinguishing research findings from hypotheses and speculative synthesis.

    109 GitHub stars~1.5k tokensUpdated 15 days ago
    Auto-check passed
  • Get Youtube Transcript

    Abilityai/cornelius

    Extract the transcript from a YouTube video by URL or video ID.

    109 GitHub stars~525 tokensUpdated 15 days ago
    Auto-check: notes
  • Insight Capture Format

    Abilityai/cornelius

    Standard format for capturing and documenting insights in the knowledge base.

    109 GitHub stars~616 tokensUpdated 15 days ago
    Auto-check passed

Questions about Multi Format Book Extractor

What does Multi Format Book Extractor do?

Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files. Multi Format Book Extractor is an agent skill from Abilityai/cornelius. Extract chapters from PDF, MOBI, and AZW3 book files into individual markdown files.

When should I use Multi Format Book Extractor?

Multi Format Book Extractor fits situations like: extracting non-EPUB books; tasks that involve PDF.

How do I install Multi Format Book Extractor in Claude Code?

Run `npx skills add Abilityai/cornelius --skill multi-format-book-extractor -a claude-code`. Or copy the skill folder (.claude/skills/multi-format-book-extractor in Abilityai/cornelius) into .claude/skills/multi-format-book-extractor in your project. Claude Code loads it when a task matches its description.

How do I install Multi Format Book Extractor in Codex?

Run `npx skills add Abilityai/cornelius --skill multi-format-book-extractor -a codex`. Or copy the skill folder (.claude/skills/multi-format-book-extractor in Abilityai/cornelius) into .agents/skills/multi-format-book-extractor in your project. Codex loads it when a task matches its description.

Can I use Multi Format Book Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Abilityai/cornelius --skill multi-format-book-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-format-book-extractor, .gemini/skills/multi-format-book-extractor, .github/skills/multi-format-book-extractor and .opencode/skills/multi-format-book-extractor in your project.

What does Multi Format Book Extractor need to run?

Going by SKILL.md and its folder, Multi Format Book Extractor needs the command-line tools its instructions call (uv and brew) and credentials named GEMINI_API_KEY and GOOGLE_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in GOOGLE_API_KEY. Its frontmatter pre-approves these tools: Bash.

Does Multi Format Book Extractor access the network?

SKILL.md names 1 domain. As links in the text: calibre-ebook.com. This is read from the text; nothing was executed.

Is Multi Format Book Extractor safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Multi Format Book Extractor use?

Multi Format Book Extractor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multi Format Book Extractor use?

About 872 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multi Format Book Extractor?

Skills that share tags, products or a category with Multi Format Book Extractor: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Format Book Extractor?

Abilityai (a GitHub organization) maintains it in Abilityai/cornelius, which has 109 GitHub stars. The repository holds 56 skills in this directory. The repository was last updated on September 22, 2026.

Source: Abilityai/cornelius on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.