Agent skill

Gutenberg

by magnus919 in magnus919/agent-skills

Search, download, and extract public-domain books from Project Gutenberg.

MITAuto-check passedDocuments & Office

Install Gutenberg

skills CLI
$ npx skills add magnus919/agent-skills --skill gutenberg -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills gutenberg --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/gutenberg .claude/skills/gutenberg && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gutenberg
GitHub stars
116
Token cost
~2.3k tokens
SKILL.md length
713 words
Files
5 (incl. scripts, references)
Skills in repo
130
Repo updated
First seen
Licence
MIT

At a glance

Search, download, and extract public-domain books from Project Gutenberg.

  • The user says gutenberg
  • SKILL.md covers Quick Start, How It Works, CLI Reference and Fiction vs Non-Fiction Handling, plus 2 more sections
  • Calls python3; reaches gutenberg.org
  • Download a book

What it does

Gutenberg is an agent skill from magnus919/agent-skills. Search, download, and extract public-domain books from Project Gutenberg. Look up books by ID or keyword via gutendex, download plain-text and EPUB editions, strip licensing boilerplate, extract clean text from EPUB for illustrated works, and classify fiction vs non-fiction. Ships a portable CLI script with zero external dependencies. Use when the user says "gutenberg", "public domain", "download a book", "classic literature", "free ebook", "gutenberg.org", or names any public-domain title or author. Do not use…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `references/epub-extraction.md`). Compatibility notes: Python 3.8+ with zero external dependencies. The CLI uses only the Python standard library (urllib.request, json, html.parser, zipfile, re, sys). For EPUB…

It sits in Documents & Office, covering Creative writing and fiction and Project scaffolding. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • The user says gutenberg
  • Download a book
  • Classic literature
  • Names any public-domain title

Example prompts

  • “gutenberg”
  • “public domain”
  • “download a book”
  • “/gutenberg”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Python 3.8+ with zero external dependencies. The CLI uses only the Python standard library (urllib.request, json, html.parser, zipfile, re, sys). For EPUB extraction, Python 3.8+ with only stdlib is required (zipfile + html.parser). The gutendex API (https://gutendex.com) requires no API key or registration. No env vars needed for basic operation.

What it can do on your machine

Read from SKILL.md and the folder at commit c545c2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gutenberg.org

    Also links to:

    • gutendex.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Python 3.8+ with zero external dependencies. The CLI uses only the Python standard library (urllib.request, json, html.parser, zipfile, re, sys). For EPUB extraction, Python 3.8+ with only stdlib is required (zipfile + html.parser). The gutendex API (https://gutendex.com) requires no API key or registration. No env vars needed for basic operation.

    From compatibility in the SKILL.md frontmatter.

Context cost

Gutenberg loads about 2.3k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 150 tokens; SKILL.md has 713 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit c545c2b, republished under its MIT licence (© magnus919). 713 words, ~2,290 tokens.

Download SKILL.mdSave it as .claude/skills/gutenberg/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
gutenberg
description
Search, download, and extract public-domain books from Project Gutenberg. Look up books by ID or keyword via gutendex, download plain-text and EPUB editions, strip licensing boilerplate, extract clean text from EPUB for illustrated works, and classify fiction vs non-fiction. Ships a portable CLI script with zero external dependencies. Use when the user says "gutenberg", "public domain", "download a book", "classic literature", "free ebook", "gutenberg.org", or names any public-domain title or author. Do not use this skill for unrelated requests; route to the nearest named specialist.
compatibility
Python 3.8+ with zero external dependencies. The CLI uses only the Python standard library (urllib.request, json, html.parser, zipfile, re, sys). For EPUB extraction, Python 3.8+ with only stdlib is required (zipfile + html.parser). The gutendex API (https://gutendex.com) requires no API key or registration. No env vars needed for basic operation.
license
MIT
metadata.tags
gutenberg, project-gutenberg, books, public-domain, literature, classics, ebooks, text-extraction, epub
metadata.sources
https://gutendex.com, https://www.gutenberg.org
metadata.skills
books, public-domain, literature, text-mining, ebooks

Gutenberg — Public Domain Book Toolkit

Search, download, and extract clean text from Project Gutenberg — 70,000+ free public-domain ebooks. Ships a portable Python CLI with zero external dependencies.

Quick Start

bash
# Search for books
python3 scripts/gutenberg search "Moby Dick"

# Download by Gutenberg ID (plain text)
python3 scripts/gutenberg download 2701 --format txt

# Download EPUB (for illustrated books)
python3 scripts/gutenberg download 2701 --format epub

# Extract clean text (strips PG boilerplate)
python3 scripts/gutenberg extract 2701

# Classify fiction vs non-fiction
python3 scripts/gutenberg classify 2701

# Full pipeline: search → download → extract
python3 scripts/gutenberg pipeline "Alice's Adventures in Wonderland"

How It Works

Project Gutenberg provides 70,000+ free public-domain ebooks in multiple formats. The gutendex API (https://gutendex.com) offers a free, unauthenticated JSON catalog. No API key required — just curl or this CLI.

Data Flow
User provides title/ID/author
       ↓
gutendex API search → pick book by ID
       ↓
Download plain text (preferred) or EPUB (fallback for illustrated books)
       ↓
Strip PG boilerplate → clean text
       ↓
Classify fiction/non-fiction → extract content

CLI Reference

search — Find books by keyword
bash
python3 scripts/gutenberg search "Moby Dick"
python3 scripts/gutenberg search "Dracula" --limit 5
python3 scripts/gutenberg search "Sherlock Holmes" --json
python3 scripts/gutenberg search "Alice" --language en

Returns: ID, title, author (with life dates), language, subjects, download count. Results sorted by download count (most popular first).

metadata — Get full metadata for a book by ID
bash
python3 scripts/gutenberg metadata 2701          # Moby Dick
python3 scripts/gutenberg metadata 11            # Alice's Adventures
python3 scripts/gutenberg metadata 1342          # Pride and Prejudice
python3 scripts/gutenberg metadata 1342 --json   # JSON-only output

Returns: title, author(s), language(s), subjects, bookshelves, summaries, copyright status, download count, and all available format URLs.

download — Download a book by Gutenberg ID
bash
# Plain text (UTF-8, preferred — works for most books)
python3 scripts/gutenberg download 2701 --format txt

# EPUB with images (for illustrated/scientific books)
python3 scripts/gutenberg download 2701 --format epub

# HTML (alternative fallback)
python3 scripts/gutenberg download 2701 --format html

# Specify output directory
python3 scripts/gutenberg download 2701 --format txt --output ./books/

The file is saved to ./gutenberg-<id>.<ext> (or --output path). Large books may take a moment.

extract — Strip PG boilerplate and produce clean text
bash
python3 scripts/gutenberg extract 2701            # from downloaded txt
python3 scripts/gutenberg extract 2701 --input ./gutenberg-2701.txt
python3 scripts/gutenberg extract 2701 --format epub  # extract from EPUB

Output: clean text without the Project Gutenberg license header/footer. For EPUB extraction (illustrated books), extracts text from all XHTML files and merges them into a single cleaned document.

Size detection: if a plain-text download is under 50KB for a known substantial book, warns that the text may be truncated and recommends EPUB mode.

classify — Classify fiction vs non-fiction
bash
python3 scripts/gutenberg classify 2701
python3 scripts/gutenberg classify 2701 --json

Uses the book's subjects and bookshelves to classify:

  • Fiction signals: "Fiction", "novels", "short stories", "poetry", "drama", "fantasy", "horror"
  • Non-fiction signals: "Essays", "History", "Philosophy", "Biography", "Science", "Religion"

Returns: fiction, non-fiction, or ambiguous (with explanation of why).

pipeline — Full fetch pipeline
bash
python3 scripts/gutenberg pipeline "Moby Dick"                           # search first
python3 scripts/gutenberg pipeline 2701                                   # by known ID
python3 scripts/gutenberg pipeline 2701 --clean /tmp/pipeline-output/     # save cleaned text

Runs: search (if title) → metadata → download (txt) → check size → extract (or EPUB fallback) → classify. Prints an executive summary at the end.

Global Flags
FlagEffect
--jsonOutput machine-readable JSON instead of human-readable text
--quietSuppress diagnostic output
--dry-runShow what would be done without executing
--output ./dirSave downloads to a specific directory
--timeout 30Override API timeout (default 15s)

Fiction vs Non-Fiction Handling

When the classified result is fiction, the extracted text comes from an authored imagination. Consider splitting analysis into two tracks:

TrackWhat it coversExample claims
CanonFacts within the fictional world — named entities, quoted lines, story events, world rules"In Stoker's text, Dracula can assume wolf, bat, and mist forms"
CraftReal-world technique — how the author achieves effect, publication history, literary influence"Stoker's epistolary form forces the reader to piece together the narrative like an investigator"
Negative spaceDeliberate omissions — what the author notably leaves unspecified"Dracula is never granted interior voice in the novel"

When classified as non-fiction, claims can be treated as real-world factual assertions about the subject matter.

Show full SKILL.md (296 more words)Show less

Known Gotchas

  • Plain text truncation for illustrated books — Books with diagrams, figures, or equations (geometry texts, scientific works, art books) may have plain-text downloads silently cut to 5-10KB (just the PG header). Always check file size. Under 50KB for a known substantial book → switch to EPUB extraction. The pipeline command does this check automatically.
  • Gutendex can be slow or timeout — The API is a free service and can be slow for less popular books. The CLI uses a 15-second default timeout. Use --timeout 30 for slow responses, or navigate directly to https://www.gutenberg.org/ebooks/<id> as a fallback.
  • HTML downloads include navigation markup — HTML downloads contain site navigation and formatting. Prefer plain text or EPUB for clean text extraction.
  • Rare books may 404 on certain format URLs — Not every book has every format. The CLI tries UTF-8 plain text first, falls back to US-ASCII, then to the -0.txt file path, then to EPUB, then to HTML. The download command reports which format was actually retrieved.
  • Rate limiting — Gutendex is unauthenticated but rate-limited. Batch requests with sleep 1 between calls for more than 10 rapid-fire requests.
  • utf-8 vs us-ascii — Gutendex returns both a text/plain; charset=utf-8 and a text/plain; charset=us-ascii URL. Prefer UTF-8; fall back to US-ASCII if the UTF-8 URL returns a 404.
  • Fiction classification ambiguity — Books with both fiction and non-fiction subjects (e.g. "Historical Fiction" + "History") are marked ambiguous. Use --json to inspect the subject list and decide manually.

References

  • scripts/gutenberg — Portable Python CLI. Zero external dependencies (stdlib only). Covers all major Gutenberg workflows: search, download (txt/epub/html), boilerplate stripping, EPUB text extraction, fiction classification, and the full pipeline.
  • references/epub-extraction.md — EPUB text extraction details for illustrated books, with expanded Python walkthrough and format detection tips.
  • Project Gutenberg — 70,000+ free ebooks.
  • Gutendex API — JSON web API for the Project Gutenberg catalog.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in gutenberg of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/evals.json
  • references/epub-extraction.md
  • scripts/gutenberg

Open the folder on GitHubat commit c545c2b

Compare with similar skills

Gutenberg next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gutenberg compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gutenberg this skillmagnus919/agent-skills116—~2.3kAutomated safety check: PassMIT
Story Setup for Web Novel Writingzenstory-ai/oh-story-claudecode7.4k—~2.8kAutomated safety check: PassMIT
Hermes3000 WritingHybridAIOne/hybridclaw158—~2.7kAutomated safety check: PassMIT
Venue TemplatesK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: PassMIT
Story Maintenancedanjdewhurst/story-skills2831 repos~5.4kAutomated safety check: NotesMIT
Weak Agent Testkklimuk/docx-cli216—~6.1kAutomated safety check: NotesMIT

Similar skills

  • Story Setup for Web Novel Writing

    zenstory-ai/oh-story-claudecode

    Deploys and checks a web-novel writing toolkit in a project folder for supported coding agents, merging into existing configuration instead of overwriting it.

    7.4k GitHub stars~2.8k tokensUpdated 6 days ago
    Writing & ContentAuto-check passed
  • Hermes3000 Writing

    HybridAIOne/hybridclaw

    Use Hermes3000 to plan, draft, revise, save, check consistency, and export long-form manuscripts through the Hermes3000 AI writing portal API.

    158 GitHub stars~2.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Venue Templates

    K-Dense-AI/claude-scientific-writer

    Prepare journal manuscripts, conference papers, research posters, and grant documents using venue-specific formatting guidance and bundled LaTeX scaffolds.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • Story Maintenance

    danjdewhurst/story-skills

    This skill should be used when the user asks to "validate", "reindex", "repair registries", "check links", "run the continuity, pacing, clue, voice, or name checks", "count words", "summarize a…

    283 GitHub starsUsed in 1 repo~5.4k tokens
    Writing & ContentAuto-check: notes
  • Weak Agent Test

    kklimuk/docx-cli

    Run the weak-agent adversarial test harness against docx-cli.

    216 GitHub stars~6.1k tokensUpdated today
    Documents & OfficeAuto-check: notes
  • Forge Codegen Crud

    yaomindong1996/forge-admin

    Generate or review Forge project code-generation output for CRUD modules.

    125 GitHub stars~1.3k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from magnus919/agent-skills

All 130 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    116 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    116 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    116 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    116 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    116 GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    116 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed

Questions about Gutenberg

What does Gutenberg do?

Search, download, and extract public-domain books from Project Gutenberg. Gutenberg is an agent skill from magnus919/agent-skills. Search, download, and extract public-domain books from Project Gutenberg.

When should I use Gutenberg?

Gutenberg fits situations like: the user says gutenberg; download a book; classic literature; names any public-domain title.

How do I install Gutenberg in Claude Code?

Run `npx skills add magnus919/agent-skills --skill gutenberg -a claude-code`. Or copy the skill folder (gutenberg in magnus919/agent-skills) into .claude/skills/gutenberg in your project. Claude Code loads it when a task matches its description.

How do I install Gutenberg in Codex?

Run `npx skills add magnus919/agent-skills --skill gutenberg -a codex`. Or copy the skill folder (gutenberg in magnus919/agent-skills) into .agents/skills/gutenberg in your project. Codex loads it when a task matches its description.

Can I use Gutenberg in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill gutenberg -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gutenberg, .gemini/skills/gutenberg, .github/skills/gutenberg and .opencode/skills/gutenberg in your project.

What does Gutenberg need to run?

Going by SKILL.md and its folder, Gutenberg needs the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Python 3.8+ with zero external dependencies. The CLI uses only the Python standard library (urllib.request, json, html.parser, zipfile, re, sys). For EPUB extraction, Python 3.8+ with only stdlib is required (zipfile + html.parser). The gutendex API (https://gutendex.com) requires no API key or registration. No env vars needed for basic operation..

Does Gutenberg access the network?

SKILL.md names 2 domains. In commands or code: gutenberg.org; the agent is likely to contact it when it follows the instructions. As links in the text: gutendex.com. This is read from the text; nothing was executed.

Is Gutenberg safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gutenberg use?

Gutenberg is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gutenberg use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 760 tokens, read only when the agent opens those files.

What are the alternatives to Gutenberg?

Skills that share tags, products or a category with Gutenberg: Story Setup for Web Novel Writing (zenstory-ai/oh-story-claudecode, 7.4k stars), Hermes3000 Writing (HybridAIOne/hybridclaw, 158 stars), Venue Templates (K-Dense-AI/claude-scientific-writer, 2.4k stars) and Story Maintenance (danjdewhurst/story-skills, 283 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gutenberg?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 116 GitHub stars. The repository holds 130 skills in this directory. The repository was last updated on October 8, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.