Agent skill

PDF

by axoviq-ai in axoviq-ai/synthadoc

“Extract text from PDF documents”

— description from SKILL.md by axoviq-ai
AGPL-3.0-or-laterAuto-check passedDocuments & Office

Install PDF

skills CLI
$ npx skills add axoviq-ai/synthadoc --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install axoviq-ai/synthadoc pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/axoviq-ai/synthadoc.git skills-src && mkdir -p .claude/skills && cp -r skills-src/synthadoc/skills/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
1.6k
Token cost
~293 tokens
SKILL.md length
67 words
Files
5 (incl. scripts, references)
Skills in repo
9
Repo updated
First seen
Licence
AGPL-3.0-or-later

At a glance

  • SKILL.md covers Setup, Standalone usage, When this skill is used and Scripts, plus 1 more section
  • Runs Python scripts from its folder; calls pip

About this skill

PDF is a skill in axoviq-ai/synthadoc (1.6k stars). Its SKILL.md is about 293 tokens, with 4 other files in the folder (scripts, references). Licence: AGPL-3.0-or-later.

What it can do on your machine

Read from SKILL.md and the folder at commit 90d0db7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF loads about 293 tokens when it runs, and up to ~393 if it reads all its reference files. Until then it costs about 9 tokens; SKILL.md has 67 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~9
When it runs · the whole SKILL.md, loaded when a task matches
~293
With references · SKILL.md plus every file in references/, read only if the agent opens them
~393

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from axoviq-ai/synthadoc at commit 90d0db7, republished under its AGPL-3.0-or-later licence (© axoviq-ai). 67 words, ~293 tokens.

Download SKILL.mdSave it as .claude/skills/pdf/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
pdf
description
Extract text from PDF documents
version
1.0
entry.script
scripts/main.py
entry.class
PdfSkill
triggers.extensions
.pdf
triggers.intents
pdf, research paper
requires
pypdf, pdfminer.six
author
axoviq.com
license
AGPL-3.0-or-later

PDF Skill

Extracts text from PDF files using pypdf as the primary parser, with pdfminer.six as a fallback for CJK fonts that pypdf cannot decode (detected when pypdf yields fewer than 50 characters per page on average).

Setup

bash
pip install pypdf pdfminer.six

Standalone usage

python
import asyncio
from synthadoc.skills.pdf.scripts.main import PdfSkill

skill = PdfSkill()

async def main():
    result = await skill.extract("/path/to/paper.pdf")
    print(result.text)          # extracted text from all pages
    print(result.metadata)      # {"pages": N, "cjk_fallback": bool, ...}

asyncio.run(main())

When this skill is used

  • Source path ends with .pdf
  • User intent contains: pdf, research paper

Scripts

  • scripts/main.py — PdfSkill class

References

  • references/cjk-notes.md — notes on CJK font handling

© axoviq-ai, AGPL-3.0-or-later. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in synthadoc/skills/pdf of axoviq-ai/synthadoc.

  • SKILL.md
  • references/cjk-notes.md
  • requirements.txt
  • scripts/__init__.py
  • scripts/main.py

Open the folder on GitHubat commit 90d0db7

Compare with similar skills

PDF next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF this skillaxoviq-ai/synthadoc1.6k—~293Automated safety check: PassAGPL-3.0-or-later
Lov Any2pdflovstudio/any2pdf211—~2.4kAutomated safety check: NotesMIT
PDFzai-org/ZCode7.7k—~18kAutomated safety check: NotesProprietary
PDFnuoyimanaituling/manus-x830—~985Automated safety check: PassNone
PDFeinverne/dotfiles12147 repos~1.8kAutomated safety check: PassProprietary
MineruNebutra/MinerU-Skill123—~504Automated safety check: PassMIT

Similar skills

  • Lov Any2pdf

    lovstudio/any2pdf

    Convert Markdown documents to professionally typeset PDF files with reportlab.

    211 GitHub stars~2.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • PDF

    zai-org/ZCode

    Professional PDF toolkit covering four production workflows: reports, creative visuals, academic LaTeX, and existing PDF processing.

    7.7k GitHub stars~18k tokensUpdated today
    Documents & OfficeAuto-check: notes
  • PDF

    nuoyimanaituling/manus-x

    Process PDF files - extract text, read content, create PDFs, merge or split documents.

    830 GitHub stars~985 tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    123 GitHub stars~504 tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed
  • Summarize

    reysu/ai-life-skills

    Summarize any content (YouTube video, article, whitepaper/PDF, podcast episode, book chapter, etc.) into a rich Obsidian note with section-by-section breakdowns, wikilinks to all technical concepts…

    270 GitHub stars~7.7k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from axoviq-ai/synthadoc

All 9 skills in this repo
  • Image

    axoviq-ai/synthadoc

    Extract text from images using a vision LLM. An agent skill from axoviq-ai/synthadoc.

    1.6k GitHub stars~544 tokensUpdated today
    Auto-check passed
  • Session

    axoviq-ai/synthadoc

    Extract conversation turns from AI session history files (.jsonl)

    1.6k GitHub stars~902 tokensUpdated today
    Auto-check passed
  • Youtube

    axoviq-ai/synthadoc

    Extract transcripts from YouTube videos via the YouTube caption system

    1.6k GitHub stars~634 tokensUpdated today
    Auto-check passed
  • PPTX

    axoviq-ai/synthadoc

    Extract text from Microsoft PowerPoint presentations. An agent skill from axoviq-ai/synthadoc.

    1.6k GitHub stars~248 tokensUpdated today
    Auto-check passed
  • DOCX

    axoviq-ai/synthadoc

    Extract text from Microsoft Word documents. An agent skill from axoviq-ai/synthadoc.

    1.6k GitHub stars~207 tokensUpdated today
    Auto-check passed
  • URL

    axoviq-ai/synthadoc

    Fetch and extract text from web URLs

    1.6k GitHub stars~371 tokensUpdated today
    Auto-check passed

Works with

Questions about PDF

How do I install PDF in Claude Code?

Run `npx skills add axoviq-ai/synthadoc --skill pdf -a claude-code`. Or copy the skill folder (synthadoc/skills/pdf in axoviq-ai/synthadoc) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF in Codex?

Run `npx skills add axoviq-ai/synthadoc --skill pdf -a codex`. Or copy the skill folder (synthadoc/skills/pdf in axoviq-ai/synthadoc) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add axoviq-ai/synthadoc --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF need to run?

Going by SKILL.md and its folder, PDF needs Python for the scripts in its folder and the command-line tools its instructions call (pip).

Does PDF access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF use?

PDF is published under the AGPL-3.0-or-later licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF use?

About 293 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 100 tokens, read only when the agent opens those files.

What are the alternatives to PDF?

Skills that share tags, products or a category with PDF: Lov Any2pdf (lovstudio/any2pdf, 211 stars), PDF (zai-org/ZCode, 7.7k stars), PDF (nuoyimanaituling/manus-x, 830 stars) and PDF (einverne/dotfiles, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF?

axoviq-ai (a GitHub organization) maintains it in axoviq-ai/synthadoc, which has 1,578 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 10, 2026.

Source: axoviq-ai/synthadoc on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.