Agent skill

PDF Processing Guide

by agentscope-ai in agentscope-ai/QwenPaw

Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

ProprietaryAuto-check passedDocuments & Office

SKILL.md written in Chinese; this summary is our English description.

Install PDF Processing Guide

skills CLI
$ npx skills add agentscope-ai/QwenPaw --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/QwenPaw pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/QwenPaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/qwenpaw/agents/skills/pdf-zh .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
35k
Token cost
~1.8k tokens
SKILL.md length
157 words
Files
12 (incl. scripts)
Skills in repo
20
Repo updated
First seen
Licence
Proprietary

At a glance

Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

  • Extracting text or tables from a PDF
  • SKILL.md covers 前置要求, 概述, 快速入门 and Python 库, plus 4 more sections
  • Runs Python scripts from its folder; calls qpdf, pdftotext and python
  • Merging, splitting or rotating PDF pages

What it does

The skill is a Chinese-language PDF guide built on pypdf for reading, merging, splitting, rotating and metadata, pdfplumber for text and table extraction, reportlab for creating PDFs, and command-line tools from poppler-utils (`pdftotext`, `pdftoppm`) and qpdf. Scripts are run from the skill directory. For advanced features and JavaScript libraries it points to a reference file, and for filling forms it points to `forms.md`. The description also lists watermarks, encryption and OCR for scanned files.

Bundled scripts check fillable fields, extract form field information and structure, fill fillable fields, fill forms through annotations, convert PDFs to images, check bounding boxes and create validation images. One warning: never use Unicode subscript or superscript characters in ReportLab output, because the built-in fonts render them as black boxes, and use XML markup tags in Paragraph objects instead. The licence is listed as proprietary, so check its terms before reuse.

When your agent uses it

  • Extracting text or tables from a PDF
  • Merging, splitting or rotating PDF pages
  • Filling in a PDF form
  • Creating a new PDF report

Example prompts

  • “Extract all the tables from invoice-batch.pdf into CSV files.”
  • “Merge cover.pdf and body.pdf into one document and rotate the last page.”
  • “Fill in the fillable fields of application-form.pdf with the details in applicant.json.”
  • “Run OCR on scan.pdf so the text becomes searchable.”

Requirements

  • Python with pypdf, pdfplumber and reportlab
  • poppler-utils (`pdftotext`, `pdftoppm`) and qpdf for command-line work

What it can do on your machine

Read from SKILL.md and the folder at commit f389534. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • qpdf
    • pdftotext
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processing Guide loads about 1.8k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 157 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Proprietary) doesn't allow us to republish the file, so here is its outline and opening line. It has 157 words (~1,750 tokens).

“本指南涵盖了使用 Python 库和命令行工具进行 PDF 处理的基本操作。有关高级功能、JavaScript 库和详细示例,请参阅 REFERENCE.md。如果需要填写 PDF 表单,请阅读 FORMS.md 并按照其中的说明操作。”

— opening of SKILL.md by agentscope-ai, Proprietary
name
pdf
license
Proprietary. LICENSE.txt has complete terms
metadata.builtin_skill_version
1.1

Read the full SKILL.md on GitHub

Files

SKILL.md and 11 other files (scripts) in src/qwenpaw/agents/skills/pdf-zh of agentscope-ai/QwenPaw.

  • SKILL.md
  • LICENSE.txt
  • forms.md
  • reference.md
  • scripts/check_bounding_boxes.py
  • scripts/check_fillable_fields.py
  • scripts/convert_pdf_to_images.py
  • scripts/create_validation_image.py
  • scripts/extract_form_field_info.py
  • scripts/extract_form_structure.py
  • scripts/fill_fillable_fields.py
  • scripts/fill_pdf_form_with_annotations.py

Open the folder on GitHubat commit f389534

Compare with similar skills

PDF Processing Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processing Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processing Guide this skillagentscope-ai/QwenPaw35k—~1.8kAutomated safety check: PassProprietary
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Toolkit

    telagod/code-abyss

    Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

    243 GitHub stars~532 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from agentscope-ai/QwenPaw

All 20 skills in this repo
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    35k GitHub starsUsed in 6 repos~4.2k tokens
    Auto-check passed
  • Make Skill

    agentscope-ai/QwenPaw

    Turns reusable decisions, templates or workflows from the current conversation into a new workspace skill through plan, approval, draft, validation and publication.

    35k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • QwenPaw Make Skill

    agentscope-ai/QwenPaw

    Creates a focused workspace skill from the current conversation in QwenPaw, moving through a planning, approval, drafting, validation and publishing script pipeline.

    35k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • QwenPaw Scheduled Tasks

    agentscope-ai/QwenPaw

    Creates and manages scheduled or recurring jobs with the qwenpaw cron commands, always tied to an explicit agent ID and a confirmed target channel.

    35k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • DOCX Creation and Editing

    agentscope-ai/QwenPaw

    Creates, reads and edits Word .docx files, including tracked changes and comments, using docx-js for new files and XML editing for existing ones.

    35k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Mailbox Operations Hub

    agentscope-ai/QwenPaw

    Connects, registers, and operates a personal mailbox, reading, searching, sending, and organizing, through a managed mail server for nine domains.

    35k GitHub stars~2.6k tokensUpdated today
    Auto-check passed

Works with

Questions about PDF Processing Guide

What does PDF Processing Guide do?

Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more. The skill is a Chinese-language PDF guide built on pypdf for reading, merging, splitting, rotating and metadata, pdfplumber for text and table extraction, reportlab for creating PDFs, and command-line tools from poppler-utils (`pdftotext`, `pdftoppm`) and qpdf. Scripts are run from the skill directory.

When should I use PDF Processing Guide?

PDF Processing Guide fits situations like: extracting text or tables from a PDF; merging, splitting or rotating PDF pages; filling in a PDF form; creating a new PDF report.

How do I install PDF Processing Guide in Claude Code?

Run `npx skills add agentscope-ai/QwenPaw --skill pdf -a claude-code`. Or copy the skill folder (src/qwenpaw/agents/skills/pdf-zh in agentscope-ai/QwenPaw) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processing Guide in Codex?

Run `npx skills add agentscope-ai/QwenPaw --skill pdf -a codex`. Or copy the skill folder (src/qwenpaw/agents/skills/pdf-zh in agentscope-ai/QwenPaw) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF Processing Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/QwenPaw --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF Processing Guide need to run?

Going by SKILL.md and its folder, PDF Processing Guide needs Python for the scripts in its folder and the command-line tools its instructions call (qpdf, pdftotext and python). Our summary lists: Python with pypdf, pdfplumber and reportlab; poppler-utils (`pdftotext`, `pdftoppm`) and qpdf for command-line work.

Does PDF Processing Guide access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Processing Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Processing Guide use?

PDF Processing Guide carries a proprietary licence (declared in SKILL.md). It is published on GitHub, but it is not open source: read the licence file before using or sharing it.

How many tokens does PDF Processing Guide use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Processing Guide?

Skills that share tags, products or a category with PDF Processing Guide: PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processing Guide?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/QwenPaw, which has 35,493 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: agentscope-ai/QwenPaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.