ALWAYS use this skill instead of the Read tool for PDF files.

Apache-2.0Auto-check passedDocuments & Office

Install PDF

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/sales-pivot-analysis/environment/skills/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
1.8k
Token cost
~1.6k tokens
SKILL.md length
257 words
Files
1
Skills in repo
180
Repo updated
First seen
Licence
Apache-2.0

At a glance

ALWAYS use this skill instead of the Read tool for PDF files.

  • Reading ANY PDF file
  • SKILL.md covers IMPORTANT: Do NOT Use the Read…, Overview, Table Extraction with pdfplumber and Text Extraction, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Extracting tables from PDFs

What it does

PDF is an agent skill from benchflow-ai/skillsbench. ALWAYS use this skill instead of the Read tool for PDF files. The Read tool cannot extract PDF tables properly. Use this skill when: (1) Reading ANY PDF file, (2) Extracting tables from PDFs, (3) Converting PDF tables to pandas DataFrames, (4) Processing multi-page PDFs

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF and DataFrames. It works with pandas and pypdf. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Reading ANY PDF file
  • Extracting tables from PDFs
  • Converting PDF tables to pandas DataFrames
  • Processing multi-page PDFs

Example prompts

  • “Use the pdf skill to alway use this skill instead of the Read tool for PDF files”
  • “/pdf”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF loads about 1.6k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 257 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 257 words, ~1,613 tokens.

Download SKILL.mdSave it as .claude/skills/pdf/SKILL.md (or your agent's skills folder).
name
pdf
description
ALWAYS use this skill instead of the Read tool for PDF files. The Read tool cannot extract PDF tables properly. Use this skill when: (1) Reading ANY PDF file, (2) Extracting tables from PDFs, (3) Converting PDF tables to pandas DataFrames, (4) Processing multi-page PDFs

PDF Processing Guide

IMPORTANT: Do NOT Use the Read Tool for PDFs

The Read tool cannot properly extract tabular data from PDFs. It will only show you a limited preview of the first page's text content, missing most of the data.

For PDF files, especially multi-page PDFs with tables:

  • DO: Use pdfplumber as shown in this skill
  • DO NOT: Use the Read tool to view PDF contents

If you need to extract data from a PDF, write Python code using pdfplumber. This is the only reliable way to get complete table data from all pages.

Overview

Extract text and tables from PDF documents using Python libraries.

Table Extraction with pdfplumber

pdfplumber is the recommended library for extracting tables from PDFs.

Basic Table Extraction
python
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    for page in pdf.pages:
        tables = page.extract_tables()
        for table in tables:
            print(table)  # List of lists
Convert to pandas DataFrame
python
import pdfplumber
import pandas as pd

with pdfplumber.open("document.pdf") as pdf:
    page = pdf.pages[0]  # First page
    tables = page.extract_tables()

    if tables:
        # First row is usually headers
        table = tables[0]
        df = pd.DataFrame(table[1:], columns=table[0])
        print(df)
Extract All Tables from Multi-Page PDF
python
import pdfplumber
import pandas as pd

all_tables = []

with pdfplumber.open("document.pdf") as pdf:
    for page in pdf.pages:
        tables = page.extract_tables()
        for table in tables:
            if table and len(table) > 1:  # Has data
                df = pd.DataFrame(table[1:], columns=table[0])
                all_tables.append(df)

# Combine if same structure
if all_tables:
    combined = pd.concat(all_tables, ignore_index=True)

Text Extraction

Full Page Text
python
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    text = ""
    for page in pdf.pages:
        text += page.extract_text() + "\n"
Text with Layout Preserved
python
with pdfplumber.open("document.pdf") as pdf:
    page = pdf.pages[0]
    text = page.extract_text(layout=True)

Common Patterns

Extract Multiple Tables and Create Mappings

When a PDF contains multiple related tables (possibly spanning multiple pages), extract from ALL pages and build lookup dictionaries:

python
import pdfplumber
import pandas as pd

category_map = {}      # CategoryID -> CategoryName
product_map = {}       # ProductID -> (ProductName, CategoryID)

with pdfplumber.open("catalog.pdf") as pdf:
    for page_num, page in enumerate(pdf.pages):
        tables = page.extract_tables()

        for table in tables:
            if not table or len(table) < 2:
                continue

            # Check first row to identify table type
            header = [str(cell).strip() if cell else '' for cell in table[0]]

            # Determine if first row is header or data (continuation page)
            if 'CategoryID' in header and 'CategoryName' in header:
                # Categories table with header
                for row in table[1:]:
                    if row and len(row) >= 2 and row[0]:
                        cat_id = int(row[0].strip())
                        cat_name = row[1].strip()
                        category_map[cat_id] = cat_name

            elif 'ProductID' in header and 'ProductName' in header:
                # Products table with header
                for row in table[1:]:
                    if row and len(row) >= 3 and row[0]:
                        try:
                            prod_id = int(row[0].strip())
                            prod_name = row[1].strip()
                            cat_id = int(row[2].strip())
                            product_map[prod_id] = (prod_name, cat_id)
                        except (ValueError, AttributeError):
                            continue
            else:
                # Continuation page (no header) - check if it's product data
                for row in table:
                    if row and len(row) >= 3 and row[0]:
                        try:
                            prod_id = int(row[0].strip())
                            prod_name = row[1].strip()
                            cat_id = int(row[2].strip())
                            product_map[prod_id] = (prod_name, cat_id)
                        except (ValueError, AttributeError):
                            continue

# Build final mapping: ProductID -> CategoryName
product_to_category_name = {
    pid: category_map[cat_id]
    for pid, (name, cat_id) in product_map.items()
}
# {1: 'Beverages', 2: 'Beverages', 3: 'Condiments', ...}

Important: Always iterate over ALL pages (for page in pdf.pages) - tables often span multiple pages, and continuation pages may not have headers.

Clean Extracted Table Data
python
import pdfplumber
import pandas as pd

with pdfplumber.open("document.pdf") as pdf:
    table = pdf.pages[0].extract_tables()[0]
    df = pd.DataFrame(table[1:], columns=table[0])

    # Clean whitespace
    df = df.apply(lambda x: x.str.strip() if x.dtype == "object" else x)

    # Remove empty rows
    df = df.dropna(how='all')

    # Rename columns if needed
    df.columns = df.columns.str.strip()
Handle Missing Headers
python
# If table has no header row
table = pdf.pages[0].extract_tables()[0]
df = pd.DataFrame(table)
df.columns = ['Col1', 'Col2', 'Col3', 'Col4']  # Assign manually

Quick Reference

TaskCode
Open PDFpdfplumber.open("file.pdf")
Get pagespdf.pages
Extract tablespage.extract_tables()
Extract textpage.extract_text()
Table to DataFramepd.DataFrame(table[1:], columns=table[0])

Troubleshooting

  • Empty table: Try page.extract_tables(table_settings={...}) with custom settings
  • Merged cells: Tables with merged cells may extract incorrectly
  • Scanned PDFs: Use OCR (pytesseract) for image-based PDFs
  • No tables found: Check if content is actually a table or styled text

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/sales-pivot-analysis/environment/skills/pdf of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

PDF next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF this skillbenchflow-ai/skillsbench1.8k—~1.6kAutomated safety check: PassApache-2.0
File Format Conversionpipeshub-ai/pipeshub-ai3.8k—~865Automated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
CSV Data Summarizerzrt-ai-lab/opencode-skills287—~577Automated safety check: PassNone
Local Knowledge Base RetrieverConardLi/rag-skill7142 repos~1.7kAutomated safety check: PassNone
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • File Format Conversion

    pipeshub-ai/pipeshub-ai

    Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.

    3.8k GitHub stars~865 tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • CSV Data Summarizer

    zrt-ai-lab/opencode-skills

    CSV数据分析技能。使用Python和pandas分析CSV文件,生成统计摘要和快速可视化图表。当用户上传或提到CSV文件、需要分析表格数据时自动使用。

    287 GitHub stars~577 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Answers questions from a local knowledge base folder by walking hierarchical index files and searching with grep, pdfplumber and pandas instead of loading whole files.

    714 GitHub starsUsed in 2 repos~1.7k tokens
    Knowledge ManagementAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed

More from benchflow-ai/skillsbench

All 180 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about PDF

What does PDF do?

ALWAYS use this skill instead of the Read tool for PDF files. PDF is an agent skill from benchflow-ai/skillsbench. ALWAYS use this skill instead of the Read tool for PDF files.

When should I use PDF?

PDF fits situations like: reading ANY PDF file; extracting tables from PDFs; converting PDF tables to pandas DataFrames; processing multi-page PDFs.

How do I install PDF in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill pdf -a claude-code`. Or copy the skill folder (tasks/sales-pivot-analysis/environment/skills/pdf in benchflow-ai/skillsbench) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill pdf -a codex`. Or copy the skill folder (tasks/sales-pivot-analysis/environment/skills/pdf in benchflow-ai/skillsbench) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF need to run?

SKILL.md names no scripts, command-line tools or credentials: PDF is instructions for the agent only. Our summary lists: Python 3.

Does PDF access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF use?

PDF is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF?

Skills that share tags, products or a category with PDF: File Format Conversion (pipeshub-ai/pipeshub-ai, 3.8k stars), PDF Processing (anthropics/skills, 180k stars), CSV Data Summarizer (zrt-ai-lab/opencode-skills, 287 stars) and Local Knowledge Base Retriever (ConardLi/rag-skill, 714 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 180 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.