Agent skill

PDF Analysis

by OpenSenseNova in OpenSenseNova/SenseNova-Skills

PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。

MITAuto-check passedDocuments & Office

Install PDF Analysis

skills CLI
$ npx skills add OpenSenseNova/SenseNova-Skills --skill pdf-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenSenseNova/SenseNova-Skills pdf-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenSenseNova/SenseNova-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sn-da-non-spreadsheet-analysis/capability/pdf-analysis .claude/skills/pdf-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-analysis
GitHub stars
5.7k
Token cost
~2.5k tokens
SKILL.md length
228 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。

  • Tasks that involve PDF
  • SKILL.md covers Step 0 — Detect PDF type (text…, Core Method 1: Text PDF — Full…, Core Method 2: Text PDF —… and Core Method 3: Scanned PDF —…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

PDF Analysis is an agent skill from OpenSenseNova/SenseNova-Skills. PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: Modular SenseNova skills for building AI-powered office assistants and productivity workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-analysis”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 5abde96. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Analysis loads about 2.5k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 228 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenSenseNova/SenseNova-Skills at commit 5abde96, republished under its MIT licence (© OpenSenseNova). 228 words, ~2,460 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-analysis/SKILL.md (or your agent's skills folder).
name
pdf-analysis
description
PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。

PDF Analysis

Step 0 — Detect PDF type (text vs scanned)

Critical first step: determine whether the PDF has extractable text or is a scanned image. Never skip this — using the wrong parser wastes time and produces empty results.

python
import fitz  # PyMuPDF

def detect_pdf_type(pdf_path, sample_pages=3):
    """
    Returns 'text' if PDF has extractable text, 'scanned' if image-based.
    Checks first N pages (or all if fewer).
    """
    doc = fitz.open(pdf_path)
    total_chars = 0
    pages_checked = min(sample_pages, len(doc))

    for i in range(pages_checked):
        page = doc[i]
        text = page.get_text("text")
        total_chars += len(text.strip())

    doc.close()
    avg_chars = total_chars / max(pages_checked, 1)
    pdf_type = 'text' if avg_chars > 50 else 'scanned'
    print(f"PDF type: {pdf_type} (avg {avg_chars:.0f} chars/page, checked {pages_checked} pages)")
    return pdf_type

Core Method 1: Text PDF — Full Text Extraction (ALL pages)

python
import fitz

def extract_text_pdf(pdf_path):
    """Extract text from all pages of a text-based PDF."""
    doc = fitz.open(pdf_path)
    total_pages = len(doc)
    print(f"Total pages: {total_pages}")

    all_text = []
    for i, page in enumerate(doc):
        text = page.get_text("text").strip()
        if text:
            all_text.append(f"=== Page {i+1} ===\n{text}")
        else:
            print(f"  Page {i+1}: no text (may be image — will caption later)")

    doc.close()
    return '\n\n'.join(all_text)

# ⚠️ MUST iterate ALL pages — never stop at page 1
full_text = extract_text_pdf(pdf_path)
print(f"Total text length: {len(full_text)} chars")

Core Method 2: Text PDF — Table Extraction

For PDFs with tables, pdfplumber gives better table structure than fitz:

python
import pdfplumber
import pandas as pd

def extract_tables_pdf(pdf_path):
    """Extract all tables from all pages as DataFrames."""
    all_tables = []
    with pdfplumber.open(pdf_path) as pdf:
        print(f"Total pages: {len(pdf.pages)}")
        for i, page in enumerate(pdf.pages):
            tables = page.extract_tables()
            for j, tbl in enumerate(tables):
                if not tbl:
                    continue
                # First row as header
                df = pd.DataFrame(tbl[1:], columns=tbl[0])
                # Clean: strip whitespace, replace None
                df = df.applymap(lambda x: x.strip() if isinstance(x, str) else x)
                df = df.dropna(how='all').reset_index(drop=True)
                all_tables.append({'page': i+1, 'table_idx': j, 'df': df})
                print(f"  Page {i+1}, Table {j}: {df.shape[0]}r × {df.shape[1]}c")
                print(df.head(3))
    return all_tables

# Verify table alignment after extraction:
# Print column headers and first 3 rows to confirm row/col mapping is correct

Core Method 3: Scanned PDF — OCR via Caption

For scanned PDFs (image-based pages), render each page as PNG and caption:

python
import fitz
import subprocess, json, os

CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"

def extract_scanned_pdf(pdf_path, prompt=None, dpi=150):
    """Render each page as image, then caption for text extraction."""
    doc = fitz.open(pdf_path)
    total_pages = len(doc)
    print(f"Scanned PDF: {total_pages} pages, captioning each...")

    all_text = []
    for i, page in enumerate(doc):
        # Render page to PNG
        mat = fitz.Matrix(dpi/72, dpi/72)
        pix = page.get_pixmap(matrix=mat)
        img_path = f"/tmp/pdf_page_{i+1}.png"
        pix.save(img_path)

        # Caption the page image
        cmd = ["python3", CAPTION, img_path, "--json"]
        if prompt:
            cmd += ["--prompt", prompt]
        else:
            cmd += ["--prompt", "提取页面中所有文字和表格内容,保持原始结构,Markdown格式输出。"]

        r = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
        if r.returncode == 0:
            desc = json.loads(r.stdout).get("description", "")
            all_text.append(f"=== Page {i+1} ===\n{desc}")
            print(f"  Page {i+1}: {len(desc)} chars extracted")
        else:
            print(f"  Page {i+1}: caption failed — {r.stderr[:100]}")

    doc.close()
    return '\n\n'.join(all_text)

# Usage for scanned invoice PDFs, bank statements, org charts, etc.
text = extract_scanned_pdf(pdf_path)

Core Method 4: Hybrid PDF (mixed text + image pages)

python
def extract_hybrid_pdf(pdf_path, text_prompt=None, image_prompt=None):
    """Handle PDFs where some pages have text, others are scanned."""
    doc_fitz = fitz.open(pdf_path)
    all_text = []

    for i, page in enumerate(doc_fitz):
        raw_text = page.get_text("text").strip()

        if len(raw_text) > 50:
            # Text page — use directly
            all_text.append(f"=== Page {i+1} (text) ===\n{raw_text}")
        else:
            # Image page — render and caption
            mat = fitz.Matrix(150/72, 150/72)
            pix = page.get_pixmap(matrix=mat)
            img_path = f"/tmp/hybrid_page_{i+1}.png"
            pix.save(img_path)

            cmd = ["python3", CAPTION, img_path, "--json"]
            prompt = image_prompt or "提取页面中所有文字和表格内容,Markdown格式输出。"
            cmd += ["--prompt", prompt]

            r = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
            if r.returncode == 0:
                desc = json.loads(r.stdout).get("description", "")
                all_text.append(f"=== Page {i+1} (image→caption) ===\n{desc}")
            else:
                all_text.append(f"=== Page {i+1} (caption failed) ===")

    doc_fitz.close()
    return '\n\n'.join(all_text)

Core Method 5: Extract Embedded Images / Charts from PDF

python
import fitz

def extract_pdf_images(pdf_path, min_width=100, min_height=100):
    """Extract all embedded images from a PDF (charts, diagrams, photos)."""
    doc = fitz.open(pdf_path)
    image_paths = []

    for page_num, page in enumerate(doc):
        for img_idx, img in enumerate(page.get_images(full=True)):
            xref = img[0]
            base = doc.extract_image(xref)
            img_bytes = base["image"]
            ext = base["ext"]

            img_path = f"/tmp/pdf_img_p{page_num+1}_{img_idx}.{ext}"
            with open(img_path, 'wb') as f:
                f.write(img_bytes)

            # Only keep images above size threshold (skip icons/logos)
            from PIL import Image
            with Image.open(img_path) as im:
                w, h = im.size
            if w >= min_width and h >= min_height:
                image_paths.append({'page': page_num+1, 'path': img_path, 'size': (w, h)})
                print(f"  Page {page_num+1}, img {img_idx}: {w}×{h} → {img_path}")

    doc.close()
    return image_paths

# After extracting, caption each image:
# for img_info in image_paths:
#     caption_image(img_info['path'], prompt="提取图表数据,Markdown 表格输出。")

Common Patterns

Multi-invoice / multi-document PDF (发票汇总)
python
# When PDF contains multiple invoices (one per page):
tables_by_page = extract_tables_pdf(pdf_path)
invoices = []
for item in tables_by_page:
    df = item['df']
    # Find key fields (flexible column name matching)
    for col in df.columns:
        if '金额' in str(col) or 'amount' in str(col).lower():
            invoices.append({'page': item['page'], 'amount_col': col, 'data': df})
            break
print(f"Found {len(invoices)} pages with amount data")
Numeric extraction with unit awareness
python
import re

def extract_number_with_unit(text_snippet):
    """
    Extract value and unit from text like '1,760 千港元' or '95,975,196,217.52元'.
    Returns (numeric_value, unit_string).
    """
    # Remove thousands separator
    text_snippet = text_snippet.replace(',', '')
    match = re.search(r'([\d\.]+)\s*(千|万|亿|百万)?\s*(元|港元|美元|人民币|%|percent)?', text_snippet)
    if not match:
        return None, None
    value = float(match.group(1))
    multiplier_map = {'千': 1000, '万': 10000, '亿': 1e8, '百万': 1e6}
    mult = multiplier_map.get(match.group(2), 1)
    unit = match.group(3) or ''
    return value * mult, f"{match.group(2) or ''}{unit}"

# Always verify unit matches what the question asks:
# "多几多" in HKD → answer in 千港元 if source says 千港元
python
def find_in_pdf(pdf_path, keyword, context_chars=200):
    """Search for keyword across all pages, return context snippets."""
    text = extract_text_pdf(pdf_path)
    results = []
    start = 0
    while True:
        idx = text.find(keyword, start)
        if idx < 0:
            break
        snippet = text[max(0, idx-context_chars//2): idx+context_chars]
        results.append({'pos': idx, 'context': snippet})
        start = idx + 1
    print(f"Found '{keyword}' {len(results)} times")
    return results

Pitfalls

PitfallFix
Use pdfplumber on scanned PDF → empty resultDetect type first (Method 0); use OCR path for scanned
Only read page 1, miss remaining invoices/dataAlways for page in doc — never index [0] only
Table columns misaligned after extractionPrint headers + first 3 rows to verify before computing
Report number as % when question asks absolute valueRead question carefully; extract_number_with_unit() preserves context
Chart data embedded as image → pdfplumber returns nothingExtract images (Method 5), then caption each
Long doc loses cross-page contextUse find_in_pdf() for keyword search across full text
.pdf contains multiple scanned docs (zip of PDFs)Check if input is dir or archive; unzip first

© OpenSenseNova, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/sn-da-non-spreadsheet-analysis/capability/pdf-analysis of OpenSenseNova/SenseNova-Skills.

Open the folder on GitHubat commit 5abde96

Compare with similar skills

PDF Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Analysis this skillOpenSenseNova/SenseNova-Skills5.7k—~2.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.8k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.8k GitHub stars~19k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from OpenSenseNova/SenseNova-Skills

All 36 skills in this repo
  • SN Motion HTML

    OpenSenseNova/SenseNova-Skills

    Builds HTML stories where one continuous camera journey advances with page progress, using researched structure, AI stills, Seedance video clips and browser QA.

    5.7k GitHub stars~2.2k tokensUpdated 20 days ago
    Auto-check: notes
  • SenseNova PPT Fallback Tools

    OpenSenseNova/SenseNova-Skills

    Fallback scripts for web search, image search and download, and image generation that PPT skills use only when the host agent lacks or fails its own tools.

    5.7k GitHub stars~575 tokensUpdated 20 days ago
    Auto-check: notes
  • SenseNova PPT Workbench

    OpenSenseNova/SenseNova-Skills

    Opens the PPT Workbench web editor for an existing SenseNova HTML slide deck so you can preview, inspect and visually edit it without regenerating.

    5.7k GitHub stars~2.5k tokensUpdated 20 days ago
    Auto-check: notes
  • SenseNova PPT Creative Renderer

    OpenSenseNova/SenseNova-Skills

    Turns an approved slide outline into a full-page image for every slide, one 16:9 PNG per page, and optionally packages the set into a PPTX.

    5.7k GitHub stars~1.2k tokensUpdated 20 days ago
    Auto-check passed
  • SenseNova PPT Entry

    OpenSenseNova/SenseNova-Skills

    Entry point for SenseNova presentation generation: creates a task folder, picks depth, output format and design richness, and routes to the right PPT skill.

    5.7k GitHub stars~2.7k tokensUpdated 20 days ago
    Auto-check: notes
  • China Market Open Data Search

    OpenSenseNova/SenseNova-Skills

    Researches Chinese market, macro, trade, procurement, listed-company and regulatory information from free official sources that need no sign-up or API key.

    5.7k GitHub stars~954 tokensUpdated 20 days ago
    Auto-check: notes

Questions about PDF Analysis

What does PDF Analysis do?

PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。. PDF Analysis is an agent skill from OpenSenseNova/SenseNova-Skills.

When should I use PDF Analysis?

PDF Analysis fits situations like: tasks that involve PDF.

How do I install PDF Analysis in Claude Code?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill pdf-analysis -a claude-code`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/pdf-analysis in OpenSenseNova/SenseNova-Skills) into .claude/skills/pdf-analysis in your project. Claude Code loads it when a task matches its description.

How do I install PDF Analysis in Codex?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill pdf-analysis -a codex`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/pdf-analysis in OpenSenseNova/SenseNova-Skills) into .agents/skills/pdf-analysis in your project. Codex loads it when a task matches its description.

Can I use PDF Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenSenseNova/SenseNova-Skills --skill pdf-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-analysis, .gemini/skills/pdf-analysis, .github/skills/pdf-analysis and .opencode/skills/pdf-analysis in your project.

What does PDF Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: PDF Analysis is instructions for the agent only. Our summary lists: Python 3.

Does PDF Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Analysis use?

PDF Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Analysis use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Analysis?

Skills that share tags, products or a category with PDF Analysis: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Analysis?

OpenSenseNova (a GitHub organization) maintains it in OpenSenseNova/SenseNova-Skills, which has 5,743 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 18, 2026.

Source: OpenSenseNova/SenseNova-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.