Agent skill

PDF Processor

by gooseworks-ai in gooseworks-ai/goose-skills

Process PDFs - extract text, tables, and structured data from documents

MITAuto-check passedDocuments & Office

Install PDF Processor

skills CLI
$ npx skills add gooseworks-ai/goose-skills --skill pdf-processor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gooseworks-ai/goose-skills pdf-processor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gooseworks-ai/goose-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-tools/capabilities/pdf-processor .claude/skills/pdf-processor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-processor
GitHub stars
1.2k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
122 words
Files
2
Skills in repo
273
Repo updated
First seen
Licence
MIT

At a glance

Process PDFs - extract text, tables, and structured data from documents

  • Works in 4 steps: Fetch PDF Content → Extract with AI → Extract Tables → …
  • Tasks that involve PDF
  • SKILL.md covers Setup, Workflow, Example Usage and Tips, plus 1 more section
  • Calls curl, python3 and npx; reaches api.gooseworks.ai; needs GOOSEWORKS_API_KEY

What it does

PDF Processor is an agent skill from gooseworks-ai/goose-skills. Process PDFs - extract text, tables, and structured data from documents

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.meta.json`).

It sits in Documents & Office, covering PDF, Schema markup and Document parsing. The repository describes itself as: Library of Growth & GTM skills + data APIs for Claude Code, Codex, Cursor to run ads, social, content, lead gen, seo and data scraping. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF
  • Tasks that involve Schema markup
  • Tasks that involve Document parsing

Example prompts

  • “/pdf-processor”

Requirements

  • Python 3
  • Node.js
  • A credential in GOOSEWORKS_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Fetch PDF Content
  2. Extract with AI
  3. Extract Tables
  4. Convert to Markdown

What it can do on your machine

Read from SKILL.md and the folder at commit c650c6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • python3
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.gooseworks.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOSEWORKS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processor loads about 1.1k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 122 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gooseworks-ai/goose-skills at commit c650c6d, republished under its MIT licence (© gooseworks-ai). 122 words, ~1,050 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-processor/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-processor
description
Process PDFs - extract text, tables, and structured data from documents
source
orthogonal

PDF Processor - Extract Data from PDFs

Setup

Read your credentials from ~/.gooseworks/credentials.json:

bash
export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])")
export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))")

If ~/.gooseworks/credentials.json does not exist, tell the user to run: npx gooseworks login

All endpoints use Bearer auth: -H "Authorization: Bearer $GOOSEWORKS_API_KEY"

Extract text, tables, and structured data from PDF documents.

Workflow

Step 1: Fetch PDF Content

Use Linkup to fetch PDF URLs:

bash
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"linkup","path":"/fetch","body":{"url":"https://example.com/document.pdf"}}'
Step 2: Extract with AI

Use ScrapeGraph to extract specific content:

bash
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"scrapegraph","path":"/v1/smartscraper"}'
  "website_url": "https://example.com/report.pdf",
  "user_prompt": "Extract all financial figures, tables, and key metrics from this document"
}'
Step 3: Extract Tables

Get structured table data:

bash
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/run"}'
  "input": {
    "urls": ["https://example.com/report.pdf"]
  },
  "output": {
    "tables": {"prompt": "Extract all tables with titles, headers, and rows", "contexts": ["urls"]}
  }
}'
Step 4: Convert to Markdown

Get readable markdown output:

bash
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"scrapegraph","path":"/v1/markdownify","body":{"website_url":"https://example.com/document.pdf"}}'

Example Usage

bash
# Extract data from financial report
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"scrapegraph","path":"/v1/smartscraper"}'
  "website_url": "https://example.com/annual-report.pdf",
  "user_prompt": "Extract revenue, profit, and key business metrics with their values"
}'

# Extract invoice data
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"riveter","path":"/v1/run"}'
  "input": {"urls": ["https://example.com/invoice.pdf"]},
  "output": {
    "vendor": {"prompt": "Vendor name", "contexts": ["urls"]},
    "amount": {"prompt": "Total amount", "contexts": ["urls"]},
    "date": {"prompt": "Invoice date", "contexts": ["urls"]}
  }
}'

Tips

  • Specify exact data you need for better extraction
  • Use schemas for consistent structured output
  • Handle multi-page documents in chunks
  • Verify extracted numbers against source

Discover More

List all endpoints, or add a path for parameter details:

bash
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"linkup API endpoints"}' api show riveter
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/search \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"scrapegraph API endpoints"}'

Example: `curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/details \
  -H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"api":"olostep","path":"/v1/scrapes`"}' for endpoint parameters.

© gooseworks-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/research-tools/capabilities/pdf-processor of gooseworks-ai/goose-skills.

  • SKILL.md
  • skill.meta.json

Open the folder on GitHubat commit c650c6d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in gooseworks-ai/goose-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

PDF Processor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processor this skillgooseworks-ai/goose-skills1.2k1 repos~1.1kAutomated safety check: PassMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html907—~4.8kAutomated safety check: PassMIT
Lt2mdlibnyx/LT2MD109—~4.4kAutomated safety check: PassAGPL-3.0
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT

Similar skills

  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    907 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Lt2md

    libnyx/LT2MD

    Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…

    109 GitHub stars~4.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Quant Paper Extractor

    CamusGIT/EvoQuant

    Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL…

    151 GitHub stars~2.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from gooseworks-ai/goose-skills

All 273 skills in this repo
  • Reddit Post Finder

    gooseworks-ai/goose-skills

    Scrape and search Reddit posts using Apify. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Image Fal

    gooseworks-ai/goose-skills

    Generate or edit an image via any FAL image model (nano-banana edit, gpt-image, flux, ...), ROUTED THROUGH THE fal-proxy so it bills the Ads agent.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Render Hook Replacement

    gooseworks-ai/goose-skills

    Replace an existing video's opening with a supplied clip or free kinetic text hook while retaining and verifying every original body frame, audio, captions and ending.

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Blog Feed Monitor

    gooseworks-ai/goose-skills

    Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites.

    1.2k GitHub starsUsed in 1 repo~578 tokens
    Auto-check passed
  • Competitor Post Engagers

    gooseworks-ai/goose-skills

    Find leads by scraping engagers from a competitor's top LinkedIn posts.

    1.2k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Render Chatgpt Chat

    gooseworks-ai/goose-skills

    Assemble a ChatGPT chat-reveal video ad from a thread + timeline JSON — one continuous Playwright recording of a ChatGPT mobile chat (user types with the iOS keyboard up → taps send → keyboard…

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about PDF Processor

What does PDF Processor do?

Process PDFs - extract text, tables, and structured data from documents. PDF Processor is an agent skill from gooseworks-ai/goose-skills.

When should I use PDF Processor?

PDF Processor fits situations like: tasks that involve PDF; tasks that involve Schema markup; tasks that involve Document parsing.

How do I install PDF Processor in Claude Code?

Run `npx skills add gooseworks-ai/goose-skills --skill pdf-processor -a claude-code`. Or copy the skill folder (skills/research-tools/capabilities/pdf-processor in gooseworks-ai/goose-skills) into .claude/skills/pdf-processor in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processor in Codex?

Run `npx skills add gooseworks-ai/goose-skills --skill pdf-processor -a codex`. Or copy the skill folder (skills/research-tools/capabilities/pdf-processor in gooseworks-ai/goose-skills) into .agents/skills/pdf-processor in your project. Codex loads it when a task matches its description.

Can I use PDF Processor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gooseworks-ai/goose-skills --skill pdf-processor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-processor, .gemini/skills/pdf-processor, .github/skills/pdf-processor and .opencode/skills/pdf-processor in your project.

What does PDF Processor need to run?

Going by SKILL.md and its folder, PDF Processor needs the command-line tools its instructions call (curl, python3 and npx) and credentials named GOOSEWORKS_API_KEY. Our summary lists: Python 3; Node.js; A credential in GOOSEWORKS_API_KEY.

Does PDF Processor access the network?

SKILL.md names 1 domain. In commands or code: api.gooseworks.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is PDF Processor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Processor use?

PDF Processor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Processor use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Processor?

Skills that share tags, products or a category with PDF Processor: Markitdown (jimmc414/Kosmos, 595 stars), Markitdown (ImCa0/just-laws, 782 stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 907 stars) and Lt2md (libnyx/LT2MD, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processor?

gooseworks-ai (a GitHub organization) maintains it in gooseworks-ai/goose-skills, which has 1,239 GitHub stars. The repository holds 273 skills in this directory. The repository was last updated on October 8, 2026.

Source: gooseworks-ai/goose-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.