Agent skill

Edgeparse

by raphaelmansuy in raphaelmansuy/edgeparse

Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.

Apache-2.0Auto-check passedDocuments & Office

Install Edgeparse

skills CLI
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/edgeparse .claude/skills/edgeparse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
edgeparse
GitHub stars
143
Token cost
~1.8k tokens
SKILL.md length
281 words
Files
3 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.

  • The user wants to read
  • SKILL.md covers When to reach for this skill, Quick start, Core API and Common patterns, plus 4 more sections
  • Calls pip and npm
  • Reason about a PDF document

What it does

Edgeparse is an agent skill from raphaelmansuy/edgeparse. Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills. Use this skill whenever the user wants to read, analyze, or reason about a PDF document; needs to feed document content to an LLM; mentions PDF extraction, parsing, or conversion; wants tables, headings, or bounding boxes from a PDF; is building a RAG pipeline; or asks an agent to process a document. Install with: pip install edgeparse

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api.md` and `references/patterns.md`).

It sits in Documents & Office, covering PDF and Retrieval-augmented generation. It works with Node.js and MarkItDown. The repository describes itself as: EdgeParse converts any digital PDF into Markdown, JSON (with bounding boxes), HTML, or plain text — deterministically, without a JVM, without a GPU, and with best-in-class… The licence is Apache-2.0.

When your agent uses it

  • The user wants to read
  • Reason about a PDF document
  • Needs to feed document content to an LLM
  • Mentions PDF extraction

Example prompts

  • “/edgeparse”

Requirements

  • Python 3
  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit ccd10f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Edgeparse loads about 1.8k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 281 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from raphaelmansuy/edgeparse at commit ccd10f0, republished under its Apache-2.0 licence (© raphaelmansuy). 281 words, ~1,770 tokens.

Download SKILL.mdSave it as .claude/skills/edgeparse/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
edgeparse
description
Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills. Use this skill whenever the user wants to read, analyze, or reason about a PDF document; needs to feed document content to an LLM; mentions PDF extraction, parsing, or conversion; wants tables, headings, or bounding boxes from a PDF; is building a RAG pipeline; or asks an agent to process a document. Install with: pip install edgeparse
license
Apache-2.0
metadata
authors: "EdgeParse Contributors" version: "0.1.0" package: "edgeparse" install_python: "pip install edgeparse" install_node: "npm install edgeparse" source…

EdgeParse Skill

Enables AI agents to extract clean, structured content from any PDF — headings, tables, paragraphs, lists, bounding boxes — deterministically, without ML dependencies or GPU requirements.

Install: pip install edgeparse · Node.js: npm install edgeparse
Speed: ~0.023 s/doc (Apple M4 Max, 200-doc benchmark)


When to reach for this skill

Activate when the workflow involves:

  • Reading or analyzing a PDF document on behalf of a user
  • Building a RAG pipeline that ingests PDFs
  • Feeding PDF content to an LLM for summarization, Q&A, or synthesis
  • Extracting tables from financial reports, research papers, or invoices
  • Processing a batch of documents for indexing or search
  • An agent tool that must "open" a PDF and return its contents

Quick start

python
import edgeparse

# Convert any PDF to Markdown — best for LLM context windows
text = edgeparse.convert("report.pdf", format="markdown")

# Convert to JSON with bounding boxes and full structure
import json
doc = json.loads(edgeparse.convert("report.pdf", format="json"))

# Plain text (fast, minimal)
plain = edgeparse.convert("report.pdf", format="text")

The format parameter controls output:

ValueBest for
"markdown"LLM context — headings, tables, lists in Markdown
"json"Bounding boxes, citations, structured element metadata
"html"Web rendering, semantic HTML5
"text"Simple full-text search, minimal output

Core API

edgeparse.convert()
python
result: str = edgeparse.convert(
    input_path,             # str or Path — required
    format="markdown",      # output format (see table above)
    pages=None,             # e.g. "1-5" or "1,3,7-10" — specific pages only
    password=None,          # for password-protected PDFs
    reading_order="xycut",  # "xycut" (spatial sort, default) or "off"
    table_method="default", # "default" (ruling-line) or "cluster" (borderless)
    image_output="off",     # "off", "embedded" (base64), "external" (files)
)

Returns the extracted content as a string. Raises FileNotFoundError for missing files and ValueError for corrupt PDFs or bad options.

edgeparse.convert_file()
python
out_path: str = edgeparse.convert_file(
    input_path,
    output_dir="output",    # write output file to this directory
    format="markdown",
    pages=None,
    password=None,
)

Writes the output file and returns its path.


Common patterns

Feed a PDF to an LLM
python
import edgeparse
import anthropic

doc = edgeparse.convert("report.pdf", format="markdown")

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-opus-4-5",
    max_tokens=4096,
    messages=[{
        "role": "user",
        "content": f"Analyze this document and summarize the key findings:\n\n{doc}"
    }]
)
print(response.content[0].text)
RAG pipeline — chunk with metadata
python
import edgeparse, json

raw = edgeparse.convert("paper.pdf", format="json")
doc = json.loads(raw)

chunks = []
for el in doc["elements"]:
    if el["type"] in ("paragraph", "heading", "table"):
        chunks.append({
            "text": el["text"],
            "metadata": {
                "page":    el["page_number"],
                "type":    el["type"],
                "bbox":    el["bounding_box"],   # for citation highlights
                "order":   el["reading_order"],
            }
        })

# Now embed chunks["text"] and store chunks["metadata"] in your vector store
Batch processing
python
import edgeparse
from pathlib import Path

results = {}
for pdf in Path("documents/").glob("*.pdf"):
    try:
        results[pdf.name] = edgeparse.convert(str(pdf), format="markdown")
    except Exception as e:
        results[pdf.name] = f"ERROR: {e}"
Extract specific pages only
python
# Pages 1–5
text = edgeparse.convert("report.pdf", format="markdown", pages="1-5")

# Non-contiguous pages
text = edgeparse.convert("report.pdf", format="markdown", pages="1,3,7-10")
Borderless table extraction

Many financial reports and invoices use tables without ruling lines. Use table_method="cluster" to handle them:

python
text = edgeparse.convert(
    "earnings.pdf",
    format="markdown",
    table_method="cluster"   # spatial clustering for borderless tables
)
Password-protected PDF
python
text = edgeparse.convert("secure.pdf", format="markdown", password="mypassword")

Node.js usage

js
import { convert } from 'edgeparse';

const markdown = convert('report.pdf', { format: 'markdown' });
const json     = convert('report.pdf', { format: 'json' });

// With options
const result = convert('report.pdf', {
    format:       'markdown',
    pages:        '1-5',
    readingOrder: 'xycut',
    tableMethod:  'cluster',
});

JSON output schema

When format="json", the output is a JSON string with shape:

json
{
  "page_count": 10,
  "title": "Document Title",
  "elements": [
    {
      "type": "heading",
      "level": 1,
      "text": "Introduction",
      "page_number": 1,
      "reading_order": 0,
      "bounding_box": { "x0": 72, "y0": 144, "x1": 540, "y1": 180 }
    },
    {
      "type": "table",
      "text": "| Col A | Col B |\n|-------|-------|\n| val1  | val2  |",
      "page_number": 2,
      "bounding_box": { "x0": 72, "y0": 200, "x1": 540, "y1": 350 }
    },
    {
      "type": "paragraph",
      "text": "This is body text...",
      "page_number": 1,
      "reading_order": 2,
      "bounding_box": { "x0": 72, "y0": 190, "x1": 540, "y1": 220 }
    }
  ]
}

Element type values: heading, paragraph, table, list, list_item, figure, caption, header, footer.


Error handling

python
import edgeparse

try:
    text = edgeparse.convert("report.pdf", format="markdown")
except FileNotFoundError:
    # PDF file not found — check the path
    pass
except ValueError as e:
    # Invalid format, corrupt PDF, wrong password, or bad page range
    print(f"Extraction failed: {e}")

For more detail

Read these reference files when the SKILL.md body isn't enough:

  • references/api.md — complete Python + Node.js API with all parameters and types
  • references/patterns.md — LangChain, LlamaIndex, MCP tool, CrewAI, and async batch patterns

© raphaelmansuy, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/edgeparse of raphaelmansuy/edgeparse.

  • SKILL.md
  • references/api.md
  • references/patterns.md

Open the folder on GitHubat commit ccd10f0

Compare with similar skills

Edgeparse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Edgeparse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Edgeparse this skillraphaelmansuy/edgeparse143—~1.8kAutomated safety check: PassApache-2.0
Markitdownaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
Md2pdf Exportwentorai/Research-Claw858—~2.1kAutomated safety check: NotesCustom licence
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Markdropshoryasethia/markdrop211—~1.4kAutomated safety check: NotesGPL-3.0
Streaming Export Safetydoccker/cc-use-exp1.1k—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Markitdown

    aipoch/medical-research-skills

    Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

    2k GitHub stars~1.3k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed
  • Md2pdf Export

    wentorai/Research-Claw

    Convert Markdown files to PDF, PNG, or JPEG using headless Chrome (Puppeteer), following the same methodology as VS Code's Markdown Preview Enhanced extension.

    858 GitHub stars~2.1k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdrop

    shoryasethia/markdrop

    Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

    211 GitHub stars~1.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Streaming Export Safety

    doccker/cc-use-exp

    当代码涉及 Excel/CSV/JSON/PDF 大文件导出、批量序列化、内存里构建大对象时触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程等内存安全陷阱。

    1.1k GitHub stars~1.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Opendataloader PDF

    chujianyun/skills

    PDF 数据提取工具。当用户提到"PDF 提取"、"PDF 转 Markdown"、"PDF 解析"、"提取 PDF 内容"、"PDF 转 JSON"、"RAG PDF"时使用。OpenDataLoader PDF 是目前基准测试第一的 PDF 解析器,支持本地模式(快速、确定)和混合 AI 模式(复杂表格、扫描件、公式),输出 Markdown、JSON(带边界框)、HTML。适用于需要从…

    740 GitHub stars~827 tokensUpdated 12 days ago
    Documents & OfficeAuto-check passed

Questions about Edgeparse

What does Edgeparse do?

Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills. Edgeparse is an agent skill from raphaelmansuy/edgeparse. Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.

When should I use Edgeparse?

Edgeparse fits situations like: the user wants to read; reason about a PDF document; needs to feed document content to an LLM; mentions PDF extraction.

How do I install Edgeparse in Claude Code?

Run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a claude-code`. Or copy the skill folder (skills/edgeparse in raphaelmansuy/edgeparse) into .claude/skills/edgeparse in your project. Claude Code loads it when a task matches its description.

How do I install Edgeparse in Codex?

Run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a codex`. Or copy the skill folder (skills/edgeparse in raphaelmansuy/edgeparse) into .agents/skills/edgeparse in your project. Codex loads it when a task matches its description.

Can I use Edgeparse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/edgeparse, .gemini/skills/edgeparse, .github/skills/edgeparse and .opencode/skills/edgeparse in your project.

What does Edgeparse need to run?

Going by SKILL.md and its folder, Edgeparse needs the command-line tools its instructions call (pip and npm). Our summary lists: Python 3; Node.js.

Does Edgeparse access the network?

SKILL.md contains no URLs. Its commands use pip and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Edgeparse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Edgeparse use?

Edgeparse is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Edgeparse use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Edgeparse?

Skills that share tags, products or a category with Edgeparse: Markitdown (aipoch/medical-research-skills, 2k stars), Md2pdf Export (wentorai/Research-Claw, 858 stars), Markdown Converter (Team-Commonly/commonly, 1.4k stars) and Markdrop (shoryasethia/markdrop, 211 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Edgeparse?

raphaelmansuy (a GitHub user) maintains it in raphaelmansuy/edgeparse, which has 143 GitHub stars. The repository was last updated on October 2, 2026.

Source: raphaelmansuy/edgeparse on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.