Markitdown
aipoch/medical-research-skills
Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).
Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/edgeparse .claude/skills/edgeparse && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .claude/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/edgeparse .agents/skills/edgeparse && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .agents/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/edgeparse .cursor/skills/edgeparse && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .cursor/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/raphaelmansuy/edgeparse.git --path skills/edgeparse--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/edgeparse .gemini/skills/edgeparse && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .gemini/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install raphaelmansuy/edgeparse edgeparseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/edgeparse .github/skills/edgeparse && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .github/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add raphaelmansuy/edgeparse --skill edgeparse -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install raphaelmansuy/edgeparse edgeparse --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/raphaelmansuy/edgeparse.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/edgeparse .opencode/skills/edgeparse && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "edgeparse" agent skill from https://github.com/raphaelmansuy/edgeparse/tree/main/skills/edgeparse into .opencode/skills/edgeparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "edgeparse", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
edgeparseExtract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.
Edgeparse is an agent skill from raphaelmansuy/edgeparse. Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills. Use this skill whenever the user wants to read, analyze, or reason about a PDF document; needs to feed document content to an LLM; mentions PDF extraction, parsing, or conversion; wants tables, headings, or bounding boxes from a PDF; is building a RAG pipeline; or asks an agent to process a document. Install with: pip install edgeparse
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api.md` and `references/patterns.md`).
It sits in Documents & Office, covering PDF and Retrieval-augmented generation. It works with Node.js and MarkItDown. The repository describes itself as: EdgeParse converts any digital PDF into Markdown, JSON (with bounding boxes), HTML, or plain text — deterministically, without a JVM, without a GPU, and with best-in-class… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit ccd10f0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip and npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Edgeparse loads about 1.8k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 281 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from raphaelmansuy/edgeparse at commit ccd10f0, republished under its Apache-2.0 licence (© raphaelmansuy). 281 words, ~1,770 tokens.
.claude/skills/edgeparse/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Enables AI agents to extract clean, structured content from any PDF — headings, tables, paragraphs, lists, bounding boxes — deterministically, without ML dependencies or GPU requirements.
Install: pip install edgeparse · Node.js: npm install edgeparse
Speed: ~0.023 s/doc (Apple M4 Max, 200-doc benchmark)
Activate when the workflow involves:
import edgeparse
# Convert any PDF to Markdown — best for LLM context windows
text = edgeparse.convert("report.pdf", format="markdown")
# Convert to JSON with bounding boxes and full structure
import json
doc = json.loads(edgeparse.convert("report.pdf", format="json"))
# Plain text (fast, minimal)
plain = edgeparse.convert("report.pdf", format="text")The format parameter controls output:
| Value | Best for |
|---|---|
"markdown" | LLM context — headings, tables, lists in Markdown |
"json" | Bounding boxes, citations, structured element metadata |
"html" | Web rendering, semantic HTML5 |
"text" | Simple full-text search, minimal output |
edgeparse.convert()result: str = edgeparse.convert(
input_path, # str or Path — required
format="markdown", # output format (see table above)
pages=None, # e.g. "1-5" or "1,3,7-10" — specific pages only
password=None, # for password-protected PDFs
reading_order="xycut", # "xycut" (spatial sort, default) or "off"
table_method="default", # "default" (ruling-line) or "cluster" (borderless)
image_output="off", # "off", "embedded" (base64), "external" (files)
)Returns the extracted content as a string. Raises FileNotFoundError for missing files and ValueError for corrupt PDFs or bad options.
edgeparse.convert_file()out_path: str = edgeparse.convert_file(
input_path,
output_dir="output", # write output file to this directory
format="markdown",
pages=None,
password=None,
)Writes the output file and returns its path.
import edgeparse
import anthropic
doc = edgeparse.convert("report.pdf", format="markdown")
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=4096,
messages=[{
"role": "user",
"content": f"Analyze this document and summarize the key findings:\n\n{doc}"
}]
)
print(response.content[0].text)import edgeparse, json
raw = edgeparse.convert("paper.pdf", format="json")
doc = json.loads(raw)
chunks = []
for el in doc["elements"]:
if el["type"] in ("paragraph", "heading", "table"):
chunks.append({
"text": el["text"],
"metadata": {
"page": el["page_number"],
"type": el["type"],
"bbox": el["bounding_box"], # for citation highlights
"order": el["reading_order"],
}
})
# Now embed chunks["text"] and store chunks["metadata"] in your vector storeimport edgeparse
from pathlib import Path
results = {}
for pdf in Path("documents/").glob("*.pdf"):
try:
results[pdf.name] = edgeparse.convert(str(pdf), format="markdown")
except Exception as e:
results[pdf.name] = f"ERROR: {e}"# Pages 1–5
text = edgeparse.convert("report.pdf", format="markdown", pages="1-5")
# Non-contiguous pages
text = edgeparse.convert("report.pdf", format="markdown", pages="1,3,7-10")Many financial reports and invoices use tables without ruling lines.
Use table_method="cluster" to handle them:
text = edgeparse.convert(
"earnings.pdf",
format="markdown",
table_method="cluster" # spatial clustering for borderless tables
)text = edgeparse.convert("secure.pdf", format="markdown", password="mypassword")import { convert } from 'edgeparse';
const markdown = convert('report.pdf', { format: 'markdown' });
const json = convert('report.pdf', { format: 'json' });
// With options
const result = convert('report.pdf', {
format: 'markdown',
pages: '1-5',
readingOrder: 'xycut',
tableMethod: 'cluster',
});When format="json", the output is a JSON string with shape:
{
"page_count": 10,
"title": "Document Title",
"elements": [
{
"type": "heading",
"level": 1,
"text": "Introduction",
"page_number": 1,
"reading_order": 0,
"bounding_box": { "x0": 72, "y0": 144, "x1": 540, "y1": 180 }
},
{
"type": "table",
"text": "| Col A | Col B |\n|-------|-------|\n| val1 | val2 |",
"page_number": 2,
"bounding_box": { "x0": 72, "y0": 200, "x1": 540, "y1": 350 }
},
{
"type": "paragraph",
"text": "This is body text...",
"page_number": 1,
"reading_order": 2,
"bounding_box": { "x0": 72, "y0": 190, "x1": 540, "y1": 220 }
}
]
}Element type values: heading, paragraph, table, list, list_item, figure, caption, header, footer.
import edgeparse
try:
text = edgeparse.convert("report.pdf", format="markdown")
except FileNotFoundError:
# PDF file not found — check the path
pass
except ValueError as e:
# Invalid format, corrupt PDF, wrong password, or bad page range
print(f"Extraction failed: {e}")Read these reference files when the SKILL.md body isn't enough:
references/api.md — complete Python + Node.js API with all parameters and typesreferences/patterns.md — LangChain, LlamaIndex, MCP tool, CrewAI, and async batch patterns© raphaelmansuy, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/edgeparse of raphaelmansuy/edgeparse.
Open the folder on GitHubat commit ccd10f0
Edgeparse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Edgeparse this skillraphaelmansuy/edgeparse | 143 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Markitdownaipoch/medical-research-skills | 2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Md2pdf Exportwentorai/Research-Claw | 858 | — | ~2.1k | Automated safety check: Notes | Custom licence | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| Markdropshoryasethia/markdrop | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | |
| Streaming Export Safetydoccker/cc-use-exp | 1.1k | — | ~1.9k | Automated safety check: Pass | Custom licence |
aipoch/medical-research-skills
Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).
wentorai/Research-Claw
Convert Markdown files to PDF, PNG, or JPEG using headless Chrome (Puppeteer), following the same methodology as VS Code's Markdown Preview Enhanced extension.
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
shoryasethia/markdrop
Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
doccker/cc-use-exp
当代码涉及 Excel/CSV/JSON/PDF 大文件导出、批量序列化、内存里构建大对象时触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程等内存安全陷阱。
chujianyun/skills
PDF 数据提取工具。当用户提到"PDF 提取"、"PDF 转 Markdown"、"PDF 解析"、"提取 PDF 内容"、"PDF 转 JSON"、"RAG PDF"时使用。OpenDataLoader PDF 是目前基准测试第一的 PDF 解析器,支持本地模式(快速、确定)和混合 AI 模式(复杂表格、扫描件、公式),输出 Markdown、JSON(带边界框)、HTML。适用于需要从…
Works with
Categories
Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills. Edgeparse is an agent skill from raphaelmansuy/edgeparse. Extract structured content from any PDF for AI agents, RAG pipelines, and Copilot Skills.
Edgeparse fits situations like: the user wants to read; reason about a PDF document; needs to feed document content to an LLM; mentions PDF extraction.
Run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a claude-code`. Or copy the skill folder (skills/edgeparse in raphaelmansuy/edgeparse) into .claude/skills/edgeparse in your project. Claude Code loads it when a task matches its description.
Run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a codex`. Or copy the skill folder (skills/edgeparse in raphaelmansuy/edgeparse) into .agents/skills/edgeparse in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add raphaelmansuy/edgeparse --skill edgeparse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/edgeparse, .gemini/skills/edgeparse, .github/skills/edgeparse and .opencode/skills/edgeparse in your project.
Going by SKILL.md and its folder, Edgeparse needs the command-line tools its instructions call (pip and npm). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use pip and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Edgeparse is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Edgeparse: Markitdown (aipoch/medical-research-skills, 2k stars), Md2pdf Export (wentorai/Research-Claw, 858 stars), Markdown Converter (Team-Commonly/commonly, 1.4k stars) and Markdrop (shoryasethia/markdrop, 211 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
raphaelmansuy (a GitHub user) maintains it in raphaelmansuy/edgeparse, which has 143 GitHub stars. The repository was last updated on October 2, 2026.
Source: raphaelmansuy/edgeparse on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.