Markdown Converter
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately.
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/convert-pdf-to-md .claude/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .claude/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-mdType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/convert-pdf-to-md .agents/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .agents/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/convert-pdf-to-md .cursor/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .cursor/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/github/awesome-copilot.git --path skills/convert-pdf-to-md--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/convert-pdf-to-md .gemini/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .gemini/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install github/awesome-copilot convert-pdf-to-mdInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/convert-pdf-to-md .github/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .github/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/convert-pdf-to-md .opencode/skills/convert-pdf-to-md && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "convert-pdf-to-md" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/convert-pdf-to-md into .opencode/skills/convert-pdf-to-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "convert-pdf-to-md", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
convert-pdf-to-mdConverts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately.
PDF is a layout format that cannot be read reliably as plain text, so the agent always runs scripts/convert_pdf_to_md.py first instead of parsing the file itself or writing extraction code. MarkItDown supplies the text and tables, while PyMuPDF extracts real embedded images into a folder per document with an img subfolder. Because inline image positions cannot be recovered safely, the images are listed in an Extracted Images section at the end of the Markdown, with a page subheading for each page that has some.
Setup happens once per environment by following references/setup.md, which installs Python, pip, markitdown and pymupdf. The skill handles only .pdf files and also covers whole folders of PDFs. When a folder mixes PDF, Word and Excel files, the agent must also invoke the sibling convert-word-to-md and convert-excel-to-md skills in parallel so no supported type is skipped.
Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF to Markdown Converter loads about 1.7k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 235 tokens; SKILL.md has 813 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 813 words, ~1,715 tokens.
.claude/skills/convert-pdf-to-md/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Trigger this skill any time there is a .pdf file that needs to be
understood or processed — for example, a user attaches a PDF and asks
questions about it, wants a summary, wants specific data or tables pulled
out, or wants multiple PDFs in a folder processed together. PDF is a
layout/print format, not reliably readable as plain text, so always convert
it to Markdown first using the script in this skill rather than trying to
open or parse the file directly.
This skill only supports .pdf — that's MarkItDown's only PDF-family
format, so there's no legacy format to worry about here (unlike Word's
.doc or Excel's .xls).
Mixed file types: When the user references a folder or set of documents
containing multiple supported file types (.pdf, .docx, .xlsx), this
skill handles only .pdf files. The agent MUST also invoke the sibling
skills in parallel:
convert-word-to-md for any .docx filesconvert-excel-to-md for any .xlsx filesNever process a folder and silently skip a supported file type. All three skills must be invoked together when mixed types are present.
Before the first conversion in a given environment, follow
references/setup.md step by step to ensure Python,
pip, markitdown, and pymupdf (for image extraction) are installed. Do
this proactively rather than guessing whether the environment is ready — the
script itself will also fail with a clear pointer back to that file if a
dependency turns out to be missing, so it's safe to just try the conversion
first if you're reasonably confident setup was already done.
The conversion script lives at scripts/convert_pdf_to_md.py.
Output structure: MarkItDown's PDF converter extracts text and tables only — it has no concept of embedded images at all. This script separately extracts real embedded images via PyMuPDF and writes a self-contained folder per document:
<name>/
img/
page001_img001.<ext>
page002_img001.<ext>
...
<name>.mdBecause MarkItDown's PDF text does not preserve reliable per-page markers,
there's no safe way to know exactly where inline an image belongs. Rather
than risk misplacing images next to the wrong paragraph, the script appends
a ## Extracted Images section at the end of the Markdown, with a
### Page N subheading per page that has images — read this section
separately from the main body text. If the document has no embedded images,
no img/ folder or Extracted Images section is created.
Single file:
python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf"This creates a document\ folder next to the source file (containing
document.md and, if present, document\img\). To control the destination
folder explicitly:
python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf" -o "C:\path\to\output_folder"A folder of PDFs (batch mode):
python scripts\convert_pdf_to_md.py "C:\path\to\folder"Add --recursive to also include subfolders:
python scripts\convert_pdf_to_md.py "C:\path\to\folder" --recursiveEach .pdf found gets its own <name>\ output folder next to it by
default. Pass -o "C:\path\to\output_parent" to collect all the generated
<name>\ folders under a separate parent directory instead (subfolder
structure is preserved when combined with --recursive).
After conversion, read the resulting .md file(s) to perform the actual
analysis the user asked for — the script's job is only to produce accurate
Markdown (and images), not to interpret the content.
Default — always output next to the source file. The <name>/ folder
is created in the same directory as the source .pdf. This is the required
default for every case. Do NOT override it unless the user explicitly asks
for a different location.
Only use -o when the user explicitly provides an output path (e.g.,
"save the output to C:\output", "put the results in D:\work"). Do NOT
pass -o based on the agent's current working directory, the session state
folder, or any implied location.
If the source file path cannot be fully resolved — for example, the
user provides only a filename with no directory, or the path is ambiguous —
use ask_user to confirm the full absolute path before running the
conversion. Never guess or assume the directory.
| Symptom | Likely cause | Fix |
|---|---|---|
ModuleNotFoundError: No module named 'markitdown' or 'fitz' / exit code 2 | MarkItDown or PyMuPDF not installed | Follow references/setup.md |
ERROR: Unsupported file type '...' / exit code 3 | Not a .pdf file | Ask the user for the correct file, or if it's .doc/.docx/.xlsx, use the matching sibling skill instead |
ERROR: Input path not found / exit code 3 | Wrong path, or file moved | Confirm the correct path with the user |
FAILED <file> -> ... in batch output | That specific file is corrupt, password-protected, or otherwise unreadable | Report which file(s) failed; other files in the batch still succeed |
NOTE: skipped N non-.pdf file(s) | Folder contains non-PDF files | Expected — those files are intentionally ignored |
| Markdown body is empty or near-empty despite images being extracted | The PDF is scanned/image-only with no embedded text layer; MarkItDown does not perform OCR | Tell the user OCR isn't supported — the extracted page images are still available for them to view |
| Images appear in an appendix instead of inline with the text | Deliberate limitation — MarkItDown's PDF text has no reliable per-page markers to place images inline | Expected behavior; cross-reference the ### Page N heading with the surrounding text context if needed |
© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts, references) in skills/convert-pdf-to-md of github/awesome-copilot.
Open the folder on GitHubat commit 727ff2e
PDF to Markdown Converter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF to Markdown Converter this skillgithub/awesome-copilot | 40k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| Markdropshoryasethia/markdrop | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | |
| Markitdownaipoch/medical-research-skills | 2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| PDF Processinganthropics/skills | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | |
| PDF Processing GuideshareAI-lab/learn-claude-code | 78k | 5 repos | ~646 | Automated safety check: Pass | MIT |
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
shoryasethia/markdrop
Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
aipoch/medical-research-skills
Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).
anthropics/skills
Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.
shareAI-lab/learn-claude-code
Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents.
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
github/awesome-copilot
Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
github/awesome-copilot
Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
Works with
Categories
Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately. py first instead of parsing the file itself or writing extraction code. MarkItDown supplies the text and tables, while PyMuPDF extracts real embedded images into a folder per document with an img subfolder.
PDF to Markdown Converter fits situations like: A user attaches a PDF and asks for a summary or specific answers; extracting tables or data from an invoice, report or contract in PDF form; processing a whole folder of PDF documents in one go; comparing several PDF documents against each other.
Run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a claude-code`. Or copy the skill folder (skills/convert-pdf-to-md in github/awesome-copilot) into .claude/skills/convert-pdf-to-md in your project. Claude Code loads it when a task matches its description.
Run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a codex`. Or copy the skill folder (skills/convert-pdf-to-md in github/awesome-copilot) into .agents/skills/convert-pdf-to-md in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/convert-pdf-to-md, .gemini/skills/convert-pdf-to-md, .github/skills/convert-pdf-to-md and .opencode/skills/convert-pdf-to-md in your project.
Going by SKILL.md and its folder, PDF to Markdown Converter needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python with pip; The markitdown and pymupdf packages.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PDF to Markdown Converter is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 669 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with PDF to Markdown Converter: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Markdrop (shoryasethia/markdrop, 211 stars), Markitdown (aipoch/medical-research-skills, 2k stars) and PDF Processing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.
Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.