Official agent skill

PDF to Markdown Converter

by github in github/awesome-copilot

Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately.

OfficialMITAuto-check passedDocuments & Office

Install PDF to Markdown Converter

skills CLI
$ npx skills add github/awesome-copilot --skill convert-pdf-to-md -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot convert-pdf-to-md --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/convert-pdf-to-md .claude/skills/convert-pdf-to-md && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
convert-pdf-to-md
GitHub stars
40k
Token cost
~1.7k tokens
SKILL.md length
813 words
Files
4 (incl. scripts, references)
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately.

  • A user attaches a PDF and asks for a summary or specific answers
  • SKILL.md covers When to use this skill, Setup (once per environment), Usage and Deciding where output goes, plus 1 more section
  • Runs Python scripts from its folder; calls python
  • Extracting tables or data from an invoice, report or contract in PDF form

What it does

PDF is a layout format that cannot be read reliably as plain text, so the agent always runs scripts/convert_pdf_to_md.py first instead of parsing the file itself or writing extraction code. MarkItDown supplies the text and tables, while PyMuPDF extracts real embedded images into a folder per document with an img subfolder. Because inline image positions cannot be recovered safely, the images are listed in an Extracted Images section at the end of the Markdown, with a page subheading for each page that has some.

Setup happens once per environment by following references/setup.md, which installs Python, pip, markitdown and pymupdf. The skill handles only .pdf files and also covers whole folders of PDFs. When a folder mixes PDF, Word and Excel files, the agent must also invoke the sibling convert-word-to-md and convert-excel-to-md skills in parallel so no supported type is skipped.

When your agent uses it

  • A user attaches a PDF and asks for a summary or specific answers
  • Extracting tables or data from an invoice, report or contract in PDF form
  • Processing a whole folder of PDF documents in one go
  • Comparing several PDF documents against each other

Example prompts

  • “Summarize the attached annual-report.pdf in five bullet points.”
  • “Extract the line items from every invoice PDF in ./invoices.”
  • “Compare the two contract PDFs in ./legal and list the differences.”

Requirements

  • Python with pip
  • The markitdown and pymupdf packages

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF to Markdown Converter loads about 1.7k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 235 tokens; SKILL.md has 813 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~235
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 813 words, ~1,715 tokens.

Download SKILL.mdSave it as .claude/skills/convert-pdf-to-md/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
convert-pdf-to-md
description
Converts PDF (.pdf) documents into Markdown so their contents can be accurately analyzed, summarized, searched, or extracted from. Use this skill whenever the user shares, references, or asks about a .pdf file — even if they don't say "convert" or "markdown" explicitly. This includes requests to "read", "summarize", "review", "extract data from", "compare", or "analyze" a PDF report, paper, invoice, form, contract, or scanned document. Always run the bundled conversion script to produce Markdown first; do not attempt to parse PDF content directly or write ad-hoc extraction code. Also use this skill for batch requests involving a whole folder of PDF documents. IMPORTANT: When the user references a folder or set of documents containing multiple file types (.pdf, .docx, .xlsx), invoke ALL three sibling skills — convert-pdf-to-md, convert-word-to-md, and convert-excel-to-md — so no file type is silently skipped.

Convert PDF to Markdown

When to use this skill

Trigger this skill any time there is a .pdf file that needs to be understood or processed — for example, a user attaches a PDF and asks questions about it, wants a summary, wants specific data or tables pulled out, or wants multiple PDFs in a folder processed together. PDF is a layout/print format, not reliably readable as plain text, so always convert it to Markdown first using the script in this skill rather than trying to open or parse the file directly.

This skill only supports .pdf — that's MarkItDown's only PDF-family format, so there's no legacy format to worry about here (unlike Word's .doc or Excel's .xls).

Mixed file types: When the user references a folder or set of documents containing multiple supported file types (.pdf, .docx, .xlsx), this skill handles only .pdf files. The agent MUST also invoke the sibling skills in parallel:

  • convert-word-to-md for any .docx files
  • convert-excel-to-md for any .xlsx files

Never process a folder and silently skip a supported file type. All three skills must be invoked together when mixed types are present.

Setup (once per environment)

Before the first conversion in a given environment, follow references/setup.md step by step to ensure Python, pip, markitdown, and pymupdf (for image extraction) are installed. Do this proactively rather than guessing whether the environment is ready — the script itself will also fail with a clear pointer back to that file if a dependency turns out to be missing, so it's safe to just try the conversion first if you're reasonably confident setup was already done.

Usage

The conversion script lives at scripts/convert_pdf_to_md.py.

Output structure: MarkItDown's PDF converter extracts text and tables only — it has no concept of embedded images at all. This script separately extracts real embedded images via PyMuPDF and writes a self-contained folder per document:

<name>/
    img/
        page001_img001.<ext>
        page002_img001.<ext>
        ...
    <name>.md

Because MarkItDown's PDF text does not preserve reliable per-page markers, there's no safe way to know exactly where inline an image belongs. Rather than risk misplacing images next to the wrong paragraph, the script appends a ## Extracted Images section at the end of the Markdown, with a ### Page N subheading per page that has images — read this section separately from the main body text. If the document has no embedded images, no img/ folder or Extracted Images section is created.

Single file:

powershell
python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf"

This creates a document\ folder next to the source file (containing document.md and, if present, document\img\). To control the destination folder explicitly:

powershell
python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf" -o "C:\path\to\output_folder"

A folder of PDFs (batch mode):

powershell
python scripts\convert_pdf_to_md.py "C:\path\to\folder"

Add --recursive to also include subfolders:

powershell
python scripts\convert_pdf_to_md.py "C:\path\to\folder" --recursive

Each .pdf found gets its own <name>\ output folder next to it by default. Pass -o "C:\path\to\output_parent" to collect all the generated <name>\ folders under a separate parent directory instead (subfolder structure is preserved when combined with --recursive).

After conversion, read the resulting .md file(s) to perform the actual analysis the user asked for — the script's job is only to produce accurate Markdown (and images), not to interpret the content.

Show full SKILL.md (320 more words)Show less

Deciding where output goes

Default — always output next to the source file. The <name>/ folder is created in the same directory as the source .pdf. This is the required default for every case. Do NOT override it unless the user explicitly asks for a different location.

Only use -o when the user explicitly provides an output path (e.g., "save the output to C:\output", "put the results in D:\work"). Do NOT pass -o based on the agent's current working directory, the session state folder, or any implied location.

If the source file path cannot be fully resolved — for example, the user provides only a filename with no directory, or the path is ambiguous — use ask_user to confirm the full absolute path before running the conversion. Never guess or assume the directory.

Troubleshooting

SymptomLikely causeFix
ModuleNotFoundError: No module named 'markitdown' or 'fitz' / exit code 2MarkItDown or PyMuPDF not installedFollow references/setup.md
ERROR: Unsupported file type '...' / exit code 3Not a .pdf fileAsk the user for the correct file, or if it's .doc/.docx/.xlsx, use the matching sibling skill instead
ERROR: Input path not found / exit code 3Wrong path, or file movedConfirm the correct path with the user
FAILED <file> -> ... in batch outputThat specific file is corrupt, password-protected, or otherwise unreadableReport which file(s) failed; other files in the batch still succeed
NOTE: skipped N non-.pdf file(s)Folder contains non-PDF filesExpected — those files are intentionally ignored
Markdown body is empty or near-empty despite images being extractedThe PDF is scanned/image-only with no embedded text layer; MarkItDown does not perform OCRTell the user OCR isn't supported — the extracted page images are still available for them to view
Images appear in an appendix instead of inline with the textDeliberate limitation — MarkItDown's PDF text has no reliable per-page markers to place images inlineExpected behavior; cross-reference the ### Page N heading with the surrounding text context if needed

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/convert-pdf-to-md of github/awesome-copilot.

  • SKILL.md
  • references/setup.md
  • scripts/convert_pdf_to_md.py
  • scripts/requirements.txt

Open the folder on GitHubat commit 727ff2e

Compare with similar skills

PDF to Markdown Converter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF to Markdown Converter compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF to Markdown Converter this skillgithub/awesome-copilot40k—~1.7kAutomated safety check: PassMIT
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Markdropshoryasethia/markdrop211—~1.4kAutomated safety check: NotesGPL-3.0
Markitdownaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing GuideshareAI-lab/learn-claude-code78k5 repos~646Automated safety check: PassMIT

Similar skills

  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdrop

    shoryasethia/markdrop

    Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

    211 GitHub stars~1.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Markitdown

    aipoch/medical-research-skills

    Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

    2k GitHub stars~1.3k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    shareAI-lab/learn-claude-code

    Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents.

    78k GitHub starsUsed in 5 repos~646 tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about PDF to Markdown Converter

What does PDF to Markdown Converter do?

Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately. py first instead of parsing the file itself or writing extraction code. MarkItDown supplies the text and tables, while PyMuPDF extracts real embedded images into a folder per document with an img subfolder.

When should I use PDF to Markdown Converter?

PDF to Markdown Converter fits situations like: A user attaches a PDF and asks for a summary or specific answers; extracting tables or data from an invoice, report or contract in PDF form; processing a whole folder of PDF documents in one go; comparing several PDF documents against each other.

How do I install PDF to Markdown Converter in Claude Code?

Run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a claude-code`. Or copy the skill folder (skills/convert-pdf-to-md in github/awesome-copilot) into .claude/skills/convert-pdf-to-md in your project. Claude Code loads it when a task matches its description.

How do I install PDF to Markdown Converter in Codex?

Run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a codex`. Or copy the skill folder (skills/convert-pdf-to-md in github/awesome-copilot) into .agents/skills/convert-pdf-to-md in your project. Codex loads it when a task matches its description.

Can I use PDF to Markdown Converter in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill convert-pdf-to-md -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/convert-pdf-to-md, .gemini/skills/convert-pdf-to-md, .github/skills/convert-pdf-to-md and .opencode/skills/convert-pdf-to-md in your project.

What does PDF to Markdown Converter need to run?

Going by SKILL.md and its folder, PDF to Markdown Converter needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python with pip; The markitdown and pymupdf packages.

Does PDF to Markdown Converter access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF to Markdown Converter safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF to Markdown Converter use?

PDF to Markdown Converter is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF to Markdown Converter use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 669 tokens, read only when the agent opens those files.

What are the alternatives to PDF to Markdown Converter?

Skills that share tags, products or a category with PDF to Markdown Converter: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Markdrop (shoryasethia/markdrop, 211 stars), Markitdown (aipoch/medical-research-skills, 2k stars) and PDF Processing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF to Markdown Converter?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.