Agent skill

PDF To Markdown

by claesbackman in claesbackman/AI-research-feedback

Split a PDF into chunks and convert it to readable markdown text.

MITAuto-check passedDocuments & Office

Install PDF To Markdown

skills CLI
$ npx skills add claesbackman/AI-research-feedback --skill pdf-to-markdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install claesbackman/AI-research-feedback pdf-to-markdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/claesbackman/AI-research-feedback.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills/pdf-to-markdown .claude/skills/pdf-to-markdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-to-markdown
GitHub stars
495
Token cost
~1.8k tokens
SKILL.md length
648 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Split a PDF into chunks and convert it to readable markdown text.

  • Works in 2 steps: If the path is relative, resolve it… → Run mdls -name kMDItemNumberOfPages ""…
  • The user wants to read
  • SKILL.md covers Input, Choose the method first, Resolve the path and get page… and Method A — pdftotext (preferred), plus 3 more sections
  • Calls pdftotext

What it does

PDF To Markdown is an agent skill from claesbackman/AI-research-feedback. Split a PDF into chunks and convert it to readable markdown text. Use when the user wants to read, extract, or convert a PDF document.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: A collection of Claude Code skills for academic research review. These tools were developed by Claes Bäckman. The licence is MIT.

When your agent uses it

  • The user wants to read
  • Convert a PDF document

Example prompts

  • “/pdf-to-markdown”

Requirements

  • Pre-approved tools (allowed-tools): Read, Bash(mdls *), Bash(pdftotext *), Bash(which *), Write

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. If the path is relative, resolve it relative to the current working directory.
  2. Run mdls -name kMDItemNumberOfPages "" to get the total page count. If mdls is unavailable or returns (null), use pdftotext or Read to…

What it can do on your machine

Read from SKILL.md and the folder at commit d129756. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(mdls *)
    • Bash(pdftotext *)
    • Bash(which *)
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF To Markdown loads about 1.8k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 648 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from claesbackman/AI-research-feedback at commit d129756, republished under its MIT licence (© claesbackman). 648 words, ~1,797 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-to-markdown/SKILL.md (or your agent's skills folder).
name
pdf-to-markdown
description
Split a PDF into chunks and convert it to readable markdown text. Use when the user wants to read, extract, or convert a PDF document.
allowed-tools
Read, Bash(mdls *), Bash(pdftotext *), Bash(which *), Write
user-invocable
true
argument-hint
path/to/file.pdf

PDF Split & Convert

Convert a PDF file to readable markdown text. Handles large PDFs efficiently.

Input

  • $ARGUMENTS[0] — Path to the PDF file (required)

Choose the method first

Before reading anything, check whether pdftotext (part of poppler) is available:

bash
which pdftotext
  • If available → use the pdftotext path below. It is 10–100× faster than the Read tool, uses no model context for the text content, and doesn't suffer from stream idle timeouts. Always prefer this for PDFs longer than ~30 pages.
  • If not available → fall back to the Read-tool path. Warn the user that large PDFs (>40 pages) may hit stream idle timeouts when run in subagents (~12 min cap). Prefer running in the main conversation for large files.

Resolve the path and get page count

  1. If the path is relative, resolve it relative to the current working directory.
  2. Run mdls -name kMDItemNumberOfPages "<pdf_path>" to get the total page count. If mdls is unavailable or returns (null), use pdftotext or Read to probe.

Method A — pdftotext (preferred)

Extract the full PDF to text in one shot, then trim at references/appendix and add page markers. pdftotext emits a form-feed character (\f) at every page break — use that for pagination.

Reference implementation (bash + awk):

bash
PDF="$1"
OUT="${PDF%.pdf}.md"
TITLE=$(basename "${PDF%.pdf}")
TMPTXT=$(mktemp)

pdftotext -layout "$PDF" "$TMPTXT"

awk -v title="$TITLE" '
BEGIN {
    print "# " title
    print ""
    page = 1
    printf "---\n## Pages %d-%d\n---\n", page, page+19
    next_marker = page + 20
}
{
    # Convert form-feed page breaks to newlines and count pages
    n = gsub(/\f/, "\n")
    if (n > 0) {
        page += n
        if (page >= next_marker) {
            printf "\n---\n## Pages %d-%d\n---\n", next_marker, next_marker+19
            next_marker += 20
        }
    }

    # Build a stripped copy for heading detection.
    # CRITICAL: strip both form feeds AND embedded newlines — gsub above inserts
    # newlines into $0, which will defeat regex anchors like ^ and $ if you skip this.
    stripped = $0
    gsub(/[\f\n]/, "", stripped)
    sub(/^[ \t]+/, "", stripped)
    sub(/[ \t]+$/, "", stripped)

    if (length(stripped) > 0) {
        # References / Bibliography — standalone word, optionally numbered, short, no prose punctuation
        if (length(stripped) < 50 && stripped !~ /[(),;]/) {
            if (stripped ~ /^([0-9]+\.?[ \t]+)?(References|REFERENCES|Bibliography|BIBLIOGRAPHY)$/) exit
            if (stripped == "Works Cited") exit
        }

        # Appendix — length up to ~120 chars (some titles are long), no prose punctuation
        if (length(stripped) < 120 && stripped !~ /[(),;]/) {
            # "Appendix A" alone (bare letter, no title)
            if (stripped ~ /^Appendix[ \t]+[A-Z][0-9]*$/) exit
            # "Appendix A. Title" or "Appendix A: Title" — punctuation REQUIRED to avoid
            # matching body-text references like "Appendix H examines the effect..."
            if (stripped ~ /^Appendix[ \t]+[A-Z][0-9]*[.:][ \t]+[A-Z].*$/) exit
            # "APPENDIX A" variants
            if (stripped ~ /^APPENDIX[ \t]+[A-Z][0-9]*([ \t]+.*)?$/) exit
            # "Online Appendix [A]"
            if (stripped ~ /^Online[ \t]+Appendix([ \t]+[A-Z].*)?$/) exit
            # "Supplemental/Supplementary/Internet Appendix"
            if (stripped ~ /^(Supplement(al|ary)|Internet)[ \t]+Appendix([ \t]+.*)?$/) exit
        }
    }

    print
}
' "$TMPTXT" > "$OUT"

rm -f "$TMPTXT"
echo "Wrote $OUT ($(wc -l < "$OUT") lines)"

After running, sanity-check the output:

bash
# Should print nothing if trimming worked
grep -cE "^[[:space:]]*(References|Bibliography|REFERENCES|BIBLIOGRAPHY)[[:space:]]*$" "$OUT"

# Eyeball last few lines — should be prose/conclusion, not body-of-table or mid-paragraph
tail -5 "$OUT"

If the tail looks truncated mid-paragraph, the heading detection likely fired on a false positive. If the tail shows references or appendix content, the detection missed the heading — inspect the PDF text around that area and extend the regex.

Method B — Read tool (fallback when pdftotext unavailable)

  1. Read the PDF in chunks of up to 20 pages at a time using the Read tool's pages parameter: 1-20, 21-40, 41-60, etc.
  2. Focus on the MAIN TEXT ONLY. Stop including content once you hit "References", "Bibliography", "Works Cited", or an appendix section. If references appear mid-chunk, keep everything before them and drop the rest.
  3. Compile output:
    • Save alongside the PDF with a .md extension.
    • Header: # [Original Filename]
    • Page markers between chunks: ---\n## Pages X-Y\n---
    • Preserve extracted text as-is.

Warning: The Read tool is slow for large PDFs (roughly 30–60 seconds per 20-page chunk). A 60-page paper can take 3–4 minutes, and sub-agents have a ~13-minute stream idle timeout that this can hit. When running a batch of conversions, run them sequentially in the main conversation or use Method A.

Show full SKILL.md (263 more words)Show less

Heading-detection pitfalls (hard-won lessons)

These false positives broke earlier attempts — keep them in mind whether you use Method A or B:

  • Parenthetical references in body text like (see Appendix B.5) — exclude lines containing (, ), ,, or ;.
  • Body text starting with "Appendix X ..." like Appendix H examines the differential effect... — require punctuation (. or :) immediately after the appendix letter when a title follows. Bare "Appendix A" alone on a line is still valid.
  • Line-wrapped headings — pdftotext can wrap Online Appendix across two lines if the PDF's layout is unusual. You'll see Online on one line and Appendix on the next. The regex above matches the joined form; if you see false trims at a lone Appendix line, inspect and tighten.
  • Form-feed at start of page — pdftotext emits \f as the first character on every new page. After gsub(/\f/, "\n") on $0, the line has an embedded \n that defeats ^ / $ anchors unless you also gsub(/[\f\n]/, "", stripped) on your detection copy.
  • Length thresholds — simple "< 50 chars" is too tight for appendix titles. "Appendix A. Merging Mortgages with the Real Estate Database" is 59 chars. Use ~120 for appendix patterns, ~50 for bare References.
  • Numbered section headings — some papers format as 7 References or 7. References. Allow an optional leading number.

Report results

Tell the user:

  • Method used (pdftotext vs Read tool)
  • Total pages processed
  • Output file path and final line count
  • Where trimming occurred (last section / page number included)
  • Any pages that were unreadable or empty

For batch conversions, print a summary table and note any files whose trim point looks suspicious (very short output, or output ending mid-sentence).

© claesbackman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in Skills/pdf-to-markdown of claesbackman/AI-research-feedback.

Open the folder on GitHubat commit d129756

Compare with similar skills

PDF To Markdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF To Markdown compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF To Markdown this skillclaesbackman/AI-research-feedback495—~1.8kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3151 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    315 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from claesbackman/AI-research-feedback

All 10 skills in this repo
  • Explorable Deck

    claesbackman/AI-research-feedback

    Build a Quarto reveal.js slide deck in the explorable-explanation style (Nicky Case) — one idea per slide, assertion titles, a concrete running example, run-time SVG stages the presenter drives…

    495 GitHub stars~2.9k tokensUpdated 14 days ago
    Auto-check passed
  • Review Paper Light

    claesbackman/AI-research-feedback

    Run a fast 2-agent pre-submission check for an economics paper — focuses on contribution, identification, and causal overclaiming.

    495 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check: notes
  • Paper Version

    claesbackman/AI-research-feedback

    Convert a LaTeX research paper into a policy brief, 1-page summary, or 5-page summary for a general audience, with factual review and a standalone HTML page for GitHub Pages.

    495 GitHub stars~4k tokensUpdated 14 days ago
    Auto-check: notes
  • Review Grant

    claesbackman/AI-research-feedback

    Run a 6-agent pre-submission panel review for a grant proposal targeting a specified funder or program

    495 GitHub starsUsed in 1 repo~5.6k tokens
    Auto-check: notes
  • Review Paper Checks

    claesbackman/AI-research-feedback

    Run a fast 3-agent mechanical check of an economics paper — spelling and grammar, internal consistency and cross-references, and unsupported claims.

    495 GitHub stars~4.3k tokensUpdated 14 days ago
    Auto-check: notes
  • Review Pap

    claesbackman/AI-research-feedback

    Run a 6-agent pre-submission review of a pre-analysis plan (PAP) for a specified registration target or journal

    495 GitHub starsUsed in 1 repo~6.4k tokens
    Auto-check: notes

Questions about PDF To Markdown

What does PDF To Markdown do?

Split a PDF into chunks and convert it to readable markdown text. PDF To Markdown is an agent skill from claesbackman/AI-research-feedback. Split a PDF into chunks and convert it to readable markdown text.

When should I use PDF To Markdown?

PDF To Markdown fits situations like: the user wants to read; convert a PDF document.

How do I install PDF To Markdown in Claude Code?

Run `npx skills add claesbackman/AI-research-feedback --skill pdf-to-markdown -a claude-code`. Or copy the skill folder (Skills/pdf-to-markdown in claesbackman/AI-research-feedback) into .claude/skills/pdf-to-markdown in your project. Claude Code loads it when a task matches its description.

How do I install PDF To Markdown in Codex?

Run `npx skills add claesbackman/AI-research-feedback --skill pdf-to-markdown -a codex`. Or copy the skill folder (Skills/pdf-to-markdown in claesbackman/AI-research-feedback) into .agents/skills/pdf-to-markdown in your project. Codex loads it when a task matches its description.

Can I use PDF To Markdown in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add claesbackman/AI-research-feedback --skill pdf-to-markdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-to-markdown, .gemini/skills/pdf-to-markdown, .github/skills/pdf-to-markdown and .opencode/skills/pdf-to-markdown in your project.

What does PDF To Markdown need to run?

Going by SKILL.md and its folder, PDF To Markdown needs the command-line tools its instructions call (pdftotext). Its frontmatter pre-approves these tools: Read, Bash(mdls *), Bash(pdftotext *), Bash(which *), Write.

Does PDF To Markdown access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF To Markdown safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF To Markdown use?

PDF To Markdown is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF To Markdown use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF To Markdown?

Skills that share tags, products or a category with PDF To Markdown: Markitdown (ImCa0/just-laws, 782 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF To Markdown?

claesbackman (a GitHub user) maintains it in claesbackman/AI-research-feedback, which has 495 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 25, 2026.

Source: claesbackman/AI-research-feedback on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.