Agent skill

Markitdown

by aipoch in aipoch/medical-research-skills

Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

MITAuto-check passedDocuments & Office

Install Markitdown

skills CLI
$ npx skills add aipoch/medical-research-skills --skill markitdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills markitdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scientific-skills/Other/markitdown .claude/skills/markitdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
markitdown
GitHub stars
2k
Token cost
~1.3k tokens
SKILL.md length
419 words
Files
8 (incl. scripts, references, assets)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

  • Tasks that involve Document parsing
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Runs Python scripts from its folder; calls pip and markitdown; reaches openrouter.ai
  • Tasks that involve Summarization

What it does

Markitdown is an agent skill from aipoch/medical-research-skills. Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts, reference files and assets (for example `assets/example_usage.md`, `markitdown_audit_result_v1.json` and `references/api_reference.md`).

It sits in Documents & Office, covering Document parsing, Summarization and PDF. It works with MarkItDown, Python, OpenAI and OpenRouter. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Document parsing
  • Tasks that involve Summarization
  • Tasks that involve PDF

Example prompts

  • “/markitdown”

Requirements

  • Python 3
  • A credential in YOUR_OPENROUTER_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • markitdown

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Markitdown loads about 1.3k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 47 tokens; SKILL.md has 419 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 419 words, ~1,258 tokens.

Download SKILL.mdSave it as .claude/skills/markitdown/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
markitdown
description
Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

  • Converting research papers or reports (PDF/DOCX/EPUB/HTML) into Markdown for LLM summarization, Q&A, or RAG indexing.
  • Extracting tables and structured content from spreadsheets (XLSX/CSV) into Markdown for analysis or documentation.
  • Turning slide decks (PPTX) into Markdown notes, including speaker notes and (optionally) AI-generated image descriptions.
  • Processing images or scanned documents with OCR to obtain searchable, editable Markdown text.
  • Transcribing audio (WAV/MP3) or pulling YouTube transcripts into Markdown for meeting notes, content analysis, or knowledge bases.

Key Features

  • Converts many formats to structured Markdown (PDF, DOCX, PPTX, XLSX, images, audio, HTML, CSV, JSON, XML, ZIP, EPUB, YouTube URLs, etc.).
  • Produces token-efficient output suitable for LLM pipelines (summarization, chunking, embedding).
  • OCR support for images/scans (when OCR dependencies are installed).
  • Audio transcription support (when transcription dependencies are installed).
  • Optional AI-enhanced image/slide descriptions via an OpenAI-compatible client (e.g., OpenRouter).
  • Plugin system to extend format support and custom behaviors.
  • Stream-based conversion API for large files.

Dependencies

  • Python: >=3.9 (recommended)
  • Package:
    • markitdown[all] (installs all optional format handlers)

Optional system dependencies (feature-dependent):

  • Tesseract OCR: tesseract-ocr (for image/scanned-text OCR)

Optional external services (feature-dependent):

  • Azure Document Intelligence endpoint (for enhanced PDF extraction)
  • OpenAI-compatible LLM endpoint (e.g., OpenRouter) for AI image descriptions

Example Usage

Install
bash
pip install 'markitdown[all]'
CLI: Convert a PDF to Markdown
bash
markitdown document.pdf -o output.md
Python: Convert multiple formats (PDF/XLSX/PPTX/DOCX) and save outputs
python
from pathlib import Path
from markitdown import MarkItDown

md = MarkItDown()

files = [
    "document.pdf",
    "spreadsheet.xlsx",
    "presentation.pptx",
    "notes.docx",
]

for path in files:
    result = md.convert(path)
    out = Path(path).with_suffix(".md")
    out.write_text(result.text_content, encoding="utf-8")
    print(f"Converted {path} -> {out}")
Python: Stream conversion (useful for large files)
python
from markitdown import MarkItDown

md = MarkItDown()

with open("large_file.pdf", "rb") as f:
    result = md.convert_stream(f, file_extension=".pdf")

with open("large_file.md", "w", encoding="utf-8") as out:
    out.write(result.text_content)
Python: AI-enhanced image/slide descriptions (OpenAI-compatible, e.g., OpenRouter)
python
from markitdown import MarkItDown
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_OPENROUTER_API_KEY",
    base_url="https://openrouter.ai/api/v1",
)

md = MarkItDown(
    llm_client=client,
    llm_model="anthropic/claude-opus-4.5",
    llm_prompt="Describe this image in detail for scientific documentation.",
)

result = md.convert("presentation.pptx")
print(result.text_content)
Show full SKILL.md (192 more words)Show less

Implementation Details

  • Conversion entry points

    • MarkItDown().convert(path) converts a file by path/URL and returns an object whose primary payload is result.text_content (Markdown).
    • MarkItDown().convert_stream(stream, file_extension=".pdf") converts from a binary stream; use this for large files or when data is not on disk.
  • Format handling

    • Format support is provided by optional extras (e.g., pdf, docx, pptx, xlsx, audio-transcription, youtube-transcription) or all.
    • ZIP inputs are typically processed by iterating through contained files and converting each supported entry.
  • OCR

    • For images/scanned documents, OCR is enabled when OCR tooling is available (commonly Tesseract). Ensure the OS-level OCR binary is installed and accessible in PATH.
  • AI image descriptions

    • When llm_client, llm_model, and llm_prompt are provided, MarkItDown can request model-generated descriptions for images (including slide images), then inject those descriptions into the Markdown output.
    • Any OpenAI-compatible client can be used (e.g., OpenRouter) by setting base_url and api_key.
  • Enhanced PDF extraction (Azure Document Intelligence)

    • When configured with a Document Intelligence endpoint, PDF extraction can be improved for complex layouts (tables, multi-column text, scanned PDFs), producing more faithful Markdown structure.
  • Plugins

    • Plugins can be listed and enabled from the CLI (e.g., --list-plugins, --use-plugins) to extend conversion behavior or add new format handlers.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references, assets) in scientific-skills/Other/markitdown of aipoch/medical-research-skills.

  • SKILL.md
  • assets/example_usage.md
  • markitdown_audit_result_v1.json
  • references/api_reference.md
  • references/file_formats.md
  • scripts/batch_convert.py
  • scripts/convert_literature.py
  • scripts/convert_with_ai.py

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Markitdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Markitdown compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Markitdown this skillaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Document ConverterBlackBeltTechnology/pi-agent-dashboard315—~999Automated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    315 GitHub stars~999 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Office File Process

    xstongxue/best-skills

    处理 Office 文档的一站式 skill:Word(.doc/.docx/.dotx)、Excel(.xls/.xlsx/.xlsm/.csv)、PowerPoint(.ppt/.pptx/.potx) 的创建、读取、编辑、提取、转换、校验。触发:『读取 word 文档』『提取 excel 内容』『看 ppt 讲了什么』、.doc 老格式打不开、生成/编辑 Word…

    3k GitHub stars~1.8k tokensUpdated 25 days ago
    Documents & OfficeAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Markitdown

What does Markitdown do?

Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation). Markitdown is an agent skill from aipoch/medical-research-skills., for summarization, search, RAG ingestion, or dataset preparation).

When should I use Markitdown?

Markitdown fits situations like: tasks that involve Document parsing; tasks that involve Summarization; tasks that involve PDF.

How do I install Markitdown in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill markitdown -a claude-code`. Or copy the skill folder (scientific-skills/Other/markitdown in aipoch/medical-research-skills) into .claude/skills/markitdown in your project. Claude Code loads it when a task matches its description.

How do I install Markitdown in Codex?

Run `npx skills add aipoch/medical-research-skills --skill markitdown -a codex`. Or copy the skill folder (scientific-skills/Other/markitdown in aipoch/medical-research-skills) into .agents/skills/markitdown in your project. Codex loads it when a task matches its description.

Can I use Markitdown in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill markitdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/markitdown, .gemini/skills/markitdown, .github/skills/markitdown and .opencode/skills/markitdown in your project.

What does Markitdown need to run?

Going by SKILL.md and its folder, Markitdown needs Python for the scripts in its folder and the command-line tools its instructions call (pip and markitdown). Our summary lists: Python 3; A credential in YOUR_OPENROUTER_API_KEY.

Does Markitdown access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Markitdown safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Markitdown use?

Markitdown is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Markitdown use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Markitdown?

Skills that share tags, products or a category with Markitdown: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Document Converter (BlackBeltTechnology/pi-agent-dashboard, 315 stars), Markitdown (ImCa0/just-laws, 781 stars) and Markitdown (jimmc414/Kosmos, 594 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Markitdown?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.