Agent skill

Markitdown

by ImCa0 in ImCa0/just-laws

Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

MITAuto-check: notesDocuments & Office

Install Markitdown

skills CLI
$ npx skills add ImCa0/just-laws --skill markitdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ImCa0/just-laws markitdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ImCa0/just-laws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/markitdown .claude/skills/markitdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
markitdown
GitHub stars
781
Used in
14 other repos
Token cost
~3.2k tokens
SKILL.md length
601 words
Files
13 (incl. scripts, references, assets)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

  • Works in 12 steps: AI-Enhanced Image Descriptions → Azure Document Intelligence → Plugin System → …
  • Tasks that involve Document parsing
  • SKILL.md covers Overview, Visual Enhancement with…, Supported Formats and Quick Start, plus 7 more sections
  • Runs Python scripts from its folder; calls markitdown, pip and docker; reaches openrouter.ai and github.com

What it does

Markitdown is an agent skill from ImCa0/just-laws. Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts, reference files and assets (for example `INSTALLATION_GUIDE.md`, `OPENROUTER_INTEGRATION.md` and `QUICK_REFERENCE.md`).

It sits in Documents & Office, covering Document parsing, PowerPoint presentations and Transcription. It works with MarkItDown, Microsoft PowerPoint, Microsoft Excel and YouTube. The repository describes itself as: 一个简洁、便捷的中国法律文库 | A Simple and Convenient Laws Library of China. The licence is MIT.

When your agent uses it

  • Tasks that involve Document parsing
  • Tasks that involve PowerPoint presentations
  • Tasks that involve Transcription

Example prompts

  • “/markitdown”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. AI-Enhanced Image Descriptions
  2. Azure Document Intelligence
  3. Plugin System
  4. Convert Scientific Papers to Markdown
  5. Extract Data from Excel for Analysis
  6. Process Multiple Documents
  7. Convert PowerPoint with AI Descriptions
  8. Batch Convert with Different Formats
  9. Extract YouTube Video Transcription
  10. Choose the Right Conversion Method
  11. Handle Errors Gracefully
  12. Process Large Files Efficiently

What it can do on your machine

Read from SKILL.md and the folder at commit b947e4e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • markitdown
    • pip
    • docker
    • python
    • git
    • brew
    • apt-get

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai
    • github.com
    • youtube.com

    Also links to:

    • pypi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Markitdown loads about 3.2k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 601 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:460
    sudo apt-get install tesseract-ocr
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ImCa0/just-laws at commit b947e4e, republished under its MIT licence (© ImCa0). 601 words, ~3,160 tokens.

Download SKILL.mdSave it as .claude/skills/markitdown/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
markitdown
description
Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.
allowed-tools
Read, Write, Edit, Bash
license
MIT
source
https://github.com/microsoft/markitdown

MarkItDown - File to Markdown Conversion

Overview

MarkItDown is a Python tool developed by Microsoft for converting various file formats to Markdown. It's particularly useful for converting documents into LLM-friendly text format, as Markdown is token-efficient and well-understood by modern language models.

Key Benefits:

  • Convert documents to clean, structured Markdown
  • Token-efficient format for LLM processing
  • Supports 15+ file formats
  • Optional AI-enhanced image descriptions
  • OCR for images and scanned documents
  • Speech transcription for audio files

Visual Enhancement with Scientific Schematics

When creating documents with this skill, always consider adding scientific diagrams and schematics to enhance visual communication.

If your document does not already contain schematics or diagrams:

  • Use the scientific-schematics skill to generate AI-powered publication-quality diagrams
  • Simply describe your desired diagram in natural language
  • Nano Banana Pro will automatically generate, review, and refine the schematic

For new documents: Scientific schematics should be generated by default to visually represent key concepts, workflows, architectures, or relationships described in the text.

How to generate schematics:

bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png

The AI will automatically:

  • Create publication-quality images with proper formatting
  • Review and refine through multiple iterations
  • Ensure accessibility (colorblind-friendly, high contrast)
  • Save outputs in the figures/ directory

When to add schematics:

  • Document conversion workflow diagrams
  • File format architecture illustrations
  • OCR processing pipeline diagrams
  • Integration workflow visualizations
  • System architecture diagrams
  • Data flow diagrams
  • Any complex concept that benefits from visualization

For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.


Supported Formats

FormatDescriptionNotes
PDFPortable Document FormatFull text extraction
DOCXMicrosoft WordTables, formatting preserved
PPTXPowerPointSlides with notes
XLSXExcel spreadsheetsTables and data
ImagesJPEG, PNG, GIF, WebPEXIF metadata + OCR
AudioWAV, MP3Metadata + transcription
HTMLWeb pagesClean conversion
CSVComma-separated valuesTable format
JSONJSON dataStructured representation
XMLXML documentsStructured format
ZIPArchive filesIterates contents
EPUBE-booksFull text extraction
YouTubeVideo URLsFetch transcriptions

Quick Start

Installation
bash
# Install with all features
pip install 'markitdown[all]'

# Or from source
git clone https://github.com/microsoft/markitdown.git
cd markitdown
pip install -e 'packages/markitdown[all]'
Command-Line Usage
bash
# Basic conversion
markitdown document.pdf > output.md

# Specify output file
markitdown document.pdf -o output.md

# Pipe content
cat document.pdf | markitdown > output.md

# Enable plugins
markitdown --list-plugins  # List available plugins
markitdown --use-plugins document.pdf -o output.md
Python API
python
from markitdown import MarkItDown

# Basic usage
md = MarkItDown()
result = md.convert("document.pdf")
print(result.text_content)

# Convert from stream
with open("document.pdf", "rb") as f:
    result = md.convert_stream(f, file_extension=".pdf")
    print(result.text_content)

Advanced Features

1. AI-Enhanced Image Descriptions

Use LLMs via OpenRouter to generate detailed image descriptions (for PPTX and image files):

python
from markitdown import MarkItDown
from openai import OpenAI

# Initialize OpenRouter client (OpenAI-compatible API)
client = OpenAI(
    api_key="your-openrouter-api-key",
    base_url="https://openrouter.ai/api/v1"
)

md = MarkItDown(
    llm_client=client,
    llm_model="anthropic/claude-sonnet-4.5",  # recommended for scientific vision
    llm_prompt="Describe this image in detail for scientific documentation"
)

result = md.convert("presentation.pptx")
print(result.text_content)
2. Azure Document Intelligence

For enhanced PDF conversion with Microsoft Document Intelligence:

bash
# Command line
markitdown document.pdf -o output.md -d -e "<document_intelligence_endpoint>"
python
# Python API
from markitdown import MarkItDown

md = MarkItDown(docintel_endpoint="<document_intelligence_endpoint>")
result = md.convert("complex_document.pdf")
print(result.text_content)
3. Plugin System

MarkItDown supports 3rd-party plugins for extending functionality:

bash
# List installed plugins
markitdown --list-plugins

# Enable plugins
markitdown --use-plugins file.pdf -o output.md

Find plugins on GitHub with hashtag: #markitdown-plugin

Optional Dependencies

Control which file formats you support:

bash
# Install specific formats
pip install 'markitdown[pdf, docx, pptx]'

# All available options:
# [all]                  - All optional dependencies
# [pptx]                 - PowerPoint files
# [docx]                 - Word documents
# [xlsx]                 - Excel spreadsheets
# [xls]                  - Older Excel files
# [pdf]                  - PDF documents
# [outlook]              - Outlook messages
# [az-doc-intel]         - Azure Document Intelligence
# [audio-transcription]  - WAV and MP3 transcription
# [youtube-transcription] - YouTube video transcription

Common Use Cases

Show full SKILL.md (255 more words)Show less
1. Convert Scientific Papers to Markdown
python
from markitdown import MarkItDown

md = MarkItDown()

# Convert PDF paper
result = md.convert("research_paper.pdf")
with open("paper.md", "w") as f:
    f.write(result.text_content)
2. Extract Data from Excel for Analysis
python
from markitdown import MarkItDown

md = MarkItDown()
result = md.convert("data.xlsx")

# Result will be in Markdown table format
print(result.text_content)
3. Process Multiple Documents
python
from markitdown import MarkItDown
import os
from pathlib import Path

md = MarkItDown()

# Process all PDFs in a directory
pdf_dir = Path("papers/")
output_dir = Path("markdown_output/")
output_dir.mkdir(exist_ok=True)

for pdf_file in pdf_dir.glob("*.pdf"):
    result = md.convert(str(pdf_file))
    output_file = output_dir / f"{pdf_file.stem}.md"
    output_file.write_text(result.text_content)
    print(f"Converted: {pdf_file.name}")
4. Convert PowerPoint with AI Descriptions
python
from markitdown import MarkItDown
from openai import OpenAI

# Use OpenRouter for access to multiple AI models
client = OpenAI(
    api_key="your-openrouter-api-key",
    base_url="https://openrouter.ai/api/v1"
)

md = MarkItDown(
    llm_client=client,
    llm_model="anthropic/claude-sonnet-4.5",  # recommended for presentations
    llm_prompt="Describe this slide image in detail, focusing on key visual elements and data"
)

result = md.convert("presentation.pptx")
with open("presentation.md", "w") as f:
    f.write(result.text_content)
5. Batch Convert with Different Formats
python
from markitdown import MarkItDown
from pathlib import Path

md = MarkItDown()

# Files to convert
files = [
    "document.pdf",
    "spreadsheet.xlsx",
    "presentation.pptx",
    "notes.docx"
]

for file in files:
    try:
        result = md.convert(file)
        output = Path(file).stem + ".md"
        with open(output, "w") as f:
            f.write(result.text_content)
        print(f"✓ Converted {file}")
    except Exception as e:
        print(f"✗ Error converting {file}: {e}")
6. Extract YouTube Video Transcription
python
from markitdown import MarkItDown

md = MarkItDown()

# Convert YouTube video to transcript
result = md.convert("https://www.youtube.com/watch?v=VIDEO_ID")
print(result.text_content)

Docker Usage

bash
# Build image
docker build -t markitdown:latest .

# Run conversion
docker run --rm -i markitdown:latest < ~/document.pdf > output.md

Best Practices

1. Choose the Right Conversion Method
  • Simple documents: Use basic MarkItDown()
  • Complex PDFs: Use Azure Document Intelligence
  • Visual content: Enable AI image descriptions
  • Scanned documents: Ensure OCR dependencies are installed
2. Handle Errors Gracefully
python
from markitdown import MarkItDown

md = MarkItDown()

try:
    result = md.convert("document.pdf")
    print(result.text_content)
except FileNotFoundError:
    print("File not found")
except Exception as e:
    print(f"Conversion error: {e}")
3. Process Large Files Efficiently
python
from markitdown import MarkItDown

md = MarkItDown()

# For large files, use streaming
with open("large_file.pdf", "rb") as f:
    result = md.convert_stream(f, file_extension=".pdf")
    
    # Process in chunks or save directly
    with open("output.md", "w") as out:
        out.write(result.text_content)
4. Optimize for Token Efficiency

Markdown output is already token-efficient, but you can:

  • Remove excessive whitespace
  • Consolidate similar sections
  • Strip metadata if not needed
python
from markitdown import MarkItDown
import re

md = MarkItDown()
result = md.convert("document.pdf")

# Clean up extra whitespace
clean_text = re.sub(r'\n{3,}', '\n\n', result.text_content)
clean_text = clean_text.strip()

print(clean_text)

Integration with Scientific Workflows

Convert Literature for Review
python
from markitdown import MarkItDown
from pathlib import Path

md = MarkItDown()

# Convert all papers in literature folder
papers_dir = Path("literature/pdfs")
output_dir = Path("literature/markdown")
output_dir.mkdir(exist_ok=True)

for paper in papers_dir.glob("*.pdf"):
    result = md.convert(str(paper))
    
    # Save with metadata
    output_file = output_dir / f"{paper.stem}.md"
    content = f"# {paper.stem}\n\n"
    content += f"**Source**: {paper.name}\n\n"
    content += "---\n\n"
    content += result.text_content
    
    output_file.write_text(content)

# For AI-enhanced conversion with figures
from openai import OpenAI

client = OpenAI(
    api_key="your-openrouter-api-key",
    base_url="https://openrouter.ai/api/v1"
)

md_ai = MarkItDown(
    llm_client=client,
    llm_model="anthropic/claude-sonnet-4.5",
    llm_prompt="Describe scientific figures with technical precision"
)
Extract Tables for Analysis
python
from markitdown import MarkItDown
import re

md = MarkItDown()
result = md.convert("data_tables.xlsx")

# Markdown tables can be parsed or used directly
print(result.text_content)

Troubleshooting

Common Issues
  1. Missing dependencies: Install feature-specific packages

    bash
    pip install 'markitdown[pdf]'  # For PDF support
  2. Binary file errors: Ensure files are opened in binary mode

    python
    with open("file.pdf", "rb") as f:  # Note the "rb"
        result = md.convert_stream(f, file_extension=".pdf")
  3. OCR not working: Install tesseract

    bash
    # macOS
    brew install tesseract
    
    # Ubuntu
    sudo apt-get install tesseract-ocr

Performance Considerations

  • PDF files: Large PDFs may take time; consider page ranges if supported
  • Image OCR: OCR processing is CPU-intensive
  • Audio transcription: Requires additional compute resources
  • AI image descriptions: Requires API calls (costs may apply)

Next Steps

  • See references/api_reference.md for complete API documentation
  • Check references/file_formats.md for format-specific details
  • Review scripts/batch_convert.py for automation examples
  • Explore scripts/convert_with_ai.py for AI-enhanced conversions

Resources

© ImCa0, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references, assets) in .agents/skills/markitdown of ImCa0/just-laws.

  • SKILL.md
  • INSTALLATION_GUIDE.md
  • LICENSE.txt
  • OPENROUTER_INTEGRATION.md
  • QUICK_REFERENCE.md
  • README.md
  • SKILL_SUMMARY.md
  • assets/example_usage.md
  • references/api_reference.md
  • references/file_formats.md
  • scripts/batch_convert.py
  • scripts/convert_literature.py
  • scripts/convert_with_ai.py

Open the folder on GitHubat commit b947e4e

Used in 14 other repositories

We found 19 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 14 other GitHub owners. This page covers the copy in ImCa0/just-laws, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Markitdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Markitdown compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Markitdown this skillImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Markdown Converterintellectronica/agent-skills2954 repos~492Automated safety check: PassCC0-1.0
To MarkdownMathews-Tom/armory328—~2kAutomated safety check: PassMIT
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Markdown Converter

    intellectronica/agent-skills

    Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.

    295 GitHub starsUsed in 4 repos~492 tokens
    Documents & OfficeAuto-check passed
  • To Markdown

    Mathews-Tom/armory

    Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube.

    328 GitHub stars~2k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    315 GitHub stars~999 tokensUpdated today
    Documents & OfficeAuto-check passed

More from ImCa0/just-laws

  • Law Pipeline

    ImCa0/just-laws

    默认法律收录与更新流程:将 .temp/lawsmd 中的法律 Markdown 规范化为 Just Laws VuePress 站点交付件,可配对读取 Word 原件,同步 category、sidebar 与 LAWSPROGRESS.md;更新已有法律时识别当前、未来和历史版本,维护 versions.json、多版本页面及生效日轮换。

    781 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Addlaws

    ImCa0/just-laws

    已弃用的旧法律手工收录入口;默认改用 law-pipeline。

    781 GitHub stars~157 tokensUpdated today
    Auto-check: notes

Questions about Markitdown

What does Markitdown do?

Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws. Markitdown is an agent skill from ImCa0/just-laws. Convert files and office documents to Markdown.

When should I use Markitdown?

Markitdown fits situations like: tasks that involve Document parsing; tasks that involve PowerPoint presentations; tasks that involve Transcription.

How do I install Markitdown in Claude Code?

Run `npx skills add ImCa0/just-laws --skill markitdown -a claude-code`. Or copy the skill folder (.agents/skills/markitdown in ImCa0/just-laws) into .claude/skills/markitdown in your project. Claude Code loads it when a task matches its description.

How do I install Markitdown in Codex?

Run `npx skills add ImCa0/just-laws --skill markitdown -a codex`. Or copy the skill folder (.agents/skills/markitdown in ImCa0/just-laws) into .agents/skills/markitdown in your project. Codex loads it when a task matches its description.

Can I use Markitdown in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ImCa0/just-laws --skill markitdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/markitdown, .gemini/skills/markitdown, .github/skills/markitdown and .opencode/skills/markitdown in your project.

What does Markitdown need to run?

Going by SKILL.md and its folder, Markitdown needs Python for the scripts in its folder and the command-line tools its instructions call (markitdown, pip, docker, python, git and brew). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does Markitdown access the network?

SKILL.md names 4 domains. In commands or code: openrouter.ai, github.com and youtube.com; the agent is likely to contact these when it follows the instructions. As links in the text: pypi.org. This is read from the text; nothing was executed.

Is Markitdown safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Markitdown use?

Markitdown is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Markitdown use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Markitdown?

Skills that share tags, products or a category with Markitdown: Markitdown (jimmc414/Kosmos, 594 stars), Markdown Converter (intellectronica/agent-skills, 295 stars), To Markdown (Mathews-Tom/armory, 328 stars) and Markdown Converter (Team-Commonly/commonly, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Markitdown?

ImCa0 (a GitHub user) maintains it in ImCa0/just-laws, which has 781 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: ImCa0/just-laws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.