Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install daymade/claude-code-skills doc-to-markdown --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/daymade-docs/doc-to-markdown .claude/skills/doc-to-markdown && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .claude/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdownType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install daymade/claude-code-skills doc-to-markdown --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/daymade-docs/doc-to-markdown .agents/skills/doc-to-markdown && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .agents/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install daymade/claude-code-skills doc-to-markdown --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/daymade-docs/doc-to-markdown .cursor/skills/doc-to-markdown && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .cursor/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/daymade/claude-code-skills.git --path daymade-docs/doc-to-markdown--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install daymade/claude-code-skills doc-to-markdown --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/daymade-docs/doc-to-markdown .gemini/skills/doc-to-markdown && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .gemini/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install daymade/claude-code-skills doc-to-markdownInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/daymade-docs/doc-to-markdown .github/skills/doc-to-markdown && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .github/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill doc-to-markdown -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install daymade/claude-code-skills doc-to-markdown --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/daymade-docs/doc-to-markdown .opencode/skills/doc-to-markdown && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "doc-to-markdown" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/doc-to-markdown into .opencode/skills/doc-to-markdown/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-to-markdown", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
doc-to-markdownConverts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
Doc To Markdown is an agent skill from daymade/claude-code-skills. Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing. Fixes pandoc grid tables, simple tables, image paths, CJK bold spacing, attribute noise, and code blocks; for PDFs also strips OCR garbage blocks, repeated headers/footers/watermarks, and absolute image paths from pymupdf4llm output. Benchmarked best-in-class (7.6/10) against Docling, MarkItDown, Pandoc raw, and Mammoth. Trigger on "convert document", "docx to markdown", "parse word", "doc to markdown", "解析word"…
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts, reference files and assets (for example `assets/obsidian-links/fixed.md`, `assets/obsidian-links/legacy.md` and `assets/reader-pilot-evidence-template.json`).
It sits in Documents & Office, covering Document parsing, Word documents and Markdown. It works with Microsoft Word, Pandoc, MarkItDown and Microsoft PowerPoint. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 91bed2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvpythonpipbrewFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Doc To Markdown loads about 2.5k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 808 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from daymade/claude-code-skills at commit 91bed2b, republished under its MIT licence (© daymade). 808 words, ~2,514 tokens.
.claude/skills/doc-to-markdown/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.Convert documents to high-quality markdown with intelligent multi-tool orchestration and automatic DOCX post-processing.
Architecture: Pandoc (best-in-class extraction) + 8 post-processing fixes (our value-add).
# DOCX → Markdown (one command, zero manual fixes)
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.docx -o output.md --assets-dir ./media
# PDF → Markdown
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.pdf -o output.md
# Saved HTML → Markdown (Pandoc required; no remote fetching)
uv run scripts/convert.py page.html -o page.md
# Run tests
uv run --with pytest pytest scripts/test_convert.py -v| Mode | Speed | Quality | Use Case |
|---|---|---|---|
| Quick (default) | Fast | Good | Drafts, simple documents |
| Heavy | Slower | Best | Final documents, complex layouts |
| Format | Quick Mode | Heavy Mode |
|---|---|---|
| pymupdf4llm | pymupdf4llm + markitdown | |
| DOCX | pandoc + post-processing | pandoc + markitdown |
| PPTX | markitdown | markitdown + pandoc |
| XLSX | markitdown | markitdown |
| HTML/HTM | pandoc + source-href retention check | unsupported; use the HTML quick path |
For HTML/HTM, read references/html-conversion.md
before converting. Use scripts/convert.py; the default converts the whole body
and does not trim navigation. An explicit --html-selector selects exactly one
tag, #id or .class. --html-heading-offset shifts parsed headings for assembly
and rejects overflow beyond H6. Relative assets are not downloaded or copied.
Verify source href occurrences against the emitted Markdown AST, then rerun
scripts/html_to_markdown.py after cleanup or merging. For manuals, map the book,
chapters, lessons and internal headings before assembly; preserve code fences and
reconcile rewritten anchors. Conversion success does not certify that figures
are readable inside the recipient's actual Markdown reader.
When converting DOCX via pandoc, 8 cleanups are applied automatically:
| Problem | Fix | Test coverage |
|---|---|---|
Grid tables (+:---+) | Single-column → blockquote, multi-column → pipe table | TestPostprocessPipeline |
Simple tables ( ---- ----) | Multi-column images → pipe table with captions | TestSimpleTable |
Image path nesting (media/media/) | Flatten to media/, absolute → relative | test_stats_tracking |
Pandoc attributes ({width="..."}) | Removed | test_pandoc_attributes_removed |
CJK bold spacing (**粗体**中文) | Add space around ** for CJK bold spans | TestCjkBoldSpacing (15 cases) |
| Indented dashed code blocks | → fenced ``` with language detection | test_code_block_with_language |
Escaped brackets (\[...\]) | → [...] | test_escaped_brackets_fixed |
Double-bracket links ([[text]](url)) | → [text](url) | test_double_bracket_links_fixed |
When converting PDF via pymupdf4llm, 3 cleanups are applied automatically (skip with --no-postprocess):
| Problem | Fix | Test coverage |
|---|---|---|
Tesseract OCR garbage on image regions (<!-- Start of picture text -->...) | Block removed; images themselves kept | TestStripOcrPictureText |
| Repeated header/footer/watermark lines (same normalized line on ≥60% of pages, incl. diagonal watermarks) | Detected via pymupdf cross-page scan, removed from markdown; bold-wrapped and merged-with-page-number variants also caught | TestRepeatingLines |
Absolute image paths () | Rewritten relative to the output markdown file (portable output) | TestImagePathsRelative |
Heavy mode additionally prints a loud ⚠️ HEAVY MODE DEGRADED warning on stderr when one engine fails and the merge would otherwise silently degrade to single-engine output.
Known limits (learned from a 62-page Chinese research-report conversion, 2026-08-30):
dn, uFE, ...); the repeating-line stripper removes full lines only, not intra-cell shards. Watermark-heavy PDFs need cell-level rebuild (collect non-watermark spans per cell bbox).DOCX uses run-level styling (no spaces between bold/normal runs in CJK text). Markdown renderers need whitespace around ** to recognize bold boundaries.
Rule: if a **content** span contains any CJK character, ensure both sides have a space — unless already spaced or at line boundary. This handles CJK punctuation, emoji adjacency, and mixed content.
Before: 打开**飞书**,就可以 → some renderers fail to bold
After: 打开 **飞书** ,就可以 → universally renders correctlyHeavy Mode runs multiple tools in parallel and selects the best segments:
| Segment Type | Selection Criteria |
|---|---|
| Tables | More rows/columns, proper header separator |
| Images | Alt text present, local paths preferred |
| Headings | Proper hierarchy, appropriate length |
| Lists | More items, nested structure preserved |
| Paragraphs | Content completeness |
# Extract images with metadata
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf -o ./extracted-images
# Generate markdown references file
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf --markdown refs.mdOutput:
extracted-images/img_page1_1.png, extracted-images/img_page2_1.jpgextracted-images/images_metadata.json (page, position, dimensions)# Validate conversion quality
uv run --with pymupdf scripts/validate_output.py document.pdf output.md
# Generate HTML report
uv run --with pymupdf scripts/validate_output.py document.pdf output.md --report report.html| Metric | Pass | Warn | Fail |
|---|---|---|---|
| Text Retention | >95% | 85-95% | <85% |
| Table Retention | 100% | 90-99% | <90% |
| Image Retention | 100% | 80-99% | <80% |
# Merge multiple markdown files
python scripts/merge_outputs.py output1.md output2.md -o merged.md
# Show segment attribution
python scripts/merge_outputs.py output1.md output2.md -o merged.md --verbose# Windows to WSL conversion
python scripts/convert_path.py "C:\Users\<windows-user>\Documents\file.pdf"
# Output: /mnt/c/Users/<windows-user>/Documents/file.pdf"No conversion tools available"
# Install all tools
pip install pymupdf4llm
uv tool install "markitdown[pdf]"
brew install pandocFontBBox warnings during PDF conversion
Images missing from output
scripts/extract_pdf_images.pyTables broken in output
scripts/validate_output.py| Script | Purpose |
|---|---|
convert.py | Main orchestrator with Quick/Heavy mode + DOCX post-processing |
html_to_markdown.py | Pandoc HTML adapter and saved-output source-href retention verifier |
test_convert.py | 31 tests covering all post-processing functions |
merge_outputs.py | Merge multiple markdown outputs |
validate_output.py | Quality validation with HTML report |
extract_pdf_images.py | PDF image extraction with metadata |
convert_path.py | Windows to WSL path converter |
references/benchmark-2026-03-22.md - 5-tool benchmark (Docling/MarkItDown/Pandoc/Mammoth/ours)references/heavy-mode-guide.md - Detailed Heavy Mode documentationreferences/tool-comparison.md - Tool capabilities comparisonreferences/conversion-examples.md - Batch operation examplesreferences/html-conversion.md - Saved HTML scope, link retention, assets and manual heading assemblyAfter converting documents to markdown, suggest cleanup:
Conversion complete: [N] files converted to markdown.
Options:
A) Clean up docs — run /daymade-docs:docs-cleaner to consolidate redundant content (Recommended if multiple files)
B) Check facts — run /fact-checker to verify claims in the converted content
C) No thanks — the markdown conversion is sufficient© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files (scripts, references, assets) in daymade-docs/doc-to-markdown of daymade/claude-code-skills.
Open the folder on GitHubat commit 91bed2b
Doc To Markdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Doc To Markdown this skilldaymade/claude-code-skills | 1.4k | — | ~2.5k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html | 908 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| Markdown Converterintellectronica/agent-skills | 295 | 4 repos | ~492 | Automated safety check: Pass | CC0-1.0 |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
alchaincyf/huashu-md-html
Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
intellectronica/agent-skills
Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.
wentorai/Research-Claw
Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.
daymade/claude-code-skills
This skill should be used when comparing two videos to analyze compression results or quality differences.
daymade/claude-code-skills
Generates professional animated CLI demos as GIFs using VHS terminal recordings.
daymade/claude-code-skills
Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.
daymade/claude-code-skills
Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.
daymade/claude-code-skills
Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…
daymade/claude-code-skills
Fetches real, citable Bilibili (B站) video data — stats, metadata, tags, and full danmaku text — via login-free API calls, never hand-typed or estimated.
Categories
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing. Doc To Markdown is an agent skill from daymade/claude-code-skills. Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
Doc To Markdown fits situations like: convert document; docx to markdown; doc to markdown; HTML to Markdown.
Run `npx skills add daymade/claude-code-skills --skill doc-to-markdown -a claude-code`. Or copy the skill folder (daymade-docs/doc-to-markdown in daymade/claude-code-skills) into .claude/skills/doc-to-markdown in your project. Claude Code loads it when a task matches its description.
Run `npx skills add daymade/claude-code-skills --skill doc-to-markdown -a codex`. Or copy the skill folder (daymade-docs/doc-to-markdown in daymade/claude-code-skills) into .agents/skills/doc-to-markdown in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill doc-to-markdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-to-markdown, .gemini/skills/doc-to-markdown, .github/skills/doc-to-markdown and .opencode/skills/doc-to-markdown in your project.
Going by SKILL.md and its folder, Doc To Markdown needs Python for the scripts in its folder and the command-line tools its instructions call (uv, python, pip and brew). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Doc To Markdown is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Doc To Markdown: Markitdown (ImCa0/just-laws, 781 stars), Markdown Converter (Team-Commonly/commonly, 1.4k stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 908 stars) and Markitdown (jimmc414/Kosmos, 594 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,444 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 8, 2026.
Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.