Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure.
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install daymade/claude-code-skills pdf-to-html --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/daymade-docs/pdf-to-html .claude/skills/pdf-to-html && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .claude/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-htmlType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install daymade/claude-code-skills pdf-to-html --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/daymade-docs/pdf-to-html .agents/skills/pdf-to-html && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .agents/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install daymade/claude-code-skills pdf-to-html --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/daymade-docs/pdf-to-html .cursor/skills/pdf-to-html && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .cursor/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/daymade/claude-code-skills.git --path daymade-docs/pdf-to-html--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install daymade/claude-code-skills pdf-to-html --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/daymade-docs/pdf-to-html .gemini/skills/pdf-to-html && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .gemini/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install daymade/claude-code-skills pdf-to-htmlInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/daymade-docs/pdf-to-html .github/skills/pdf-to-html && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .github/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill pdf-to-html -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install daymade/claude-code-skills pdf-to-html --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/daymade-docs/pdf-to-html .opencode/skills/pdf-to-html && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-to-html" agent skill from https://github.com/daymade/claude-code-skills/tree/main/daymade-docs/pdf-to-html into .opencode/skills/pdf-to-html/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-to-html", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-to-htmlConverts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure.
PDF To HTML is an agent skill from daymade/claude-code-skills. Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure. Uses structured extraction (PyMuPDF), font-size-driven layout, compressed base64-inlined images (a single portable file), and mandatory headless-Chrome visual verification. Use whenever someone wants to READ a PDF as a web page or clean document, turn a PDF into HTML, or translate a PDF into another language while keeping…
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/failure_cases.md`, `references/translation_workflow.md` and `scripts/build_html.py`).
It sits in Documents & Office, covering PDF, Translation and Document parsing. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 91bed2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF To HTML loads about 1.8k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 219 tokens; SKILL.md has 719 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from daymade/claude-code-skills at commit 91bed2b, republished under its MIT licence (© daymade). 719 words, ~1,828 tokens.
.claude/skills/pdf-to-html/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Turn a PDF into a single, self-contained, readable HTML file — images, tables, charts and reading order preserved — and optionally translate it, keeping every figure in place.
The pipeline is extract → look → (translate) → build → verify. The middle "look" and final "verify" steps are where faithfulness actually comes from: a PDF is a layout, not just a text stream, so you read the rendered pages before building and the rendered HTML before delivering.
This skill runs inline (no context: fork): translation orchestrates a
Dynamic Workflow, and a subagent cannot spawn one.
ocrmypdf), then use this.uv (runs Python with inline deps), Google Chrome or Chromium (visual
verification). Python packages come via uv run --with: PyMuPDF, Pillow, numpy.
Nothing to pre-install beyond Chrome and uv.
Copy this checklist and tick as you go:
- [ ] 1. Extract structure + render pages (extract_pdf.py)
- [ ] 2. Read pages/*.png — SEE the layout, find content vs decorative images
- [ ] 3. (only if translating) run the translation workflow
- [ ] 4. Build the single-file HTML (build_html.py)
- [ ] 5. Verify visually (verify_render.py → Read every segment)
- [ ] 6. Deliver the .htmluv run --with pymupdf python scripts/extract_pdf.py input.pdfWrites input-build/ with structure.json (text blocks with font sizes + image
blocks flagged decorative), images/, and pages/ (one PNG per page).
Read input-build/pages/*.png. This is not optional: you need to see the real
layout, confirm which images are content vs decoration, and spot tables/charts.
For a long PDF, read every page; for a short one it's quick. This is also where
you understand the document well enough to translate it well.
Only if the user asked for another language. Read
references/translation_workflow.md and
follow it: a Dynamic Workflow translates pages in parallel, captions data charts,
and reconciles terminology. It produces two overlay files (units.json,
caps.json) that step 4 consumes. Do not hand-translate inline for anything
longer than a page — the workflow keeps terminology consistent and is far faster.
# original-language HTML
uv run --with Pillow python scripts/build_html.py input-build/structure.json --out output.html
# translated HTML (overlays from step 3)
uv run --with Pillow python scripts/build_html.py input-build/structure.json --out output.html \
--translation input-build/units.json --captions input-build/caps.json --lang zh-CNbuild_html.py is data-driven: it infers heading levels from font size (most
common size = body; larger steps up to h3/h2/h1), drops decorative images, and
inlines content images as compressed base64 → one portable file. It is not
hand-tuned to any document. If a particular PDF has an unusual structure (e.g.
multi-column, sidebars, a figure the size heuristic misreads), read the script and
adjust — it's short and meant to be edited per document.
uv run --with Pillow --with numpy python scripts/verify_render.py output.htmlThen Read every seg-*.png and check: fonts render (no tofu boxes), no
clipped tables/figures, headings/lists look right, all expected images present.
Text being correct does not mean the render is correct (failure_cases #7). Fix and
re-verify until it's clean.
A quick structural cross-check is fine too, but count occurrences correctly:
grep -o '<figure>' output.html | wc -l — not grep -c (failure_cases #1).
Hand over the single .html. It's self-contained (images inlined), so it opens
with a double-click and nothing can go missing.
| Script | Run with | Purpose |
|---|---|---|
scripts/extract_pdf.py | uv run --with pymupdf | PDF → structure.json + images/ + page renders |
scripts/build_html.py | uv run --with Pillow | structure.json (+ optional translation/captions) → single-file HTML |
scripts/verify_render.py | uv run --with Pillow --with numpy | headless-Chrome render → readable PNG segments |
The deliverable looks authoritative, so wrong content is worse than ugly content. The non-negotiable rules — and the specific ways this has gone wrong before — are in references/failure_cases.md. The one that bites hardest: never give a real person an inferred translated name, and copy every number/proper-noun verbatim (failure_cases #6). Read that file before any translation run; skim it before any run.
After producing the HTML, suggest the natural follow-up:
Conversion complete: output.html (single self-contained file).
Options:
A) Make a PDF of it — run /daymade-docs:pdf-creator if you want a print/share copy (Recommended if they need to send it)
B) Extract the text as Markdown instead — run /daymade-docs:doc-to-markdown (if they wanted editable text, not a reading page)
C) No thanks — the HTML is what I wanted© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in daymade-docs/pdf-to-html of daymade/claude-code-skills.
Open the folder on GitHubat commit 91bed2b
PDF To HTML next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF To HTML this skilldaymade/claude-code-skills | 1.4k | — | ~1.8k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html | 908 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Lt2mdlibnyx/LT2MD | 109 | — | ~4.4k | Automated safety check: Pass | AGPL-3.0 | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Mineru PDF Parserstaruhub/ClaudeSkills | 727 | — | ~671 | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
alchaincyf/huashu-md-html
Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.
libnyx/LT2MD
Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
staruhub/ClaudeSkills
用 MinerU 将复杂PDF文档转换为LLM友好的Markdown/JSON格式。适用于:(1) PDF转Markdown/JSON,(2) 提取PDF中的文本、表格、公式、图像,(3) 解析学术论文、技术文档、商业报告,(4) 为RAG应用准备文档数据,(5) 批量处理PDF。触发关键词:"PDF解析"、"PDF转Markdown"、"提取PDF表格/公式"、"MinerU"、"parse…
mitsuhiko/agent-stuff
Fetch a URL or convert a local file (PDF/DOCX/HTML/etc.) into Markdown using uvx markitdown, optionally it can summarize
daymade/claude-code-skills
This skill should be used when comparing two videos to analyze compression results or quality differences.
daymade/claude-code-skills
Generates professional animated CLI demos as GIFs using VHS terminal recordings.
daymade/claude-code-skills
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
daymade/claude-code-skills
Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.
daymade/claude-code-skills
Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.
daymade/claude-code-skills
Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…
Categories
Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure. PDF To HTML is an agent skill from daymade/claude-code-skills. Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure.
PDF To HTML fits situations like: someone wants to READ a PDF as a web page; turn a PDF into HTML; translate a PDF into another language while keeping its images/tables/charts intact — e.g.
Run `npx skills add daymade/claude-code-skills --skill pdf-to-html -a claude-code`. Or copy the skill folder (daymade-docs/pdf-to-html in daymade/claude-code-skills) into .claude/skills/pdf-to-html in your project. Claude Code loads it when a task matches its description.
Run `npx skills add daymade/claude-code-skills --skill pdf-to-html -a codex`. Or copy the skill folder (daymade-docs/pdf-to-html in daymade/claude-code-skills) into .agents/skills/pdf-to-html in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill pdf-to-html -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-to-html, .gemini/skills/pdf-to-html, .github/skills/pdf-to-html and .opencode/skills/pdf-to-html in your project.
Going by SKILL.md and its folder, PDF To HTML needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PDF To HTML is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with PDF To HTML: Markitdown (ImCa0/just-laws, 781 stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 908 stars), Lt2md (libnyx/LT2MD, 109 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,444 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 8, 2026.
Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.