Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Read any web page, document, or YouTube video as clean Markdown using PullMD.
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install AeternaLabsHQ/pullmd pullmd --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill/pullmd/skills/pullmd .claude/skills/pullmd && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .claude/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmdType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install AeternaLabsHQ/pullmd pullmd --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skill/pullmd/skills/pullmd .agents/skills/pullmd && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .agents/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install AeternaLabsHQ/pullmd pullmd --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skill/pullmd/skills/pullmd .cursor/skills/pullmd && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .cursor/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/AeternaLabsHQ/pullmd.git --path skill/pullmd/skills/pullmd--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install AeternaLabsHQ/pullmd pullmd --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skill/pullmd/skills/pullmd .gemini/skills/pullmd && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .gemini/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install AeternaLabsHQ/pullmd pullmdInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .github/skills && cp -r skills-src/skill/pullmd/skills/pullmd .github/skills/pullmd && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .github/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AeternaLabsHQ/pullmd --skill pullmd -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install AeternaLabsHQ/pullmd pullmd --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AeternaLabsHQ/pullmd.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skill/pullmd/skills/pullmd .opencode/skills/pullmd && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pullmd" agent skill from https://github.com/AeternaLabsHQ/pullmd/tree/main/skill/pullmd/skills/pullmd into .opencode/skills/pullmd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pullmd", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pullmdRead any web page, document, or YouTube video as clean Markdown using PullMD.
Pullmd is an agent skill from AeternaLabsHQ/pullmd. Read any web page, document, or YouTube video as clean Markdown using PullMD. Use this skill whenever you need to fetch, read, extract, or summarize content from a URL — web articles, Reddit threads, PDF/Word/PowerPoint/Excel/EPUB documents, or YouTube transcripts. This includes when the user says 'read this page', 'what does this URL say', 'fetch this article', 'summarize this PDF', 'get the transcript of this video', or when you need web content as context for another task. Also use this when WebFetch fails or…
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering Video and podcast notes, PDF and PowerPoint presentations. It works with YouTube, Reddit, Microsoft Excel and Microsoft PowerPoint. The repository describes itself as: Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware… The licence is AGPL-3.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 64abf90. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
reddit.comyoutube.commistral.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pullmd loads about 2.6k tokens when it runs. Until then it costs about 174 tokens; SKILL.md has 970 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from AeternaLabsHQ/pullmd at commit 64abf90, republished under its AGPL-3.0 licence (© AeternaLabsHQ). 970 words, ~2,630 tokens.
.claude/skills/pullmd/SKILL.md (or your agent's skills folder).Read web pages, documents, and YouTube videos as clean, structured Markdown via the self-hosted PullMD service. Falls back gracefully to WebFetch if PullMD is unavailable.
PullMD routes each URL through the extraction path that fits it:
/news, /newest, /ask, /show, /jobs, /best) go through the HN API and come back as a clean nested comment tree.Accept: text/markdown get native Markdown directly.The result is much cleaner than the raw HTML that WebFetch returns, and it works on JavaScript-heavy sites and binary formats that WebFetch can't handle at all.
Use Bash to curl the PullMD API. This is preferred over WebFetch because it returns clean Markdown directly:
curl -s "__PULLMD_URL__/api?url=<URL>"The response is text/markdown — ready to use as-is.
Available parameters:
| Param | Default | Notes |
|---|---|---|
url | — | Required. |
comments | true | Include Reddit / Hacker News comments. Ignored for other URLs. |
comment_depth | 3 | Comment nesting depth (1–10), Reddit and Hacker News. |
comment_limit | none | Max top-level Reddit comments (Reddit returns ~200 without a cap). |
frontmatter | false | Prepend YAML metadata (title, source, quality, share id, …). |
format | md | text strips Markdown; json returns a structured response with metadata. |
nocache | false | Bypass the 1-hour cache and refetch from source. |
render | auto | force → always render via Playwright. skip → never render. Bypasses cache. |
extractor | auto | Force readability / trafilatura / playwright, skipping the quality pick. Bypasses cache. |
pdf | — | ocr → high-quality OCR conversion for PDFs (table-grade output; needs a server-side OCR key). Bypasses cache. |
yt_timecodes | links | YouTube transcripts: links (clickable timestamps), plain ([MM:SS]), none. |
yt_chunk | 30 | YouTube transcript block size in seconds; 0 = per original snippet. |
query | — | Set this when you need specific information from a page rather than the whole document: pass the question you are trying to answer, in natural language, and get back only the matching sections - typically 70-95% fewer tokens on long pages. No LLM involved. Empty/absent = full page, unchanged. |
max_tokens | 600 | Token budget for query (64–20000). No effect without query. Raise it when the answer likely spans several sections; leave the default for single-fact lookups. Only validated when query is set. |
lang | de | Language for the comments-section header (de or en). |
Response headers worth checking:
X-Source — reddit · hackernews · cloudflare · readability · readability-fallback · trafilatura · playwright · recipe-content · coverage-guard · markitdown · youtube · image-caption · audio-transcript · pdf-ocrX-Quality — 0.0–1.0 extraction confidence (low values mean the static extraction was thin or noisy)X-Share-Id — 8-hex permalink, openable as __PULLMD_URL__/s/<id> (absent for /api/html — local conversions are never cached or shared)X-Suggested-Filename — a ready-made filename for this conversion (e.g. YT-some-talk-dQw4w9WgXcQ.md); use it when you save the output to a file instead of inventing a name.X-Transcript-Status — YouTube only: ok / none / blocked / error. blocked and error are transient (rate limit) and not cached — retry later; none means the video has no transcript at all.X-Extracted / X-Extract-Confidence / X-Extract-Sections / X-Extract-Original-Tokens / X-Extract-Returned-Tokens — only when query is active; the last two show how much context the extraction saved.Example calls:
# Read an article
curl -s "__PULLMD_URL__/api?url=https://example.com/article"
# Read a Reddit post with comments
curl -s "__PULLMD_URL__/api?url=https://reddit.com/r/node/comments/abc/title/&comments=true"
# Convert a PDF / Office document by URL
curl -s "__PULLMD_URL__/api?url=https://example.com/report.pdf"
# Table-heavy PDF via the OCR tier (if enabled on the instance)
curl -s "__PULLMD_URL__/api?url=https://example.com/report.pdf&pdf=ocr"
# YouTube transcript with clickable timecodes (if enabled on the instance)
curl -s "__PULLMD_URL__/api?url=https://www.youtube.com/watch?v=dQw4w9WgXcQ"
# Long page, but you only need one thing: get just the relevant sections
curl -s "__PULLMD_URL__/api?url=https://example.com/long-doc&query=rate+limit+headers&max_tokens=800"
# Get fresh (uncached) content
curl -s "__PULLMD_URL__/api?url=https://example.com/news&nocache=true"
# Force the Playwright fallback for a JS-rendered page that didn't trigger
# the auto-detection (or where you want to be sure)
curl -s "__PULLMD_URL__/api?url=https://mistral.ai/pricing&render=force"
# Convert a local HTML file you already have (never cached, no share link; X-Filename keeps the name out of access logs)
curl -s -X POST --data-binary @page.html -H 'Content-Type: text/html' -H 'X-Filename: page.html' "__PULLMD_URL__/api/html"
# Upload a local document (PDF/DOCX/…, max 25 MB)
curl -s -X POST --data-binary @report.pdf -H 'Content-Type: application/pdf' -H 'X-Filename: report.pdf' "__PULLMD_URL__/api/file"If curl returns valid Markdown (starts with # or contains readable text), use that content. The X-Source response header tells you which extraction method was used. If X-Source: playwright, the page needed JavaScript rendering — that's normal for SPAs (Next.js, React, Vue dashboards, …).
If PullMD fails (network error, timeout, empty response), fall back to the built-in WebFetch tool:
WebFetch(url="<URL>", prompt="Extract the main content of this page")This still works but produces noisier output since it processes raw HTML. (For document and YouTube URLs there is no WebFetch equivalent — report the failure instead.)
Need to read a URL?
├── Is it a GitHub URL? → Use `gh` CLI instead
├── Is it a JSON API? → Use curl/fetch directly
└── Anything else (web page, PDF/Office doc, YouTube, image, audio):
├── Try: curl PullMD API
│ ├── Success (got Markdown) → Use it
│ └── Failed (error/timeout/empty) → Fallback below
└── Fallback: WebFetch tool (web pages only)nocache=true if you need the latest version. render=force|skip, extractor=, pdf=ocr, and explicit yt_* params also bypass the cache.comments=true to include the discussion below the post. Reddit and Hacker News URLs are auto-detected and use dedicated pipelines; comment_depth controls how deep the tree goes.query=<the question you are trying to answer>, phrased in natural language - it returns just the matching sections (typically 70-95% fewer tokens on long pages) and reports the saving in X-Extract-*. It falls back to the full page when nothing matches, so it is safe to try. Omit it only when you genuinely need the complete document - summarizing, translating, archiving.render=force re-extracts via headless Chromium.redd.it short links and /r/<sub>/s/<id> share links) and use a specialized extraction pipeline that handles posts, comments, galleries, and videos.frontmatter=true when you want metadata: extraction source and quality always; for Reddit posts also subreddit, author, upvotes, and publish date; for media/YouTube/OCR results duration, image size, and LLM token usage (cost tracking)./api/history endpoint shows recent conversions — useful for checking what's been fetched: curl -s "__PULLMD_URL__/api/history?limit=5".share_id. GET __PULLMD_URL__/s/<id> returns the cached markdown and re-fetches from source if older than one hour — useful as a stable URL that always returns fresh content.© AeternaLabsHQ, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skill/pullmd/skills/pullmd of AeternaLabsHQ/pullmd.
Open the folder on GitHubat commit 64abf90
Pullmd next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pullmd this skillAeternaLabsHQ/pullmd | 486 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| To MarkdownMathews-Tom/armory | 328 | — | ~2k | Automated safety check: Pass | MIT | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| Document ConverterBlackBeltTechnology/pi-agent-dashboard | 315 | — | ~999 | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
Mathews-Tom/armory
Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube.
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
BlackBeltTechnology/pi-agent-dashboard
Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…
affaan-m/ECC
Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.
Categories
Read any web page, document, or YouTube video as clean Markdown using PullMD. Pullmd is an agent skill from AeternaLabsHQ/pullmd. Read any web page, document, or YouTube video as clean Markdown using PullMD.
Pullmd fits situations like: you need to fetch; summarize content from a URL — web articles; PDF/Word/PowerPoint/Excel/EPUB documents; youTube transcripts.
Run `npx skills add AeternaLabsHQ/pullmd --skill pullmd -a claude-code`. Or copy the skill folder (skill/pullmd/skills/pullmd in AeternaLabsHQ/pullmd) into .claude/skills/pullmd in your project. Claude Code loads it when a task matches its description.
Run `npx skills add AeternaLabsHQ/pullmd --skill pullmd -a codex`. Or copy the skill folder (skill/pullmd/skills/pullmd in AeternaLabsHQ/pullmd) into .agents/skills/pullmd in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AeternaLabsHQ/pullmd --skill pullmd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pullmd, .gemini/skills/pullmd, .github/skills/pullmd and .opencode/skills/pullmd in your project.
Going by SKILL.md and its folder, Pullmd needs the command-line tools its instructions call (curl).
SKILL.md names 3 domains. In commands or code: reddit.com, youtube.com and mistral.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pullmd is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pullmd: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), To Markdown (Mathews-Tom/armory, 328 stars) and Markdown Converter (Team-Commonly/commonly, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
AeternaLabsHQ (a GitHub organization) maintains it in AeternaLabsHQ/pullmd, which has 486 GitHub stars. The repository was last updated on September 21, 2026.
Source: AeternaLabsHQ/pullmd on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.