Markdown Converter
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
$ npx skills add shoryasethia/markdrop --skill markdrop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install shoryasethia/markdrop markdrop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .claude/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shoryasethia/markdrop --skill markdrop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install shoryasethia/markdrop markdrop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .agents/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shoryasethia/markdrop --skill markdrop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install shoryasethia/markdrop markdrop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .cursor/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shoryasethia/markdrop --skill markdrop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install shoryasethia/markdrop markdrop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .gemini/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install shoryasethia/markdrop markdropInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add shoryasethia/markdrop --skill markdrop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .github/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shoryasethia/markdrop --skill markdrop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install shoryasethia/markdrop markdrop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "markdrop" agent skill from https://github.com/shoryasethia/markdrop/tree/main into .opencode/skills/markdrop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "markdrop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
markdropProfessional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
Markdrop is an agent skill from shoryasethia/markdrop. Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 54 other files (for example `CHANGELOG.md`, `CODE_OF_CONDUCT.md` and `CONTRIBUTING.md`).
It sits in Documents & Office, covering PDF, Document parsing and Markdown. It works with Python and MarkItDown. The repository describes itself as: A Python package for converting PDFs to markdown while extracting images and tables, generate descriptive text descriptions for extracted tables/images using several LLM clients… The licence is GPL-3.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 2a1c475. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and python).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
url.toFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYOPENAI_API_KEYANTHROPIC_API_KEYGROQ_API_KEYOPENROUTER_API_KEYLITELLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Markdrop loads about 1.4k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 367 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
API keys must be available in the root `.env` file or environment variables.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from shoryasethia/markdrop at commit 2a1c475, republished under its GPL-3.0 licence (© shoryasethia). 367 words, ~1,375 tokens.
.claude/skills/markdrop/SKILL.md (or your agent's skills folder). This skill also uses 51 other files; get the full folder from GitHub.Welcome to the markdrop skill. markdrop is a powerful Python package and CLI tool used to convert PDF documents into structured Markdown and interactive HTML, while natively leveraging AI vision models to interpret and describe extracted images and tables.
If you are an AI agent or a user aiming to process PDFs and augment them with text or image descriptions, this document serves as your complete guide on utilizing markdrop efficiently and accurately.
Before using AI features, API keys must be available in the root .env file or environment variables.
If deploying programmatically, you can run the built-in CLI command, or inject them into os.environ:
markdrop setup gemini # -> GEMINI_API_KEY
markdrop setup openai # -> OPENAI_API_KEY
markdrop setup anthropic # -> ANTHROPIC_API_KEY
markdrop setup groq # -> GROQ_API_KEY
markdrop setup openrouter # -> OPENROUTER_API_KEY
markdrop setup litellm # -> LITELLM_API_KEYThe Python API is the recommended way to embed markdrop into applications.
Use markdrop function combined with add_downloadable_tables:
from markdrop import markdrop, MarkDropConfig, add_downloadable_tables
from pathlib import Path
import logging
# Configuration block
config = MarkDropConfig(
image_resolution_scale=2.0,
download_button_color="#444444",
log_level=logging.INFO,
log_dir="logs",
excel_dir="markdrop-excel-tables",
)
# 1. Convert PDF to base HTML (and markdown locally). URL supported here too: "https://url.to/pdf"
html_path = markdrop("path/to/document.pdf", "output_directory", config)
# 2. Enrich HTML to allow downloading tables as Excel sheets
enhanced_html_path = add_downloadable_tables(html_path, config)If you have a Markdown file containing image/table links, process_markdown automatically routes vision requests to the chosen provider and inserts contextual descriptions.
from markdrop import process_markdown, ProcessorConfig, AIProvider
config = ProcessorConfig(
input_path="output_directory/document.md",
output_dir="output_directory",
ai_provider=AIProvider.GEMINI, # Available: GEMINI, OPENAI, ANTHROPIC, GROQ, OPENROUTER, LITELLM
# Target configurations
remove_images=False,
remove_tables=False,
table_descriptions=True,
image_descriptions=True,
# Provider-Specific overrides (Optional)
# Allows granular decoupling of vision parsing vs table text-parsing
model_name_override="gemini-2.0-flash", # Primary vision analysis model
text_model_name_override="gemini-2.0-flash", # Lean text-only model for generic parsing
)
# Executes AI processing and saves the enriched document
output_path = process_markdown(config)For standalone image directories or files:
from markdrop import generate_descriptions
generate_descriptions(
input_path="images_folder/",
output_dir="descriptions_output/",
prompt="Analyze this image and describe all textual and structural elements.",
llm_client=["gemini", "openai"],
)As an agent, you can also trigger markdrop workflows via Bash.
Convert PDF to MD/HTML (including tables):
markdrop convert <input_path_or_url> --output_dir <dir> --add_tablesRun AI Provider over the Markdown Output with exact models:
markdrop describe <markdown_file> \
--ai_provider anthropic \
--model claude-opus-4-6 \
--text-model claude-sonnet-4-5 \
--remove_imagesOnly Analyze / Extract Images:
# Also accepts URLs directly
markdrop analyze https://domain.com/report.pdf --output_dir pdf_analysis --save_imagesBatch Image Description:
markdrop generate images/ --output_dir descriptions/ \
--prompt "Describe in detail." \
--llm_client gemini openaigemini (Gemini 2.0 Flash) is frequently the fastest and cheapest for large scale document evaluation.anthropic with the latest Claude models (claude-opus-4-6 or claude-sonnet-4-5) excel in reasoning and formatting.groq using LLaMA models.Whenever instantiating ProcessorConfig, be exact about paths—use absolute paths if the current working directory is dynamically changing.
© shoryasethia, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 51 other files in the repository root of shoryasethia/markdrop.
Open the folder on GitHubat commit 2a1c475
Markdrop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Markdrop this skillshoryasethia/markdrop | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | |
| Markdown ConverterTeam-Commonly/commonly | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | |
| MineruNebutra/MinerU-Skill | 122 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Markitdownaipoch/medical-research-skills | 2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Document ConverterBlackBeltTechnology/pi-agent-dashboard | 315 | — | ~999 | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT |
Team-Commonly/commonly
Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.
aipoch/medical-research-skills
Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).
BlackBeltTechnology/pi-agent-dashboard
Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
alchaincyf/huashu-md-html
Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.
Works with
Categories
Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions. Markdrop is an agent skill from shoryasethia/markdrop. Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.
Markdrop fits situations like: tasks that involve PDF; tasks that involve Document parsing; tasks that involve Markdown.
Run `npx skills add shoryasethia/markdrop --skill markdrop -a claude-code`. Or copy the skill folder (the shoryasethia/markdrop repository) into .claude/skills/markdrop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add shoryasethia/markdrop --skill markdrop -a codex`. Or copy the skill folder (the shoryasethia/markdrop repository) into .agents/skills/markdrop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shoryasethia/markdrop --skill markdrop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/markdrop, .gemini/skills/markdrop, .github/skills/markdrop and .opencode/skills/markdrop in your project.
Going by SKILL.md and its folder, Markdrop needs credentials named GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY and GROQ_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in OPENAI_API_KEY.
SKILL.md names 1 domain. In commands or code: url.to; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Markdrop is published under the GPL-3.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Markdrop: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Mineru (Nebutra/MinerU-Skill, 122 stars), Markitdown (aipoch/medical-research-skills, 2k stars) and Document Converter (BlackBeltTechnology/pi-agent-dashboard, 315 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
shoryasethia (a GitHub user) maintains it in shoryasethia/markdrop, which has 211 GitHub stars. The repository was last updated on August 9, 2026.
Source: shoryasethia/markdrop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.