Agent skill

Markdrop

by shoryasethia in shoryasethia/markdrop

Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

GPL-3.0Auto-check: notesDocuments & Office

Install Markdrop

skills CLI
$ npx skills add shoryasethia/markdrop --skill markdrop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shoryasethia/markdrop markdrop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
markdrop
GitHub stars
211
Token cost
~1.4k tokens
SKILL.md length
367 words
Files
52
Skills in repo
1
Repo updated
First seen
Licence
GPL-3.0

At a glance

Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

  • Works in 5 steps: Capabilities → API Keys Setup → Python API Integration → …
  • Tasks that involve PDF
  • SKILL.md covers 1. Capabilities, 2. API Keys Setup, 3. Python API Integration and 4. CLI Execution Best Practices, plus 1 more section
  • Reaches url.to; needs GEMINI_API_KEY and OPENAI_API_KEY

What it does

Markdrop is an agent skill from shoryasethia/markdrop. Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 54 other files (for example `CHANGELOG.md`, `CODE_OF_CONDUCT.md` and `CONTRIBUTING.md`).

It sits in Documents & Office, covering PDF, Document parsing and Markdown. It works with Python and MarkItDown. The repository describes itself as: A Python package for converting PDFs to markdown while extracting images and tables, generate descriptive text descriptions for extracted tables/images using several LLM clients… The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve PDF
  • Tasks that involve Document parsing
  • Tasks that involve Markdown

Example prompts

  • “/markdrop”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Capabilities
  2. API Keys Setup
  3. Python API Integration
  4. CLI Execution Best Practices
  5. Typical Model Fallbacks & Suggestions

What it can do on your machine

Read from SKILL.md and the folder at commit 2a1c475. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • url.to

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY
    • GROQ_API_KEY
    • OPENROUTER_API_KEY
    • LITELLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Markdrop loads about 1.4k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 367 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:21
    API keys must be available in the root `.env` file or environment variables.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shoryasethia/markdrop at commit 2a1c475, republished under its GPL-3.0 licence (© shoryasethia). 367 words, ~1,375 tokens.

Download SKILL.mdSave it as .claude/skills/markdrop/SKILL.md (or your agent's skills folder). This skill also uses 51 other files; get the full folder from GitHub.
name
markdrop
description
Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

Markdrop Skill

Welcome to the markdrop skill. markdrop is a powerful Python package and CLI tool used to convert PDF documents into structured Markdown and interactive HTML, while natively leveraging AI vision models to interpret and describe extracted images and tables.

If you are an AI agent or a user aiming to process PDFs and augment them with text or image descriptions, this document serves as your complete guide on utilizing markdrop efficiently and accurately.

1. Capabilities

  • PDF to Markdown/HTML: Retains formatting, extracts images, and detects tables via Microsoft Table Transformer and Docling. Supports processing both local file paths and direct PDF URLs.
  • AI Vision Descriptions: Query GEMINI, OPENAI, ANTHROPIC, GROQ, OPENROUTER, or LITELLM to generate rich descriptions of images and tables.
  • Batch Processing: Describe entire directories of images in single commands using multiple LLM backends simultaneously.
  • Extensible Configuration: Precise override control over which structural text-models vs vision-models are used, as well as prompts, resolution scales, and output features.

2. API Keys Setup

Before using AI features, API keys must be available in the root .env file or environment variables.

If deploying programmatically, you can run the built-in CLI command, or inject them into os.environ:

bash
markdrop setup gemini     # -> GEMINI_API_KEY
markdrop setup openai     # -> OPENAI_API_KEY
markdrop setup anthropic  # -> ANTHROPIC_API_KEY
markdrop setup groq       # -> GROQ_API_KEY
markdrop setup openrouter # -> OPENROUTER_API_KEY
markdrop setup litellm    # -> LITELLM_API_KEY

3. Python API Integration

The Python API is the recommended way to embed markdrop into applications.

3.1 PDF Conversion to Interactive HTML

Use markdrop function combined with add_downloadable_tables:

python
from markdrop import markdrop, MarkDropConfig, add_downloadable_tables
from pathlib import Path
import logging

# Configuration block
config = MarkDropConfig(
    image_resolution_scale=2.0,
    download_button_color="#444444",
    log_level=logging.INFO,
    log_dir="logs",
    excel_dir="markdrop-excel-tables",
)

# 1. Convert PDF to base HTML (and markdown locally). URL supported here too: "https://url.to/pdf"
html_path = markdrop("path/to/document.pdf", "output_directory", config)

# 2. Enrich HTML to allow downloading tables as Excel sheets
enhanced_html_path = add_downloadable_tables(html_path, config)
3.2 Injecting AI Descriptions into Markdown

If you have a Markdown file containing image/table links, process_markdown automatically routes vision requests to the chosen provider and inserts contextual descriptions.

python
from markdrop import process_markdown, ProcessorConfig, AIProvider

config = ProcessorConfig(
    input_path="output_directory/document.md",
    output_dir="output_directory",
    ai_provider=AIProvider.GEMINI,  # Available: GEMINI, OPENAI, ANTHROPIC, GROQ, OPENROUTER, LITELLM
    # Target configurations
    remove_images=False,
    remove_tables=False,
    table_descriptions=True,
    image_descriptions=True,
    # Provider-Specific overrides (Optional)
    # Allows granular decoupling of vision parsing vs table text-parsing
    model_name_override="gemini-2.0-flash",  # Primary vision analysis model
    text_model_name_override="gemini-2.0-flash",  # Lean text-only model for generic parsing
)

# Executes AI processing and saves the enriched document
output_path = process_markdown(config)
Show full SKILL.md (158 more words)Show less
3.3 Batch Image Description

For standalone image directories or files:

python
from markdrop import generate_descriptions

generate_descriptions(
    input_path="images_folder/",
    output_dir="descriptions_output/",
    prompt="Analyze this image and describe all textual and structural elements.",
    llm_client=["gemini", "openai"],
)

4. CLI Execution Best Practices

As an agent, you can also trigger markdrop workflows via Bash.

  1. Convert PDF to MD/HTML (including tables):

    bash
    markdrop convert <input_path_or_url> --output_dir <dir> --add_tables
  2. Run AI Provider over the Markdown Output with exact models:

    bash
    markdrop describe <markdown_file> \
        --ai_provider anthropic \
        --model claude-opus-4-6 \
        --text-model claude-sonnet-4-5 \
        --remove_images
  3. Only Analyze / Extract Images:

    bash
    # Also accepts URLs directly
    markdrop analyze https://domain.com/report.pdf --output_dir pdf_analysis --save_images
  4. Batch Image Description:

    bash
    markdrop generate images/ --output_dir descriptions/ \
        --prompt "Describe in detail." \
        --llm_client gemini openai

5. Typical Model Fallbacks & Suggestions

  • Default / Cost-Effective: gemini (Gemini 2.0 Flash) is frequently the fastest and cheapest for large scale document evaluation.
  • High Complexity / Intricate Tables: anthropic with the latest Claude models (claude-opus-4-6 or claude-sonnet-4-5) excel in reasoning and formatting.
  • Maximum Speed: groq using LLaMA models.

Whenever instantiating ProcessorConfig, be exact about paths—use absolute paths if the current working directory is dynamically changing.

© shoryasethia, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 51 other files in the repository root of shoryasethia/markdrop.

  • SKILL.md
  • .agent/skills
  • .gitignore
  • CHANGELOG.md
  • CODE_OF_CONDUCT.md
  • CONTRIBUTING.md
  • LICENSE
  • OLD-DOCUMENTATION.md
  • README.md
  • dist/markdrop-4.0.1-py3-none-any.whl
  • dist/markdrop-4.0.1.tar.gz
  • docs/api_reference.md
  • docs/architecture.md
  • docs/benchmarking.md
  • docs/cli_reference.md
  • docs/cpu-guide.md
  • docs/getting_started.md
  • docs/providers.md
  • … and 34 more

Open the folder on GitHubat commit 2a1c475

Compare with similar skills

Markdrop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Markdrop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Markdrop this skillshoryasethia/markdrop211—~1.4kAutomated safety check: NotesGPL-3.0
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
MineruNebutra/MinerU-Skill122—~1.4kAutomated safety check: PassMIT
Markitdownaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
Document ConverterBlackBeltTechnology/pi-agent-dashboard315—~999Automated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT

Similar skills

  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed
  • Markitdown

    aipoch/medical-research-skills

    Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    315 GitHub stars~999 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    908 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

Questions about Markdrop

What does Markdrop do?

Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions. Markdrop is an agent skill from shoryasethia/markdrop. Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions.

When should I use Markdrop?

Markdrop fits situations like: tasks that involve PDF; tasks that involve Document parsing; tasks that involve Markdown.

How do I install Markdrop in Claude Code?

Run `npx skills add shoryasethia/markdrop --skill markdrop -a claude-code`. Or copy the skill folder (the shoryasethia/markdrop repository) into .claude/skills/markdrop in your project. Claude Code loads it when a task matches its description.

How do I install Markdrop in Codex?

Run `npx skills add shoryasethia/markdrop --skill markdrop -a codex`. Or copy the skill folder (the shoryasethia/markdrop repository) into .agents/skills/markdrop in your project. Codex loads it when a task matches its description.

Can I use Markdrop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shoryasethia/markdrop --skill markdrop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/markdrop, .gemini/skills/markdrop, .github/skills/markdrop and .opencode/skills/markdrop in your project.

What does Markdrop need to run?

Going by SKILL.md and its folder, Markdrop needs credentials named GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY and GROQ_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in OPENAI_API_KEY.

Does Markdrop access the network?

SKILL.md names 1 domain. In commands or code: url.to; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Markdrop safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Markdrop use?

Markdrop is published under the GPL-3.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Markdrop use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Markdrop?

Skills that share tags, products or a category with Markdrop: Markdown Converter (Team-Commonly/commonly, 1.4k stars), Mineru (Nebutra/MinerU-Skill, 122 stars), Markitdown (aipoch/medical-research-skills, 2k stars) and Document Converter (BlackBeltTechnology/pi-agent-dashboard, 315 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Markdrop?

shoryasethia (a GitHub user) maintains it in shoryasethia/markdrop, which has 211 GitHub stars. The repository was last updated on August 9, 2026.

Source: shoryasethia/markdrop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.