Agent skill

Document Conversion

by Harryoung in Harryoung/efka

Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.

Apache-2.0Auto-check passedDocuments & Office

Install Document Conversion

skills CLI
$ npx skills add Harryoung/efka --skill document-conversion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Harryoung/efka document-conversion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/document-conversion .claude/skills/document-conversion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-conversion
GitHub stars
104
Token cost
~505 tokens
SKILL.md length
155 words
Files
3 (incl. scripts)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.

  • Works in 4 steps: Execute conversion command (must use… → Parse JSON output, check success field → If success: false, report error and end → …
  • Administrator onboards non-Markdown documents
  • SKILL.md covers Supported Formats, Usage, Output Format and Processing Flow, plus 2 more sections
  • Runs Python scripts from its folder; calls python

What it does

Document Conversion is an agent skill from Harryoung/efka. Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. Automatically detect PDF type (electronic/scanned), extract images to separate directory. Use this Skill when administrator onboards non-Markdown documents. Trigger condition: Onboard DOC/DOCX/PDF/PPT/PPTX format files.

Its SKILL.md is about 510 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `FORMATS.md` and `scripts/smart_convert.py`).

It sits in Documents & Office, covering PowerPoint presentations, Word documents and Document parsing. It works with Microsoft PowerPoint and Microsoft Word. The repository describes itself as: AI-powered knowledge management without vector embeddings. Built upon Claude Agent SDK, File system based, Agent driven. Maybe slower, but results are much more reliable! The licence is Apache-2.0.

When your agent uses it

  • Administrator onboards non-Markdown documents
  • Condition: Onboard DOC/DOCX/PDF/PPT/PPTX format files

Example prompts

  • “/document-conversion”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Execute conversion command (must use --original-name and --json-output)
  2. Parse JSON output, check success field
  3. If success: false, report error and end
  4. If success: true, record generated file path and image directory

What it can do on your machine

Read from SKILL.md and the folder at commit 9e6d3a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Conversion loads about 505 tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 155 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~505

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Harryoung/efka at commit 9e6d3a4, republished under its Apache-2.0 licence (© Harryoung). 155 words, ~505 tokens.

Download SKILL.mdSave it as .claude/skills/document-conversion/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
document-conversion
description
Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. Automatically detect PDF type (electronic/scanned), extract images to separate directory. Use this Skill when administrator onboards non-Markdown documents. Trigger condition: Onboard DOC/DOCX/PDF/PPT/PPTX format files.

Document Format Conversion

Convert various document formats to Markdown for knowledge base onboarding.

Supported Formats

FormatProcessing Method
DOCXPandoc conversion, preserve formatting and images
DOCLibreOffice → DOCX → Pandoc
PDF ElectronicPyMuPDF4LLM fast conversion
PDF ScannedPaddleOCR-VL online OCR
PPTXpptx2md professional conversion
PPTLibreOffice → PPTX → pptx2md

Usage

bash
python .claude/skills/document-conversion/scripts/smart_convert.py \
    <temp_path> \
    --original-name "<original_filename>" \
    --json-output

Parameters:

  • <temp_path>: Temporary file path (e.g. /tmp/kb_upload_xxx.pptx)
  • --original-name: Must pass original filename, used to generate correct image directory name
  • --json-output: Output JSON format result

Output Format

json
{
  "success": true,
  "markdown_file": "/path/to/output.md",
  "images_dir": "original_filename_images",
  "image_count": 5,
  "input_file": "/path/to/input.pptx"
}

Processing Flow

  1. Execute conversion command (must use --original-name and --json-output)
  2. Parse JSON output, check success field
  3. If success: false, report error and end
  4. If success: true, record generated file path and image directory

Important Notes

  • Image directory uses original filename naming (e.g. 培训资料_images/)
  • Not passing --original-name will cause incorrect image reference paths
  • PDF type is automatically detected, scanned version processing is slower (tens of seconds to minutes)

Format Details

Detailed processing instructions for each format, see FORMATS.md

© Harryoung, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/document-conversion of Harryoung/efka.

  • SKILL.md
  • FORMATS.md
  • scripts/smart_convert.py

Open the folder on GitHubat commit 9e6d3a4

Compare with similar skills

Document Conversion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Conversion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Conversion this skillHarryoung/efka104—~505Automated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Nutrient Document Processingaffaan-m/ECC274k4 repos~1.5kAutomated safety check: PassMIT
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence
LiteparsePrismer-AI/PrismerCloud1.6k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.

    274k GitHub starsUsed in 4 repos~1.5k tokens
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Liteparse

    Prismer-AI/PrismerCloud

    Parse local document files into LLM-ready content on the daemon itself — PDFs → Markdown / structured JSON (with bounding boxes) / page screenshots via the bundled lit CLI.

    1.6k GitHub stars~2.1k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed
  • To Markdown

    Mathews-Tom/armory

    Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube.

    327 GitHub stars~2k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from Harryoung/efka

  • Excel Parser

    Harryoung/efka

    Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

    104 GitHub stars~2.3k tokensUpdated 6 mo ago
    Auto-check passed
  • Batch Notification

    Harryoung/efka

    Send IM messages to users in batch. An agent skill from Harryoung/efka.

    104 GitHub stars~512 tokensUpdated 6 mo ago
    Auto-check passed
  • Expert Routing

    Harryoung/efka

    Domain expert routing. An agent skill from Harryoung/efka.

    104 GitHub stars~325 tokensUpdated 6 mo ago
    Auto-check passed
  • Large File Toc

    Harryoung/efka

    Generate table of contents overview for large files. An agent skill from Harryoung/efka.

    104 GitHub stars~293 tokensUpdated 6 mo ago
    Auto-check passed
  • Satisfaction Feedback

    Harryoung/efka

    Handle user satisfaction feedback. An agent skill from Harryoung/efka.

    104 GitHub stars~346 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Document Conversion

What does Document Conversion do?

Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka. Document Conversion is an agent skill from Harryoung/efka. Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format.

When should I use Document Conversion?

Document Conversion fits situations like: administrator onboards non-Markdown documents; condition: Onboard DOC/DOCX/PDF/PPT/PPTX format files.

How do I install Document Conversion in Claude Code?

Run `npx skills add Harryoung/efka --skill document-conversion -a claude-code`. Or copy the skill folder (skills/document-conversion in Harryoung/efka) into .claude/skills/document-conversion in your project. Claude Code loads it when a task matches its description.

How do I install Document Conversion in Codex?

Run `npx skills add Harryoung/efka --skill document-conversion -a codex`. Or copy the skill folder (skills/document-conversion in Harryoung/efka) into .agents/skills/document-conversion in your project. Codex loads it when a task matches its description.

Can I use Document Conversion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Harryoung/efka --skill document-conversion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-conversion, .gemini/skills/document-conversion, .github/skills/document-conversion and .opencode/skills/document-conversion in your project.

What does Document Conversion need to run?

Going by SKILL.md and its folder, Document Conversion needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Document Conversion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Document Conversion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Document Conversion use?

Document Conversion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Conversion use?

About 505 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Document Conversion?

Skills that share tags, products or a category with Document Conversion: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), Nutrient Document Processing (affaan-m/ECC, 274k stars) and Document Converter (wentorai/Research-Claw, 858 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Conversion?

Harryoung (a GitHub user) maintains it in Harryoung/efka, which has 104 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on March 16, 2026.

Source: Harryoung/efka on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.