Batch convert documents between multiple formats using a unified pipeline

MITAuto-check passedDocuments & Office

Install Batch Convert

skills CLI
$ npx skills add claude-office-skills/skills --skill batch-convert -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install claude-office-skills/skills batch-convert --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/claude-office-skills/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/batch-convert .claude/skills/batch-convert && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-convert
GitHub stars
505
Token cost
~3.8k tokens
SKILL.md length
237 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Batch convert documents between multiple formats using a unified pipeline

  • Works in 4 steps: Specify the source folder or files → Choose target format(s) → Optionally configure conversion options → …
  • Tasks that involve Document parsing
  • SKILL.md covers Overview, How to Use, Domain Knowledge and Best Practices, plus 5 more sections
  • Calls brew, apt and pip

What it does

Batch Convert is an agent skill from claude-office-skills/skills. Batch convert documents between multiple formats using a unified pipeline

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Document parsing. The repository describes itself as: A curated collection of practical Claude Skills for real-world office tasks. The licence is MIT.

When your agent uses it

  • Tasks that involve Document parsing

Example prompts

  • “/batch-convert”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Specify the source folder or files
  2. Choose target format(s)
  3. Optionally configure conversion options
  4. I'll process all files with progress tracking

What it can do on your machine

Read from SKILL.md and the folder at commit 9c4c7d5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • brew
    • apt
    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pandoc.org
    • github.com
    • pdf2docx.readthedocs.io
    • help.libreoffice.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Batch Convert loads about 3.8k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 237 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from claude-office-skills/skills at commit 9c4c7d5, republished under its MIT licence (© claude-office-skills). 237 words, ~3,843 tokens.

Download SKILL.mdSave it as .claude/skills/batch-convert/SKILL.md (or your agent's skills folder).
name
batch-convert
description
Batch convert documents between multiple formats using a unified pipeline
version
1.0
author
claude-office-skills
license
MIT
category
conversion
tags
batch, conversion, bulk, automation
department
All
models.recommended
claude-sonnet-4, claude-opus-4
models.compatible
claude-3-5-sonnet, gpt-4, gpt-4o
mcp.server
office-mcp
mcp.tools
batch_convert

Batch Convert Skill

Overview

This skill enables batch conversion of documents between multiple formats using a unified pipeline. Convert hundreds of files at once with consistent settings, automatic format detection, and parallel processing for maximum efficiency.

How to Use

  1. Specify the source folder or files
  2. Choose target format(s)
  3. Optionally configure conversion options
  4. I'll process all files with progress tracking

Example prompts:

  • "Convert all PDFs in this folder to Word documents"
  • "Batch convert these markdown files to PDF and HTML"
  • "Process all Office files and convert to Markdown"
  • "Convert this folder of images to a single PDF"

Domain Knowledge

Supported Format Matrix
FromTo: DOCXTo: PDFTo: MDTo: HTMLTo: PPTX
DOCX-✅✅✅-
PDF✅-✅✅-
MD✅✅-✅✅
HTML✅✅✅--
XLSX-✅✅✅-
PPTX-✅✅✅-
Core Pipeline
python
from pathlib import Path
from concurrent.futures import ThreadPoolExecutor, as_completed
import subprocess
import os

class DocumentConverter:
    """Unified document conversion pipeline."""
    
    def __init__(self, max_workers=4):
        self.max_workers = max_workers
        self.converters = {
            ('md', 'docx'): self._md_to_docx,
            ('md', 'pdf'): self._md_to_pdf,
            ('md', 'html'): self._md_to_html,
            ('md', 'pptx'): self._md_to_pptx,
            ('docx', 'pdf'): self._docx_to_pdf,
            ('docx', 'md'): self._docx_to_md,
            ('pdf', 'docx'): self._pdf_to_docx,
            ('pdf', 'md'): self._pdf_to_md,
            ('xlsx', 'pdf'): self._xlsx_to_pdf,
            ('xlsx', 'md'): self._xlsx_to_md,
            ('pptx', 'pdf'): self._pptx_to_pdf,
            ('pptx', 'md'): self._pptx_to_md,
            ('html', 'md'): self._html_to_md,
            ('html', 'pdf'): self._html_to_pdf,
        }
    
    def convert(self, input_path, output_format, output_dir=None):
        """Convert single file to target format."""
        input_path = Path(input_path)
        input_format = input_path.suffix[1:].lower()
        
        if output_dir:
            output_path = Path(output_dir) / f"{input_path.stem}.{output_format}"
        else:
            output_path = input_path.with_suffix(f".{output_format}")
        
        converter_key = (input_format, output_format)
        if converter_key not in self.converters:
            raise ValueError(f"Conversion not supported: {input_format} -> {output_format}")
        
        converter = self.converters[converter_key]
        return converter(input_path, output_path)
    
    def batch_convert(self, input_dir, output_format, output_dir=None, 
                      file_pattern="*", recursive=False):
        """Batch convert all matching files."""
        input_path = Path(input_dir)
        output_path = Path(output_dir) if output_dir else input_path / "converted"
        output_path.mkdir(exist_ok=True)
        
        # Find files
        if recursive:
            files = list(input_path.rglob(file_pattern))
        else:
            files = list(input_path.glob(file_pattern))
        
        # Filter to supported formats
        supported_ext = ['.md', '.docx', '.pdf', '.xlsx', '.pptx', '.html']
        files = [f for f in files if f.suffix.lower() in supported_ext]
        
        results = []
        
        with ThreadPoolExecutor(max_workers=self.max_workers) as executor:
            future_to_file = {
                executor.submit(self.convert, f, output_format, output_path): f
                for f in files
            }
            
            for future in as_completed(future_to_file):
                file = future_to_file[future]
                try:
                    result = future.result()
                    results.append({'file': str(file), 'status': 'success', 'output': str(result)})
                except Exception as e:
                    results.append({'file': str(file), 'status': 'error', 'error': str(e)})
        
        return results
Converter Implementations
python
# Markdown conversions (using Pandoc)
def _md_to_docx(self, input_path, output_path):
    subprocess.run(['pandoc', str(input_path), '-o', str(output_path)], check=True)
    return output_path

def _md_to_pdf(self, input_path, output_path):
    subprocess.run(['pandoc', str(input_path), '-o', str(output_path)], check=True)
    return output_path

def _md_to_html(self, input_path, output_path):
    subprocess.run(['pandoc', str(input_path), '-s', '-o', str(output_path)], check=True)
    return output_path

def _md_to_pptx(self, input_path, output_path):
    subprocess.run(['marp', str(input_path), '-o', str(output_path)], check=True)
    return output_path

# Office to Markdown (using markitdown)
def _docx_to_md(self, input_path, output_path):
    from markitdown import MarkItDown
    md = MarkItDown()
    result = md.convert(str(input_path))
    with open(output_path, 'w') as f:
        f.write(result.text_content)
    return output_path

def _xlsx_to_md(self, input_path, output_path):
    from markitdown import MarkItDown
    md = MarkItDown()
    result = md.convert(str(input_path))
    with open(output_path, 'w') as f:
        f.write(result.text_content)
    return output_path

def _pptx_to_md(self, input_path, output_path):
    from markitdown import MarkItDown
    md = MarkItDown()
    result = md.convert(str(input_path))
    with open(output_path, 'w') as f:
        f.write(result.text_content)
    return output_path

# PDF conversions
def _pdf_to_docx(self, input_path, output_path):
    from pdf2docx import Converter
    cv = Converter(str(input_path))
    cv.convert(str(output_path))
    cv.close()
    return output_path

def _pdf_to_md(self, input_path, output_path):
    from markitdown import MarkItDown
    md = MarkItDown()
    result = md.convert(str(input_path))
    with open(output_path, 'w') as f:
        f.write(result.text_content)
    return output_path

# Office to PDF (using LibreOffice)
def _docx_to_pdf(self, input_path, output_path):
    subprocess.run([
        'soffice', '--headless', '--convert-to', 'pdf',
        '--outdir', str(output_path.parent), str(input_path)
    ], check=True)
    return output_path

def _xlsx_to_pdf(self, input_path, output_path):
    subprocess.run([
        'soffice', '--headless', '--convert-to', 'pdf',
        '--outdir', str(output_path.parent), str(input_path)
    ], check=True)
    return output_path

def _pptx_to_pdf(self, input_path, output_path):
    subprocess.run([
        'soffice', '--headless', '--convert-to', 'pdf',
        '--outdir', str(output_path.parent), str(input_path)
    ], check=True)
    return output_path
Progress Tracking
python
from tqdm import tqdm

def batch_convert_with_progress(converter, input_dir, output_format, output_dir=None):
    """Batch convert with progress bar."""
    input_path = Path(input_dir)
    files = list(input_path.glob('*'))
    
    results = []
    for file in tqdm(files, desc=f"Converting to {output_format}"):
        try:
            result = converter.convert(file, output_format, output_dir)
            results.append({'file': str(file), 'status': 'success'})
        except Exception as e:
            results.append({'file': str(file), 'status': 'error', 'error': str(e)})
    
    return results

Best Practices

  1. Test Sample First: Convert a few files before batch processing
  2. Check Disk Space: Ensure sufficient space for output
  3. Use Parallel Processing: Speed up with multiple workers
  4. Handle Errors Gracefully: Log failures, continue processing
  5. Verify Output: Spot-check converted files

Common Patterns

Format Detection Pipeline
python
def detect_and_convert(file_path, target_format):
    """Automatically detect format and convert."""
    import mimetypes
    
    mime_type, _ = mimetypes.guess_type(str(file_path))
    
    format_map = {
        'application/pdf': 'pdf',
        'application/vnd.openxmlformats-officedocument.wordprocessingml.document': 'docx',
        'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet': 'xlsx',
        'application/vnd.openxmlformats-officedocument.presentationml.presentation': 'pptx',
        'text/markdown': 'md',
        'text/html': 'html',
    }
    
    source_format = format_map.get(mime_type, Path(file_path).suffix[1:])
    
    converter = DocumentConverter()
    return converter.convert(file_path, target_format)
Multi-Format Output
python
def convert_to_multiple_formats(input_file, output_formats, output_dir):
    """Convert one file to multiple formats."""
    converter = DocumentConverter()
    results = {}
    
    for fmt in output_formats:
        try:
            output = converter.convert(input_file, fmt, output_dir)
            results[fmt] = {'status': 'success', 'path': str(output)}
        except Exception as e:
            results[fmt] = {'status': 'error', 'error': str(e)}
    
    return results

# Convert README to multiple formats
results = convert_to_multiple_formats(
    'README.md',
    ['docx', 'pdf', 'html'],
    './exports'
)

Examples

Example 1: Documentation Export
python
from pathlib import Path
import json

def export_documentation(docs_dir, export_dir):
    """Export all documentation to multiple formats."""
    
    converter = DocumentConverter(max_workers=8)
    docs_path = Path(docs_dir)
    export_path = Path(export_dir)
    
    # Create format directories
    for fmt in ['pdf', 'docx', 'html']:
        (export_path / fmt).mkdir(parents=True, exist_ok=True)
    
    all_results = {}
    
    # Find all markdown files
    md_files = list(docs_path.rglob('*.md'))
    
    for md_file in md_files:
        file_results = {}
        
        for fmt in ['pdf', 'docx', 'html']:
            output_dir = export_path / fmt
            try:
                output = converter.convert(md_file, fmt, output_dir)
                file_results[fmt] = 'success'
            except Exception as e:
                file_results[fmt] = f'error: {e}'
        
        all_results[str(md_file)] = file_results
        print(f"Processed: {md_file.name}")
    
    # Save report
    with open(export_path / 'export_report.json', 'w') as f:
        json.dump(all_results, f, indent=2)
    
    return all_results

results = export_documentation('./docs', './exports')
Example 2: Legacy Document Migration
python
def migrate_legacy_docs(source_dir, target_dir):
    """Migrate legacy documents to modern formats."""
    
    converter = DocumentConverter(max_workers=4)
    
    # Migration rules
    migrations = [
        ('*.doc', 'docx'),   # Old Word to new
        ('*.xls', 'xlsx'),   # Old Excel to new
        ('*.ppt', 'pptx'),   # Old PowerPoint to new
        ('*.rtf', 'docx'),   # RTF to Word
    ]
    
    source_path = Path(source_dir)
    target_path = Path(target_dir)
    target_path.mkdir(exist_ok=True)
    
    total_migrated = 0
    errors = []
    
    for pattern, target_format in migrations:
        files = list(source_path.glob(pattern))
        
        for file in files:
            try:
                # Use LibreOffice for legacy formats
                subprocess.run([
                    'soffice', '--headless',
                    '--convert-to', target_format,
                    '--outdir', str(target_path),
                    str(file)
                ], check=True)
                
                total_migrated += 1
                print(f"Migrated: {file.name}")
                
            except Exception as e:
                errors.append({'file': str(file), 'error': str(e)})
    
    print(f"\nMigration complete: {total_migrated} files")
    print(f"Errors: {len(errors)}")
    
    return {'migrated': total_migrated, 'errors': errors}
Example 3: Report Generation Pipeline
python
def generate_reports_pipeline(data_files, template_dir, output_dir):
    """Generate reports from data files using templates."""
    
    from datetime import datetime
    
    converter = DocumentConverter()
    output_path = Path(output_dir)
    output_path.mkdir(exist_ok=True)
    
    timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
    
    reports = []
    
    for data_file in data_files:
        # Load data
        data_path = Path(data_file)
        
        # Generate markdown report
        md_content = f"""---
title: Report - {data_path.stem}
date: {datetime.now().strftime('%Y-%m-%d')}
---

# {data_path.stem} Report

Generated: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}

## Data Summary

"""
        
        # Add data content (simplified)
        if data_path.suffix == '.xlsx':
            from markitdown import MarkItDown
            md = MarkItDown()
            result = md.convert(str(data_path))
            md_content += result.text_content
        
        # Save markdown
        md_file = output_path / f"{data_path.stem}_{timestamp}.md"
        with open(md_file, 'w') as f:
            f.write(md_content)
        
        # Convert to PDF and DOCX
        for fmt in ['pdf', 'docx']:
            try:
                output = converter.convert(md_file, fmt, output_path)
                reports.append({'source': str(data_file), 'output': str(output), 'format': fmt})
            except Exception as e:
                print(f"Error converting {data_file} to {fmt}: {e}")
    
    return reports

Limitations

  • Some format combinations not supported
  • Complex formatting may be lost in conversion
  • Large files may require more time
  • Some conversions need external tools (LibreOffice, Pandoc)
  • Quality varies by source document complexity

Installation

bash
# Core dependencies
pip install pdf2docx markitdown python-docx openpyxl

# Pandoc (for MD conversions)
brew install pandoc  # macOS
apt install pandoc   # Ubuntu

# Marp (for PPTX)
npm install -g @marp-team/marp-cli

# LibreOffice (for Office formats)
brew install libreoffice  # macOS
apt install libreoffice   # Ubuntu

Resources

© claude-office-skills, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in batch-convert of claude-office-skills/skills.

Open the folder on GitHubat commit 9c4c7d5

Compare with similar skills

Batch Convert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Batch Convert compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Batch Convert this skillclaude-office-skills/skills505—~3.8kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
DOCX ToolkitXiaomiMiMo/MiMo-Code14k—~2.4kAutomated safety check: PassApache-2.0
Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html910—~4.8kAutomated safety check: PassMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
Liteparsebastani-inc/atomic855—~1.4kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    910 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Liteparse

    bastani-inc/atomic

    A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

    855 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Beautiful Article

    ConardLi/garden-skills

    Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow.

    13k GitHub stars~4.7k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from claude-office-skills/skills

All 12 skills in this repo
  • Airtable Automation

    claude-office-skills/skills

    Airtable database automation - views, automations, integrations, and workflow triggers

    505 GitHub stars~2.7k tokensUpdated 8 mo ago
    Auto-check passed
  • Academic Search

    claude-office-skills/skills

    Search and analyze academic literature. An agent skill from claude-office-skills/skills.

    505 GitHub stars~2.2k tokensUpdated 8 mo ago
    Auto-check passed
  • Ads Copywriter

    claude-office-skills/skills

    Multi-platform ad copy generation for Google Ads, Meta/Facebook, TikTok, LinkedIn with A/B testing variants

    505 GitHub stars~2.1k tokensUpdated 8 mo ago
    Auto-check passed
  • AI Agent Builder

    claude-office-skills/skills

    Build AI agents with tools, memory, and multi-step reasoning - ChatGPT, Claude, Gemini integration patterns

    505 GitHub stars~3.1k tokensUpdated 8 mo ago
    Auto-check passed
  • AI Slides

    claude-office-skills/skills

    Generate complete presentations with AI - from outline to polished slides

    505 GitHub stars~1.4k tokensUpdated 8 mo ago
    Auto-check passed
  • Batch Processor

    claude-office-skills/skills

    Process multiple documents in bulk with parallel execution. An agent skill from claude-office-skills/skills.

    505 GitHub stars~1.1k tokensUpdated 8 mo ago
    Auto-check passed

Questions about Batch Convert

What does Batch Convert do?

Batch convert documents between multiple formats using a unified pipeline. Batch Convert is an agent skill from claude-office-skills/skills.

When should I use Batch Convert?

Batch Convert fits situations like: tasks that involve Document parsing.

How do I install Batch Convert in Claude Code?

Run `npx skills add claude-office-skills/skills --skill batch-convert -a claude-code`. Or copy the skill folder (batch-convert in claude-office-skills/skills) into .claude/skills/batch-convert in your project. Claude Code loads it when a task matches its description.

How do I install Batch Convert in Codex?

Run `npx skills add claude-office-skills/skills --skill batch-convert -a codex`. Or copy the skill folder (batch-convert in claude-office-skills/skills) into .agents/skills/batch-convert in your project. Codex loads it when a task matches its description.

Can I use Batch Convert in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add claude-office-skills/skills --skill batch-convert -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-convert, .gemini/skills/batch-convert, .github/skills/batch-convert and .opencode/skills/batch-convert in your project.

What does Batch Convert need to run?

Going by SKILL.md and its folder, Batch Convert needs the command-line tools its instructions call (brew, apt, pip and npm). Our summary lists: Python 3.

Does Batch Convert access the network?

SKILL.md names 4 domains. As links in the text: pandoc.org, github.com, pdf2docx.readthedocs.io and help.libreoffice.org. This is read from the text; nothing was executed.

Is Batch Convert safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Batch Convert use?

Batch Convert is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Batch Convert use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Batch Convert?

Skills that share tags, products or a category with Batch Convert: Markitdown (ImCa0/just-laws, 781 stars), DOCX Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 910 stars) and Markitdown (jimmc414/Kosmos, 595 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Batch Convert?

claude-office-skills (a GitHub organization) maintains it in claude-office-skills/skills, which has 505 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on January 31, 2026.

Source: claude-office-skills/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.