Agent skill

Doc Cleaner

by notoriouslab in notoriouslab/doc-cleaner

Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

MITAuto-check passedDocuments & Office

Install Doc Cleaner

skills CLI
$ npx skills add notoriouslab/doc-cleaner --skill doc-cleaner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notoriouslab/doc-cleaner doc-cleaner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doc-cleaner
GitHub stars
309
Token cost
~712 tokens
SKILL.md length
241 words
Files
109 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

  • Tasks that involve Word documents
  • SKILL.md covers When to use, Commands, Options and Supported formats, plus 2 more sections
  • Runs Python scripts from its folder; calls python3
  • Tasks that involve Excel spreadsheets

What it does

Doc Cleaner is an agent skill from notoriouslab/doc-cleaner. Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown. CJK-friendly, table-friendly, privacy-first.

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. The skill folder holds 111 other files, including scripts (for example `.github/workflows/build-windows.yml`, `.github/workflows/tests.yml` and `AGENTS.md`).

It sits in Documents & Office, covering Word documents, Excel spreadsheets and PDF. It works with Microsoft Excel, Microsoft Word and Python. The repository describes itself as: 日常文件轉 Markdown —— 涵蓋 PDF、Office、Apple Keynote/Numbers、EPUB 電子書等 16 種格式。中文友好、表格保留、隱私優先、全程本地。 The licence is MIT.

When your agent uses it

  • Tasks that involve Word documents
  • Tasks that involve Excel spreadsheets
  • Tasks that involve PDF

Example prompts

  • “/doc-cleaner”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 55e3c47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doc Cleaner loads about 712 tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 241 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~712

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from notoriouslab/doc-cleaner at commit 55e3c47, republished under its MIT licence (© notoriouslab). 241 words, ~712 tokens.

Download SKILL.mdSave it as .claude/skills/doc-cleaner/SKILL.md (or your agent's skills folder). This skill also uses 108 other files; get the full folder from GitHub.
name
doc-cleaner
description
Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown. CJK-friendly, table-friendly, privacy-first.
version
1.0.0

doc-cleaner

Convert documents (PDF, DOCX, XLSX, TXT) to clean, structured Markdown.

When to use

  • User asks to convert a document to Markdown
  • User wants to extract text or tables from PDF/DOCX/XLSX files
  • User wants to clean up bank statements or financial documents
  • User asks to process a batch of documents in a directory

Commands

Convert a single file (no AI, fastest)
bash
python3 {baseDir}/cleaner.py --input "{{file_path}}" --ai none
Convert a single file with AI structuring
bash
python3 {baseDir}/cleaner.py --input "{{file_path}}" --ai gemini
Convert a single file with Groq structuring
bash
python3 {baseDir}/cleaner.py --input "{{file_path}}" --ai groq
Convert all files in a directory
bash
python3 {baseDir}/cleaner.py --input "{{directory}}" --ai none --output-dir "{{output_dir}}"
Preview without writing (dry run)
bash
python3 {baseDir}/cleaner.py --input "{{file_path}}" --dry-run --verbose
Get machine-readable result summary
bash
python3 {baseDir}/cleaner.py --input "{{file_path}}" --ai none --summary

The --summary flag prints a JSON summary to stdout after processing:

json
{"version":"1.0.0","total":3,"success":2,"failed":1,"files":[{"file":"report.pdf","output":"./output/report.md","status":"ok"},{"file":"scan.pdf","output":null,"status":"no_content"},{"file":"data.xlsx","output":"./output/data.md","status":"ok"}]}

Options

FlagDescription
--input, -iFile or directory to process (required, non-recursive)
--output-dir, -oOutput directory (default: ./output)
--aigemini, groq, ollama, or none (default: from config or gemini)
--passwordPDF decryption password
--configPath to config JSON
--summaryPrint JSON summary to stdout after processing
--dry-runPreview without writing files
--verboseEnable debug logging

Supported formats

PDF (native, scanned, encrypted), DOCX, XLSX, XLS, CSV, TXT, MD

Exit codes

CodeMeaning
0All files processed successfully
1Some files failed (partial success)
2No processable files found or config error

Notes

  • Output defaults to ./output/ relative to current directory
  • For scanned PDFs, AI mode (gemini, groq, or ollama) gives much better results
  • --ai none requires zero API keys and zero network access
  • CJK encoding (Big5, CP950, UTF-16) is auto-detected
  • Tables in DOCX and XLSX are preserved as Markdown pipe tables

© notoriouslab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 108 other files (scripts) in the repository root of notoriouslab/doc-cleaner.

  • SKILL.md
  • .env.example
  • .github/workflows/build-windows.yml
  • .github/workflows/tests.yml
  • .gitignore
  • AGENTS.md
  • CHANGELOG.md
  • CONTRIBUTING.md
  • LICENSE
  • README.en.md
  • README.md
  • ReadMe.txt
  • SECURITY.md
  • ai/__init__.py
  • ai/base.py
  • ai/gemini.py
  • ai/groq.py
  • ai/mlx.py
  • … and 91 more

Open the folder on GitHubat commit 55e3c47

Compare with similar skills

Doc Cleaner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doc Cleaner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doc Cleaner this skillnotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~1.4kAutomated safety check: PassMIT
Markdown Exporterbowenliang123/markdown-exporter2721 repos~5.3kAutomated safety check: PassApache-2.0
Markdown ConverterTeam-Commonly/commonly1.4k—~557Automated safety check: PassApache-2.0
Document ConverterBlackBeltTechnology/pi-agent-dashboard315—~999Automated safety check: PassMIT

Similar skills

  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    272 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed
  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    315 GitHub stars~999 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes

Questions about Doc Cleaner

What does Doc Cleaner do?

Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown. Doc Cleaner is an agent skill from notoriouslab/doc-cleaner. Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

When should I use Doc Cleaner?

Doc Cleaner fits situations like: tasks that involve Word documents; tasks that involve Excel spreadsheets; tasks that involve PDF.

How do I install Doc Cleaner in Claude Code?

Run `npx skills add notoriouslab/doc-cleaner --skill doc-cleaner -a claude-code`. Or copy the skill folder (the notoriouslab/doc-cleaner repository) into .claude/skills/doc-cleaner in your project. Claude Code loads it when a task matches its description.

How do I install Doc Cleaner in Codex?

Run `npx skills add notoriouslab/doc-cleaner --skill doc-cleaner -a codex`. Or copy the skill folder (the notoriouslab/doc-cleaner repository) into .agents/skills/doc-cleaner in your project. Codex loads it when a task matches its description.

Can I use Doc Cleaner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notoriouslab/doc-cleaner --skill doc-cleaner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-cleaner, .gemini/skills/doc-cleaner, .github/skills/doc-cleaner and .opencode/skills/doc-cleaner in your project.

What does Doc Cleaner need to run?

Going by SKILL.md and its folder, Doc Cleaner needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Doc Cleaner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Doc Cleaner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Doc Cleaner use?

Doc Cleaner is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doc Cleaner use?

About 712 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doc Cleaner?

Skills that share tags, products or a category with Doc Cleaner: Mineru (Nebutra/MinerU-Skill, 122 stars), Mineru (Nebutra/MinerU-Skill, 122 stars), Markdown Exporter (bowenliang123/markdown-exporter, 272 stars) and Markdown Converter (Team-Commonly/commonly, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doc Cleaner?

notoriouslab (a GitHub user) maintains it in notoriouslab/doc-cleaner, which has 309 GitHub stars. The repository was last updated on August 20, 2026.

Source: notoriouslab/doc-cleaner on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.