Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
PDF 处理工具,支持扫描件预处理、OCR 双层 PDF、页码添加、PDF 合并、解密、水印去除和压缩。本技能应在用户需要一键处理、优化或整理 PDF 文档时使用。不要用于:纯文本 PDF 内容编辑、PDF 阅读与批注、电子签名、非压缩目的的格式转换。
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cat-xierluo/legal-skills pdf-processor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf-processor .claude/skills/pdf-processor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .claude/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cat-xierluo/legal-skills pdf-processor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/pdf-processor .agents/skills/pdf-processor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .agents/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cat-xierluo/legal-skills pdf-processor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/pdf-processor .cursor/skills/pdf-processor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .cursor/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cat-xierluo/legal-skills.git --path skills/pdf-processor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cat-xierluo/legal-skills pdf-processor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/pdf-processor .gemini/skills/pdf-processor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .gemini/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cat-xierluo/legal-skills pdf-processorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/pdf-processor .github/skills/pdf-processor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .github/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cat-xierluo/legal-skills --skill pdf-processor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cat-xierluo/legal-skills pdf-processor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/pdf-processor .opencode/skills/pdf-processor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-processor" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/pdf-processor into .opencode/skills/pdf-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-processor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-processorPDF 处理工具,支持扫描件预处理、OCR 双层 PDF、页码添加、PDF 合并、解密、水印去除和压缩。本技能应在用户需要一键处理、优化或整理 PDF 文档时使用。不要用于:纯文本 PDF 内容编辑、PDF 阅读与批注、电子签名、非压缩目的的格式转换。
PDF Processor is an agent skill from cat-xierluo/legal-skills. PDF 处理工具,支持扫描件预处理、OCR 双层 PDF、页码添加、PDF 合并、解密、水印去除和压缩。本技能应在用户需要一键处理、优化或整理 PDF 文档时使用。不要用于:纯文本 PDF 内容编辑、PDF 阅读与批注、电子签名、非压缩目的的格式转换。
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 53 other files, including scripts and reference files (for example `CHANGELOG.md`, `references/mineru-api-guide.md` and `references/ocr-backend-guide.md`).
It sits in Documents & Office, covering PDF. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c077fcc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3pipbrewapt-getpdftotextFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Processor loads about 2.2k tokens when it runs, and up to ~787k if it reads all its reference files. Until then it costs about 35 tokens; SKILL.md has 338 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo apt-get install poppler-utilssudo apt-get install tesseract-ocr tesseract-ocr-chi-simAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from cat-xierluo/legal-skills at commit c077fcc, republished under its MIT licence (© cat-xierluo). 338 words, ~2,175 tokens.
.claude/skills/pdf-processor/SKILL.md (or your agent's skills folder). This skill also uses 49 other files; get the full folder from GitHub.本技能是 PDF 处理的统一入口,覆盖扫描件预处理、OCR 双层 PDF 生成、页码添加、PDF 合并、解密、水印去除和压缩。优先保护原始文件,按用户意图选择最短可用流程。
核心职责:
本技能不做纯文本 PDF 内容编辑、PDF 阅读批注、电子签名、非压缩目的的格式转换。
_1、_2 等序号。ocrmypdf 默认直接保留原 PDF,只有 MinerU 或显式图像处理请求才走统一栅格预处理。auto 在已配置时默认优先 PaddleOCR API,再尝试 MinerU,最后回退本地引擎(已安装 RapidOCR 时优先 RapidOCR,否则 ocrmypdf);明确禁止外传的材料必须使用 --local-only。--force-raster-preprocess。--preprocess-only。python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdfauto 选中已配置的 PaddleOCR 时,默认把原 PDF 直接送给 PP-OCRv6,跳过统一栅格化和 OCR 前压缩,以保留扫描分辨率、图像层和 API 坐标空间;实际后端确定为本地 ocrmypdf 时同样保留原扫描页,并使用 OCRmyPDF 自带的方向检测、纠偏和清理。MinerU 或显式预处理路径仍以 medium 的 200 DPI、JPEG 质量 72、色度子采样 1 为目标,并以 25MP 保护异常大画布。确需裁剪或统一重栅格化时使用 --enable-crop 或 --force-raster-preprocess。电子/混合 PDF 自动保留原有层。文件大小限制很严时使用:
python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdf --compress-level high确需保留超大栅格时可显式调高上限;--max-preprocess-megapixels 0 会关闭保护,但可能显著增加内存占用和 OCR 跳页风险。
敏感材料用统一入口但禁止外传:
python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdf --local-only页面方向已正确的大批量扫描件可提速:
python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdf \
--skip-coarse-rotation --preprocess-jobs 6 --preprocess-chunk-pages 80python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdf --preprocess-only只做页面矫正、不压缩、不 OCR:
python3 scripts/pdf-preprocess-ocr.py --input input.pdf --output output.pdf \
--preprocess-only --no-compresspython3 scripts/pdf-ocr.py --input input.pdf --output output.pdf默认后端为 auto:已配置 PaddleOCR 时先用 PP-OCRv6 的行级坐标生成文字层,再按 OCR_API_ORDER 尝试 MinerU;外部服务失败或未配置时回退本地引擎——已安装 RapidOCR 时优先本地 RapidOCR(中文行级识别质量好、不出本机),否则回退本地 ocrmypdf。该默认路径会上传完整 PDF,不允许外传时使用:
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf --local-only--local-only 同样走本地优先链(RapidOCR → ocrmypdf)。敏感材料的本地中文 OCR 推荐先安装 RapidOCR:
pip install rapidocr--allow-external-upload 仅为旧命令兼容参数,不再控制后端选择。Paddle 的服务端方向矫正和去畸变默认关闭,避免 OCR 坐标与原图空间不一致。
# 强制本地兜底
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf --backend local_ocrmypdf
# 强制本地 RapidOCR(onnx 本地推理,中文质量好且不出本机)
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf --backend rapidocr_local
# 强制 PaddleOCR API;PP-OCRv6 是双层 PDF 默认模型
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf \
--backend paddle_api --paddle-model PP-OCRv6
# 干净扫描件/表格可试 PP-StructureV3 的 overall_ocr_res 行级坐标
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf \
--backend paddle_api --paddle-model PP-StructureV3
# 强制 MinerU API
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf --backend mineru_api
# 显式保存 OCR 可读文本和运行元数据;默认不归档案件材料
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf --archive-results后端选择、API 配置和协议细节见 references/ocr-backend-guide.md、references/paddleocr-api-guide.md、references/mineru-api-guide.md。
PaddleOCR-VL-1.5/1.6 也可解析,但只提供块级坐标,文字层定位粒度低于 PP-OCRv5/v6 / PP-StructureV3。本技能不接入 Qwen/GLM 等视觉识别链路,也不宣称云端结果含字符级坐标。
PP-OCRv6 文字模型默认开启 --actualtext:文字层生成后再调一次 PP-StructureV3 拿版面,融合出自然段并以 /ActualText marked-content 写入 PDF。Paddle 文字层还会用本地几何规则识别正文多行段落,把段内字号向下统一到现有最小安全值,同时保留原行坐标和横向框宽,用于改善 PDF Expert 的段落推断。--no-actualtext 可关闭;--layout-dump FILE 复用已有版面 dump 避免二次 API 调用。阅读器兼容性见下文第 4 节。
双层 PDF 的文字层按行叠层(保护选区坐标精度),直接从 PDF 复制会按物理行断行。解决方式有两条:
默认:/ActualText + 独立 Markdown。 --actualtext(默认开启)把自然段写入 PDF 的 /ActualText marked-content,同时保留行级字形坐标。实测各阅读器支持:
| 提取方式 | 从 PDF 复制的段落连续性 |
|---|---|
Poppler(pdftotext -raw、pdftotext 默认) | ✅ 整段连续,无换行 |
| PDF Expert | ⚠️ 忽略 /ActualText,但会按字号和几何自行推断段落;Paddle 正文段落样式归一可减少换行,不保证完全消除 |
| PyMuPDF、pypdf、macOS 预览(PDFKit) | ❌ 仍按物理行断行 |
因此 Poppler 用户能直接从 PDF 拿到段落级文本;PDF Expert 只能做阅读器启发式改善;macOS 预览用户复制仍会断行。跨阅读器需要稳定段落文本时,显式输出独立 Markdown;默认不在案件材料旁生成正文副本:
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf \
--backend paddle_api --clean-text-output output-clean.md独立 Markdown 也可由 pdf_ocr_paragraphs.py 生成。个别自然段若跳过被排除的行(如印章碎片),因物理行不连续无法包裹,会降级为行级(文字不丢失,仅该段复制按行断行)。
需要解决复制文本中的段内回车、多余空行和印章文字时,采用两模型分工:PP-OCRv6 提供主要行级文字及坐标,PP-StructureV3 提供阅读顺序、区域类型和 seal 区域。只有 v6 行置信度低于 0.80、Structure 对应行置信度不低于 0.90、双方坐标高度重合且后者至少高 0.10 时,才采用 Structure 的行级文字兜底;始终不读取版面块文字,也不调用大语言模型。
Structure 的块边界只作为候选而非强制段界:融合器先合并同一视觉行的碎片,在 text 区域内按行距、缩进和右边界恢复物理换行,再依据句末标点和条款编号跨相邻文本块连接正文;table 区域按视觉行保留单元格次序并用 | 分隔。这样可处理 Structure 高覆盖但正文块过度切碎的页面。
需要独立 Markdown、自定义 dump 审查或复用已有版面 dump 时,用以下四步法(主流程已默认自动完成等价工作):
# 1. 获取文字真值(--dump-and-pdf 同时出 dump 和双层 PDF)
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf \
--backend paddle_api --paddle-model PP-OCRv6 \
--ocr-dump /tmp/text.json --dump-and-pdf
# 2. 获取版面结构
python3 scripts/pdf-ocr.py -i input.pdf -o unused.pdf \
--backend paddle_api --paddle-model PP-StructureV3 --ocr-dump /tmp/layout.json
# 3. 生成自然段文本,同时输出不含正文的诊断和已过滤文字层 dump
python3 scripts/pdf_ocr_paragraphs.py \
--text-dump /tmp/text.json --layout-dump /tmp/layout.json \
--output /tmp/clean.md --diagnostics /tmp/paragraphs.json \
--filtered-dump /tmp/text-filtered.json
# 4. 用过滤后的行级坐标生成双层 PDF
python3 scripts/pdf-ocr.py -i input.pdf -o output.pdf \
--backend paddle_api --ocr-resume /tmp/text-filtered.json诊断中的 layout_coverage 低于 0.75 时自动退回纯几何规则;同时查看 structure_text_fallbacks 和 layout_boundary_merges,确认低置信替换与跨块合并均可审计。只有 Structure 明显漏块、区域类型错误或阅读顺序错误,且纯几何回退仍不能恢复时,才另测 PaddleOCR-VL-1.5 作为 --layout-dump;VL-1.6 不作为默认兜底。全文、dump 与账号等敏感信息继续只放临时目录,除非用户明确要求归档。
# 手动旋转
python3 scripts/pdf-rotate.py --input input.pdf --output output.pdf --angle 90
# 解密
python3 scripts/pdf-decrypt.py --input input.pdf --output output.pdf
python3 scripts/pdf-decrypt.py --input input.pdf --output output.pdf --password 123456
# 去水印
python3 scripts/pdf-remove-watermark.py --input input.pdf --output output.pdf
# 压缩
python3 scripts/pdf-compress.py -i input.pdf -o output.pdf --level medium
# 加页码
python3 scripts/pdf-add-page-numbers.py -i input.pdf -o output.pdf
# 合并
python3 scripts/pdf-merge.py -i file1.pdf file2.pdf file3.pdf -o merged.pdf
python3 scripts/pdf-merge.py -i file1.pdf file2.pdf -o merged.pdf --add-numbers --continuous页码、合并、压缩等详细参数见 references/pdf-workflows.md。
pip install pymupdf pypdf pillow numpy opencv-python pdf2imagemacOS:
brew install popplerLinux:
sudo apt-get install poppler-utilspip install ocrmypdfmacOS:
brew install tesseract tesseract-langLinux:
sudo apt-get install tesseract-ocr tesseract-ocr-chi-simpip install rapidocr安装后 auto 的本地回退与 --local-only 都会优先使用 RapidOCR(onnx 本地推理,中文行级识别质量明显高于 tesseract,且全程不出本机);首次运行自动下载检测/识别模型(各约 5-16MB,仅一次)。也可显式指定 --backend rapidocr_local。
完整可选依赖清单见 references/optional-dependencies.txt。历史保留的本地 Paddle 双层实现已拆到 scripts/pdf_ocr_paddle_local.py,不属于默认生产链路;需要实验时再安装 paddleocr paddlepaddle 并单独接入。
python3 scripts/pdf-ocr-quality-check.py \
-i input.pdf -o output.pdf --keywords 合同,法院
python3 scripts/pdf-ocr-benchmark.py \
-i input.pdf \
--backend local_ocrmypdf \
--sample-pages 5 \
--skip-coarse-rotation \
--preprocess-jobs 6 \
--preprocess-chunk-pages 80关键词门禁会先做 NFKC、大小写和空白归一化,避免中文 OCR 在汉字间插入空格后被误判为未命中;CER 仍按独立参考文本计算。
常见问题见 references/troubleshooting.md。
PaddleOCR API 的行级 poly 使用正向渲染图的左上原点像素坐标;PyMuPDF 页面坐标同样使用左上原点。不要按文字分布猜测 y 原点,也不要执行 page_height - y 翻转。
输入 PDF 若依赖 rotation 元数据装正,Paddle 路径会先用 PyMuPDF 无损移除 rotation,把原页面内容变换为 rotation=0 的正存页面,再按原有缩放和文字排版逻辑叠层。该过程不重采样扫描图,不改变可见页面像素。MinerU 等尚未验证坐标契约的后端继续保持原行为,不复用 Paddle 的坐标假设。
© cat-xierluo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 49 other files (scripts, references) in skills/pdf-processor of cat-xierluo/legal-skills.
Open the folder on GitHubat commit c077fcc
PDF Processor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Processor this skillcat-xierluo/legal-skills | 713 | — | ~2.2k | Automated safety check: Notes | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Gzh Designisjiamu/gzh-design-skill | 3.9k | 1 repos | ~2.2k | Automated safety check: Pass | AGPL-3.0 | |
| GenOffice Document CLIgenspark-ai/genoffice | 8.8k | — | ~19k | Automated safety check: Pass | Apache-2.0 | |
| Harness Book Best Practicewquguru/harness-books | 3.2k | — | ~4.1k | Automated safety check: Pass | None | |
| Bookforge Korean Ebook PDF Makergongnyang/bookforge | 314 | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
isjiamu/gzh-design-skill
微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…
genspark-ai/genoffice
Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.
wquguru/harness-books
Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.
gongnyang/bookforge
Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
cat-xierluo/legal-skills
Converts a lawyer's ordinary complaint or a described case into the Supreme People's Court's elements-style Word template, with layout checks on the result.
cat-xierluo/legal-skills
Analyzes raw lecture transcripts for verbal tics, pacing, time use and promise follow-through, with optional slide-by-slide comparison and cross-session tracking.
cat-xierluo/legal-skills
Detects and rewrites machine-sounding patterns in the body text of Chinese articles while keeping the author's facts, headings and legal terms intact.
cat-xierluo/legal-skills
Finds GitHub projects mentioned in articles or screenshots and stars them, tracks updates to your starred repos, and builds an HTML dashboard to browse them.
cat-xierluo/legal-skills
Sets up or incrementally updates AGENTS.md and CLAUDE.md for legal professionals, with a minimal safety baseline and a check that a new session loads and follows the rules.
cat-xierluo/legal-skills
Chinese-language skill that organizes a case file into a multi-role mock trial with judge, parties and clerk, producing a transcript, issue review and a to-strengthen list.
Categories
PDF 处理工具,支持扫描件预处理、OCR 双层 PDF、页码添加、PDF 合并、解密、水印去除和压缩。本技能应在用户需要一键处理、优化或整理 PDF 文档时使用。不要用于:纯文本 PDF 内容编辑、PDF 阅读与批注、电子签名、非压缩目的的格式转换。. PDF Processor is an agent skill from cat-xierluo/legal-skills.
PDF Processor fits situations like: tasks that involve PDF.
Run `npx skills add cat-xierluo/legal-skills --skill pdf-processor -a claude-code`. Or copy the skill folder (skills/pdf-processor in cat-xierluo/legal-skills) into .claude/skills/pdf-processor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cat-xierluo/legal-skills --skill pdf-processor -a codex`. Or copy the skill folder (skills/pdf-processor in cat-xierluo/legal-skills) into .agents/skills/pdf-processor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cat-xierluo/legal-skills --skill pdf-processor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-processor, .gemini/skills/pdf-processor, .github/skills/pdf-processor and .opencode/skills/pdf-processor in your project.
Going by SKILL.md and its folder, PDF Processor needs Python for the scripts in its folder and the command-line tools its instructions call (python3, pip, brew, apt-get and pdftotext). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PDF Processor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 785k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with PDF Processor: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cat-xierluo (a GitHub user) maintains it in cat-xierluo/legal-skills, which has 713 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 6, 2026.
Source: cat-xierluo/legal-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.