Agent skill

Doc OCR

by davepoon in davepoon/buildwithclaude

文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via…

MITAuto-check passedDocuments & Office

Install Doc OCR

skills CLI
$ npx skills add davepoon/buildwithclaude --skill doc-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davepoon/buildwithclaude doc-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/all-skills/skills/doc-ocr .claude/skills/doc-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doc-ocr
GitHub stars
3.6k
Token cost
~244 tokens
SKILL.md length
47 words
Files
2 (incl. scripts)
Skills in repo
246
Repo updated
First seen
Licence
MIT

At a glance

文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via…

  • Works in 3 steps: 单个文件 → 批量目录 → Markdown 输出
  • Tasks that involve PDF
  • SKILL.md covers 触发条件, 使用步骤, 依赖(首次使用时安装) and 已知陷阱
  • Runs Python scripts from its folder; calls python3 and pip3

What it does

Doc OCR is an agent skill from davepoon/buildwithclaude. 文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via macOS Vision.

Its SKILL.md is about 240 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/dococr.py`).

It sits in Documents & Office, covering PDF and Forms and invoices. It works with macOS. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF
  • Tasks that involve Forms and invoices

Example prompts

  • “/doc-ocr”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 单个文件
  2. 批量目录
  3. Markdown 输出

What it can do on your machine

Read from SKILL.md and the folder at commit 10bfc43. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip3, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doc OCR loads about 244 tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 47 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~244

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davepoon/buildwithclaude at commit 10bfc43, republished under its MIT licence (© davepoon). 47 words, ~244 tokens.

Download SKILL.mdSave it as .claude/skills/doc-ocr/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
doc-ocr
description
文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via macOS Vision.
category
document-processing
license
MIT

Doc-OCR 文档文字识别

PDF / 扫描件 / 图片 → 可编辑文字。有文字层的 PDF 直接提取,扫描件自动 OCR(macOS Vision 自带,中英文)。

触发条件

用户提供 PDF/图片文件,要求:

  • "提取文字""转文字""OCR"
  • 处理扫描件、合同、发票、书页、截图

使用步骤

1. 单个文件
bash
python3 scripts/dococr.py 合同.pdf
python3 scripts/dococr.py 发票.jpg

输出保存为 <输入名>_ocr.txt。

2. 批量目录
bash
python3 scripts/dococr.py ./扫描件/ -o 全部.txt
3. Markdown 输出
bash
python3 scripts/dococr.py 书.pdf --md

依赖(首次使用时安装)

bash
pip3 install pymupdf pyobjc-framework-Vision

注意:OCR 依赖 macOS Vision(仅 macOS 可用)。Linux 需另装 tesseract 等引擎。

已知陷阱

  • 扫描件判定:PDF 文字层 <20 字自动走 OCR,正常 PDF 直接提取。
  • 手写体:Vision 对印刷体/清晰手写效果好,潦草手写不保证。
  • 隐私卖点:文件在本机处理,不上传第三方。

© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in plugins/all-skills/skills/doc-ocr of davepoon/buildwithclaude.

  • SKILL.md
  • scripts/dococr.py

Open the folder on GitHubat commit 10bfc43

Compare with similar skills

Doc OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doc OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doc OCR this skilldavepoon/buildwithclaude3.6k—~244Automated safety check: PassMIT
Code2mediatsingliuwin/autoclaw280—~956Automated safety check: PassMIT
Reportlabjimmc414/Kosmos5941 repos~4.2kAutomated safety check: PassNone
PDF ReadingWide-Moat/open-computer-use1261 repos~2.7kAutomated safety check: PassProprietary
PDF Processing Toolkittelagod/code-abyss244—~532Automated safety check: NotesMIT
Document Converter Suitedkyazzentwatwa/chatgpt-skills112—~386Automated safety check: PassNone

Similar skills

  • Code2media

    tsingliuwin/autoclaw

    Universal media renderer with no fixed template — any custom layout, size or style becomes pixel-perfect images (PNG/JPEG/WebP), vector SVG, paged PDFs or animations (WebP/GIF).

    280 GitHub stars~956 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Reportlab

    jimmc414/Kosmos

    PDF generation toolkit. An agent skill from jimmc414/Kosmos.

    594 GitHub starsUsed in 1 repo~4.2k tokens
    Documents & OfficeAuto-check passed
  • PDF Reading

    Wide-Moat/open-computer-use

    A skill your agent uses when you need to read, inspect, or extract content from PDF files — especially when file content is NOT in your context and you need to read it from disk.

    126 GitHub starsUsed in 1 repo~2.7k tokens
    Documents & OfficeAuto-check passed
  • PDF Processing Toolkit

    telagod/code-abyss

    Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

    244 GitHub stars~532 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Document Converter Suite

    dkyazzentwatwa/chatgpt-skills

    Convert PDFs, Office docs, markdown, HTML, and tables between editable formats.

    112 GitHub stars~386 tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Create PDF

    theexperiencecompany/gaia

    Generate a polished, printable PDF — reports, letters, invoices, resumes, one-pagers.

    308 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed

More from davepoon/buildwithclaude

All 246 skills in this repo
  • Qwen Vision

    davepoon/buildwithclaude

    A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

    3.6k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Hard Predict Future

    davepoon/buildwithclaude

    Activate this agent for any future-oriented question that requires deep quantitative analysis, historical precedents, and structured scenario planning.

    3.6k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • iOS Hig Design Guide

    davepoon/buildwithclaude

    Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.

    3.6k GitHub stars~735 tokensUpdated yesterday
    Auto-check passed
  • Video Downloader

    davepoon/buildwithclaude

    Download YouTube videos with customizable quality and format options.

    3.6k GitHub starsUsed in 1 repo~871 tokens
    Auto-check passed
  • Atlas Cloud Media

    davepoon/buildwithclaude

    Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.

    3.6k GitHub stars~852 tokensUpdated yesterday
    Auto-check passed
  • Slack Gif Creator

    davepoon/buildwithclaude

    Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.

    3.6k GitHub starsUsed in 12 repos~4.3k tokens
    Auto-check passed

Works with

Questions about Doc OCR

What does Doc OCR do?

文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via…. Doc OCR is an agent skill from davepoon/buildwithclaude. 文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via macOS Vision.

When should I use Doc OCR?

Doc OCR fits situations like: tasks that involve PDF; tasks that involve Forms and invoices.

How do I install Doc OCR in Claude Code?

Run `npx skills add davepoon/buildwithclaude --skill doc-ocr -a claude-code`. Or copy the skill folder (plugins/all-skills/skills/doc-ocr in davepoon/buildwithclaude) into .claude/skills/doc-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Doc OCR in Codex?

Run `npx skills add davepoon/buildwithclaude --skill doc-ocr -a codex`. Or copy the skill folder (plugins/all-skills/skills/doc-ocr in davepoon/buildwithclaude) into .agents/skills/doc-ocr in your project. Codex loads it when a task matches its description.

Can I use Doc OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill doc-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-ocr, .gemini/skills/doc-ocr, .github/skills/doc-ocr and .opencode/skills/doc-ocr in your project.

What does Doc OCR need to run?

Going by SKILL.md and its folder, Doc OCR needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and pip3). Our summary lists: Python 3.

Does Doc OCR access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Doc OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Doc OCR use?

Doc OCR is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doc OCR use?

About 244 tokens (SKILL.md is roughly 976 characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doc OCR?

Skills that share tags, products or a category with Doc OCR: Code2media (tsingliuwin/autoclaw, 280 stars), Reportlab (jimmc414/Kosmos, 594 stars), PDF Reading (Wide-Moat/open-computer-use, 126 stars) and PDF Processing Toolkit (telagod/code-abyss, 244 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doc OCR?

davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,601 GitHub stars. The repository holds 246 skills in this directory. The repository was last updated on October 6, 2026.

Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.