Agent skill

Yao OCR Tools

by YaoApp in YaoApp/yao

Extracts text from images and PDFs, including invoices, receipts, ID cards and tables, with plain text, JSON or Markdown output and a choice of OCR providers.

Custom licenceAuto-check passedDocuments & Office

Install Yao OCR Tools

skills CLI
$ npx skills add YaoApp/yao --skill yao-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install YaoApp/yao yao-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/YaoApp/yao.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tools/skills/yao-ocr .claude/skills/yao-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
yao-ocr
GitHub stars
8.1k
Token cost
~1.6k tokens
SKILL.md length
508 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Custom licence

At a glance

Extracts text from images and PDFs, including invoices, receipts, ID cards and tables, with plain text, JSON or Markdown output and a choice of OCR providers.

  • Pulling structured fields out of an invoice or receipt PDF
  • SKILL.md covers ocr_recognize, ocr_providers, PDF support and Multi-page PDF response, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Converting a photographed table into Markdown

What it does

The skill wraps the ocr_recognize tool, called through the tai command line. It takes a file path or URL for an image or PDF and can use either a vision-language model or a traditional OCR service such as Baidu, Google, Azure or PaddleOCR, choosing one automatically when you do not name a provider.

Options control the recognition type (general by default, with presets such as table and invoice), the output format (text, JSON with coordinates and fields, or Markdown), accurate or standard mode, a language hint and a PDF page range. A custom prompt can steer vision-language recognition only; traditional OCR ignores it.

When your agent uses it

  • Pulling structured fields out of an invoice or receipt PDF
  • Converting a photographed table into Markdown
  • Reading text from ID cards, bank cards or business licenses
  • Recognizing handwritten documents or a page range of a scanned PDF

Example prompts

  • “Run OCR on ./scans/invoice.pdf and return the fields as JSON.”
  • “Extract the table in table.png as Markdown.”
  • “Read pages 1-5 of report.pdf and give me the text in Markdown.”
  • “Use the Baidu provider to recognize the Chinese text in doc.png.”

Requirements

  • The tai command line tool
  • An OCR provider or LLM connector set up for it

What it can do on your machine

Read from SKILL.md and the folder at commit 0b6ee5f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Yao OCR Tools loads about 1.6k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 508 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 508 words (~1,635 tokens).

“Two tools for optical character recognition, supporting both VLM-OCR (vision language models) and traditional OCR APIs (Baidu, Google, Azure, PaddleOCR).”

— opening of SKILL.md by YaoApp, Custom licence
name
yao-ocr

Read the full SKILL.md on GitHub

Files

Just SKILL.md in tools/skills/yao-ocr of YaoApp/yao.

Open the folder on GitHubat commit 0b6ee5f

Compare with similar skills

Yao OCR Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Yao OCR Tools compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Yao OCR Tools this skillYaoApp/yao8.1k—~1.6kAutomated safety check: PassCustom licence
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed

More from YaoApp/yao

All 14 skills in this repo
  • Yao Image Tools

    YaoApp/yao

    Lets an agent read and describe images through a vision model and generate or edit images from text prompts, using the tai command-line tool.

    8.1k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Lists, downloads, references and deploys agents on a Yao host through five bash-invoked tools, keeping edits to the smith namespace while allowing read-only study of any agent.

    8.1k GitHub stars~873 tokensUpdated 2 days ago
    Auto-check passed
  • Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers.

    8.1k GitHub stars~416 tokensUpdated 2 days ago
    Auto-check passed
  • Starts, lists, checks, restarts and stops long-running background services through Yao daemon tools, reporting the verified port and URL of each.

    8.1k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Starts, lists, monitors, waits on and stops long-running background commands through a dedicated set of job-management tools.

    8.1k GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Lists workspaces and reads or writes their files on a remote node through five tai tool commands, for cross-node file work the local filesystem cannot do.

    8.1k GitHub stars~939 tokensUpdated 2 days ago
    Auto-check passed

Questions about Yao OCR Tools

What does Yao OCR Tools do?

Extracts text from images and PDFs, including invoices, receipts, ID cards and tables, with plain text, JSON or Markdown output and a choice of OCR providers. The skill wraps the ocr_recognize tool, called through the tai command line. It takes a file path or URL for an image or PDF and can use either a vision-language model or a traditional OCR service such as Baidu, Google, Azure or PaddleOCR, choosing one automatically when you do not name a provider.

When should I use Yao OCR Tools?

Yao OCR Tools fits situations like: pulling structured fields out of an invoice or receipt PDF; converting a photographed table into Markdown; reading text from ID cards, bank cards or business licenses; recognizing handwritten documents or a page range of a scanned PDF.

How do I install Yao OCR Tools in Claude Code?

Run `npx skills add YaoApp/yao --skill yao-ocr -a claude-code`. Or copy the skill folder (tools/skills/yao-ocr in YaoApp/yao) into .claude/skills/yao-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Yao OCR Tools in Codex?

Run `npx skills add YaoApp/yao --skill yao-ocr -a codex`. Or copy the skill folder (tools/skills/yao-ocr in YaoApp/yao) into .agents/skills/yao-ocr in your project. Codex loads it when a task matches its description.

Can I use Yao OCR Tools in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add YaoApp/yao --skill yao-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/yao-ocr, .gemini/skills/yao-ocr, .github/skills/yao-ocr and .opencode/skills/yao-ocr in your project.

What does Yao OCR Tools need to run?

SKILL.md names no scripts, command-line tools or credentials: Yao OCR Tools is instructions for the agent only. Our summary lists: The tai command line tool; An OCR provider or LLM connector set up for it.

Does Yao OCR Tools access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Yao OCR Tools safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Yao OCR Tools use?

Yao OCR Tools has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Yao OCR Tools use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Yao OCR Tools?

Skills that share tags, products or a category with Yao OCR Tools: PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Yao OCR Tools?

YaoApp (a GitHub organization) maintains it in YaoApp/yao, which has 8,090 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: YaoApp/yao on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.