Agent skill

OCR Document Processor

by dkyazzentwatwa in dkyazzentwatwa/chatgpt-skills

Extract text and structure from scans, images, and scanned PDFs.

No licenceAuto-check passedDocuments & Office

Install OCR Document Processor

skills CLI
$ npx skills add dkyazzentwatwa/chatgpt-skills --skill ocr-document-processor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dkyazzentwatwa/chatgpt-skills ocr-document-processor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dkyazzentwatwa/chatgpt-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ocr-document-processor .claude/skills/ocr-document-processor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ocr-document-processor
GitHub stars
114
Token cost
~310 tokens
SKILL.md length
136 words
Files
7 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
None found

At a glance

Extract text and structure from scans, images, and scanned PDFs.

  • Works in 5 steps: Decide whether plain OCR, structured… → Preprocess noisy inputs before… → Use scripts/ocr_processor.py for core… → …
  • Searchable PDFs
  • SKILL.md covers Use This For, Workflow and Guardrails
  • Runs Python scripts from its folder

What it does

OCR Document Processor is an agent skill from dkyazzentwatwa/chatgpt-skills. Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.

Its SKILL.md is about 310 tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `agents/openai.yaml`, `scripts/business_card_scanner.py` and `scripts/ocr_processor.py`).

It sits in Documents & Office, covering Document parsing. The repository describes itself as: My comprehensive, tested + audited, library of skills to use for ChatGPT.

When your agent uses it

  • Searchable PDFs
  • Table extraction
  • Receipt parsing
  • Business card parsing

Example prompts

  • “/ocr-document-processor”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Decide whether plain OCR, structured extraction, or document-specific parsing is needed.
  2. Preprocess noisy inputs before extraction when skew, blur, or shadows are present.
  3. Use scripts/ocr_processor.py for core OCR tasks.
  4. Use the focused helpers when the input is specialized
  5. Return confidence caveats when the source is low quality, rotated, handwritten, or multilingual.

What it can do on your machine

Read from SKILL.md and the folder at commit 103b430. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OCR Document Processor loads about 310 tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 136 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~310

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 136 words (~310 tokens).

“Handle OCR-heavy inputs where text must be recovered from images or scanned pages.”

— opening of SKILL.md by dkyazzentwatwa
name
ocr-document-processor

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts) in ocr-document-processor of dkyazzentwatwa/chatgpt-skills.

  • SKILL.md
  • .DS_Store
  • agents/openai.yaml
  • scripts/business_card_scanner.py
  • scripts/ocr_processor.py
  • scripts/receipt_scanner.py
  • scripts/requirements.txt

Open the folder on GitHubat commit 103b430

Compare with similar skills

OCR Document Processor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OCR Document Processor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OCR Document Processor this skilldkyazzentwatwa/chatgpt-skills114—~310Automated safety check: PassNone
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
DOCX ToolkitXiaomiMiMo/MiMo-Code14k—~2.4kAutomated safety check: PassApache-2.0
Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html908—~4.8kAutomated safety check: PassMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Liteparsebastani-inc/atomic846—~1.4kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    908 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Liteparse

    bastani-inc/atomic

    A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

    846 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Beautiful Article

    ConardLi/garden-skills

    Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow.

    13k GitHub stars~4.7k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from dkyazzentwatwa/chatgpt-skills

All 12 skills in this repo
  • Skill Creator

    dkyazzentwatwa/chatgpt-skills

    Create or update agent skills with concise SKILL.md instructions, bundled resources, agent metadata, validation, and packaging.

    114 GitHub stars~444 tokensUpdated 6 mo ago
    Auto-check passed
  • Crypto Ta Analyzer

    dkyazzentwatwa/chatgpt-skills

    Run multi-indicator technical analysis on crypto or market OHLCV data.

    114 GitHub stars~213 tokensUpdated 6 mo ago
    Auto-check passed
  • Document Converter Suite

    dkyazzentwatwa/chatgpt-skills

    Convert PDFs, Office docs, markdown, HTML, and tables between editable formats.

    114 GitHub stars~386 tokensUpdated 6 mo ago
    Auto-check passed
  • MCP Builder

    dkyazzentwatwa/chatgpt-skills

    Plan and build MCP servers with agent-friendly tools, schemas, error handling, and evaluation.

    114 GitHub stars~376 tokensUpdated 6 mo ago
    Auto-check passed
  • SVG Precision Skill

    dkyazzentwatwa/chatgpt-skills

    Generate deterministic SVGs from structured specs with validation and rendering.

    114 GitHub stars~223 tokensUpdated 6 mo ago
    Auto-check passed
  • Image Enhancement Suite

    dkyazzentwatwa/chatgpt-skills

    Process images for cleanup, conversion, metadata, comparison, icons, palettes, collages, and sprite sheets.

    114 GitHub stars~332 tokensUpdated 6 mo ago
    Auto-check passed

Questions about OCR Document Processor

What does OCR Document Processor do?

Extract text and structure from scans, images, and scanned PDFs. OCR Document Processor is an agent skill from dkyazzentwatwa/chatgpt-skills. Extract text and structure from scans, images, and scanned PDFs.

When should I use OCR Document Processor?

OCR Document Processor fits situations like: searchable PDFs; table extraction; receipt parsing; business card parsing.

How do I install OCR Document Processor in Claude Code?

Run `npx skills add dkyazzentwatwa/chatgpt-skills --skill ocr-document-processor -a claude-code`. Or copy the skill folder (ocr-document-processor in dkyazzentwatwa/chatgpt-skills) into .claude/skills/ocr-document-processor in your project. Claude Code loads it when a task matches its description.

How do I install OCR Document Processor in Codex?

Run `npx skills add dkyazzentwatwa/chatgpt-skills --skill ocr-document-processor -a codex`. Or copy the skill folder (ocr-document-processor in dkyazzentwatwa/chatgpt-skills) into .agents/skills/ocr-document-processor in your project. Codex loads it when a task matches its description.

Can I use OCR Document Processor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dkyazzentwatwa/chatgpt-skills --skill ocr-document-processor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocr-document-processor, .gemini/skills/ocr-document-processor, .github/skills/ocr-document-processor and .opencode/skills/ocr-document-processor in your project.

What does OCR Document Processor need to run?

Going by SKILL.md and its folder, OCR Document Processor needs Python for the scripts in its folder. Our summary lists: Python 3.

Does OCR Document Processor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is OCR Document Processor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does OCR Document Processor use?

No licence was found for OCR Document Processor or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does OCR Document Processor use?

About 310 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OCR Document Processor?

Skills that share tags, products or a category with OCR Document Processor: Markitdown (ImCa0/just-laws, 781 stars), DOCX Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 908 stars) and Markitdown (jimmc414/Kosmos, 594 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OCR Document Processor?

dkyazzentwatwa (a GitHub user) maintains it in dkyazzentwatwa/chatgpt-skills, which has 114 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on April 8, 2026.

Source: dkyazzentwatwa/chatgpt-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.