Agent skill

OCR And Documents

by taracodlabs in taracodlabs/aiden

Extract text from PDFs, images, scans, Word docs (Python). An agent skill from taracodlabs/aiden.

Apache-2.0Auto-check passedDocuments & Office

Install OCR And Documents

skills CLI
$ npx skills add taracodlabs/aiden --skill ocr-and-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install taracodlabs/aiden ocr-and-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/taracodlabs/aiden.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ocr-and-documents .claude/skills/ocr-and-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ocr-and-documents
GitHub stars
851
Token cost
~934 tokens
SKILL.md length
277 words
Files
2
Skills in repo
63
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract text from PDFs, images, scans, Word docs (Python). An agent skill from taracodlabs/aiden.

  • Works in 7 steps: Extract text from a PDF (pymupdf —… → Extract text from a PDF (pdf-parse via… → OCR a scanned image (Tesseract) → …
  • Tasks that involve Word documents
  • SKILL.md covers When to Use, How to Use, Examples and Cautions
  • Calls winget

What it does

OCR And Documents is an agent skill from taracodlabs/aiden. Extract text from PDFs, images, scans, Word docs (Python)

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

It sits in Documents & Office, covering Word documents and PDF. It works with Python and Microsoft Word. The repository describes itself as: Aiden — an autonomous AI agent and work engine built solo. It can operate your browser, terminal, files, apps, APIs, skills and tools, remember context, recover from failures… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Word documents
  • Tasks that involve PDF

Example prompts

  • “/ocr-and-documents”

Requirements

  • Python 3
  • Node.js

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Extract text from a PDF (pymupdf — fastest)
  2. Extract text from a PDF (pdf-parse via Node.js)
  3. OCR a scanned image (Tesseract)
  4. OCR with preprocessing for better accuracy
  5. Extract text from a Word .docx file
  6. Extract a specific page range from a PDF
  7. Extract tables from a PDF

What it can do on your machine

Read from SKILL.md and the folder at commit 3704204. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • winget

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OCR And Documents loads about 934 tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 277 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~934

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from taracodlabs/aiden at commit 3704204, republished under its Apache-2.0 licence (© taracodlabs). 277 words, ~934 tokens.

Download SKILL.mdSave it as .claude/skills/ocr-and-documents/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ocr-and-documents
description
Extract text from PDFs, images, scans, Word docs (Python)
category
productivity
version
1.0.0
origin
aiden
license
Apache-2.0
tags
ocr, pdf, image, text-extraction, documents, docx, scan, pymupdf, tesseract, pdf-parse

OCR and Document Text Extraction

Extract readable text from PDFs, scanned images, and Word documents using Python libraries available in most environments. No cloud API required.

When to Use

  • User wants to read text from a PDF file
  • User wants to extract text from a scanned image or photo of a document
  • User wants to read a .docx Word document programmatically
  • User wants to convert a multi-page document to plain text for analysis
  • User wants to extract specific pages or sections from a PDF

How to Use

1. Extract text from a PDF (pymupdf — fastest)
python
import fitz  # pip install pymupdf

doc  = fitz.open("document.pdf")
text = "\n\n".join(page.get_text() for page in doc)
print(text[:2000])  # preview first 2000 chars
doc.close()
2. Extract text from a PDF (pdf-parse via Node.js)
javascript
// requires: npm install pdf-parse (already in DevOS dependencies)
const pdfParse = require('pdf-parse')
const fs       = require('fs')
const data     = await pdfParse(fs.readFileSync('document.pdf'))
console.log(data.text.slice(0, 2000))
console.log(`Pages: ${data.numpages}`)
3. OCR a scanned image (Tesseract)

Requires Tesseract installed: winget install UB-Mannheim.TesseractOCR

python
import pytesseract          # pip install pytesseract
from PIL import Image       # pip install Pillow

img  = Image.open("scan.png")
text = pytesseract.image_to_string(img, lang="eng")
print(text)
4. OCR with preprocessing for better accuracy
python
import pytesseract
from PIL import Image, ImageFilter, ImageOps

img = Image.open("scan.jpg")
img = ImageOps.grayscale(img)
img = img.filter(ImageFilter.SHARPEN)
img = img.point(lambda p: 255 if p > 128 else 0)  # binarize
text = pytesseract.image_to_string(img, config="--psm 6")
print(text)
5. Extract text from a Word .docx file
python
from docx import Document   # pip install python-docx

doc   = Document("report.docx")
paras = [p.text for p in doc.paragraphs if p.text.strip()]
text  = "\n".join(paras)
print(text)
6. Extract a specific page range from a PDF
python
import fitz

doc    = fitz.open("big_report.pdf")
pages  = range(4, 9)   # pages 5-9 (0-indexed)
text   = "\n\n".join(doc[i].get_text() for i in pages)
print(text)
7. Extract tables from a PDF
python
import pdfplumber   # pip install pdfplumber

with pdfplumber.open("financial_report.pdf") as pdf:
  for page in pdf.pages:
    for table in page.extract_tables():
      for row in table:
        print("\t".join(str(cell or "") for cell in row))

Examples

"Read the text from this PDF contract" → Use step 1 (pymupdf) or step 2 (pdf-parse) depending on whether Python or Node is preferred.

"Extract the table from page 3 of this quarterly report PDF" → Use step 7 (pdfplumber) targeting pdf.pages[2] for page 3.

"Read the text from this scanned invoice image" → Use step 3 or 4 (Tesseract). For low-quality scans, use step 4 with preprocessing.

Cautions

  • Scanned PDFs (image-only) have no embedded text — Tesseract OCR is required
  • Tesseract accuracy drops on handwriting, decorative fonts, or low-resolution images (< 150 DPI)
  • pymupdf (fitz) extracts only programmatically embedded text — it won't OCR scanned pages
  • Large PDFs can use significant memory — process page by page for files > 100 MB
  • For non-English text, specify the language code in Tesseract: lang="hin" for Hindi, "deu" for German

© taracodlabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ocr-and-documents of taracodlabs/aiden.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit 3704204

Compare with similar skills

OCR And Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OCR And Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OCR And Documents this skilltaracodlabs/aiden851—~934Automated safety check: PassApache-2.0
Software Certificate SkillIvanCodesDev/software-certificate-skill156—~1.6kAutomated safety check: PassMIT
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Exam IngestZeKaiNie/universal-examprep-skill303—~5.6kAutomated safety check: PassMIT
Writeraiskillstore/marketplace4303 repos~1.3kAutomated safety check: PassNone

Similar skills

  • Software Certificate Skill

    IvanCodesDev/software-certificate-skill

    面向普通用户,从真实软件项目全自动生成中国软件著作权申请资料:一次收集登记事实,自动分析业务、选择可追溯源码、取得真实界面证据,生成申请表信息、规范黑白灰操作手册、代码前后30页或全部材料及真实 DOCX/PDF;内部验证、渲染、哈希与备份只进入系统临时运行区,项目最终仅保留正式资料。适配 Codex、Claude…

    156 GitHub stars~1.6k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed
  • Exam Ingest

    ZeKaiNie/universal-examprep-skill

    从学生上传的课件/大纲/老师勾的重点/真题,一键初始化并验证备考工作区:解析 PDF、DOCX、PPTX、 XLSX、常见独立图片与 txt/md,建立分章节 LLM Wiki、标准题库、结构化接管队列与进度状态;仅在 Python 确实无法运行时 明确降级为手动写盘。当工作区尚未建立、资料发生变化、或建库 readiness 被阻断时使用。

    303 GitHub stars~5.6k tokensUpdated 10 days ago
    Documents & OfficeAuto-check passed
  • Writer

    aiskillstore/marketplace

    Document creation, format conversion (ODT/DOCX/PDF), mail merge, and automation with LibreOffice Writer.

    430 GitHub starsUsed in 3 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed

More from taracodlabs/aiden

All 63 skills in this repo
  • Google Flights Search

    taracodlabs/aiden

    Searches Google Flights for prices, schedules and availability through browser automation with URL-based queries, and stops short of booking.

    851 GitHub stars~1.5k tokensUpdated 25 days ago
    Auto-check passed
  • OpenAI Codex CLI Bridge

    taracodlabs/aiden

    Delegates code generation, editing and explanation tasks to the OpenAI Codex CLI, with commands for interactive, auto-edit, question-only and model-specific runs.

    851 GitHub starsUsed in 1 repo~826 tokens
    Auto-check passed
  • Google Hotels Search

    taracodlabs/aiden

    Searches Google Hotels through the agent-browser tool for prices, ratings, amenities and availability, building a search URL from the location and dates and reporting a results table.

    851 GitHub stars~1.5k tokensUpdated 25 days ago
    Auto-check passed
  • Generates dark-themed architecture, component, data-flow and network diagrams as self-contained HTML and SVG files that open in any browser.

    851 GitHub stars~1.3k tokensUpdated 25 days ago
    Auto-check passed
  • Aggregates holdings across Zerodha, Upstox and Angel One and normalizes order parameters into one format, with confirmation required before any routing.

    851 GitHub stars~1.1k tokensUpdated 25 days ago
    Auto-check passed
  • arXiv Paper Search

    taracodlabs/aiden

    Searches arXiv by keyword, category, author or paper ID through its public API and downloads PDFs, with no API key needed.

    851 GitHub stars~1k tokensUpdated 25 days ago
    Auto-check passed

Questions about OCR And Documents

What does OCR And Documents do?

Extract text from PDFs, images, scans, Word docs (Python). An agent skill from taracodlabs/aiden. OCR And Documents is an agent skill from taracodlabs/aiden.

When should I use OCR And Documents?

OCR And Documents fits situations like: tasks that involve Word documents; tasks that involve PDF.

How do I install OCR And Documents in Claude Code?

Run `npx skills add taracodlabs/aiden --skill ocr-and-documents -a claude-code`. Or copy the skill folder (skills/ocr-and-documents in taracodlabs/aiden) into .claude/skills/ocr-and-documents in your project. Claude Code loads it when a task matches its description.

How do I install OCR And Documents in Codex?

Run `npx skills add taracodlabs/aiden --skill ocr-and-documents -a codex`. Or copy the skill folder (skills/ocr-and-documents in taracodlabs/aiden) into .agents/skills/ocr-and-documents in your project. Codex loads it when a task matches its description.

Can I use OCR And Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add taracodlabs/aiden --skill ocr-and-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocr-and-documents, .gemini/skills/ocr-and-documents, .github/skills/ocr-and-documents and .opencode/skills/ocr-and-documents in your project.

What does OCR And Documents need to run?

Going by SKILL.md and its folder, OCR And Documents needs the command-line tools its instructions call (winget). Our summary lists: Python 3; Node.js.

Does OCR And Documents access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is OCR And Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OCR And Documents use?

OCR And Documents is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OCR And Documents use?

About 934 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OCR And Documents?

Skills that share tags, products or a category with OCR And Documents: Software Certificate Skill (IvanCodesDev/software-certificate-skill, 156 stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Exam Ingest (ZeKaiNie/universal-examprep-skill, 303 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OCR And Documents?

taracodlabs (a GitHub organization) maintains it in taracodlabs/aiden, which has 851 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on September 13, 2026.

Source: taracodlabs/aiden on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.