Agent skill

Smart OCR

by huangruiteng in huangruiteng/CS-Notes

Extract text from images and scanned documents using PaddleOCR - supports 100+ languages

MITAuto-check passedDocuments & Office

Install Smart OCR

skills CLI
$ npx skills add huangruiteng/CS-Notes --skill smart-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huangruiteng/CS-Notes smart-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huangruiteng/CS-Notes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.trae/openclaw-skills/smart-ocr .claude/skills/smart-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
smart-ocr
GitHub stars
4k
Used in
1 other repo
Token cost
~3.1k tokens
SKILL.md length
210 words
Files
2
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Extract text from images and scanned documents using PaddleOCR - supports 100+ languages

  • Works in 3 steps: Provide the image or scanned document → Optionally specify language(s) to detect → I'll extract text with position and…
  • Tasks that involve PDF
  • SKILL.md covers Overview, How to Use, Domain Knowledge and Best Practices, plus 5 more sections
  • Calls pip

What it does

Smart OCR is an agent skill from huangruiteng/CS-Notes. Extract text from images and scanned documents using PaddleOCR - supports 100+ languages

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_meta.json`).

It sits in Documents & Office, covering PDF. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/smart-ocr”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Provide the image or scanned document
  2. Optionally specify language(s) to detect
  3. I'll extract text with position and confidence data

What it can do on your machine

Read from SKILL.md and the folder at commit f7b4e92. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Smart OCR loads about 3.1k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huangruiteng/CS-Notes at commit f7b4e92, republished under its MIT licence (© huangruiteng). 210 words, ~3,110 tokens.

Download SKILL.mdSave it as .claude/skills/smart-ocr/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
smart-ocr
description
Extract text from images and scanned documents using PaddleOCR - supports 100+ languages
author
claude-office-skills
version
1.0
tags
ocr, paddleocr, text-extraction, multilingual, image
models
claude-sonnet-4, claude-opus-4
tools
computer, code_execution, file_operations
library.name
PaddleOCR
library.url
https://github.com/PaddlePaddle/PaddleOCR
library.stars
69k

Smart OCR Skill

Overview

This skill enables intelligent text extraction from images and scanned documents using PaddleOCR - a leading OCR engine supporting 100+ languages. Extract text from photos, screenshots, scanned PDFs, and handwritten documents with high accuracy.

How to Use

  1. Provide the image or scanned document
  2. Optionally specify language(s) to detect
  3. I'll extract text with position and confidence data

Example prompts:

  • "Extract all text from this screenshot"
  • "OCR this scanned PDF document"
  • "Read the text from this business card photo"
  • "Extract Chinese and English text from this image"

Domain Knowledge

PaddleOCR Fundamentals
python
from paddleocr import PaddleOCR

# Initialize OCR engine
ocr = PaddleOCR(use_angle_cls=True, lang='en')

# Run OCR on image
result = ocr.ocr('image.png', cls=True)

# Result structure: [[box, (text, confidence)], ...]
for line in result[0]:
    box = line[0]      # [[x1,y1], [x2,y2], [x3,y3], [x4,y4]]
    text = line[1][0]  # Extracted text
    conf = line[1][1]  # Confidence score
    print(f"{text} ({conf:.2f})")
Supported Languages
python
# Common language codes
languages = {
    'en': 'English',
    'ch': 'Chinese (Simplified)',
    'cht': 'Chinese (Traditional)',
    'japan': 'Japanese',
    'korean': 'Korean',
    'french': 'French',
    'german': 'German',
    'spanish': 'Spanish',
    'russian': 'Russian',
    'arabic': 'Arabic',
    'hindi': 'Hindi',
    'vi': 'Vietnamese',
    'th': 'Thai',
    # ... 100+ languages supported
}

# Use specific language
ocr = PaddleOCR(lang='ch')  # Chinese
ocr = PaddleOCR(lang='japan')  # Japanese
ocr = PaddleOCR(lang='multilingual')  # Auto-detect
Configuration Options
python
from paddleocr import PaddleOCR

ocr = PaddleOCR(
    # Detection settings
    det_model_dir=None,         # Custom detection model
    det_limit_side_len=960,     # Max side length for detection
    det_db_thresh=0.3,          # Binarization threshold
    det_db_box_thresh=0.5,      # Box score threshold
    
    # Recognition settings
    rec_model_dir=None,         # Custom recognition model
    rec_char_dict_path=None,    # Custom character dictionary
    
    # Angle classification
    use_angle_cls=True,         # Enable angle classification
    cls_model_dir=None,         # Custom classification model
    
    # Language
    lang='en',                  # Language code
    
    # Performance
    use_gpu=True,               # Use GPU if available
    gpu_mem=500,                # GPU memory limit (MB)
    enable_mkldnn=True,         # CPU optimization
    
    # Output
    show_log=False,             # Suppress logs
)
Processing Different Sources
Image Files
python
# Single image
result = ocr.ocr('image.png')

# Multiple images
images = ['img1.png', 'img2.png', 'img3.png']
for img in images:
    result = ocr.ocr(img)
    process_result(result)
PDF Files (Scanned)
python
from pdf2image import convert_from_path

def ocr_pdf(pdf_path):
    """OCR a scanned PDF."""
    # Convert PDF pages to images
    images = convert_from_path(pdf_path)
    
    all_text = []
    for i, img in enumerate(images):
        # Save temp image
        temp_path = f'temp_page_{i}.png'
        img.save(temp_path)
        
        # OCR the image
        result = ocr.ocr(temp_path)
        
        # Extract text
        page_text = '\n'.join([line[1][0] for line in result[0]])
        all_text.append(f"--- Page {i+1} ---\n{page_text}")
        
        os.remove(temp_path)
    
    return '\n\n'.join(all_text)
URLs and Bytes
python
import requests
from io import BytesIO

# From URL
response = requests.get('https://example.com/image.png')
result = ocr.ocr(BytesIO(response.content))

# From bytes
with open('image.png', 'rb') as f:
    img_bytes = f.read()
result = ocr.ocr(BytesIO(img_bytes))
Result Processing
python
def process_ocr_result(result):
    """Process OCR result into structured data."""
    
    lines = []
    for line in result[0]:
        box = line[0]
        text = line[1][0]
        confidence = line[1][1]
        
        # Calculate bounding box
        x_coords = [p[0] for p in box]
        y_coords = [p[1] for p in box]
        
        lines.append({
            'text': text,
            'confidence': confidence,
            'bbox': {
                'left': min(x_coords),
                'top': min(y_coords),
                'right': max(x_coords),
                'bottom': max(y_coords),
            },
            'raw_box': box
        })
    
    return lines

# Sort by position (top to bottom, left to right)
def sort_by_position(lines):
    return sorted(lines, key=lambda x: (x['bbox']['top'], x['bbox']['left']))
Text Layout Reconstruction
python
def reconstruct_layout(result, line_threshold=10):
    """Reconstruct text layout from OCR results."""
    
    lines = process_ocr_result(result)
    lines = sort_by_position(lines)
    
    # Group into logical lines
    text_lines = []
    current_line = []
    current_y = None
    
    for line in lines:
        y = line['bbox']['top']
        
        if current_y is None or abs(y - current_y) < line_threshold:
            current_line.append(line)
            current_y = y
        else:
            # New line
            text_lines.append(' '.join([l['text'] for l in current_line]))
            current_line = [line]
            current_y = y
    
    # Add last line
    if current_line:
        text_lines.append(' '.join([l['text'] for l in current_line]))
    
    return '\n'.join(text_lines)

Best Practices

  1. Preprocess Images: Improve quality before OCR
  2. Choose Correct Language: Specify language for better accuracy
  3. Handle Multi-column: Process columns separately
  4. Filter Low Confidence: Skip results below threshold
  5. Batch Processing: Process multiple images efficiently

Common Patterns

Image Preprocessing
python
from PIL import Image, ImageEnhance, ImageFilter

def preprocess_image(image_path):
    """Preprocess image for better OCR."""
    img = Image.open(image_path)
    
    # Convert to grayscale
    img = img.convert('L')
    
    # Enhance contrast
    enhancer = ImageEnhance.Contrast(img)
    img = enhancer.enhance(2.0)
    
    # Sharpen
    img = img.filter(ImageFilter.SHARPEN)
    
    # Save preprocessed
    preprocessed_path = 'preprocessed.png'
    img.save(preprocessed_path)
    
    return preprocessed_path
Batch OCR with Progress
python
from tqdm import tqdm
from concurrent.futures import ThreadPoolExecutor

def batch_ocr(image_paths, max_workers=4):
    """OCR multiple images in parallel."""
    
    results = {}
    
    def process_single(img_path):
        result = ocr.ocr(img_path)
        return img_path, result
    
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        futures = [executor.submit(process_single, p) for p in image_paths]
        
        for future in tqdm(futures, desc="Processing OCR"):
            path, result = future.result()
            results[path] = result
    
    return results

Examples

Example 1: Business Card Reader
python
from paddleocr import PaddleOCR
import re

def read_business_card(image_path):
    """Extract contact info from business card."""
    
    ocr = PaddleOCR(use_angle_cls=True, lang='en')
    result = ocr.ocr(image_path)
    
    # Extract all text
    all_text = []
    for line in result[0]:
        all_text.append(line[1][0])
    
    full_text = '\n'.join(all_text)
    
    # Parse contact info
    contact = {
        'name': None,
        'email': None,
        'phone': None,
        'company': None,
        'title': None,
        'raw_text': full_text
    }
    
    # Email pattern
    email_match = re.search(r'[\w\.-]+@[\w\.-]+\.\w+', full_text)
    if email_match:
        contact['email'] = email_match.group()
    
    # Phone pattern
    phone_match = re.search(r'[\+\d][\d\s\-\(\)]{8,}', full_text)
    if phone_match:
        contact['phone'] = phone_match.group().strip()
    
    # Name is usually the largest/first text
    if all_text:
        contact['name'] = all_text[0]
    
    return contact

card_info = read_business_card('business_card.jpg')
print(f"Name: {card_info['name']}")
print(f"Email: {card_info['email']}")
print(f"Phone: {card_info['phone']}")
Example 2: Receipt Scanner
python
from paddleocr import PaddleOCR
import re

def scan_receipt(image_path):
    """Extract items and total from receipt."""
    
    ocr = PaddleOCR(use_angle_cls=True, lang='en')
    result = ocr.ocr(image_path)
    
    lines = []
    for line in result[0]:
        text = line[1][0]
        y_pos = line[0][0][1]
        lines.append({'text': text, 'y': y_pos})
    
    # Sort by vertical position
    lines.sort(key=lambda x: x['y'])
    
    receipt = {
        'items': [],
        'subtotal': None,
        'tax': None,
        'total': None
    }
    
    for line in lines:
        text = line['text']
        
        # Look for total
        if 'total' in text.lower():
            amount = re.search(r'\$?([\d,]+\.?\d*)', text)
            if amount:
                if 'sub' in text.lower():
                    receipt['subtotal'] = float(amount.group(1).replace(',', ''))
                else:
                    receipt['total'] = float(amount.group(1).replace(',', ''))
        
        # Look for tax
        elif 'tax' in text.lower():
            amount = re.search(r'\$?([\d,]+\.?\d*)', text)
            if amount:
                receipt['tax'] = float(amount.group(1).replace(',', ''))
        
        # Look for items (line with price)
        else:
            item_match = re.search(r'(.+?)\s+\$?([\d,]+\.?\d+)$', text)
            if item_match:
                receipt['items'].append({
                    'name': item_match.group(1).strip(),
                    'price': float(item_match.group(2).replace(',', ''))
                })
    
    return receipt

receipt_data = scan_receipt('receipt.jpg')
print(f"Items: {len(receipt_data['items'])}")
print(f"Total: ${receipt_data['total']}")
Example 3: Multi-language Document
python
from paddleocr import PaddleOCR

def ocr_multilingual(image_path, languages=['en', 'ch']):
    """OCR document with multiple languages."""
    
    all_results = {}
    
    for lang in languages:
        ocr = PaddleOCR(use_angle_cls=True, lang=lang)
        result = ocr.ocr(image_path)
        
        texts = []
        for line in result[0]:
            texts.append({
                'text': line[1][0],
                'confidence': line[1][1]
            })
        
        all_results[lang] = texts
    
    # Merge results, keeping highest confidence
    merged = {}
    for lang, texts in all_results.items():
        for item in texts:
            text = item['text']
            conf = item['confidence']
            
            if text not in merged or merged[text]['confidence'] < conf:
                merged[text] = {'confidence': conf, 'language': lang}
    
    return merged

result = ocr_multilingual('bilingual_document.png')
for text, info in result.items():
    print(f"[{info['language']}] {text} ({info['confidence']:.2f})")

Limitations

  • Handwritten text accuracy varies
  • Very small text may not be detected
  • Complex backgrounds reduce accuracy
  • Rotated text needs angle classification
  • GPU recommended for best performance

Installation

bash
# CPU version
pip install paddlepaddle paddleocr

# GPU version (CUDA 11.x)
pip install paddlepaddle-gpu paddleocr

# Additional dependencies
pip install pdf2image Pillow

Resources

© huangruiteng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .trae/openclaw-skills/smart-ocr of huangruiteng/CS-Notes.

  • SKILL.md
  • _meta.json

Open the folder on GitHubat commit f7b4e92

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huangruiteng/CS-Notes, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Smart OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Smart OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Smart OCR this skillhuangruiteng/CS-Notes4k1 repos~3.1kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3151 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    315 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from huangruiteng/CS-Notes

All 39 skills in this repo
  • CLI Creator

    huangruiteng/CS-Notes

    Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script.

    4k GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Codex Thread Heartbeat

    huangruiteng/CS-Notes

    Inspect and manage guarded Codex App-native or launchd heartbeats for Codex main control threads.

    4k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Slack

    huangruiteng/CS-Notes

    A skill your agent uses when you need to control Slack from Clawdbot via the slack tool, including reacting to messages or pinning/unpinning items in Slack channels or DMs.

    4k GitHub starsUsed in 10 repos~578 tokens
    Auto-check passed
  • Codex Thread Reader

    huangruiteng/CS-Notes

    Locate and read a Codex thread by a codex thread link, thread id, or rollout path across all local CODEXHOME directories (~/.codex, ~/.codex-gpt, ...).

    4k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Research Material Scout

    huangruiteng/CS-Notes

    A skill your agent uses when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information…

    4k GitHub stars~8.3k tokensUpdated yesterday
    Auto-check passed
  • GitHub

    huangruiteng/CS-Notes

    Interact with GitHub using the gh CLI. An agent skill from huangruiteng/CS-Notes.

    4k GitHub starsUsed in 27 repos~279 tokens
    Auto-check passed

Questions about Smart OCR

What does Smart OCR do?

Extract text from images and scanned documents using PaddleOCR - supports 100+ languages. Smart OCR is an agent skill from huangruiteng/CS-Notes.

When should I use Smart OCR?

Smart OCR fits situations like: tasks that involve PDF.

How do I install Smart OCR in Claude Code?

Run `npx skills add huangruiteng/CS-Notes --skill smart-ocr -a claude-code`. Or copy the skill folder (.trae/openclaw-skills/smart-ocr in huangruiteng/CS-Notes) into .claude/skills/smart-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Smart OCR in Codex?

Run `npx skills add huangruiteng/CS-Notes --skill smart-ocr -a codex`. Or copy the skill folder (.trae/openclaw-skills/smart-ocr in huangruiteng/CS-Notes) into .agents/skills/smart-ocr in your project. Codex loads it when a task matches its description.

Can I use Smart OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huangruiteng/CS-Notes --skill smart-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/smart-ocr, .gemini/skills/smart-ocr, .github/skills/smart-ocr and .opencode/skills/smart-ocr in your project.

What does Smart OCR need to run?

Going by SKILL.md and its folder, Smart OCR needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Smart OCR access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Smart OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Smart OCR use?

Smart OCR is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Smart OCR use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Smart OCR?

Skills that share tags, products or a category with Smart OCR: Markitdown (ImCa0/just-laws, 782 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Smart OCR?

huangruiteng (a GitHub user) maintains it in huangruiteng/CS-Notes, which has 4,000 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.

Source: huangruiteng/CS-Notes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.