Agent skill

Image OCR

by benchflow-ai in benchflow-ai/skillsbench

Extract text content from images using Tesseract OCR via Python

Apache-2.0Auto-check passedDocuments & Office

Install Image OCR

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench image-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .claude/skills/image-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-ocr
GitHub stars
1.8k
Token cost
~2.8k tokens
SKILL.md length
622 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract text content from images using Tesseract OCR via Python

  • Works in 5 steps: Grayscale + Autocontrast - Basic… → Inverted - Use ImageOps.invert() for… → Scaling - Upscale small images (e.g.,… → …
  • Documents & Office work in your project
  • SKILL.md covers Purpose, When to Use, Required Libraries and Input Requirements, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Image OCR is an agent skill from benchflow-ai/skillsbench. Extract text content from images using Tesseract OCR via Python

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office. It works with Python. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/image-ocr”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Grayscale + Autocontrast - Basic enhancement for most images
  2. Inverted - Use ImageOps.invert() for dark backgrounds with light text
  3. Scaling - Upscale small images (e.g., 2x) before OCR to improve character recognition
  4. Thresholding - Convert to binary using img.point(lambda p: 255 if p > threshold else 0) with different threshold values (e.g., 100, 128)
  5. Sharpening - Apply ImageFilter.SHARPEN to improve edge clarity

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image OCR loads about 2.8k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 622 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 622 words, ~2,837 tokens.

Download SKILL.mdSave it as .claude/skills/image-ocr/SKILL.md (or your agent's skills folder).
name
image-ocr
description
Extract text content from images using Tesseract OCR via Python

Image OCR Skill

Purpose

This skill enables accurate text extraction from image files (JPG, PNG, etc.) using Tesseract OCR via the pytesseract Python library. It is suitable for scanned documents, screenshots, photos of text, receipts, forms, and other visual content containing text.

When to Use

  • Extracting text from scanned documents or photos
  • Reading text from screenshots or image captures
  • Processing batch image files that contain textual information
  • Converting visual documents to machine-readable text
  • Extracting structured data from forms, receipts, or tables in images

Required Libraries

The following Python libraries are required:

python
import pytesseract
from PIL import Image
import json
import os

Input Requirements

  • File formats: JPG, JPEG, PNG, WEBP
  • Image quality: Minimum 300 DPI recommended for printed text; clear and legible text
  • File size: Under 5MB per image (resize if necessary)
  • Text language: Specify if non-English to improve accuracy

Output Schema

All extracted content must be returned as valid JSON conforming to this schema:

json
{
  "success": true,
  "filename": "example.jpg",
  "extracted_text": "Full raw text extracted from the image...",
  "confidence": "high|medium|low",
  "metadata": {
    "language_detected": "en",
    "text_regions": 3,
    "has_tables": false,
    "has_handwriting": false
  },
  "warnings": [
    "Text partially obscured in bottom-right corner",
    "Low contrast detected in header section"
  ]
}
Field Descriptions
  • success: Boolean indicating whether text extraction completed
  • filename: Original image filename
  • extracted_text: Complete text content in reading order (top-to-bottom, left-to-right)
  • confidence: Overall OCR confidence level based on image quality and text clarity
  • metadata.language_detected: ISO 639-1 language code
  • metadata.text_regions: Number of distinct text blocks identified
  • metadata.has_tables: Whether tabular data structures were detected
  • metadata.has_handwriting: Whether handwritten text was detected
  • warnings: Array of quality issues or potential errors

Code Examples

Basic OCR Extraction
python
import pytesseract
from PIL import Image

def extract_text_from_image(image_path):
    """Extract text from a single image using Tesseract OCR."""
    img = Image.open(image_path)
    text = pytesseract.image_to_string(img)
    return text.strip()
OCR with Confidence Data
python
import pytesseract
from PIL import Image

def extract_with_confidence(image_path):
    """Extract text with per-word confidence scores."""
    img = Image.open(image_path)

    # Get detailed OCR data including confidence
    data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

    words = []
    confidences = []

    for i, word in enumerate(data['text']):
        if word.strip():  # Skip empty strings
            words.append(word)
            confidences.append(data['conf'][i])

    # Calculate average confidence
    avg_confidence = sum(c for c in confidences if c > 0) / len([c for c in confidences if c > 0]) if confidences else 0

    return {
        'text': ' '.join(words),
        'average_confidence': avg_confidence,
        'word_count': len(words)
    }
Full OCR with JSON Output
python
import pytesseract
from PIL import Image
import json
import os

def ocr_to_json(image_path):
    """Perform OCR and return results as JSON."""
    filename = os.path.basename(image_path)
    warnings = []

    try:
        img = Image.open(image_path)

        # Get detailed OCR data
        data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

        # Extract text preserving structure
        text = pytesseract.image_to_string(img)

        # Calculate confidence
        confidences = [c for c in data['conf'] if c > 0]
        avg_conf = sum(confidences) / len(confidences) if confidences else 0

        # Determine confidence level
        if avg_conf >= 80:
            confidence = "high"
        elif avg_conf >= 50:
            confidence = "medium"
        else:
            confidence = "low"
            warnings.append(f"Low OCR confidence: {avg_conf:.1f}%")

        # Count text regions (blocks)
        block_nums = set(data['block_num'])
        text_regions = len([b for b in block_nums if b > 0])

        result = {
            "success": True,
            "filename": filename,
            "extracted_text": text.strip(),
            "confidence": confidence,
            "metadata": {
                "language_detected": "en",
                "text_regions": text_regions,
                "has_tables": False,
                "has_handwriting": False
            },
            "warnings": warnings
        }

    except Exception as e:
        result = {
            "success": False,
            "filename": filename,
            "extracted_text": "",
            "confidence": "low",
            "metadata": {
                "language_detected": "unknown",
                "text_regions": 0,
                "has_tables": False,
                "has_handwriting": False
            },
            "warnings": [f"OCR failed: {str(e)}"]
        }

    return result

# Usage
result = ocr_to_json("document.jpg")
print(json.dumps(result, indent=2))
Batch Processing Multiple Images
python
import pytesseract
from PIL import Image
import json
import os
from pathlib import Path

def process_image_directory(directory_path, output_file):
    """Process all images in a directory and save results."""
    image_extensions = {'.jpg', '.jpeg', '.png', '.webp'}
    results = []

    for file_path in sorted(Path(directory_path).iterdir()):
        if file_path.suffix.lower() in image_extensions:
            result = ocr_to_json(str(file_path))
            results.append(result)
            print(f"Processed: {file_path.name}")

    # Save results
    with open(output_file, 'w') as f:
        json.dump(results, f, indent=2)

    return results

Tesseract Configuration Options

Language Selection
python
# Specify language (default is English)
text = pytesseract.image_to_string(img, lang='eng')

# Multiple languages
text = pytesseract.image_to_string(img, lang='eng+fra+deu')
Page Segmentation Modes (PSM)

Use --psm to control how Tesseract segments the image:

python
# PSM 3: Fully automatic page segmentation (default)
text = pytesseract.image_to_string(img, config='--psm 3')

# PSM 4: Assume single column of text
text = pytesseract.image_to_string(img, config='--psm 4')

# PSM 6: Assume uniform block of text
text = pytesseract.image_to_string(img, config='--psm 6')

# PSM 11: Sparse text - find as much text as possible
text = pytesseract.image_to_string(img, config='--psm 11')

Common PSM values:

  • 0: Orientation and script detection (OSD) only
  • 3: Fully automatic page segmentation (default)
  • 4: Single column of text of variable sizes
  • 6: Uniform block of text
  • 7: Single text line
  • 11: Sparse text
  • 13: Raw line

Image Preprocessing

For better OCR accuracy, preprocess images:

python
from PIL import Image, ImageFilter, ImageOps

def preprocess_image(image_path):
    """Preprocess image for better OCR results."""
    img = Image.open(image_path)

    # Convert to grayscale
    img = img.convert('L')

    # Increase contrast
    img = ImageOps.autocontrast(img)

    # Apply slight sharpening
    img = img.filter(ImageFilter.SHARPEN)

    return img

# Use preprocessed image for OCR
img = preprocess_image("document.jpg")
text = pytesseract.image_to_string(img)
Advanced Preprocessing Strategies

For difficult images (low contrast, faded text, dark backgrounds), try multiple preprocessing approaches:

  1. Grayscale + Autocontrast - Basic enhancement for most images
  2. Inverted - Use ImageOps.invert() for dark backgrounds with light text
  3. Scaling - Upscale small images (e.g., 2x) before OCR to improve character recognition
  4. Thresholding - Convert to binary using img.point(lambda p: 255 if p > threshold else 0) with different threshold values (e.g., 100, 128)
  5. Sharpening - Apply ImageFilter.SHARPEN to improve edge clarity
Show full SKILL.md (252 more words)Show less

Multi-Pass OCR Strategy

For challenging images, a single OCR pass may miss text. Use multiple passes with different configurations:

  1. Try multiple PSM modes - Different page segmentation modes work better for different layouts (e.g., --psm 6 for blocks, --psm 4 for columns, --psm 11 for sparse text)

  2. Try multiple preprocessing variants - Run OCR on several preprocessed versions of the same image

  3. Combine results - Aggregate text from all passes to maximize extraction coverage

python
def multi_pass_ocr(image_path):
    """Run OCR with multiple strategies and combine results."""
    img = Image.open(image_path)
    gray = ImageOps.grayscale(img)

    # Generate preprocessing variants
    variants = [
        ImageOps.autocontrast(gray),
        ImageOps.invert(ImageOps.autocontrast(gray)),
        gray.filter(ImageFilter.SHARPEN),
    ]

    # PSM modes to try
    psm_modes = ['--psm 6', '--psm 4', '--psm 11']

    all_text = []
    for variant in variants:
        for psm in psm_modes:
            try:
                text = pytesseract.image_to_string(variant, config=psm)
                if text.strip():
                    all_text.append(text)
            except Exception:
                pass

    # Combine all extracted text
    return "\n".join(all_text)

This approach improves extraction for receipts, faded documents, and images with varying quality.

Error Handling

Common Issues and Solutions

Issue: Tesseract not found

python
# Verify Tesseract is installed
try:
    pytesseract.get_tesseract_version()
except pytesseract.TesseractNotFoundError:
    print("Tesseract is not installed or not in PATH")

Issue: Poor OCR quality

  • Preprocess image (grayscale, contrast, sharpen)
  • Use appropriate PSM mode for the document type
  • Ensure image resolution is sufficient (300+ DPI)

Issue: Empty or garbage output

  • Check if image contains actual text
  • Try different PSM modes
  • Verify image is not corrupted

Quality Self-Check

Before returning results, verify:

  • Output is valid JSON (use json.loads() to validate)
  • All required fields are present (success, filename, extracted_text, confidence, metadata)
  • Text preserves logical reading order
  • Confidence level reflects actual OCR quality
  • Warnings array includes all detected issues
  • Special characters are properly escaped in JSON

Limitations

  • Tesseract works best with printed text; handwriting recognition is limited
  • Accuracy decreases with decorative fonts, artistic text, or extreme stylization
  • Mathematical equations and special notation may not extract accurately
  • Redacted or watermarked text cannot be recovered
  • Severe image degradation (blur, noise, low resolution) reduces accuracy
  • Complex multi-column layouts may require custom PSM configuration

Version History

  • 1.0.0 (2026-01-13): Initial release with Tesseract/pytesseract OCR

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/jpg-ocr-stat/environment/skills/image-ocr of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Image OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image OCR this skillbenchflow-ai/skillsbench1.8k—~2.8kAutomated safety check: PassApache-2.0
Word Document Reader and WriterHKUDS/DeepTutor41k—~2.5kAutomated safety check: PassApache-2.0
Google WorkspaceNousResearch/hermes-agent252k3 repos~3.5kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Word DOCX ToolkitTokenRhythm/opensquilla7.1k—~1.7kAutomated safety check: PassApache-2.0
Markdown to HTML ReportMegaSuperKitty/WeClaw370—~456Automated safety check: PassMIT

Similar skills

  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Google Workspace

    NousResearch/hermes-agent

    Gmail, Calendar, Drive, Docs, Sheets via gws CLI or Python. An agent skill from NousResearch/hermes-agent.

    252k GitHub starsUsed in 3 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Word DOCX Toolkit

    TokenRhythm/opensquilla

    Inspects, edits in place or creates Word .docx files with bundled Python scripts, keeping existing styles intact when content changes.

    7.1k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Markdown to HTML Report

    MegaSuperKitty/WeClaw

    Drafts a report in Markdown with numbered inline citations and a references section, then renders it to a styled HTML file through a Jinja2 template on Windows.

    370 GitHub stars~456 tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Mathmodel Skill

    handsomeZR-netizen/mathmodel-skill

    CUMCM 国赛、MCM/ICM 美赛与电工杯数学建模竞赛的端到端协作工作流。Use when a user explicitly works on one of these modeling contests or asks to run/review a modeling-competition paper from problem selection through modeling…

    292 GitHub stars~2.5k tokensUpdated 12 days ago
    Documents & OfficeAuto-check passed

More from benchflow-ai/skillsbench

All 178 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Image OCR

What does Image OCR do?

Extract text content from images using Tesseract OCR via Python. Image OCR is an agent skill from benchflow-ai/skillsbench.

When should I use Image OCR?

Image OCR fits situations like: documents & Office work in your project.

How do I install Image OCR in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a claude-code`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/image-ocr in benchflow-ai/skillsbench) into .claude/skills/image-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Image OCR in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a codex`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/image-ocr in benchflow-ai/skillsbench) into .agents/skills/image-ocr in your project. Codex loads it when a task matches its description.

Can I use Image OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-ocr, .gemini/skills/image-ocr, .github/skills/image-ocr and .opencode/skills/image-ocr in your project.

What does Image OCR need to run?

SKILL.md names no scripts, command-line tools or credentials: Image OCR is instructions for the agent only. Our summary lists: Python 3.

Does Image OCR access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image OCR use?

Image OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image OCR use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image OCR?

Skills that share tags, products or a category with Image OCR: Word Document Reader and Writer (HKUDS/DeepTutor, 41k stars), Google Workspace (NousResearch/hermes-agent, 252k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars) and Word DOCX Toolkit (TokenRhythm/opensquilla, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image OCR?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.