Word Document Reader and Writer
HKUDS/DeepTutor
Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.
Extract text content from images using Tesseract OCR via Python
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benchflow-ai/skillsbench image-ocr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .claude/skills/image-ocr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .claude/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocrType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benchflow-ai/skillsbench image-ocr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .agents/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .agents/skills/image-ocr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .agents/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benchflow-ai/skillsbench image-ocr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .cursor/skills/image-ocr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .cursor/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benchflow-ai/skillsbench.git --path tasks/jpg-ocr-stat/environment/skills/image-ocr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benchflow-ai/skillsbench image-ocr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .gemini/skills/image-ocr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .gemini/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benchflow-ai/skillsbench image-ocrInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .github/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .github/skills/image-ocr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .github/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill image-ocr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benchflow-ai/skillsbench image-ocr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/image-ocr .opencode/skills/image-ocr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "image-ocr" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/image-ocr into .opencode/skills/image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-ocr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
image-ocrExtract text content from images using Tesseract OCR via Python
Image OCR is an agent skill from benchflow-ai/skillsbench. Extract text content from images using Tesseract OCR via Python
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office. It works with Python. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Image OCR loads about 2.8k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 622 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 622 words, ~2,837 tokens.
.claude/skills/image-ocr/SKILL.md (or your agent's skills folder).This skill enables accurate text extraction from image files (JPG, PNG, etc.) using Tesseract OCR via the pytesseract Python library. It is suitable for scanned documents, screenshots, photos of text, receipts, forms, and other visual content containing text.
The following Python libraries are required:
import pytesseract
from PIL import Image
import json
import osAll extracted content must be returned as valid JSON conforming to this schema:
{
"success": true,
"filename": "example.jpg",
"extracted_text": "Full raw text extracted from the image...",
"confidence": "high|medium|low",
"metadata": {
"language_detected": "en",
"text_regions": 3,
"has_tables": false,
"has_handwriting": false
},
"warnings": [
"Text partially obscured in bottom-right corner",
"Low contrast detected in header section"
]
}success: Boolean indicating whether text extraction completedfilename: Original image filenameextracted_text: Complete text content in reading order (top-to-bottom, left-to-right)confidence: Overall OCR confidence level based on image quality and text claritymetadata.language_detected: ISO 639-1 language codemetadata.text_regions: Number of distinct text blocks identifiedmetadata.has_tables: Whether tabular data structures were detectedmetadata.has_handwriting: Whether handwritten text was detectedwarnings: Array of quality issues or potential errorsimport pytesseract
from PIL import Image
def extract_text_from_image(image_path):
"""Extract text from a single image using Tesseract OCR."""
img = Image.open(image_path)
text = pytesseract.image_to_string(img)
return text.strip()import pytesseract
from PIL import Image
def extract_with_confidence(image_path):
"""Extract text with per-word confidence scores."""
img = Image.open(image_path)
# Get detailed OCR data including confidence
data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)
words = []
confidences = []
for i, word in enumerate(data['text']):
if word.strip(): # Skip empty strings
words.append(word)
confidences.append(data['conf'][i])
# Calculate average confidence
avg_confidence = sum(c for c in confidences if c > 0) / len([c for c in confidences if c > 0]) if confidences else 0
return {
'text': ' '.join(words),
'average_confidence': avg_confidence,
'word_count': len(words)
}import pytesseract
from PIL import Image
import json
import os
def ocr_to_json(image_path):
"""Perform OCR and return results as JSON."""
filename = os.path.basename(image_path)
warnings = []
try:
img = Image.open(image_path)
# Get detailed OCR data
data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)
# Extract text preserving structure
text = pytesseract.image_to_string(img)
# Calculate confidence
confidences = [c for c in data['conf'] if c > 0]
avg_conf = sum(confidences) / len(confidences) if confidences else 0
# Determine confidence level
if avg_conf >= 80:
confidence = "high"
elif avg_conf >= 50:
confidence = "medium"
else:
confidence = "low"
warnings.append(f"Low OCR confidence: {avg_conf:.1f}%")
# Count text regions (blocks)
block_nums = set(data['block_num'])
text_regions = len([b for b in block_nums if b > 0])
result = {
"success": True,
"filename": filename,
"extracted_text": text.strip(),
"confidence": confidence,
"metadata": {
"language_detected": "en",
"text_regions": text_regions,
"has_tables": False,
"has_handwriting": False
},
"warnings": warnings
}
except Exception as e:
result = {
"success": False,
"filename": filename,
"extracted_text": "",
"confidence": "low",
"metadata": {
"language_detected": "unknown",
"text_regions": 0,
"has_tables": False,
"has_handwriting": False
},
"warnings": [f"OCR failed: {str(e)}"]
}
return result
# Usage
result = ocr_to_json("document.jpg")
print(json.dumps(result, indent=2))import pytesseract
from PIL import Image
import json
import os
from pathlib import Path
def process_image_directory(directory_path, output_file):
"""Process all images in a directory and save results."""
image_extensions = {'.jpg', '.jpeg', '.png', '.webp'}
results = []
for file_path in sorted(Path(directory_path).iterdir()):
if file_path.suffix.lower() in image_extensions:
result = ocr_to_json(str(file_path))
results.append(result)
print(f"Processed: {file_path.name}")
# Save results
with open(output_file, 'w') as f:
json.dump(results, f, indent=2)
return results# Specify language (default is English)
text = pytesseract.image_to_string(img, lang='eng')
# Multiple languages
text = pytesseract.image_to_string(img, lang='eng+fra+deu')Use --psm to control how Tesseract segments the image:
# PSM 3: Fully automatic page segmentation (default)
text = pytesseract.image_to_string(img, config='--psm 3')
# PSM 4: Assume single column of text
text = pytesseract.image_to_string(img, config='--psm 4')
# PSM 6: Assume uniform block of text
text = pytesseract.image_to_string(img, config='--psm 6')
# PSM 11: Sparse text - find as much text as possible
text = pytesseract.image_to_string(img, config='--psm 11')Common PSM values:
0: Orientation and script detection (OSD) only3: Fully automatic page segmentation (default)4: Single column of text of variable sizes6: Uniform block of text7: Single text line11: Sparse text13: Raw lineFor better OCR accuracy, preprocess images:
from PIL import Image, ImageFilter, ImageOps
def preprocess_image(image_path):
"""Preprocess image for better OCR results."""
img = Image.open(image_path)
# Convert to grayscale
img = img.convert('L')
# Increase contrast
img = ImageOps.autocontrast(img)
# Apply slight sharpening
img = img.filter(ImageFilter.SHARPEN)
return img
# Use preprocessed image for OCR
img = preprocess_image("document.jpg")
text = pytesseract.image_to_string(img)For difficult images (low contrast, faded text, dark backgrounds), try multiple preprocessing approaches:
ImageOps.invert() for dark backgrounds with light textimg.point(lambda p: 255 if p > threshold else 0) with different threshold values (e.g., 100, 128)ImageFilter.SHARPEN to improve edge clarityFor challenging images, a single OCR pass may miss text. Use multiple passes with different configurations:
Try multiple PSM modes - Different page segmentation modes work better for different layouts (e.g., --psm 6 for blocks, --psm 4 for columns, --psm 11 for sparse text)
Try multiple preprocessing variants - Run OCR on several preprocessed versions of the same image
Combine results - Aggregate text from all passes to maximize extraction coverage
def multi_pass_ocr(image_path):
"""Run OCR with multiple strategies and combine results."""
img = Image.open(image_path)
gray = ImageOps.grayscale(img)
# Generate preprocessing variants
variants = [
ImageOps.autocontrast(gray),
ImageOps.invert(ImageOps.autocontrast(gray)),
gray.filter(ImageFilter.SHARPEN),
]
# PSM modes to try
psm_modes = ['--psm 6', '--psm 4', '--psm 11']
all_text = []
for variant in variants:
for psm in psm_modes:
try:
text = pytesseract.image_to_string(variant, config=psm)
if text.strip():
all_text.append(text)
except Exception:
pass
# Combine all extracted text
return "\n".join(all_text)This approach improves extraction for receipts, faded documents, and images with varying quality.
Issue: Tesseract not found
# Verify Tesseract is installed
try:
pytesseract.get_tesseract_version()
except pytesseract.TesseractNotFoundError:
print("Tesseract is not installed or not in PATH")Issue: Poor OCR quality
Issue: Empty or garbage output
Before returning results, verify:
json.loads() to validate)success, filename, extracted_text, confidence, metadata)© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in tasks/jpg-ocr-stat/environment/skills/image-ocr of benchflow-ai/skillsbench.
Open the folder on GitHubat commit 9a1f4dd
Image OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Image OCR this skillbenchflow-ai/skillsbench | 1.8k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Word Document Reader and WriterHKUDS/DeepTutor | 41k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Google WorkspaceNousResearch/hermes-agent | 252k | 3 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences | 274 | 2 repos | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Word DOCX ToolkitTokenRhythm/opensquilla | 7.1k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Markdown to HTML ReportMegaSuperKitty/WeClaw | 370 | — | ~456 | Automated safety check: Pass | MIT |
HKUDS/DeepTutor
Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.
NousResearch/hermes-agent
Gmail, Calendar, Drive, Docs, Sheets via gws CLI or Python. An agent skill from NousResearch/hermes-agent.
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
TokenRhythm/opensquilla
Inspects, edits in place or creates Word .docx files with bundled Python scripts, keeping existing styles intact when content changes.
MegaSuperKitty/WeClaw
Drafts a report in Markdown with numbered inline citations and a references section, then renders it to a styled HTML file through a Jinja2 template on Windows.
handsomeZR-netizen/mathmodel-skill
CUMCM 国赛、MCM/ICM 美赛与电工杯数学建模竞赛的端到端协作工作流。Use when a user explicitly works on one of these modeling contests or asks to run/review a modeling-competition paper from problem selection through modeling…
benchflow-ai/skillsbench
This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
benchflow-ai/skillsbench
AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.
benchflow-ai/skillsbench
Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.
benchflow-ai/skillsbench
Build deterministic, verifiable data visualizations with D3.js (v6).
benchflow-ai/skillsbench
DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.
Works with
Categories
Extract text content from images using Tesseract OCR via Python. Image OCR is an agent skill from benchflow-ai/skillsbench.
Image OCR fits situations like: documents & Office work in your project.
Run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a claude-code`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/image-ocr in benchflow-ai/skillsbench) into .claude/skills/image-ocr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a codex`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/image-ocr in benchflow-ai/skillsbench) into .agents/skills/image-ocr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill image-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-ocr, .gemini/skills/image-ocr, .github/skills/image-ocr and .opencode/skills/image-ocr in your project.
SKILL.md names no scripts, command-line tools or credentials: Image OCR is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Image OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Image OCR: Word Document Reader and Writer (HKUDS/DeepTutor, 41k stars), Google Workspace (NousResearch/hermes-agent, 252k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars) and Word DOCX Toolkit (TokenRhythm/opensquilla, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.
Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.