Agent skill

Local OCR

by arcships in arcships/light-ocr

Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI.

Apache-2.0Auto-check passedDocuments & Office

Install Local OCR

skills CLI
$ npx skills add arcships/light-ocr --skill local-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install arcships/light-ocr local-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/arcships/light-ocr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/local-ocr .claude/skills/local-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
local-ocr
GitHub stars
610
Token cost
~1.4k tokens
SKILL.md length
462 words
Files
2
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI.

  • Works in 5 steps: Never fabricate OCR text. Only use the… → Cite coordinates when relevant. Box… → Use --schema-version 1 for reproducible… → …
  • Needing to read small/dense text in screenshots
  • SKILL.md covers Scenarios, Decision flow, Output schema and Exit codes, plus 1 more section
  • Calls python

What it does

Local OCR is an agent skill from arcships/light-ocr. Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI. Use when needing to read small/dense text in screenshots, receipts, labels, forms, or documents that a multimodal model may misread; when exact text plus bounding box coordinates are needed for field extraction, redaction, counting, or layout analysis; or when deterministic offline OCR is required without network or API calls.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Documents & Office. It works with Node.js, NVIDIA AI Platform and TypeScript. The repository describes itself as: Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr. The licence is Apache-2.0.

When your agent uses it

  • Needing to read small/dense text in screenshots
  • Documents that a multimodal model may misread
  • Exact text plus bounding box coordinates are needed for field extraction
  • Layout analysis

Example prompts

  • “/local-ocr”

Requirements

  • Python 3
  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Never fabricate OCR text. Only use the text field from results. If confidence < 0.5, state it.
  2. Cite coordinates when relevant. Box coordinates are in pageSpace.
  3. Use --schema-version 1 for reproducible output. Do not parse help text.
  4. Prefer detect first on large images, then recognize --region on areas of interest.
  5. Check exit codes before parsing stdout. Non-zero exit means stdout may be empty; read stderr.

What it can do on your machine

Read from SKILL.md and the folder at commit d2d558c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Local OCR loads about 1.4k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 462 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from arcships/light-ocr at commit d2d558c, republished under its Apache-2.0 licence (© arcships). 462 words, ~1,415 tokens.

Download SKILL.mdSave it as .claude/skills/local-ocr/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
local-ocr
description
Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI. Use when needing to read small/dense text in screenshots, receipts, labels, forms, or documents that a multimodal model may misread; when exact text plus bounding box coordinates are needed for field extraction, redaction, counting, or layout analysis; or when deterministic offline OCR is required without network or API calls.

local-ocr

light-ocr is a local OCR engine. It runs offline, returns text with coordinates and confidence, and follows a strict stdout/stderr contract for scripting.

Scenarios

Screenshot with small text

A user shares a screenshot and asks about specific text that is too small or dense to read visually.

bash
# Step 1: recognize the full image
light-ocr screenshot.png --format json

If the result has low confidence or missing text in a region:

bash
# Step 2: re-run on the specific region (coordinates from step 1 boxes)
light-ocr recognize screenshot.png --region 100,80,640,320 --format json
Form or receipt field extraction

Need to extract specific fields (names, amounts, dates) from a form or receipt image.

bash
# Step 1: detect where text regions are (fast, no recognition)
light-ocr detect receipt.png

# Step 2: recognize only the region containing the target field
light-ocr recognize receipt.png --region 50,200,300,80 --format json

This two-step pattern saves time on large images: detect first, then recognize only the regions of interest.

Counting text regions

Need to count how many text lines or regions exist in an image.

bash
light-ocr detect image.png | python -c "import json,sys; print(len(json.load(sys.stdin)['pages'][0]['detections']))"
Verifying multimodal model output

A multimodal model claims to read text from an image. Verify the claim against deterministic OCR.

bash
light-ocr image.png --format text

Compare the text output with the model's claim. If they differ, trust the OCR text field — do not fabricate.

Batch processing via shell

Process multiple images sequentially with JSONL output.

bash
for f in *.png; do
  light-ocr recognize "$f" --format jsonl
done

Each line is one page record. Check exit codes: a non-zero exit for one image does not stop the loop, but stdout for that image may be empty.

PDF and multi-page documents

Need to extract text from a PDF file or process multiple images as one document.

bash
# Full PDF with JSON output
light-ocr report.pdf --format json

# Page range with streaming JSONL
light-ocr report.pdf --pages 1-10 --format jsonl

# Multiple images as one document
light-ocr document scan1.png scan2.png --format text

# Check if PDF support is available
light-ocr document info

PDF support is part of @arcships/light-ocr. The matching PDFium binary is inside the npm platform package; do not install another package and do not download a renderer at runtime.

System diagnostics

Need to check hardware or execution provider status.

bash
light-ocr doctor --json

Returns system info (Node.js version, OS, CPU, memory), native runtime status, and available providers. No user content, hostname, username, path, or stable device identifier is collected.

Decision flow

Need text from an image or document?
├── Know which region? → recognize --region x,y,w,h --format json
├── Need full text only? → recognize --format text
├── Need text + coordinates? → recognize --format json
├── Only need where text is? → detect
├── Large image, unsure where text is? → detect first, then recognize --region
├── PDF? → light-ocr <file.pdf> --format json
├── Multiple images? → light-ocr document <files> --format json
├── Need system/hardware info? → doctor --json
└── Need engine info or version? → info --model-info / info --version
Show full SKILL.md (186 more words)Show less

Output schema

json
{
  "schemaVersion": 1,
  "source": { "kind": "image", "mediaType": "...", "identity": {}, "appliedTransforms": {} },
  "pages": [{
    "index": 0,
    "width": 640, "height": 480,
    "coordinateSpace": "pageSpace",
    "structure": "ocr-order",
    "lines": [{ "id": "L0", "text": "HELLO", "confidence": 0.99, "box": [4 points] }]
  }]
}
  • box is 4 points in pageSpace (top-left origin, x right, y down, post-EXIF pixels)
  • detect replaces lines with detections[] ({ id, score, box }) and sets structure: "detect"
  • --format text: recognized text only, one line per line, no coordinates
  • --format jsonl: one page record per line (for streaming/batch)

Exit codes

CodeMeaningAction
0SuccessParse stdout
64Usage errorFix command syntax
65Invalid argument (bad region, unsupported format/schema)Fix input
66Invalid imageTry different image
67Unsupported capabilityCheck info --model-info
68Model/bundle errorReinstall package
69Resource limit exceededUse smaller image or --region
70Environment/package failureCheck native addon
71Inference failureRetry or report bug
72Internal errorReport bug

Rules

  1. Never fabricate OCR text. Only use the text field from results. If confidence < 0.5, state it.
  2. Cite coordinates when relevant. Box coordinates are in pageSpace.
  3. Use --schema-version 1 for reproducible output. Do not parse help text.
  4. Prefer detect first on large images, then recognize --region on areas of interest.
  5. Check exit codes before parsing stdout. Non-zero exit means stdout may be empty; read stderr.

© arcships, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/local-ocr of arcships/light-ocr.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit d2d558c

Compare with similar skills

Local OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Local OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Local OCR this skillarcships/light-ocr610—~1.4kAutomated safety check: PassApache-2.0
Spreadsheet Managementfreema/mcp-gsheets100—~486Automated safety check: PassMIT
M365 Agents TSmicrosoft/skills3.1k6 repos~1.7kAutomated safety check: PassMIT
Kreuzberghashgraph-online/awesome-codex-plugins1.2k—~4.2kAutomated safety check: PassElastic-2.0
Node Backend Development Guidelinesdiet103/claude-code-infrastructure-showcase10k2 repos~2kAutomated safety check: PassMIT
Generate Release Notesteambit/bit18k—~2.2kAutomated safety check: PassCustom licence

Similar skills

  • Spreadsheet Management

    freema/mcp-gsheets

    This skill should be used when the user asks about Google Sheets, spreadsheets, cells, rows, columns, charts, or data tables.

    100 GitHub stars~486 tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • M365 Agents TS

    microsoft/skills

    Official

    Microsoft 365 Agents SDK for TypeScript/Node.js. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 6 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Kreuzberg

    hashgraph-online/awesome-codex-plugins

    Extract text, tables, metadata, and images from 91+ document formats (PDF, Office, images, HTML, email, archives, academic) using Kreuzberg.

    1.2k GitHub stars~4.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Node Backend Development Guidelines

    diet103/claude-code-infrastructure-showcase

    Sets layered architecture and coding rules for Node.js, Express and TypeScript microservices, covering routes, controllers, services, repositories, Prisma, Sentry and Zod.

    10k GitHub starsUsed in 2 repos~2k tokens
    Backend & APIsAuto-check passed
  • Generate comprehensive release notes for Bit from git commits and pull requests.

    18k GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Compromise NLP Library

    spencermountain/compromise

    Helps write and debug JavaScript or TypeScript that uses the compromise English NLP library for matching, entity extraction, tagging and sentence transforms.

    12k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

Questions about Local OCR

What does Local OCR do?

Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI. Local OCR is an agent skill from arcships/light-ocr. Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI.

When should I use Local OCR?

Local OCR fits situations like: needing to read small/dense text in screenshots; documents that a multimodal model may misread; exact text plus bounding box coordinates are needed for field extraction; layout analysis.

How do I install Local OCR in Claude Code?

Run `npx skills add arcships/light-ocr --skill local-ocr -a claude-code`. Or copy the skill folder (.agents/skills/local-ocr in arcships/light-ocr) into .claude/skills/local-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Local OCR in Codex?

Run `npx skills add arcships/light-ocr --skill local-ocr -a codex`. Or copy the skill folder (.agents/skills/local-ocr in arcships/light-ocr) into .agents/skills/local-ocr in your project. Codex loads it when a task matches its description.

Can I use Local OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add arcships/light-ocr --skill local-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/local-ocr, .gemini/skills/local-ocr, .github/skills/local-ocr and .opencode/skills/local-ocr in your project.

What does Local OCR need to run?

Going by SKILL.md and its folder, Local OCR needs the command-line tools its instructions call (python). Our summary lists: Python 3; Node.js.

Does Local OCR access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Local OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Local OCR use?

Local OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Local OCR use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Local OCR?

Skills that share tags, products or a category with Local OCR: Spreadsheet Management (freema/mcp-gsheets, 100 stars), M365 Agents TS (microsoft/skills, 3.1k stars), Kreuzberg (hashgraph-online/awesome-codex-plugins, 1.2k stars) and Node Backend Development Guidelines (diet103/claude-code-infrastructure-showcase, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Local OCR?

arcships (a GitHub organization) maintains it in arcships/light-ocr, which has 610 GitHub stars. The repository was last updated on October 7, 2026.

Source: arcships/light-ocr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.