Agent skill

OCR Parser Optimizer

by lets-mica in lets-mica/mica-ppocr

Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes.

Apache-2.0Auto-check passedDocuments & Office

Install OCR Parser Optimizer

skills CLI
$ npx skills add lets-mica/mica-ppocr --skill ocr-parser-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lets-mica/mica-ppocr ocr-parser-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lets-mica/mica-ppocr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ocr-parser-optimizer .claude/skills/ocr-parser-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ocr-parser-optimizer
GitHub stars
110
Token cost
~3.2k tokens
SKILL.md length
1,250 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes.

  • Works in 8 steps: Prepare → Baseline → Diagnose (Failure Pattern taxonomy) → …
  • Provides image/PDF batches to improve a parsers accuracy
  • SKILL.md covers When to invoke, Inputs the agent should ask…, Workflow (the optimization loop) and Code references (canonical), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

OCR Parser Optimizer is an agent skill from lets-mica/mica-ppocr. Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes. Invoke when user provides image/PDF batches to improve a parser's accuracy or fix OCR-corrupted labels.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. It works with Java. The repository describes itself as: PP-OCRv6 纯 Java 图片 OCR(ONNX Runtime,零 PaddlePaddle 依赖,bit-exact 对位 Python,Spring Boot Starter,结构化识别行驶证/身份证/银行卡/驾驶证/营业执照等). The licence is Apache-2.0.

When your agent uses it

  • Provides image/PDF batches to improve a parsers accuracy
  • Fix OCR-corrupted labels

Example prompts

  • “Use the ocr-parser-optimizer skill to optimiz mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes”
  • “/ocr-parser-optimizer”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Prepare
  2. Baseline
  3. Diagnose (Failure Pattern taxonomy)
  4. Propose a fix
  5. Apply
  6. Re-verify
  7. Iterate
  8. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 17f7bd0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are java).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OCR Parser Optimizer loads about 3.2k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,250 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lets-mica/mica-ppocr at commit 17f7bd0, republished under its Apache-2.0 licence (© lets-mica). 1,250 words, ~3,193 tokens.

Download SKILL.mdSave it as .claude/skills/ocr-parser-optimizer/SKILL.md (or your agent's skills folder).
name
ocr-parser-optimizer
description
Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes. Invoke when user provides image/PDF batches to improve a parser's accuracy or fix OCR-corrupted labels.

OCR Parser Optimizer

Iterative, batch-driven optimization loop for the 10 built-in mica-ppocr-structured parsers. The skill is not a metric calculator (use ocr-accuracy-checker for that) and not a parser authoring kit (use mica-ppocr-custom-parser for new types). It sits between the two: given a real batch of samples, it diagnoses why a field-level check fails, proposes a minimal, localized fix, and re-verifies until the target is met.

When to invoke

  • User provides an image directory or PDF directory and asks to "调优 / 优化 / 修一下 / 看哪些样本解析不对" a specific parser (e.g. 行驶证 / 身份证 / 增值税发票 / 出租车票).
  • User reports a regression ("升级后 X 字段错了") and wants to reproduce + diagnose on a batch.
  • User wants to extend a parser's tolerance to a new document variant (e.g. 横向打印的营业执照, 异形发票) by feeding a few real samples.

Do not invoke for: one-off single-image debugging (use BaseTest directly), pure metric calculation on labeled data (ocr-accuracy-checker), or building a brand-new parser from scratch (mica-ppocr-custom-parser).

Inputs the agent should ask for up front

SlotRequiredDefault
parser keyyesone of idcard vehicle driver bankcard business invoice train taxi household pdd
input kindyesimage-dir (png/jpg/jpeg/bmp/webp) or pdf-dir (recursively walked)
input pathyesabsolute path to the directory
ground truthnopath to a *.json / *.jsonl map {filename: {field: value}}; if absent, run in diagnosis-only mode (visual diff only)
model tiernotiny (default) / small / medium
target fieldnoif set, optimize only this field; otherwise optimize all fields

Always confirm these slots in one round with AskUserQuestion before doing any work.

Workflow (the optimization loop)

1.  prepare   --  collect input, ground truth, model tier
2.  baseline  --  run BatchOcrMain-style batch, compute per-field pass-rate
3.  diagnose  --  classify each failure into a Failure Pattern (see taxonomy)
4.  propose   --  emit minimal fix proposal (LabelMatcher / regex / heuristic)
5.  apply     --  edit the parser file in-place; preserve Java 8 + Lombok @Value constraints
6.  re-verify --  re-run baseline; compare; record metrics
7.  iterate   --  loop 3-6 until target hit or no fix improves metrics
8.  report    --  write optimization-reports/<parser>-<timestamp>.md
1. Prepare
  • Walk the input directory:
    • image-dir: CollUtil.listOf("png","jpg","jpeg","bmp","webp")
    • pdf-dir: recursive, every *.pdf; remind the user that PDF is a dual channel (text layer + OCR fallback) and tuning the parser only affects the OCR fallback path.
  • Copy a few samples (3-5) into test_images/<parser>/ if not already present so BaseTest<Parser, Result> single-image debug is available.
  • If ground truth provided, load it once into a Map<String, Map<String,String>>.
2. Baseline
  • For images, leverage the existing BatchOcrMain entry. Either invoke it as a subprocess (Maven exec:java) or, in agent mode, replicate its core loop using PPOcrV6Engine.run(Path) / run(byte[]) and a ParserSpec instance.
  • For PDFs, use PPOcrV6Engine.run(Path) directly (it auto-sniffs %PDF- and flattens all pages). One file may produce N per-page results; the diagnosis needs to be page-aware.
  • Persist raw outputs to optimization-reports/<parser>-<timestamp>/raw/<file>.txt so each iteration diffs cleanly.
  • If ground truth exists, compute field-level precision / recall / F1 with the ocr-accuracy-checker conventions (string equality + whitespace-trimmed compare).
3. Diagnose (Failure Pattern taxonomy)

Always classify every failure into exactly one of these buckets before proposing a fix. The mapping is the heart of this skill.

#PatternSymptomTypical cause
F1label-not-foundmatchValue returns null for a field that existsOCR corrupted the label; need findLabelBox fallback or label variants
F2label-value-row-splitvalue is on a different row / far belowwrong y-overlap threshold; multi-line label handling
F3value-truncatedregex matched only a prefix, e.g. 1\*\*\*0regex too strict; missing optional chars / spacing
F4wrong-candidate-pickedvalue is from a similar label nearby (e.g. 登记 vs 抵押)missing disambiguation; need nearest-right + y-overlap combo
F5value-merged-in-labellabel and value are in a single OCR boxuse matchValueFromPrefix instead of matchValue
F6value-absent-in-ocrthe field is genuinely missing from OCR outputdetection failure; try small/medium tier, or add a region pre-crop
F7garbled-valuevalue present but unreadable (¥ -> 羊, 1 -> l)model tier too low; switch to small / medium
F8side-mis-routede.g. idcard 正面字段跑到反面IdCardSide logic bug; usually cross-line bleed
F9field-box-wrongfield extracted correctly but getFieldBoxes() points to a nearby boxre-use of LabeledMatch.box after position math

When in doubt, render the failure sample with BaseTest.saveVis (green boxes for OCR regions; overlay the field's matched box) and look at the picture. Most F1/F4/F5 cases are obvious once visualized.

4. Propose a fix

Always emit a minimal, localized fix proposal. Match the pattern to one of these templates:

  • F1 → add the OCR-mangled form to the label variant list:

    java
    // was: matchValue(results, "号牌号码")
    // after:
    findLabelBox(results, "号牌号码", "号牌编码", "号牌号吗")  // if such helper exists

    or relax the matcher. If the parser uses a literal string, list all observed OCR variants seen in the raw output of step 2 and feed them as alternates.

  • F2 → widen the right-overlap tolerance:

    java
    matchValue(results, label, LabelMatcher.DEFAULT_RIGHT_OVERLAP_TOLERANCE) // 5 -> 8

    or add a manual y-bias when the label and value rows are visually adjacent.

  • F3 → loosen the regex: [\\d\\s-]{6,} instead of \\d{6,8}; add optional spaces / dashes that OCR inserts; allow · for * masking.

  • F4 → prefer the value box with the largest y-overlap to the label box, not the leftmost. If two labels are visually adjacent, anchor on the label's y-center and pick the value box whose y-center is closest within ±N px.

  • F5 → swap matchValue → matchValueFromPrefix.

  • F6 → not a parser bug; recommend tier bump or region pre-crop. Do not change the parser for this.

  • F7 → recommend tier bump. Do not change the parser for this.

  • F8 → re-think IdCardSide / InvoiceVersion heuristics. Common fix: tighten top-region keyword list (e.g. add 中华人民共和国 for the 户口本 cover side).

  • F9 → preserve the matched box from matchValueWithBox / matchSubstringWithBox instead of re-deriving it from text. Use the box the matcher actually returned.

Output the proposal as a unified diff against the parser file. Never rewrite the whole parser; only the failing field's match site.

Show full SKILL.md (405 more words)Show less
5. Apply
  • Edit the parser file in place. Honor project hard constraints:
    • Java 8 source level (no record, no List.of, no instanceof pattern, no String.strip). Use CollUtil.listOf / CollUtil.repeat / CollUtil.stripTrailing etc.
    • Lombok @Value + @Accessors(fluent = true) for value objects; @UtilityClass for static helpers; @Slf4j for logging.
    • Keep BaseStructuredParser SPI intact: parseResults(List<PPOcrV6Result>) -> R.
  • Add the new label variant as a constant if it's reused, or inline as a CollUtil.listOf(...) if single-use. Do not introduce new helper classes unless two fields share the same fix.
6. Re-verify
  • Re-run step 2 on the same batch and the same tier.
  • Diff raw outputs directory-by-directory: diff -ru shows exactly which samples flipped.
  • Re-compute field-level P/R/F1 if ground truth exists.
  • Only merge the fix if metrics improve and no other field regresses (regression budget = 0 unless user explicitly waives it).
7. Iterate

Stop when any of:

  • All fields ≥ target P/R/F1 (default 0.95 if user did not specify).
  • No proposed fix improves both target metric and overall accuracy.
  • Iteration count ≥ 5 (default budget; raise on user request).

Do not loop blindly; if 3 consecutive iterations have no metric gain, switch diagnosis to visualization review and ask the user for a hint.

8. Report

Write a single Markdown report to optimization-reports/<parser>-<timestamp>.md with this structure:

# Parser Optimization Report — <parser> @ <tier>
## Summary
- Baseline: P=…, R=…, F1=…
- Final:    P=…, R=…, F1=…
- Iterations: N, Net Δ: …

## Per-field metrics
| field  | baseline F1 | final F1 | Δ      | dominant pattern |
| ------ | ----------- | -------- | ------ | ---------------- |
| …      | …           | …        | …      | F1 / F3 / F6 …   |

## Applied fixes (chronological)
### Iteration 1 — F3 on `plateNo`
- Diff: `parser/vehicle/VehicleLicenseParser.java` L412-L418
- Rationale: …
- Impact: F1 0.81 → 0.93 on 12 samples

## Regressions
- (none / list)

## Remaining failures
| sample | field | pattern | next step |
| ------ | ----- | ------- | --------- |
| …      | …     | …       | tier bump |

Delete the per-iteration raw directories on completion (keep only the final baseline + final raw snapshot inside the report folder).

Code references (canonical)

  • BaseStructuredParser — SPI entry, do not break
  • LabelMatcher — matchValue* / matchSubstring* / matchPattern / geometry helpers
  • BaseTest — single-image debug + vis
  • BatchOcrMain — batch run baseline
  • ParserSpec — parser registry (key → class)

Cross-references

  • For pure metric computation (P/R/F1) on labeled data → ocr-accuracy-checker
  • For building a brand-new parser from a single sample → mica-ppocr-custom-parser
  • For running OCR on a directory without parser involvement → BatchOcrMain directly

Conventions & guardrails

  • One parser at a time. Do not touch other parsers in the same iteration even if the batch crosses types.
  • No tier-3 by default. If the fix is "switch to medium", ask the user — medium adds 130 MB and 10× latency.
  • No new dependencies. All fixes must live within mica-ppocr-structured.
  • No silent regressions. A fix that improves one field but breaks another must be reverted, not weakened.
  • Preserve bit-exactness of OCR; the parser must never alter the OCR pipeline (engine, preprocessor, postprocessor, decoder). It only consumes List<PPOcrV6Result>.
  • Java 8 source level: no Java 11+ APIs, no record, no List.of. Use CollUtil shims and Lombok @Value style.
  • Do not add Co-Authored-By: trailer to any commit the skill may produce (project memory rule).

© lets-mica, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ocr-parser-optimizer of lets-mica/mica-ppocr.

Open the folder on GitHubat commit 17f7bd0

Compare with similar skills

OCR Parser Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OCR Parser Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OCR Parser Optimizer this skilllets-mica/mica-ppocr110—~3.2kAutomated safety check: PassApache-2.0
Example Validationaspose-pdf/Aspose.PDF-for-Java144—~628Automated safety check: PassMIT
Canvas Ics33X-isdoingreat/canvas-pilot125—~8.8kAutomated safety check: NotesAGPL-3.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.9k—~19kAutomated safety check: PassApache-2.0

Similar skills

  • Example Validation

    aspose-pdf/Aspose.PDF-for-Java

    Validate Aspose.PDF Java example changes by compiling, running example runners, and checking expected output files.

    144 GitHub stars~628 tokensUpdated 28 days ago
    Documents & OfficeAuto-check passed
  • Canvas Ics33

    X-isdoingreat/canvas-pilot

    Generic code-course handler — programming assignments where the spec lives outside Canvas (instructor's external site, attached PDF, or referenced textbook), starter code is downloaded as a…

    125 GitHub stars~8.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check: notes
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed

More from lets-mica/mica-ppocr

  • Mica Ppocr Custom Parser

    lets-mica/mica-ppocr

    在 mica-ppocr 项目中新增自定义结构化解析器(证件 / 票据 / 卡证 OCR → 业务字段)时加载本 skill。

    110 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about OCR Parser Optimizer

What does OCR Parser Optimizer do?

Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes. OCR Parser Optimizer is an agent skill from lets-mica/mica-ppocr. Optimizes mica-ppocr Java parsers via batch run, failure diagnosis, and LabelMatcher/regex fixes.

When should I use OCR Parser Optimizer?

OCR Parser Optimizer fits situations like: provides image/PDF batches to improve a parsers accuracy; fix OCR-corrupted labels.

How do I install OCR Parser Optimizer in Claude Code?

Run `npx skills add lets-mica/mica-ppocr --skill ocr-parser-optimizer -a claude-code`. Or copy the skill folder (.agents/skills/ocr-parser-optimizer in lets-mica/mica-ppocr) into .claude/skills/ocr-parser-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install OCR Parser Optimizer in Codex?

Run `npx skills add lets-mica/mica-ppocr --skill ocr-parser-optimizer -a codex`. Or copy the skill folder (.agents/skills/ocr-parser-optimizer in lets-mica/mica-ppocr) into .agents/skills/ocr-parser-optimizer in your project. Codex loads it when a task matches its description.

Can I use OCR Parser Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lets-mica/mica-ppocr --skill ocr-parser-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocr-parser-optimizer, .gemini/skills/ocr-parser-optimizer, .github/skills/ocr-parser-optimizer and .opencode/skills/ocr-parser-optimizer in your project.

What does OCR Parser Optimizer need to run?

SKILL.md names no scripts, command-line tools or credentials: OCR Parser Optimizer is instructions for the agent only.

Does OCR Parser Optimizer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is OCR Parser Optimizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OCR Parser Optimizer use?

OCR Parser Optimizer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OCR Parser Optimizer use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OCR Parser Optimizer?

Skills that share tags, products or a category with OCR Parser Optimizer: Example Validation (aspose-pdf/Aspose.PDF-for-Java, 144 stars), Canvas Ics33 (X-isdoingreat/canvas-pilot, 125 stars), Markitdown (ImCa0/just-laws, 781 stars) and Gzh Design (isjiamu/gzh-design-skill, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OCR Parser Optimizer?

lets-mica (a GitHub organization) maintains it in lets-mica/mica-ppocr, which has 110 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 1, 2026.

Source: lets-mica/mica-ppocr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.