Agent skill

PDF OCR Feedback

by tokenbender in tokenbender/agent-guides

High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.

Apache-2.0Auto-check passedDocuments & Office

Install PDF OCR Feedback

skills CLI
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .claude/skills/pdf-ocr-feedback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-ocr-feedback
GitHub stars
367
Token cost
~1.7k tokens
SKILL.md length
823 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.

  • Works in 5 steps: Initial Transcription → Self-Evaluation → Maj@K Consensus Voting → …
  • Tasks that involve Transcription
  • SKILL.md covers Objective, When to Use, Output Contract and Pipeline Overview, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

PDF OCR Feedback is an agent skill from tokenbender/agent-guides. High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Transcription and PDF. The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve PDF

Example prompts

  • “/pdf-ocr-feedback”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Initial Transcription
  2. Self-Evaluation
  3. Maj@K Consensus Voting
  4. Targeted Span Repair
  5. Final Merge and Summary

What it can do on your machine

Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF OCR Feedback loads about 1.7k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 823 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 823 words, ~1,724 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-ocr-feedback/SKILL.md (or your agent's skills folder).
name
pdf-ocr-feedback
description
High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.

PDF OCR Feedback

Use this skill when transcribing PDF pages through a vision model and a single OCR pass is not reliable enough.

Objective

Produce page-faithful OCR with exact page boundaries, explicit uncertainty, and a practical target of at least 95/100 quality whenever the source allows it.

When to Use

Escalate to this workflow when any of the following are true:

  • equations or mathematical notation matter,
  • tables have nontrivial structure,
  • the page is multi-column,
  • the scan is noisy, low-resolution, or artifact-heavy,
  • the document mixes languages, scripts, or handwriting,
  • or a single OCR pass leaves meaningful uncertainty.

Output Contract

For every page:

  1. Preserve reading order.
  2. Capture all visible regions that matter: headers, footers, footnotes, captions, margin notes, table content, equation text, figure labels, and code blocks.
  3. Emit exact page delimiters:
text
===== PAGE N =====
<page text>
  1. Keep page order unchanged.
  2. Preserve equations, units, and table semantics.
  3. Never silently drop unknown symbols.
  4. If a tie cannot be resolved, mark the span explicitly as [uncertain: "A" | "B"].

Pipeline Overview

text
For each page:
  1. Pass-1 OCR
  2. Self-evaluate on a 0-100 rubric
  3. If score >= 95 and no red flags -> ACCEPT
  4. Else run Maj@K escalation:
     a. Generate K-1 additional independent passes
     b. Vote at the smallest reliable unit
     c. Re-score the merged result
     d. If still weak, repair only flagged spans
  5. Stop when accepted, capped, or no longer improving

Phase 1: Initial Transcription

For the first pass on every page:

  1. Transcribe the full page faithfully.
  2. Preserve top-to-bottom, left-to-right reading order. For multi-column pages, process column-by-column.
  3. Do not skip difficult regions; capture them or mark them uncertain.
  4. Keep formatting structure when it carries meaning, such as headings, lists, table rows, and code blocks.

Phase 2: Self-Evaluation

Switch into evaluator mode. In this phase you do not edit text; you only score, flag, and decide whether the page is accepted or escalated.

Scoring Rubric

Score each page on a 0-100 scale across five dimensions:

  • Structural fidelity: 0-25
  • Completeness: 0-25
  • Character and numeric accuracy: 0-20
  • Layout-sensitive content: 0-20
  • Noise and garbling: 0-10
What to Check
  • Structural fidelity: headings, paragraph breaks, list structure, column order, and section boundaries.
  • Completeness: no dropped text regions, no truncated lines, no missing footnotes or captions.
  • Character and numeric accuracy: symbols, digits, citation numbers, units, and OCR confusions such as 0/O, 1/l, or rn/m.
  • Layout-sensitive content: table cells, equation operators, superscripts, subscripts, code tokens, and figure labels.
  • Noise and garbling: repeated fragments, hallucinated text, gibberish, or broken words.
Red Flags

Any red flag forces escalation even if the numeric score is high:

  • an acknowledged unreadable region with no transcription attempt,
  • a suspected skipped column,
  • an ambiguous table grid or cell assignment,
  • an equation with uncertain structure or operators,
  • more than two unresolved uncertainty markers on the page,
  • conflicting variants that remain unresolved after voting.
Mandatory Spot-Check

For every scored page:

  1. Pick 3-5 high-risk snippets, such as equations, numbers, citations, table cells, or proper nouns.
  2. Explain why each snippet is risky.
  3. Record confidence for each snippet.
  4. Flag any snippet below 80 percent confidence for retry or repair.
Evaluation Format
text
PAGE N - Score: XX/100
  Structural Fidelity: XX/25 - [notes]
  Completeness: XX/25 - [notes]
  Character/Numeric Accuracy: XX/20 - [notes]
  Layout-Sensitive Content: XX/20 - [notes]
  Noise/Garbling: XX/10 - [notes]
  Red Flags: [list or "none"]
  Spot-Check:
    1. "snippet text" - risk: [reason] - confidence: [high/medium/low]
    2. ...
  Decision: ACCEPT / ESCALATE (reason)

Phase 3: Maj@K Consensus Voting

Use this phase for pages scoring below 95 or pages with any red flag.

Show full SKILL.md (349 more words)Show less
Additional Passes
  1. Generate K-1 additional independent passes.
  2. Use K=3 by default.
  3. Use K=5 for hard pages: equations, dense tables, multi-column layouts, handwriting, mixed scripts, or poor scans.
  4. Keep the passes independent. Do not anchor later passes to the first pass.
Voting Rules
  • Default to line-level voting.
  • For a disagreement, shrink the dispute to the smallest meaningful span.
  • Majority wins when a clear majority exists.
  • If the vote is tied, choose the most contextually consistent variant only when the evidence is strong.
  • If no clear winner exists, keep the span explicit as [uncertain: ...].
Special Cases
  • Numbers and equations: vote at character level when needed.
  • Tables: vote per cell, not per line.
  • Proper nouns and citations: cross-check against other occurrences in the document.
Consensus Output

Produce one merged transcription per page:

  • majority-agreed content passes through unchanged,
  • disputed spans are either resolved or marked uncertain,
  • remaining uncertain spans are queued for targeted repair.

Phase 4: Targeted Span Repair

Only repair flagged spans. Do not regenerate accepted text.

  1. Identify unresolved or low-confidence spans.
  2. Re-read only those regions from the source.
  3. Replace the uncertain span if the repair is clearly better.
  4. Re-score the page after repair.

Phase 5: Final Merge and Summary

Assemble the final document in original page order with exact delimiters preserved.

Append a short refinement summary:

text
## OCR Refinement Summary

Total pages: N
Pass-1 accepted: [pages]
Maj@K escalated: [pages]
Targeted repair needed: [pages]
Final scores: [page -> score]
Remaining uncertain spans: [count and pages]
Iterations used: [count]

Stopping Criteria

Stop when the first applicable rule fires:

  1. Accept: score at least 95, no red flags, and output contract satisfied.
  2. Diminishing returns: less than 2 points of improvement across consecutive rounds.
  3. Hard cap per page: 3 total iterations.
  4. Hard cap global: 5 refinement iterations across the full document.

If a page hits the cap below 95, accept it only with an explicit note about the remaining uncertain spans.

Anti-Patterns

  • Never invent text that appears in none of the OCR passes.
  • Never skip reading-order validation on multi-column pages.
  • Never score a page without filling the rubric and the spot-check.
  • Never regenerate a whole page when only a few spans are weak.
  • Never exceed K=5.
  • Never round up uncertain work into a confident score.

© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in claude-skills/pdf-ocr-feedback of tokenbender/agent-guides.

Open the folder on GitHubat commit a74dd9d

Compare with similar skills

PDF OCR Feedback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF OCR Feedback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF OCR Feedback this skilltokenbender/agent-guides367—~1.7kAutomated safety check: PassApache-2.0
Ag2 Multimodal Inputag2ai/build-with-ag2252—~1.7kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Lt2mdlibnyx/LT2MD109—~4.4kAutomated safety check: PassAGPL-3.0
Timestamped Video Summarycnfjlhj/ai-collab-playbook452—~608Automated safety check: PassNone

Similar skills

  • Ag2 Multimodal Input

    ag2ai/build-with-ag2

    Send images, audio, video, or documents into an AG2 beta Agent alongside text.

    252 GitHub stars~1.7k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Lt2md

    libnyx/LT2MD

    Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…

    109 GitHub stars~4.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Timestamped Video Summary

    cnfjlhj/ai-collab-playbook

    Generate a detailed, professional video content summary from timestamped subtitles/transcripts (e.g., lines starting with 00:00 / 1:23:45).

    452 GitHub stars~608 tokensUpdated 24 days ago
    Knowledge ManagementAuto-check passed
  • Glmocr SDK

    zai-org/GLM-skills

    Trigger when: (1) User wants to extract text, tables, formulas, or structured data from images/PDFs/scanned documents, (2) User mentions "OCR", "文字识别", "文档解析", (3) User has a document (screenshot…

    476 GitHub stars~2.7k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check: notes

More from tokenbender/agent-guides

All 11 skills in this repo
  • Capability Horizon Estimator

    tokenbender/agent-guides

    Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.

    367 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Technical Writing Workflow

    tokenbender/agent-guides

    A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.

    367 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Epistemic Libido

    tokenbender/agent-guides

    Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.

    367 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Manim Math Explainer

    tokenbender/agent-guides

    Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.

    367 GitHub stars~799 tokensUpdated 2 mo ago
    Auto-check passed
  • Worklog

    tokenbender/agent-guides

    Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.

    367 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • X Thread Reader

    tokenbender/agent-guides

    A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…

    367 GitHub stars~506 tokensUpdated 2 mo ago
    Auto-check passed

Questions about PDF OCR Feedback

What does PDF OCR Feedback do?

High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription. PDF OCR Feedback is an agent skill from tokenbender/agent-guides. High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.

When should I use PDF OCR Feedback?

PDF OCR Feedback fits situations like: tasks that involve Transcription; tasks that involve PDF.

How do I install PDF OCR Feedback in Claude Code?

Run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a claude-code`. Or copy the skill folder (claude-skills/pdf-ocr-feedback in tokenbender/agent-guides) into .claude/skills/pdf-ocr-feedback in your project. Claude Code loads it when a task matches its description.

How do I install PDF OCR Feedback in Codex?

Run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a codex`. Or copy the skill folder (claude-skills/pdf-ocr-feedback in tokenbender/agent-guides) into .agents/skills/pdf-ocr-feedback in your project. Codex loads it when a task matches its description.

Can I use PDF OCR Feedback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-ocr-feedback, .gemini/skills/pdf-ocr-feedback, .github/skills/pdf-ocr-feedback and .opencode/skills/pdf-ocr-feedback in your project.

What does PDF OCR Feedback need to run?

SKILL.md names no scripts, command-line tools or credentials: PDF OCR Feedback is instructions for the agent only.

Does PDF OCR Feedback access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF OCR Feedback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF OCR Feedback use?

PDF OCR Feedback is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF OCR Feedback use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF OCR Feedback?

Skills that share tags, products or a category with PDF OCR Feedback: Ag2 Multimodal Input (ag2ai/build-with-ag2, 252 stars), Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars) and Lt2md (libnyx/LT2MD, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF OCR Feedback?

tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.

Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.