Ag2 Multimodal Input
ag2ai/build-with-ag2
Send images, audio, video, or documents into an AG2 beta Agent alongside text.
High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .claude/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .claude/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedbackType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .agents/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .agents/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .agents/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .cursor/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .cursor/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tokenbender/agent-guides.git --path claude-skills/pdf-ocr-feedback--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .gemini/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .gemini/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedbackInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .github/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .github/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .github/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tokenbender/agent-guides pdf-ocr-feedback --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/claude-skills/pdf-ocr-feedback .opencode/skills/pdf-ocr-feedback && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-ocr-feedback" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/pdf-ocr-feedback into .opencode/skills/pdf-ocr-feedback/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-ocr-feedback", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-ocr-feedbackHigh-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.
PDF OCR Feedback is an agent skill from tokenbender/agent-guides. High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering Transcription and PDF. The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF OCR Feedback loads about 1.7k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 823 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 823 words, ~1,724 tokens.
.claude/skills/pdf-ocr-feedback/SKILL.md (or your agent's skills folder).Use this skill when transcribing PDF pages through a vision model and a single OCR pass is not reliable enough.
Produce page-faithful OCR with exact page boundaries, explicit uncertainty, and a practical target of at least 95/100 quality whenever the source allows it.
Escalate to this workflow when any of the following are true:
For every page:
===== PAGE N =====
<page text>[uncertain: "A" | "B"].For each page:
1. Pass-1 OCR
2. Self-evaluate on a 0-100 rubric
3. If score >= 95 and no red flags -> ACCEPT
4. Else run Maj@K escalation:
a. Generate K-1 additional independent passes
b. Vote at the smallest reliable unit
c. Re-score the merged result
d. If still weak, repair only flagged spans
5. Stop when accepted, capped, or no longer improvingFor the first pass on every page:
Switch into evaluator mode. In this phase you do not edit text; you only score, flag, and decide whether the page is accepted or escalated.
Score each page on a 0-100 scale across five dimensions:
0/O, 1/l, or rn/m.Any red flag forces escalation even if the numeric score is high:
For every scored page:
PAGE N - Score: XX/100
Structural Fidelity: XX/25 - [notes]
Completeness: XX/25 - [notes]
Character/Numeric Accuracy: XX/20 - [notes]
Layout-Sensitive Content: XX/20 - [notes]
Noise/Garbling: XX/10 - [notes]
Red Flags: [list or "none"]
Spot-Check:
1. "snippet text" - risk: [reason] - confidence: [high/medium/low]
2. ...
Decision: ACCEPT / ESCALATE (reason)Use this phase for pages scoring below 95 or pages with any red flag.
K-1 additional independent passes.K=3 by default.K=5 for hard pages: equations, dense tables, multi-column layouts, handwriting, mixed scripts, or poor scans.[uncertain: ...].Produce one merged transcription per page:
Only repair flagged spans. Do not regenerate accepted text.
Assemble the final document in original page order with exact delimiters preserved.
Append a short refinement summary:
## OCR Refinement Summary
Total pages: N
Pass-1 accepted: [pages]
Maj@K escalated: [pages]
Targeted repair needed: [pages]
Final scores: [page -> score]
Remaining uncertain spans: [count and pages]
Iterations used: [count]Stop when the first applicable rule fires:
If a page hits the cap below 95, accept it only with an explicit note about the remaining uncertain spans.
K=5.© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in claude-skills/pdf-ocr-feedback of tokenbender/agent-guides.
Open the folder on GitHubat commit a74dd9d
PDF OCR Feedback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF OCR Feedback this skilltokenbender/agent-guides | 367 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Ag2 Multimodal Inputag2ai/build-with-ag2 | 252 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| Lt2mdlibnyx/LT2MD | 109 | — | ~4.4k | Automated safety check: Pass | AGPL-3.0 | |
| Timestamped Video Summarycnfjlhj/ai-collab-playbook | 452 | — | ~608 | Automated safety check: Pass | None |
ag2ai/build-with-ag2
Send images, audio, video, or documents into an AG2 beta Agent alongside text.
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
libnyx/LT2MD
Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…
cnfjlhj/ai-collab-playbook
Generate a detailed, professional video content summary from timestamped subtitles/transcripts (e.g., lines starting with 00:00 / 1:23:45).
zai-org/GLM-skills
Trigger when: (1) User wants to extract text, tables, formulas, or structured data from images/PDFs/scanned documents, (2) User mentions "OCR", "文字识别", "文档解析", (3) User has a document (screenshot…
tokenbender/agent-guides
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.
tokenbender/agent-guides
A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.
tokenbender/agent-guides
Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.
tokenbender/agent-guides
Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.
tokenbender/agent-guides
Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.
tokenbender/agent-guides
A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…
Categories
High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription. PDF OCR Feedback is an agent skill from tokenbender/agent-guides. High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.
PDF OCR Feedback fits situations like: tasks that involve Transcription; tasks that involve PDF.
Run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a claude-code`. Or copy the skill folder (claude-skills/pdf-ocr-feedback in tokenbender/agent-guides) into .claude/skills/pdf-ocr-feedback in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a codex`. Or copy the skill folder (claude-skills/pdf-ocr-feedback in tokenbender/agent-guides) into .agents/skills/pdf-ocr-feedback in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill pdf-ocr-feedback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-ocr-feedback, .gemini/skills/pdf-ocr-feedback, .github/skills/pdf-ocr-feedback and .opencode/skills/pdf-ocr-feedback in your project.
SKILL.md names no scripts, command-line tools or credentials: PDF OCR Feedback is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
PDF OCR Feedback is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with PDF OCR Feedback: Ag2 Multimodal Input (ag2ai/build-with-ag2, 252 stars), Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars) and Lt2md (libnyx/LT2MD, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.
Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.