Agent skill

Paper OCR Notes Pipeline

by tokenbender in tokenbender/agent-guides

Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md.

Apache-2.0Auto-check passedDocuments & Office

Install Paper OCR Notes Pipeline

skills CLI
$ npx skills add tokenbender/agent-guides --skill paper-ocr-notes-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tokenbender/agent-guides paper-ocr-notes-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/paper-ocr-notes-pipeline .claude/skills/paper-ocr-notes-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paper-ocr-notes-pipeline
GitHub stars
367
Token cost
~1.8k tokens
SKILL.md length
888 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md.

  • Works in 8 steps: Intake and setup → Extraction and OCR feedback loop → 5 - Identity gate (mandatory) → …
  • Tasks that involve PDF
  • SKILL.md covers Objective, Inputs, Output Root Resolution and Output Contract, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Paper OCR Notes Pipeline is an agent skill from tokenbender/agent-guides. Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/paper-ocr-notes-pipeline”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Intake and setup
  2. Extraction and OCR feedback loop
  3. 5 - Identity gate (mandatory)
  4. Ground-truth extraction (facts only)
  5. Compose notes.md in strict order
  6. Exploration and consistency checks
  7. Revision passes (non-negotiable)
  8. Quality gates (must pass all)

What it can do on your machine

Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paper OCR Notes Pipeline loads about 1.8k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 888 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 888 words, ~1,809 tokens.

Download SKILL.mdSave it as .claude/skills/paper-ocr-notes-pipeline/SKILL.md (or your agent's skills folder).
name
paper-ocr-notes-pipeline
description
Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md.

Paper OCR Notes Pipeline

This skill is a reusable policy for paper ingestion and note production. It must stay self-contained and work in any workspace.

Do not require any companion README.md or local notes file to run this workflow. If one exists, treat it only as supplemental context.

Objective

Generate a self-contained notes package from a paper source while preserving provenance, factual grounding, OCR quality, and teachability.

Inputs

  • One of: a PDF path, a paper URL, or an arXiv abs URL.
  • Optional: preferred folder slug.
  • Optional: output root directory.

Output Root Resolution

Use this priority order:

  1. If the user provides an output root, use it.
  2. Else create and use ./papers-articles in the current working directory.

Output Contract

Create a dedicated folder:

text
<output-root>/<paper-slug>/

Required files:

  • source file: <paper>.pdf or link.txt for URL-only sources,
  • notes.md as the primary deliverable,
  • README.md as a paper-local artifact and provenance index,
  • <paper>.ocr.feedback-loop.txt as the final OCR output,
  • <paper>.ocr.feedback-loop-comparison.md as the OCR replacement report.

notes.md must be sufficient for a new reader to learn the paper without opening extra files.

Canonical Notes Sections

Use these headings in order:

  1. # <Paper Title>
  2. ## TL;DR
  3. ## Beginner's Guide
  4. ## Paper Summary
  5. ## Deep Dive
  6. ## Critical Analysis
  7. ## Learning Path
  8. ## Key Equations Summary
  9. ## References

Nested Skills

  • Use pdf-ocr-feedback for pagewise OCR refinement.
  • Use technical-writing-workflow while drafting and revising notes.md.

Core Principles

  1. Reader intuition comes before technical detail.
  2. Do not blur distinct concepts.
  3. Evidence-first writing: claims point to table, figure, section, or page evidence.
  4. Revision is mandatory.
  5. Identity verification is mandatory before trusting extracted content.
  6. Uncertainty and limitations must be stated explicitly.

Workflow

Follow the phases in order.

Phase 0 - Intake and setup
  1. Resolve canonical identity: title, authors, date/version, identifier (arXiv/DOI when available).
  2. Create output folder and store source.
  3. Initialize notes.md skeleton with required section order.
  4. Create paper-local README.md listing source and produced artifacts.
Phase 1 - Extraction and OCR feedback loop

Use nested skill pdf-ocr-feedback for pagewise OCR refinement.

Required OCR behavior:

  1. Pass-1 extraction across all pages.
  2. Per-page quality scoring and weak-page detection.
  3. Pass-2 retries only on weak pages.
  4. Merge pages in original order.

Required OCR format:

  • Preserve ===== PAGE N ===== delimiters.
  • Preserve page order.
  • Record replaced pages and rationale in comparison report.
Phase 1.5 - Identity gate (mandatory)

Before trusting extracted content, verify anchors:

  • title matches target,
  • author block matches reasonably,
  • identifier/date lines are consistent when available.

If any anchor mismatches:

  1. reject extraction,
  2. rerun extraction,
  3. revalidate before proceeding.
Phase 2 - Ground-truth extraction (facts only)

Collect source-backed facts only (no interpretation yet):

  • problem statement and failure mode,
  • contributions as stated by authors,
  • method overview,
  • key equations and definitions,
  • experimental setup: tasks, datasets, models, baselines, metrics, ablations/stress tests,
  • headline quantitative results,
  • limitations stated by authors.

Stop condition: answer "what did they do" without guessing.

Phase 3 - Compose notes.md in strict order
  1. TL;DR (problem, key idea, 1-2 evidence numbers, caveat)
  2. Beginner's Guide (college-level, define symbols once, one concrete toy example)
  3. Paper Summary (problem recap, contributions, results, experiments table)
  4. Deep Dive (definitions, derivation steps, implementation sketch, optional pseudocode)
  5. Critical Analysis (assumptions, objections, failure modes, validation/refutation experiments)
  6. Learning Path (prereqs, next study steps, diagnostics)
  7. Key Equations Summary (equation + one-line meaning)

While drafting, apply technical-writing-workflow:

  • anchor explanations in familiar ground,
  • define terms at point of need,
  • keep ontology and score ownership clear,
  • cut implementation-operator detail unless it matters scientifically,
  • keep claims tied to evidence.
Show full SKILL.md (314 more words)Show less
Phase 4 - Exploration and consistency checks

While reading and drafting, run these checks:

Problem framing checks:

  • what exactly breaks and under what conditions,
  • where boundary conditions are stated.

Method checks:

  • what estimator/objective is actually optimized,
  • which terms are additive vs multiplicative,
  • what comes from probabilities vs gradients vs external signals.

Claims checks:

  • theoretical vs empirical claims separation,
  • strongest evidence source for each major claim,
  • missing ablations/OOD/calibration/compute analyses.

Consistency checks:

  • terminology consistency across sections,
  • equation-to-implementation narrative alignment,
  • baseline independence assumptions where claimed,
  • proxy computability with stated costs.
Phase 5 - Revision passes (non-negotiable)
  1. Clarity pass:
    • remove duplicates,
    • ensure definitions appear before equations,
    • soften overconfident wording where uncertainty exists.
  2. Correctness pass:
    • detect concept conflations,
    • remove unsupported claims,
    • fix misleading pseudocode assumptions.
  3. Objection stress test:
    • state strongest objection,
    • provide rebuttal conditions,
    • propose falsifiable tests.
  4. Update TL;DR caveat after all revisions.
Phase 6 - Quality gates (must pass all)
  • Accuracy: symbols and claims are correct and evidence-backed.
  • Completeness: experiments table includes tasks, baselines, metrics.
  • Teachability: beginner section is understandable by a college-level reader.
  • Critical rigor: at least one strong objection and a test plan.
  • Implementation sanity: pseudocode and grouping assumptions are coherent.

Inference and process optimization defaults

  • Prefer source-of-truth hierarchy:
    1. official source text (arXiv source/HTML when available),
    2. OCR feedback-loop output,
    3. targeted visual rechecks for disputed pages.
  • Use selective rereads: only revisit sections needed for unresolved claims.
  • Maintain a claim-to-evidence ledger while writing to reduce drift.
  • Minimize full-document rescans unless identity gate or consistency checks fail.
  • Keep output deterministic: fixed section order and explicit caveats.

Common pitfalls to avoid

  • Over-summarizing without experimental context.
  • Concept conflation (baseline vs weighting, confidence vs correctness).
  • Pseudocode that violates grouping assumptions.
  • Omitting caveats/limitations.

Must not do

  • Do not depend on external README files for policy.
  • Do not rely on memory summaries over source evidence.
  • Do not finalize if identity gate fails.
  • Do not skip revision passes or quality gates.

© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in claude-skills/paper-ocr-notes-pipeline of tokenbender/agent-guides.

Open the folder on GitHubat commit a74dd9d

Compare with similar skills

Paper OCR Notes Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paper OCR Notes Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paper OCR Notes Pipeline this skilltokenbender/agent-guides367—~1.8kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3151 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    315 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from tokenbender/agent-guides

All 11 skills in this repo
  • Capability Horizon Estimator

    tokenbender/agent-guides

    Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.

    367 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Technical Writing Workflow

    tokenbender/agent-guides

    A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.

    367 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Epistemic Libido

    tokenbender/agent-guides

    Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.

    367 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Manim Math Explainer

    tokenbender/agent-guides

    Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.

    367 GitHub stars~799 tokensUpdated 2 mo ago
    Auto-check passed
  • Worklog

    tokenbender/agent-guides

    Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.

    367 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • X Thread Reader

    tokenbender/agent-guides

    A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…

    367 GitHub stars~506 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Paper OCR Notes Pipeline

What does Paper OCR Notes Pipeline do?

Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md. Paper OCR Notes Pipeline is an agent skill from tokenbender/agent-guides.md.

When should I use Paper OCR Notes Pipeline?

Paper OCR Notes Pipeline fits situations like: tasks that involve PDF.

How do I install Paper OCR Notes Pipeline in Claude Code?

Run `npx skills add tokenbender/agent-guides --skill paper-ocr-notes-pipeline -a claude-code`. Or copy the skill folder (claude-skills/paper-ocr-notes-pipeline in tokenbender/agent-guides) into .claude/skills/paper-ocr-notes-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Paper OCR Notes Pipeline in Codex?

Run `npx skills add tokenbender/agent-guides --skill paper-ocr-notes-pipeline -a codex`. Or copy the skill folder (claude-skills/paper-ocr-notes-pipeline in tokenbender/agent-guides) into .agents/skills/paper-ocr-notes-pipeline in your project. Codex loads it when a task matches its description.

Can I use Paper OCR Notes Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill paper-ocr-notes-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paper-ocr-notes-pipeline, .gemini/skills/paper-ocr-notes-pipeline, .github/skills/paper-ocr-notes-pipeline and .opencode/skills/paper-ocr-notes-pipeline in your project.

What does Paper OCR Notes Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Paper OCR Notes Pipeline is instructions for the agent only.

Does Paper OCR Notes Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Paper OCR Notes Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paper OCR Notes Pipeline use?

Paper OCR Notes Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paper OCR Notes Pipeline use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Paper OCR Notes Pipeline?

Skills that share tags, products or a category with Paper OCR Notes Pipeline: Markitdown (ImCa0/just-laws, 782 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paper OCR Notes Pipeline?

tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.

Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.