Agent skill

Ocring Pdfs

by oaustegard in oaustegard/claude-skills

Adds a searchable text layer to a scanned PDF with ocrmypdf.

MITAuto-check passedDocuments & Office

Install Ocring Pdfs

skills CLI
$ npx skills add oaustegard/claude-skills --skill ocring-pdfs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills ocring-pdfs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ocring-pdfs .claude/skills/ocring-pdfs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ocring-pdfs
GitHub stars
150
Token cost
~1.2k tokens
SKILL.md length
621 words
Files
3 (incl. scripts)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Adds a searchable text layer to a scanned PDF with ocrmypdf.

  • A PDFs pages are images and the deliverable is a file to keep
  • SKILL.md covers Probe before installing, Install, Run and Which text-layer mode, plus 3 more sections
  • Runs Shell scripts from its folder; calls pdftotext, sh and apt-get
  • Hand to another tool: make this PDF searchable

What it does

Ocring Pdfs is an agent skill from oaustegard/claude-skills. Adds a searchable text layer to a scanned PDF with ocrmypdf. Installs the toolchain at runtime in ~18s. Use when a PDF's pages are images and the deliverable is a file to keep, grep, or hand to another tool: 'make this PDF searchable', 'OCR this scan', 'I can't select the text in this PDF', 'search across these scanned pages', 'the pdf skill returned nothing'. Handles non-English scans via tesseract language packs. NOT for charts, diagrams, handwriting, or slide layouts. Route those to transcribing-images.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `CHANGELOG.md` and `scripts/ensure_ocr.sh`).

It sits in Documents & Office, covering PDF, Slides and decks and Diagrams. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • A PDFs pages are images and the deliverable is a file to keep
  • Hand to another tool: make this PDF searchable
  • I cant select the text in this PDF
  • Search across these scanned pages

Example prompts

  • “make this PDF searchable”
  • “OCR this scan”
  • “t select the text in this PDF”
  • “/ocring-pdfs”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pdftotext
    • sh
    • apt-get
    • pdftoppm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ocring Pdfs loads about 1.2k tokens when it runs. Until then it costs about 131 tokens; SKILL.md has 621 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 621 words, ~1,226 tokens.

Download SKILL.mdSave it as .claude/skills/ocring-pdfs/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ocring-pdfs
description
Adds a searchable text layer to a scanned PDF with ocrmypdf. Installs the toolchain at runtime in ~18s. Use when a PDF's pages are images and the deliverable is a file to keep, grep, or hand to another tool: 'make this PDF searchable', 'OCR this scan', 'I can't select the text in this PDF', 'search across these scanned pages', 'the pdf skill returned nothing'. Handles non-English scans via tesseract language packs. NOT for charts, diagrams, handwriting, or slide layouts. Route those to transcribing-images.
metadata.version
0.1.0

OCRing PDFs

ocrmypdf writes an invisible text layer over the original page images, so the output file is both the scan you can look at and a document pdftotext, grep, and pdfplumber can read. Rasterize-then-tesseract gives you a .txt divorced from the pages; page numbers and coordinates are gone.

The toolchain is not in the base container. It installs in 18 seconds (measured 2026-09-12: apt 3s, pip 15s), so install it when a scan shows up rather than carrying it in a container layer.

Probe before installing

bash
pdftotext in.pdf - | tr -d '\f \n' | wc -c

Nonzero means the PDF already has a text layer and is not a scan. Extract with pdftotext or pdfplumber and stop. Running OCR on it wastes a minute, and with --force-ocr it replaces exact embedded text with a lossy reading of a raster of itself.

A small nonzero count (tens of characters across many pages) is the mixed case: a born-digital cover page in front of scanned body pages, or a scan whose producer stamped a header. --skip-text handles it.

Install

bash
sh scripts/ensure_ocr.sh              # English
sh scripts/ensure_ocr.sh nor deu      # plus Norwegian and German

Idempotent: 0.8s when everything is already present, 2.6s to add one more language pack. Installs ghostscript, pngquant, poppler-utils, tesseract and its language packs via apt, then ocrmypdf via pip.

Run

bash
ocrmypdf --skip-text --deskew --rotate-pages --output-type pdf in.pdf out.pdf
pdftotext out.pdf - | wc -w        # verify: zero words means it failed quietly

About 2s per page for a single dense page at 200 DPI on one core. A 300-page scan is therefore a background job, not a single bash call — launch it detached with a sentinel file per the external-call pattern in bash-tool-timeout.

Which text-layer mode

flaguse it when
--skip-textDefault. Pages that already carry text are passed through untouched; image-only pages get OCR. The safe choice for anything mixed.
--force-ocrEvery page is rasterized and re-OCRed, discarding any existing text. Correct for a scan carrying a junk text layer, and for pages with text-over-image that --skip-text would skip. Destroys real embedded text, so probe first.
--redo-ocrReplaces a previous OCR layer while leaving born-digital text alone. Narrower than --force-ocr and slower to fail on odd inputs.

--output-type pdf skips PDF/A conversion. Drop it when the output is going into an archive that requires PDF/A; ghostscript does the conversion either way.

Show full SKILL.md (273 more words)Show less

Languages

-l eng+nor for a mixed-language document, -l nor for a monolingual one. Order does not matter. Every code needs its tesseract-ocr-<code> pack installed. Pass the codes to ensure_ocr.sh and it handles them. Accuracy drops noticeably when the language is wrong, and tesseract will not tell you; it returns confident garbage instead.

Container facts (measured 2026-09-12)

  • apt-get update exits 100 here. A preconfigured nodesource repo is off the egress allowlist and returns 403, and the nonzero exit aborts any && chain behind it. The Ubuntu mirrors are reachable without an update. Run apt-get install directly.
  • unpaper is absent, so --clean and --clean-final fail. Don't pass them.
  • One core, so --jobs buys nothing on claude.ai. CCotw has four.
  • ocrmypdf --version prints to stderr. Capture with 2>&1 or a version check reads as empty.
  • jbig2 is absent; output uses CCITT/JPEG instead, which costs some file size and nothing else.

When to use transcribing-images instead

This skill produces glyphs. It does not read a chart, describe a diagram, or recover handwriting. Tesseract on those pages returns nothing useful and gives no sign that it lost anything.

Route to transcribing-images when the meaningful content is a picture, or when the deliverable is a reading rather than a file. Both is a normal answer: OCR the document so it is greppable, then send the pages that carry figures to a vision model.

In an interactive session, native vision beats both for a handful of pages: rasterize with pdftoppm -r 200 -png and view the images. Reach for OCR when the document is longer than context will hold, or when the text has to outlive the conversation as a file.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in ocring-pdfs of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/ensure_ocr.sh

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Ocring Pdfs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ocring Pdfs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ocring Pdfs this skilloaustegard/claude-skills150—~1.2kAutomated safety check: PassMIT
Ky Markdown RebuilderKyrieCheungYep/ky-markdown-rebuilder117—~5.7kAutomated safety check: PassNone
Extract Tikzbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~498Automated safety check: NotesCustom licence
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Paper Deckzsyggg/paper-craft-skills1.3k—~1.4kAutomated safety check: PassNone
Research To Diagramwshuyi/research-to-diagram1561 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Ky Markdown Rebuilder

    KyrieCheungYep/ky-markdown-rebuilder

    Rebuild visual documents into reliable Markdown by combining text extraction with page or screenshot alignment.

    117 GitHub stars~5.7k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Extract Tikz

    brycewang-stanford/Auto-Empirical-Research-Skills

    Extract TikZ diagrams from Beamer source, compile to PDF, convert to SVG with 0-based indexing.

    4.5k GitHub stars~498 tokensUpdated 4 days ago
    Documents & OfficeAuto-check: notes
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Paper Deck

    zsyggg/paper-craft-skills

    将论文、技术文章或知识内容制作成高真实感的 AIGC 幻灯片。先做叙事结构和逐页视觉导演,再调用生图模型生成每一页 16:9 slide image,最后合成为 PPTX/PDF。适合论文汇报、组会、公开课、技术分享、商业化研究展示;当用户提到“论文PPT”“AI生成PPT”“不像AI的PPT”“高质感幻灯片”“逐页生图PPT”时使用。

    1.3k GitHub stars~1.4k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Research To Diagram

    wshuyi/research-to-diagram

    深度调研主题并自动生成知识关系图谱PDF。接收研究主题后自动进行网络调研、信息收集、知识整理,最终生成专业的可视化关系图谱。适用于"研究...并做图"、"深度分析...并可视化"、"生成知识图谱"等场景。

    156 GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check passed
  • Paper2slides

    QuZhan51496/paper2anything

    Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.

    469 GitHub stars~3.8k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Questions about Ocring Pdfs

What does Ocring Pdfs do?

Adds a searchable text layer to a scanned PDF with ocrmypdf. Ocring Pdfs is an agent skill from oaustegard/claude-skills. Adds a searchable text layer to a scanned PDF with ocrmypdf.

When should I use Ocring Pdfs?

Ocring Pdfs fits situations like: A PDFs pages are images and the deliverable is a file to keep; hand to another tool: make this PDF searchable; I cant select the text in this PDF; search across these scanned pages.

How do I install Ocring Pdfs in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill ocring-pdfs -a claude-code`. Or copy the skill folder (ocring-pdfs in oaustegard/claude-skills) into .claude/skills/ocring-pdfs in your project. Claude Code loads it when a task matches its description.

How do I install Ocring Pdfs in Codex?

Run `npx skills add oaustegard/claude-skills --skill ocring-pdfs -a codex`. Or copy the skill folder (ocring-pdfs in oaustegard/claude-skills) into .agents/skills/ocring-pdfs in your project. Codex loads it when a task matches its description.

Can I use Ocring Pdfs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill ocring-pdfs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocring-pdfs, .gemini/skills/ocring-pdfs, .github/skills/ocring-pdfs and .opencode/skills/ocring-pdfs in your project.

What does Ocring Pdfs need to run?

Going by SKILL.md and its folder, Ocring Pdfs needs a shell for the scripts in its folder and the command-line tools its instructions call (pdftotext, sh, apt-get and pdftoppm). Our summary lists: A Bash shell.

Does Ocring Pdfs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ocring Pdfs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ocring Pdfs use?

Ocring Pdfs is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ocring Pdfs use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ocring Pdfs?

Skills that share tags, products or a category with Ocring Pdfs: Ky Markdown Rebuilder (KyrieCheungYep/ky-markdown-rebuilder, 117 stars), Extract Tikz (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Harness Book Best Practice (wquguru/harness-books, 3.2k stars) and Paper Deck (zsyggg/paper-craft-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ocring Pdfs?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.