Markdown Article Formatter
JimLiu/baoyu-skills
Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.
A skill your agent uses when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .claude/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .claude/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocrType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .agents/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .agents/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .cursor/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .cursor/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/hashgraph-online/awesome-codex-plugins.git --path plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .gemini/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .gemini/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocrInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .github/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .github/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-with-ocr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr .opencode/skills/extracting-with-ocr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "extracting-with-ocr" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr into .opencode/skills/extracting-with-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extracting-with-ocr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
extracting-with-ocrA skill your agent uses when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer.
Extracting With OCR is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 9e7b281. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
aptbrewpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Extracting With OCR loads about 1.2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 474 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo apt install tesseract-ocr-deu tesseract-ocr-jpn tesseract-ocr-frasudo apt install tesseract-ocr-<iso639-2>Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from hashgraph-online/awesome-codex-plugins at commit 9e7b281, republished under its Apache-2.0 licence (© hashgraph-online). 474 words, ~1,249 tokens.
.claude/skills/extracting-with-ocr/SKILL.md (or your agent's skills folder).Use this when a document is image-based: scanned PDFs, photographed pages, screenshots, JPEG/PNG/TIFF with text. Kreuzberg auto-OCRs raster images and auto-detects PDFs that lack a text layer. Force it on when extraction returned empty/garbled text from a PDF that "looks" textual.
content field, but the file opens visually.kreuzberg extract scan.pdf --force-ocr=true
kreuzberg extract scan.pdf --ocr=true --ocr-language engIf a page has an unreliable text layer, --force-ocr=true re-rasterizes
and runs OCR on every page.
Tesseract is the default and ships with the CLI — no extra install. Other backends are opt-in:
| Backend | Flag | Install | Notes |
|---|---|---|---|
| Tesseract | --ocr-backend tesseract (default) | bundled | Best general-purpose, 100+ languages via tessdata. |
| PaddleOCR | --ocr-backend paddle-ocr | bundled (ONNX Runtime) | Strong on Asian scripts. Not available on WASM or Windows. |
| EasyOCR | --ocr-backend easyocr | Python binding (pip install kreuzberg[easyocr]) | Heavier model. CUDA accel via easyocr_kwargs={"gpu": True}. |
| VLM (vision) | layout + a multimodal LLM via config | configured per backend | Use when OCR fails on dense or handwritten layouts. |
Pick Tesseract first. Switch only when accuracy is unacceptable.
Tesseract uses ISO 639-2 codes. Default is eng. Combine with +:
kreuzberg extract menu.jpg --ocr=true --ocr-language "eng+deu"
kreuzberg extract bilingual.pdf --ocr-language "eng+jpn"
kreuzberg extract any.pdf --ocr-language all # all installed packsInstall missing packs at the OS level:
# macOS
brew install tesseract-lang
# Debian/Ubuntu
sudo apt install tesseract-ocr-deu tesseract-ocr-jpn tesseract-ocr-fra
# Specific lang only
sudo apt install tesseract-ocr-<iso639-2>Kreuzberg fails fast with a helpful error if you request a language pack that is not installed. Read the error — it names the missing file.
--ocr=true — enable OCR (auto-enabled for images and scanned PDFs).--force-ocr=true — OCR every page even if a text layer exists.--disable-ocr=true — never OCR (extract embedded text only or fail).--ocr-language <lang> — single code or +-joined list, or all.--ocr-backend <tesseract|paddle-ocr|easyocr> — pick backend.--ocr-auto-rotate=true — pre-rotate via the auto-rotate model.--acceleration <cpu|coreml|cuda|tensorrt|auto> — ONNX accelerator for
paddle-ocr / auto-rotate / layout models.--no-cache=true unless you have a reason.kreuzberg batch *.pdf --ocr=true — internal worker
pool parallelizes across CPU cores. Cap with --max-concurrent N if
memory is tight.--target-dpi (default 300) only for low-resolution scans. Higher
DPI is slower; 200 is usually enough for printed text.--ocr-auto-rotate=true only when pages may be rotated; the
classifier adds latency.--acceleration coreml typically beats CPU for
paddle-ocr and layout detection.Long flag chains belong in kreuzberg.toml — auto-discovered from cwd
upward.
force_ocr = true
output_format = "markdown"
[ocr]
backend = "tesseract"
language = "eng+deu"
auto_rotate = trueThen just run:
kreuzberg extract document.pdf--force-ocr — the file has a
bogus zero-width text layer. Re-run with --force-ocr=true.--ocr-auto-rotate=true or pre-rotate.--ocr-language; consider paddle-ocr for Chinese/Japanese.See references/cli-reference.md and references/configuration.md in the
sibling kreuzberg skill for the full flag and config schema.
© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr of hashgraph-online/awesome-codex-plugins.
Open the folder on GitHubat commit 9e7b281
Extracting With OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Extracting With OCR this skillhashgraph-online/awesome-codex-plugins | 1.3k | — | ~1.2k | Automated safety check: Notes | Apache-2.0 | |
| Markdown Article FormatterJimLiu/baoyu-skills | 26k | 6 repos | ~3.5k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 782 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Obsidian MarkdownAtmosphere/atmosphere | 3.8k | 20 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| DOCXrvdbreemen/OTGW-firmware | 207 | 33 repos | ~4.3k | Automated safety check: Pass | Proprietary | |
| Gzh Designisjiamu/gzh-design-skill | 3.9k | 1 repos | ~2.2k | Automated safety check: Pass | AGPL-3.0 |
JimLiu/baoyu-skills
Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Atmosphere/atmosphere
Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.
rvdbreemen/OTGW-firmware
A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).
isjiamu/gzh-design-skill
微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…
HKUDS/DeepTutor
Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.
hashgraph-online/awesome-codex-plugins
Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.
hashgraph-online/awesome-codex-plugins
Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).
hashgraph-online/awesome-codex-plugins
A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…
hashgraph-online/awesome-codex-plugins
Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…
hashgraph-online/awesome-codex-plugins
Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.
hashgraph-online/awesome-codex-plugins
Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.
Categories
A skill your agent uses when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Extracting With OCR is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer.
Extracting With OCR fits situations like: extracting text from scanned PDFs; photographed pages; images that have no embedded text layer.
Run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr in hashgraph-online/awesome-codex-plugins) into .claude/skills/extracting-with-ocr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-with-ocr in hashgraph-online/awesome-codex-plugins) into .agents/skills/extracting-with-ocr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-with-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-with-ocr, .gemini/skills/extracting-with-ocr, .github/skills/extracting-with-ocr and .opencode/skills/extracting-with-ocr in your project.
Going by SKILL.md and its folder, Extracting With OCR needs the command-line tools its instructions call (apt, brew and pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Extracting With OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Extracting With OCR: Markdown Article Formatter (JimLiu/baoyu-skills, 26k stars), Markitdown (ImCa0/just-laws, 782 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,255 GitHub stars. The repository holds 714 skills in this directory. The repository was last updated on October 9, 2026.
Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.