Agent skill

Transcribing Images

by oaustegard in oaustegard/claude-skills

Reads the visual content of slides, pages, and images the way a human would, not just their embedded text.

MITAuto-check passedDocuments & Office

Install Transcribing Images

skills CLI
$ npx skills add oaustegard/claude-skills --skill transcribing-images -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills transcribing-images --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/transcribing-images .claude/skills/transcribing-images && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
transcribing-images
GitHub stars
150
Token cost
~1.3k tokens
SKILL.md length
643 words
Files
3 (incl. scripts)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Reads the visual content of slides, pages, and images the way a human would, not just their embedded text.

  • PDF has image slides
  • SKILL.md covers When to reach for this vs. the…, The pipeline, Choosing the model and OCR fallback (tesseract), plus 2 more sections
  • Runs Python scripts from its folder; calls python3
  • Scanned figures

What it does

Transcribing Images is an agent skill from oaustegard/claude-skills. Reads the visual content of slides, pages, and images the way a human would, not just their embedded text. Use when a PPTX or PDF has image slides, screenshots, charts, scanned figures, or flattened-to-image layouts that the built-in pptx/pdf skills read as empty; when asked to transcribe, describe, OCR, or extract what is shown in an image, slide deck, or document page; or when embedded-text extraction returned little or nothing from a visually rich file. Triggers on 'read this deck', 'what's on these slides'…

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `CHANGELOG.md` and `scripts/transcribe_pages.py`).

It sits in Documents & Office, covering Slides and decks, Transcription and PowerPoint presentations. It works with Microsoft PowerPoint. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • PDF has image slides
  • Scanned figures
  • Flattened-to-image layouts that the built-in pptx/pdf skills read as empty
  • Asked to transcribe

Example prompts

  • “read this deck”
  • “s on these slides”
  • “transcribe”
  • “/transcribing-images”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Transcribing Images loads about 1.3k tokens when it runs. Until then it costs about 164 tokens; SKILL.md has 643 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 643 words, ~1,307 tokens.

Download SKILL.mdSave it as .claude/skills/transcribing-images/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
transcribing-images
description
Reads the visual content of slides, pages, and images the way a human would, not just their embedded text. Use when a PPTX or PDF has image slides, screenshots, charts, scanned figures, or flattened-to-image layouts that the built-in pptx/pdf skills read as empty; when asked to transcribe, describe, OCR, or extract what is shown in an image, slide deck, or document page; or when embedded-text extraction returned little or nothing from a visually rich file. Triggers on 'read this deck', 'what's on these slides', 'transcribe', 'OCR', 'extract text from image', 'describe this chart/diagram', .pptx/.pdf/.png/.jpg with visual content.
metadata.version
0.2.0

Transcribing Images

Read what a slide, page, or image actually shows — text plus charts, diagrams, screenshots, and layout — by rasterizing it and sending the picture to a vision model. This is the fix for the gap the built-in pptx and pdf skills leave: they extract embedded text only, so an image slide, a chart, or a scanned figure reads as empty. Visual transcription reads it the way a person looking at the slide would.

When to reach for this vs. the built-in skills

Use the pptx / pdf skills first for text-native documents — a normal deck or report where the content is real text boxes. They are faster and exact.

Switch to this skill when text extraction comes back thin or empty on a file you can see is visually rich, or whenever the meaningful content is a picture: chart, graph, diagram, screenshot, photo, scanned page, or a slide exported as one flat image. Don't guess which case you're in — if pptx/pdf returned little from a file that clearly has content, that is the signal.

The pipeline

Everything routes through scripts/transcribe_pages.py, which handles all three ingress paths and one bad page never aborts the rest:

  • .pptx / .ppt → LibreOffice headless → PDF → pdftoppm → one PNG per slide
  • .pdf → pdftoppm → one PNG per page
  • image file → used directly as a single page

Each page image is then transcribed by a vision model. Run it directly:

bash
python3 scripts/transcribe_pages.py deck.pptx                  # all slides
python3 scripts/transcribe_pages.py report.pdf --pages 3-7     # subset
python3 scripts/transcribe_pages.py slide.png --model opus     # one image
python3 scripts/transcribe_pages.py deck.pptx --json out.json  # structured

Or import transcribe_file(...) for programmatic use; it returns a list of {page, image, text, error} dicts.

Choosing the model

The transcription core and its empirical cost/recall data are reused from browsing-bluesky/scripts/image_transcribe.py — same registry, kept in sync. Pick with --model:

  • gemini-lite (default) — cheapest and fastest, ~95% token recall on dense screenshots. Right for routine deck reading.
  • gemini-flash — token-perfect, ~3x the cost. Use when exact text matters.
  • gemini-3.5-flash — heavier reasoning alongside transcription, ~19x cost. Use when a page needs interpretation, not just reading. Check invoking-gemini's model table for the current frontier Flash before assuming this is still the strongest reasoning tier available.
  • opus — for interactive sessions where you want the reading in your own context anyway.
  • haiku — only if constrained to single-vendor Anthropic; weak at dense transcription (tends to summarize instead of transcribe). That finding is from Haiku 4.5; the alias now resolves to claude-haiku-5-5 and has not been re-measured.

Default to gemini-lite and escalate only when recall or reasoning demands it.

Show full SKILL.md (252 more words)Show less

OCR fallback (tesseract)

Tesseract 5.x is installed (eng + osd language packs only) and is exposed as --engine tesseract. It returns glyphs, not a reading — no chart interpretation, no diagram description, no layout meaning. Use it only for pages you already know are plain scanned text, when you want a zero-cost, fully-offline pass. For anything with a chart, diagram, or visual layout, the vision path is the correct tool; tesseract on those pages will quietly lose the content that mattered.

When the deliverable is a file rather than a reading, use ocring-pdfs instead. It installs ocrmypdf at runtime (~18s) and writes the OCR back as an invisible text layer over the original page images, so the output is a PDF that pdftotext and grep can read. The --engine tesseract path here returns loose text with no page anchoring.

Interactive shortcut

In an interactive session you can often skip the model call entirely: rasterize with scripts/transcribe_pages.py … --json to get the page PNGs, or just convert and view each page image yourself — Claude reads images natively. The script's vision-model path exists for batch and autonomous runs where no human-in-loop reader is available, or when a deck has more pages than is practical to view one by one.

DPI

Default raster is 150 DPI — legible for a vision model and safely under the 5 MB/image base64 ceiling. Bump to --dpi 200–300 only for pages with dense small fonts; higher DPI risks exceeding the per-image size limit and costs more tokens for no gain on normal slides.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in transcribing-images of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/transcribe_pages.py

Open the folder on GitHubat commit 90b0f1b

Compare with similar skills

Transcribing Images next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Transcribing Images compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Transcribing Images this skilloaustegard/claude-skills150—~1.3kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Paper Deckzsyggg/paper-craft-skills1.3k—~1.4kAutomated safety check: PassNone
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
Paper2slidesQuZhan51496/paper2anything450—~3.8kAutomated safety check: NotesApache-2.0
Ky Markdown RebuilderKyrieCheungYep/ky-markdown-rebuilder117—~5.7kAutomated safety check: PassNone

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Paper Deck

    zsyggg/paper-craft-skills

    将论文、技术文章或知识内容制作成高真实感的 AIGC 幻灯片。先做叙事结构和逐页视觉导演,再调用生图模型生成每一页 16:9 slide image,最后合成为 PPTX/PDF。适合论文汇报、组会、公开课、技术分享、商业化研究展示;当用户提到“论文PPT”“AI生成PPT”“不像AI的PPT”“高质感幻灯片”“逐页生图PPT”时使用。

    1.3k GitHub stars~1.4k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Paper2slides

    QuZhan51496/paper2anything

    Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.

    450 GitHub stars~3.8k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Ky Markdown Rebuilder

    KyrieCheungYep/ky-markdown-rebuilder

    Rebuild visual documents into reliable Markdown by combining text extraction with page or screenshot alignment.

    117 GitHub stars~5.7k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Sci HTML

    ShZhao27208/Aut_Sci_Write

    Generate academic presentation-style HTML slide decks and browser reports from PDFs, structured text, Markdown, paper summaries, outlines, or research notes.

    209 GitHub stars~1.4k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Deciding With Confidence

    oaustegard/claude-skills

    Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

    150 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Transcribing Images

What does Transcribing Images do?

Reads the visual content of slides, pages, and images the way a human would, not just their embedded text. Transcribing Images is an agent skill from oaustegard/claude-skills. Reads the visual content of slides, pages, and images the way a human would, not just their embedded text.

When should I use Transcribing Images?

Transcribing Images fits situations like: PDF has image slides; scanned figures; flattened-to-image layouts that the built-in pptx/pdf skills read as empty; asked to transcribe.

How do I install Transcribing Images in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill transcribing-images -a claude-code`. Or copy the skill folder (transcribing-images in oaustegard/claude-skills) into .claude/skills/transcribing-images in your project. Claude Code loads it when a task matches its description.

How do I install Transcribing Images in Codex?

Run `npx skills add oaustegard/claude-skills --skill transcribing-images -a codex`. Or copy the skill folder (transcribing-images in oaustegard/claude-skills) into .agents/skills/transcribing-images in your project. Codex loads it when a task matches its description.

Can I use Transcribing Images in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill transcribing-images -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcribing-images, .gemini/skills/transcribing-images, .github/skills/transcribing-images and .opencode/skills/transcribing-images in your project.

What does Transcribing Images need to run?

Going by SKILL.md and its folder, Transcribing Images needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Transcribing Images access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Transcribing Images safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Transcribing Images use?

Transcribing Images is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Transcribing Images use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Transcribing Images?

Skills that share tags, products or a category with Transcribing Images: Markitdown (ImCa0/just-laws, 781 stars), Paper Deck (zsyggg/paper-craft-skills, 1.3k stars), Markitdown (jimmc414/Kosmos, 595 stars) and Paper2slides (QuZhan51496/paper2anything, 450 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Transcribing Images?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.