Agent skill

Image OCR

by Tai609 in Tai609/NebulaMat

Read and transcribe the text visible inside an image file (PNG/JPG/JPEG/WebP/BMP/GIF) using the built-in Windows Media.Ocr OCR engine, with NO external credential or network required.

Custom licenceAuto-check passedMedia & Creative

Install Image OCR

skills CLI
$ npx skills add Tai609/NebulaMat --skill image-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Tai609/NebulaMat image-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Tai609/NebulaMat.git skills-src && mkdir -p .claude/skills && cp -r skills-src/runtime/skills-bundle/image-ocr .claude/skills/image-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-ocr
GitHub stars
100
Token cost
~1.1k tokens
SKILL.md length
521 words
Files
2 (incl. scripts)
Skills in repo
15
Repo updated
First seen
Licence
Custom licence

At a glance

Read and transcribe the text visible inside an image file (PNG/JPG/JPEG/WebP/BMP/GIF) using the built-in Windows Media.Ocr OCR engine, with NO external credential or network required.

  • Works in 4 steps: glob for the image path (e.g.… → If it is outside the current allowed… → Run ocr-image.ps1 on the resolved path. → …
  • Tasks that involve Transcription
  • SKILL.md covers When to use, What it CAN and CANNOT do, How to run and Getting the image into the…, plus 3 more sections
  • Runs PowerShell scripts from its folder; calls pwsh; needs VISION_API_KEY

What it does

Image OCR is an agent skill from Tai609/NebulaMat. Read and transcribe the text visible inside an image file (PNG/JPG/JPEG/WebP/BMP/GIF) using the built-in Windows Media.Ocr OCR engine, with NO external credential or network required. Use this whenever the model cannot read an image directly, the VISIONAPIKEY credential is missing, or the user pastes/attaches a screenshot and asks what it says or what is in it. Also computes image dimensions/format so the agent can describe scene context. English and Chinese (and other installed Windows OCR languages) supported.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts.

It sits in Media & Creative, covering Transcription. It works with PowerShell. The repository describes itself as: NebulaMat scientific materials research workbench.

When your agent uses it

  • Tasks that involve Transcription

Example prompts

  • “/image-ocr”

Requirements

  • PowerShell
  • A credential in VISION_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. glob for the image path (e.g. **/pasted.png).
  2. If it is outside the current allowed workspace, Copy-Item it into the CWD.
  3. Run ocr-image.ps1 on the resolved path.
  4. Merge the OCR lines into a clean reading, fixing obvious character splits

What it can do on your machine

Read from SKILL.md and the folder at commit c906ed5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (PowerShell), which the agent can run.

    Shell commands in SKILL.md call:

    • pwsh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VISION_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image OCR loads about 1.1k tokens when it runs. Until then it costs about 132 tokens; SKILL.md has 521 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 521 words (~1,054 tokens).

“This skill gives a text-only agent a fully local, credential-free way to "see" the text inside an image, using the OCR engine that ships with Windows 10/11 (Windows.Media.Ocr). It does not need VISION_API_KEY or any network call.”

— opening of SKILL.md by Tai609, Custom licence
name
image-ocr

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (scripts) in runtime/skills-bundle/image-ocr of Tai609/NebulaMat.

  • SKILL.md
  • scripts/ocr-image.ps1

Open the folder on GitHubat commit c906ed5

Compare with similar skills

Image OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image OCR this skillTai609/NebulaMat100—~1.1kAutomated safety check: PassCustom licence
Bilibili Page ReaderMisaka-Mikoto-Tech/agent-skills275—~3.6kAutomated safety check: PassMIT
Videodbaffaan-m/ECC275k3 repos~3.5kAutomated safety check: NotesMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Bilibili Transcribechubbyguan/chubbyskills1.2k1 repos~578Automated safety check: NotesMIT
Video Assemblezenstory-ai/video-recap-skills5551 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Bilibili Page Reader

    Misaka-Mikoto-Tech/agent-skills

    Get content from Bilibili videos: official subtitles, danmaku (density/peaks/sample), comments.

    275 GitHub stars~3.6k tokensUpdated 17 days ago
    Media & CreativeAuto-check passed
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Bilibili Transcribe

    chubbyguan/chubbyskills

    哔哩哔哩视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub starsUsed in 1 repo~578 tokens
    Media & CreativeAuto-check: notes
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    555 GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • Douyin Transcribe

    chubbyguan/chubbyskills

    抖音视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub starsUsed in 1 repo~551 tokens
    Media & CreativeAuto-check: notes

More from Tai609/NebulaMat

All 15 skills in this repo
  • Nature Figure

    Tai609/NebulaMat

    Create, revise, audit, and export submission-grade scientific figures for Nature-family and other high-impact venues in Python (matplotlib/seaborn) or R (ggplot2/patchwork/ComplexHeatmap), including…

    100 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Aris Proof Orchestrator

    Tai609/NebulaMat

    Manage a stateful, run-directory-based proof project with Codex: continuation across runs, run-local source bookkeeping, manual GPT Pro handoff packages when a local attempt stalls, and an optional…

    100 GitHub stars~5k tokensUpdated 1 mo ago
    Auto-check passed
  • Build provenance-controlled facet-specific electrochemical adsorption model cohorts from MatterGen candidates after MatterSim relaxation, including slab terminations, adsorption sites and…

    100 GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Generate auditable candidate crystal structures with the project's pinned MatterGen runtime, including request construction, model/conditioning selection, GPU-cost bounds, workspace-safe output…

    100 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Run governed MatterSim-v1.0.0-5M energy, force, stress, and optional ASE FIRE relaxation for workspace-local standardized structures.

    100 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Nature Academic Search

    Tai609/NebulaMat

    Multi-source literature search, citation verification, strict independent other-citation audits, article-level citation metric tables, influential citer profiling with citation-context extraction…

    100 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Image OCR

What does Image OCR do?

Read and transcribe the text visible inside an image file (PNG/JPG/JPEG/WebP/BMP/GIF) using the built-in Windows Media.Ocr OCR engine, with NO external credential or network required. Image OCR is an agent skill from Tai609/NebulaMat.Ocr OCR engine, with NO external credential or network required.

When should I use Image OCR?

Image OCR fits situations like: tasks that involve Transcription.

How do I install Image OCR in Claude Code?

Run `npx skills add Tai609/NebulaMat --skill image-ocr -a claude-code`. Or copy the skill folder (runtime/skills-bundle/image-ocr in Tai609/NebulaMat) into .claude/skills/image-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Image OCR in Codex?

Run `npx skills add Tai609/NebulaMat --skill image-ocr -a codex`. Or copy the skill folder (runtime/skills-bundle/image-ocr in Tai609/NebulaMat) into .agents/skills/image-ocr in your project. Codex loads it when a task matches its description.

Can I use Image OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Tai609/NebulaMat --skill image-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-ocr, .gemini/skills/image-ocr, .github/skills/image-ocr and .opencode/skills/image-ocr in your project.

What does Image OCR need to run?

Going by SKILL.md and its folder, Image OCR needs PowerShell for the scripts in its folder, the command-line tools its instructions call (pwsh) and credentials named VISION_API_KEY. Our summary lists: PowerShell; A credential in VISION_API_KEY.

Does Image OCR access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Image OCR use?

Image OCR has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Image OCR use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image OCR?

Skills that share tags, products or a category with Image OCR: Bilibili Page Reader (Misaka-Mikoto-Tech/agent-skills, 275 stars), Videodb (affaan-m/ECC, 275k stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Bilibili Transcribe (chubbyguan/chubbyskills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image OCR?

Tai609 (a GitHub user) maintains it in Tai609/NebulaMat, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on September 6, 2026.

Source: Tai609/NebulaMat on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.