Agent skill

Image To Text

by godot-fun in godot-fun/gai

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…

MITAuto-check passedMedia & Creative

Install Image To Text

skills CLI
$ npx skills add godot-fun/gai --skill image-to-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install godot-fun/gai image-to-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/image-to-text .claude/skills/image-to-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-to-text
GitHub stars
183
Token cost
~813 tokens
SKILL.md length
266 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…

  • The user asks to describe
  • SKILL.md covers Rules, Usage, Output guidance and Troubleshooting
  • Calls python
  • Extract visible details and text from PNG

What it does

Image To Text is an agent skill from godot-fun/gai. Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back to CPU. Use when the user asks to describe, caption, compare, transcribe, inspect, or extract visible details and text from PNG, JPEG, WebP, GIF, BMP, or TIFF images without a GPU or remote vision API.

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering LLM inference and serving and Transcription. It works with llama.cpp. The repository describes itself as: A lightweight AI agent and skill workflow framework built with Godot. The licence is MIT.

When your agent uses it

  • The user asks to describe
  • Extract visible details and text from PNG
  • TIFF images without a GPU
  • Remote vision API

Example prompts

  • “Use the image-to-text skill to convert one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model…”
  • “/image-to-text”

What it can do on your machine

Read from SKILL.md and the folder at commit 6394686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image To Text loads about 813 tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 266 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~813

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from godot-fun/gai at commit 6394686, republished under its MIT licence (© godot-fun). 266 words, ~813 tokens.

Download SKILL.mdSave it as .claude/skills/image-to-text/SKILL.md (or your agent's skills folder).
name
image-to-text
description
Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back to CPU. Use when the user asks to describe, caption, compare, transcribe, inspect, or extract visible details and text from PNG, JPEG, WebP, GIF, BMP, or TIFF images without a GPU or remote vision API.

Image to Text

Run the official openbmb/MiniCPM-V-4.6-gguf Q4_K_M model locally with llama.cpp. Do not send images to remote APIs.

Rules

Read and follow skill-dependency-manager before running commands.

  • Run .ai/image-to-text/image_to_text.py through the default python manifest entry only.
  • Load the language model and multimodal projector only from .dependency/minicpm-v-4.6/model.
  • Before inference, check the llama-cpp-gpu runtime with --list-devices. Prefer it when a Vulkan device is reported; otherwise use llama-cpp-cpu.
  • Use --gpu or -GPU to require GPU inference and fail when unavailable.
  • Use --cpu to bypass detection and force CPU inference for troubleshooting or comparison.
  • Preserve observable facts and clearly distinguish uncertainty from fact.
  • Pass related images after one --images or --image option; both names are equivalent.

Usage

powershell
.dependency/python/python.exe .ai/image-to-text/image_to_text.py --images C:\path\photo.png
powershell
.dependency/python/python.exe .ai/image-to-text/image_to_text.py --images before.png after.png --prompt "Compare these images and list only visible changes."
powershell
.dependency/python/python.exe .ai/image-to-text/image_to_text.py --image C:\path\photo.png -GPU
powershell
.dependency/python/python.exe .ai/image-to-text/image_to_text.py --image C:\path\photo.png --cpu

Use --output description.md to also save UTF-8 text. See cli/image-to-text.md for copy-paste commands.

Output guidance

  • General description: mention scene, subjects, actions, composition, colors, and legible text.
  • OCR: preserve reading order and line breaks where practical; mark uncertain characters with [unclear].
  • Accessibility alt text: be concise and omit decorative speculation.
  • Structured extraction: request JSON explicitly and validate it before presenting it as machine-readable.
  • Never identify a real person from appearance alone or infer sensitive traits not explicitly visible as text.

Troubleshooting

  • Missing CPU runtime: verify .dependency/llama-cpp-cpu/llama-mtmd-cli.exe exists.
  • GPU unavailable: verify .dependency/llama-cpp-gpu/llama-mtmd-cli.exe exists and --list-devices reports a Vulkan device. Update the graphics driver if necessary.
  • Missing model: verify MiniCPM-V-4_6-Q4_K_M.gguf and mmproj-model-f16.gguf exist under .dependency/minicpm-v-4.6/model.
  • Out of memory: close other applications or process fewer/smaller images; the official GGUF listing states approximately 2GB memory, but working memory varies with image size and context.
  • Slow inference: CPU speed, core count, memory bandwidth, and image resolution directly affect latency.

© godot-fun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/image-to-text of godot-fun/gai.

Open the folder on GitHubat commit 6394686

Compare with similar skills

Image To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image To Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image To Text this skillgodot-fun/gai183—~813Automated safety check: PassMIT
Local AI App Integrationamd/skills408—~6kAutomated safety check: PassMIT
Teach A ModelAseiel/VideoHighlighter165—~839Automated safety check: PassAGPL-3.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Helmor Bump Vendorsdohooo/helmor1.3k—~2.1kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT

Similar skills

  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    408 GitHub stars~6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Teach A Model

    Aseiel/VideoHighlighter

    Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

    165 GitHub stars~839 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Helmor Bump Vendors

    dohooo/helmor

    Bump or upgrade the pinned versions of Helmor's bundled agent CLIs, SDKs, and supporting binaries — Claude Code + claude-agent-sdk (lockstep), Codex, Cursor SDK, OpenCode, Kimi, Pi, and gh / glab /…

    1.3k GitHub stars~2.1k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from godot-fun/gai

All 36 skills in this repo
  • AI Text To Speech

    godot-fun/gai

    Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).

    183 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audio Denoise

    godot-fun/gai

    Reduces background noise in a single audio file using FFmpeg afftdn.

    183 GitHub stars~465 tokensUpdated today
    Auto-check passed
  • Audio Fade

    godot-fun/gai

    Applies fade-in and fade-out at the start and end of a single audio file using FFmpeg.

    183 GitHub stars~652 tokensUpdated today
    Auto-check passed
  • Normalizes a single audio file to consistent LUFS loudness with true-peak limiting using FFmpeg.

    183 GitHub stars~684 tokensUpdated today
    Auto-check passed
  • Standardizes a single audio file to 44100 or 48000 Hz and exports 16-bit PCM WAV using FFmpeg.

    183 GitHub stars~638 tokensUpdated today
    Auto-check passed
  • Audio Split

    godot-fun/gai

    Splits a single audio file into two segments (part 1 before the split point, part 2 after) using FFmpeg.

    183 GitHub stars~554 tokensUpdated today
    Auto-check passed

Works with

Questions about Image To Text

What does Image To Text do?

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…. Image To Text is an agent skill from godot-fun/gai.cpp, automatically preferring an available Vulkan GPU and falling back to CPU.

When should I use Image To Text?

Image To Text fits situations like: the user asks to describe; extract visible details and text from PNG; TIFF images without a GPU; remote vision API.

How do I install Image To Text in Claude Code?

Run `npx skills add godot-fun/gai --skill image-to-text -a claude-code`. Or copy the skill folder (.agents/skills/image-to-text in godot-fun/gai) into .claude/skills/image-to-text in your project. Claude Code loads it when a task matches its description.

How do I install Image To Text in Codex?

Run `npx skills add godot-fun/gai --skill image-to-text -a codex`. Or copy the skill folder (.agents/skills/image-to-text in godot-fun/gai) into .agents/skills/image-to-text in your project. Codex loads it when a task matches its description.

Can I use Image To Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add godot-fun/gai --skill image-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-to-text, .gemini/skills/image-to-text, .github/skills/image-to-text and .opencode/skills/image-to-text in your project.

What does Image To Text need to run?

Going by SKILL.md and its folder, Image To Text needs the command-line tools its instructions call (python).

Does Image To Text access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image To Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image To Text use?

Image To Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image To Text use?

About 813 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image To Text?

Skills that share tags, products or a category with Image To Text: Local AI App Integration (amd/skills, 408 stars), Teach A Model (Aseiel/VideoHighlighter, 165 stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars) and Helmor Bump Vendors (dohooo/helmor, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image To Text?

godot-fun (a GitHub organization) maintains it in godot-fun/gai, which has 183 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 10, 2026.

Source: godot-fun/gai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.