Agent skill

Image

by axoviq-ai in axoviq-ai/synthadoc

Extract text from images using a vision LLM. An agent skill from axoviq-ai/synthadoc.

AGPL-3.0-or-laterAuto-check passedKnowledge Management

Install Image

skills CLI
$ npx skills add axoviq-ai/synthadoc --skill image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install axoviq-ai/synthadoc image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/axoviq-ai/synthadoc.git skills-src && mkdir -p .claude/skills && cp -r skills-src/synthadoc/skills/image .claude/skills/image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image
GitHub stars
1.6k
Token cost
~544 tokens
SKILL.md length
116 words
Files
3 (incl. scripts)
Skills in repo
9
Repo updated
First seen
Licence
AGPL-3.0-or-later

At a glance

Extract text from images using a vision LLM. An agent skill from axoviq-ai/synthadoc.

  • Knowledge Management work in your project
  • SKILL.md covers Setup, Standalone usage and When this skill is used
  • Runs Python scripts from its folder

What it does

Image is an agent skill from axoviq-ai/synthadoc. Extract text from images using a vision LLM

Its SKILL.md is about 540 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/__init__.py` and `scripts/main.py`).

It sits in Knowledge Management. It works with Obsidian. The repository describes itself as: Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional… The licence is AGPL-3.0-or-later.

When your agent uses it

  • Knowledge Management work in your project

Example prompts

  • “/image”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 66c9089. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image loads about 544 tokens when it runs. Until then it costs about 12 tokens; SKILL.md has 116 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~12
When it runs · the whole SKILL.md, loaded when a task matches
~544

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from axoviq-ai/synthadoc at commit 66c9089, republished under its AGPL-3.0-or-later licence (© axoviq-ai). 116 words, ~544 tokens.

Download SKILL.mdSave it as .claude/skills/image/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
image
description
Extract text from images using a vision LLM
version
1.0
entry.script
scripts/main.py
entry.class
ImageSkill
triggers.extensions
.png, .jpg, .jpeg, .webp, .gif, .tiff
triggers.intents
image, screenshot, diagram, photo
author
axoviq.com
license
AGPL-3.0-or-later

Image Skill

Base64-encodes the image and passes it to a vision-capable LLM that extracts all text and key information. Returns the LLM's response as result.text.

Setup

No pip dependency — the skill uses only the Python standard library plus a LLM provider you supply at construction time. The provider can be any object that implements the complete() interface (see below).

Standalone usage

python
import asyncio
from synthadoc.skills.image.scripts.main import ImageSkill

# ImageSkill REQUIRES a vision-capable provider — calling extract() without
# one raises ValueError immediately.
skill = ImageSkill(provider=my_provider)

async def main():
    result = await skill.extract("/path/to/screenshot.png")
    print(result.text)          # extracted text from the image
    print(result.metadata)      # {"tokens_input": N, "tokens_output": N}

asyncio.run(main())

Provider interface — any object with this async method:

python
async def complete(
    messages: list,             # list of Message objects from synthadoc.skills.base
    system: str | None = None,
    temperature: float = 0.0,
    max_tokens: int = 4096,
) -> object                     # must have .text (str), .input_tokens (int), .output_tokens (int)

Build the provider with any vision-capable model. Message is importable from synthadoc.skills.base — no dependency on synthadoc.providers:

python
from synthadoc.skills.base import Message

Supported image formats: .png, .jpg/.jpeg, .webp, .gif, .tiff

When this skill is used

  • Source path ends with .png, .jpg, .jpeg, .webp, .gif, or .tiff
  • User intent contains: image, screenshot, diagram, photo

© axoviq-ai, AGPL-3.0-or-later. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in synthadoc/skills/image of axoviq-ai/synthadoc.

  • SKILL.md
  • scripts/__init__.py
  • scripts/main.py

Open the folder on GitHubat commit 66c9089

Compare with similar skills

Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image this skillaxoviq-ai/synthadoc1.6k—~544Automated safety check: PassAGPL-3.0-or-later
Obsidian CLIAtmosphere/atmosphere3.8k13 repos~795Automated safety check: PassApache-2.0
Obsidian Canvas BoardsAgriciDaniel/claude-obsidian15k—~1.4kAutomated safety check: PassMIT
LLM Wikilewislulu/llm-wiki-skill655—~3.7kAutomated safety check: PassNone
Second BrainNicholasSpisak/second-brain737—~1.5kAutomated safety check: NotesNone
Codex History IngestAr9av/obsidian-wiki3.5k—~2.2kAutomated safety check: NotesMIT

Similar skills

  • Obsidian CLI

    Atmosphere/atmosphere

    Interact with Obsidian vaults using the Obsidian CLI to read, create, search, and manage notes, tasks, properties, and more.

    3.8k GitHub starsUsed in 13 repos~795 tokens
    Knowledge ManagementAuto-check passed
  • Obsidian Canvas Boards

    AgriciDaniel/claude-obsidian

    Creates, inspects and updates Obsidian JSON Canvas boards in a vault, with text, file, link, group and edge nodes, using safe recoverable edits.

    15k GitHub stars~1.4k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • LLM Wiki

    lewislulu/llm-wiki-skill

    Build and maintain a Karpathy-style LLM knowledge base — a self-compiling Obsidian markdown wiki where an Agent ingests raw sources, compiles cross-linked concept/entity/summary pages, answers…

    655 GitHub stars~3.7k tokensUpdated 5 mo ago
    Knowledge ManagementAuto-check passed
  • Second Brain

    NicholasSpisak/second-brain

    Set up a new Obsidian knowledge base with the LLM Wiki pattern.

    737 GitHub stars~1.5k tokensUpdated 6 mo ago
    Knowledge ManagementAuto-check: notes
  • Codex History Ingest

    Ar9av/obsidian-wiki

    Ingest Codex CLI conversation/session history into Obsidian as distilled knowledge.

    3.5k GitHub stars~2.2k tokensUpdated 2 days ago
    Knowledge ManagementAuto-check: notes
  • Autograph

    smixs/agent-second-brain

    Schema-as-code enforcement for any Obsidian vault. An agent skill from smixs/agent-second-brain.

    393 GitHub stars~3.7k tokensUpdated 2 mo ago
    Knowledge ManagementAuto-check passed

More from axoviq-ai/synthadoc

All 9 skills in this repo
  • Session

    axoviq-ai/synthadoc

    Extract conversation turns from AI session history files (.jsonl)

    1.6k GitHub stars~902 tokensUpdated today
    Auto-check passed
  • Youtube

    axoviq-ai/synthadoc

    Extract transcripts from YouTube videos via the YouTube caption system

    1.6k GitHub stars~634 tokensUpdated today
    Auto-check passed
  • PPTX

    axoviq-ai/synthadoc

    Extract text from Microsoft PowerPoint presentations. An agent skill from axoviq-ai/synthadoc.

    1.6k GitHub stars~248 tokensUpdated today
    Auto-check passed
  • DOCX

    axoviq-ai/synthadoc

    Extract text from Microsoft Word documents. An agent skill from axoviq-ai/synthadoc.

    1.6k GitHub stars~207 tokensUpdated today
    Auto-check passed
  • PDF

    axoviq-ai/synthadoc

    Extract text from PDF documents

    1.6k GitHub stars~293 tokensUpdated today
    Auto-check passed
  • URL

    axoviq-ai/synthadoc

    Fetch and extract text from web URLs

    1.6k GitHub stars~371 tokensUpdated today
    Auto-check passed

Works with

Questions about Image

What does Image do?

Extract text from images using a vision LLM. An agent skill from axoviq-ai/synthadoc. Image is an agent skill from axoviq-ai/synthadoc.

When should I use Image?

Image fits situations like: knowledge Management work in your project.

How do I install Image in Claude Code?

Run `npx skills add axoviq-ai/synthadoc --skill image -a claude-code`. Or copy the skill folder (synthadoc/skills/image in axoviq-ai/synthadoc) into .claude/skills/image in your project. Claude Code loads it when a task matches its description.

How do I install Image in Codex?

Run `npx skills add axoviq-ai/synthadoc --skill image -a codex`. Or copy the skill folder (synthadoc/skills/image in axoviq-ai/synthadoc) into .agents/skills/image in your project. Codex loads it when a task matches its description.

Can I use Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add axoviq-ai/synthadoc --skill image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image, .gemini/skills/image, .github/skills/image and .opencode/skills/image in your project.

What does Image need to run?

Going by SKILL.md and its folder, Image needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Image access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Image use?

Image is published under the AGPL-3.0-or-later licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image use?

About 544 tokens (SKILL.md is roughly 2.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image?

Skills that share tags, products or a category with Image: Obsidian CLI (Atmosphere/atmosphere, 3.8k stars), Obsidian Canvas Boards (AgriciDaniel/claude-obsidian, 15k stars), LLM Wiki (lewislulu/llm-wiki-skill, 655 stars) and Second Brain (NicholasSpisak/second-brain, 737 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image?

axoviq-ai (a GitHub organization) maintains it in axoviq-ai/synthadoc, which has 1,576 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 10, 2026.

Source: axoviq-ai/synthadoc on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.