Unified tool for analyzing and displaying images (vision analysis and image display).

MITAuto-check passed

Install Image

skills CLI
$ npx skills add PhyAgentOS/PhyAgentOS-core --skill image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PhyAgentOS/PhyAgentOS-core image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PhyAgentOS/PhyAgentOS-core.git skills-src && mkdir -p .claude/skills && cp -r skills-src/PhyAgentOS/skills/image .claude/skills/image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image
GitHub stars
2.8k
Token cost
~899 tokens
SKILL.md length
308 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Unified tool for analyzing and displaying images (vision analysis and image display).

  • Works in 3 steps: ALWAYS use image tool - Never attempt… → Absolute paths only - Convert all paths… → Keep text unchanged in generate mode -…
  • SKILL.md covers Features, Tools, Examples and Important Rules
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Image is an agent skill from PhyAgentOS/PhyAgentOS-core. Unified tool for analyzing and displaying images (vision analysis and image display).

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: PhyAgentOS is a Recursive Self-Improving (RSI) physical agent operating system that enables agents to recursively self-improve through agentic workflows. The licence is MIT.

Example prompts

  • “/image”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. ALWAYS use image tool - Never attempt direct LLM API calls
  2. Absolute paths only - Convert all paths to absolute before calling
  3. Keep text unchanged in generate mode - Do not change the text when generating image

What it can do on your machine

Read from SKILL.md and the folder at commit 466e38b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image loads about 899 tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~899

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PhyAgentOS/PhyAgentOS-core at commit 466e38b, republished under its MIT licence (© PhyAgentOS). 308 words, ~899 tokens.

Download SKILL.mdSave it as .claude/skills/image/SKILL.md (or your agent's skills folder).
name
image
description
Unified tool for analyzing and displaying images (vision analysis and image display).

Image Tool

Unified tool for analyzing and displaying images. Supports two modes: vision analysis using multimodal LLM models and displaying images to users through the frontend.

Features

  • Analyze images using multimodal LLM models (OCR, description, visual QA)
  • Display images to users by sending them to the frontend
  • Generate images from text prompts using AI models
  • Support multiple image formats: PNG, JPG, JPEG, GIF, WEBP, BMP
  • Base64 encoding for image processing and transmission

Tools

This skill provides the following tool:

image

Unified tool for image analysis, display, and generation.

Parameters:

  • mode (string, required): The operation mode - vision for image analysis, display for showing images to users, generate for creating images from text prompts
  • image_path (string, required):
    • In vision/display mode: Absolute path to the image file (e.g., "/Users/photo.png")
    • In generate mode: File name where the generated image will be saved. This parameter should contain few words which generalize the text
  • text (string, optional):
    • In vision mode: User's request or question about the image (e.g., "Describe this image", "Extract text from this image")
    • In display mode: Caption to display with the image (appears above the image in the message box)
    • In generate mode: Text prompt describing the desired image content, style, and composition (supports Chinese and English, max 800 characters)

Examples

Example for vision - Describe an image at image.png:

<tool>image</tool>
<parameter name="mode">vision</parameter>
<parameter name="text">Describe this image</parameter>
<parameter name="image_path">image.png</parameter>

Example for vision - Extract text from image image.jpeg:

<tool>image</tool>
<parameter name="mode">vision</parameter>
<parameter name="text">Get words in this image</parameter>
<parameter name="image_path">image.jpeg</parameter>

Example for vision - How many birds in image image.png:

<tool>image</tool>
<parameter name="mode">vision</parameter>
<parameter name="text">How many birds in image?</parameter>
<parameter name="image_path">image.png</parameter>

Example for display - Display image at image.png with title "The image":

<tool>image</tool>
<parameter name="mode">display</parameter>
<parameter name="text">The image</parameter>
<parameter name="image_path">image.png</parameter>

Example for display - Display image at image.png:

<tool>image</tool>
<parameter name="mode">display</parameter>
<parameter name="image_path">image.png</parameter>

Example for generate - Create an image of a cute orange cat:

<tool>image</tool>
<parameter name="mode">generate</parameter>
<parameter name="text">A sitting orange cat with happy expression</parameter>
<parameter name="image_path">sitting_orange_cat.png</parameter>

Example for generate - Create a landscape painting:

<tool>image</tool>
<parameter name="mode">generate</parameter>
<parameter name="text">A beautiful sunset over mountains in oil painting style</parameter>
<parameter name="image_path">sunset_painting_style.png</parameter>

Important Rules

  1. ALWAYS use image tool - Never attempt direct LLM API calls
  2. Absolute paths only - Convert all paths to absolute before calling
  3. Keep text unchanged in generate mode - Do not change the text when generating image

© PhyAgentOS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in PhyAgentOS/skills/image of PhyAgentOS/PhyAgentOS-core.

Open the folder on GitHubat commit 466e38b

Compare with similar skills

Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image this skillPhyAgentOS/PhyAgentOS-core2.8k—~899Automated safety check: PassMIT
Fal Visionnexu-io/open-design100k—~295Automated safety check: PassApache-2.0
Vision Sftwshobson/agents40k—~2kAutomated safety check: PassMIT
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Senior Computer Visiondavila7/claude-code-templates32k3 repos~1.4kAutomated safety check: PassMIT
Agent Code Analyzerruvnet/ruflo74k3 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Fal Vision

    nexu-io/open-design

    Analyze images — segment objects, detect, run OCR, describe, and answer visual questions via fal.ai vision models.

    100k GitHub stars~295 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vision Sft

    wshobson/agents

    Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    davila7/claude-code-templates

    World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems.

    32k GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent skill for code-analyzer - invoke with $agent-code-analyzer

    74k GitHub starsUsed in 3 repos~1.5k tokens
    DevelopmentAuto-check passed
  • Agent skill for pagerank-analyzer - invoke with $agent-pagerank-analyzer

    74k GitHub starsUsed in 3 repos~2.9k tokens
    Auto-check passed

More from PhyAgentOS/PhyAgentOS-core

  • Clawhub

    PhyAgentOS/PhyAgentOS-core

    Search and install agent skills from ClawHub, the public skill registry.

    2.8k GitHub starsUsed in 3 repos~351 tokens
    Auto-check passed
  • Agent Mode

    PhyAgentOS/PhyAgentOS-core

    Unified tool for managing agent LLM modes (add, remove, update, list, switch).

    2.8k GitHub stars~711 tokensUpdated today
    Auto-check passed
  • Memory

    PhyAgentOS/PhyAgentOS-core

    Two-layer memory system with grep-based recall. An agent skill from PhyAgentOS/PhyAgentOS-core.

    2.8k GitHub starsUsed in 1 repo~365 tokens
    Auto-check passed
  • Cron

    PhyAgentOS/PhyAgentOS-core

    Schedule reminders and recurring tasks.

    2.8k GitHub starsUsed in 2 repos~377 tokens
    Auto-check passed

Questions about Image

What does Image do?

Unified tool for analyzing and displaying images (vision analysis and image display). Image is an agent skill from PhyAgentOS/PhyAgentOS-core. Unified tool for analyzing and displaying images (vision analysis and image display).

How do I install Image in Claude Code?

Run `npx skills add PhyAgentOS/PhyAgentOS-core --skill image -a claude-code`. Or copy the skill folder (PhyAgentOS/skills/image in PhyAgentOS/PhyAgentOS-core) into .claude/skills/image in your project. Claude Code loads it when a task matches its description.

How do I install Image in Codex?

Run `npx skills add PhyAgentOS/PhyAgentOS-core --skill image -a codex`. Or copy the skill folder (PhyAgentOS/skills/image in PhyAgentOS/PhyAgentOS-core) into .agents/skills/image in your project. Codex loads it when a task matches its description.

Can I use Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PhyAgentOS/PhyAgentOS-core --skill image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image, .gemini/skills/image, .github/skills/image and .opencode/skills/image in your project.

What does Image need to run?

SKILL.md names no scripts, command-line tools or credentials: Image is instructions for the agent only.

Does Image access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image use?

Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image use?

About 899 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image?

Skills that share tags, products or a category with Image: Fal Vision (nexu-io/open-design, 100k stars), Vision Sft (wshobson/agents, 40k stars), Vision (gridaco/grida, 2.7k stars) and Senior Computer Vision (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image?

PhyAgentOS (a GitHub organization) maintains it in PhyAgentOS/PhyAgentOS-core, which has 2,758 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: PhyAgentOS/PhyAgentOS-core on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.