Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

MITAuto-check passedAI & LLM Engineering

Install Vision

skills CLI
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xiincs/claude-code-vision-skill vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/vision .claude/skills/vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vision
GitHub stars
170
Token cost
~1.2k tokens
SKILL.md length
407 words
Files
2
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

  • You need to understand screenshots
  • SKILL.md covers When to use this tool, Quick start, Providers and Configuration, plus 1 more section
  • Runs Python scripts from its folder; calls python and pip; needs DOUBAO_API_KEY and DASHSCOPE_API_KEY
  • Any image content

What it does

Vision is an agent skill from xiincs/claude-code-vision-skill. Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Use when you need to understand screenshots, UI layouts, diagrams, or any image content. Supports png/jpg/webp/gif.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `vision.py`).

It sits in AI & LLM Engineering, covering Computer vision and Diagrams. It works with DeepSeek, OpenAI, Qwen and Python. The repository describes itself as: 为 Claude Code 赋能多模态视觉能力,适配 纯文本 LLM 底座,用于截图 / UI / 图表分析;搭配 browser-harness 可做前端布局自动化检查。 The licence is MIT.

When your agent uses it

  • You need to understand screenshots
  • Any image content

Example prompts

  • “/vision”

Requirements

  • Python 3
  • A credential in DOUBAO_API_KEY
  • A credential in DASHSCOPE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 32ec684. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DOUBAO_API_KEY
    • DASHSCOPE_API_KEY
    • DEEPSEEK_API_KEY
    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY
    • MYAPI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vision loads about 1.2k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 407 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xiincs/claude-code-vision-skill at commit 32ec684, republished under its MIT licence (© xiincs). 407 words, ~1,208 tokens.

Download SKILL.mdSave it as .claude/skills/vision/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vision
description
Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Use when you need to understand screenshots, UI layouts, diagrams, or any image content. Supports png/jpg/webp/gif.

vision

Multi-provider vision tool. Call various vision models to describe images. Feed it a prompt + image path, get back a text description.

When to use this tool

If you can already see and understand the image yourself (native multimodal model), skip this tool — analyze it directly.

A SessionStart hook normally announces this session's routing status up front. If that context isn't visible (e.g. compacted out of a long conversation, or the hook isn't installed), check before calling this tool:

bash
python vision.py --check-routing
  • native → you already have native image understanding this session; don't call this tool.
  • external (default) → proceed with the quick start below.

Quick start

bash
python vision.py [--provider <name>] <image_path> <prompt>

When --provider is omitted, the provider is resolved by: --provider flag > VISION_PROVIDER env > first API key found.

Providers

doubao (Volcengine Ark)
  • API key: DOUBAO_API_KEY
  • Default model: doubao-seed-2-0-pro-260215
  • Custom endpoint: DOUBAO_BASE_URL
qwen (DashScope)
  • API key: DASHSCOPE_API_KEY
  • Default model: qwen-vl-max
  • Custom endpoint: DASHSCOPE_BASE_URL
  • Available models: qwen-vl-max, qwen-vl-plus, qvq-max
deepseek (DeepSeek)
  • API key: DEEPSEEK_API_KEY
  • Default model: deepseek-v4-flash-vision-exp
  • Custom endpoint: DEEPSEEK_BASE_URL
  • Only deepseek-v4-flash-vision-exp accepts images — deepseek-v4-flash and deepseek-v4-pro are text-only and reject image input with an error.
openai (GPT-4o)
  • API key: OPENAI_API_KEY
  • Default model: gpt-4o
  • Custom endpoint: OPENAI_BASE_URL
  • Also works with any OpenAI-compatible endpoint.
anthropic (Claude)
  • API key: ANTHROPIC_API_KEY
  • Default model: claude-sonnet-5
  • Custom endpoint: ANTHROPIC_BASE_URL
  • Requires the anthropic package (pip install anthropic); it's imported lazily so other providers work without it.
Show full SKILL.md (188 more words)Show less
any custom provider

Any --provider name outside the built-in ones is resolved dynamically from environment variables named after it — no code changes needed:

Env VarRequiredNotes
{NAME}_API_KEYyeschecked at request time, same as built-ins
{NAME}_BASE_URLyesno default — arbitrary endpoint
{NAME}_MODELyesno default (or set global VISION_MODEL instead)
{NAME}_PROTOCOLnoopenai (default) or anthropic — picks the request shape

openai covers essentially every OpenAI-compatible endpoint (vLLM, Ollama, LiteLLM, OpenRouter, Azure OpenAI, self-hosted proxies, ...). Use {NAME}_PROTOCOL=anthropic only if the endpoint speaks the Anthropic Messages API shape.

bash
export MYAPI_API_KEY="sk-xxx"
export MYAPI_BASE_URL="https://my-endpoint.example.com/v1"
export MYAPI_MODEL="my-vision-model"
python vision.py --provider myapi "screenshot.png" "describe this"

If {NAME}_BASE_URL or {NAME}_MODEL is missing, the tool prints exactly which variables to set instead of a generic "unknown provider" error.

Configuration

Env VarScopeDefault
VISION_PROVIDERDefault provider (built-in or custom name)auto-detect (built-ins only)
VISION_MODELOverride model (all providers)provider default
{PROVIDER}_MODELOverride model (per provider)—
{PROVIDER}_BASE_URLOverride/define endpoint (per provider)built-in default, or required for custom
{PROVIDER}_PROTOCOLRequest shape for a custom provider: openai | anthropicopenai
VISION_TEMPERATUREResponse creativity 0–10
VISION_MAX_TOKENSMax response tokens4096

Note: auto-detect (no --provider / VISION_PROVIDER set) only scans the built-in providers' API keys — a custom provider must always be named explicitly.

Examples

bash
# Auto-detect provider from API keys
python vision.py "screenshot.png" "Describe the page layout and any visible UI issues."

# Explicit provider
python vision.py --provider qwen "mockup.png" "List all components, colors, and spacing patterns."

# Custom model
QWEN_MODEL=qvq-max python vision.py --provider qwen "diagram.png" "Explain the architecture."

# GPT-4o for visual regression
python vision.py -p openai "after.png" "Compare with app design spec, flag differences."

# Fully custom provider (self-hosted, third-party proxy, any OpenAI-compatible endpoint)
MYAPI_API_KEY=sk-xxx MYAPI_BASE_URL=https://host/v1 MYAPI_MODEL=my-model \
  python vision.py --provider myapi "ui.png" "Analyze layout issues"

© xiincs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in vision of xiincs/claude-code-vision-skill.

  • SKILL.md
  • vision.py

Open the folder on GitHubat commit 32ec684

Compare with similar skills

Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vision this skillxiincs/claude-code-vision-skill170—~1.2kAutomated safety check: PassMIT
Vllm Ascendascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNone
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Pocketmen With Yousix-nut/PocketMen-with-you310—~2.6kAutomated safety check: PassMIT
Bridgic LLMsbitsky-tech/bridgic155—~839Automated safety check: NotesMIT

Similar skills

  • Vllm Ascend

    ascend-ai-coding/awesome-ascend-skills

    vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

    174 GitHub stars~2.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Pocketmen With You

    six-nut/PocketMen-with-you

    Turn 2+ user reference images into a high-fidelity animated Codex companion using PocketMen's own local stack.

    310 GitHub stars~2.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Bridgic LLMs

    bitsky-tech/bridgic

    LLM provider initialization for bridgic projects. An agent skill from bitsky-tech/bridgic.

    155 GitHub stars~839 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Mmsp Python

    Prism-Shadow/model-message-stream-protocol

    Guidance for using the MMSP Python SDK (mmsp). An agent skill from Prism-Shadow/model-message-stream-protocol.

    113 GitHub stars~1.4k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

Questions about Vision

What does Vision do?

Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Vision is an agent skill from xiincs/claude-code-vision-skill. Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

When should I use Vision?

Vision fits situations like: you need to understand screenshots; any image content.

How do I install Vision in Claude Code?

Run `npx skills add xiincs/claude-code-vision-skill --skill vision -a claude-code`. Or copy the skill folder (vision in xiincs/claude-code-vision-skill) into .claude/skills/vision in your project. Claude Code loads it when a task matches its description.

How do I install Vision in Codex?

Run `npx skills add xiincs/claude-code-vision-skill --skill vision -a codex`. Or copy the skill folder (vision in xiincs/claude-code-vision-skill) into .agents/skills/vision in your project. Codex loads it when a task matches its description.

Can I use Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiincs/claude-code-vision-skill --skill vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision, .gemini/skills/vision, .github/skills/vision and .opencode/skills/vision in your project.

What does Vision need to run?

Going by SKILL.md and its folder, Vision needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named DOUBAO_API_KEY, DASHSCOPE_API_KEY, DEEPSEEK_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in DOUBAO_API_KEY; A credential in DASHSCOPE_API_KEY.

Does Vision access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vision use?

Vision is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vision use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vision?

Skills that share tags, products or a category with Vision: Vllm Ascend (ascend-ai-coding/awesome-ascend-skills, 174 stars), ModLens Image Vision Bridge (liustack/modlens, 4.2k stars), Dingo Verify (MigoXLab/dingo, 757 stars) and Pocketmen With You (six-nut/PocketMen-with-you, 310 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vision?

xiincs (a GitHub user) maintains it in xiincs/claude-code-vision-skill, which has 170 GitHub stars. The repository was last updated on August 25, 2026.

Source: xiincs/claude-code-vision-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.