Vllm Ascend
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install xiincs/claude-code-vision-skill vision --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/vision .claude/skills/vision && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .claude/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/xiincs/claude-code-vision-skill/tree/main/visionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install xiincs/claude-code-vision-skill vision --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/vision .agents/skills/vision && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .agents/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install xiincs/claude-code-vision-skill vision --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/vision .cursor/skills/vision && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .cursor/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/xiincs/claude-code-vision-skill.git --path vision--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install xiincs/claude-code-vision-skill vision --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/vision .gemini/skills/vision && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .gemini/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install xiincs/claude-code-vision-skill visionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/vision .github/skills/vision && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .github/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xiincs/claude-code-vision-skill --skill vision -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install xiincs/claude-code-vision-skill vision --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiincs/claude-code-vision-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/vision .opencode/skills/vision && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/xiincs/claude-code-vision-skill/tree/main/vision into .opencode/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
visionCall vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.
Vision is an agent skill from xiincs/claude-code-vision-skill. Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Use when you need to understand screenshots, UI layouts, diagrams, or any image content. Supports png/jpg/webp/gif.
Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `vision.py`).
It sits in AI & LLM Engineering, covering Computer vision and Diagrams. It works with DeepSeek, OpenAI, Qwen and Python. The repository describes itself as: 为 Claude Code 赋能多模态视觉能力,适配 纯文本 LLM 底座,用于截图 / UI / 图表分析;搭配 browser-harness 可做前端布局自动化检查。 The licence is MIT.
Read from SKILL.md and the folder at commit 32ec684. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DOUBAO_API_KEYDASHSCOPE_API_KEYDEEPSEEK_API_KEYOPENAI_API_KEYANTHROPIC_API_KEYMYAPI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vision loads about 1.2k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 407 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from xiincs/claude-code-vision-skill at commit 32ec684, republished under its MIT licence (© xiincs). 407 words, ~1,208 tokens.
.claude/skills/vision/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Multi-provider vision tool. Call various vision models to describe images. Feed it a prompt + image path, get back a text description.
If you can already see and understand the image yourself (native multimodal model), skip this tool — analyze it directly.
A SessionStart hook normally announces this session's routing status up front. If that context isn't visible (e.g. compacted out of a long conversation, or the hook isn't installed), check before calling this tool:
python vision.py --check-routingnative → you already have native image understanding this session; don't call this tool.external (default) → proceed with the quick start below.python vision.py [--provider <name>] <image_path> <prompt>When --provider is omitted, the provider is resolved by: --provider flag > VISION_PROVIDER env > first API key found.
DOUBAO_API_KEYdoubao-seed-2-0-pro-260215DOUBAO_BASE_URLDASHSCOPE_API_KEYqwen-vl-maxDASHSCOPE_BASE_URLqwen-vl-max, qwen-vl-plus, qvq-maxDEEPSEEK_API_KEYdeepseek-v4-flash-vision-expDEEPSEEK_BASE_URLdeepseek-v4-flash-vision-exp accepts images — deepseek-v4-flash and deepseek-v4-pro are text-only and reject image input with an error.OPENAI_API_KEYgpt-4oOPENAI_BASE_URLANTHROPIC_API_KEYclaude-sonnet-5ANTHROPIC_BASE_URLanthropic package (pip install anthropic); it's imported lazily so other providers work without it.Any --provider name outside the built-in ones is resolved dynamically from
environment variables named after it — no code changes needed:
| Env Var | Required | Notes |
|---|---|---|
{NAME}_API_KEY | yes | checked at request time, same as built-ins |
{NAME}_BASE_URL | yes | no default — arbitrary endpoint |
{NAME}_MODEL | yes | no default (or set global VISION_MODEL instead) |
{NAME}_PROTOCOL | no | openai (default) or anthropic — picks the request shape |
openai covers essentially every OpenAI-compatible endpoint (vLLM, Ollama,
LiteLLM, OpenRouter, Azure OpenAI, self-hosted proxies, ...). Use
{NAME}_PROTOCOL=anthropic only if the endpoint speaks the Anthropic Messages
API shape.
export MYAPI_API_KEY="sk-xxx"
export MYAPI_BASE_URL="https://my-endpoint.example.com/v1"
export MYAPI_MODEL="my-vision-model"
python vision.py --provider myapi "screenshot.png" "describe this"If {NAME}_BASE_URL or {NAME}_MODEL is missing, the tool prints exactly which
variables to set instead of a generic "unknown provider" error.
| Env Var | Scope | Default |
|---|---|---|
VISION_PROVIDER | Default provider (built-in or custom name) | auto-detect (built-ins only) |
VISION_MODEL | Override model (all providers) | provider default |
{PROVIDER}_MODEL | Override model (per provider) | — |
{PROVIDER}_BASE_URL | Override/define endpoint (per provider) | built-in default, or required for custom |
{PROVIDER}_PROTOCOL | Request shape for a custom provider: openai | anthropic | openai |
VISION_TEMPERATURE | Response creativity 0–1 | 0 |
VISION_MAX_TOKENS | Max response tokens | 4096 |
Note: auto-detect (no --provider / VISION_PROVIDER set) only scans the
built-in providers' API keys — a custom provider must always be named explicitly.
# Auto-detect provider from API keys
python vision.py "screenshot.png" "Describe the page layout and any visible UI issues."
# Explicit provider
python vision.py --provider qwen "mockup.png" "List all components, colors, and spacing patterns."
# Custom model
QWEN_MODEL=qvq-max python vision.py --provider qwen "diagram.png" "Explain the architecture."
# GPT-4o for visual regression
python vision.py -p openai "after.png" "Compare with app design spec, flag differences."
# Fully custom provider (self-hosted, third-party proxy, any OpenAI-compatible endpoint)
MYAPI_API_KEY=sk-xxx MYAPI_BASE_URL=https://host/v1 MYAPI_MODEL=my-model \
python vision.py --provider myapi "ui.png" "Analyze layout issues"© xiincs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in vision of xiincs/claude-code-vision-skill.
Open the folder on GitHubat commit 32ec684
Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vision this skillxiincs/claude-code-vision-skill | 170 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Vllm Ascendascend-ai-coding/awesome-ascend-skills | 174 | — | ~2.7k | Automated safety check: Pass | None | |
| ModLens Image Vision Bridgeliustack/modlens | 4.2k | — | ~1.3k | Automated safety check: Notes | MIT | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~833 | Automated safety check: Pass | Apache-2.0 | |
| Pocketmen With Yousix-nut/PocketMen-with-you | 310 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Bridgic LLMsbitsky-tech/bridgic | 155 | — | ~839 | Automated safety check: Notes | MIT |
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
liustack/modlens
Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
six-nut/PocketMen-with-you
Turn 2+ user reference images into a high-fidelity animated Codex companion using PocketMen's own local stack.
bitsky-tech/bridgic
LLM provider initialization for bridgic projects. An agent skill from bitsky-tech/bridgic.
Prism-Shadow/model-message-stream-protocol
Guidance for using the MMSP Python SDK (mmsp). An agent skill from Prism-Shadow/model-message-stream-protocol.
Categories
Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Vision is an agent skill from xiincs/claude-code-vision-skill. Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.
Vision fits situations like: you need to understand screenshots; any image content.
Run `npx skills add xiincs/claude-code-vision-skill --skill vision -a claude-code`. Or copy the skill folder (vision in xiincs/claude-code-vision-skill) into .claude/skills/vision in your project. Claude Code loads it when a task matches its description.
Run `npx skills add xiincs/claude-code-vision-skill --skill vision -a codex`. Or copy the skill folder (vision in xiincs/claude-code-vision-skill) into .agents/skills/vision in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiincs/claude-code-vision-skill --skill vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision, .gemini/skills/vision, .github/skills/vision and .opencode/skills/vision in your project.
Going by SKILL.md and its folder, Vision needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named DOUBAO_API_KEY, DASHSCOPE_API_KEY, DEEPSEEK_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in DOUBAO_API_KEY; A credential in DASHSCOPE_API_KEY.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vision is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vision: Vllm Ascend (ascend-ai-coding/awesome-ascend-skills, 174 stars), ModLens Image Vision Bridge (liustack/modlens, 4.2k stars), Dingo Verify (MigoXLab/dingo, 757 stars) and Pocketmen With You (six-nut/PocketMen-with-you, 310 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
xiincs (a GitHub user) maintains it in xiincs/claude-code-vision-skill, which has 170 GitHub stars. The repository was last updated on August 25, 2026.
Source: xiincs/claude-code-vision-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.