Computer Use
TheSyart/emperor-agent
Operate web pages in Emperor's built-in browser with the browser tools, and macOS app windows with the desktop tools — open or bind a target, read it, fill, click or type, and verify the result.
Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Anionex/agent-vision-toolkit vision-skills --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vision-skills .claude/skills/vision-skills && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .claude/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skillsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Anionex/agent-vision-toolkit vision-skills --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/vision-skills .agents/skills/vision-skills && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .agents/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Anionex/agent-vision-toolkit vision-skills --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/vision-skills .cursor/skills/vision-skills && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .cursor/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Anionex/agent-vision-toolkit.git --path skills/vision-skills--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Anionex/agent-vision-toolkit vision-skills --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/vision-skills .gemini/skills/vision-skills && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .gemini/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Anionex/agent-vision-toolkit vision-skillsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/vision-skills .github/skills/vision-skills && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .github/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Anionex/agent-vision-toolkit vision-skills --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/vision-skills .opencode/skills/vision-skills && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vision-skills" agent skill from https://github.com/Anionex/agent-vision-toolkit/tree/main/skills/vision-skills into .opencode/skills/vision-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-skills", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vision-skillsLocal vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…
Vision Skills is an agent skill from Anionex/agent-vision-toolkit. Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/htmlshot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operating a GUI from screenshots — and to re-check an…
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/gui.md` and `references/long-screenshot-ocr.md`).
It sits in Productivity & Automation, covering Desktop control. It works with DeepSeek. The repository describes itself as: 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend…. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1384ef4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VISION_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vision Skills loads about 4k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 149 tokens; SKILL.md has 1,874 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Anionex/agent-vision-toolkit at commit 1384ef4, republished under its MIT licence (© Anionex). 1,874 words, ~4,047 tokens.
.claude/skills/vision-skills/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Five local CLIs that give a text-only agent eyes. They read one shared
vision config (VISION_API_KEY / VISION_BASE_URL / VISION_MODEL /
LANG), plus the optional Python-client settings VISION_API_PROTOCOL,
VISION_REASONING_EFFORT, and VISION_USER_AGENT — no extra credentials.
Pick the tool by the question you are answering:
| Question | Tool |
|---|---|
| "What does this image show / say?" | glance |
| "Where is X?" — a thing you can name | ground |
| "Where are all the Xs?" — every instance of a kind | detect |
| "What is its exact shape, size, offset?" | trace |
| "Cut this box out as its own image file" | crop |
| "OCR this long screenshot / scrolling page / chat history" | scripts/long_screenshot_ocr.py |
| "Extract the icon/logo foreground as transparent PNG — manual region or auto (cropped+scaled screenshots)" | scripts/extract_fg.py |
| "Turn this HTML file into a viewport or full-page screenshot" | scripts/html_shot.py |
| "Which colours dominate a region, and which palette value fits it?" | scripts/dominant_colors.py |
| A relation none of them return — a gap, a distance between two located things | code over the pixels (Pillow) |
glance answers what something is; ground and detect answer where.
You give ground a description of a particular thing; you give detect a
kind and it enumerates the instances.
Both give real coordinates, but they are not pixel-exact: the box arrives
on a 0-1000 grid and is scaled to your image, so the last pixel or few are
not reliable. That is accurate enough to crop with, to click, to compare
positions against. When a number has to be exact, trace derives it from
the actual pixels — offsets, sizes, shapes.
Everything this toolkit ships a tool for, call the tool — do not rewrite it with Pillow in the middle of a task. The CLIs exist so the same pixel work is not hand-coded differently every time:
crop, not Image.open(...).crop(...)scripts/dominant_colors.pyscripts/pixel_diff.pytraceground / detectglancescripts/long_screenshot_ocr.pyscripts/html_shot.pyHand-written Pillow is only for what none of them return: a relation
between two things you already located (a gap, a distance), a resize or
overlay, drawing. If you catch yourself writing .crop(), .convert(),
or histogram code where one of the tools above fits, replace it with the
tool call — same coordinates, same box format, and the output feeds the
next tool directly.
glance <image> # detailed description
glance <image> -q "<question>" # targeted question (qualitative only)
glance <image> --ocr # verbatim OCR
glance <image> --region X1,Y1,X2,Y2 -q "..." # zoom into a crop
glance <img1> <img2> -q "..." # compare in ONE callWhen you do compare with glance, pass all paths to one call — separate
calls cannot see both images, so two descriptions compared afterwards are
two hallucination surfaces, not a comparison. --region uploads only the
crop, so small text and icons become readable.
But "what changed between these two?" is not a glance question. A one-word
badge or a small shift is a rounding error to a vision model and exact to
scripts/pixel_diff.py. Diff first to get the box, then glance --region
that box to read what the change actually is.
For a tall scrolling screenshot, do not send the whole image through one OCR
call and accept the model's downscaling loss. Run the long-screenshot workflow,
which finds low-content cut bands, invokes glance on each chunk, uses
structured extraction for chat histories, merges only duplicated overlap, and
writes a boundary audit:
python3 scripts/long_screenshot_ocr.py work/page.png -o work/page.ocr.md
python3 scripts/long_screenshot_ocr.py work/chat.png --mode chat --resume -o work/chat.ocr.mdRead references/long-screenshot-ocr.md before using it. It defines the
verification pass for unsafe cuts and chat-message boundaries.
ground <image> "<target description>"
ground <image> "<target>" --region X1,Y1,X2,Y2Output: x1: .., y1: .., x2: .., y2: .. in original-image pixels — with
--region too (crop hits are mapped back).
Provider-native 0-1000 boxes do not all use the same array order: Gemini uses
[y0, x0, y1, x1], while Qwen3-VL, Qwen3.5, and Qwen3.6 use
[x0, y0, x1, y1]. Grounding code must select the order by model family (or
an explicit override) before scaling to pixels; never parse every provider as
Gemini-style yxyx.
If several boxes come back numbered, your description matched more than one element rather than picking out a single thing. Narrow it with what distinguishes the one you mean — its text, its position, the block it sits in — and ask again.
The box is a handle, not just an answer — it feeds the next call:
$ ground screenshot.png "the send button"
x1: 1067, y1: 841, x2: 1108, y2: 881
$ glance screenshot.png --region 1067,841,1108,881 -q "is it enabled or greyed out?"That two-step is how you inspect anything too small to survive a full-image pass.
detect <image> # every UI element
detect <image> "buttons" # one kind only
detect <image> --region X1,Y1,X2,Y2 # inside one box
detect <image> "text" --fail-on-empty # fail if no text is detectedThe default category targets UI screenshots. For photographs or video frames,
pass an explicit category such as "objects" or "text". An empty inventory
normally prints no elements detected and exits 0; --fail-on-empty instead
exits 1 with a stderr diagnostic and no stdout. This reports an empty model
result, not proof that the image has no matching elements.
You name a particular thing for ground; you name a kind for detect and
it enumerates the instances. Output is a numbered list with each item's
visible text and box. A full-screen
pass is a fast first draft — counts vary run to run on dense screens. For
completeness, detect the layout blocks first, then detect --region each
block.
trace <image> # b/w spline SVG to stdout
trace <image> --polygon # boxy diagrams/wireframes
trace <image> --region X1,Y1,X2,Y2 -o out.svg # crop firstCoordinates come from the actual pixels, not a model's estimate. Flat,
high-contrast graphics only; text becomes curves (pair with --ocr when
the text matters). Small images are upscaled automatically before tracing,
so a 30px icon traces as readily as a screenshot — size is not a reason to
skip the tool. Before shipping or reusing a traced SVG, read
references/restore-graphic.md — it holds the reuse traps and the
ship-vs-hand-write call.
crop <image> --region X1,Y1,X2,Y2 # writes <image-stem>.crop.png next to the input
crop <image> --region X1,Y1,X2,Y2 -o out.png
crop <image> --region X1,Y1,X2,Y2 --scale 4 # upscale the cut-out 4x (LANCZOS) firstThe same X1,Y1,X2,Y2 pixel boxes ground/detect print, clamped to the
image bounds. Once a box is worth keeping — the same crop is about to feed
pixel_diff, dominant_colors, and trace in turn — cut it to a file
once and reuse it, instead of re-cropping in memory on every call.
--scale N upscales the cut-out before writing (default output name becomes
<image-stem>.crop@Nx.png): for icons too small for ground/trace to see
clearly, crop with --scale 4, then run ground/trace on the upscaled
file — coordinates it returns are in the upscaled grid, divide by N to map
back to the original image. Requires the optional pillow.
# manual: you know the region (and optionally the background colour)
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 -o icon.png
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --mode dark # grey/black line logos
python3 scripts/extract_fg.py shot.png --region X1,Y1,X2,Y2 --exclude-color '#E6E6E6'
# auto: `crop --scale` cut-outs with the icon centred — no region needed
crop shot.png --region X1,Y1,X2,Y2 --scale 4 -o d/icon1.png
python3 scripts/extract_fg.py d/icon1.png d/icon2.png # writes <stem>.clean.png next to each input
python3 scripts/extract_fg.py d/icon1.png --disc-radius 60
python3 scripts/extract_fg.py d/icon1.png --boxes "101,84,184,171"Manual mode keeps every sufficiently large connected component of the
region (separate logo sub-shapes stay together; specks drop out). Auto mode
takes a crop --scale cut-out with the icon centred (disc + glyph): the
disc centre is the image centre, the disc radius defaults to
min(w,h)/2 * 0.6, and the disc colour is sampled from a ring around the
centre; that colour is excluded and the glyph is picked as the most
saturated among the three largest coloured components (white rings,
ripples, and text fall away), output as a 1:1 transparent PNG. When auto
inference fails, override the radius with --disc-radius, or pass a
ground box (in the upscaled grid) as --boxes to recentre and re-filter
by overlap. Multiple images may be passed at once (auto mode).
Requires the optional pillow (and numpy for auto mode).
python3 scripts/html_shot.py page.html # writes page.png, 1280x800
python3 scripts/html_shot.py page.html --width 1440 --height 900 -o page.png
python3 scripts/html_shot.py page.html --scale 2 # 2x pixels: small text stays readable
python3 scripts/html_shot.py page.html --full-page # complete scroll height, same layout viewport
python3 scripts/html_shot.py page.html --full-page --max-pixels 40000000The visual-alignment loop: write HTML, screenshot it at the reference
viewport, then compare it with the design. Use pixel_diff to locate
material differences, not to chase a zero-difference score. Rendering
happens in headless Chrome/Chromium/Edge — no Python dependencies. The default
captures only the viewport. Use --full-page for the complete document while
keeping --width and --height as the layout viewport, so vh/svh and
responsive breakpoints do not change. Add --max-pixels N when the page height
is untrusted. --wait-ms N pauses for fonts, images, or animation before
capturing. Paths are relative to this skill's own directory.
python3 scripts/pixel_diff.py <a> <b> # path is relative to this skill dirPrints an overall difference percentage plus the worst regions as x1: ..
boxes you can feed straight into glance --region. Exact where a vision
model rounds off.
python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2 # top colour clusters + shares
python3 scripts/dominant_colors.py <image> --region X1,Y1,X2,Y2 \
--candidates '#F9FAFA,#F5F5F5,#F3F3F3,#EDEDED' # pick the best candidateA vision model names a colour ("light gray") but not its value. The first
mode downsamples, quantizes, and merges near-duplicates to list the region's
significant colours with the share each owns — the histogram shows which
colour is the background and which is the accent. Given the candidate palette
your label implies, the second mode scores each candidate by how close the
region's pixels are to it and prints the winner. Take the value from here,
never from glance's prose. Paths are relative to this skill's own
directory.
If the image lives in a temp directory, before your first tool call on one, copy it somewhere durable and run everything against the copy — that is what keeps the image reachable later:
cp "<the temp path>" work/shot.png
glance work/shot.png -q "..."Exception: the user asked for the image to stay in a temp folder.
If an image reached you only as text — a description written by a person, a tool, or another model — and the image's file path is visible in the conversation, do not reason past a missing detail. Look again yourself:
glance <path> -q "<the specific detail>" — one qualitative follow-up.ground <path> "<target>" then glance <path> --region <that box> -q "..." —
locate, then zoom. The reliable way to inspect one element closely.If the file no longer exists, say so instead of guessing.
For a single question about an image, glance is the whole answer. For
anything multi-step, work outside-in:
glance, or a description you already have) for
the layout and an inventory of what is where.ground it, then zoom with
glance --region <box> -q "...". Full-image passes routinely miss small
text and icons; a crop puts all the pixels on one detail, so the model
sees it at effectively higher resolution. When the same box will be
checked more than once, cut it to a file first with crop.trace, from a ground box, or
from pixel_diff; sample the pixels yourself only for what those cannot
return.Each file below is one job, start to finish: when it applies, the call sequence, and how to tell you got it right.
| The job | Read |
|---|---|
| OCR a long screenshot, scrolling page, or chat history without losing text at chunk boundaries | references/long-screenshot-ocr.md |
| Rebuild a page or component as HTML/CSS, including a roughly three-minute fast approximation mode, or align an existing UI with its reference image | references/restore-ui.md |
| Extract or rebuild an icon, logo, illustration, or other isolated graphic as transparent PNG/SVG | references/restore-graphic.md |
| Turn a sketch, diagram, or whiteboard into Mermaid, Graphviz, or another structured representation | references/restore-structure.md |
| Operate a GUI from screenshots — locate, act, verify each step | references/gui.md |
Source repository: https://github.com/Anionex/agent-vision-toolkit
Installation guide: https://github.com/Anionex/agent-vision-toolkit/blob/main/AGENT_INSTALL.md
© Anionex, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in skills/vision-skills of Anionex/agent-vision-toolkit.
Open the folder on GitHubat commit 1384ef4
Vision Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vision Skills this skillAnionex/agent-vision-toolkit | 1.2k | — | ~4k | Automated safety check: Pass | MIT | |
| Computer UseTheSyart/emperor-agent | 195 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Mac Computer UseTo3akaRin/mac-computer-use | 1.1k | — | ~495 | Automated safety check: Pass | MIT | |
| Manage Taskboardshengsheng90/DSH-taskboard | 332 | — | ~885 | Automated safety check: Pass | Apache-2.0 | |
| Cloud Computer Usedavidondrej/cloudroom-core | 285 | — | ~881 | Automated safety check: Notes | Apache-2.0 | |
| Verify Reelrselbach/reel | 118 | — | ~2k | Automated safety check: Pass | Unlicense |
TheSyart/emperor-agent
Operate web pages in Emperor's built-in browser with the browser tools, and macOS app windows with the desktop tools — open or bind a target, read it, fill, click or type, and verify the result.
To3akaRin/mac-computer-use
操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。
shengsheng90/DSH-taskboard
Manage work in the native DeepSeek Harness Taskboard with exact task ids and optimistic versions.
davidondrej/cloudroom-core
See and control desktop apps on this Cloud sandbox’s virtual Linux screen with cloudroom computer-use: launch GUI apps you build or install, read their UI, click, type, and take screenshots.
rselbach/reel
Verify Reel, the macOS menu-bar screen recorder, by launching a disposable app bundle and driving its real UI with Computer Use.
KevPH2026/muse-catch
Muse · Catch — AI 灵感捕手。浏览器插件 + 任意 Agent(Telegram/WhatsApp/Discord…)一键捕获灵感,AI自动提炼,Web Dashboard浏览。分级LLM路由:Agent内置LLM隐私分析 + TokenRouter云端分布式调用。当用户说「灵感」「捕获」「Muse」「Catch」「记下来」「收藏」「稍后读」时自动触发。
Works with
Categories
Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…. Vision Skills is an agent skill from Anionex/agent-vision-toolkit.py (HTML file to image).
Vision Skills fits situations like: any task involving an image — questions; splitting and transcribing long screenshots; locating elements; rebuilding as HTML/SVG.
Run `npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a claude-code`. Or copy the skill folder (skills/vision-skills in Anionex/agent-vision-toolkit) into .claude/skills/vision-skills in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a codex`. Or copy the skill folder (skills/vision-skills in Anionex/agent-vision-toolkit) into .agents/skills/vision-skills in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision-skills, .gemini/skills/vision-skills, .github/skills/vision-skills and .opencode/skills/vision-skills in your project.
Going by SKILL.md and its folder, Vision Skills needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named VISION_API_KEY. Our summary lists: Python 3; A credential in VISION_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Vision Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Vision Skills: Computer Use (TheSyart/emperor-agent, 195 stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars), Manage Taskboard (shengsheng90/DSH-taskboard, 332 stars) and Cloud Computer Use (davidondrej/cloudroom-core, 285 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Anionex (a GitHub user) maintains it in Anionex/agent-vision-toolkit, which has 1,220 GitHub stars. The repository was last updated on October 8, 2026.
Source: Anionex/agent-vision-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.