Fal Vision
nexu-io/open-design
Analyze images — segment objects, detect, run OCR, describe, and answer visual questions via fal.ai vision models.
Augmented vision tools for analyzing images beyond native visual capabilities.
$ npx skills add oaustegard/claude-skills --skill seeing-images -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install oaustegard/claude-skills seeing-images --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/seeing-images .claude/skills/seeing-images && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .claude/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/oaustegard/claude-skills/tree/main/seeing-imagesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add oaustegard/claude-skills --skill seeing-images -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install oaustegard/claude-skills seeing-images --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/seeing-images .agents/skills/seeing-images && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .agents/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oaustegard/claude-skills --skill seeing-images -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install oaustegard/claude-skills seeing-images --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/seeing-images .cursor/skills/seeing-images && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .cursor/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/oaustegard/claude-skills.git --path seeing-images--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add oaustegard/claude-skills --skill seeing-images -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install oaustegard/claude-skills seeing-images --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/seeing-images .gemini/skills/seeing-images && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .gemini/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install oaustegard/claude-skills seeing-imagesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add oaustegard/claude-skills --skill seeing-images -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/seeing-images .github/skills/seeing-images && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .github/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add oaustegard/claude-skills --skill seeing-images -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install oaustegard/claude-skills seeing-images --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/seeing-images .opencode/skills/seeing-images && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "seeing-images" agent skill from https://github.com/oaustegard/claude-skills/tree/main/seeing-images into .opencode/skills/seeing-images/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "seeing-images", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
seeing-imagesAugmented vision tools for analyzing images beyond native visual capabilities.
Seeing Images is an agent skill from oaustegard/claude-skills. Augmented vision tools for analyzing images beyond native visual capabilities. Use when tasked with describing images in detail, reproducing images as SVGs, identifying subtle features, comparing image regions, reading degraded text, or any task requiring careful visual inspection. Also use when the image-to-svg skill needs ground truth about colors, shapes, or boundaries.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `CHANGELOG.md` and `scripts/see.py`).
The repository describes itself as: My collection of Claude skills. The licence is MIT.
Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Seeing Images loads about 1.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 532 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 532 words, ~1,372 tokens.
.claude/skills/seeing-images/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Compensatory vision tools based on blindspots measured by vision diagnostic v1-v4 on 2026-03-25. Which model it ran on is not recorded here and it has not been re-run since, so read the thresholds as measured-then: they say a tool exists for each failure, not that the current model fails at exactly that number. Re-run the diagnostic before relying on a specific threshold.
Activate this skill when:
Measured, not guessed:
| Blindspot | Threshold | Compensatory Tool |
|---|---|---|
| Luminance contrast | ~15-20 RGB steps invisible | enhance, histogram, sample |
| Gradients | <30-step range invisible | gradient_map, enhance |
| Context color bias | Dress effect, simultaneous contrast | isolate, sample |
| Small elements | <15px effectively invisible | crop, grid |
| Dense counting | Degrades >15 items, ~50% error at 30 | count_elements |
| Subtle atmospherics | Steam, faint reflections lost in noise | enhance, denoise |
import sys; sys.path.insert(0, '/mnt/skills/user/seeing-images/scripts')
from see import grid, sample, enhance, edges, histogram, isolate, palette, compare, count_elements, gradient_map, denoise, cropgrid(path, rows=2, cols=2) # → view the output
sample(path, [(x1,y1), ...]) # → verify colors at points of interestgrid(path, rows=3, cols=3) # 1. Overview
palette(path, n=10) # 2. Dominant colors
edges(path, threshold=30) # 3. Shape boundaries
sample(path, [(x1,y1), (x2,y2), ...]) # 4. Exact RGB at points
enhance(path, region=(x,y,w,h), mode='auto') # 5. Reveal low-contrast areas
isolate(path, region=(x,y,w,h)) # 6. Remove context biasAll functions in scripts/see.py. Every function that produces an image saves to /home/claude/see_*.png and returns the path. Use view tool on the returned path.
Splits image into labeled cells for systematic inspection. Call it first: it reduces attentional competition.
Returns exact RGB values at specified pixel coordinates. Use to verify what you think you see. Averages over a small radius to handle noise.
Color histogram showing value distribution. Reveals bimodal distributions (hidden gradients), dominant colors, and contrast range. With region=(x,y,w,h), analyzes only that area.
Boosts contrast in the image or a region. Modes: 'contrast', 'brightness', 'color', 'sharpness'. Use factor=3-5 for near-threshold features.
Sobel edge detection revealing shape boundaries invisible at low contrast. Lower threshold = more edges (noisier). Output is a white-on-black edge map.
Computes local gradient magnitude across the image. Bright = high gradient, dark = flat. Reveals gradients below the 30-step detection threshold.
Extracts a region and places it on a neutral gray background. Removes surrounding context that causes simultaneous contrast and Dress-type illusions. The bg parameter defaults to mid-gray to minimize context bias.
Side-by-side comparison of two regions with diff overlay. Highlights pixel-level differences with amplification. Use for spot-the-difference tasks.
Programmatic element counting using connected component analysis. Specify approximate color_range as ((r_min,g_min,b_min), (r_max,g_max,b_max)) to count specific colored elements.
Median filter to reduce photographic noise, revealing subtle features hidden in the noise floor (like steam, faint reflections).
Extracts the n most dominant colors using k-means clustering. Returns RGB values and their proportions. Essential for SVG reproduction.
Call grid() first on a complex image. Verify colors near context boundaries
with sample() or isolate(), counts above 15 with count_elements(),
gradients with gradient_map(), and faint features with enhance() before
describing them. Each of these is a row of the blindspot table, so the tool
call is the evidence — perception alone is not.
© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in seeing-images of oaustegard/claude-skills.
Open the folder on GitHubat commit 90b0f1b
Seeing Images next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Seeing Images this skilloaustegard/claude-skills | 150 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Fal Visionnexu-io/open-design | 100k | — | ~295 | Automated safety check: Pass | Apache-2.0 | |
| Vision Sftwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Visiongridaco/grida | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Senior Computer Visiondavila7/claude-code-templates | 33k | 2 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Agent Code Analyzerruvnet/ruflo | 74k | 2 repos | ~1.5k | Automated safety check: Pass | MIT |
nexu-io/open-design
Analyze images — segment objects, detect, run OCR, describe, and answer visual questions via fal.ai vision models.
wshobson/agents
Fine-tune vision-language models (VLMs) with supervised learning on image+text data.
gridaco/grida
Query images with a local Ollama vision model without loading the image into the main agent context.
davila7/claude-code-templates
World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems.
ruvnet/ruflo
Agent skill for code-analyzer - invoke with $agent-code-analyzer
ruvnet/ruflo
Agent skill for pagerank-analyzer - invoke with $agent-pagerank-analyzer
oaustegard/claude-skills
Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.
oaustegard/claude-skills
Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.
oaustegard/claude-skills
Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…
oaustegard/claude-skills
Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.
oaustegard/claude-skills
Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.
oaustegard/claude-skills
Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.
Augmented vision tools for analyzing images beyond native visual capabilities. Seeing Images is an agent skill from oaustegard/claude-skills. Augmented vision tools for analyzing images beyond native visual capabilities.
Seeing Images fits situations like: tasked with describing images in detail; reproducing images as SVGs; identifying subtle features; comparing image regions.
Run `npx skills add oaustegard/claude-skills --skill seeing-images -a claude-code`. Or copy the skill folder (seeing-images in oaustegard/claude-skills) into .claude/skills/seeing-images in your project. Claude Code loads it when a task matches its description.
Run `npx skills add oaustegard/claude-skills --skill seeing-images -a codex`. Or copy the skill folder (seeing-images in oaustegard/claude-skills) into .agents/skills/seeing-images in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill seeing-images -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/seeing-images, .gemini/skills/seeing-images, .github/skills/seeing-images and .opencode/skills/seeing-images in your project.
Going by SKILL.md and its folder, Seeing Images needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Seeing Images is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Seeing Images: Fal Vision (nexu-io/open-design, 100k stars), Vision Sft (wshobson/agents, 40k stars), Vision (gridaco/grida, 2.7k stars) and Senior Computer Vision (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.
Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.