Agent skill

Seeing Images

by oaustegard in oaustegard/claude-skills

Augmented vision tools for analyzing images beyond native visual capabilities.

MITAuto-check passed

Install Seeing Images

skills CLI
$ npx skills add oaustegard/claude-skills --skill seeing-images -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills seeing-images --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/seeing-images .claude/skills/seeing-images && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
seeing-images
GitHub stars
150
Token cost
~1.4k tokens
SKILL.md length
532 words
Files
3 (incl. scripts)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Augmented vision tools for analyzing images beyond native visual capabilities.

  • Tasked with describing images in detail
  • SKILL.md covers When to Use, Known Blindspots (from…, Workflow and Tool Reference, plus 1 more section
  • Runs Python scripts from its folder
  • Reproducing images as SVGs

What it does

Seeing Images is an agent skill from oaustegard/claude-skills. Augmented vision tools for analyzing images beyond native visual capabilities. Use when tasked with describing images in detail, reproducing images as SVGs, identifying subtle features, comparing image regions, reading degraded text, or any task requiring careful visual inspection. Also use when the image-to-svg skill needs ground truth about colors, shapes, or boundaries.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `CHANGELOG.md` and `scripts/see.py`).

The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Tasked with describing images in detail
  • Reproducing images as SVGs
  • Identifying subtle features
  • Comparing image regions

Example prompts

  • “/seeing-images”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Seeing Images loads about 1.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 532 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 532 words, ~1,372 tokens.

Download SKILL.mdSave it as .claude/skills/seeing-images/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
seeing-images
description
Augmented vision tools for analyzing images beyond native visual capabilities. Use when tasked with describing images in detail, reproducing images as SVGs, identifying subtle features, comparing image regions, reading degraded text, or any task requiring careful visual inspection. Also use when the image-to-svg skill needs ground truth about colors, shapes, or boundaries.
metadata.version
1.0.1

Seeing Images

Compensatory vision tools based on blindspots measured by vision diagnostic v1-v4 on 2026-03-25. Which model it ran on is not recorded here and it has not been re-run since, so read the thresholds as measured-then: they say a tool exists for each failure, not that the current model fails at exactly that number. Re-run the diagnostic before relying on a specific threshold.

When to Use

Activate this skill when:

  • Describing an uploaded image in detail
  • Reproducing an image as SVG (use BEFORE drawing to establish ground truth)
  • Comparing two images or regions for differences
  • Reading text in degraded/compressed/low-contrast images
  • Identifying subtle features (gradients, faint overlays, reflections)
  • Any image task where accuracy matters more than speed

Known Blindspots (from diagnostics)

Measured, not guessed:

BlindspotThresholdCompensatory Tool
Luminance contrast~15-20 RGB steps invisibleenhance, histogram, sample
Gradients<30-step range invisiblegradient_map, enhance
Context color biasDress effect, simultaneous contrastisolate, sample
Small elements<15px effectively invisiblecrop, grid
Dense countingDegrades >15 items, ~50% error at 30count_elements
Subtle atmosphericsSteam, faint reflections lost in noiseenhance, denoise

Workflow

Setup (one line, every time)
python
import sys; sys.path.insert(0, '/mnt/skills/user/seeing-images/scripts')
from see import grid, sample, enhance, edges, histogram, isolate, palette, compare, count_elements, gradient_map, denoise, crop
Quick Analysis (2-3 tool calls)
python
grid(path, rows=2, cols=2)   # → view the output
sample(path, [(x1,y1), ...]) # → verify colors at points of interest
Deep Analysis (for SVG reproduction, spot-the-difference, etc.)
python
grid(path, rows=3, cols=3)                    # 1. Overview
palette(path, n=10)                           # 2. Dominant colors
edges(path, threshold=30)                     # 3. Shape boundaries
sample(path, [(x1,y1), (x2,y2), ...])        # 4. Exact RGB at points
enhance(path, region=(x,y,w,h), mode='auto')  # 5. Reveal low-contrast areas
isolate(path, region=(x,y,w,h))              # 6. Remove context bias

Tool Reference

All functions in scripts/see.py. Every function that produces an image saves to /home/claude/see_*.png and returns the path. Use view tool on the returned path.

grid(path, rows=3, cols=3, labels=True)

Splits image into labeled cells for systematic inspection. Call it first: it reduces attentional competition.

sample(path, points, radius=3)

Returns exact RGB values at specified pixel coordinates. Use to verify what you think you see. Averages over a small radius to handle noise.

histogram(path, region=None)

Color histogram showing value distribution. Reveals bimodal distributions (hidden gradients), dominant colors, and contrast range. With region=(x,y,w,h), analyzes only that area.

enhance(path, region=None, factor=2.0, mode='contrast')

Boosts contrast in the image or a region. Modes: 'contrast', 'brightness', 'color', 'sharpness'. Use factor=3-5 for near-threshold features.

Show full SKILL.md (218 more words)Show less
edges(path, threshold=50)

Sobel edge detection revealing shape boundaries invisible at low contrast. Lower threshold = more edges (noisier). Output is a white-on-black edge map.

gradient_map(path, region=None)

Computes local gradient magnitude across the image. Bright = high gradient, dark = flat. Reveals gradients below the 30-step detection threshold.

isolate(path, region, padding=20, bg=(128,128,128))

Extracts a region and places it on a neutral gray background. Removes surrounding context that causes simultaneous contrast and Dress-type illusions. The bg parameter defaults to mid-gray to minimize context bias.

compare(path, r1, r2)

Side-by-side comparison of two regions with diff overlay. Highlights pixel-level differences with amplification. Use for spot-the-difference tasks.

count_elements(path, region=None, color_range=None, min_size=3)

Programmatic element counting using connected component analysis. Specify approximate color_range as ((r_min,g_min,b_min), (r_max,g_max,b_max)) to count specific colored elements.

denoise(path, region=None, strength=3)

Median filter to reduce photographic noise, revealing subtle features hidden in the noise floor (like steam, faint reflections).

palette(path, n=8)

Extracts the n most dominant colors using k-means clustering. Returns RGB values and their proportions. Essential for SVG reproduction.

Accuracy Notes

Call grid() first on a complex image. Verify colors near context boundaries with sample() or isolate(), counts above 15 with count_elements(), gradients with gradient_map(), and faint features with enhance() before describing them. Each of these is a row of the blindspot table, so the tool call is the evidence — perception alone is not.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in seeing-images of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/see.py

Open the folder on GitHubat commit 90b0f1b

Compare with similar skills

Seeing Images next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Seeing Images compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Seeing Images this skilloaustegard/claude-skills150—~1.4kAutomated safety check: PassMIT
Fal Visionnexu-io/open-design100k—~295Automated safety check: PassApache-2.0
Vision Sftwshobson/agents40k—~2kAutomated safety check: PassMIT
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Senior Computer Visiondavila7/claude-code-templates33k2 repos~1.4kAutomated safety check: PassMIT
Agent Code Analyzerruvnet/ruflo74k2 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Fal Vision

    nexu-io/open-design

    Analyze images — segment objects, detect, run OCR, describe, and answer visual questions via fal.ai vision models.

    100k GitHub stars~295 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vision Sft

    wshobson/agents

    Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    davila7/claude-code-templates

    World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems.

    33k GitHub starsUsed in 2 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent skill for code-analyzer - invoke with $agent-code-analyzer

    74k GitHub starsUsed in 2 repos~1.5k tokens
    DevelopmentAuto-check passed
  • Agent skill for pagerank-analyzer - invoke with $agent-pagerank-analyzer

    74k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Deciding With Confidence

    oaustegard/claude-skills

    Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

    150 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Seeing Images

What does Seeing Images do?

Augmented vision tools for analyzing images beyond native visual capabilities. Seeing Images is an agent skill from oaustegard/claude-skills. Augmented vision tools for analyzing images beyond native visual capabilities.

When should I use Seeing Images?

Seeing Images fits situations like: tasked with describing images in detail; reproducing images as SVGs; identifying subtle features; comparing image regions.

How do I install Seeing Images in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill seeing-images -a claude-code`. Or copy the skill folder (seeing-images in oaustegard/claude-skills) into .claude/skills/seeing-images in your project. Claude Code loads it when a task matches its description.

How do I install Seeing Images in Codex?

Run `npx skills add oaustegard/claude-skills --skill seeing-images -a codex`. Or copy the skill folder (seeing-images in oaustegard/claude-skills) into .agents/skills/seeing-images in your project. Codex loads it when a task matches its description.

Can I use Seeing Images in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill seeing-images -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/seeing-images, .gemini/skills/seeing-images, .github/skills/seeing-images and .opencode/skills/seeing-images in your project.

What does Seeing Images need to run?

Going by SKILL.md and its folder, Seeing Images needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Seeing Images access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Seeing Images safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Seeing Images use?

Seeing Images is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Seeing Images use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Seeing Images?

Skills that share tags, products or a category with Seeing Images: Fal Vision (nexu-io/open-design, 100k stars), Vision Sft (wshobson/agents, 40k stars), Vision (gridaco/grida, 2.7k stars) and Senior Computer Vision (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Seeing Images?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.