Agent skill

Vision

by gridaco in gridaco/grida

Query images with a local Ollama vision model without loading the image into the main agent context.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vision

skills CLI
$ npx skills add gridaco/grida --skill vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gridaco/grida vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vision .claude/skills/vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vision
GitHub stars
2.7k
Token cost
~1.5k tokens
SKILL.md length
478 words
Files
2 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Query images with a local Ollama vision model without loading the image into the main agent context.

  • Works in 3 steps: A tool (browser automation, screenshot… → Call ask.py with a targeted prompt… → Parse the text response to decide the…
  • You need to describe a screenshot
  • SKILL.md covers When to Use This Skill, Quick Reference, Prerequisites and Model Selection, plus 4 more sections
  • Runs Python scripts from its folder; calls uv, ollama and pip

What it does

Vision is an agent skill from gridaco/grida. Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (gemma4 on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/ask.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Computer vision and Codebase knowledge for agents. It works with Ollama. The licence is Apache-2.0.

When your agent uses it

  • You need to describe a screenshot
  • Check whether rendered content is present
  • Detect overlapping elements
  • Ask any visual question about a PNG/JPEG/WebP file

Example prompts

  • “describe image”
  • “what does this screenshot show”
  • “does the canvas contain content”
  • “/vision”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A tool (browser automation, screenshot capture, golden renderer) writes
  2. Call ask.py with a targeted prompt suited to the task.
  3. Parse the text response to decide the next action.

What it can do on your machine

Read from SKILL.md and the folder at commit 165496f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • ollama
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ollama.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vision loads about 1.5k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 478 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gridaco/grida at commit 165496f, republished under its Apache-2.0 licence (© gridaco). 478 words, ~1,473 tokens.

Download SKILL.mdSave it as .claude/skills/vision/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vision
description
Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (`gemma4` on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check screenshot visually", "look at this image", "any overlapping elements", "vision query".

Vision — Local Image Querying via Ollama

Ask natural-language questions about images without passing them to the main agent as visual input. Useful for verifying screenshots, annotating assets, or building automated checks around visual output.

When to Use This Skill

  • Describing a screenshot for a PR description or user-facing document
  • Checking whether an automated browser run produced visible canvas content
  • Asking "do any elements overlap?" on a rendered output
  • Any question where the answer is in the pixels but you don't want to use vision tokens in the main context

Quick Reference

All commands use uv run — dependencies are installed automatically.

sh
SCRIPT=.agents/skills/vision/scripts/ask.py

# health check (fast, no image, confirms Ollama + model respond)
uv run $SCRIPT --ping

# system info — memory, storage, installed models
uv run $SCRIPT --info
uv run $SCRIPT --memory
uv run $SCRIPT --storage

# describe an image (default prompt)
uv run $SCRIPT path/to/image.png

# explicit shortcut
uv run $SCRIPT path/to/image.png describe

# custom question
uv run $SCRIPT path/to/image.png \
  --prompt "Do you see any overlapping UI elements?"

uv run $SCRIPT canvas.png \
  --prompt "Does this canvas contain any designed content, or is it empty?"

# optional: pin a specific Gemma 4 tag (default is any installed gemma4)
uv run $SCRIPT image.png --model gemma4:e4b

# list installed Gemma 4 vision models
uv run $SCRIPT --list-models

Prerequisites

Ollama must be running locally. The script connects to http://localhost:11434 and fails immediately if it cannot reach it.

sh
# start Ollama (if not already running)
ollama serve

# install Gemma 4 (multimodal — required for this skill)
ollama pull gemma4

The script does not install models. If Gemma 4 is not installed it prints the list of installed models and a pull suggestion, then exits.

uv is required to run the script (handles dependency installation automatically). No requirements.txt or manual pip install needed.


Model Selection

This skill uses only Gemma 4 on Ollama (gemma4 and tags such as gemma4:latest, gemma4:e4b). Other multimodal models are ignored so agents do not silently fall back to a different family.

When --model is omitted, the script picks any installed gemma4 tag (for example gemma4:latest). Use --model gemma4:e4b (or another tag) to pin a specific variant.


System Info

Before running a heavy query, check whether the machine has enough resources. This is optional — the script does not enforce limits — but useful context for deciding whether to proceed or skip.

sh
uv run $SCRIPT --info      # memory + storage + model list
uv run $SCRIPT --memory    # just memory
uv run $SCRIPT --storage   # just storage

Tip: on machines with ≤8 GB RAM, large vision models may cause swapping or OOM. Consider a smaller Gemma 4 variant (for example gemma4:e2b) or skip the query.


Show full SKILL.md (197 more words)Show less

Behavior

  • Fails fast if Ollama is unreachable or Gemma 4 is not installed. Exit code is non-zero; the error message includes a hint or pull command.
  • Sequential only — Ollama is a single-worker process. Never call ask.py in parallel (e.g. two concurrent tool calls). Queue calls one at a time.
  • No side effects beyond the local Ollama process.
  • Auto-installs deps via uv inline script metadata (PEP 723). Only dependency is the ollama Python package.
  • Supported formats: .png, .jpg, .jpeg, .webp, .gif, .bmp.

Typical Agent Workflow

  1. A tool (browser automation, screenshot capture, golden renderer) writes an image to disk.
  2. Call ask.py with a targeted prompt suited to the task.
  3. Parse the text response to decide the next action.
sh
# Quick sanity check first
uv run $SCRIPT --ping

# Verify a browser screenshot has content before including it in a doc
uv run $SCRIPT /tmp/preview.png \
  --prompt "Answer with YES or NO: does this screenshot show any visible UI content, shapes, or text?"

# Describe a captured screenshot for a PR description
uv run $SCRIPT /tmp/canvas-screenshot.png \
  --prompt "Describe what visual effect is shown. Be specific about blur, colors, and shapes."

Troubleshooting

SymptomCauseFix
cannot reach OllamaOllama not runningollama serve
no Gemma 4 vision model foundGemma 4 not installedollama pull gemma4
model 'X' is not availableModel name typo or not installed--list-models to see what's installed
Slow responseLarge model on CPUTry a smaller tag (e.g. gemma4:e2b)
Vague or wrong answerGeneric promptWrite a more specific --prompt
'ollama' package not foundNot using uv runRun with uv run ask.py instead

© gridaco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .agents/skills/vision of gridaco/grida.

  • SKILL.md
  • scripts/ask.py

Open the folder on GitHubat commit 165496f

Compare with similar skills

Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vision this skillgridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Teach A ModelAseiel/VideoHighlighter157—~839Automated safety check: PassAGPL-3.0
Subwave LLM Benchperminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMIT
Pii Safe Documentsdanyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT
Add Vlm Modelintel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0
Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool240—~17kAutomated safety check: PassNone

Similar skills

  • Teach A Model

    Aseiel/VideoHighlighter

    Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

    157 GitHub stars~839 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Subwave LLM Bench

    perminder-klair/subwave

    Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

    1.4k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Pii Safe Documents

    danyuchn/pii-guard

    Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

    249 GitHub stars~2.7k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Domodomo Local AI Maintenance

    darknecrocities/DomoDomo---All-in-one-Tool

    Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

    240 GitHub stars~17k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cc Ollama

    mathruffian-dot/claude-code-lazy-packs

    Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

    253 GitHub stars~118 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from gridaco/grida

All 29 skills in this repo
  • Desktop

    gridaco/grida

    Grida Desktop Electron shell and release-impact work: BrowserWindow, preload, window.grida, menus, protocol/deep links, file associations, Forge, path-scoped bridge security, Electron-only UI bugs…

    2.7k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check: notes
  • Io Figma

    gridaco/grida

    Guides work on the Figma I/O package (@grida/io-figma, packages/grida-canvas-io-figma/).

    2.7k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Opt Library

    gridaco/grida

    Set up, download, verify, and seed the optional Grida Library developer corpus into local Supabase.

    2.7k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • AI Models

    gridaco/grida

    Research, compare, and update shared AI model JSON for TypeScript, web, and Rust consumers.

    2.7k GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Agent System

    gridaco/grida

    Grida AI agent system work: @grida/daemon (DaemonServer, loopback HTTP perimeter, files/workspaces, secrets store, daemon discovery) and @grida/agent (the agent tenant: sessions, providers/BYOK…

    2.7k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Database

    gridaco/grida

    Use BEFORE editing any file in supabase/migrations/ or supabase/schemas/, OR when the user runs a /database subcommand (compact local migration, rls scenarios, align).

    2.7k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Vision

What does Vision do?

Query images with a local Ollama vision model without loading the image into the main agent context. Vision is an agent skill from gridaco/grida. Query images with a local Ollama vision model without loading the image into the main agent context.

When should I use Vision?

Vision fits situations like: you need to describe a screenshot; check whether rendered content is present; detect overlapping elements; ask any visual question about a PNG/JPEG/WebP file.

How do I install Vision in Claude Code?

Run `npx skills add gridaco/grida --skill vision -a claude-code`. Or copy the skill folder (.agents/skills/vision in gridaco/grida) into .claude/skills/vision in your project. Claude Code loads it when a task matches its description.

How do I install Vision in Codex?

Run `npx skills add gridaco/grida --skill vision -a codex`. Or copy the skill folder (.agents/skills/vision in gridaco/grida) into .agents/skills/vision in your project. Codex loads it when a task matches its description.

Can I use Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gridaco/grida --skill vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision, .gemini/skills/vision, .github/skills/vision and .opencode/skills/vision in your project.

What does Vision need to run?

Going by SKILL.md and its folder, Vision needs Python for the scripts in its folder and the command-line tools its instructions call (uv, ollama and pip). Our summary lists: Python 3.

Does Vision access the network?

SKILL.md names 1 domain. As links in the text: ollama.com. This is read from the text; nothing was executed.

Is Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vision use?

Vision is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vision use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vision?

Skills that share tags, products or a category with Vision: Teach A Model (Aseiel/VideoHighlighter, 157 stars), Subwave LLM Bench (perminder-klair/subwave, 1.4k stars), Pii Safe Documents (danyuchn/pii-guard, 249 stars) and Add Vlm Model (intel/auto-round, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vision?

gridaco (a GitHub organization) maintains it in gridaco/grida, which has 2,657 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.

Source: gridaco/grida on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.