Teach A Model
Aseiel/VideoHighlighter
Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.
Query images with a local Ollama vision model without loading the image into the main agent context.
$ npx skills add gridaco/grida --skill vision -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install gridaco/grida vision --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vision .claude/skills/vision && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .claude/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/gridaco/grida/tree/main/.agents/skills/visionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add gridaco/grida --skill vision -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install gridaco/grida vision --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/vision .agents/skills/vision && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .agents/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gridaco/grida --skill vision -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install gridaco/grida vision --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/vision .cursor/skills/vision && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .cursor/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/gridaco/grida.git --path .agents/skills/vision--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add gridaco/grida --skill vision -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install gridaco/grida vision --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/vision .gemini/skills/vision && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .gemini/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install gridaco/grida visionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add gridaco/grida --skill vision -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/vision .github/skills/vision && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .github/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gridaco/grida --skill vision -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install gridaco/grida vision --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gridaco/grida.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/vision .opencode/skills/vision && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vision" agent skill from https://github.com/gridaco/grida/tree/main/.agents/skills/vision into .opencode/skills/vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
visionQuery images with a local Ollama vision model without loading the image into the main agent context.
Vision is an agent skill from gridaco/grida. Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (gemma4 on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check…
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/ask.py`).
It sits in AI & LLM Engineering, covering LLM inference and serving, Computer vision and Codebase knowledge for agents. It works with Ollama. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 165496f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvollamapipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
ollama.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vision loads about 1.5k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 478 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from gridaco/grida at commit 165496f, republished under its Apache-2.0 licence (© gridaco). 478 words, ~1,473 tokens.
.claude/skills/vision/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Ask natural-language questions about images without passing them to the main agent as visual input. Useful for verifying screenshots, annotating assets, or building automated checks around visual output.
All commands use uv run — dependencies are installed automatically.
SCRIPT=.agents/skills/vision/scripts/ask.py
# health check (fast, no image, confirms Ollama + model respond)
uv run $SCRIPT --ping
# system info — memory, storage, installed models
uv run $SCRIPT --info
uv run $SCRIPT --memory
uv run $SCRIPT --storage
# describe an image (default prompt)
uv run $SCRIPT path/to/image.png
# explicit shortcut
uv run $SCRIPT path/to/image.png describe
# custom question
uv run $SCRIPT path/to/image.png \
--prompt "Do you see any overlapping UI elements?"
uv run $SCRIPT canvas.png \
--prompt "Does this canvas contain any designed content, or is it empty?"
# optional: pin a specific Gemma 4 tag (default is any installed gemma4)
uv run $SCRIPT image.png --model gemma4:e4b
# list installed Gemma 4 vision models
uv run $SCRIPT --list-modelsOllama must be running locally. The script connects to http://localhost:11434
and fails immediately if it cannot reach it.
# start Ollama (if not already running)
ollama serve
# install Gemma 4 (multimodal — required for this skill)
ollama pull gemma4The script does not install models. If Gemma 4 is not installed it prints
the list of installed models and a pull suggestion, then exits.
uv is required to run the script (handles dependency installation
automatically). No requirements.txt or manual pip install needed.
This skill uses only Gemma 4 on
Ollama (gemma4 and tags such as gemma4:latest, gemma4:e4b). Other
multimodal models are ignored so agents do not silently fall back to a
different family.
When --model is omitted, the script picks any installed gemma4 tag (for
example gemma4:latest). Use --model gemma4:e4b (or another tag) to pin a
specific variant.
Before running a heavy query, check whether the machine has enough resources. This is optional — the script does not enforce limits — but useful context for deciding whether to proceed or skip.
uv run $SCRIPT --info # memory + storage + model list
uv run $SCRIPT --memory # just memory
uv run $SCRIPT --storage # just storageTip: on machines with ≤8 GB RAM, large vision models may cause swapping or
OOM. Consider a smaller Gemma 4 variant (for example gemma4:e2b) or skip
the query.
hint or pull command.ask.py
in parallel (e.g. two concurrent tool calls). Queue calls one at a time.uv inline script metadata (PEP 723). Only
dependency is the ollama Python package..png, .jpg, .jpeg, .webp, .gif, .bmp.ask.py with a targeted prompt suited to the task.# Quick sanity check first
uv run $SCRIPT --ping
# Verify a browser screenshot has content before including it in a doc
uv run $SCRIPT /tmp/preview.png \
--prompt "Answer with YES or NO: does this screenshot show any visible UI content, shapes, or text?"
# Describe a captured screenshot for a PR description
uv run $SCRIPT /tmp/canvas-screenshot.png \
--prompt "Describe what visual effect is shown. Be specific about blur, colors, and shapes."| Symptom | Cause | Fix |
|---|---|---|
cannot reach Ollama | Ollama not running | ollama serve |
no Gemma 4 vision model found | Gemma 4 not installed | ollama pull gemma4 |
model 'X' is not available | Model name typo or not installed | --list-models to see what's installed |
| Slow response | Large model on CPU | Try a smaller tag (e.g. gemma4:e2b) |
| Vague or wrong answer | Generic prompt | Write a more specific --prompt |
'ollama' package not found | Not using uv run | Run with uv run ask.py instead |
© gridaco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in .agents/skills/vision of gridaco/grida.
Open the folder on GitHubat commit 165496f
Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vision this skillgridaco/grida | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Teach A ModelAseiel/VideoHighlighter | 157 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | |
| Subwave LLM Benchperminder-klair/subwave | 1.4k | — | ~2.4k | Automated safety check: Notes | MIT | |
| Pii Safe Documentsdanyuchn/pii-guard | 249 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Add Vlm Modelintel/auto-round | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool | 240 | — | ~17k | Automated safety check: Pass | None |
Aseiel/VideoHighlighter
Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.
perminder-klair/subwave
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…
danyuchn/pii-guard
Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.
intel/auto-round
Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.
darknecrocities/DomoDomo---All-in-one-Tool
Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.
mathruffian-dot/claude-code-lazy-packs
Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.
gridaco/grida
Grida Desktop Electron shell and release-impact work: BrowserWindow, preload, window.grida, menus, protocol/deep links, file associations, Forge, path-scoped bridge security, Electron-only UI bugs…
gridaco/grida
Guides work on the Figma I/O package (@grida/io-figma, packages/grida-canvas-io-figma/).
gridaco/grida
Set up, download, verify, and seed the optional Grida Library developer corpus into local Supabase.
gridaco/grida
Research, compare, and update shared AI model JSON for TypeScript, web, and Rust consumers.
gridaco/grida
Grida AI agent system work: @grida/daemon (DaemonServer, loopback HTTP perimeter, files/workspaces, secrets store, daemon discovery) and @grida/agent (the agent tenant: sessions, providers/BYOK…
gridaco/grida
Use BEFORE editing any file in supabase/migrations/ or supabase/schemas/, OR when the user runs a /database subcommand (compact local migration, rls scenarios, align).
Works with
Categories
Query images with a local Ollama vision model without loading the image into the main agent context. Vision is an agent skill from gridaco/grida. Query images with a local Ollama vision model without loading the image into the main agent context.
Vision fits situations like: you need to describe a screenshot; check whether rendered content is present; detect overlapping elements; ask any visual question about a PNG/JPEG/WebP file.
Run `npx skills add gridaco/grida --skill vision -a claude-code`. Or copy the skill folder (.agents/skills/vision in gridaco/grida) into .claude/skills/vision in your project. Claude Code loads it when a task matches its description.
Run `npx skills add gridaco/grida --skill vision -a codex`. Or copy the skill folder (.agents/skills/vision in gridaco/grida) into .agents/skills/vision in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gridaco/grida --skill vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision, .gemini/skills/vision, .github/skills/vision and .opencode/skills/vision in your project.
Going by SKILL.md and its folder, Vision needs Python for the scripts in its folder and the command-line tools its instructions call (uv, ollama and pip). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: ollama.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Vision is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vision: Teach A Model (Aseiel/VideoHighlighter, 157 stars), Subwave LLM Bench (perminder-klair/subwave, 1.4k stars), Pii Safe Documents (danyuchn/pii-guard, 249 stars) and Add Vlm Model (intel/auto-round, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
gridaco (a GitHub organization) maintains it in gridaco/grida, which has 2,657 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.
Source: gridaco/grida on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.