Agent skill

Configuring Vision

by oxbshw in oxbshw/watch-skill

The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

MITAuto-check: notesAI & LLM Engineering

Install Configuring Vision

skills CLI
$ npx skills add oxbshw/watch-skill --skill configuring-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oxbshw/watch-skill configuring-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oxbshw/watch-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/configuring-vision .claude/skills/configuring-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
configuring-vision
GitHub stars
469
Token cost
~509 tokens
SKILL.md length
167 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers Supported providers, Route bulk work and… and No provider is also valid
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Model routing and gateways

What it does

Configuring Vision is an agent skill from oxbshw/watch-skill. The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models. Use this to configure provider-neutral visual understanding without tying Watch Skill to one agent or model vendor.

Its SKILL.md is about 510 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Model routing and gateways. It works with Ollama, OpenAI, OpenRouter and DeepSeek. The repository describes itself as: Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Model routing and gateways

Example prompts

  • “can I use OpenAI/Anthropic/Gemini/OpenRouter”
  • “/configuring-vision”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read

What it can do on your machine

Read from SKILL.md and the folder at commit f1317c8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Configuring Vision loads about 509 tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 167 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~509

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oxbshw/watch-skill at commit f1317c8, republished under its MIT licence (© oxbshw). 167 words, ~509 tokens.

Download SKILL.mdSave it as .claude/skills/configuring-vision/SKILL.md (or your agent's skills folder).
name
configuring-vision
description
The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models. Use this to configure provider-neutral visual understanding without tying Watch Skill to one agent or model vendor.
allowed-tools
Bash, Read
version
1.4.3
license
MIT
user-invocable
true

Configuring vision

Watch Skill's agent surface and model backend are separate choices. Claude Code, Codex, Cursor, OpenClaw, framework agents, and REST clients all call the same engine; the engine can send selected frames to any supported vision provider.

Supported providers

bash
watch-skill setup-vision --provider anthropic --api-key <KEY>
watch-skill setup-vision --provider openai --api-key <KEY>
watch-skill setup-vision --provider gemini --api-key <KEY>
watch-skill setup-vision --provider openrouter --api-key <KEY>
watch-skill setup-vision --provider ollama

Prefer a key the user already has. Do not claim Ollama is required, and do not ask the user to reveal a secret in chat. They can set the matching environment variable or run the command privately in their terminal.

Route bulk work and verification separately

One model can serve both tiers:

bash
watch-skill setup-vision --provider openai --api-key <KEY> --model <vision-model>

Or use a cheaper model for scene descriptions and a stronger model for uncertain answers and loop critiques:

bash
watch-skill setup-vision --provider openrouter --api-key <KEY> \
  --cheap-model <fast-vision-model> --strong-model <strong-vision-model>

Add --verify to make one live probe call. If it fails, report the structured error and its fix; never echo the key.

No provider is also valid

Without a vision API, Watch Skill still acquires video, reads captions, runs local transcription and OCR, indexes evidence, and searches it. Visual synthesis degrades to timestamped evidence instead of guessing.

© oxbshw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/configuring-vision of oxbshw/watch-skill.

Open the folder on GitHubat commit f1317c8

Compare with similar skills

Configuring Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Configuring Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Configuring Vision this skilloxbshw/watch-skill469—~509Automated safety check: NotesMIT
New Providerfinch-xu/cc-router272—~1.6kAutomated safety check: PassMIT
QuorumDetrol/quorum-cli119—~807Automated safety check: NotesCustom licence
Claudish UsageMadAppGang/claudish1k—~9kAutomated safety check: PassNone
Page AgentTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT
Summarize Anythingswyxio/skills175—~6.3kAutomated safety check: PassMIT

Similar skills

  • New Provider

    finch-xu/cc-router

    用于在 cc-router 仓库新增一个 LLM provider(即在 src-tauri/providers/ 下添加 YAML 描述符并完成配套的同步改动)。当用户说「加 provider」「接入 XX 厂商」「新增订阅源」「provider YAML」「让 cc-router 支持 OpenRouter/Together/Groq/Ollama 之类」时必须触发本…

    272 GitHub stars~1.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quorum

    Detrol/quorum-cli

    Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum…

    119 GitHub stars~807 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check: notes
  • Claudish Usage

    MadAppGang/claudish

    CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

    1k GitHub stars~9k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Page Agent

    Tommy-yw/RunbookHermes

    Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

    546 GitHub starsUsed in 3 repos~2.3k tokens
    Productivity & AutomationAuto-check: notes
  • Summarize Anything

    swyxio/skills

    Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend.

    175 GitHub stars~6.3k tokensUpdated 4 days ago
    Writing & ContentAuto-check passed
  • Mesh API

    mr-tbot/mesh-api

    Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.

    180 GitHub stars~1.8k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from oxbshw/watch-skill

All 10 skills in this repo
  • Asking With Evidence

    oxbshw/watch-skill

    The user asks a question about a video that was already watched or indexed — "what did they say about X", "what error code appears", "what happens at 2:30", "does the video show Y".

    469 GitHub stars~550 tokensUpdated 24 days ago
    Auto-check: notes
  • The Loop

    oxbshw/watch-skill

    The user built or changed something visual — a UI, an animation, a game, a generated video — and wants it verified, or asks "why does my UI look wrong", "check that the fix actually worked", "does…

    469 GitHub stars~622 tokensUpdated 24 days ago
    Auto-check: notes
  • Video Memory

    oxbshw/watch-skill

    The user asks about videos watched in the past or across sessions — "have we watched anything about X", "which video showed that error", "what did that meeting decide", "search my videos", or a…

    469 GitHub stars~569 tokensUpdated 24 days ago
    Auto-check: notes
  • Watching Videos

    oxbshw/watch-skill

    The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…

    469 GitHub stars~599 tokensUpdated 24 days ago
    Auto-check: notes
  • Extracting Structure

    oxbshw/watch-skill

    The user wants structure pulled out of a watched video — "make chapters for this video", "where does the bug appear in this recording", "turn this screen recording into a bug report", "how strong is…

    469 GitHub stars~490 tokensUpdated 24 days ago
    Auto-check: notes
  • Learning From Mistakes

    oxbshw/watch-skill

    The user corrected an answer about a video — "no, it actually says X", "that's the wrong timestamp", "you misread the error code" — or asks why a video answer was wrong.

    469 GitHub stars~380 tokensUpdated 24 days ago
    Auto-check: notes

Questions about Configuring Vision

What does Configuring Vision do?

The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models. Configuring Vision is an agent skill from oxbshw/watch-skill. The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

When should I use Configuring Vision?

Configuring Vision fits situations like: tasks that involve LLM inference and serving; tasks that involve Model routing and gateways.

How do I install Configuring Vision in Claude Code?

Run `npx skills add oxbshw/watch-skill --skill configuring-vision -a claude-code`. Or copy the skill folder (skills/configuring-vision in oxbshw/watch-skill) into .claude/skills/configuring-vision in your project. Claude Code loads it when a task matches its description.

How do I install Configuring Vision in Codex?

Run `npx skills add oxbshw/watch-skill --skill configuring-vision -a codex`. Or copy the skill folder (skills/configuring-vision in oxbshw/watch-skill) into .agents/skills/configuring-vision in your project. Codex loads it when a task matches its description.

Can I use Configuring Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oxbshw/watch-skill --skill configuring-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/configuring-vision, .gemini/skills/configuring-vision, .github/skills/configuring-vision and .opencode/skills/configuring-vision in your project.

What does Configuring Vision need to run?

SKILL.md names no scripts, command-line tools or credentials: Configuring Vision is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Read.

Does Configuring Vision access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Configuring Vision safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Configuring Vision use?

Configuring Vision is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Configuring Vision use?

About 509 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Configuring Vision?

Skills that share tags, products or a category with Configuring Vision: New Provider (finch-xu/cc-router, 272 stars), Quorum (Detrol/quorum-cli, 119 stars), Claudish Usage (MadAppGang/claudish, 1k stars) and Page Agent (Tommy-yw/RunbookHermes, 546 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Configuring Vision?

oxbshw (a GitHub user) maintains it in oxbshw/watch-skill, which has 469 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 14, 2026.

Source: oxbshw/watch-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.