Agent skill

Add Image Vision

by sbusso in sbusso/claudeclaw

Add image vision to ClaudeClaw agents. An agent skill from sbusso/claudeclaw.

MITAuto-check passed

Install Add Image Vision

skills CLI
$ npx skills add sbusso/claudeclaw --skill add-image-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sbusso/claudeclaw add-image-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sbusso/claudeclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/add-image-vision .claude/skills/add-image-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-image-vision
GitHub stars
194
Token cost
~787 tokens
SKILL.md length
304 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Add image vision to ClaudeClaw agents. An agent skill from sbusso/claudeclaw.

  • Works in 4 steps: Pre-flight → Apply Code Changes → Configure → …
  • SKILL.md covers Phase 1: Pre-flight, Phase 2: Apply Code Changes, Phase 3: Configure and Phase 4: Verify, plus 1 more section
  • Calls git, npm and npx; reaches github.com

What it does

Add Image Vision is an agent skill from sbusso/claudeclaw. Add image vision to ClaudeClaw agents. Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks.

Its SKILL.md is about 790 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with WhatsApp and npm. The repository describes itself as: Use Claude to orchestrate agents like OpenClaw. The licence is MIT.

Example prompts

  • “/add-image-vision”

Requirements

  • Node.js
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Pre-flight
  2. Apply Code Changes
  3. Configure
  4. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 1395af4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npm
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Image Vision loads about 787 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~787

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sbusso/claudeclaw at commit 1395af4, republished under its MIT licence (© sbusso). 304 words, ~787 tokens.

Download SKILL.mdSave it as .claude/skills/add-image-vision/SKILL.md (or your agent's skills folder).
name
add-image-vision
description
Add image vision to ClaudeClaw agents. Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks.

Image Vision Skill

Adds the ability for ClaudeClaw agents to see and understand images sent via WhatsApp. Images are downloaded, resized with sharp, saved to the group workspace, and passed to the agent as base64-encoded multimodal content blocks.

Phase 1: Pre-flight

  1. Check if src/image.ts exists — skip to Phase 3 if already applied
  2. Confirm sharp is installable (native bindings require build tools)

Prerequisite: WhatsApp must be installed first (skill/whatsapp merged). This skill modifies WhatsApp channel files.

Phase 2: Apply Code Changes

Ensure WhatsApp fork remote
bash
git remote -v

If whatsapp is missing, add it:

bash
git remote add whatsapp https://github.com/qwibitai/claudeclaw-whatsapp.git
Merge the skill branch
bash
git fetch whatsapp skill/image-vision
git merge whatsapp/skill/image-vision || {
  git checkout --theirs package-lock.json
  git add package-lock.json
  git merge --continue
}

This merges in:

  • src/image.ts (image download, resize via sharp, base64 encoding)
  • src/image.test.ts (8 unit tests)
  • Image attachment handling in src/channels/whatsapp.ts
  • Image passing to agent in src/index.ts and src/orchestrator/container-runner.ts
  • Image content block support in agent/runner/src/index.ts
  • sharp npm dependency in package.json

If the merge reports conflicts, resolve them by reading the conflicted files and understanding the intent of both sides.

Validate code changes
bash
npm install
npm run build
npx vitest run src/image.test.ts

All tests must pass and build must be clean before proceeding.

Phase 3: Configure

  1. Rebuild the container (agent-runner changes need a rebuild):

    bash
    ./src/runtimes/docker/build.sh
  2. Sync agent-runner source to group caches:

    bash
    for dir in data/sessions/*/agent-runner-src/; do
      cp agent/runner/src/*.ts "$dir"
    done

Service name: Derived from the directory name: com.claudeclaw.<dirname> (macOS) / claudeclaw-<dirname> (Linux). For example, if cwd is my-assistant, the service is com.claudeclaw.my-assistant. Determine the correct service name before running service commands below.

  1. Restart the service:
    bash
    launchctl kickstart -k gui/$(id -u)/com.claudeclaw

Phase 4: Verify

  1. Send an image in a registered WhatsApp group
  2. Check the agent responds with understanding of the image content
  3. Check logs for "Processed image attachment":
    bash
    tail -50 groups/*/logs/container-*.log

Troubleshooting

  • "Image - download failed": Check WhatsApp connection stability. The download may timeout on slow connections.
  • "Image - processing failed": Sharp may not be installed correctly. Run npm ls sharp to verify.
  • Agent doesn't mention image content: Check container logs for "Loaded image" messages. If missing, ensure agent-runner source was synced to group caches.

© sbusso, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/add-image-vision of sbusso/claudeclaw.

Open the folder on GitHubat commit 1395af4

Compare with similar skills

Add Image Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Image Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Image Vision this skillsbusso/claudeclaw194—~787Automated safety check: PassMIT
Defuddlekepano/obsidian-skills49k11 repos~208Automated safety check: PassMIT
Vercel Deploybytedance/deer-flow84k10 repos~797Automated safety check: PassMIT
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Knap Markdown Templateskepano/obsidian-skills49k2 repos~986Automated safety check: PassMIT
Nx Run Tasksnomcopter/react-mosaic4.8k8 repos~613Automated safety check: PassCustom licence

Similar skills

  • Defuddle

    kepano/obsidian-skills

    Uses the Defuddle CLI to pull clean, readable Markdown, JSON or metadata from web pages, stripping navigation, ads and clutter to save tokens.

    49k GitHub starsUsed in 11 repos~208 tokens
    Knowledge ManagementAuto-check passed
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    84k GitHub starsUsed in 10 repos~797 tokens
    DevOps & CloudAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Knap Markdown Templates

    kepano/obsidian-skills

    Renders Markdown notes from Knap templates and JSON data on the command line, including notes built from Defuddle web page output.

    49k GitHub starsUsed in 2 repos~986 tokens
    Documents & OfficeAuto-check passed
  • Nx Run Tasks

    nomcopter/react-mosaic

    Helps with running tasks in an Nx workspace. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 8 repos~613 tokens
    DevelopmentAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes

More from sbusso/claudeclaw

All 23 skills in this repo
  • Debug

    sbusso/claudeclaw

    Debug container agent issues. An agent skill from sbusso/claudeclaw.

    194 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes
  • X Integration

    sbusso/claudeclaw

    X (Twitter) integration for ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Gmail

    sbusso/claudeclaw

    Add Gmail integration to ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Qmd

    sbusso/claudeclaw

    Add QMD (Query Markup Documents) as an advanced memory search backend.

    194 GitHub stars~629 tokensUpdated 1 mo ago
    Auto-check passed
  • Add Telegram

    sbusso/claudeclaw

    Add Telegram as a channel. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Telegram Swarm

    sbusso/claudeclaw

    Add Agent Swarm (Teams) support to Telegram. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3.7k tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Questions about Add Image Vision

What does Add Image Vision do?

Add image vision to ClaudeClaw agents. An agent skill from sbusso/claudeclaw. Add Image Vision is an agent skill from sbusso/claudeclaw. Add image vision to ClaudeClaw agents.

How do I install Add Image Vision in Claude Code?

Run `npx skills add sbusso/claudeclaw --skill add-image-vision -a claude-code`. Or copy the skill folder (skills/add-image-vision in sbusso/claudeclaw) into .claude/skills/add-image-vision in your project. Claude Code loads it when a task matches its description.

How do I install Add Image Vision in Codex?

Run `npx skills add sbusso/claudeclaw --skill add-image-vision -a codex`. Or copy the skill folder (skills/add-image-vision in sbusso/claudeclaw) into .agents/skills/add-image-vision in your project. Codex loads it when a task matches its description.

Can I use Add Image Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sbusso/claudeclaw --skill add-image-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-image-vision, .gemini/skills/add-image-vision, .github/skills/add-image-vision and .opencode/skills/add-image-vision in your project.

What does Add Image Vision need to run?

Going by SKILL.md and its folder, Add Image Vision needs the command-line tools its instructions call (git, npm and npx). Our summary lists: Node.js; Docker.

Does Add Image Vision access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Add Image Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Image Vision use?

Add Image Vision is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Image Vision use?

About 787 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Image Vision?

Skills that share tags, products or a category with Add Image Vision: Defuddle (kepano/obsidian-skills, 49k stars), Vercel Deploy (bytedance/deer-flow, 84k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars) and Knap Markdown Templates (kepano/obsidian-skills, 49k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Image Vision?

sbusso (a GitHub user) maintains it in sbusso/claudeclaw, which has 194 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 12, 2026.

Source: sbusso/claudeclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.