Agent skill

Image

by guaardvark in guaardvark/guaardvark

Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch…

MITAuto-check passedMedia & Creative

Install Image

skills CLI
$ npx skills add guaardvark/guaardvark --skill image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guaardvark/guaardvark image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guaardvark/guaardvark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/image .claude/skills/image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image
GitHub stars
255
Token cost
~1.8k tokens
SKILL.md length
901 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch…

  • The user asks to create
  • SKILL.md covers One image: MCP generate_image, Photo edits over MCP run as jobs, Edit an existing image: MCP… and New scene from a face: MCP…, plus 4 more sections
  • Calls curl
  • Batch-generate images locally

What it does

Image is an agent skill from guaardvark/guaardvark. Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch runs of many prompts. Use when the user asks to create, draw, render, visualize, edit, or batch-generate images locally.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image editing, Image generation and Diffusion and image models. It works with Model Context Protocol, ComfyUI and Stable Diffusion. The repository describes itself as: The self-hosted AI studio: local video, image, music, voice, LoRA training, coding swarms, RAG and screen agents on one GPU, driven from the Studio or by your coding agent… The licence is MIT.

When your agent uses it

  • The user asks to create
  • Batch-generate images locally

Example prompts

  • “/image”

What it can do on your machine

Read from SKILL.md and the folder at commit f022ceb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image loads about 1.8k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 901 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guaardvark/guaardvark at commit f022ceb, republished under its MIT licence (© guaardvark). 901 words, ~1,755 tokens.

Download SKILL.mdSave it as .claude/skills/image/SKILL.md (or your agent's skills folder).
name
image
description
Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch runs of many prompts. Use when the user asks to create, draw, render, visualize, edit, or batch-generate images locally.

Images with Guaardvark

Read setup first if the backend or the comfyui plugin state is unknown.

One image: MCP generate_image

  • prompt is scene, pose, lighting, setting. Plain prose. Do not paste JSON or tag soup; the default model (Z-Image Turbo) reads prompts as language, and SD-era tag lists hurt it.
  • model default auto picks the best downloaded model. Only override when the user names one: zimage-turbo, krea2-turbo, krea2-raw, flux-dev, sd-xl, sdxl-turbo, realistic-vision, epic-realism.
  • width / height: 512, 768 or 1024. style: realistic, artistic, anime, photographic, digital-art.
  • Consistent character: pass subject_ids=[<cast id>] as its own array. Never put the trigger word alone in the prompt and expect the LoRA to load. Find ids with GET /api/cast-library (see the cast skill).
  • On-image text: quote the exact words in double quotes inside the prompt.
  • The tool returns the image URL (/api/outputs/generated_images/<file>.png, relative to the backend), the model that ran, steps, seed and whether a Cast LoRA was applied. Show the URL and the prompt you used. Measured: 768x768 on Z-Image Turbo in ~20 s on a free 16 GB card.
  • Over MCP the call queues by default (wait_for_result defaults to false there) and returns Image queued as batch ImageBatch_... at once. Poll get_generation_status(batch_id=...) every few seconds until completed; it returns the file URL. Pass wait_for_result: true to block for the render instead (allowed up to 30 minutes). A call that exceeds the server's timeout answers with an error that says the render is still running; it is not lost.
  • A failed call carries the backend's reason (plugin off, out of memory, bad model). Read it and act on it; inspect_gpu and GET /api/plugins/status are the two checks that resolve most.

Photo edits over MCP run as jobs

edit_image, inpaint_image, outpaint_image and remove_background answer at once over MCP with ... queued as job tooljob_...: the work runs in the Guaardvark backend and keeps going if the client disconnects. Poll get_generation_status(batch_id="tooljob_...") every few seconds until done (it gives the file URL) or failed (it gives the error). Pass wait_for_result: true to wait up to 60 s (half the server's MCP timeout when that is shorter) for the result; a longer edit still comes back as the job id. A Qwen-Image-Edit takes about three minutes on a 16 GB card, and a job waits its turn while another render holds the GPU.

Edit an existing image: MCP edit_image

  • instruction is the change ("put a cowboy hat on him", "make the shirt red"). The image the user just attached is used automatically; otherwise pass image as a path or URL. Optional reference_image_2 / reference_image_3 (Qwen-Image-Edit only) for extra people or style; the call is refused when a reference cannot be read or the edit would run on another backend.
  • model auto uses Qwen-Image-Edit when installed, else FLUX.1 Kontext. With neither installed the call is refused and names the pack; relay that instead of retrying. Override with qwen-image-edit or kontext. Install those packs from Manage Image Models → Image editing (qwen-image-edit, flux-kontext-dev) — do not Install unless the user asked. Naming another downloaded image model runs a light img2img pass that keeps most of the picture.
  • Same canvas, same pose. For a brand-new picture use generate_image. For a new scene that keeps a face use generate_identity.
Show full SKILL.md (368 more words)Show less

New scene from a face: MCP generate_identity (off by default)

  • Not exposed unless the server runs with GUAARDVARK_IDENTITY_TOOL=1: the likeness it keeps has not passed verification yet. If the tool is absent, say so and offer edit_image instead.
  • Attach a likeness the user has the right to use (their photo or a Cast subject they uploaded). consented must be true. Refuse if they have not confirmed that.
  • prompt is the new scene. Needs the PuLID identity pack (pulid-flux; Manage Image Models → Image editing installs it with its face files, EVA02-CLIP and flux-dev). Comfy must have been restarted after the PuLID-Flux custom node was added.
  • This is not a face swap onto an existing poster, and not an instruction edit of the same photo.

Background remove: MCP remove_background

  • Cuts the subject out of the attached photo (transparent PNG). ONNX matting, no diffusion; needs a background-removal model from Manage Image Models → Image editing.
  • For a new background, remove first then edit_image / generate_identity, or describe the new scene in edit_image if Qwen-Image-Edit is installed.

Inpaint / outpaint

  • inpaint_image: change or remove something ("remove the coffee cup"). Runs on Qwen-Image-Edit or FLUX.1 Kontext; refused when neither is installed.
  • outpaint_image: extend the canvas (left/right/top/bottom pixels) and fill. Prefers Qwen.

Many images: REST batch

bash
B=${GUAARDVARK_URL:-http://localhost:5000}
curl -s -X POST $B/api/batch-image/generate/prompts -H 'Content-Type: application/json' -d '{
  "prompts": ["prompt one", "prompt two"],
  "model": "auto",
  "subject_ids": []
}'
  • prompts may be strings or {"prompt": "..."} objects. There is a per-batch maximum; if the server answers 400 "Too many prompts", split the list.
  • Optional adapters (user LoRAs from the models skill) and subject_ids (Cast Library).
  • The response is data.batch_id (ImageBatch_<date>_<n>). Poll GET $B/api/batch-image/status/<batch_id>?include_results=true: status goes running → completed, with completed_images / total_images, output_dir, and one results[] entry per prompt (success, image_path, thumbnail_path, generation_time, metadata.model_used). A contact sheet: GET $B/api/batch-image/preview/<batch_id>; one file: GET $B/api/batch-image/image/<batch_id>/<image_name> (the basename of image_path). Cancel with POST $B/api/batch-image/cancel/<batch_id>. Measured: one 1024x1024 prompt completed in ~30 s.
  • Helpers: POST /api/batch-image/enhance-prompt, /analyze-prompt, /expand-concept (JSON body with the prompt) when the user wants prompt help before spending GPU time.

Rules

  • Say which model actually ran (the response names it). Do not promise a model that is not installed.
  • Generation time depends on the GPU; a first image after Ollama held the card can take longer because the orchestrator swaps models. That is normal.
  • Never upload the user's images anywhere. Everything here is local.

© guaardvark, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/image of guaardvark/guaardvark.

Open the folder on GitHubat commit f022ceb

Compare with similar skills

Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image this skillguaardvark/guaardvark255—~1.8kAutomated safety check: PassMIT
Comfyui Skill OpenclawHuangYuChuh/ComfyUI_Skills_OpenClaw411—~2.7kAutomated safety check: PassApache-2.0
Workflow Template BuilderMooshieblob1/MooshieUI207—~640Automated safety check: PassAGPL-3.0
Iibzanllp/infinite-image-browsing1.4k—~3.3kAutomated safety check: PassMIT
Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.2kAutomated safety check: PassMIT
Anima Baseartokun/comfyui-mcp795—~4kAutomated safety check: PassMIT

Similar skills

  • Comfyui Skill Openclaw

    HuangYuChuh/ComfyUI_Skills_OpenClaw

    Run registered ComfyUI workflows through the fast comfyui-skill CLI, and use the official local Comfy MCP for live template, node, model, validation, and orchestration capabilities.

    411 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Workflow Template Builder

    Mooshieblob1/MooshieUI

    Builds or modifies ComfyUI workflow JSON templates in MooshieUI's Rust backend (src-tauri/src/templates).

    207 GitHub stars~640 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Iib

    zanllp/infinite-image-browsing

    Interact with IIB (Infinite Image Browsing) service for searching, browsing, tagging, and organizing AI-generated images.

    1.4k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Stable Diffusion with Diffusers

    Orchestra-Research/AI-Research-SKILLs

    Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.

    13k GitHub starsUsed in 6 repos~3.2k tokens
    Media & CreativeAuto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    795 GitHub stars~4k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from guaardvark/guaardvark

All 15 skills in this repo
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    255 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Voice

    guaardvark/guaardvark

    Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip.

    255 GitHub starsUsed in 1 repo~798 tokens
    Auto-check passed
  • Cast

    guaardvark/guaardvark

    Build consistent characters, environments and props in Guaardvark's Cast Library and train LoRAs for them locally (reference photos → vision bible → sample plan → approved samples → training).

    255 GitHub stars~673 tokensUpdated today
    Auto-check passed
  • Music

    guaardvark/guaardvark

    Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.

    255 GitHub stars~710 tokensUpdated today
    Auto-check passed
  • Ops

    guaardvark/guaardvark

    Operate a running Guaardvark: GPU and VRAM state, plugin start/stop, logs, Celery tasks, the Interconnector sync to other machines, overnight RAG autoresearch, and infographics.

    255 GitHub stars~779 tokensUpdated today
    Auto-check passed
  • Swarm

    guaardvark/guaardvark

    Launch and watch Guaardvark's Swarm Orchestrator: parallel coding agents, each in its own git worktree, working a markdown plan and merging back deterministically.

    255 GitHub stars~746 tokensUpdated today
    Auto-check passed

Questions about Image

What does Image do?

Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch…. Image is an agent skill from guaardvark/guaardvark. Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch runs of many prompts.

When should I use Image?

Image fits situations like: the user asks to create; batch-generate images locally.

How do I install Image in Claude Code?

Run `npx skills add guaardvark/guaardvark --skill image -a claude-code`. Or copy the skill folder (.agents/skills/image in guaardvark/guaardvark) into .claude/skills/image in your project. Claude Code loads it when a task matches its description.

How do I install Image in Codex?

Run `npx skills add guaardvark/guaardvark --skill image -a codex`. Or copy the skill folder (.agents/skills/image in guaardvark/guaardvark) into .agents/skills/image in your project. Codex loads it when a task matches its description.

Can I use Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guaardvark/guaardvark --skill image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image, .gemini/skills/image, .github/skills/image and .opencode/skills/image in your project.

What does Image need to run?

Going by SKILL.md and its folder, Image needs the command-line tools its instructions call (curl).

Does Image access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image use?

Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image?

Skills that share tags, products or a category with Image: Comfyui Skill Openclaw (HuangYuChuh/ComfyUI_Skills_OpenClaw, 411 stars), Workflow Template Builder (Mooshieblob1/MooshieUI, 207 stars), Iib (zanllp/infinite-image-browsing, 1.4k stars) and Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image?

guaardvark (a GitHub user) maintains it in guaardvark/guaardvark, which has 255 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: guaardvark/guaardvark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.