Agent skill

Image Studio

by vellum-ai in vellum-ai/vellum-assistant

Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint).

MITAuto-check passedMedia & Creative

Install Image Studio

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill image-studio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant image-studio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assistant/src/config/bundled-skills/image-studio .claude/skills/image-studio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-studio
GitHub stars
1.4k
Token cost
~1.6k tokens
SKILL.md length
771 words
Files
3
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint).

  • Works in 2 steps: Configuration errors (missing API key,… → Generation failures (any other error:…
  • Wants options to choose from
  • SKILL.md covers Modes, Models, Example calls and Source images for edit mode, plus 5 more sections
  • Runs TypeScript scripts from its folder

What it does

Image Studio is an agent skill from vellum-ai/vellum-assistant. Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint). Can produce multiple variants when the user wants options to choose from.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `TOOLS.json` and `tools/media-generate-image.ts`). Compatibility notes: Designed for Vellum personal assistants

It sits in Media & Creative, covering Image editing. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

When your agent uses it

  • Wants options to choose from
  • Tasks that involve Image editing

Example prompts

  • “/image-studio”

Requirements

  • Node.js
  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Configuration errors (missing API key, provider not set up): report the error to the user as-is. Do NOT change service configuration…
  2. Generation failures (any other error: "invalid", content policy, safety rejection, provider error). Do not diagnose the cause; switch…

What it can do on your machine

Read from SKILL.md and the folder at commit 33cc983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Image Studio loads about 1.6k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 771 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit 33cc983, republished under its MIT licence (© vellum-ai). 771 words, ~1,576 tokens.

Download SKILL.mdSave it as .claude/skills/image-studio/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
image-studio
description
Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint). Can produce multiple variants when the user wants options to choose from.
compatibility
Designed for Vellum personal assistants
metadata.emoji
🎨

Use the media_generate_image tool via skill_execute to create or edit images.

Modes

  • generate (default): Create a new image from a text prompt.
  • edit: Modify an existing image. Requires one or more source images via source_paths.

Models

Do not pass the model parameter unless you need a specific tier. Omitting it uses the configured default, which is correct for most requests.

When you do need to choose, use an alias, not a concrete model ID. Aliases always resolve to the current model for that tier:

  • fast: quickest, good quality (default tier)
  • quality: higher fidelity, slower
  • openai: OpenAI's model; most permissive on photo edits

Pass a concrete model ID only if the user names one explicitly. If the tool rejects an unknown model ID, the error lists the currently available models and aliases.

Example calls

Generate (no model parameter, default is correct):

json
{
  "tool": "media_generate_image",
  "input": {
    "prompt": "A sunset over the ocean, golden hour, soft haze, 35mm photo style",
    "variants": 2
  }
}

Edit:

json
{
  "tool": "media_generate_image",
  "input": {
    "prompt": "Remove the watermark text from the background. Keep the subject, framing, lighting, and colors exactly identical. Change nothing else.",
    "mode": "edit",
    "source_paths": ["conversations/<conv-id>/attachments/photo.jpeg"],
    "model": "openai"
  }
}

source_paths is a flat array of file path strings. Do NOT pass objects:

  • Wrong: "source_paths": [{ "path": "img.jpeg" }] → schema validation error
  • Right: "source_paths": ["img.jpeg"]

Source images for edit mode

  • Paths resolve inside the workspace. Conversation attachments live under conversations/<conversation-id>/attachments/; prefer that path for images the user attached.
  • Host paths (e.g. ~/Desktop/photo.jpg) only work if the file arrived as an attachment; the tool falls back to the stored workspace copy. If the user references a host file that was never attached, pull it into the workspace first, then pass the workspace path.

Prompting

  • Generate: describe style, composition, lighting, and mood, not just the subject.
  • Edit: name the change AND what must stay the same. Models re-render the whole image, so without preservation language ("keep subject, framing, and lighting identical; only change X") they drift on crop and color.
  • Aspect ratio and size have no parameter today. State them in the prompt ("16:9 widescreen banner") and verify the output.
  • Use variants (1 to 4) when the user wants options. In edit mode always use variants: 1: edits run 60-90 seconds per variant, and two variants can exceed the tool execution timeout (timeouts.toolExecutionTimeoutSec, default 120s). If the user wants multiple edit options, make separate sequential calls.

Timing

Edits on large photos are slow (1 to 2 minutes). If the tool reports a timeout ("timed out after Ns"), the result is lost; do not wait for it to appear. Retry with variants: 1, or if it already was 1, fall back to the CLI which writes files to disk: assistant image-generation generate --prompt "..." --mode edit --source <path> --model openai --output-dir <dir>.

Show full SKILL.md (364 more words)Show less

Output handling

Each generated image is saved into the workspace under media/generated/ and the tool result lists the saved paths. The images also come back as inline content blocks so you can judge the result before presenting it.

  • Present an image to the user by embedding its saved path in your reply: ![short description](vellum://workspace/media/generated/<file>.png). The app renders it inline where your text refers to it, and chat channels (Slack, Telegram, WhatsApp) deliver it as a native image upload.
  • Label a plain image link "View image", for example [View image](vellum://workspace/media/generated/<file>.png). The link opens a preview; do not call it "Download" or promise a download in the surrounding text. To save the image, the user can choose Download in the preview.
  • If you do not embed it, the image is still auto-attached to your reply as a file, so it is never lost. Prefer embedding: an attachment chip at the end of the message is a worse presentation than the image inline.
  • To iterate on a result, pass its saved path via source_paths with mode: "edit".

Error handling

Two kinds of failure. Treat them differently:

  1. Configuration errors (missing API key, provider not set up): report the error to the user as-is. Do NOT change service configuration (managed vs your-own mode, default provider, or default model in Settings). Configuration changes happen only at the user's explicit request.
  2. Generation failures (any other error: "invalid", content policy, safety rejection, provider error). Do not diagnose the cause; switch providers. The error message names the model that failed:
    • If the error names a gemini-* model (or no model) → retry ONCE with model: "openai".
    • If the error names gpt-image-2 or another gpt-* model → retry ONCE with model: "quality".
    • If the retry also fails → stop and report both errors to the user.

Do NOT rephrase the prompt and retry on the same model, even if the error suggests checking the prompt. One provider switch, then stop.

Complete when

The tool has returned at least one image and your reply presents it to the user, preferably as an inline ![description](vellum://workspace/...) embed of the saved path. An error report counts as complete only after the retry path in Error handling has been exhausted.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in assistant/src/config/bundled-skills/image-studio of vellum-ai/vellum-assistant.

  • SKILL.md
  • TOOLS.json
  • tools/media-generate-image.ts

Open the folder on GitHubat commit 33cc983

Compare with similar skills

Image Studio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Studio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Studio this skillvellum-ai/vellum-assistant1.4k—~1.6kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Generate Imageynulihao/AgentSkillOS61810 repos~1.7kAutomated safety check: NotesNone
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Media Useedenfunf/reelmimic1.9k1 repos~2kAutomated safety check: PassMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Generate Image

    ynulihao/AgentSkillOS

    Generate or edit images using AI models (FLUX, Gemini). An agent skill from ynulihao/AgentSkillOS.

    618 GitHub starsUsed in 10 repos~1.7k tokens
    Media & CreativeAuto-check: notes
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Media Use

    edenfunf/reelmimic

    Agent Media OS, the single skill for every media need in a HyperFrames project.

    1.9k GitHub starsUsed in 1 repo~2k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated today
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check: warnings

Questions about Image Studio

What does Image Studio do?

Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint). Image Studio is an agent skill from vellum-ai/vellum-assistant. Create images from a text description, or edit photos and graphics the user provides (remove backgrounds or watermarks, retouch, restyle, in-paint).

When should I use Image Studio?

Image Studio fits situations like: wants options to choose from; tasks that involve Image editing.

How do I install Image Studio in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill image-studio -a claude-code`. Or copy the skill folder (assistant/src/config/bundled-skills/image-studio in vellum-ai/vellum-assistant) into .claude/skills/image-studio in your project. Claude Code loads it when a task matches its description.

How do I install Image Studio in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill image-studio -a codex`. Or copy the skill folder (assistant/src/config/bundled-skills/image-studio in vellum-ai/vellum-assistant) into .agents/skills/image-studio in your project. Codex loads it when a task matches its description.

Can I use Image Studio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill image-studio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-studio, .gemini/skills/image-studio, .github/skills/image-studio and .opencode/skills/image-studio in your project.

What does Image Studio need to run?

Going by SKILL.md and its folder, Image Studio needs TypeScript for the scripts in its folder. Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Image Studio access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Studio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image Studio use?

Image Studio is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Studio use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Studio?

Skills that share tags, products or a category with Image Studio: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Generate Image (ynulihao/AgentSkillOS, 618 stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars) and HyperFrames Media Use (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Studio?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,408 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.