Agent skill

Image Gen

by amd in amd/gaia

Turn a description into an image file with local Stable Diffusion, then iterate on it.

MITAuto-check passedMedia & Creative

Install Image Gen

skills CLI
$ npx skills add amd/gaia --skill image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/skills/image-gen .claude/skills/image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-gen
GitHub stars
1.6k
Token cost
~1.4k tokens
SKILL.md length
794 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Turn a description into an image file with local Stable Diffusion, then iterate on it.

  • The user says draw
  • SKILL.md covers Check what is loaded before…, Build the prompt for them, The default is a few-step… and Iterate instead of starting over, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Make a picture of

What it does

Image Gen is an agent skill from amd/gaia. Turn a description into an image file with local Stable Diffusion, then iterate on it. Use when the user says draw, sketch, paint, render, "make a picture of", "generate an image", or asks for concept art, a thumbnail, a logo idea, a wallpaper, or a mockup — and when they want the last image changed rather than replaced.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image generation and Diffusion and image models. It works with Stable Diffusion. The repository describes itself as: Build AI agents for your PC. The licence is MIT.

When your agent uses it

  • The user says draw
  • Make a picture of
  • Generate an image
  • Asks for concept art

Example prompts

  • “make a picture of”
  • “generate an image”
  • “/image-gen”

What it can do on your machine

Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Gen loads about 1.4k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 794 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 794 words, ~1,358 tokens.

Download SKILL.mdSave it as .claude/skills/image-gen/SKILL.md (or your agent's skills folder).
name
image-gen
description
Turn a description into an image file with local Stable Diffusion, then iterate on it. Use when the user says draw, sketch, paint, render, "make a picture of", "generate an image", or asks for concept art, a thumbnail, a logo idea, a wallpaper, or a mockup — and when they want the last image changed rather than replaced.
license
MIT
version
1.0.0

Image Generation

Generation runs locally and is slow — tens of seconds to minutes per image, and the first call for a model downloads gigabytes. That changes the job: you get few attempts, so spend the thinking before the call rather than firing off four variations and picking one.

It also costs the conversation. Drawing loads the image model in place of the chat model, so the reply after an image pauses while the chat model comes back. Generate when the user actually asked for a picture — not to illustrate an answer they did not ask to have illustrated.

Check what is loaded before you promise anything

Call list_sd_models() first. It tells you which models exist and what each costs, and the reported default_model is the one you get if you pass no model. Do not assume a specific model is resident — naming one the machine has not pulled turns a 20-second request into a multi-gigabyte download the user did not agree to.

Tell the user the estimate before a slow model, not after: SDXL-Base-1.0 at 1024x1024 is on the order of minutes, the Turbo models are seconds.

Build the prompt for them

A user asking for "a cat" has a picture in their head that "a cat" will not produce. Expand it yourself rather than interrogating them — one round of questions is fine, three is a worse experience than a decent first image.

A usable prompt names, roughly in this order: subject, what it is doing or how it is arranged, setting, style, lighting or mood. So "a red bicycle" becomes "a red bicycle leaning against a brick wall, morning sunlight, shallow depth of field, photographic".

Then say the expanded prompt back to the user with the result. They cannot correct a prompt they never saw, and "make it warmer" is only meaningful if they know what you asked for.

The default is a few-step model — do not over-tune it

SDXL-Turbo is the default and it is distilled to converge in about 4 steps with CFG around 1.0. The knobs that matter on a normal model do nothing useful here:

  • Raising steps to 30 costs seven times the wall clock and does not improve the image.
  • Raising cfg_scale degrades it — Turbo models are trained for guidance-free sampling.
  • Long negative-prompt boilerplate ("blurry, low quality, watermark, extra fingers…") is wasted. Spend those words describing what you do want.

Leave steps, cfg_scale, and size unset unless you have a reason; the tool fills in the right values per model. Reach for SDXL-Base-1.0 only when the user explicitly wants photorealism and has accepted the wait.

Show full SKILL.md (363 more words)Show less

Iterate instead of starting over

get_generation_history() returns this session's generations with the exact prompt, model, size and seed of each. When the user says "same but at sunset" or "make it wider", read the previous entry, change the one thing they asked about, and keep everything else — including the seed. Reusing the seed is what makes the second image recognisably the same picture rather than an unrelated one that happens to match the words.

Rewriting the prompt from scratch throws away everything that was already working, and the user has to re-explain the parts they liked.

When it fails, say what failed

generate_image returns {"status": "error", "error": ...} rather than raising. Read it and pass the actual message to the user.

Do not quietly retry with a different model, a smaller size, or fewer steps. A user who asked for a photorealistic 1024px render and silently received a 512px Turbo sketch has been given the wrong thing and told nothing. If a fallback would genuinely help, propose it and let them choose.

The common failures and what to say:

  • Cannot reach Lemonade Server — inference is not running. Tell them to start it; nothing here works until it is up.
  • Timed out — usually the first use of a model, downloading several GB. The server is fine. Tell them to pre-fetch it with the pull command the error names for their install, then retry, rather than restarting anything.
  • Invalid model or size — you passed something outside the supported set. Call list_sd_models() and pick from what it returned.

Reporting a generated image

Give them the path. It is the only part of the result they can act on:

Saved to ~/.gaia/cache/sd/images/a_red_bicycle_..._SDXL-Turbo_....png (18s).

Prompt used: "a red bicycle leaning against a brick wall, morning sunlight, shallow depth of field, photographic" — say the word if you want it warmer, wider, or at a different time of day.

Never describe an image you did not generate, and never claim a file exists because the call was made — check status first.

Fork this

Pin the style clause in step two to your own house look (brand palette, flat vector, isometric) and the skill stops needing to be told it every time.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hub/skills/image-gen of amd/gaia.

Open the folder on GitHubat commit 6c3bb5c

Compare with similar skills

Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Gen this skillamd/gaia1.6k—~1.4kAutomated safety check: PassMIT
Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.2kAutomated safety check: PassMIT
Stable DiffusionLuciole-Studio/Misaka-Agent1251 repos~3.2kAutomated safety check: PassMIT
Iibzanllp/infinite-image-browsing1.4k—~3.3kAutomated safety check: PassMIT
ImageNexus-JPF/note-companion8692 repos~3.9kAutomated safety check: PassMIT
Anima Baseartokun/comfyui-mcp793—~4kAutomated safety check: PassMIT

Similar skills

  • Stable Diffusion with Diffusers

    Orchestra-Research/AI-Research-SKILLs

    Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.

    13k GitHub starsUsed in 6 repos~3.2k tokens
    Media & CreativeAuto-check passed
  • Stable Diffusion

    Luciole-Studio/Misaka-Agent

    Text-to-image generation, inpainting, and img2img. An agent skill from Luciole-Studio/Misaka-Agent.

    125 GitHub starsUsed in 1 repo~3.2k tokens
    Media & CreativeAuto-check passed
  • Iib

    zanllp/infinite-image-browsing

    Interact with IIB (Infinite Image Browsing) service for searching, browsing, tagging, and organizing AI-generated images.

    1.4k GitHub stars~3.3k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Image

    Nexus-JPF/note-companion

    When the user wants to create, generate, edit, or optimize images for marketing — blog heroes, social graphics, product mockups, profile banners, listing visuals, or brand assets.

    869 GitHub starsUsed in 2 repos~3.9k tokens
    Media & CreativeAuto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    793 GitHub stars~4k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Add Tauri Command

    Mooshieblob1/MooshieUI

    Adds a MooshieUI Tauri command end-to-end — Rust handler, lib.rs registration, and TypeScript ipcInvoke wrapper.

    207 GitHub stars~467 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Image Gen

What does Image Gen do?

Turn a description into an image file with local Stable Diffusion, then iterate on it. Image Gen is an agent skill from amd/gaia. Turn a description into an image file with local Stable Diffusion, then iterate on it.

When should I use Image Gen?

Image Gen fits situations like: the user says draw; make a picture of; generate an image; asks for concept art.

How do I install Image Gen in Claude Code?

Run `npx skills add amd/gaia --skill image-gen -a claude-code`. Or copy the skill folder (hub/skills/image-gen in amd/gaia) into .claude/skills/image-gen in your project. Claude Code loads it when a task matches its description.

How do I install Image Gen in Codex?

Run `npx skills add amd/gaia --skill image-gen -a codex`. Or copy the skill folder (hub/skills/image-gen in amd/gaia) into .agents/skills/image-gen in your project. Codex loads it when a task matches its description.

Can I use Image Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-gen, .gemini/skills/image-gen, .github/skills/image-gen and .opencode/skills/image-gen in your project.

What does Image Gen need to run?

SKILL.md names no scripts, command-line tools or credentials: Image Gen is instructions for the agent only.

Does Image Gen access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image Gen use?

Image Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Gen use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Gen?

Skills that share tags, products or a category with Image Gen: Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Stable Diffusion (Luciole-Studio/Misaka-Agent, 125 stars), Iib (zanllp/infinite-image-browsing, 1.4k stars) and Image (Nexus-JPF/note-companion, 869 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Gen?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.