Stable Diffusion with Diffusers
Orchestra-Research/AI-Research-SKILLs
Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.
Turn a description into an image file with local Stable Diffusion, then iterate on it.
$ npx skills add amd/gaia --skill image-gen -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/gaia image-gen --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/skills/image-gen .claude/skills/image-gen && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .claude/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/gaia/tree/main/hub/skills/image-genType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/gaia --skill image-gen -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/gaia image-gen --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .agents/skills && cp -r skills-src/hub/skills/image-gen .agents/skills/image-gen && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .agents/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill image-gen -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/gaia image-gen --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/hub/skills/image-gen .cursor/skills/image-gen && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .cursor/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/gaia.git --path hub/skills/image-gen--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/gaia --skill image-gen -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/gaia image-gen --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/hub/skills/image-gen .gemini/skills/image-gen && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .gemini/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/gaia image-genInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/gaia --skill image-gen -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .github/skills && cp -r skills-src/hub/skills/image-gen .github/skills/image-gen && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .github/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill image-gen -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/gaia image-gen --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/hub/skills/image-gen .opencode/skills/image-gen && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "image-gen" agent skill from https://github.com/amd/gaia/tree/main/hub/skills/image-gen into .opencode/skills/image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-gen", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
image-genTurn a description into an image file with local Stable Diffusion, then iterate on it.
Image Gen is an agent skill from amd/gaia. Turn a description into an image file with local Stable Diffusion, then iterate on it. Use when the user says draw, sketch, paint, render, "make a picture of", "generate an image", or asks for concept art, a thumbnail, a logo idea, a wallpaper, or a mockup — and when they want the last image changed rather than replaced.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Image generation and Diffusion and image models. It works with Stable Diffusion. The repository describes itself as: Build AI agents for your PC. The licence is MIT.
Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Image Gen loads about 1.4k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 794 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 794 words, ~1,358 tokens.
.claude/skills/image-gen/SKILL.md (or your agent's skills folder).Generation runs locally and is slow — tens of seconds to minutes per image, and the first call for a model downloads gigabytes. That changes the job: you get few attempts, so spend the thinking before the call rather than firing off four variations and picking one.
It also costs the conversation. Drawing loads the image model in place of the chat model, so the reply after an image pauses while the chat model comes back. Generate when the user actually asked for a picture — not to illustrate an answer they did not ask to have illustrated.
Call list_sd_models() first. It tells you which models exist and what each
costs, and the reported default_model is the one you get if you pass no
model. Do not assume a specific model is resident — naming one the machine
has not pulled turns a 20-second request into a multi-gigabyte download the
user did not agree to.
Tell the user the estimate before a slow model, not after: SDXL-Base-1.0 at 1024x1024 is on the order of minutes, the Turbo models are seconds.
A user asking for "a cat" has a picture in their head that "a cat" will not produce. Expand it yourself rather than interrogating them — one round of questions is fine, three is a worse experience than a decent first image.
A usable prompt names, roughly in this order: subject, what it is doing or how it is arranged, setting, style, lighting or mood. So "a red bicycle" becomes "a red bicycle leaning against a brick wall, morning sunlight, shallow depth of field, photographic".
Then say the expanded prompt back to the user with the result. They cannot correct a prompt they never saw, and "make it warmer" is only meaningful if they know what you asked for.
SDXL-Turbo is the default and it is distilled to converge in about 4 steps with CFG around 1.0. The knobs that matter on a normal model do nothing useful here:
steps to 30 costs seven times the wall clock and does not improve
the image.cfg_scale degrades it — Turbo models are trained for guidance-free
sampling.Leave steps, cfg_scale, and size unset unless you have a reason; the tool
fills in the right values per model. Reach for SDXL-Base-1.0 only when the
user explicitly wants photorealism and has accepted the wait.
get_generation_history() returns this session's generations with the exact
prompt, model, size and seed of each. When the user says "same but at sunset"
or "make it wider", read the previous entry, change the one thing they asked
about, and keep everything else — including the seed. Reusing the seed is
what makes the second image recognisably the same picture rather than an
unrelated one that happens to match the words.
Rewriting the prompt from scratch throws away everything that was already working, and the user has to re-explain the parts they liked.
generate_image returns {"status": "error", "error": ...} rather than
raising. Read it and pass the actual message to the user.
Do not quietly retry with a different model, a smaller size, or fewer steps. A user who asked for a photorealistic 1024px render and silently received a 512px Turbo sketch has been given the wrong thing and told nothing. If a fallback would genuinely help, propose it and let them choose.
The common failures and what to say:
list_sd_models() and pick from what it returned.Give them the path. It is the only part of the result they can act on:
Saved to
~/.gaia/cache/sd/images/a_red_bicycle_..._SDXL-Turbo_....png(18s).Prompt used: "a red bicycle leaning against a brick wall, morning sunlight, shallow depth of field, photographic" — say the word if you want it warmer, wider, or at a different time of day.
Never describe an image you did not generate, and never claim a file exists
because the call was made — check status first.
Pin the style clause in step two to your own house look (brand palette, flat vector, isometric) and the skill stops needing to be told it every time.
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in hub/skills/image-gen of amd/gaia.
Open the folder on GitHubat commit 6c3bb5c
Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Image Gen this skillamd/gaia | 1.6k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Stable DiffusionLuciole-Studio/Misaka-Agent | 125 | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Iibzanllp/infinite-image-browsing | 1.4k | — | ~3.3k | Automated safety check: Pass | MIT | |
| ImageNexus-JPF/note-companion | 869 | 2 repos | ~3.9k | Automated safety check: Pass | MIT | |
| Anima Baseartokun/comfyui-mcp | 793 | — | ~4k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.
Luciole-Studio/Misaka-Agent
Text-to-image generation, inpainting, and img2img. An agent skill from Luciole-Studio/Misaka-Agent.
zanllp/infinite-image-browsing
Interact with IIB (Infinite Image Browsing) service for searching, browsing, tagging, and organizing AI-generated images.
Nexus-JPF/note-companion
When the user wants to create, generate, edit, or optimize images for marketing — blog heroes, social graphics, product mockups, profile banners, listing visuals, or brand assets.
artokun/comfyui-mcp
Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).
Mooshieblob1/MooshieUI
Adds a MooshieUI Tauri command end-to-end — Rust handler, lib.rs registration, and TypeScript ipcInvoke wrapper.
amd/gaia
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
amd/gaia
Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
amd/gaia
Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.
amd/gaia
Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.
Works with
Categories
Turn a description into an image file with local Stable Diffusion, then iterate on it. Image Gen is an agent skill from amd/gaia. Turn a description into an image file with local Stable Diffusion, then iterate on it.
Image Gen fits situations like: the user says draw; make a picture of; generate an image; asks for concept art.
Run `npx skills add amd/gaia --skill image-gen -a claude-code`. Or copy the skill folder (hub/skills/image-gen in amd/gaia) into .claude/skills/image-gen in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/gaia --skill image-gen -a codex`. Or copy the skill folder (hub/skills/image-gen in amd/gaia) into .agents/skills/image-gen in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-gen, .gemini/skills/image-gen, .github/skills/image-gen and .opencode/skills/image-gen in your project.
SKILL.md names no scripts, command-line tools or credentials: Image Gen is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Image Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Image Gen: Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Stable Diffusion (Luciole-Studio/Misaka-Agent, 125 stars), Iib (zanllp/infinite-image-browsing, 1.4k stars) and Image (Nexus-JPF/note-companion, 869 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.
Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.