AI Image Generation and Editing
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpt-image .claude/skills/gpt-image && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .claude/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-imageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpt-image .agents/skills/gpt-image && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .agents/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpt-image .cursor/skills/gpt-image && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .cursor/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wuyoscar/GPT-Image2-Skill.git --path skills/gpt-image--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpt-image .gemini/skills/gpt-image && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .gemini/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-imageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpt-image .github/skills/gpt-image && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .github/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpt-image .opencode/skills/gpt-image && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpt-image" agent skill from https://github.com/wuyoscar/GPT-Image2-Skill/tree/main/skills/gpt-image into .opencode/skills/gpt-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpt-imageGenerates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
Requests are sorted into generate, edit, inpaint or multi-reference jobs, with the asset type, exact text, aspect ratio, references and budget noted, and the model is settled before any API call. Three choices are listed: gpt-image-2.5-flare for fast drafts, gpt-image-2.5-sunburst for precise reference edits, and gpt-image-2 for existing Image 2 workflows. Image 2 keeps a gallery-first routine, while a precise 2.5 brief can go straight to generation.
The agent calls the gpt-image command or scripts/generate.py with an explicit --model and never writes its own API wrapper. Before costly or ambiguous calls it offers one to three matched directions with planned size and quality, asking at most one short question at a time. It does not reinstall, overwrite skill folders or write keys or .env files unless you ask for setup, and it ends by reporting file paths, key flags and one refinement idea.
A large reference set holds a craft guide and galleries by theme, such as anime and manga, architecture, brand identity, character design, data visualization, fashion, fine art and infographics. Calls read OPENAI_API_KEY and may incur OpenAI API charges.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9f8aa1a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
uvuvxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges.
From compatibility in the SKILL.md frontmatter.
GPT Image Generation CLI loads about 2.5k tokens when it runs, and up to ~79k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 1,185 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
overwrite skill folders, create/modify `.env`, or write API keys unless the user explicitly requested setup. Global/sha`OPENAI_API_KEY` from process env, then `.env`, then `~/.env` without overriding existing env; successful API calls mayset OPENAI_API_KEY`; if a key exists in `.env`/`~/.env`, tell them to remove/rename it for the session rather than workiAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from wuyoscar/GPT-Image2-Skill at commit 9f8aa1a, republished under its MIT licence (© wuyoscar). 1,185 words, ~2,550 tokens.
.claude/skills/gpt-image/SKILL.md (or your agent's skills folder). This skill also uses 44 other files; get the full folder from GitHub.Agent runbook for GPT Image 2 / 2.5 generation/editing. Use the prompt library + packaged CLI. Do not reimplement image API code.
generate, edit, inpaint, or multi-reference; identify asset type, exact text, aspect ratio, references, safety constraints, and budget/quality. Apply the model-choice rules below before any API call.command -v gpt-image), installed tool lists when the tool manager exists, or the runtime’s own skill registry when available. Do not assume a local home path in cloud/hosted runtimes..env, or write API keys unless the user explicitly requested setup. Global/shared installs are opt-in only.gpt-image or scripts/generate.py with an explicit --model. Do not create a new generate.py, SDK wrapper, or ad-hoc script for normal image requests.Fast path: confirmed 2.5 model + precise prompt + “generate now” → preflight and CLI, without a mandatory reference/craft pass. Do not reconfirm an exact valid model.
| Choice | API model ID | Suggested use |
|---|---|---|
| Flare | gpt-image-2.5-flare | Fast general generation and drafts |
| Sunburst | gpt-image-2.5-sunburst | Precise reference edits and detailed control |
| Image 2 | gpt-image-2 | Existing Image 2 workflows and compatibility |
gpt-image-2.5 as an API model ID.--model. The CLI retains gpt-image-2 as its backward-compatible default, but that default is not a substitute for the agent resolving the user's choice.references/models.md when parameter support or validation status needs checking. On an invalid-model, access/403, quota, or policy failure, report the failure and stop; do not switch models or rewrite the prompt to retry automatically.Preferred call order:
# Existing CLI on PATH
gpt-image --model MODEL_ID -p "PROMPT" [-f OUT] [-i REF...] [-m MASK] [options]
# Installed skill folder; use runtime-provided skill path when available
uv run "$SKILL_DIR/scripts/generate.py" --model MODEL_ID -p "PROMPT" [-f OUT] [-i REF...] [-m MASK] [options]
# Direct transient CLI when the user requested setup/one-off CLI execution
uvx --from git+https://github.com/wuyoscar/gpt_image_2_skill gpt-image --model MODEL_ID -p "PROMPT" [options]scripts/generate.py is a launcher: repo-local src/gpt_image_cli → installed gpt-image → PATH gpt-image → transient uvx/uv fallback.
OPENAI_API_KEY from process env, then .env, then ~/.env without overriding existing env; successful API calls may bill the user’s OpenAI account.OPENAI_API_KEY is unset, report missing key or use host-native generation when requested; do not write secrets.unset OPENAI_API_KEY; if a key exists in .env/~/.env, tell them to remove/rename it for the session rather than working around it.| Flag | Values | Use |
|---|---|---|
-p, --prompt | string | Required prompt/edit instruction |
-f, --file | path | Output path; auto-named if omitted |
-i, --image | repeatable path | Use edits endpoint; supports multiple references |
-m, --mask | PNG path | Inpaint with alpha mask; requires -i |
--model | gpt-image-2, gpt-image-2.5-flare, gpt-image-2.5-sunburst | Agent must pass the resolved choice explicitly |
--size | 1k, 2k, 4k, portrait, landscape, square, wide, tall, or literal | Canvas size |
--quality | low, medium, high, auto; 2.5 also xhigh, max | Cost/quality dial; check model-specific limits |
-n, --n | integer | Number of images |
--background | auto, opaque; 2.5 also transparent | Transparency requires PNG or WebP, not JPEG |
--input-fidelity | low, high; omitted by default | Edit-only; explicit 2.5 values are forwarded to the API, not assumed supported |
--moderation | auto, low | Generation moderation setting |
--format | png, jpeg, webp | Output encoding |
--compression | 0-100 | JPEG/WebP compression |
--user | string | Optional end-user identifier |
Quality starting points (not guarantees; keep the user's agreed setting). For 2.5, these take precedence over fixed quality advice in older craft references:
low: cheap drafts and broad exploration; multiple variants require user authorization.medium: normal exploration, style probing, balanced cost.high: CLI default and a candidate for final assets, dense text, diagrams and UI. On 2.5, evaluate against the task requirements; do not assume medium fails or a higher setting always wins.xhigh / max: 2.5-only options for higher-quality work; discuss the cost trade-off before increasing an already agreed quality. Do not use them automatically for budget-conscious requests.Size policy:
1k / 1024x1024portraitlandscape2k4ktall| Mode | Trigger | Endpoint |
|---|---|---|
| Text-to-image | no -i | /v1/images/generations |
| Reference edit | one or more -i | /v1/images/edits |
| Inpaint | -i + -m | /v1/images/edits with mask |
Surface API errors verbatim enough for debugging; exit codes: 0 success, 1 API/refusal, 2 bad args/missing key.
references/gallery.md, then one matching references/gallery-*.md category and its actual prompt text. Read only relevant sections of references/craft.md or the historical references/openai-cookbook.md when needed.references/openai-image-2.5-generation.md (photo/product/illustration), references/openai-image-2.5-layout-and-text.md (text/UI/diagrams/panels), or references/openai-image-2.5-editing.md (references/masks/translation).references/openai-image-2.5.md is an optional index; references/openai-image-2.5-migration.md covers migration/comparison; references/models.md owns current API parameters and validation status.references/templates-gpt-image-2.5.md (community adaptations, not verified outputs). Do not preload them or the old Cookbook for 2.5.Load the smallest useful slice, not both model routes. Add a second task slice only for a genuine hybrid. Historical examples do not override current API notes or the user's model/settings.
--model, endpoint mode, size, quality, output path, and required reference/mask files. Omit --input-fidelity unless explicitly needed; do not assume 2.5 always uses or accepts high.-i paths exist; verify -m exists when used.Preserve Curated vs Author + Source metadata when adapting examples. Add new collected prompts to the Reference Gallery before README promotion.
© wuyoscar, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 44 other files (scripts, references) in skills/gpt-image of wuyoscar/GPT-Image2-Skill.
Open the folder on GitHubat commit 9f8aa1a
GPT Image Generation CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPT Image Generation CLI this skillwuyoscar/GPT-Image2-Skill | 5.7k | — | ~2.5k | Automated safety check: Notes | MIT | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| BlockRun Image GenerationBlockRunAI/ClawRouter | 6.6k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Codex API Image Generatoryc-duan/api-image | 101 | — | ~4.9k | Automated safety check: Pass | MIT | |
| ImagegenJetBrains/skills | 366 | 3 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| BananaTape Image Editor CLINomaDamas/bananatape | 188 | — | ~547 | Automated safety check: Pass | None |
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
BlockRunAI/ClawRouter
Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.
yc-duan/api-image
Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.
JetBrains/skills
A skill your agent uses when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent…
NomaDamas/bananatape
Drives the BananaTape CLI to create, launch, list and delete local AI image editing projects from an agent, including headless smoke tests.
intellectronica/agent-skills
Generate and edit images using OpenAI's GPT Image 1.5 model.
wuyoscar/GPT-Image2-Skill
Analyzes a reference image and writes a prompt that could recreate it in an AI image generator, focusing on the visual traits that most affect similarity.
Works with
Categories
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request. Requests are sorted into generate, edit, inpaint or multi-reference jobs, with the asset type, exact text, aspect ratio, references and budget noted, and the model is settled before any API call.5-sunburst for precise reference edits, and gpt-image-2 for existing Image 2 workflows.
GPT Image Generation CLI fits situations like: generating a poster or typography-heavy image; editing an existing picture from one or more reference images; inpainting part of an image; choosing between GPT Image 2 and the 2.5 models before generating.
Run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a claude-code`. Or copy the skill folder (skills/gpt-image in wuyoscar/GPT-Image2-Skill) into .claude/skills/gpt-image in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a codex`. Or copy the skill folder (skills/gpt-image in wuyoscar/GPT-Image2-Skill) into .agents/skills/gpt-image in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpt-image, .gemini/skills/gpt-image, .github/skills/gpt-image and .opencode/skills/gpt-image in your project.
Going by SKILL.md and its folder, GPT Image Generation CLI needs the command-line tools its instructions call (uv and uvx) and credentials named OPENAI_API_KEY. Our summary lists: Python 3.11 or newer; The gpt-image CLI, or uv or uvx to run it; An OPENAI_API_KEY, since calls may incur OpenAI API charges. Compatibility (from SKILL.md): Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges..
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
GPT Image Generation CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 77k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with GPT Image Generation CLI: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars), Codex API Image Generator (yc-duan/api-image, 101 stars) and Imagegen (JetBrains/skills, 366 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wuyoscar (a GitHub user) maintains it in wuyoscar/GPT-Image2-Skill, which has 5,701 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 30, 2026.
Source: wuyoscar/GPT-Image2-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.