Agent skill

GPT Image Generation CLI

by wuyoscar in wuyoscar/GPT-Image2-Skill

Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

MITAuto-check: notesMedia & Creative

Install GPT Image Generation CLI

skills CLI
$ npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wuyoscar/GPT-Image2-Skill gpt-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wuyoscar/GPT-Image2-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpt-image .claude/skills/gpt-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpt-image
GitHub stars
5.7k
Token cost
~2.5k tokens
SKILL.md length
1,185 words
Files
45 (incl. scripts, references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

  • Works in 8 steps: Classify request and resolve model:… → Choose the reference path: Image 2 keeps… → Refine only as needed: preserve the… → …
  • Generating a poster or typography-heavy image
  • SKILL.md covers Operating loop, Model choice and prompt…, CLI resolution and Key and cost rules, plus 4 more sections
  • Calls uv and uvx; reaches github.com; needs OPENAI_API_KEY

What it does

Requests are sorted into generate, edit, inpaint or multi-reference jobs, with the asset type, exact text, aspect ratio, references and budget noted, and the model is settled before any API call. Three choices are listed: gpt-image-2.5-flare for fast drafts, gpt-image-2.5-sunburst for precise reference edits, and gpt-image-2 for existing Image 2 workflows. Image 2 keeps a gallery-first routine, while a precise 2.5 brief can go straight to generation.

The agent calls the gpt-image command or scripts/generate.py with an explicit --model and never writes its own API wrapper. Before costly or ambiguous calls it offers one to three matched directions with planned size and quality, asking at most one short question at a time. It does not reinstall, overwrite skill folders or write keys or .env files unless you ask for setup, and it ends by reporting file paths, key flags and one refinement idea.

A large reference set holds a craft guide and galleries by theme, such as anime and manga, architecture, brand identity, character design, data visualization, fashion, fine art and infographics. Calls read OPENAI_API_KEY and may incur OpenAI API charges.

When your agent uses it

  • Generating a poster or typography-heavy image
  • Editing an existing picture from one or more reference images
  • Inpainting part of an image
  • Choosing between GPT Image 2 and the 2.5 models before generating

Example prompts

  • “Make a poster for a jazz night with the exact title Blue Hour Sessions.”
  • “Inpaint the sky in photo.png so it shows a sunset.”
  • “Use the 2.5 sunburst model to match the style of ref1.png and ref2.png in a new product shot.”
  • “Which GPT Image model should I use for quick concept drafts?”

Requirements

  • Python 3.11 or newer
  • The gpt-image CLI, or uv or uvx to run it
  • An OPENAI_API_KEY, since calls may incur OpenAI API charges
  • Compatibility (from SKILL.md): Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges.

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Classify request and resolve model: generate, edit, inpaint, or multi-reference; identify asset type, exact text, aspect ratio…
  2. Choose the reference path: Image 2 keeps the gallery-first workflow below. For 2.5, a precise brief needs no reference loading; otherwise…
  3. Refine only as needed: preserve the brief. Add a specific gallery case, craft section or template only to fill a concrete gap; do not load…
  4. Confer when useful: before costly/ambiguous/high-polish calls, present 1–3 matched directions plus planned size/quality; ask at most one…
  5. Preflight, no side effects: use existing CLI/skill if present. Check command availability (command -v gpt-image), installed tool lists…
  6. No blind setup: do not reinstall, overwrite skill folders, create/modify .env, or write API keys unless the user explicitly requested…
  7. Execute via CLI only: call gpt-image or scripts/generate.py with an explicit --model. Do not create a new generate.py, SDK wrapper, or…
  8. Report: output file path(s), key flags, and one concise refinement suggestion if useful.

What it can do on your machine

Read from SKILL.md and the folder at commit 9f8aa1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges.

    From compatibility in the SKILL.md frontmatter.

Context cost

GPT Image Generation CLI loads about 2.5k tokens when it runs, and up to ~79k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 1,185 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~79k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:19
    overwrite skill folders, create/modify `.env`, or write API keys unless the user explicitly requested setup. Global/sha
  • NoteMentions a .env fileSKILL.md:59
    `OPENAI_API_KEY` from process env, then `.env`, then `~/.env` without overriding existing env; successful API calls may
  • NoteMentions a .env fileSKILL.md:62
    set OPENAI_API_KEY`; if a key exists in `.env`/`~/.env`, tell them to remove/rename it for the session rather than worki

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from wuyoscar/GPT-Image2-Skill at commit 9f8aa1a, republished under its MIT licence (© wuyoscar). 1,185 words, ~2,550 tokens.

Download SKILL.mdSave it as .claude/skills/gpt-image/SKILL.md (or your agent's skills folder). This skill also uses 44 other files; get the full folder from GitHub.
name
gpt-image
description
Generate or edit images with GPT Image 2 or 2.5 through the packaged CLI and Reference Gallery. Use for image requests including imprecise 'GPT 2.5' model names, posters, typography, reference edits, and inpainting; resolve the model choice before generation.
compatibility
Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges.

gpt-image

Agent runbook for GPT Image 2 / 2.5 generation/editing. Use the prompt library + packaged CLI. Do not reimplement image API code.

Operating loop

  1. Classify request and resolve model: generate, edit, inpaint, or multi-reference; identify asset type, exact text, aspect ratio, references, safety constraints, and budget/quality. Apply the model-choice rules below before any API call.
  2. Choose the reference path: Image 2 keeps the gallery-first workflow below. For 2.5, a precise brief needs no reference loading; otherwise choose one short task slice.
  3. Refine only as needed: preserve the brief. Add a specific gallery case, craft section or template only to fill a concrete gap; do not load them as a bundle for 2.5.
  4. Confer when useful: before costly/ambiguous/high-polish calls, present 1–3 matched directions plus planned size/quality; ask at most one concise question at a time. Skip long discussion for precise “generate now” requests with a resolved model.
  5. Preflight, no side effects: use existing CLI/skill if present. Check command availability (command -v gpt-image), installed tool lists when the tool manager exists, or the runtime’s own skill registry when available. Do not assume a local home path in cloud/hosted runtimes.
  6. No blind setup: do not reinstall, overwrite skill folders, create/modify .env, or write API keys unless the user explicitly requested setup. Global/shared installs are opt-in only.
  7. Execute via CLI only: call gpt-image or scripts/generate.py with an explicit --model. Do not create a new generate.py, SDK wrapper, or ad-hoc script for normal image requests.
  8. Report: output file path(s), key flags, and one concise refinement suggestion if useful.

Fast path: confirmed 2.5 model + precise prompt + “generate now” → preflight and CLI, without a mandatory reference/craft pass. Do not reconfirm an exact valid model.

Model choice and prompt adaptation

ChoiceAPI model IDSuggested use
Flaregpt-image-2.5-flareFast general generation and drafts
Sunburstgpt-image-2.5-sunburstPrecise reference edits and detailed control
Image 2gpt-image-2Existing Image 2 workflows and compatibility
  • If the model is absent, ambiguous (such as “GPT 2.5”), or misspelled, ask one clear question offering Flare, Sunburst, and Image 2 with these trade-offs, then wait. For a typo, suggest the likely intended choice without silently correcting it. Do not treat gpt-image-2.5 as an API model ID.
  • Use an exact supported model ID, an unambiguous choice from this menu, or the user's already confirmed choice for the current task without asking again. If the user explicitly says “you choose,” explain the pick briefly and proceed; consider their task and budget rather than always selecting the most expensive settings.
  • Always pass the chosen ID through --model. The CLI retains gpt-image-2 as its backward-compatible default, but that default is not a substitute for the agent resolving the user's choice.
  • Keep model confirmation, cost discussion, and API flags separate from the final image prompt. Use the model-specific reference path below only when it helps the task. Preserve the user's exact text, intended content, reference identity, and edit invariants; obtain confirmation before material prompt changes.
  • Generate one image unless the user requested more. Do not silently switch models, change material prompt content, or run multiple model/prompt variants or comparisons. Ask before additional paid variants. For an authorized same-prompt comparison, keep the prompt, size, and quality identical unless the user asks to vary them.
  • Consult references/models.md when parameter support or validation status needs checking. On an invalid-model, access/403, quota, or policy failure, report the failure and stop; do not switch models or rewrite the prompt to retry automatically.

CLI resolution

Preferred call order:

bash
# Existing CLI on PATH
gpt-image --model MODEL_ID -p "PROMPT" [-f OUT] [-i REF...] [-m MASK] [options]

# Installed skill folder; use runtime-provided skill path when available
uv run "$SKILL_DIR/scripts/generate.py" --model MODEL_ID -p "PROMPT" [-f OUT] [-i REF...] [-m MASK] [options]

# Direct transient CLI when the user requested setup/one-off CLI execution
uvx --from git+https://github.com/wuyoscar/gpt_image_2_skill gpt-image --model MODEL_ID -p "PROMPT" [options]

scripts/generate.py is a launcher: repo-local src/gpt_image_cli → installed gpt-image → PATH gpt-image → transient uvx/uv fallback.

Key and cost rules

  • CLI reads OPENAI_API_KEY from process env, then .env, then ~/.env without overriding existing env; successful API calls may bill the user’s OpenAI account.
  • If host/runtime has native platform-managed image generation and the user wants that path, use the host tool instead of this CLI.
  • If OPENAI_API_KEY is unset, report missing key or use host-native generation when requested; do not write secrets.
  • If user wants to avoid local-key use, respect unset OPENAI_API_KEY; if a key exists in .env/~/.env, tell them to remove/rename it for the session rather than working around it.
  • Never print secret values.
Show full SKILL.md (496 more words)Show less

Flags

FlagValuesUse
-p, --promptstringRequired prompt/edit instruction
-f, --filepathOutput path; auto-named if omitted
-i, --imagerepeatable pathUse edits endpoint; supports multiple references
-m, --maskPNG pathInpaint with alpha mask; requires -i
--modelgpt-image-2, gpt-image-2.5-flare, gpt-image-2.5-sunburstAgent must pass the resolved choice explicitly
--size1k, 2k, 4k, portrait, landscape, square, wide, tall, or literalCanvas size
--qualitylow, medium, high, auto; 2.5 also xhigh, maxCost/quality dial; check model-specific limits
-n, --nintegerNumber of images
--backgroundauto, opaque; 2.5 also transparentTransparency requires PNG or WebP, not JPEG
--input-fidelitylow, high; omitted by defaultEdit-only; explicit 2.5 values are forwarded to the API, not assumed supported
--moderationauto, lowGeneration moderation setting
--formatpng, jpeg, webpOutput encoding
--compression0-100JPEG/WebP compression
--userstringOptional end-user identifier

Quality starting points (not guarantees; keep the user's agreed setting). For 2.5, these take precedence over fixed quality advice in older craft references:

  • low: cheap drafts and broad exploration; multiple variants require user authorization.
  • medium: normal exploration, style probing, balanced cost.
  • high: CLI default and a candidate for final assets, dense text, diagrams and UI. On 2.5, evaluate against the task requirements; do not assume medium fails or a higher setting always wins.
  • xhigh / max: 2.5-only options for higher-quality work; discuss the cost trade-off before increasing an already agreed quality. Do not use them automatically for budget-conscious requests.

Size policy:

  • default/social square: 1k / 1024x1024
  • poster/mobile/beauty: portrait
  • landscape/gameplay/photo: landscape
  • print/paper figure: 2k
  • widescreen hero: 4k
  • vertical story/banner: tall

Endpoint routing

ModeTriggerEndpoint
Text-to-imageno -i/v1/images/generations
Reference editone or more -i/v1/images/edits
Inpaint-i + -m/v1/images/edits with mask

Surface API errors verbatim enough for debugging; exit codes: 0 success, 1 API/refusal, 2 bad args/missing key.

Reference loading

  • Image 2: open references/gallery.md, then one matching references/gallery-*.md category and its actual prompt text. Read only relevant sections of references/craft.md or the historical references/openai-cookbook.md when needed.
  • Image 2.5: no mandatory references for a precise brief. If guidance is needed, choose one directly: references/openai-image-2.5-generation.md (photo/product/illustration), references/openai-image-2.5-layout-and-text.md (text/UI/diagrams/panels), or references/openai-image-2.5-editing.md (references/masks/translation).
  • Other questions: references/openai-image-2.5.md is an optional index; references/openai-image-2.5-migration.md covers migration/comparison; references/models.md owns current API parameters and validation status.
  • Extra inspiration only: gallery cases, a targeted craft section, or references/templates-gpt-image-2.5.md (community adaptations, not verified outputs). Do not preload them or the old Cookbook for 2.5.

Load the smallest useful slice, not both model routes. Add a second task slice only for a genuine hybrid. Historical examples do not override current API notes or the user's model/settings.

Verification

  • Before API call: check the resolved model and explicit --model, endpoint mode, size, quality, output path, and required reference/mask files. Omit --input-fidelity unless explicitly needed; do not assume 2.5 always uses or accepts high.
  • After CLI call: report path(s) printed by the CLI and surface stderr on failure.
  • For edits/inpaints: verify -i paths exist; verify -m exists when used.

Preserve Curated vs Author + Source metadata when adapting examples. Add new collected prompts to the Reference Gallery before README promotion.

© wuyoscar, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 44 other files (scripts, references) in skills/gpt-image of wuyoscar/GPT-Image2-Skill.

  • SKILL.md
  • agents/openai.yaml
  • references/craft.md
  • references/gallery-anime-and-manga.md
  • references/gallery-architecture-and-interior.md
  • references/gallery-beauty-and-lifestyle.md
  • references/gallery-brand-systems-and-identity.md
  • references/gallery-character-design.md
  • references/gallery-cinematic-and-animation.md
  • references/gallery-cinematic-film-references.md
  • references/gallery-data-visualization.md
  • references/gallery-edit-endpoint-showcase.md
  • references/gallery-events-and-experience.md
  • references/gallery-fashion-editorial.md
  • references/gallery-fine-art-painting.md
  • references/gallery-gaming.md
  • references/gallery-illustration.md
  • references/gallery-infographics-and-field-guides.md
  • references/gallery-ink-and-chinese.md
  • … and 26 more

Open the folder on GitHubat commit 9f8aa1a

Compare with similar skills

GPT Image Generation CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPT Image Generation CLI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPT Image Generation CLI this skillwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT
Codex API Image Generatoryc-duan/api-image101—~4.9kAutomated safety check: PassMIT
ImagegenJetBrains/skills3663 repos~2.5kAutomated safety check: PassApache-2.0
BananaTape Image Editor CLINomaDamas/bananatape188—~547Automated safety check: PassNone

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Codex API Image Generator

    yc-duan/api-image

    Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.

    101 GitHub stars~4.9k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Imagegen

    JetBrains/skills

    Official

    A skill your agent uses when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent…

    366 GitHub starsUsed in 3 repos~2.5k tokens
    Media & CreativeAuto-check passed
  • BananaTape Image Editor CLI

    NomaDamas/bananatape

    Drives the BananaTape CLI to create, launch, list and delete local AI image editing projects from an agent, including headless smoke tests.

    188 GitHub stars~547 tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Gpt Image 1 5

    intellectronica/agent-skills

    Generate and edit images using OpenAI's GPT Image 1.5 model.

    295 GitHub stars~1.5k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed

More from wuyoscar/GPT-Image2-Skill

  • Image to Prompt Reverse Engineering

    wuyoscar/GPT-Image2-Skill

    Analyzes a reference image and writes a prompt that could recreate it in an AI image generator, focusing on the visual traits that most affect similarity.

    5.7k GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about GPT Image Generation CLI

What does GPT Image Generation CLI do?

Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request. Requests are sorted into generate, edit, inpaint or multi-reference jobs, with the asset type, exact text, aspect ratio, references and budget noted, and the model is settled before any API call.5-sunburst for precise reference edits, and gpt-image-2 for existing Image 2 workflows.

When should I use GPT Image Generation CLI?

GPT Image Generation CLI fits situations like: generating a poster or typography-heavy image; editing an existing picture from one or more reference images; inpainting part of an image; choosing between GPT Image 2 and the 2.5 models before generating.

How do I install GPT Image Generation CLI in Claude Code?

Run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a claude-code`. Or copy the skill folder (skills/gpt-image in wuyoscar/GPT-Image2-Skill) into .claude/skills/gpt-image in your project. Claude Code loads it when a task matches its description.

How do I install GPT Image Generation CLI in Codex?

Run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a codex`. Or copy the skill folder (skills/gpt-image in wuyoscar/GPT-Image2-Skill) into .agents/skills/gpt-image in your project. Codex loads it when a task matches its description.

Can I use GPT Image Generation CLI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wuyoscar/GPT-Image2-Skill --skill gpt-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpt-image, .gemini/skills/gpt-image, .github/skills/gpt-image and .opencode/skills/gpt-image in your project.

What does GPT Image Generation CLI need to run?

Going by SKILL.md and its folder, GPT Image Generation CLI needs the command-line tools its instructions call (uv and uvx) and credentials named OPENAI_API_KEY. Our summary lists: Python 3.11 or newer; The gpt-image CLI, or uv or uvx to run it; An OPENAI_API_KEY, since calls may incur OpenAI API charges. Compatibility (from SKILL.md): Requires Python 3.11+ and either `gpt-image`, `uv`, or `uvx`. CLI/API calls read `OPENAI_API_KEY` and may incur OpenAI API charges..

Does GPT Image Generation CLI access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is GPT Image Generation CLI safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does GPT Image Generation CLI use?

GPT Image Generation CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPT Image Generation CLI use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 77k tokens, read only when the agent opens those files.

What are the alternatives to GPT Image Generation CLI?

Skills that share tags, products or a category with GPT Image Generation CLI: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars), Codex API Image Generator (yc-duan/api-image, 101 stars) and Imagegen (JetBrains/skills, 366 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPT Image Generation CLI?

wuyoscar (a GitHub user) maintains it in wuyoscar/GPT-Image2-Skill, which has 5,701 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 30, 2026.

Source: wuyoscar/GPT-Image2-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.