Agent skill

GPT Image Prompt Guide

by code-yeongyu in code-yeongyu/senpi

Prompt-crafting guide for gpt-image-2.5: which image tool to call, which model to pick, and how to write prompts, edit with references and refine over turns.

MITAuto-check passedMedia & Creative

Install GPT Image Prompt Guide

skills CLI
$ npx skills add code-yeongyu/senpi --skill gpt-image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install code-yeongyu/senpi gpt-image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/coding-agent/src/core/extensions/builtin/imagegen/skill .claude/skills/gpt-image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpt-image-gen
GitHub stars
472
Token cost
~2.3k tokens
SKILL.md length
1,304 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Prompt-crafting guide for gpt-image-2.5: which image tool to call, which model to pick, and how to write prompts, edit with references and refine over turns.

  • Generating a logo, product image or diagram where text must render correctly
  • SKILL.md covers Which tool, Model selection, Prompt crafting and Quality, size, and format, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Editing an image from reference photos while keeping the subject faithful

What it does

Before generating, the agent checks which image surface actually exists in the current tool set. A native image_generation server tool is preferred when present; otherwise the generate_image tool sends the prompt and any reference images to a configured OpenAI-compatible endpoint and saves a png, jpeg or webp file. Tool state can change mid-session, so the tools visible now win over what the page said at startup.

Three models are described: gpt-image-2.5-sunburst, the default, aimed at quality, precise edits, reference fidelity and dense text or diagrams; gpt-image-2.5-flare, a smaller and faster one for latency or volume; and the previous gpt-image-2, whose quality tiers stop at high. For latency-sensitive work the same prompt is run on Flare and kept only if the result still meets the bar.

Prompting starts from the deliverable and describes what is visible. A specific request is normalized without padding, keeping every stated requirement and adding no objects, people, brands or slogans, while a generic request such as a bakery logo gets a concrete medium, composition, lighting, palette and materials. The description also lists exact text rendering, transparent assets, output formats and multi-turn refinement as covered topics.

When your agent uses it

  • Generating a logo, product image or diagram where text must render correctly
  • Editing an image from reference photos while keeping the subject faithful
  • Choosing between the quality model and the faster model for a batch of images

Example prompts

  • “Generate a flat logo for a bakery called Rye and Co, with the name spelled exactly.”
  • “Edit the attached product photo onto a plain white background and keep the label sharp.”
  • “Make twenty quick concept thumbnails for a landing page using the faster model.”

Requirements

  • The native image_generation tool, or generate_image with an OpenAI-compatible endpoint

What it can do on your machine

Read from SKILL.md and the folder at commit 6073042. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GPT Image Prompt Guide loads about 2.3k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from code-yeongyu/senpi at commit 6073042, republished under its MIT licence (© code-yeongyu). 1,304 words, ~2,335 tokens.

Download SKILL.mdSave it as .claude/skills/gpt-image-gen/SKILL.md (or your agent's skills folder).
name
gpt-image-gen
description
MUST read before generating images. Prompt-crafting guide for gpt-image-2.5 covering tool routing (native image_generation server tool vs the generate_image tool), model and quality selection, prompt structure, exact text rendering, reference-image editing, transparent assets, output formats, and multi-turn refinement.

GPT Image Generation

How to get the image right the first time, and how to fix it fast when it is not. Read this before your first generation call.

Which tool

When image generation tooling is present in your tool set, pick the surface that actually exists right now:

  • If a native image_generation server tool is available, use it. The provider runs generation server-side and returns the image in the response stream.
  • Otherwise, call the generate_image tool. It sends your prompt and optional reference images to the configured OpenAI-compatible endpoint and saves the result as a png, jpeg, or webp file. The model, quality, size, background, and format controls below apply to generate_image, not to the native server tool.

Check your current tool set before choosing. Tool state can change mid-session (model switch, credential change), so trust the tools you can see over what this page said at startup. If both surfaces ever appear, prefer the native server tool. When the request is clear, generate directly instead of asking for confirmation; ask only when a required reference image is missing.

Model selection

  • gpt-image-2.5-sunburst (default): the base model, optimized for quality, above gpt-image-2. Use it for precise edits, fidelity to reference subjects and products, dense text or diagrams, and final production assets.
  • gpt-image-2.5-flare: the small model, optimized for speed, with quality comparable to gpt-image-2. Reach for it when latency or volume matters more than the last increment of quality.
  • gpt-image-2: previous generation; quality tiers stop at high, transparency is preview-level.

The default is the better model on purpose. When a workflow turns out to be latency-sensitive, run the same prompt and inputs on Flare and switch only if the result still meets the bar.

Prompt crafting

Start from the deliverable, then describe what is visible. Match the prompt's specificity to the user's request:

  • A specific request (exact text, a named layout, a product photo brief) is normalized, not padded. Keep every stated requirement; do not add objects, people, brands, slogans, or story beats the user did not ask for.
  • A generic request ("a logo for a bakery") gets tasteful concreteness: medium, composition, lighting, palette, materials. Stay within what the request implies.

Cover these when they matter: the deliverable and its use (poster, product shot, UI mockup, diagram, icon); the subject with concrete physical detail; one medium or style, stated plainly; composition and camera (framing, angle, focal length or its visual equivalent); lighting and color; mood; background and how much of it is in focus. For people, describe body framing, gaze, and interaction ("full body visible, feet included", "looking down at the open book").

A short specific prompt is fine. For complex requests, organize the prompt into labeled sections (scene, subject, details, constraints); a descriptive paragraph and a labeled spec express the same intent, so choose whichever is easier to read and update. Request "photorealistic" or "real photograph" explicitly when that is the goal; camera specifications are cues for appearance, not a guarantee of exact optics.

Rendering text in the image

Put the exact string in double quotes, say how many times it appears, and describe its position and typography: A weathered wooden sign above the door reads "OPEN TIL LATE" once, in hand-painted white serif letters, centered, slightly faded. Spell unusual words or brand names letter by letter when they matter. Add "no other text" so the model does not invent captions. Keep on-image text short; long passages smear. Use quality medium or high for small text, dense labels, or multiple fonts, then check spelling and legibility in the result.

Anti-patterns
  • Contradictory instructions. "Photorealistic watercolor" or "minimalist scene packed with detail" averages two opposing goals, and you get neither.
  • Element overcrowding. Past roughly five or six distinct elements, small ones get dropped or mangled. Cut before you add.
  • Style-list collisions. "In the style of anime, oil painting, and pixel art" is three prompts. Choose one style per image and generate variants separately.
  • Invented content. Extra characters, logos, watermarks, or slogans the user never asked for. State exclusions when the model tends to add them ("no watermark, no logo").

Quality, size, and format

quality accepts low, medium, high, xhigh, max, or auto (default). Explore with auto or a lower tier; raise it for the final image only while it fixes an unmet requirement, since a higher tier does not guarantee a better image and costs more latency. xhigh and max exist only on GPT Image 2.5; the API rejects them on gpt-image-2.

size defaults to auto. Presets: 1024x1024, 1536x1024 (landscape), 1024x1536 (portrait). Custom WIDTHxHEIGHT works when both edges are multiples of 16, the aspect ratio stays within 1:3 to 3:1, neither edge exceeds 3840, and the total pixel count is between 655,360 and 8,294,400 (4K is 3840x2160 or 2160x3840, not 3840x3840). Outputs above 2560x1440 are experimental; square images are usually fastest.

output_format: png (default) is lossless and keeps alpha; jpeg is the fastest and smallest for photos with no transparency; webp is small and keeps alpha. output_compression (0-100) applies to jpeg and webp only. The output_path extension must match the format or be omitted.

Show full SKILL.md (462 more words)Show less
Transparent assets

For logos, stickers, cutouts, and UI icons set background: "transparent" with output_format png or webp, and also ask for an isolated subject in the prompt ("isolated on a transparent background, no drop shadow"). A drawn checkerboard is not transparency: after saving, confirm the file has an alpha channel (the result reports Background: transparent when the provider confirms it) and inspect edges, hair, glass, and shadows. Repeat the transparent requirement on every later edit of that asset. On gpt-image-2 transparency is a preview feature; prefer a 2.5 model and verify the alpha channel either way.

Editing with reference images

Pass reference_image_paths to edit an existing image or to use images as references: 1-5 local png, jpeg, or webp files, each at most 50 MB, absolute or relative to the working directory. With references the tool sends an image-edit request; without them it generates from text. Assign each input a role by position ("image 1 is the product to keep unchanged; image 2 is the style reference") and say how they combine.

Separate the change from the constraints. Say "change only X" and list what must stay fixed: identity, geometry, layout, lighting, labels, colors, surrounding objects. Describe the desired END STATE, not only the delta.

For a local repaint, add mask_image_path: a png with an alpha channel the same size as the first reference, where transparent pixels mark the area to repaint. Masking is prompt-guided, so restate the edit and the preserved regions in words too.

json
{
  "prompt": "The same red fox from image 1, now wearing a knitted blue scarf, in an eye-level wildlife photograph. Change only the scarf. Preserve its face, fur markings, pose, and forest background; soft overcast light, muted woodland colors, shallow depth of field.",
  "model": "gpt-image-2.5-sunburst",
  "reference_image_paths": ["art/fox.png"],
  "quality": "high",
  "size": "2048x1152",
  "output_path": "art/fox-with-scarf.png"
}

Use a new output_path so the source stays intact. Only the native server tool returns a revised_prompt (the mainline model's rewrite); the client tool reports one in details.revisedPrompts only when the provider sends it, so judge the saved image against your request rather than waiting for a rewrite.

Multi-turn refinement

  • Generate one image first and inspect it against every clause of the request before spending more.
  • Refine one thing per turn: pass the previous output as the next reference_image_paths entry, request the single change, and restate the constraints that must survive. "Same style as before" carries context, but restate critical details if the result drifts.
  • For a recurring character or product, establish one reference image and reuse it in every scene, repeating its defining details.
  • n produces variants of one prompt; distinct assets need distinct calls. Use n only after the prompt is proven.
  • Keep the full prompt text in the conversation as the reproducibility record.

Check the result

Before handing the image over: required text is spelled correctly and legible; identities, product shapes, and labels survived; the edit changed only what was requested; a transparent asset has real alpha rather than a painted background; diagram labels and relationships are correct. Fix a miss by editing the prompt or masking the region, not by regenerating blindly.

© code-yeongyu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/coding-agent/src/core/extensions/builtin/imagegen/skill of code-yeongyu/senpi.

Open the folder on GitHubat commit 6073042

Compare with similar skills

GPT Image Prompt Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPT Image Prompt Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPT Image Prompt Guide this skillcode-yeongyu/senpi472—~2.3kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
Openai Image Gentrpc-group/trpc-agent-go1.8k13 repos~843Automated safety check: PassApache-2.0
Imagegentheowenyoung/home1154 repos~4.8kAutomated safety check: PassApache-2.0
Image Generationonyx-dot-app/onyx32k1 repos~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 8 days ago
    Media & CreativeAuto-check: notes
  • Openai Image Gen

    trpc-group/trpc-agent-go

    Batch-generate images via OpenAI Images API. An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 13 repos~843 tokens
    Media & CreativeAuto-check passed
  • Imagegen

    theowenyoung/home

    Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts.

    115 GitHub starsUsed in 4 repos~4.8k tokens
    Media & CreativeAuto-check passed
  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from code-yeongyu/senpi

  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    472 GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Merge Upstream into Fork

    code-yeongyu/senpi

    Syncs a fork branch with its upstream remote using a history-preserving merge commit, with no rebase and no force push.

    472 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Bun 1.4 Builtins Guide

    code-yeongyu/senpi

    Points the agent at Bun 1.4 built-in APIs before it installs an npm package, so image, browser, markdown, cron, PTY and test work uses what Bun already ships.

    472 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Worker brief for implementing one pre-assigned feature in the senpi todotools built-in extension, with strict scope, typing, testing and git-safety rules.

    472 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Senpi Release Publishing

    code-yeongyu/senpi

    Walks the canonical CalVer release flow for senpi, from a clean main checkout through changelog audit, checks, tag push, GitHub Release and npm publishing.

    472 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Tmux Manual QA Worker

    code-yeongyu/senpi

    Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.

    472 GitHub stars~1.6k tokensUpdated today
    Auto-check passed

Works with

Questions about GPT Image Prompt Guide

What does GPT Image Prompt Guide do?

Prompt-crafting guide for gpt-image-2.5: which image tool to call, which model to pick, and how to write prompts, edit with references and refine over turns. Before generating, the agent checks which image surface actually exists in the current tool set. A native image_generation server tool is preferred when present; otherwise the generate_image tool sends the prompt and any reference images to a configured OpenAI-compatible endpoint and saves a png, jpeg or webp file.

When should I use GPT Image Prompt Guide?

GPT Image Prompt Guide fits situations like: generating a logo, product image or diagram where text must render correctly; editing an image from reference photos while keeping the subject faithful; choosing between the quality model and the faster model for a batch of images.

How do I install GPT Image Prompt Guide in Claude Code?

Run `npx skills add code-yeongyu/senpi --skill gpt-image-gen -a claude-code`. Or copy the skill folder (packages/coding-agent/src/core/extensions/builtin/imagegen/skill in code-yeongyu/senpi) into .claude/skills/gpt-image-gen in your project. Claude Code loads it when a task matches its description.

How do I install GPT Image Prompt Guide in Codex?

Run `npx skills add code-yeongyu/senpi --skill gpt-image-gen -a codex`. Or copy the skill folder (packages/coding-agent/src/core/extensions/builtin/imagegen/skill in code-yeongyu/senpi) into .agents/skills/gpt-image-gen in your project. Codex loads it when a task matches its description.

Can I use GPT Image Prompt Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add code-yeongyu/senpi --skill gpt-image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpt-image-gen, .gemini/skills/gpt-image-gen, .github/skills/gpt-image-gen and .opencode/skills/gpt-image-gen in your project.

What does GPT Image Prompt Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: GPT Image Prompt Guide is instructions for the agent only. Our summary lists: The native image_generation tool, or generate_image with an OpenAI-compatible endpoint.

Does GPT Image Prompt Guide access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is GPT Image Prompt Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GPT Image Prompt Guide use?

GPT Image Prompt Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPT Image Prompt Guide use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPT Image Prompt Guide?

Skills that share tags, products or a category with GPT Image Prompt Guide: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars), Openai Image Gen (trpc-group/trpc-agent-go, 1.8k stars) and Imagegen (theowenyoung/home, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPT Image Prompt Guide?

code-yeongyu (a GitHub user) maintains it in code-yeongyu/senpi, which has 472 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: code-yeongyu/senpi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.