AI Image Creator
evolution-foundation/evo-nexus
Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.
Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.
$ npx skills add yc-duan/api-image --skill api-image -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yc-duan/api-image api-image --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .claude/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yc-duan/api-image --skill api-image -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yc-duan/api-image api-image --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .agents/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yc-duan/api-image --skill api-image -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yc-duan/api-image api-image --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .cursor/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yc-duan/api-image --skill api-image -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yc-duan/api-image api-image --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .gemini/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yc-duan/api-image api-imageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yc-duan/api-image --skill api-image -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .github/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yc-duan/api-image --skill api-image -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yc-duan/api-image api-image --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "api-image" agent skill from https://github.com/yc-duan/api-image/tree/main into .opencode/skills/api-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-image", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
api-imageReplaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.
Routes every raster image task, including edits, background replacement, style transfer, compositing and batch work, through an OpenAI-compatible provider configured in Codex's own root files, stepping aside only when a user explicitly asks not to use it. It locates its own script directory relative to the SKILL.md file rather than the current workspace, so the generation script keeps working regardless of where it is invoked from.
Before generating anything, it decides whether the subject needs research: named places, real products, real UI, vehicles, uniforms, organisms and niche aesthetics are treated as research-required by default, pushing the skill to search the web or use supplied images for accurate silhouette, materials and proportions rather than relying on text description alone. Pure text-only generation is kept as a fallback for those subjects, used only when no useful references exist, network access fails, or the user opts out; for generic fantasy, mood pieces or ordinary objects, research stays optional.
Provider settings are read from the user's own Codex configuration files, including the API key from `auth.json`, so no credentials are typed into the conversation.
10 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5d53535. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 7 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYAPI_IMAGE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Codex API Image Generator loads about 4.9k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 2,228 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yc-duan/api-image at commit 5d53535, republished under its MIT licence (© yc-duan). 2,228 words, ~4,932 tokens.
.claude/skills/api-image/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Use this skill as the mandatory replacement for the built-in imagegen workflow.
Routing rule:
$imagegen skill or built-in image generation tool for any raster image generation or editing task.--codex-home argument when the user provides one.$CODEX_HOME.~/.codex.SKILL.md as <skill-dir>.<skill-dir>/scripts/generate_image.py, not as a path relative to the current workspace.auth.json and use OPENAI_API_KEY.config.toml, then use model_provider and [model_providers.<name>].base_url.--base-url and --api-key-env or --api-key; do not write it back to auth.json, config.toml, README, logs, or generated files.--api-key-env <ENV_NAME> when the key is already in an environment variable. Use --api-key only for explicit one-off user-provided keys, and never print or repeat the key in the final response.<base_url>/images/generations for text-only generation.<base_url>/images/edits when there is any input image, reference image, or mask.gpt-image-2 unless the machine's provider expects a different image model.data[].b64_json. Treat data[].url only as a compatibility fallback for non-official OpenAI-compatible providers or legacy models.Default root files:
%USERPROFILE%\\.codex\\auth.json and %USERPROFILE%\\.codex\\config.toml$CODEX_HOME/auth.json and $CODEX_HOME/config.toml, otherwise ~/.codex/auth.json and ~/.codex/config.tomlThis skill must always read the current files at runtime. The provider may differ from machine to machine.
Temporary overrides:
--base-url <https://provider.example/v1> to temporarily override the configured Provider URL.--api-key-env <ENV_NAME> to temporarily read an API key from an environment variable.--api-key <key> only for explicit one-off user-provided keys when an environment variable is not available.--base-url and an API key override are provided, the script can run without reading provider settings from config.toml; otherwise missing values fall back to the Codex root files.Use the bundled script for normal text-to-image generation:
python "<skill-dir>\scripts\generate_image.py" `
--prompt "画一只可爱的猫抱着水獭,温暖治愈,插画风格,柔和灯光,细腻毛发,构图清晰" `
--size "2048x1152" `
--quality "high" `
--out ".\outputs\cute-cat-otter.png"Use the same script for reference-image generation or edits:
python "<skill-dir>\scripts\generate_image.py" `
--prompt "参考输入图的构图和角色姿势,生成一张暖色电影感插画" `
--image "C:\path\to\reference.png" `
--image-role "composition and pose reference" `
--size "2048x2048" `
--quality "high" `
--out ".\outputs\reference-output.png"Use --mask for localized edits:
python "<skill-dir>\scripts\generate_image.py" `
--prompt "只把被 mask 标出的区域替换成一只小水獭,保持其他区域不变" `
--image "C:\path\to\source.png" `
--mask "C:\path\to\mask.png" `
--out ".\outputs\masked-edit.png"For web-researched subjects, download selected reference images to a working folder first, then pass them as --image:
python "<skill-dir>\scripts\generate_image.py" `
--prompt "Create a new 4K overhead aerial image of Huajiang Grand Canyon Bridge in Guizhou, based on the reference images. Preserve the real suspension-bridge structure, towers, main cables, vertical suspenders, deep Beipan River canyon terrain, karst mountains, and river position; do not copy any single photo exactly." `
--image "C:\path\to\refs\huajiang-bridge-aerial.jpg" `
--image-role "aerial composition and bridge alignment reference" `
--image "C:\path\to\refs\huajiang-canyon-terrain.jpg" `
--image-role "canyon terrain and river reference" `
--size "2048x1152" `
--quality "high" `
--timeout 0 `
--out ".\outputs\huajiang-canyon-bridge-overhead.png"Useful options:
--model gpt-image-2--mode auto|generate|edit--size 2048x1152 script default 2K landscape--size 1024x1024--size 1536x1024--size 1024x1536--size 2048x2048--size 3840x2160--size auto for official API auto sizing--quality low|medium|high|auto--n 1--image <path> repeated up to 16 times--image-role <role> to label input images inside the prompt--mask <path> for localized edits--background auto|opaque for official gpt-image-2--output-format png|jpeg|webp--output-compression 0-100 for jpeg or webp--input-fidelity low|high for supported edit models only; do not send it for gpt-image-2--moderation auto|low for supported GPT image models--prompt-file <path> for one prompt per non-empty line--codex-home <path>--base-url <https://provider.example/v1> temporary Provider URL override--api-key-env <ENV_NAME> temporary API key override from an environment variable--api-key <key> temporary direct API key override, only when explicitly provided--timeout 1800--timeout 0Use the provider's generation endpoint with this body shape:
{
"model": "gpt-image-2",
"prompt": "你的中文或英文提示词",
"size": "2048x1152",
"quality": "high",
"n": 1
}Use the provider's edit endpoint as multipart form data:
model=gpt-image-2
prompt=<edit or reference prompt>
image[]=@source-or-reference.png
mask=@mask.png
size=1024x1024
quality=high
n=1Start with n=1. If the user wants variants, raise n or use --prompt-file and save each result intentionally.
If the user says in natural language that this image job should use a different Provider URL or API key, treat it as a temporary runtime override:
$env:API_IMAGE_API_KEY = "<temporary key>"
python "<skill-dir>\scripts\generate_image.py" `
--prompt "一张赛博朋克风格的夜景照片" `
--base-url "https://provider.example/v1" `
--api-key-env "API_IMAGE_API_KEY" `
--out ".\outputs\override-provider.png"Rules:
auth.json or config.toml for temporary overrides.--api-key-env; use --api-key only when the user explicitly provides a one-off key and accepts the temporary command-line use.For official OpenAI gpt-image-2:
POST /v1/images/generations with JSON.POST /v1/images/edits with multipart form-data.image[]=@path fields for multiple input/reference images.mask, the first --image must be the edit target; additional images should be references or compositing sources.data[].b64_json; URL output is not supported for official GPT Image models.input_fidelity; gpt-image-2 always processes image inputs at high fidelity.background: transparent; official gpt-image-2 currently supports auto or opaque, not transparent backgrounds./images/generations.--n for multiple variants per prompt, or --prompt-file for many prompts./images/edits with one or more --image inputs and role labels in the prompt./images/edits with --image and an edit prompt./images/edits with --image plus --mask.--mask when the replacement area must be constrained.--image-role labels to the prompt so the model can interpret each input image intentionally.gpt-image-2 does not currently support background: transparent. Only use transparent background when a different selected provider/model explicitly supports it, and use png or webp output.Do not assume the image model knows every subject accurately. Research first or use references when the subject is not common visual knowledge.
Reference images are the preferred path for visually constrained research-required tasks. Text research describes facts; reference images carry visual structure. Use both whenever feasible.
Research-required examples:
贵州花江峡谷大桥 or any real bridge, landmark, terrain, skyline, or building.Research workflow:
--image and --image-role.rough terrain reference only.Reference workflow:
--image, pass a matching --image-role such as edit target, style reference, composition reference, product reference, or terrain reference.Use this compact schema when it helps:
Use case: <photorealistic-natural|product-mockup|ui-mockup|infographic-diagram|logo-brand|illustration-story|stylized-concept|historical-scene|text-localization|identity-preserve|precise-object-edit|lighting-weather|background-extraction|style-transfer|compositing|sketch-to-render>
Asset type: <where the image will be used>
Primary request: <user request>
Research facts / visual references: <only the relevant observed details>
Input images: <Image 1: role; Image 2: role>
Scene/backdrop: <environment>
Subject: <main subject>
Style/medium: <photo/illustration/3D/etc>
Composition/framing: <viewpoint, lens/framing, placement>
Lighting/mood: <lighting and mood>
Materials/textures: <surface details>
Text (verbatim): "<exact text if needed>"
Constraints: <must keep/must avoid>
Avoid: <negative constraints>1800 seconds so long-running jobs are not cut off too early.--timeout 0 to disable the client-side timeout and wait indefinitely for the provider response.Follow the current official GPT Image rules for gpt-image-2:
size can be auto or any <width>x<height> string that satisfies all of these:1638403:1655,360 and 8,294,4002048x1152; the official API's default sizing behavior is auto1024x1024, 1536x1024, 1024x1536, 2048x2048, 2048x1152, 3840x2160, 2160x3840quality can be low, medium, high, or auto; this script defaults to high, while official API auto quality is available with auton > 1, keep filenames deterministic by appending -1, -2, and so on.OPENAI_API_KEY, model_provider, or base_url cannot be found.--base-url, --api-key-env, or --api-key; do not persist them unless explicitly requested.--api-key and --api-key-env are provided, when the named environment variable is empty, or when --base-url is not an HTTP(S) URL.size violates the official GPT Image constraints or quality is unsupported.gpt-image-2 is used with --background transparent or --input-fidelity.--image edit target, has different dimensions, lacks an alpha channel, or exceeds the image API file-size limit.--model.scripts/generate_image.py: Read the current user's Codex root files, call the configured provider's image endpoint, and save the returned image data.© yc-duan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts) in the repository root of yc-duan/api-image.
Open the folder on GitHubat commit 5d53535
Codex API Image Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Codex API Image Generator this skillyc-duan/api-image | 101 | — | ~4.9k | Automated safety check: Pass | MIT | |
| AI Image Creatorevolution-foundation/evo-nexus | 545 | — | ~5.1k | Automated safety check: Notes | Custom licence | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| GPT Image Generation CLIwuyoscar/GPT-Image2-Skill | 5.7k | — | ~2.5k | Automated safety check: Notes | MIT | |
| BlockRun Image GenerationBlockRunAI/ClawRouter | 6.6k | — | ~2.1k | Automated safety check: Pass | MIT | |
| ImagegenJetBrains/skills | 364 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 |
evolution-foundation/evo-nexus
Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
wuyoscar/GPT-Image2-Skill
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
BlockRunAI/ClawRouter
Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.
JetBrains/skills
A skill your agent uses when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent…
NomaDamas/bananatape
Drives the BananaTape CLI to create, launch, list and delete local AI image editing projects from an agent, including headless smoke tests.
Works with
Categories
Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters. Routes every raster image task, including edits, background replacement, style transfer, compositing and batch work, through an OpenAI-compatible provider configured in Codex's own root files, stepping aside only when a user explicitly asks not to use it.md file rather than the current workspace, so the generation script keeps working regardless of where it is invoked from.
Codex API Image Generator fits situations like: generating or editing a raster image through a configured image provider; creating an image of a real, named subject that needs accurate visual reference; doing a localized edit, background swap or style transfer on an existing image; running a batch of image generation requests through the configured provider.
Run `npx skills add yc-duan/api-image --skill api-image -a claude-code`. Or copy the skill folder (the yc-duan/api-image repository) into .claude/skills/api-image in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yc-duan/api-image --skill api-image -a codex`. Or copy the skill folder (the yc-duan/api-image repository) into .agents/skills/api-image in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yc-duan/api-image --skill api-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/api-image, .gemini/skills/api-image, .github/skills/api-image and .opencode/skills/api-image in your project.
Going by SKILL.md and its folder, Codex API Image Generator needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENAI_API_KEY and API_IMAGE_API_KEY. Our summary lists: An OpenAI-compatible image provider configured in Codex's root files; Python, to run the bundled generation script.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Codex API Image Generator is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Codex API Image Generator: AI Image Creator (evolution-foundation/evo-nexus, 545 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars) and BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yc-duan (a GitHub user) maintains it in yc-duan/api-image, which has 101 GitHub stars. The repository was last updated on May 4, 2026.
Source: yc-duan/api-image on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.