Agent skill

Codex API Image Generator

by yc-duan in yc-duan/api-image

Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.

MITAuto-check passedMedia & Creative

Install Codex API Image Generator

skills CLI
$ npx skills add yc-duan/api-image --skill api-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yc-duan/api-image api-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
api-image
GitHub stars
101
Token cost
~4.9k tokens
SKILL.md length
2,228 words
Files
12 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.

  • Works in 10 steps: Resolve the Codex root. → Resolve this skill's directory. → Decide whether research or references… → …
  • Generating or editing a raster image through a configured image provider
  • SKILL.md covers Workflow, Root Files, Command and Payload Shape, plus 11 more sections
  • Runs Python scripts from its folder; calls python; needs OPENAI_API_KEY and API_IMAGE_API_KEY

What it does

Routes every raster image task, including edits, background replacement, style transfer, compositing and batch work, through an OpenAI-compatible provider configured in Codex's own root files, stepping aside only when a user explicitly asks not to use it. It locates its own script directory relative to the SKILL.md file rather than the current workspace, so the generation script keeps working regardless of where it is invoked from.

Before generating anything, it decides whether the subject needs research: named places, real products, real UI, vehicles, uniforms, organisms and niche aesthetics are treated as research-required by default, pushing the skill to search the web or use supplied images for accurate silhouette, materials and proportions rather than relying on text description alone. Pure text-only generation is kept as a fallback for those subjects, used only when no useful references exist, network access fails, or the user opts out; for generic fantasy, mood pieces or ordinary objects, research stays optional.

Provider settings are read from the user's own Codex configuration files, including the API key from `auth.json`, so no credentials are typed into the conversation.

When your agent uses it

  • Generating or editing a raster image through a configured image provider
  • Creating an image of a real, named subject that needs accurate visual reference
  • Doing a localized edit, background swap or style transfer on an existing image
  • Running a batch of image generation requests through the configured provider

Example prompts

  • “Generate a product shot of this specific sneaker model using the configured provider.”
  • “Edit this screenshot to swap the background, keeping the subject untouched.”
  • “Create a few variations of this logo-free mascot illustration for a pitch deck.”

Requirements

  • An OpenAI-compatible image provider configured in Codex's root files
  • Python, to run the bundled generation script

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Resolve the Codex root.
  2. Resolve this skill's directory.
  3. Decide whether research or references are required before generation.
  4. Read the active provider settings from the user's root files.
  5. Decide the intent and input image roles.
  6. Choose the endpoint.
  7. Build a structured prompt.
  8. For official GPT Image models, decode data[].b64_json. Treat data[].url only as a compatibility fallback for non-official…
  9. Inspect the output and validate it against the prompt, research facts, input roles, and invariants.
  10. Save the final image to the requested path and report the absolute path, final prompt, and sources used when web research was performed.

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53535. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • API_IMAGE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codex API Image Generator loads about 4.9k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 2,228 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yc-duan/api-image at commit 5d53535, republished under its MIT licence (© yc-duan). 2,228 words, ~4,932 tokens.

Download SKILL.mdSave it as .claude/skills/api-image/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
api-image
description
Mandatory replacement for the built-in `imagegen` flow. Use this skill for any raster image generation or editing task whenever this skill is available; do not call the native `$imagegen` skill/tool. Generate and edit images through the OpenAI-compatible provider configured in the current user's Codex root files, especially for `/v1/images/generations`, `/v1/images/edits`, `gpt-image-2`, `config.toml`, `auth.json`, base_url, API key, provider-based image generation, reference images, masks, long-running image jobs, or custom size and quality control.

API Image

Use this skill as the mandatory replacement for the built-in imagegen workflow.

Routing rule:

  • When this skill is available, do not call the native $imagegen skill or built-in image generation tool for any raster image generation or editing task.
  • Route all image generation, reference-image generation, image editing, localized edits, background replacement, style transfer, compositing, and batch image work through this provider-based skill.
  • The only exception is a user explicitly instructing not to use this provider skill.

Workflow

  1. Resolve the Codex root.
    • Prefer an explicit --codex-home argument when the user provides one.
    • Otherwise use $CODEX_HOME.
    • Otherwise use ~/.codex.
  2. Resolve this skill's directory.
    • Treat the directory containing this SKILL.md as <skill-dir>.
    • Run the bundled script as <skill-dir>/scripts/generate_image.py, not as a path relative to the current workspace.
  3. Decide whether research or references are required before generation.
    • Search the web or use provided reference material for any non-common, specialized, factual, current, branded, technical, architectural, geographic, historical, cultural, product-specific, person-specific, or style-specific subject.
    • Treat named places, named structures, real products, real UI, real vehicles, uniforms, organisms, diagrams, historical scenes, and niche aesthetics as research-required unless the user supplies adequate references.
    • Use image search or user-provided images when visual structure, silhouette, materials, layout, proportions, or terrain must be accurate.
    • Be proactive about reference images. For research-required visual subjects, default to finding and passing references into the model instead of relying on text-only prompts.
    • Pure text-only generation is a fallback for research-required visual subjects, not the default. Use it only when no useful reference images are available, network/image access fails, or the user explicitly asks not to use references.
    • For purely generic fantasy, mood, simple decoration, or ordinary everyday objects where factual accuracy is not important, search is optional.
    • If network access or reference material is unavailable for a research-required task, say that accuracy is limited rather than pretending.
  4. Read the active provider settings from the user's root files.
    • Read auth.json and use OPENAI_API_KEY.
    • Read config.toml, then use model_provider and [model_providers.<name>].base_url.
    • Never hardcode a provider URL or API key into the skill.
    • Default to the root files above. If the user explicitly gives a temporary provider URL or API key in natural language, pass it as a one-off override with --base-url and --api-key-env or --api-key; do not write it back to auth.json, config.toml, README, logs, or generated files.
    • Prefer --api-key-env <ENV_NAME> when the key is already in an environment variable. Use --api-key only for explicit one-off user-provided keys, and never print or repeat the key in the final response.
  5. Decide the intent and input image roles.
    • If the user wants a new image from text only, treat it as generation.
    • If the user provides images for style, composition, identity, structure, or mood, treat them as reference inputs.
    • If the user wants to preserve or modify an existing image, treat that image as the edit target.
    • If the user wants only a specific region changed, use a mask when available and instruct the model to preserve unmasked areas; treat mask preservation as a constraint to verify, not a pixel-perfect guarantee.
    • Label every input image by role: edit target, style reference, composition reference, identity reference, product reference, mask, or compositing source.
  6. Choose the endpoint.
    • Use <base_url>/images/generations for text-only generation.
    • Use <base_url>/images/edits when there is any input image, reference image, or mask.
    • Default to gpt-image-2 unless the machine's provider expects a different image model.
  7. Build a structured prompt.
    • Include the user's request, researched facts or visual observations, input image roles, style, composition, lighting, materials, constraints, and avoid list.
    • Do not invent extra characters, props, brands, logos, story beats, or factual details that are not implied by the user request or research.
  8. For official GPT Image models, decode data[].b64_json. Treat data[].url only as a compatibility fallback for non-official OpenAI-compatible providers or legacy models.
  9. Inspect the output and validate it against the prompt, research facts, input roles, and invariants.
  10. Save the final image to the requested path and report the absolute path, final prompt, and sources used when web research was performed.

Root Files

Default root files:

  • Windows: %USERPROFILE%\\.codex\\auth.json and %USERPROFILE%\\.codex\\config.toml
  • General rule: $CODEX_HOME/auth.json and $CODEX_HOME/config.toml, otherwise ~/.codex/auth.json and ~/.codex/config.toml

This skill must always read the current files at runtime. The provider may differ from machine to machine.

Temporary overrides:

  • Default behavior reads the Codex root files above.
  • Use --base-url <https://provider.example/v1> to temporarily override the configured Provider URL.
  • Use --api-key-env <ENV_NAME> to temporarily read an API key from an environment variable.
  • Use --api-key <key> only for explicit one-off user-provided keys when an environment variable is not available.
  • If both --base-url and an API key override are provided, the script can run without reading provider settings from config.toml; otherwise missing values fall back to the Codex root files.
  • Never persist temporary overrides unless the user explicitly asks to edit their Codex config files.

Command

Use the bundled script for normal text-to-image generation:

powershell
python "<skill-dir>\scripts\generate_image.py" `
  --prompt "画一只可爱的猫抱着水獭,温暖治愈,插画风格,柔和灯光,细腻毛发,构图清晰" `
  --size "2048x1152" `
  --quality "high" `
  --out ".\outputs\cute-cat-otter.png"

Use the same script for reference-image generation or edits:

powershell
python "<skill-dir>\scripts\generate_image.py" `
  --prompt "参考输入图的构图和角色姿势,生成一张暖色电影感插画" `
  --image "C:\path\to\reference.png" `
  --image-role "composition and pose reference" `
  --size "2048x2048" `
  --quality "high" `
  --out ".\outputs\reference-output.png"

Use --mask for localized edits:

powershell
python "<skill-dir>\scripts\generate_image.py" `
  --prompt "只把被 mask 标出的区域替换成一只小水獭,保持其他区域不变" `
  --image "C:\path\to\source.png" `
  --mask "C:\path\to\mask.png" `
  --out ".\outputs\masked-edit.png"

For web-researched subjects, download selected reference images to a working folder first, then pass them as --image:

powershell
python "<skill-dir>\scripts\generate_image.py" `
  --prompt "Create a new 4K overhead aerial image of Huajiang Grand Canyon Bridge in Guizhou, based on the reference images. Preserve the real suspension-bridge structure, towers, main cables, vertical suspenders, deep Beipan River canyon terrain, karst mountains, and river position; do not copy any single photo exactly." `
  --image "C:\path\to\refs\huajiang-bridge-aerial.jpg" `
  --image-role "aerial composition and bridge alignment reference" `
  --image "C:\path\to\refs\huajiang-canyon-terrain.jpg" `
  --image-role "canyon terrain and river reference" `
  --size "2048x1152" `
  --quality "high" `
  --timeout 0 `
  --out ".\outputs\huajiang-canyon-bridge-overhead.png"

Useful options:

  • --model gpt-image-2
  • --mode auto|generate|edit
  • --size 2048x1152 script default 2K landscape
  • --size 1024x1024
  • --size 1536x1024
  • --size 1024x1536
  • --size 2048x2048
  • --size 3840x2160
  • --size auto for official API auto sizing
  • --quality low|medium|high|auto
  • --n 1
  • --image <path> repeated up to 16 times
  • --image-role <role> to label input images inside the prompt
  • --mask <path> for localized edits
  • --background auto|opaque for official gpt-image-2
  • --output-format png|jpeg|webp
  • --output-compression 0-100 for jpeg or webp
  • --input-fidelity low|high for supported edit models only; do not send it for gpt-image-2
  • --moderation auto|low for supported GPT image models
  • --prompt-file <path> for one prompt per non-empty line
  • --codex-home <path>
  • --base-url <https://provider.example/v1> temporary Provider URL override
  • --api-key-env <ENV_NAME> temporary API key override from an environment variable
  • --api-key <key> temporary direct API key override, only when explicitly provided
  • --timeout 1800
  • --timeout 0

Payload Shape

Use the provider's generation endpoint with this body shape:

json
{
  "model": "gpt-image-2",
  "prompt": "你的中文或英文提示词",
  "size": "2048x1152",
  "quality": "high",
  "n": 1
}

Use the provider's edit endpoint as multipart form data:

text
model=gpt-image-2
prompt=<edit or reference prompt>
image[]=@source-or-reference.png
mask=@mask.png
size=1024x1024
quality=high
n=1

Start with n=1. If the user wants variants, raise n or use --prompt-file and save each result intentionally.

Temporary Provider Overrides

If the user says in natural language that this image job should use a different Provider URL or API key, treat it as a temporary runtime override:

powershell
$env:API_IMAGE_API_KEY = "<temporary key>"
python "<skill-dir>\scripts\generate_image.py" `
  --prompt "一张赛博朋克风格的夜景照片" `
  --base-url "https://provider.example/v1" `
  --api-key-env "API_IMAGE_API_KEY" `
  --out ".\outputs\override-provider.png"

Rules:

  • Default to the Codex root configuration when no override is requested.
  • Do not modify auth.json or config.toml for temporary overrides.
  • Do not echo API keys back to the user.
  • Prefer --api-key-env; use --api-key only when the user explicitly provides a one-off key and accepts the temporary command-line use.

gpt-image-2 API Notes

For official OpenAI gpt-image-2:

  • Text-only generation uses POST /v1/images/generations with JSON.
  • Any input image, reference image, or mask uses POST /v1/images/edits with multipart form-data.
  • Use repeated image[]=@path fields for multiple input/reference images.
  • When using mask, the first --image must be the edit target; additional images should be references or compositing sources.
  • Read image bytes from data[].b64_json; URL output is not supported for official GPT Image models.
  • Do not send input_fidelity; gpt-image-2 always processes image inputs at high fidelity.
  • Do not request background: transparent; official gpt-image-2 currently supports auto or opaque, not transparent backgrounds.
  • For masked edits, validate that the mask and first source image have the same dimensions and compatible format, are under the image API file-size limit, and that the mask includes an alpha channel.

Capability Map

  • New text-to-image generation: supported through /images/generations.
  • Batch generation: supported through --n for multiple variants per prompt, or --prompt-file for many prompts.
  • Reference-image generation: supported through /images/edits with one or more --image inputs and role labels in the prompt.
  • Image editing: supported through /images/edits with --image and an edit prompt.
  • Localized edits / inpainting: supported through /images/edits with --image plus --mask.
  • Background replacement: supported as an edit prompt, with --mask when the replacement area must be constrained.
  • Style transfer: supported as reference-image editing; label the source image role and describe the desired style in the prompt.
  • Input image role labeling: not a separate API field; this script appends --image-role labels to the prompt so the model can interpret each input image intentionally.
  • Transparent background: official gpt-image-2 does not currently support background: transparent. Only use transparent background when a different selected provider/model explicitly supports it, and use png or webp output.
Show full SKILL.md (862 more words)Show less

Research And References

Do not assume the image model knows every subject accurately. Research first or use references when the subject is not common visual knowledge.

Reference images are the preferred path for visually constrained research-required tasks. Text research describes facts; reference images carry visual structure. Use both whenever feasible.

Research-required examples:

  • A named structure or location, such as 贵州花江峡谷大桥 or any real bridge, landmark, terrain, skyline, or building.
  • A real product, vehicle, machine, UI, logo-free product silhouette, game asset from a known franchise, or fashion item.
  • A historical scene, cultural garment, uniform, heraldry, ritual object, weapon, architecture style, organism, map, technical diagram, or scientific subject.
  • A current or recently changed subject.
  • A niche visual style or artist-adjacent style where references are necessary to avoid generic output.

Research workflow:

  • Search the web for factual text details when structure, function, history, geography, or current status matters.
  • Use image search or provided images for visual details such as shape, terrain, materials, proportions, colors, and surrounding environment.
  • For named structures, real locations, real products, real vehicles, historical/cultural visual subjects, technical diagrams, and niche styles, actively collect at least one useful reference image before generating.
  • Prefer multiple complementary references when possible: one for subject structure, one for surrounding environment, and one for desired camera angle/composition.
  • Download reference images to a local working path before invoking the script, then pass them with --image and --image-role.
  • If a reference image is low quality but still useful, label its role narrowly, for example rough terrain reference only.
  • Extract only the details needed for the prompt; keep the prompt concise.
  • Include factual constraints in the prompt, for example bridge type, deck/tower/cable arrangement, canyon terrain, river position, viewpoint, and surrounding landforms.
  • Cite sources in the final response whenever web research was used.

Reference workflow:

  • Prefer provided reference images over memory when visual fidelity matters.
  • For each --image, pass a matching --image-role such as edit target, style reference, composition reference, product reference, or terrain reference.
  • Do not treat every image as an edit target; decide whether it is reference-only or should be preserved and modified.
  • For compositing, state exactly what comes from each input and how lighting, perspective, scale, and shadows should match.
  • For reference-only generation, prompt the model to create a new image based on the references, not to copy a single source photo exactly.
  • If the image model/provider rejects a reference-image edit request, surface that error and then decide whether a text-only generation is acceptable for this task. Do not silently fall back.

Prompt Structure

Use this compact schema when it helps:

text
Use case: <photorealistic-natural|product-mockup|ui-mockup|infographic-diagram|logo-brand|illustration-story|stylized-concept|historical-scene|text-localization|identity-preserve|precise-object-edit|lighting-weather|background-extraction|style-transfer|compositing|sketch-to-render>
Asset type: <where the image will be used>
Primary request: <user request>
Research facts / visual references: <only the relevant observed details>
Input images: <Image 1: role; Image 2: role>
Scene/backdrop: <environment>
Subject: <main subject>
Style/medium: <photo/illustration/3D/etc>
Composition/framing: <viewpoint, lens/framing, placement>
Lighting/mood: <lighting and mood>
Materials/textures: <surface details>
Text (verbatim): "<exact text if needed>"
Constraints: <must keep/must avoid>
Avoid: <negative constraints>

Waiting Behavior

  • This script is synchronous. It does not return before the provider finishes the image job or an explicit error occurs.
  • Default timeout is 1800 seconds so long-running jobs are not cut off too early.
  • Use --timeout 0 to disable the client-side timeout and wait indefinitely for the provider response.
  • When invoking this script from Codex, give the shell command a timeout that is longer than the expected image job. Do not use a short shell timeout for long generations.

Size And Quality

Follow the current official GPT Image rules for gpt-image-2:

  • size can be auto or any <width>x<height> string that satisfies all of these:
  • width and height are multiples of 16
  • the longest edge is at most 3840
  • the long-edge to short-edge ratio is at most 3:1
  • total pixels are between 655,360 and 8,294,400
  • this script defaults to 2048x1152; the official API's default sizing behavior is auto
  • common examples: 1024x1024, 1536x1024, 1024x1536, 2048x2048, 2048x1152, 3840x2160, 2160x3840
  • quality can be low, medium, high, or auto; this script defaults to high, while official API auto quality is available with auto
  • 2K and 4K requests are valid when they satisfy those constraints
  • Do not silently coerce unsupported sizes. Raise a clear error instead.

Prompting

  • Prefer concrete prompts with subject, style, lighting, and composition.
  • Use the user's language by default. Chinese prompts are valid for this workflow.
  • Change one aspect at a time when iterating.

Output Rules

  • Save project assets inside the project workspace when the user wants a usable asset.
  • Save preview images to a clear temporary or user-requested path.
  • If n > 1, keep filenames deterministic by appending -1, -2, and so on.

Failure Rules

  • Raise explicit errors when OPENAI_API_KEY, model_provider, or base_url cannot be found.
  • Allow temporary Provider overrides with --base-url, --api-key-env, or --api-key; do not persist them unless explicitly requested.
  • Raise explicit errors when both --api-key and --api-key-env are provided, when the named environment variable is empty, or when --base-url is not an HTTP(S) URL.
  • Raise explicit errors when size violates the official GPT Image constraints or quality is unsupported.
  • Raise explicit errors when edit/reference mode is requested without input images.
  • Raise explicit errors when gpt-image-2 is used with --background transparent or --input-fidelity.
  • Raise explicit errors when a mask is not compatible with the first --image edit target, has different dimensions, lacks an alpha channel, or exceeds the image API file-size limit.
  • Surface provider errors directly. Do not fabricate a success result.
  • If the provider uses a different image model name on this machine, override --model.

Resource

  • scripts/generate_image.py: Read the current user's Codex root files, call the configured provider's image endpoint, and save the returned image data.

© yc-duan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts) in the repository root of yc-duan/api-image.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/openai.yaml
  • scripts/generate_image.py
  • scripts/provider_imagegen/__init__.py
  • scripts/provider_imagegen/config.py
  • scripts/provider_imagegen/http_client.py
  • scripts/provider_imagegen/outputs.py
  • scripts/provider_imagegen/payloads.py
  • scripts/provider_imagegen/validation.py

Open the folder on GitHubat commit 5d53535

Compare with similar skills

Codex API Image Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codex API Image Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codex API Image Generator this skillyc-duan/api-image101—~4.9kAutomated safety check: PassMIT
AI Image Creatorevolution-foundation/evo-nexus545—~5.1kAutomated safety check: NotesCustom licence
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT
ImagegenJetBrains/skills3644 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • AI Image Creator

    evolution-foundation/evo-nexus

    Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.

    545 GitHub stars~5.1k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: notes
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Imagegen

    JetBrains/skills

    Official

    A skill your agent uses when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent…

    364 GitHub starsUsed in 4 repos~2.5k tokens
    Media & CreativeAuto-check passed
  • BananaTape Image Editor CLI

    NomaDamas/bananatape

    Drives the BananaTape CLI to create, launch, list and delete local AI image editing projects from an agent, including headless smoke tests.

    188 GitHub stars~547 tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

Works with

Questions about Codex API Image Generator

What does Codex API Image Generator do?

Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters. Routes every raster image task, including edits, background replacement, style transfer, compositing and batch work, through an OpenAI-compatible provider configured in Codex's own root files, stepping aside only when a user explicitly asks not to use it.md file rather than the current workspace, so the generation script keeps working regardless of where it is invoked from.

When should I use Codex API Image Generator?

Codex API Image Generator fits situations like: generating or editing a raster image through a configured image provider; creating an image of a real, named subject that needs accurate visual reference; doing a localized edit, background swap or style transfer on an existing image; running a batch of image generation requests through the configured provider.

How do I install Codex API Image Generator in Claude Code?

Run `npx skills add yc-duan/api-image --skill api-image -a claude-code`. Or copy the skill folder (the yc-duan/api-image repository) into .claude/skills/api-image in your project. Claude Code loads it when a task matches its description.

How do I install Codex API Image Generator in Codex?

Run `npx skills add yc-duan/api-image --skill api-image -a codex`. Or copy the skill folder (the yc-duan/api-image repository) into .agents/skills/api-image in your project. Codex loads it when a task matches its description.

Can I use Codex API Image Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yc-duan/api-image --skill api-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/api-image, .gemini/skills/api-image, .github/skills/api-image and .opencode/skills/api-image in your project.

What does Codex API Image Generator need to run?

Going by SKILL.md and its folder, Codex API Image Generator needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENAI_API_KEY and API_IMAGE_API_KEY. Our summary lists: An OpenAI-compatible image provider configured in Codex's root files; Python, to run the bundled generation script.

Does Codex API Image Generator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Codex API Image Generator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Codex API Image Generator use?

Codex API Image Generator is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codex API Image Generator use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Codex API Image Generator?

Skills that share tags, products or a category with Codex API Image Generator: AI Image Creator (evolution-foundation/evo-nexus, 545 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars) and BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codex API Image Generator?

yc-duan (a GitHub user) maintains it in yc-duan/api-image, which has 101 GitHub stars. The repository was last updated on May 4, 2026.

Source: yc-duan/api-image on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.