AI Image Generation and Editing
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
Generate and edit images using OpenAI's GPT Image 2 API. An agent skill from glebis/claude-skills.
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install glebis/claude-skills gpt-image-2 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/gpt-image-2 .claude/skills/gpt-image-2 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .claude/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/glebis/claude-skills/tree/main/gpt-image-2Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install glebis/claude-skills gpt-image-2 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/gpt-image-2 .agents/skills/gpt-image-2 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .agents/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install glebis/claude-skills gpt-image-2 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/gpt-image-2 .cursor/skills/gpt-image-2 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .cursor/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/glebis/claude-skills.git --path gpt-image-2--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install glebis/claude-skills gpt-image-2 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/gpt-image-2 .gemini/skills/gpt-image-2 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .gemini/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install glebis/claude-skills gpt-image-2Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/gpt-image-2 .github/skills/gpt-image-2 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .github/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill gpt-image-2 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install glebis/claude-skills gpt-image-2 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/gpt-image-2 .opencode/skills/gpt-image-2 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpt-image-2" agent skill from https://github.com/glebis/claude-skills/tree/main/gpt-image-2 into .opencode/skills/gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpt-image-2", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpt-image-2Generate and edit images using OpenAI's GPT Image 2 API. An agent skill from glebis/claude-skills.
Gpt Image 2 is an agent skill from glebis/claude-skills. Generate and edit images using OpenAI's GPT Image 2 API. Interactive skill that guides users through image creation with style presets, cost-aware draft/final workflow, thinking mode, carousels, and photo editing. This skill should be used when the user requests image generation via OpenAI/GPT Image 2, wants to create social media carousels, edit photos into artistic styles, or needs images with readable text (infographics, diagrams, posters).
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `.claude-plugin/plugin.json`, `platforms.yaml` and `presets.yaml`).
It sits in Media & Creative, covering Image generation and Image editing. It works with OpenAI. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3b88261. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gpt Image 2 loads about 2.5k tokens when it runs, and up to ~130k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,096 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from glebis/claude-skills at commit 3b88261, republished under its MIT licence (© glebis). 1,096 words, ~2,463 tokens.
.claude/skills/gpt-image-2/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Generate and edit images via OpenAI's GPT Image 2 API with an interactive, guided workflow.
When the user invokes this skill, guide them through these steps using AskUserQuestion. Do not skip steps — the interactive flow is the core experience.
Ask the user what they want to create. Offer these options:
If the user already provided a clear prompt (e.g. "generate an editorial image of a rocket"), skip to Step 3.
Show the user available presets grouped by category. Read presets.yaml and present them:
Visual styles (no text in image): editorial, blueprint, ink, risograph, wireframe, constellation, brutalist, grain
Text-heavy (leverages GPT Image 2 text rendering): infographic, slide, diagram, poster, menu, manga
Community favorites: trading-card, pixar, app-mockup, isometric, action-figure, cinematic, panorama
Reference-anchored:
vhs — 1980s late-night infomercial title card: scanline-striped gradient italic caps on pure black. It auto-attaches a bundled reference image (references/vhs-infomercial.png), so the look stays consistent batch-to-batch. Pass the ad copy as the subject; for multi-line copy separate lines with / (e.g. --preset vhs "THEY TRUSTED YOU / NOW / PROVE IT").
Custom — user describes their own style
Ask: "Which style? Or describe your own."
Ask where this will be used:
Aspect-ratio caveat: --platform does NOT change the generation size — it generates at the configured size (default 1024×1024) and resizes/stretches afterwards, which distorts non-square targets (e.g. --platform story stretches a square to 1080×1920, cropping the composition's edges). For portrait or landscape compositions, pass the API-native size directly: --size 1024x1536 (portrait) or --size 1536x1024 (landscape).
Preflight false positives: the background-conflict heuristic trips on color words applied to non-background elements (e.g. "off-white text" in a dark-background prompt reads as a second background). If the flagged conflict is spurious, re-run with --force, or rephrase ("pale gray text").
Before any generation spend, the script now composes the final prompt first
(preset + subject + style), then checks it for internal contradictions — most often
a preset that hard-codes something the subject overrides (e.g. the editorial preset
forces "on pure black background" while your subject asks for a warm off-white ground).
The check prefers a fast Haiku call via the llm CLI; if Haiku is unavailable (no
llm, no Anthropic credit) it falls back to the configured llm default model, then to a
built-in static heuristic. The resolved prompt and the verdict are printed. If a conflict
is found, generation is aborted before spending — fix the prompt or preset and re-run, or
override with --force (generate anyway) or --no-preflight (skip the check). This is what
prevents the "generated on the wrong background, now regenerate" waste.
When composing prompts that set a background/palette, don't combine a background-fixing
preset (editorial, blueprint, etc.) with a different requested background — either drop
the preset and specify the full style yourself, or accept the preset's background.
Always generate a draft first unless the user says "skip draft" or uses --draft false.
--draft (quality=low, ~$0.006/image)--quality high (~$0.21/image)--seed from the draft to maintain composition when upgrading to finalThis draft→final flow saves ~97% on iteration costs.
After generation, always:
open <path> for full-resolution previewWhen the user wants a carousel (5-10 slides):
Ask: "What's the story? Give me the key message and I'll draft a 10-slide arc."
Then propose a slide-by-slide plan like:
Slide 1: [Cover] — hook headline + hero image
Slide 2: [Problem] — bold statement
Slide 3: [Context] — illustration + explanation
...
Slide 10: [CTA] — call to action with URLAsk the user to approve or modify the plan.
Use the same preset + seed range across all slides. For carousels:
--seed to lock composition patternsGenerate all slides as drafts first ($0.006 × 10 = $0.06 total). Show them all to the user as a contact sheet or one by one. Ask which ones to regenerate or adjust.
Only generate finals for approved slides. Offer to generate all at once with -y flag.
When the user wants to transform a photo:
osascript to a temp fileUse --edit <path> for the API call.
Always communicate costs before generating:
| Quality | Per image | 10-slide carousel |
|---|---|---|
--draft (low) | $0.006 | $0.06 |
| medium | $0.05 | $0.50 |
| high (default) | $0.21 | $2.10 |
| high + thinking | $0.25-0.42 | $2.50-4.20 |
Thinking mode adds 20-100% cost. Only suggest it for text-heavy or complex compositions.
The script auto-confirms when cost < $0.50. Above that, it prompts the user.
When helping users write prompts, apply these patterns:
'with the headline "Hello World"'editorial-magazine, studio-product to converge batches--seed for iteration: lock composition, vary only the prompt details# Basic generation
scripts/gpt_image_2.py "prompt" output.png
# With preset and platform
scripts/gpt_image_2.py --preset editorial --platform square "subject" out.png
# Draft mode (~$0.006/image)
scripts/gpt_image_2.py --draft "prompt" out.png
# With thinking for complex layouts
scripts/gpt_image_2.py --thinking medium --preset diagram "OAuth flow" out.png
# Seed for reproducibility
scripts/gpt_image_2.py --seed 42 "prompt" out.png
# Edit existing photo
scripts/gpt_image_2.py --edit photo.png "transform into constellation style" out.png
# Reference-anchored preset (auto-attaches its bundled reference image)
scripts/gpt_image_2.py --preset vhs --platform youtube "THEY TRUSTED YOU / NOW / PROVE IT" ad.png
# Variants with contact sheet
scripts/gpt_image_2.py --n 4 --preset ink "mountain" out.png
# Cost estimate
scripts/gpt_image_2.py --estimate --n 10 --quality high "batch test"
# Skip confirmation
scripts/gpt_image_2.py -y --n 10 "batch" out.png
# Dry run (show prompt without API call)
scripts/gpt_image_2.py --dry-run --preset editorial "test" out.png
# Preflight runs automatically before spend; override if needed
scripts/gpt_image_2.py --force "prompt with a known conflict" out.png # generate anyway
scripts/gpt_image_2.py --no-preflight "prompt" out.png # skip the checkscripts/gpt_image_2.py — main CLI (Python, requires PyYAML)presets.yaml — style presets (visual + text-heavy + community + reference-anchored). A preset may declare a reference: path (relative to the skill dir); it auto-attaches as a style anchor unless the user passes their own --reference. See the vhs preset.platforms.yaml — 8 platform sizing presetsreferences/api_reference.md — full API documentationreferences/vhs-infomercial.png — bundled style anchor for the vhs preset~/.config/gpt-image-2/config.yaml — user defaults~/.config/gpt-image-2/history.jsonl — generation log~/.config/gpt-image-2/last.json — last run (for again)© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in gpt-image-2 of glebis/claude-skills.
Open the folder on GitHubat commit 3b88261
Gpt Image 2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gpt Image 2 this skillglebis/claude-skills | 391 | — | ~2.5k | Automated safety check: Pass | MIT | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| GPT Image Generation CLIwuyoscar/GPT-Image2-Skill | 5.7k | — | ~2.5k | Automated safety check: Notes | MIT | |
| BlockRun Image GenerationBlockRunAI/ClawRouter | 6.6k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Codex API Image Generatoryc-duan/api-image | 101 | — | ~4.9k | Automated safety check: Pass | MIT | |
| ImagegenJetBrains/skills | 366 | 3 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 |
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
wuyoscar/GPT-Image2-Skill
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
BlockRunAI/ClawRouter
Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.
yc-duan/api-image
Replaces Codex's built-in image tool with provider-based generation and editing, fetching reference images first whenever visual accuracy actually matters.
JetBrains/skills
A skill your agent uses when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent…
NomaDamas/bananatape
Drives the BananaTape CLI to create, launch, list and delete local AI image editing projects from an agent, including headless smoke tests.
glebis/claude-skills
Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.
glebis/claude-skills
Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.
glebis/claude-skills
This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.
glebis/claude-skills
This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…
glebis/claude-skills
Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.
glebis/claude-skills
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
Works with
Categories
Generate and edit images using OpenAI's GPT Image 2 API. An agent skill from glebis/claude-skills. Gpt Image 2 is an agent skill from glebis/claude-skills. Generate and edit images using OpenAI's GPT Image 2 API.
Gpt Image 2 fits situations like: requests image generation via OpenAI/GPT Image 2; wants to create social media carousels; edit photos into artistic styles; needs images with readable text (infographics.
Run `npx skills add glebis/claude-skills --skill gpt-image-2 -a claude-code`. Or copy the skill folder (gpt-image-2 in glebis/claude-skills) into .claude/skills/gpt-image-2 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add glebis/claude-skills --skill gpt-image-2 -a codex`. Or copy the skill folder (gpt-image-2 in glebis/claude-skills) into .agents/skills/gpt-image-2 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill gpt-image-2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpt-image-2, .gemini/skills/gpt-image-2, .github/skills/gpt-image-2 and .opencode/skills/gpt-image-2 in your project.
Going by SKILL.md and its folder, Gpt Image 2 needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gpt Image 2 is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 127k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gpt Image 2: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars) and Codex API Image Generator (yc-duan/api-image, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
glebis (a GitHub user) maintains it in glebis/claude-skills, which has 391 GitHub stars. The repository holds 92 skills in this directory. The repository was last updated on October 8, 2026.
Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.