AI Image Generation and Editing
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
Multi-vendor AI image generation with authentication-aware parallel dispatch.
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install first-fluke/oh-my-agent oma-image --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .claude/skills/oma-image && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .claude/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-imageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install first-fluke/oh-my-agent oma-image --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .agents/skills/oma-image && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .agents/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install first-fluke/oh-my-agent oma-image --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .cursor/skills/oma-image && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .cursor/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/first-fluke/oh-my-agent.git --path benchmarks/runs/oma/.agents/skills/oma-image--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install first-fluke/oh-my-agent oma-image --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .gemini/skills/oma-image && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .gemini/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install first-fluke/oh-my-agent oma-imageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .github/skills/oma-image && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .github/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add first-fluke/oh-my-agent --skill oma-image -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install first-fluke/oh-my-agent oma-image --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/first-fluke/oh-my-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/benchmarks/runs/oma/.agents/skills/oma-image .opencode/skills/oma-image && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "oma-image" agent skill from https://github.com/first-fluke/oh-my-agent/tree/main/benchmarks/runs/oma/.agents/skills/oma-image into .opencode/skills/oma-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oma-image", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
oma-imageMulti-vendor AI image generation with authentication-aware parallel dispatch.
Oma Image is an agent skill from first-fluke/oh-my-agent. Multi-vendor AI image generation with authentication-aware parallel dispatch. Routes to Codex (gpt-image-2 via ChatGPT OAuth) and Pollinations (flux/zimage, free with signup). Gemini provider is present but disabled by default (requires billing). Use for image generation, image creation, visual asset generation, and AI art.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `config/image-config.yaml`, `resources/checklist.md` and `resources/execution-protocol.md`).
It sits in Media & Creative, covering Image generation. It works with OpenAI. The repository describes itself as: Mechanical verification for AI coding agents — skills pack or full harness (stop-hook gates, artifact checks, independent judges). The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b364119. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
codexgeminiFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
enter.pollinations.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
POLLINATIONS_API_KEYGEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Oma Image loads about 3.7k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,608 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from first-fluke/oh-my-agent at commit b364119, republished under its MIT licence (© first-fluke). 1,608 words, ~3,699 tokens.
.claude/skills/oma-image/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Generate images and visual assets through authenticated multi-vendor routing while preserving prompt clarity, reference-image handling, cost controls, and reproducible output manifests.
.agents/results/images/ or requested output directorymanifest.json with prompt, vendor, model, and reproducibility metadata--vendor all is usedoma image generate CLI and vendor authenticationresources/vendor-matrix.md, resources/prompt-tips.md, and config/image-config.yamloma image generate with selected vendor(s), prompt, references, and options.--vendor all is requested, require every requested vendor to be available.oma update.| Action | SSL primitive | Evidence |
|---|---|---|
| Validate prompt completeness | VALIDATE | Clarification protocol |
| Select vendor strategy | SELECT | Vendor matrix and auth state |
| Read reference images | READ | --reference paths |
| Call generation CLI/API | CALL_TOOL | oma image generate |
| Write image outputs | WRITE | Image files and manifest |
| Validate result | VALIDATE | Exit code, manifest, files |
| Report output | NOTIFY | Final path summary |
oma image generate, oma image doctor, oma image list-vendorsoma image doctor
oma image generate "<prompt>" --vendor auto --size auto --quality auto --format jsonWith reference images:
oma image generate --reference "<absolute-path>" --vendor codex "<prompt>"| Scope | Resource target |
|---|---|
LOCAL_FS | Reference images, generated images, manifests |
PROCESS | Provider CLIs and image router commands |
NETWORK | Pollinations/Gemini or provider APIs |
CREDENTIALS | Provider auth and API keys |
Clarification Protocol below.--vendor all, every requested vendor must be available (strict).$0.20 (configurable). --yes / OMA_IMAGE_YES=1 bypass. Default vendor pollinations (flux/zimage) is free, so auto-triggering on keywords is safe.$PWD require --allow-external-out.manifest.json next to the images for reproducibility.n = 5 — wall-time bound.oma search fetch (0, 1, 2=safety, 3=not-found, 4=invalid-input, 5=auth-required, 6=timeout).Before invoking oma image generate, the calling agent runs this checklist against the user's request. If any answer is "no / unknown", clarify with the user first.
Required signal (must be present or inferable):
Strongly recommended (ask if absent AND not inferable from context):
1024x1024), portrait (1024x1536), landscape (1536x1024)?Amplification shortcut. For brief prompts (e.g. "a red apple"), do not pop clarifying questions if the request is genuinely that simple — instead amplify inline and show the user the expanded version before invoking:
User: "a red apple" Agent: "I'll generate this as: a single glossy red apple centered on a clean white background, soft studio lighting, photorealistic, shallow depth of field, 1024×1024. Shall I proceed, or would you like a different style/composition?"
Skip both clarification and amplification when the user has clearly authored a full creative brief (≥ 2 of: subject + style + lighting + composition). Respect their prompt verbatim.
Category-specific briefs (app mockup, poster, thumbnail, infographic, comic panel, avatar): consult resources/prompt-tips.md → External Prompt Libraries.
Output language. Generation prompts are sent to the provider in English (image models are trained predominantly on English captions). Translate the user's request if they wrote in another language, and show them the translated version during amplification so they can correct misreadings.
This skill follows oh-my-agent's CLI-first concept: whenever a vendor's native CLI can drive generation (and return raw bytes), the subprocess path is preferred over direct API keys. Direct API is only used as a fallback for vendors whose CLI can't yet emit raw image bytes.
| Vendor | Strategy | Models | Trigger |
|---|---|---|---|
codex | CLI-first — codex exec via ChatGPT OAuth (codex login), built-in image_gen | gpt-image-2 | Logged in via Codex CLI (no API key) |
pollinations | Direct HTTP — gen.pollinations.ai/v1/images/generations (free signup for key) | Free: flux, zimage. Credit-gated: qwen-image, wan-image, gpt-image-2, klein, kontext, gptimage, gptimage-large | POLLINATIONS_API_KEY set (free at https://enter.pollinations.ai). No native CLI exists. |
gemini | CLI-first fallback → direct API. gemini -p (stream) is the preferred path but currently disabled at precheck (CLI's agentic loop does not return raw inlineData bytes on stdout as of Gemini CLI 0.38). Until the CLI exposes a non-agentic image surface, the provider falls back to the direct generativelanguage.googleapis.com API. | gemini-2.5-flash-image, gemini-3.1-flash-image-preview | Preferred: gemini auth login. Fallback: GEMINI_API_KEY + billing. |
/oma-image a red apple on white background
/oma-image --vendor all --size 1536x1024 jeju coastline at sunset
/oma-image -n 3 --quality high --out ./hero "minimalist dashboard hero illustration"oma image generate "<prompt>" [--vendor auto|codex|pollinations|gemini|all] [-n 1..5] \
[--size 1024x1024|1024x1536|1536x1024|auto] \
[--quality low|medium|high|auto] \
[--out <dir>] [--allow-external-out] \
[-r <path>]... \
[--timeout 180] [-y] [--no-prompt-in-manifest] \
[--dry-run] [--format text|json]
oma image doctor
oma image list-vendorsGemini-only escalation flag: --strategy mcp,stream,api (overrides vendors.gemini.strategies).
-r, --reference)Attach up to 10 reference images (PNG/JPEG/GIF/WebP, ≤ 5MB each) to guide style, subject identity, or composition. Repeatable or comma-separated.
oma image generate -r ~/Downloads/otter.jpeg "same otter in dramatic lighting"
oma image generate -r a.png -r b.png "blend these two styles"Supported vendors:
| Vendor | Support | How |
|---|---|---|
codex (gpt-image-2) | ✅ | Passes -i <path> to codex exec |
gemini (2.5-flash-image) | ✅ | Inlines base64 inlineData parts in request |
pollinations | ❌ | Rejected with exit code 4 (requires URL hosting; see PR #2 roadmap) |
Paths: absolute or relative to $CWD. Host CLIs usually expose attached images via:
~/.claude/image-cache/<session>/N.png (surfaced in system messages as [Image: source: <path>])When ALL of the following are true, the calling agent MUST pass the attached image via --reference <path> automatically. Never describe the image in prose as a workaround.
[Image: source: <path>], or an Antigravity workspace upload path, or an explicit filesystem path in the user's message.codex or gemini).Required action: invoke oma image generate --reference <absolute-path> --vendor <codex|gemini> "<prompt>". If the user didn't specify a vendor, default to codex (CLI-first, widest availability). Do NOT:
oma image generate --help to verify.If the local CLI is outdated (--reference is missing from --help): tell the user to run oma update once, then retry. Do not silently degrade to prose.
If the reference path is from Claude Code's image-cache: note to the user that the path is session-scoped and suggest copying the file to a durable location if they want to reuse it later. Still proceed with the generation.
Other skills call oma image generate --format json and parse the JSON manifest from stdout.
.agents/results/images/
├── 20260424-143052-ab12cd/ # single-vendor run
│ └── pollinations-flux.jpg
│ (or codex-gpt-image-2.png)
│ manifest.json
└── 20260424-143122-7z9kqw-compare/ # --vendor all run
├── codex-gpt-image-2.png
├── pollinations-flux.jpg
└── manifest.jsonFollow resources/execution-protocol.md step by step.
See resources/vendor-matrix.md for strategy precheck rules.
Use resources/prompt-tips.md for writing effective prompts.
Before submitting, run resources/checklist.md.
Project-specific settings: config/image-config.yaml.
Env vars: OMA_IMAGE_DEFAULT_VENDOR, OMA_IMAGE_DEFAULT_OUT, OMA_IMAGE_YES, POLLINATIONS_API_KEY, GEMINI_API_KEY, OMA_IMAGE_GEMINI_STRATEGIES.
resources/execution-protocol.mdresources/vendor-matrix.mdresources/prompt-tips.mdresources/checklist.md../_shared/core/context-loading.md© first-fluke, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in benchmarks/runs/oma/.agents/skills/oma-image of first-fluke/oh-my-agent.
Open the folder on GitHubat commit b364119
Oma Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Oma Image this skillfirst-fluke/oh-my-agent | 1.3k | — | ~3.7k | Automated safety check: Pass | MIT | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| GPT Image Generation CLIwuyoscar/GPT-Image2-Skill | 5.7k | — | ~2.5k | Automated safety check: Notes | MIT | |
| Imagegentheowenyoung/home | 115 | 4 repos | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| Openai Image Gentrpc-group/trpc-agent-go | 1.9k | 12 repos | ~843 | Automated safety check: Pass | Apache-2.0 | |
| Image Generationonyx-dot-app/onyx | 32k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence |
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
wuyoscar/GPT-Image2-Skill
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
theowenyoung/home
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts.
trpc-group/trpc-agent-go
Batch-generate images via OpenAI Images API. An agent skill from trpc-group/trpc-agent-go.
onyx-dot-app/onyx
Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.
BlockRunAI/ClawRouter
Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.
first-fluke/oh-my-agent
Decomposes a complex feature into tasks, dispatches parallel specialist agents with durable state, and supervises verification, QA review and retries.
first-fluke/oh-my-agent
Splits a complex feature into prioritized tasks, spawns specialist CLI subagents in parallel, tracks them through shared memory and verifies each result.
first-fluke/oh-my-agent
Evaluates system boundaries and tradeoffs and writes architecture recommendations, option comparisons or ADRs, with a Mermaid diagram when structure changes.
first-fluke/oh-my-agent
Backend specialist for APIs, database work, authentication and migrations that follows clean architecture with router, service and repository layers.
first-fluke/oh-my-agent
Installs or checks the oma CLI and its runtimes (bun, uv, serena) in a fresh workspace so that oma-* skills can run their commands.
first-fluke/oh-my-agent
Explores intent, constraints and alternative approaches before any planning, working through questions one at a time and saving an approved design for later steps.
Works with
Categories
Multi-vendor AI image generation with authentication-aware parallel dispatch. Oma Image is an agent skill from first-fluke/oh-my-agent. Multi-vendor AI image generation with authentication-aware parallel dispatch.
Oma Image fits situations like: image generation; visual asset generation.
Run `npx skills add first-fluke/oh-my-agent --skill oma-image -a claude-code`. Or copy the skill folder (benchmarks/runs/oma/.agents/skills/oma-image in first-fluke/oh-my-agent) into .claude/skills/oma-image in your project. Claude Code loads it when a task matches its description.
Run `npx skills add first-fluke/oh-my-agent --skill oma-image -a codex`. Or copy the skill folder (benchmarks/runs/oma/.agents/skills/oma-image in first-fluke/oh-my-agent) into .agents/skills/oma-image in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add first-fluke/oh-my-agent --skill oma-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/oma-image, .gemini/skills/oma-image, .github/skills/oma-image and .opencode/skills/oma-image in your project.
Going by SKILL.md and its folder, Oma Image needs the command-line tools its instructions call (codex and gemini) and credentials named POLLINATIONS_API_KEY and GEMINI_API_KEY. Our summary lists: A credential in POLLINATIONS_API_KEY; A credential in GEMINI_API_KEY.
SKILL.md names 1 domain. As links in the text: enter.pollinations.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Oma Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Oma Image: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars), Imagegen (theowenyoung/home, 115 stars) and Openai Image Gen (trpc-group/trpc-agent-go, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
first-fluke (a GitHub organization) maintains it in first-fluke/oh-my-agent, which has 1,338 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on October 9, 2026.
Source: first-fluke/oh-my-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.