AI Image Generation and Editing
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install einverne/dotfiles gemini-image-gen --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-image-gen .claude/skills/gemini-image-gen && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .claude/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-genType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install einverne/dotfiles gemini-image-gen --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .agents/skills && cp -r skills-src/claude/skills/gemini-image-gen .agents/skills/gemini-image-gen && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .agents/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install einverne/dotfiles gemini-image-gen --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/claude/skills/gemini-image-gen .cursor/skills/gemini-image-gen && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .cursor/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/einverne/dotfiles.git --path claude/skills/gemini-image-gen--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install einverne/dotfiles gemini-image-gen --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/claude/skills/gemini-image-gen .gemini/skills/gemini-image-gen && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .gemini/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install einverne/dotfiles gemini-image-genInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .github/skills && cp -r skills-src/claude/skills/gemini-image-gen .github/skills/gemini-image-gen && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .github/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install einverne/dotfiles gemini-image-gen --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/claude/skills/gemini-image-gen .opencode/skills/gemini-image-gen && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gemini-image-gen" agent skill from https://github.com/einverne/dotfiles/tree/master/claude/skills/gemini-image-gen into .opencode/skills/gemini-image-gen/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-image-gen", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gemini-image-genGuide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.
Gemini Image Gen is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model. Use when generating images, creating visual content, or implementing text-to-image features. Supports text-to-image, image editing, multi-image composition, and iterative refinement.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `README.md`, `SKILL_CREATION_SUMMARY.md` and `references/api-reference.md`).
It sits in Media & Creative, covering Image generation and Image editing. It works with Google Gemini. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
aistudio.google.comai.google.devFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gemini Image Gen loads about 1.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 504 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
tory**: `.claude/skills/gemini-image-gen/.env`3. **Project directory**: `./.env` (project root)Create `.env` file with:# Verify .env file existscat .claude/skills/gemini-image-gen/.envcat .envallowed-tools: Bash, Read, WriteAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from einverne/dotfiles at commit c6c0686, republished under its MIT licence (© einverne). 504 words, ~1,801 tokens.
.claude/skills/gemini-image-gen/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Generate high-quality images using Google's Gemini 2.5 Flash Image model with text prompts, image editing, and multi-image composition capabilities.
Use this skill when you need to:
The skill automatically detects your GEMINI_API_KEY in this order:
export GEMINI_API_KEY="your-key".claude/skills/gemini-image-gen/.env./.env (project root)Get your API key: Visit Google AI Studio
Create .env file with:
GEMINI_API_KEY=your_api_key_hereInstall required package:
pip install google-genaifrom google import genai
from google.genai import types
import os
# API key detection handled automatically by helper script
client = genai.Client(api_key=os.getenv('GEMINI_API_KEY'))
response = client.models.generate_content(
model='gemini-2.5-flash-image',
contents='A serene mountain landscape at sunset with snow-capped peaks',
config=types.GenerateContentConfig(
response_modalities=['image'],
aspect_ratio='16:9'
)
)
# Save to ./docs/assets/
for i, part in enumerate(response.candidates[0].content.parts):
if part.inline_data:
with open(f'./docs/assets/generated-{i}.png', 'wb') as f:
f.write(part.inline_data.data)For convenience, use the provided helper script that handles API key detection and file saving:
# Generate single image
python .claude/skills/gemini-image-gen/scripts/generate.py \
"A futuristic city with flying cars" \
--aspect-ratio 16:9 \
--output ./docs/assets/city.png
# Generate with specific modalities
python .claude/skills/gemini-image-gen/scripts/generate.py \
"Modern architecture design" \
--response-modalities image text \
--aspect-ratio 1:1| Ratio | Resolution | Use Case | Token Cost |
|---|---|---|---|
| 1:1 | 1024×1024 | Social media, avatars | 1290 |
| 16:9 | 1344×768 | Landscapes, banners | 1290 |
| 9:16 | 768×1344 | Mobile, portraits | 1290 |
| 4:3 | 1152×896 | Traditional media | 1290 |
| 3:4 | 896×1152 | Vertical posters | 1290 |
['image']: Generate only images['text']: Generate only text descriptions['image', 'text']: Generate both images and descriptionsProvide existing image + text instructions to modify:
import PIL.Image
img = PIL.Image.open('original.png')
response = client.models.generate_content(
model='gemini-2.5-flash-image',
contents=[
'Add a red balloon floating in the sky',
img
]
)Combine up to 3 source images (recommended):
img1 = PIL.Image.open('background.png')
img2 = PIL.Image.open('foreground.png')
response = client.models.generate_content(
model='gemini-2.5-flash-image',
contents=[
'Combine these images into a cohesive scene',
img1,
img2
]
)Structure effective prompts with three elements:
Example: "A robot in a futuristic city, cyberpunk style with neon lighting and rain-slicked streets"
Quality modifiers:
Text in images:
See references/prompting-guide.md for comprehensive prompt engineering strategies.
The model includes adjustable safety filters. Configure per-request:
config = types.GenerateContentConfig(
response_modalities=['image'],
safety_settings=[
types.SafetySetting(
category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE
)
]
)See references/safety-settings.md for detailed configuration options.
All generated images should be saved to ./docs/assets/ directory:
# Create directory if needed
mkdir -p ./docs/assetsThe helper script automatically saves to this location with timestamped filenames.
Model: gemini-2.5-flash-image
Common issues and solutions:
API key not found:
# Check environment variables
echo $GEMINI_API_KEY
# Verify .env file exists
cat .claude/skills/gemini-image-gen/.env
# or
cat .envSafety filter blocking:
response.prompt_feedback.block_reasonToken limit exceeded:
For detailed information, see:
references/api-reference.md - Complete API specificationsreferences/prompting-guide.md - Advanced prompt engineeringreferences/safety-settings.md - Safety configuration detailsreferences/code-examples.md - Additional implementation examples© einverne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in claude/skills/gemini-image-gen of einverne/dotfiles.
Open the folder on GitHubat commit c6c0686
Gemini Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gemini Image Gen this skilleinverne/dotfiles | 121 | — | ~1.8k | Automated safety check: Notes | MIT | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| BlockRun Image GenerationBlockRunAI/ClawRouter | 6.6k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Antigravity Gemini ImageuluckyXH/OpenMOSS | 1.3k | — | ~730 | Automated safety check: Notes | MIT | |
| FigureMuuuun/luxas | 1.2k | — | ~1.2k | Automated safety check: Pass | MIT | |
| Gemini Image Generatordair-ai/dair-academy-plugins | 614 | 2 repos | ~3.5k | Automated safety check: Notes | MIT |
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
BlockRunAI/ClawRouter
Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.
uluckyXH/OpenMOSS
Generate or edit images using the Antigravity-hosted Gemini image model via the local gateway.
Muuuun/luxas
Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly).
dair-ai/dair-academy-plugins
Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.
ReScienceLab/opc-skills
Generate and edit images using Google Gemini 3 Pro Image (Nano Banana Pro).
einverne/dotfiles
Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.
einverne/dotfiles
Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.
einverne/dotfiles
Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.
einverne/dotfiles
Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.
einverne/dotfiles
Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.
einverne/dotfiles
Presentation creation, editing, and analysis. An agent skill from einverne/dotfiles.
Works with
Categories
Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model. Gemini Image Gen is an agent skill from einverne/dotfiles.5-flash-image model.
Gemini Image Gen fits situations like: generating images; creating visual content; implementing text-to-image features.
Run `npx skills add einverne/dotfiles --skill gemini-image-gen -a claude-code`. Or copy the skill folder (claude/skills/gemini-image-gen in einverne/dotfiles) into .claude/skills/gemini-image-gen in your project. Claude Code loads it when a task matches its description.
Run `npx skills add einverne/dotfiles --skill gemini-image-gen -a codex`. Or copy the skill folder (claude/skills/gemini-image-gen in einverne/dotfiles) into .agents/skills/gemini-image-gen in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-image-gen, .gemini/skills/gemini-image-gen, .github/skills/gemini-image-gen and .opencode/skills/gemini-image-gen in your project.
Going by SKILL.md and its folder, Gemini Image Gen needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write.
SKILL.md names 2 domains. As links in the text: aistudio.google.com and ai.google.dev. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gemini Image Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gemini Image Gen: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars), Antigravity Gemini Image (uluckyXH/OpenMOSS, 1.3k stars) and Figure (Muuuun/luxas, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.
Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.