Agent skill

Gemini Image Gen

by einverne in einverne/dotfiles

Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.

MITAuto-check: notesMedia & Creative

Install Gemini Image Gen

skills CLI
$ npx skills add einverne/dotfiles --skill gemini-image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install einverne/dotfiles gemini-image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-image-gen .claude/skills/gemini-image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-image-gen
GitHub stars
121
Token cost
~1.8k tokens
SKILL.md length
504 words
Files
9 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.

  • Works in 3 steps: Process environment: export… → Skill directory:… → Project directory: ./.env (project root)
  • Generating images
  • SKILL.md covers When to Use This Skill, Prerequisites, Quick Start and Key Features, plus 8 more sections
  • Runs Python scripts from its folder; calls python and pip; needs GEMINI_API_KEY

What it does

Gemini Image Gen is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model. Use when generating images, creating visual content, or implementing text-to-image features. Supports text-to-image, image editing, multi-image composition, and iterative refinement.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `README.md`, `SKILL_CREATION_SUMMARY.md` and `references/api-reference.md`).

It sits in Media & Creative, covering Image generation and Image editing. It works with Google Gemini. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is MIT.

When your agent uses it

  • Generating images
  • Creating visual content
  • Implementing text-to-image features

Example prompts

  • “/gemini-image-gen”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Skill directory: .claude/skills/gemini-image-gen/.env
  3. Project directory: ./.env (project root)

What it can do on your machine

Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com
    • ai.google.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Image Gen loads about 1.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 504 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:32
    tory**: `.claude/skills/gemini-image-gen/.env`
  • NoteMentions a .env fileSKILL.md:33
    3. **Project directory**: `./.env` (project root)
  • NoteMentions a .env fileSKILL.md:37
    Create `.env` file with:
  • NoteMentions a .env fileSKILL.md:224
    # Verify .env file exists
  • NoteMentions a .env fileSKILL.md:225
    cat .claude/skills/gemini-image-gen/.env
  • NoteMentions a .env fileSKILL.md:227
    cat .env
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from einverne/dotfiles at commit c6c0686, republished under its MIT licence (© einverne). 504 words, ~1,801 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-image-gen/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
gemini-image-gen
description
Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model. Use when generating images, creating visual content, or implementing text-to-image features. Supports text-to-image, image editing, multi-image composition, and iterative refinement.
allowed-tools
Bash, Read, Write
license
MIT
version
1.0.0

Gemini Image Generation Skill

Generate high-quality images using Google's Gemini 2.5 Flash Image model with text prompts, image editing, and multi-image composition capabilities.

When to Use This Skill

Use this skill when you need to:

  • Generate images from text descriptions
  • Edit existing images by adding/removing elements or changing styles
  • Combine multiple source images into new compositions
  • Iteratively refine images through conversational editing
  • Create visual content for documentation, design, or creative projects

Prerequisites

API Key Setup

The skill automatically detects your GEMINI_API_KEY in this order:

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Skill directory: .claude/skills/gemini-image-gen/.env
  3. Project directory: ./.env (project root)

Get your API key: Visit Google AI Studio

Create .env file with:

bash
GEMINI_API_KEY=your_api_key_here
Python Setup

Install required package:

bash
pip install google-genai

Quick Start

Basic Text-to-Image Generation
python
from google import genai
from google.genai import types
import os

# API key detection handled automatically by helper script
client = genai.Client(api_key=os.getenv('GEMINI_API_KEY'))

response = client.models.generate_content(
    model='gemini-2.5-flash-image',
    contents='A serene mountain landscape at sunset with snow-capped peaks',
    config=types.GenerateContentConfig(
        response_modalities=['image'],
        aspect_ratio='16:9'
    )
)

# Save to ./docs/assets/
for i, part in enumerate(response.candidates[0].content.parts):
    if part.inline_data:
        with open(f'./docs/assets/generated-{i}.png', 'wb') as f:
            f.write(part.inline_data.data)
Using the Helper Script

For convenience, use the provided helper script that handles API key detection and file saving:

bash
# Generate single image
python .claude/skills/gemini-image-gen/scripts/generate.py \
  "A futuristic city with flying cars" \
  --aspect-ratio 16:9 \
  --output ./docs/assets/city.png

# Generate with specific modalities
python .claude/skills/gemini-image-gen/scripts/generate.py \
  "Modern architecture design" \
  --response-modalities image text \
  --aspect-ratio 1:1

Key Features

Aspect Ratios
RatioResolutionUse CaseToken Cost
1:11024×1024Social media, avatars1290
16:91344×768Landscapes, banners1290
9:16768×1344Mobile, portraits1290
4:31152×896Traditional media1290
3:4896×1152Vertical posters1290
Response Modalities
  • ['image']: Generate only images
  • ['text']: Generate only text descriptions
  • ['image', 'text']: Generate both images and descriptions
Image Editing

Provide existing image + text instructions to modify:

python
import PIL.Image

img = PIL.Image.open('original.png')
response = client.models.generate_content(
    model='gemini-2.5-flash-image',
    contents=[
        'Add a red balloon floating in the sky',
        img
    ]
)
Multi-Image Composition

Combine up to 3 source images (recommended):

python
img1 = PIL.Image.open('background.png')
img2 = PIL.Image.open('foreground.png')

response = client.models.generate_content(
    model='gemini-2.5-flash-image',
    contents=[
        'Combine these images into a cohesive scene',
        img1,
        img2
    ]
)

Prompt Engineering Tips

Structure effective prompts with three elements:

  1. Subject: What to generate ("a robot")
  2. Context: Environmental setting ("in a futuristic city")
  3. Style: Artistic treatment ("cyberpunk style, neon lighting")

Example: "A robot in a futuristic city, cyberpunk style with neon lighting and rain-slicked streets"

Quality modifiers:

  • Add terms like "4K", "HDR", "high-quality", "professional photography"
  • Specify camera settings: "35mm lens", "shallow depth of field", "golden hour lighting"

Text in images:

  • Limit to 25 characters maximum
  • Use up to 3 distinct phrases
  • Specify font styles: "bold sans-serif title" or "handwritten script"

See references/prompting-guide.md for comprehensive prompt engineering strategies.

Show full SKILL.md (193 more words)Show less

Safety Settings

The model includes adjustable safety filters. Configure per-request:

python
config = types.GenerateContentConfig(
    response_modalities=['image'],
    safety_settings=[
        types.SafetySetting(
            category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
            threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE
        )
    ]
)

See references/safety-settings.md for detailed configuration options.

Output Management

All generated images should be saved to ./docs/assets/ directory:

bash
# Create directory if needed
mkdir -p ./docs/assets

The helper script automatically saves to this location with timestamped filenames.

Model Specifications

Model: gemini-2.5-flash-image

  • Input tokens: Up to 65,536
  • Output tokens: Up to 32,768
  • Supported inputs: Text and images
  • Supported outputs: Text and images
  • Knowledge cutoff: June 2025
  • Features: Image generation, structured outputs, batch API, caching

Limitations

  • Maximum 3 input images recommended for best results
  • Text rendering works best when generated separately first
  • Does not support audio/video inputs
  • Regional restrictions on child image uploads (EEA, CH, UK)
  • Optimal language support: English, Spanish (Mexico), Japanese, Mandarin, Hindi

Error Handling

Common issues and solutions:

API key not found:

bash
# Check environment variables
echo $GEMINI_API_KEY

# Verify .env file exists
cat .claude/skills/gemini-image-gen/.env
# or
cat .env

Safety filter blocking:

  • Review response.prompt_feedback.block_reason
  • Adjust safety settings if appropriate for your use case
  • Modify prompt to avoid triggering filters

Token limit exceeded:

  • Reduce prompt length
  • Use fewer input images
  • Simplify image editing instructions

Reference Documentation

For detailed information, see:

  • references/api-reference.md - Complete API specifications
  • references/prompting-guide.md - Advanced prompt engineering
  • references/safety-settings.md - Safety configuration details
  • references/code-examples.md - Additional implementation examples

Resources

© einverne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in claude/skills/gemini-image-gen of einverne/dotfiles.

  • SKILL.md
  • .env.example
  • README.md
  • SKILL_CREATION_SUMMARY.md
  • references/api-reference.md
  • references/code-examples.md
  • references/prompting-guide.md
  • references/safety-settings.md
  • scripts/generate.py

Open the folder on GitHubat commit c6c0686

Compare with similar skills

Gemini Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Image Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Image Gen this skilleinverne/dotfiles121—~1.8kAutomated safety check: NotesMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT
Antigravity Gemini ImageuluckyXH/OpenMOSS1.3k—~730Automated safety check: NotesMIT
FigureMuuuun/luxas1.2k—~1.2kAutomated safety check: PassMIT
Gemini Image Generatordair-ai/dair-academy-plugins6142 repos~3.5kAutomated safety check: NotesMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Antigravity Gemini Image

    uluckyXH/OpenMOSS

    Generate or edit images using the Antigravity-hosted Gemini image model via the local gateway.

    1.3k GitHub stars~730 tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Figure

    Muuuun/luxas

    Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly).

    1.2k GitHub stars~1.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Gemini Image Generator

    dair-ai/dair-academy-plugins

    Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

    614 GitHub starsUsed in 2 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Nanobanana

    ReScienceLab/opc-skills

    Generate and edit images using Google Gemini 3 Pro Image (Nano Banana Pro).

    1.8k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from einverne/dotfiles

All 39 skills in this repo
  • DOCX

    einverne/dotfiles

    Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.

    121 GitHub starsUsed in 35 repos~2.5k tokens
    Auto-check: notes
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Auto-check passed
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

    121 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check: notes
  • PPTX

    einverne/dotfiles

    Presentation creation, editing, and analysis. An agent skill from einverne/dotfiles.

    121 GitHub starsUsed in 38 repos~6.4k tokens
    Auto-check: notes

Works with

Questions about Gemini Image Gen

What does Gemini Image Gen do?

Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model. Gemini Image Gen is an agent skill from einverne/dotfiles.5-flash-image model.

When should I use Gemini Image Gen?

Gemini Image Gen fits situations like: generating images; creating visual content; implementing text-to-image features.

How do I install Gemini Image Gen in Claude Code?

Run `npx skills add einverne/dotfiles --skill gemini-image-gen -a claude-code`. Or copy the skill folder (claude/skills/gemini-image-gen in einverne/dotfiles) into .claude/skills/gemini-image-gen in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Image Gen in Codex?

Run `npx skills add einverne/dotfiles --skill gemini-image-gen -a codex`. Or copy the skill folder (claude/skills/gemini-image-gen in einverne/dotfiles) into .agents/skills/gemini-image-gen in your project. Codex loads it when a task matches its description.

Can I use Gemini Image Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-image-gen, .gemini/skills/gemini-image-gen, .github/skills/gemini-image-gen and .opencode/skills/gemini-image-gen in your project.

What does Gemini Image Gen need to run?

Going by SKILL.md and its folder, Gemini Image Gen needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write.

Does Gemini Image Gen access the network?

SKILL.md names 2 domains. As links in the text: aistudio.google.com and ai.google.dev. This is read from the text; nothing was executed.

Is Gemini Image Gen safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Image Gen use?

Gemini Image Gen is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Image Gen use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Image Gen?

Skills that share tags, products or a category with Gemini Image Gen: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars), Antigravity Gemini Image (uluckyXH/OpenMOSS, 1.3k stars) and Figure (Muuuun/luxas, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Image Gen?

einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.

Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.