Agent skill

Gemini Image Generator

by dair-ai in dair-ai/dair-academy-plugins

Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

MITAuto-check: notesMedia & Creative

Install Gemini Image Generator

skills CLI
$ npx skills add dair-ai/dair-academy-plugins --skill image-generator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dair-ai/dair-academy-plugins image-generator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/image-generator/skills/image-generator .claude/skills/image-generator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-generator
GitHub stars
614
Used in
2 other repos
Token cost
~3.5k tokens
SKILL.md length
597 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

  • Works in 8 steps: Read and encode the image to base64 → Send image with edit prompt (File-Based… → Extract and save the edited image → …
  • Generating a logo, mockup or illustration from a text description
  • SKILL.md covers IMPORTANT: Setup Required, Pre-flight Check, Configuration and Iterating on User-Provided…, plus 5 more sections
  • Calls curl, python3 and jq; reaches generativelanguage.googleapis.com; needs GEMINI_API_KEY

What it does

This skill generates and edits images with Google's Gemini image model named in the skill, calling the Gemini API from shell commands. It covers text-to-image generation, editing a photo you point to, and combining elements from several input images.

A GEMINI_API_KEY environment variable is required. The agent checks for it before any API call and stops with setup instructions if it is missing. For edits, the agent encodes the image in base64 and writes the request body to a file, because large inline arguments cause argument-list-too-long errors. It then extracts the returned image and saves it to disk.

A free key can be obtained from Google AI Studio and exported in the shell profile. Typical edits include adding an object to a photo, replacing a background or moving clothing from one image to another.

When your agent uses it

  • Generating a logo, mockup or illustration from a text description
  • Editing a photo, for example changing its background
  • Combining elements from two images into one
  • Iterating on an image you already saved locally

Example prompts

  • “Generate a flat-style logo for a coffee roaster called Ember.”
  • “Edit ./photos/team.png and add a santa hat to the person on the left.”
  • “Make the background of ./photos/portrait.png a sunset beach.”
  • “Put the jacket from jacket.png onto the person in model.png.”

Requirements

  • A GEMINI_API_KEY from Google AI Studio exported in the shell
  • Network access to the Gemini API
  • Pre-approved tools (allowed-tools): Read, Write, Bash, WebFetch

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Read and encode the image to base64
  2. Send image with edit prompt (File-Based Approach)
  3. Extract and save the edited image
  4. Be Descriptive, Not Keyword-Based
  5. Specify Style and Mood
  6. For Text in Images
  7. For Editing
  8. For Product/Commercial Images

What it can do on your machine

Read from SKILL.md and the folder at commit 0abffdc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Bash
    • WebFetch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • python3
    • jq
    • black
    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • generativelanguage.googleapis.com

    Also links to:

    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Image Generator loads about 3.5k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 597 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Bash, WebFetch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dair-ai/dair-academy-plugins at commit 0abffdc, republished under its MIT licence (© dair-ai). 597 words, ~3,517 tokens.

Download SKILL.mdSave it as .claude/skills/image-generator/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
image-generator
description
Generate and edit images using Gemini's Nano Banana Pro model (gemini-3-pro-image-preview). Use this skill when the user asks you to generate images, create visuals, edit photos, create logos, generate product mockups, or perform any image generation/editing task.
allowed-tools
Read, Write, Bash, WebFetch

Image Generator

This skill generates and edits images using Google's Gemini Nano Banana Pro model (gemini-3-pro-image-preview).

IMPORTANT: Setup Required

Before using this skill, the user must set the GEMINI_API_KEY environment variable:

  1. Get a free API key from Google AI Studio
  2. Export the key in your shell profile (~/.zshrc, ~/.bashrc, etc.):
    bash
    export GEMINI_API_KEY="your_api_key_here"
  3. Restart your terminal or run source ~/.zshrc (or ~/.bashrc)

The skill will not work without this configuration.

Pre-flight Check

Before making any API call, verify the key is set:

bash
if [ -z "$GEMINI_API_KEY" ]; then
  echo "ERROR: GEMINI_API_KEY is not set. Please export it in your shell profile."
  exit 1
fi

If the key is missing, stop and tell the user to set it using the instructions above.

Configuration

Model: gemini-3-pro-image-preview

API Key: Read from the GEMINI_API_KEY environment variable

Iterating on User-Provided Images

When the user provides a path to an image they want to edit or iterate on, use this workflow:

Step 1: Read and encode the image to base64
bash
# Get the image path from user
IMG_PATH="/path/to/user/image.png"

# Detect mime type
if [[ "$IMG_PATH" == *.png ]]; then
    MIME_TYPE="image/png"
elif [[ "$IMG_PATH" == *.jpg ]] || [[ "$IMG_PATH" == *.jpeg ]]; then
    MIME_TYPE="image/jpeg"
elif [[ "$IMG_PATH" == *.webp ]]; then
    MIME_TYPE="image/webp"
else
    MIME_TYPE="image/png"
fi

# Encode to base64 (works on both macOS and Linux)
if [[ "$(uname)" == "Darwin" ]]; then
    IMG_BASE64=$(base64 -i "$IMG_PATH")
else
    IMG_BASE64=$(base64 -w0 "$IMG_PATH")
fi
Step 2: Send image with edit prompt (File-Based Approach)

IMPORTANT: Always use a file-based approach for the request body. Base64-encoded images are too large for command-line arguments and will cause "argument list too long" errors.

bash
# User's edit request
EDIT_PROMPT="Add a santa hat to the person in this image"

# Write request to a JSON file (avoids command line length limits)
cat > /tmp/gemini_request.json << JSONEOF
{
  "contents": [{
    "parts": [
      {"text": "$EDIT_PROMPT"},
      {
        "inline_data": {
          "mime_type": "$MIME_TYPE",
          "data": "$IMG_BASE64"
        }
      }
    ]
  }],
  "generationConfig": {
    "responseModalities": ["TEXT", "IMAGE"]
  }
}
JSONEOF

# Call the API using the file
curl -s -X POST \
  "https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d @/tmp/gemini_request.json > /tmp/gemini_response.json
Step 3: Extract and save the edited image
bash
# Extract image from response and save
python3 -c "
import json
import base64

with open('/tmp/gemini_response.json') as f:
    data = json.load(f)

for part in data['candidates'][0]['content']['parts']:
    if 'inlineData' in part:
        img_data = part['inlineData']['data']
        mime = part['inlineData']['mimeType']
        ext = 'png' if 'png' in mime else 'jpg'
        with open('edited_image.' + ext, 'wb') as out:
            out.write(base64.b64decode(img_data))
        print(f'Saved: edited_image.{ext}')
    elif 'text' in part:
        print(part['text'])
"
Complete Example (File-Based)

For iterating on images, always use file-based requests:

bash
# Variables
IMG_PATH="/path/to/image.png"
EDIT_PROMPT="Make the background a sunset beach"
OUTPUT_PATH="edited_output.png"
# Detect mime type and encode
MIME_TYPE=$([[ "$IMG_PATH" == *.png ]] && echo "image/png" || echo "image/jpeg")
IMG_BASE64=$(base64 -i "$IMG_PATH" 2>/dev/null || base64 -w0 "$IMG_PATH")

# Write request to file (required - base64 images are too large for command line)
cat > /tmp/gemini_request.json << JSONEOF
{
  "contents": [{
    "parts": [
      {"text": "$EDIT_PROMPT"},
      {"inline_data": {"mime_type": "$MIME_TYPE", "data": "$IMG_BASE64"}}
    ]
  }],
  "generationConfig": {
    "responseModalities": ["TEXT", "IMAGE"]
  }
}
JSONEOF

# Call API and extract image
curl -s -X POST \
  "https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d @/tmp/gemini_request.json > /tmp/gemini_response.json

# Save the output image
python3 -c "
import json, base64
with open('/tmp/gemini_response.json') as f:
    data = json.load(f)
for part in data.get('candidates', [{}])[0].get('content', {}).get('parts', []):
    if 'inlineData' in part:
        with open('$OUTPUT_PATH', 'wb') as f:
            f.write(base64.b64decode(part['inlineData']['data']))
        print('Saved: $OUTPUT_PATH')
"
Multi-Image Input (Combine/Compose)

To combine elements from multiple images (also uses file-based approach):

bash
IMG1_PATH="/path/to/image1.png"
IMG2_PATH="/path/to/image2.png"
PROMPT="Put the dress from the first image on the person in the second image"
IMG1_BASE64=$(base64 -i "$IMG1_PATH" 2>/dev/null || base64 -w0 "$IMG1_PATH")
IMG2_BASE64=$(base64 -i "$IMG2_PATH" 2>/dev/null || base64 -w0 "$IMG2_PATH")

# Write request to file
cat > /tmp/gemini_request.json << JSONEOF
{
  "contents": [{
    "parts": [
      {"text": "$PROMPT"},
      {"inline_data": {"mime_type": "image/png", "data": "$IMG1_BASE64"}},
      {"inline_data": {"mime_type": "image/png", "data": "$IMG2_BASE64"}}
    ]
  }],
  "generationConfig": {"responseModalities": ["TEXT", "IMAGE"]}
}
JSONEOF

curl -s -X POST \
  "https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d @/tmp/gemini_request.json > /tmp/gemini_response.json

Capabilities

Text-to-Image Generation
  • Generate high-quality images from text descriptions
  • Support for photorealistic, stylized, and artistic outputs
  • Accurate text rendering in images (logos, infographics, diagrams)
Image Editing
  • Add or remove elements from images
  • Inpainting with semantic masking (edit specific parts)
  • Style transfer (apply artistic styles to photos)
  • Multi-image composition (combine elements from multiple images)
Advanced Features
  • High Resolution: 1K, 2K, or 4K output
  • Aspect Ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
  • Google Search Grounding: Generate images based on real-time data
  • Multi-turn Editing: Iteratively refine images through conversation
  • Up to 14 Reference Images: Combine multiple inputs for complex compositions

API Usage

Basic Text-to-Image (Python)
python
from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-3-pro-image-preview",
    contents=["Your prompt here"],
    config=types.GenerateContentConfig(
        response_modalities=['TEXT', 'IMAGE'],
        image_config=types.ImageConfig(
            aspect_ratio="16:9",  # Optional
            image_size="2K"       # Optional: "1K", "2K", "4K"
        )
    )
)

for part in response.parts:
    if part.text is not None:
        print(part.text)
    elif part.inline_data is not None:
        image = part.as_image()
        image.save("generated_image.png")
Basic Text-to-Image (JavaScript)
javascript
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({});

const response = await ai.models.generateContent({
    model: "gemini-3-pro-image-preview",
    contents: "Your prompt here",
    config: {
        responseModalities: ['TEXT', 'IMAGE'],
        imageConfig: {
            aspectRatio: "16:9",
            imageSize: "2K"
        }
    }
});

for (const part of response.candidates[0].content.parts) {
    if (part.text) {
        console.log(part.text);
    } else if (part.inlineData) {
        const buffer = Buffer.from(part.inlineData.data, "base64");
        fs.writeFileSync("generated_image.png", buffer);
    }
}
REST API (curl)
bash
curl -s -X POST \
  "https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{
      "parts": [{"text": "Your prompt here"}]
    }],
    "generationConfig": {
      "responseModalities": ["TEXT", "IMAGE"],
      "imageConfig": {
        "aspectRatio": "16:9",
        "imageSize": "2K"
      }
    }
  }' | jq -r '.candidates[0].content.parts[] | select(.inlineData) | .inlineData.data' | base64 --decode > output.png
Image Editing (with input image)
python
from google import genai
from google.genai import types
from PIL import Image

client = genai.Client()

input_image = Image.open('input.png')
prompt = "Add a wizard hat to the cat in this image"

response = client.models.generate_content(
    model="gemini-3-pro-image-preview",
    contents=[prompt, input_image],
    config=types.GenerateContentConfig(
        response_modalities=['TEXT', 'IMAGE']
    )
)

for part in response.parts:
    if part.inline_data is not None:
        image = part.as_image()
        image.save("edited_image.png")
Multi-Image Composition
python
from google import genai
from google.genai import types
from PIL import Image

client = genai.Client()

image1 = Image.open('dress.png')
image2 = Image.open('model.png')
prompt = "Put the dress from the first image on the model from the second image"

response = client.models.generate_content(
    model="gemini-3-pro-image-preview",
    contents=[image1, image2, prompt],
    config=types.GenerateContentConfig(
        response_modalities=['TEXT', 'IMAGE'],
        image_config=types.ImageConfig(
            aspect_ratio="3:4",
            image_size="2K"
        )
    )
)
With Google Search Grounding
python
from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-3-pro-image-preview",
    contents="Visualize the current weather forecast for San Francisco",
    config=types.GenerateContentConfig(
        response_modalities=['TEXT', 'IMAGE'],
        image_config=types.ImageConfig(aspect_ratio="16:9"),
        tools=[{"google_search": {}}]
    )
)

Prompting Best Practices

1. Be Descriptive, Not Keyword-Based

Instead of: cat, wizard hat, cute Write: A fluffy orange cat wearing a small knitted wizard hat, sitting on a wooden floor with soft natural lighting from a window

Show full SKILL.md (228 more words)Show less
2. Specify Style and Mood
  • Photography terms: "shot with 85mm lens", "soft bokeh background", "golden hour lighting"
  • Artistic styles: "in the style of Van Gogh", "minimalist illustration", "photorealistic"
  • Mood: "warm and cozy atmosphere", "dramatic noir lighting"
3. For Text in Images

Be explicit about:

  • The exact text to render
  • Font style (descriptively): "clean, bold, sans-serif font"
  • Placement and size
4. For Editing
  • Describe what to change and what to preserve
  • Use "keep everything else unchanged"
  • Reference specific elements clearly
5. For Product/Commercial Images

Mention:

  • Lighting setup: "three-point softbox lighting"
  • Background: "clean white studio background"
  • Camera angle: "slightly elevated 45-degree shot"

Resolution and Aspect Ratio Reference

Aspect Ratio1K Resolution2K Resolution4K Resolution
1:11024x10242048x20484096x4096
16:91376x7682752x15365504x3072
9:16768x13761536x27523072x5504
3:21264x8482528x16965056x3392
2:3848x12641696x25283392x5056

Common Use Cases

Logo Creation
Create a modern, minimalist logo for a coffee shop called 'The Daily Grind'.
The text should be in a clean, bold, sans-serif font.
Black and white color scheme. Put the logo in a circle.
Product Photography
A high-resolution, studio-lit product photograph of a minimalist ceramic
coffee mug in matte black on a polished concrete surface. Three-point
softbox lighting with soft, diffused highlights. Slightly elevated
45-degree camera angle. Sharp focus on steam rising from the coffee.
Style Transfer
Transform this photograph of a city street at night into Vincent van Gogh's
'Starry Night' style. Preserve the composition but render with swirling,
impasto brushstrokes and deep blues with bright yellows.
Infographic
Create a vibrant infographic explaining photosynthesis as a recipe.
Show "ingredients" (sunlight, water, CO2) and "finished dish" (sugar/energy).
Style like a colorful kids' cookbook, suitable for 4th graders.

Error Handling

Common issues:

  • No image returned: Check that response_modalities includes 'IMAGE'
  • Safety filters: Some prompts may be blocked; try rephrasing
  • Rate limits: Implement exponential backoff for retries
  • Large images: For 4K, ensure sufficient timeout settings

Dependencies

To use the Python SDK:

bash
pip install google-genai pillow

For JavaScript:

bash
npm install @google/genai

Important Notes

  • All generated images include a SynthID watermark
  • The model uses a "thinking" process for complex prompts
  • For best text rendering, generate text first, then request image with that text
  • Images are not stored by the API - save outputs locally

© dair-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/image-generator/skills/image-generator of dair-ai/dair-academy-plugins.

  • SKILL.md
  • .env.example

Open the folder on GitHubat commit 0abffdc

Used in 2 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in dair-ai/dair-academy-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gemini Image Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Image Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Image Generator this skilldair-ai/dair-academy-plugins6142 repos~3.5kAutomated safety check: NotesMIT
Generate ImageK-Dense-AI/claude-scientific-writer2.4k1 repos~3.8kAutomated safety check: NotesMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Logo Generatorop7418/logo-generator-skill2.2k—~1.8kAutomated safety check: NotesNone
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT
Antigravity Gemini ImageuluckyXH/OpenMOSS1.3k—~730Automated safety check: NotesMIT

Similar skills

  • Generate Image

    K-Dense-AI/claude-scientific-writer

    Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow).

    2.4k GitHub starsUsed in 1 repo~3.8k tokens
    Media & CreativeAuto-check: notes
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Logo Generator

    op7418/logo-generator-skill

    Generate professional SVG logos and high-end showcase images.

    2.2k GitHub stars~1.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Antigravity Gemini Image

    uluckyXH/OpenMOSS

    Generate or edit images using the Antigravity-hosted Gemini image model via the local gateway.

    1.3k GitHub stars~730 tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Figure

    Muuuun/luxas

    Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly).

    1.2k GitHub stars~1.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from dair-ai/dair-academy-plugins

All 8 skills in this repo
  • Research Wiki Builder

    dair-ai/dair-academy-plugins

    Creates and maintains configurable research wikis: scaffold a folder, add sources, compile pages and indexes, and file query answers back.

    614 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Survey Paper Generator

    dair-ai/dair-academy-plugins

    Builds a single-file HTML survey paper on an AI or ML topic from a research bundle the agent curates, with prose and SVG figures written by Kimi K2.6.

    614 GitHub starsUsed in 2 repos~2.1k tokens
    Auto-check: notes
  • YouTube Talk Notetaker

    dair-ai/dair-academy-plugins

    Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.

    614 GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Learn

    dair-ai/dair-academy-plugins

    Help a user learn a topic through adaptive tutoring, lesson planning, practice, retrieval checks, explanations, study guides, or exercises.

    614 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • LLM Council on Fireworks AI

    dair-ai/dair-academy-plugins

    Has several open-weight models answer a question, rank each other's anonymized answers, then lets a chairman model write the final response through Fireworks AI.

    614 GitHub stars~5k tokensUpdated 2 mo ago
    Auto-check: notes
  • X Agent Intelligence Feed

    dair-ai/dair-academy-plugins

    Builds a self-contained HTML digest of AI and agent news pulled from chosen X accounts through the official X MCP server, grouped into categories like Coding Agents and Agent Research.

    614 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Gemini Image Generator

What does Gemini Image Generator do?

Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition. This skill generates and edits images with Google's Gemini image model named in the skill, calling the Gemini API from shell commands. It covers text-to-image generation, editing a photo you point to, and combining elements from several input images.

When should I use Gemini Image Generator?

Gemini Image Generator fits situations like: generating a logo, mockup or illustration from a text description; editing a photo, for example changing its background; combining elements from two images into one; iterating on an image you already saved locally.

How do I install Gemini Image Generator in Claude Code?

Run `npx skills add dair-ai/dair-academy-plugins --skill image-generator -a claude-code`. Or copy the skill folder (plugins/image-generator/skills/image-generator in dair-ai/dair-academy-plugins) into .claude/skills/image-generator in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Image Generator in Codex?

Run `npx skills add dair-ai/dair-academy-plugins --skill image-generator -a codex`. Or copy the skill folder (plugins/image-generator/skills/image-generator in dair-ai/dair-academy-plugins) into .agents/skills/image-generator in your project. Codex loads it when a task matches its description.

Can I use Gemini Image Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dair-ai/dair-academy-plugins --skill image-generator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-generator, .gemini/skills/image-generator, .github/skills/image-generator and .opencode/skills/image-generator in your project.

What does Gemini Image Generator need to run?

Going by SKILL.md and its folder, Gemini Image Generator needs the command-line tools its instructions call (curl, python3, jq, black, pip and npm) and credentials named GEMINI_API_KEY. Our summary lists: A GEMINI_API_KEY from Google AI Studio exported in the shell; Network access to the Gemini API. Its frontmatter pre-approves these tools: Read, Write, Bash, WebFetch.

Does Gemini Image Generator access the network?

SKILL.md names 2 domains. In commands or code: generativelanguage.googleapis.com; the agent is likely to contact it when it follows the instructions. As links in the text: aistudio.google.com. This is read from the text; nothing was executed.

Is Gemini Image Generator safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gemini Image Generator use?

Gemini Image Generator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Image Generator use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini Image Generator?

Skills that share tags, products or a category with Gemini Image Generator: Generate Image (K-Dense-AI/claude-scientific-writer, 2.4k stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Logo Generator (op7418/logo-generator-skill, 2.2k stars) and BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Image Generator?

dair-ai (a GitHub organization) maintains it in dair-ai/dair-academy-plugins, which has 614 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on July 21, 2026.

Source: dair-ai/dair-academy-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.