Agent skill

Generate Nanobanana

by sickn33 in sickn33/agentic-awesome-skills

Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.

MITAuto-check: notesMedia & Creative

Install Generate Nanobanana

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill generate-nanobanana -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills generate-nanobanana --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/generate-nanobanana .claude/skills/generate-nanobanana && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-nanobanana
GitHub stars
47k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
1,338 words
Files
5 (incl. references)
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.

  • Works in 4 steps: Route to a model → Load references → Generate → …
  • Tasks that involve Image generation
  • SKILL.md covers Overview, When to Use This Skill, How It Works and Examples, plus 5 more sections
  • Needs GEMINI_API_KEY

What it does

Generate Nanobanana is an agent skill from sickn33/agentic-awesome-skills. Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/gemini-3-pro-image.md`, `references/gemini-3.1-flash-image.md` and `references/gemini-3.1-flash-lite-image.md`).

It sits in Media & Creative, covering Image generation and Image editing. It works with Google Gemini. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Image generation
  • Tasks that involve Image editing

Example prompts

  • “/generate-nanobanana”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Route to a model
  2. Load references
  3. Generate
  4. Verify and log

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ai.google.dev
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Nanobanana loads about 2.8k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,338 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:112
    ead from the environment or a workspace `.env` the user already set up; it is never logged, printed, or written into a s

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 1,338 words, ~2,750 tokens.

Download SKILL.mdSave it as .claude/skills/generate-nanobanana/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
generate-nanobanana
description
Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.
category
media
risk
critical
source
community
source_repo
AntonioCardenas/generate-nanobanana
source_type
community
date_added
2026-08-04
author
antonio
tags
nanobanana, gemini, google-ai-studio, image-generation, video-generation
tools
claude, cursor, gemini, codex, antigravity
license
MIT

Generate Nanobanana

Overview

generate-nanobanana calls Google's Gemini media models directly through the Gemini API — no third-party routing layer — to generate and edit images and video. It routes each request to the right model tier (draft, standard, quality, or video), loads real reference images instead of relying on text descriptions, gates every paid call behind explicit user approval, and writes a JSON sidecar next to every output recording the exact prompt, model, and cost. It registers a single /generate command.

This skill adapts the workflow (model routing, reference-image handling, sidecar logging) from AntonioCardenas/generate-nanobanana. The actual request shapes in references/ were independently verified against the live Gemini API docs rather than copied from that upstream repo, whose examples predate Google's migration to the Interactions API and use stale, non-functional request methods. Model IDs, request contracts, and pricing all change on Google's own schedule — re-verify against the docs linked from each reference file before relying on this skill in a new session.

When to Use This Skill

  • Use when the user asks to generate, create, or make an image or video, or wants a thumbnail.
  • Use when the user wants to animate a still image, or says "generate on brand" or "generate from reference".
  • Use when the user wants to link or import a folder of reference images (logos, faces, product shots) for reuse across generations.
  • Use when the user invokes /generate or /generate frf <set>, even without naming a specific model.

How It Works

Step 1: Route to a model

Pick the model for the job and read its reference file under references/ before calling anything — each file holds the current, verified request shape for that model.

TaskModelModel IDReference
Image (draft)Nano Banana 2 Litegemini-3.1-flash-lite-imagereferences/gemini-3.1-flash-lite-image.md
Image (standard)Nano Banana 2gemini-3.1-flash-imagereferences/gemini-3.1-flash-image.md
Image (quality, multi-image fusion)Nano Banana Progemini-3-pro-imagereferences/gemini-3-pro-image.md
VideoGemini Omni Flashgemini-omni-flash-previewreferences/gemini-omni-flash-preview.md

All four models are called through the Interactions API (client.interactions.create(...), REST POST /v1beta/interactions) — see each reference file for the exact shape, including reference-image input and, for video, large-output retrieval. Every call is billable; see Step 3.

Draft on Nano Banana 2 Lite first and rerun the picked favorite on Nano Banana 2 or Pro; reserve Pro for heavy multi-image fusion, character-consistent series, or dense on-image text.

Step 2: Load references

Pull real reference images from generations/refs/, or from a named reference set when the request says "on brand" or invokes /generate frf <set>. Never substitute a text description for a reference image (logo, face, brand mark) that already exists — stop and ask if a named reference is missing instead of approximating it.

Reference sets are registered by importing (copying files into generations/refs/<set>/, a snapshot) or linking (recording the source path in generations/refs/sets.json, read live at generation time). A set may carry a style.md whose contents are prepended verbatim to every prompt generated from that set.

Step 3: Generate

Call the Gemini API per the model's reference file. Every generation — image or video — is billable and requires an explicit approval gate: quote the current per-unit price from the live pricing page for the selected model and get explicit user go-ahead before that specific call. One approval covers exactly one call; a rerun needs its own. Run generations one at a time, never in parallel, so approval and cost tracking stay accurate.

No model in this skill documents a seed or reproducibility parameter — do not promise an identical re-roll. For "same image but change X" requests, reuse the exact original prompt and reference images (from the sidecar log) and change only the requested delta; for video, chain edits via previous_interaction_id where supported (see the Omni Flash reference).

Step 4: Verify and log

Confirm the generated file is on disk and non-empty, then write a matching .json sidecar next to it (see Examples) recording the exact model ID, prompt, references used, response id, cost, and timestamp. Never log a generation whose file isn't there, and never write a sidecar for a failed or safety-blocked call.

Examples

Example 1: On-brand thumbnail from a linked reference set
User: generate a thumbnail on brand for the new pricing page

The skill resolves the brand reference set from generations/refs/sets.json, prepends its style.md (if present), picks the relevant reference images (e.g. the logo and a style shot), quotes the current Nano Banana 2 Lite price and gets approval, then saves the result to generations/pricing_page_thumbnail_<timestamp>.png with a sidecar.

Example 2: Sidecar log written beside an output
json
{
  "model": "gemini-3.1-flash-lite-image",
  "prompt": "the exact prompt sent",
  "reference_images": ["generations/refs/brand/logo_dark.png"],
  "reference_set": "brand",
  "response_id": "v1_...",
  "params": { "aspect_ratio": "16:9", "image_size": "1K" },
  "cost": "{price quoted from the live pricing page before running}",
  "created": "2026-07-31T14:20:00Z",
  "approved_by_user": true
}

Best Practices

  • ✅ Quote the current price and get explicit approval before every paid generation — image or video, not just video. A quote is not approval, and each rerun needs its own.
  • ✅ Use real reference images for faces, logos, and brand marks instead of describing them in text.
  • ✅ Read the model's reference file in references/ before calling it — model IDs and request shapes have already changed once in this skill's lifetime (Interactions API migration, gemini-3-pro-image-preview shutdown).
  • ❌ Don't generate "on brand" from an empty or nonexistent reference set — bootstrap the folder and stop until it has at least one real image.
  • ❌ Don't claim a generation is exactly reproducible — no model here documents a seed parameter. Reuse the exact prompt and references instead of promising identical output.
  • ❌ Don't run generations in parallel or reconstruct a prompt from memory when the original's sidecar still has the exact text.
Show full SKILL.md (468 more words)Show less

Limitations

  • Covers Google Gemini models only; there is no multi-provider routing to other image/video generators.
  • Requires a Google AI Studio API key (GEMINI_API_KEY) and, outside Antigravity's native tool fallback, the google-genai Python package.
  • No model documents a seed or reproducibility guarantee; reruns are best-effort via the saved prompt and references, not identical output.
  • Model IDs and pricing are Google's to change; the reference files carry the model IDs verified at the time this skill was last updated, and each links to the live docs to re-verify against.
  • This skill does not replace environment-specific validation, testing, or expert review of generated assets.
  • Stop and ask for clarification if a required reference image, permission, or the API key is missing.

Security & Safety Notes

  • Network — Generation and file-transfer calls go to generativelanguage.googleapis.com; checking current docs or pricing contacts ai.google.dev, and an explicitly approved package install contacts the configured PyPI index. Never send prompts or reference media to any other endpoint.
  • Secrets — GEMINI_API_KEY is only ever read from the environment or a workspace .env the user already set up; it is never logged, printed, or written into a sidecar, prompt, or committed file. The skill never creates or edits .env, .env.example, or .gitignore itself.
  • File writes — skill-authored project outputs are confined to the workspace's generations/ folder (including generations/refs/, REST request/response files, and sets.json); nothing is written outside the current project except an explicitly approved package installation in its selected environment.
  • Package installs — only the official google-genai PyPI package, and only when missing; never installed silently or alongside any other package.
  • Cost — every call spends real money against the user's Google AI Studio billing; that, plus filesystem writes, is why this skill is risk: critical rather than safe.
  • Treat any change that would add a new network endpoint, a new package install, or a write outside generations/ as a design decision for the user to approve, not something to do quietly.

Common Pitfalls

  • Problem: Requesting "on brand" generation before any reference images exist. Solution: Create generations/refs/<name>/, tell the user its path, and wait for at least one image before generating.
  • Problem: Varying an existing image by re-describing it from memory. Solution: Read the original's sidecar for its exact prompt and references, and change only the requested delta.
  • Problem: Running an image or video generation without a cost quote. Solution: Always quote the current per-unit price from the live pricing page and get explicit approval before submitting any paid call.
  • Problem: Calling a model ID from memory instead of the reference file. Solution: Model IDs shift (e.g. gemini-3-pro-image-preview was shut down and replaced by gemini-3-pro-image) — always read references/<model>.md first.
  • @image-generator - Nano Banana Pro image generation and editing without the multi-model routing, reference-set library, or cost-gate workflow.
  • @nanobanana-ppt-skills - AI-powered PPT generation with document analysis and styled images.
  • @2slides-ppt-generator - Presentation generation via 2slides API.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/generate-nanobanana of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/gemini-3-pro-image.md
  • references/gemini-3.1-flash-image.md
  • references/gemini-3.1-flash-lite-image.md
  • references/gemini-omni-flash-preview.md

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Generate Nanobanana next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Nanobanana compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Nanobanana this skillsickn33/agentic-awesome-skills47k1 repos~2.8kAutomated safety check: NotesMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT
Antigravity Gemini ImageuluckyXH/OpenMOSS1.3k—~730Automated safety check: NotesMIT
FigureMuuuun/luxas1.2k—~1.2kAutomated safety check: PassMIT
Gemini Image Generatordair-ai/dair-academy-plugins6142 repos~3.5kAutomated safety check: NotesMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Antigravity Gemini Image

    uluckyXH/OpenMOSS

    Generate or edit images using the Antigravity-hosted Gemini image model via the local gateway.

    1.3k GitHub stars~730 tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Figure

    Muuuun/luxas

    Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly).

    1.2k GitHub stars~1.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Gemini Image Generator

    dair-ai/dair-academy-plugins

    Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

    614 GitHub starsUsed in 2 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Nanobanana

    ReScienceLab/opc-skills

    Generate and edit images using Google Gemini 3 Pro Image (Nano Banana Pro).

    1.8k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Generate Nanobanana

What does Generate Nanobanana do?

Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call. Generate Nanobanana is an agent skill from sickn33/agentic-awesome-skills. Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.

When should I use Generate Nanobanana?

Generate Nanobanana fits situations like: tasks that involve Image generation; tasks that involve Image editing.

How do I install Generate Nanobanana in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill generate-nanobanana -a claude-code`. Or copy the skill folder (skills/generate-nanobanana in sickn33/agentic-awesome-skills) into .claude/skills/generate-nanobanana in your project. Claude Code loads it when a task matches its description.

How do I install Generate Nanobanana in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill generate-nanobanana -a codex`. Or copy the skill folder (skills/generate-nanobanana in sickn33/agentic-awesome-skills) into .agents/skills/generate-nanobanana in your project. Codex loads it when a task matches its description.

Can I use Generate Nanobanana in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill generate-nanobanana -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-nanobanana, .gemini/skills/generate-nanobanana, .github/skills/generate-nanobanana and .opencode/skills/generate-nanobanana in your project.

What does Generate Nanobanana need to run?

Going by SKILL.md and its folder, Generate Nanobanana needs credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Generate Nanobanana access the network?

SKILL.md names 2 domains. As links in the text: ai.google.dev and github.com. This is read from the text; nothing was executed.

Is Generate Nanobanana safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Generate Nanobanana use?

Generate Nanobanana is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Nanobanana use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Generate Nanobanana?

Skills that share tags, products or a category with Generate Nanobanana: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), BlockRun Image Generation (BlockRunAI/ClawRouter, 6.6k stars), Antigravity Gemini Image (uluckyXH/OpenMOSS, 1.3k stars) and Figure (Muuuun/luxas, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Nanobanana?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.