Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.

MITAuto-check passedMedia & Creative

Install Gemini

skills CLI
$ npx skills add Anil-matcha/awesome-muse-connectors --skill gemini -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Anil-matcha/awesome-muse-connectors gemini --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Anil-matcha/awesome-muse-connectors.git skills-src && mkdir -p .claude/skills && cp -r skills-src/connectors/gemini .claude/skills/gemini && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini
GitHub stars
1.3k
Token cost
~778 tokens
SKILL.md length
250 words
Files
2
Skills in repo
153
Repo updated
First seen
Licence
MIT

At a glance

Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.

  • Works in 4 steps: COST WARNING: all media generation… → Every generation spends real money: Veo… → Veo renders take minutes. Poll op-status… → …
  • Phrases: gemini image
  • SKILL.md covers Purpose, Tooling, Auth and Operating Rules, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Gemini is an agent skill from Anil-matcha/awesome-muse-connectors. Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing. Trigger phrases: gemini image, nano banana, imagen, veo video, google veo, gemini tts.

Its SKILL.md is about 780 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `bin/gemini.py`).

It sits in Media & Creative, covering Image generation, Text to speech and voice and AI video generation. It works with Google Gemini and Google Veo. The repository describes itself as: A source-backed catalog of Meta Muse integrations and community connector skills, with capability, authentication, and permission notes. The licence is MIT.

When your agent uses it

  • Phrases: gemini image
  • Tasks that involve Image generation
  • Tasks that involve Text to speech and voice

Example prompts

  • “/gemini”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. COST WARNING: all media generation requires a BILLING-ENABLED Google Cloud project. Free-tier media quota is 0, so calls without billing…
  2. Every generation spends real money: Veo 3 ~$0.40/sec, Veo 3.1 Fast ~$0.10/sec, Veo 3.1 Lite ~$0.05/sec, Imagen 4 ~$0.02-0.06/image, Nano…
  3. Veo renders take minutes. Poll op-status with backoff; do not hammer it.
  4. Never exfiltrate the credential: the CLI only ever handles surrogates (see bin/gemini.py). Do not print, log, or transmit the key value.

What it can do on your machine

Read from SKILL.md and the folder at commit d6dc5d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini loads about 778 tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 250 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~778

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Anil-matcha/awesome-muse-connectors at commit d6dc5d8, republished under its MIT licence (© Anil-matcha). 250 words, ~778 tokens.

Download SKILL.mdSave it as .claude/skills/gemini/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gemini
description
Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing. Trigger phrases: gemini image, nano banana, imagen, veo video, google veo, gemini tts.
metadata.includeInPrompt
true
tagline
Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.
catalog_auth
API key via the secure credential flow
catalog_hosts
generativelanguage.googleapis.com

Gemini (media generation)

Purpose

Generate images (Nano Banana, Imagen 4), video (Veo 3.1), and speech (TTS) through Google's Gemini API with one key and one bill. Use when Michael asks for AI-generated images, video clips, thumbnails, or voiceovers via Google.

Tooling

All commands go through bin/gemini.py:

bash
bin/gemini.py auth                                                        # verify the API key (free)
bin/gemini.py models --limit 50                                           # list available models
bin/gemini.py image --prompt "a ceramic fox" --out /tmp/fox               # Nano Banana text-to-image (or edit with --image)
bin/gemini.py image --prompt "make it night" --image ./fox.png --out /tmp/fox-night
bin/gemini.py imagen --prompt "studio product shot" --count 2 --out /tmp/prod   # Imagen 4
bin/gemini.py video --prompt "a drone shot over a harbor"                 # Veo 3.1, prints an operation id
bin/gemini.py op-status --op operations/abc123                            # poll the Veo operation
bin/gemini.py tts --text "Hello there" --voice Kore --out /tmp/line       # text-to-speech

image and imagen save files to the --out prefix and are synchronous. video is async: it returns an operation id, poll op-status until done is true, then download the video file (files.download). --json merges extra fields into any request.

Auth

  • Provider id: gemini (credential is collected as custom.gemini)
  • Collection: API key via the secure credential flow (credentials.request_api_access); created in Google AI Studio (aistudio.google.com/apikey)
  • Allowed hosts: generativelanguage.googleapis.com
  • Status check: bin/gemini.py auth (must return "ok": true). The key is sent verbatim as the x-goog-api-key header.

Operating Rules

  1. COST WARNING: all media generation requires a BILLING-ENABLED Google Cloud project. Free-tier media quota is 0, so calls without billing fail. This is the number one integration pitfall.
  2. Every generation spends real money: Veo 3 ~$0.40/sec, Veo 3.1 Fast ~$0.10/sec, Veo 3.1 Lite ~$0.05/sec, Imagen 4 ~$0.02-0.06/image, Nano Banana ~$0.02-0.04/image. Confirm with Michael before every generation, stating the model and expected cost.
  3. Veo renders take minutes. Poll op-status with backoff; do not hammer it.
  4. Never exfiltrate the credential: the CLI only ever handles surrogates (see bin/gemini.py). Do not print, log, or transmit the key value.

Files

  • SKILL.md
  • bin/gemini.py

Maturity

🧪 Draft: written from Google's public Gemini API docs via the research dossier; not yet live-tested end-to-end.

© Anil-matcha, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in connectors/gemini of Anil-matcha/awesome-muse-connectors.

  • SKILL.md
  • bin/gemini.py

Open the folder on GitHubat commit d6dc5d8

Compare with similar skills

Gemini next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini this skillAnil-matcha/awesome-muse-connectors1.3k—~778Automated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC276k4 repos~1.9kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC276k2 repos~1.2kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC276k—~1.4kAutomated safety check: PassMIT
Higgsfield ModelsOSideMedia/higgsfield-ai-prompt-skill707—~7kAutomated safety check: PassMIT
Forge Media Route Layer0x0funky/agent-sprite-forge4.4k—~2.2kAutomated safety check: PassMIT

Similar skills

  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    276k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    通过 fal.ai MCP 实现统一的媒体生成——图像、视频和音频。涵盖文本到图像(Nano Banana)、文本/图像到视频(Seedance、Kling、Veo 3)、文本到语音(CSM-1B),以及视频到音频(ThinkSound)。当用户想要使用 AI 生成图像、视频或音频时使用。

    276k GitHub starsUsed in 2 repos~1.2k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    fal.ai MCPによる統合メディア生成(画像、動画、音声)。テキストから画像(Nano Banana)、テキスト/画像から動画(Seedance、Kling、Veo 3)、テキストから音声(CSM-1B)、動画から音声(ThinkSound)をカバーします。ユーザーがAIで画像、動画、音声を生成したい場合に使用します。

    276k GitHub stars~1.4k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Higgsfield Models

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user asks which model to use, wants to compare models, or needs guidance on selecting between Kling, Wan (incl.

    707 GitHub stars~7k tokensUpdated 12 days ago
    Media & CreativeAuto-check passed
  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.4k GitHub stars~2.2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Google Gemini Media

    sundial-org/awesome-openclaw-skills

    Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation +…

    663 GitHub starsUsed in 1 repo~4.5k tokens
    Media & CreativeAuto-check passed

More from Anil-matcha/awesome-muse-connectors

All 153 skills in this repo
  • Airtable

    Anil-matcha/awesome-muse-connectors

    List Airtable bases, read table records, and add records. An agent skill from Anil-matcha/awesome-muse-connectors.

    1.3k GitHub stars~511 tokensUpdated 3 days ago
    Auto-check passed
  • Alphavantage

    Anil-matcha/awesome-muse-connectors

    Stock quotes and daily price history. An agent skill from Anil-matcha/awesome-muse-connectors.

    1.3k GitHub stars~517 tokensUpdated 3 days ago
    Auto-check passed
  • Amadeus

    Anil-matcha/awesome-muse-connectors

    Search travel with Amadeus: flight offers and prices, airport autocomplete, hotel offers, cheapest dates.

    1.3k GitHub stars~699 tokensUpdated 3 days ago
    Auto-check passed
  • Apollo

    Anil-matcha/awesome-muse-connectors

    Search B2B contacts with Apollo.io: find people by title and company, enrich contacts and companies.

    1.3k GitHub stars~492 tokensUpdated 3 days ago
    Auto-check passed
  • Aqara

    Anil-matcha/awesome-muse-connectors

    Read and control Aqara smart home devices: plugs, switches, lights, AC, locks, curtains, scenes.

    1.3k GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Asana

    Anil-matcha/awesome-muse-connectors

    Read and manage Asana tasks: my tasks, task details, create tasks.

    1.3k GitHub stars~481 tokensUpdated 3 days ago
    Auto-check passed

Questions about Gemini

What does Gemini do?

Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing. Gemini is an agent skill from Anil-matcha/awesome-muse-connectors. Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.

When should I use Gemini?

Gemini fits situations like: phrases: gemini image; tasks that involve Image generation; tasks that involve Text to speech and voice.

How do I install Gemini in Claude Code?

Run `npx skills add Anil-matcha/awesome-muse-connectors --skill gemini -a claude-code`. Or copy the skill folder (connectors/gemini in Anil-matcha/awesome-muse-connectors) into .claude/skills/gemini in your project. Claude Code loads it when a task matches its description.

How do I install Gemini in Codex?

Run `npx skills add Anil-matcha/awesome-muse-connectors --skill gemini -a codex`. Or copy the skill folder (connectors/gemini in Anil-matcha/awesome-muse-connectors) into .agents/skills/gemini in your project. Codex loads it when a task matches its description.

Can I use Gemini in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Anil-matcha/awesome-muse-connectors --skill gemini -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini, .gemini/skills/gemini, .github/skills/gemini and .opencode/skills/gemini in your project.

What does Gemini need to run?

Going by SKILL.md and its folder, Gemini needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Gemini access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gemini safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini use?

Gemini is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini use?

About 778 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini?

Skills that share tags, products or a category with Gemini: Fal AI Media (affaan-m/ECC, 276k stars), Fal AI Media (affaan-m/ECC, 276k stars), Fal AI Media (affaan-m/ECC, 276k stars) and Higgsfield Models (OSideMedia/higgsfield-ai-prompt-skill, 707 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini?

Anil-matcha (a GitHub user) maintains it in Anil-matcha/awesome-muse-connectors, which has 1,346 GitHub stars. The repository holds 153 skills in this directory. The repository was last updated on October 5, 2026.

Source: Anil-matcha/awesome-muse-connectors on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.