Agent skill

Google Gemini Media

by sundial-org in sundial-org/awesome-openclaw-skills

Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation +…

MITAuto-check passedMedia & Creative

Install Google Gemini Media

skills CLI
$ npx skills add sundial-org/awesome-openclaw-skills --skill google-gemini-media -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sundial-org/awesome-openclaw-skills google-gemini-media --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sundial-org/awesome-openclaw-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/google-gemini-media .claude/skills/google-gemini-media && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
google-gemini-media
GitHub stars
663
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,512 words
Files
1
Skills in repo
383
Repo updated
First seen
Licence
MIT

At a glance

Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation +…

  • Works in 12 steps: Goals and scope → Quick routing (decide which capability… → Unified engineering constraints and I/O… → …
  • Tasks that involve Image generation
  • SKILL.md covers 1. Goals and scope, 2. Quick routing (decide which…, 3. Unified engineering… and 4. Model selection matrix…, plus 4 more sections
  • Calls curl, npm and jq; reaches generativelanguage.googleapis.com; needs GEMINI_API_KEY

What it does

Google Gemini Media is an agent skill from sundial-org/awesome-openclaw-skills. Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image generation, AI video generation and Text to speech and voice. It works with Google Gemini. The repository describes itself as: Top OpenClaw skills, with the most popular and useful ones. The licence is MIT.

When your agent uses it

  • Tasks that involve Image generation
  • Tasks that involve AI video generation
  • Tasks that involve Text to speech and voice

Example prompts

  • “generation + understanding”
  • “/google-gemini-media”

Requirements

  • Node.js
  • A credential in GEMINI_API_KEY

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Goals and scope
  2. Quick routing (decide which capability to use)
  3. Unified engineering constraints and I/O spec (must read)
  4. Model selection matrix (choose by scenario)
  5. Image generation (Nano Banana)
  6. Image understanding (Image Understanding)
  7. Video generation (Veo 3.1)
  8. Video understanding (Video Understanding)
  9. Speech generation (Text-to-Speech, TTS)
  10. Audio understanding (Audio Understanding)
  11. End-to-end examples (composition)
  12. Compliance and risk (must follow)

What it can do on your machine

Read from SKILL.md and the folder at commit b80cde2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • npm
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • generativelanguage.googleapis.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Google Gemini Media loads about 4.5k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,512 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sundial-org/awesome-openclaw-skills at commit b80cde2, republished under its MIT licence (© sundial-org). 1,512 words, ~4,528 tokens.

Download SKILL.mdSave it as .claude/skills/google-gemini-media/SKILL.md (or your agent's skills folder).
name
google-gemini-media
description
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
license
MIT

Gemini Multimodal Media (Image/Video/Speech) Skill

1. Goals and scope

This Skill consolidates six Gemini API capabilities into reusable workflows and implementation templates:

  • Image generation (Nano Banana: text-to-image, image editing, multi-turn iteration)
  • Image understanding (caption/VQA/classification/comparison, multi-image prompts; supports inline and Files API)
  • Video generation (Veo 3.1: text-to-video, aspect ratio/resolution control, reference-image guidance, first/last frames, video extension, native audio)
  • Video understanding (upload/inline/YouTube URL; summaries, Q&A, timestamped evidence)
  • Speech generation (Gemini native TTS: single-speaker and multi-speaker; controllable style/accent/pace/tone)
  • Audio understanding (upload/inline; description, transcription, time-range transcription, token counting)

Convention: This Skill follows the official Google Gen AI SDK (Node.js/REST) as the main line; currently only Node.js/REST examples are provided. If your project already wraps other languages or frameworks, map this Skill's request structure, model selection, and I/O spec to your wrapper layer.


2. Quick routing (decide which capability to use)

  1. Do you need to produce images?
  • Need to generate images from scratch or edit based on an image -> use Nano Banana image generation (see Section 5)
  1. Do you need to understand images?
  • Need recognition, description, Q&A, comparison, or info extraction -> use Image understanding (see Section 6)
  1. Do you need to produce video?
  • Need to generate an 8-second video (optionally with native audio) -> use Veo 3.1 video generation (see Section 7)
  1. Do you need to understand video?
  • Need summaries/Q&A/segment extraction with timestamps -> use Video understanding (see Section 8)
  1. Do you need to read text aloud?
  • Need controllable narration, podcast/audiobook style, etc. -> use Speech generation (TTS) (see Section 9)
  1. Do you need to understand audio?
  • Need audio descriptions, transcription, time-range transcription, token counting -> use Audio understanding (see Section 10)

3. Unified engineering constraints and I/O spec (must read)

3.0 Prerequisites (dependencies and tools)
  • Node.js 18+ (match your project version)
  • Install SDK (example):
bash
npm install @google/genai
  • REST examples only need curl; if you need to parse image Base64, install jq (optional).
3.1 Authentication and environment variables
  • Put your API key in GEMINI_API_KEY
  • REST requests use x-goog-api-key: $GEMINI_API_KEY
3.2 Two file input modes: Inline vs Files API

Inline (embedded bytes/Base64)

  • Pros: shorter call chain, good for small files.
  • Key constraint: total request size (text prompt + system instructions + embedded bytes) typically has a ~20MB ceiling.

Files API (upload then reference)

  • Pros: good for large files, reusing the same file, or multi-turn conversations.
  • Typical flow:
    1. files.upload(...) (SDK) or POST /upload/v1beta/files (REST resumable)
    2. Use file_data / file_uri in generateContent

Engineering suggestion: implement ensure_file_uri() so that when a file exceeds a threshold (for example 10-15MB warning) or is reused, you automatically route through the Files API.

3.3 Unified handling of binary media outputs
  • Images: usually returned as inline_data (Base64) in response parts; in the SDK use part.as_image() or decode Base64 and save as PNG/JPG.
  • Speech (TTS): usually returns PCM bytes (Base64); save as .pcm or wrap into .wav (commonly 24kHz, 16-bit, mono).
  • Video (Veo): long-running async task; poll the operation; download the file (or use the returned URI).

4. Model selection matrix (choose by scenario)

Important: model names, versions, limits, and quotas can change over time. Verify against official docs before use. Last updated: 2026-01-22.

4.1 Image generation (Nano Banana)
  • gemini-2.5-flash-image: optimized for speed/throughput; good for frequent, low-latency generation/editing.
  • gemini-3-pro-image-preview: stronger instruction following and high-fidelity text rendering; better for professional assets and complex edits.
4.2 General image/video/audio understanding
  • Docs use gemini-3-flash-preview for image, video, and audio understanding (choose stronger models as needed for quality/cost).
4.3 Video generation (Veo)
  • Example model: veo-3.1-generate-preview (generates 8-second video and can natively generate audio).
4.4 Speech generation (TTS)
  • Example model: gemini-2.5-flash-preview-tts (native TTS, currently in preview).

5. Image generation (Nano Banana)

5.1 Text-to-Image

SDK (Node.js) minimal template

js
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const response = await ai.models.generateContent({
  model: "gemini-2.5-flash-image",
  contents:
    "Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme",
});

const parts = response.candidates?.[0]?.content?.parts ?? [];
for (const part of parts) {
  if (part.text) console.log(part.text);
  if (part.inlineData?.data) {
    fs.writeFileSync("out.png", Buffer.from(part.inlineData.data, "base64"));
  }
}

REST (with imageConfig) minimal template

bash
curl -s -X POST   "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:generateContent"   -H "x-goog-api-key: $GEMINI_API_KEY"   -H "Content-Type: application/json"   -d '{
    "contents":[{"parts":[{"text":"Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme"}]}],
    "generationConfig": {"imageConfig": {"aspectRatio":"16:9"}}
  }'

REST image parsing (Base64 decode)

bash
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"contents":[{"parts":[{"text":"A minimal studio product shot of a nano banana"}]}]}' \
  | jq -r '.candidates[0].content.parts[] | select(.inline_data) | .inline_data.data' \
  | base64 --decode > out.png

# macOS can use: base64 -D > out.png
5.2 Text-and-Image-to-Image

Use case: given an image, add/remove/modify elements, change style, color grading, etc.

SDK (Node.js) minimal template

js
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const prompt =
  "Add a nano banana on the table, keep lighting consistent, cinematic tone.";
const imageBase64 = fs.readFileSync("input.png").toString("base64");

const response = await ai.models.generateContent({
  model: "gemini-2.5-flash-image",
  contents: [
    { text: prompt },
    { inlineData: { mimeType: "image/png", data: imageBase64 } },
  ],
});

const parts = response.candidates?.[0]?.content?.parts ?? [];
for (const part of parts) {
  if (part.inlineData?.data) {
    fs.writeFileSync("edited.png", Buffer.from(part.inlineData.data, "base64"));
  }
}
5.3 Multi-turn image iteration (Multi-turn editing)

Best practice: use chat for continuous iteration (for example: generate first, then "only edit a specific region/element", then "make variants in the same style").
To output mixed "text + image" results, set response_modalities to ["TEXT", "IMAGE"].

5.4 ImageConfig

You can set in generationConfig.imageConfig or the SDK config:

  • aspectRatio: e.g. 16:9, 1:1.
  • imageSize: e.g. 2K, 4K (higher resolution is usually slower/more expensive and model support can vary).

6. Image understanding (Image Understanding)

6.1 Two ways to provide input images
  • Inline image data: suitable for small files (total request size < 20MB).
  • Files API upload: better for large files or reuse across multiple requests.
6.2 Inline images (Node.js) minimal template
js
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const imageBase64 = fs.readFileSync("image.jpg").toString("base64");

const response = await ai.models.generateContent({
  model: "gemini-3-flash-preview",
  contents: [
    { inlineData: { mimeType: "image/jpeg", data: imageBase64 } },
    { text: "Caption this image, and list any visible brands." },
  ],
});

console.log(response.text);
6.3 Upload and reference with Files API (Node.js) minimal template
js
import { GoogleGenAI, createPartFromUri, createUserContent } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const uploaded = await ai.files.upload({ file: "image.jpg" });

const response = await ai.models.generateContent({
  model: "gemini-3-flash-preview",
  contents: createUserContent([
    createPartFromUri(uploaded.uri, uploaded.mimeType),
    "Caption this image.",
  ]),
});

console.log(response.text);
6.4 Multi-image prompts

Append multiple images as multiple Part entries in the same contents; you can mix uploaded references and inline bytes.


7. Video generation (Veo 3.1)

7.1 Core features (must know)
  • Generates 8-second high-fidelity video, optionally 720p / 1080p / 4k, and supports native audio generation (dialogue, ambience, SFX).
  • Supports:
    • Aspect ratio (16:9 / 9:16)
    • Video extension (extend a generated video; typically limited to 720p)
    • First/last frame control (frame-specific)
    • Up to 3 reference images (image-based direction)
7.2 SDK (Node.js) minimal template: async polling + download
js
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const prompt =
  "A cinematic shot of a cat astronaut walking on the moon. Include subtle wind ambience.";
let operation = await ai.models.generateVideos({
  model: "veo-3.1-generate-preview",
  prompt,
  config: { resolution: "1080p" },
});

while (!operation.done) {
  await new Promise((resolve) => setTimeout(resolve, 10_000));
  operation = await ai.operations.getVideosOperation({ operation });
}

const video = operation.response?.generatedVideos?.[0]?.video;
if (!video) throw new Error("No video returned");
await ai.files.download({ file: video, downloadPath: "out.mp4" });
7.3 REST minimal template: predictLongRunning + poll + download

Key point: Veo REST uses :predictLongRunning to return an operation name, then poll GET /v1beta/{operation_name}; once done, download from the video URI in the response.

7.4 Common controls (recommend a unified wrapper)
  • aspectRatio: "16:9" or "9:16"
  • resolution: "720p" | "1080p" | "4k" (higher resolutions are usually slower/more expensive)
  • When writing prompts: put dialogue in quotes; explicitly call out SFX and ambience; use cinematography language (camera position, movement, composition, lens effects, mood).
  • Negative constraints: if the API supports a negative prompt field, use it; otherwise list elements you do not want to see.
Show full SKILL.md (589 more words)Show less
7.5 Important limits (engineering fallback needed)
  • Latency can vary from seconds to minutes; implement timeouts and retries.
  • Generated videos are only retained on the server for a limited time (download promptly).
  • Outputs include a SynthID watermark.

Polling fallback (with timeout/backoff) pseudocode

js
const deadline = Date.now() + 300_000; // 5 min
let sleepMs = 2000;
while (!operation.done && Date.now() < deadline) {
  await new Promise((resolve) => setTimeout(resolve, sleepMs));
  sleepMs = Math.min(Math.floor(sleepMs * 1.5), 15_000);
  operation = await ai.operations.getVideosOperation({ operation });
}
if (!operation.done) throw new Error("video generation timed out");

8. Video understanding (Video Understanding)

8.1 Video input options
  • Files API upload: recommended when file > 100MB, video length > ~1 minute, or you need reuse.
  • Inline video data: for smaller files.
  • Direct YouTube URL: can analyze public videos.
8.2 Files API (Node.js) minimal template
js
import { GoogleGenAI, createPartFromUri, createUserContent } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const uploaded = await ai.files.upload({ file: "sample.mp4" });

const response = await ai.models.generateContent({
  model: "gemini-3-flash-preview",
  contents: createUserContent([
    createPartFromUri(uploaded.uri, uploaded.mimeType),
    "Summarize this video. Provide timestamps for key events.",
  ]),
});

console.log(response.text);
8.3 Timestamp prompting strategy
  • Ask for segmented bullets with "(mm:ss)" timestamps.
  • Require "evidence with specific time ranges" and include downstream structured extraction (JSON) in the same prompt if needed.

9. Speech generation (Text-to-Speech, TTS)

9.1 Positioning
  • Native TTS: for "precise reading + controllable style" (podcasts, audiobooks, ad voiceover, etc.).
  • Distinguish from the Live API: Live API is more interactive and non-structured audio/multimodal conversation; TTS is focused on controlled narration.
9.2 Single-speaker TTS (Node.js) minimal template
js
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const response = await ai.models.generateContent({
  model: "gemini-2.5-flash-preview-tts",
  contents: [{ parts: [{ text: "Say cheerfully: Have a wonderful day!" }] }],
  config: {
    responseModalities: ["AUDIO"],
    speechConfig: {
      voiceConfig: {
        prebuiltVoiceConfig: { voiceName: "Kore" },
      },
    },
  },
});

const data =
  response.candidates?.[0]?.content?.parts?.[0]?.inlineData?.data ?? "";
if (!data) throw new Error("No audio returned");
fs.writeFileSync("out.pcm", Buffer.from(data, "base64"));
9.3 Multi-speaker TTS (max 2 speakers)

Requirements:

  • Use multiSpeakerVoiceConfig
  • Each speaker name must match the dialogue labels in the prompt (e.g., Joe/Jane).
9.4 Voice options and language
  • voice_name supports 30 prebuilt voices (for example Zephyr, Puck, Charon, Kore, etc.).
  • The model can auto-detect input language and supports 24 languages (see docs for the list).

Provide controllable directions for style, pace, accent, etc., but avoid over-constraining.


10. Audio understanding (Audio Understanding)

10.1 Typical tasks
  • Describe audio content (including non-speech like birds, alarms, etc.)
  • Generate transcripts
  • Transcribe specific time ranges
  • Count tokens (for cost estimates/segmentation)
10.2 Files API (Node.js) minimal template
js
import { GoogleGenAI, createPartFromUri, createUserContent } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const uploaded = await ai.files.upload({ file: "sample.mp3" });

const response = await ai.models.generateContent({
  model: "gemini-3-flash-preview",
  contents: createUserContent([
    "Describe this audio clip.",
    createPartFromUri(uploaded.uri, uploaded.mimeType),
  ]),
});

console.log(response.text);
10.3 Key limits and engineering tips
  • Supports common formats: WAV/MP3/AIFF/AAC/OGG/FLAC.
  • Audio tokenization: about 32 tokens/second (about 1920 tokens per minute; values may change).
  • Total audio length per prompt is capped at 9.5 hours; multi-channel audio is downmixed; audio is resampled (see docs for exact parameters).
  • If total request size exceeds 20MB, you must use the Files API.

11. End-to-end examples (composition)

Example A: Image generation -> validation via understanding
  1. Generate product images with Nano Banana (require negative space, consistent lighting).
  2. Use image understanding for self-check: verify text clarity, brand spelling, and unsafe elements.
  3. If not satisfied, feed the generated image into text+image editing and iterate.
Example B: Video generation -> video understanding -> narration script
  1. Generate an 8-second shot with Veo (include dialogue or SFX).
  2. Download and save (respect retention window).
  3. Upload video to video understanding to produce a storyboard + timestamps + narration copy (then feed to TTS).
Example C: Audio understanding -> time-range transcription -> TTS redub
  1. Upload meeting audio and transcribe full content.
  2. Transcribe or summarize specific time ranges.
  3. Use TTS to generate a "broadcast" version of the summary.

12. Compliance and risk (must follow)

  • Ensure you have the necessary rights to upload images/video/audio; do not generate infringing, deceptive, harassing, or harmful content.
  • Generated images and videos include SynthID watermarking; videos may also have regional/person-based generation constraints.
  • Production systems must implement timeouts, retries, failure fallbacks, and human review/post-processing for generated content.

13. Quick reference (Checklist)

  • Pick the right model: image generation (Flash Image / Pro Image Preview), video generation (Veo 3.1), TTS (Gemini 2.5 TTS), understanding (Gemini Flash/Pro).
  • Pick the right input mode: inline for small files; Files API for large/reuse.
  • Parse binary outputs correctly: image/audio via inline_data decode; video via operation polling + download.
  • For video generation: set aspectRatio / resolution, and download promptly (avoid expiration).
  • For TTS: set response_modalities=["AUDIO"]; max 2 speakers; speaker names must match prompt.
  • For audio understanding: countTokens when needed; segment long audio or use Files API.

© sundial-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/google-gemini-media of sundial-org/awesome-openclaw-skills.

Open the folder on GitHubat commit b80cde2

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sundial-org/awesome-openclaw-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Google Gemini Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Google Gemini Media compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Google Gemini Media this skillsundial-org/awesome-openclaw-skills6631 repos~4.5kAutomated safety check: PassMIT
Gemini Omnicalesthio/OpenMontage66k—~2.1kAutomated safety check: NotesAGPL-3.0
Gemini Interactions APIAyuilos/Miffan225—~4.6kAutomated safety check: PassAGPL-3.0
Fal AI Mediaaffaan-m/ECC277k4 repos~1.9kAutomated safety check: PassMIT
GeminiAnil-matcha/awesome-muse-connectors1.3k—~778Automated safety check: PassMIT
Gemini API Devgoogle-gemini/gemini-skills4.3k—~5.1kAutomated safety check: PassApache-2.0

Similar skills

  • Gemini Omni

    calesthio/OpenMontage

    Generate and conversationally edit short videos with Google Gemini Omni Flash (gemini-omni-flash-preview).

    66k GitHub stars~2.1k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    225 GitHub stars~4.6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    277k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Gemini

    Anil-matcha/awesome-muse-connectors

    Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.

    1.3k GitHub stars~778 tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Wedding Video Guided Wizard

    aaronyi97/wedding-video-guided-wizard

    Guide a creator through a real couple's custom wedding video, from a shareable story intake card and Kimi writing pack through narration, external GPT image prompts, image-to-video packs, music and…

    310 GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from sundial-org/awesome-openclaw-skills

All 383 skills in this repo
  • UI UX Pro Max

    sundial-org/awesome-openclaw-skills

    UI/UX design intelligence and implementation guidance for building polished interfaces.

    663 GitHub starsUsed in 2 repos~657 tokens
    Auto-check passed
  • Web Deploy GitHub

    sundial-org/awesome-openclaw-skills

    Create and deploy single-page static websites to GitHub Pages with autonomous workflow.

    663 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Clawd Modifier

    sundial-org/awesome-openclaw-skills

    Modify Clawd, the Claude Code mascot. An agent skill from sundial-org/awesome-openclaw-skills.

    663 GitHub stars~625 tokensUpdated 7 mo ago
    Auto-check passed
  • Figma

    sundial-org/awesome-openclaw-skills

    Professional Figma design analysis and asset export. An agent skill from sundial-org/awesome-openclaw-skills.

    663 GitHub stars~1.7k tokensUpdated 7 mo ago
    Auto-check: notes
  • Habit Flow

    sundial-org/awesome-openclaw-skills

    AI-powered atomic habit tracker with natural language logging, streak tracking, smart reminders, and coaching.

    663 GitHub stars~3.1k tokensUpdated 7 mo ago
    Auto-check passed
  • Edge Tts

    sundial-org/awesome-openclaw-skills

    Text-to-speech conversion using node-edge-tts npm package for generating audio from text.

    663 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed

Works with

Questions about Google Gemini Media

What does Google Gemini Media do?

Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation +…. Google Gemini Media is an agent skill from sundial-org/awesome-openclaw-skills. Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".

When should I use Google Gemini Media?

Google Gemini Media fits situations like: tasks that involve Image generation; tasks that involve AI video generation; tasks that involve Text to speech and voice.

How do I install Google Gemini Media in Claude Code?

Run `npx skills add sundial-org/awesome-openclaw-skills --skill google-gemini-media -a claude-code`. Or copy the skill folder (skills/google-gemini-media in sundial-org/awesome-openclaw-skills) into .claude/skills/google-gemini-media in your project. Claude Code loads it when a task matches its description.

How do I install Google Gemini Media in Codex?

Run `npx skills add sundial-org/awesome-openclaw-skills --skill google-gemini-media -a codex`. Or copy the skill folder (skills/google-gemini-media in sundial-org/awesome-openclaw-skills) into .agents/skills/google-gemini-media in your project. Codex loads it when a task matches its description.

Can I use Google Gemini Media in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sundial-org/awesome-openclaw-skills --skill google-gemini-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/google-gemini-media, .gemini/skills/google-gemini-media, .github/skills/google-gemini-media and .opencode/skills/google-gemini-media in your project.

What does Google Gemini Media need to run?

Going by SKILL.md and its folder, Google Gemini Media needs the command-line tools its instructions call (curl, npm and jq) and credentials named GEMINI_API_KEY. Our summary lists: Node.js; A credential in GEMINI_API_KEY.

Does Google Gemini Media access the network?

SKILL.md names 1 domain. In commands or code: generativelanguage.googleapis.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Google Gemini Media safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Google Gemini Media use?

Google Gemini Media is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Google Gemini Media use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Google Gemini Media?

Skills that share tags, products or a category with Google Gemini Media: Gemini Omni (calesthio/OpenMontage, 66k stars), Gemini Interactions API (Ayuilos/Miffan, 225 stars), Fal AI Media (affaan-m/ECC, 277k stars) and Gemini (Anil-matcha/awesome-muse-connectors, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Google Gemini Media?

sundial-org (a GitHub organization) maintains it in sundial-org/awesome-openclaw-skills, which has 663 GitHub stars. The repository holds 383 skills in this directory. The repository was last updated on March 7, 2026.

Source: sundial-org/awesome-openclaw-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.