Official agent skill

Gemini Live API Dev

by google-gemini in google-gemini/gemini-skills

A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Gemini Live API Dev

skills CLI
$ npx skills add google-gemini/gemini-skills --skill gemini-live-api-dev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-gemini/gemini-skills gemini-live-api-dev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-gemini/gemini-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gemini-live-api-dev .claude/skills/gemini-live-api-dev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-live-api-dev
GitHub stars
4.3k
Token cost
~4.6k tokens
SKILL.md length
1,309 words
Files
2 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

  • Works in 9 steps: Use headphones when testing mic audio to… → Enable context window compression for… → Implement session resumption to handle… → …
  • Building real-time
  • SKILL.md covers Overview, Models, SDKs and Partner Integrations, plus 9 more sections
  • Calls pip and npm; reaches ai.google.dev

What it does

Gemini Live API Dev is an agent skill from google-gemini/gemini-skills, published by the product's own GitHub organization. Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), background reasoning (extended thinking), asynchronous function calling, session management, ephemeral tokens, live transcription, and live translation. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/migration.md`).

It sits in Backend & APIs, covering Realtime and WebSockets, Transcription and Authentication. It works with Google Gemini, Python, JavaScript and TypeScript. The repository describes itself as: Skills for the Gemini API, SDK and model/agent interactions. The licence is Apache-2.0.

When your agent uses it

  • Building real-time
  • Bidirectional streaming applications with the Gemini Live API
  • Migrating legacy Live models (2.0/2.5/3.
  • Gemini 3.8 Live

Example prompts

  • “/gemini-live-api-dev”

Requirements

  • Python 3
  • Node.js
  • A credential in YOUR_API_KEY

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Use headphones when testing mic audio to prevent echo/self-interruption
  2. Enable context window compression for sessions longer than 15 minutes
  3. Implement session resumption to handle connection resets gracefully
  4. Use ephemeral tokens for client-side deployments — never expose API keys in browsers
  5. Use send_realtime_input for real-time user input (audio, video, text). Use send_client_content with explicit user/model roles to inject…
  6. Send audioStreamEnd / audio_stream_end (Hybrid VAD) when the mic is paused or user finishes speaking
  7. Clear audio playback queues on interruption signals (interrupted: true)
  8. Process all parts in each server event — events can contain multiple content parts
  9. Monitor interaction_status (IN_PROGRESS vs IDLE) when using gemini-3.8-live-extended-thinking rather than relying on turn_complete alone

What it can do on your machine

Read from SKILL.md and the folder at commit 832c8f9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ai.google.dev

    Also links to:

    • docs.livekit.io
    • docs.pipecat.ai
    • docs.fishjam.io
    • visionagents.ai
    • voximplant.com
    • firebase.google.com
    • colab.research.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Live API Dev loads about 4.6k tokens when it runs, and up to ~6.9k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 1,309 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google-gemini/gemini-skills at commit 832c8f9, republished under its Apache-2.0 licence (© google-gemini). 1,309 words, ~4,645 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-live-api-dev/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gemini-live-api-dev
description
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), background reasoning (extended thinking), asynchronous function calling, session management, ephemeral tokens, live transcription, and live translation. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).

Gemini Live API Development Skill

Overview

The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.

Key capabilities:

  • Bidirectional audio streaming — real-time mic-to-speaker conversations
  • Background reasoning (extended thinking) — multi-step background reasoning with spoken conversational fillers
  • Live streaming transcription — real-time speech-to-text with interim and finalized streams
  • Video streaming — send camera/screen frames alongside audio
  • Text input/output — send and receive text within a live session
  • Audio transcriptions — get text transcripts of both input and output audio
  • Voice Activity Detection (VAD) — automatic server VAD, client-side Hybrid VAD, and manual Push-to-Talk
  • Asynchronous function calling — non-blocking tool execution while audio continues streaming
  • Full-session client content — inject and update conversation turns mid-stream
  • Session management — context compression, session resumption, GoAway signals
  • Ephemeral tokens — secure client-side authentication

[!NOTE] The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.

Models

Current Models (Use These)
  • gemini-3.8-live — Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.
  • gemini-3.8-live-extended-thinking — High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKING required) while streaming continuous spoken conversational fillers; lifecycle managed via interaction_status (IN_PROGRESS vs IDLE).
  • gemini-3.5-transcribe-live — Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.
  • gemini-3.5-live-translate-preview — Real-time speech-to-speech streaming translation across 70+ languages.

[!WARNING] Legacy Models (gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, gemini-live-2.5-flash-preview, gemini-2.0-flash-live-001): Read references/migration.md for breaking protocol changes (behavior: "NON_BLOCKING", thinking_level, interaction_status, send_client_content).

SDKs

  • Python: google-genai >= 2.3.0 — pip install -U google-genai
  • JavaScript/TypeScript: @google/genai >= 2.3.0 — npm install @google/genai

[!WARNING] Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Never use them.

Partner Integrations

To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets:

  • LiveKit — Use the Gemini Live API with LiveKit Agents.
  • Pipecat by Daily — Create a real-time AI chatbot using Gemini Live and Pipecat.
  • Fishjam by Software Mansion — Create live video and audio streaming applications with Fishjam.
  • Vision Agents by Stream — Build real-time voice and video AI applications with Vision Agents.
  • Voximplant — Connect inbound and outbound calls to Live API with Voximplant.
  • Firebase AI SDK — Get started with the Gemini Live API using Firebase AI Logic.

Audio Formats

  • Input: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type: audio/pcm;rate=16000
  • Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.

[!IMPORTANT] Use send_realtime_input / sendRealtimeInput for all real-time streaming user input (audio, video, and text). On Gemini 3.8 models, send_client_content / sendClientContent is supported across the full session lifecycle with explicit roles (user or model) to inject conversation context (turn_complete=true unconditionally interrupts active generation).

[!WARNING] Do not use media in sendRealtimeInput. Use the specific keys: audio for audio data, video for images/video frames, and text for text input.


Quick Start

Authentication
Python
python
from google import genai

client = genai.Client(api_key="YOUR_API_KEY")
JavaScript
js
import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI({ apiKey: 'YOUR_API_KEY' });
Connecting to the Live API
Python
python
from google.genai import types

config = types.LiveConnectConfig(
    response_modalities=[types.Modality.AUDIO],
    system_instruction=types.Content(
        parts=[types.Part(text="You are a helpful assistant.")]
    )
)

async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
    pass  # Session is active
JavaScript
js
const session = await ai.live.connect({
  model: 'gemini-3.8-live',
  config: {
    responseModalities: ['audio'],
    systemInstruction: { parts: [{ text: 'You are a helpful assistant.' }] }
  },
  callbacks: {
    onopen: () => console.log('Connected'),
    onmessage: (response) => console.log('Message:', response),
    onerror: (error) => console.error('Error:', error),
    onclose: () => console.log('Closed')
  }
});
Sending Text
Python
python
await session.send_realtime_input(text="Hello, how are you?")
JavaScript
js
session.sendRealtimeInput({ text: 'Hello, how are you?' });
Sending Audio
Python
python
await session.send_realtime_input(
    audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")
)
JavaScript
js
session.sendRealtimeInput({
  audio: { data: chunk.toString('base64'), mimeType: 'audio/pcm;rate=16000' }
});
Sending Video
Python
python
# frame: raw JPEG-encoded bytes
await session.send_realtime_input(
    video=types.Blob(data=frame, mime_type="image/jpeg")
)
JavaScript
js
session.sendRealtimeInput({
  video: { data: frame.toString('base64'), mimeType: 'image/jpeg' }
});
Receiving Audio and Text

[!IMPORTANT] A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content.

Python
python
async for response in session.receive():
    content = response.server_content
    if content:
        # Audio — process ALL parts in each event
        if content.model_turn:
            for part in content.model_turn.parts:
                if part.inline_data:
                    audio_data = part.inline_data.data
        # Transcription
        if content.input_transcription:
            print(f"User: {content.input_transcription.text}")
        if content.output_transcription:
            print(f"Gemini: {content.output_transcription.text}")
        # Interruption
        if content.interrupted is True:
            pass  # Stop playback, clear audio queue
JavaScript
js
// Inside the onmessage callback
const content = response.serverContent;
if (content?.modelTurn?.parts) {
  for (const part of content.modelTurn.parts) {
    if (part.inlineData) {
      const audioData = part.inlineData.data; // Base64 encoded
    }
  }
}
if (content?.inputTranscription) console.log('User:', content.inputTranscription.text);
if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text);
if (content?.interrupted) { /* Stop playback, clear audio queue */ }

Background Reasoning (Extended Thinking)

Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.

Key requirements:

  • Thinking config: Set thinking_config=types.ThinkingConfig(thinking_level="low") ("minimal" | "low" | "medium" | "high").
  • Non-blocking tools: All function declarations must set behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error.
  • Lifecycle tracking (interaction_status): Do not rely on turn_complete=True alone to detect turn completion. Monitor message.interaction_status (Python) / message.interactionStatus (JS):
    • "IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses.
    • "IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.

See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.


Live Translation (Gemini Live Translate)

The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.

Model
  • gemini-3.5-live-translate-preview — The recommended translation model for all Live Translate use cases.
Configuration (TranslationConfig)

To enable translation, specify a TranslationConfig object inside your live session setup:

  • Python SDK: Configure the connection using translation_config on LiveConnectConfig:
    python
    config = types.LiveConnectConfig(
        response_modalities=[types.Modality.AUDIO],
        translation_config=types.TranslationConfig(
            target_language_code="es",  # Target language code (e.g. es, fr, pl)
            echo_target_language=True,
        ),
        input_audio_transcription=types.AudioTranscriptionConfig(),
        output_audio_transcription=types.AudioTranscriptionConfig(),
    )
  • Raw WebSockets: Place translationConfig inside generationConfig:
    json
    {
      "setup": {
        "model": "models/gemini-3.5-live-translate-preview",
        "generationConfig": {
          "responseModalities": ["AUDIO"],
          "translationConfig": {
            "targetLanguageCode": "es",
            "echoTargetLanguage": true
          }
        }
      }
    }

Live Streaming Transcription (Gemini Live Transcribe)

The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.

Show full SKILL.md (532 more words)Show less
Model
  • gemini-3.5-transcribe-live
Modes
  • smart: cleans up filler words, resolves inline self-corrections, and structures formatting.
  • verbatim (default): exact word-for-word transcript.
Python
python
config = types.LiveConnectConfig(
    response_modalities=["TEXT"],
    input_audio_transcription=types.AudioTranscriptionConfig(),
)

async with client.aio.live.connect(model="gemini-3.5-transcribe-live", config=config) as session:
    # Stream audio
    await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000"))
    # Hybrid VAD: notify turn end on client-detected silence for zero latency
    await session.send_realtime_input(audio_stream_end=True)
JavaScript
javascript
const session = await ai.live.connect({
  model: 'gemini-3.5-transcribe-live',
  config: {
    responseModalities: ['text'],
    inputAudioTranscription: { mode: 'smart' }
  },
  callbacks: {
    onmessage: (msg) => {
      if (msg.serverContent?.interimInputTranscription) {
        console.log('Interim:', msg.serverContent.interimInputTranscription.text);
      }
      if (msg.serverContent?.inputTranscription) {
        console.log('Final:', msg.serverContent.inputTranscription.text);
      }
    }
  }
});

session.sendRealtimeInput({ audio: { data: chunkBase64, mimeType: 'audio/pcm;rate=16000' } });
session.sendRealtimeInput({ audioStreamEnd: true }); // Hybrid VAD
Raw WebSockets
json
{
  "setup": {
    "model": "models/gemini-3.5-transcribe-live",
    "generationConfig": {
      "responseModalities": ["TEXT"],
      "speechConfig": {
        "voiceConfig": {}
      }
    },
    "inputAudioTranscription": {
      "mode": "smart"
    }
  }
}

Limitations

  • Response modality — Only TEXT or AUDIO per session, not both. Native audio models output audio (response_modalities=["AUDIO"]); enable output_audio_transcription if you need text transcripts.
  • Audio-only session — 15 min without compression
  • Audio+video session — 2 min without compression
  • Connection lifetime — ~10 min (use session resumption)
  • Context window — 128k input tokens / 64k output tokens
  • Code execution / URL context — Not supported

Upgrading & Migration

For step-by-step migration checklists and protocol deltas when upgrading from gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, or gemini-2.0-flash-live-001 to Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, read references/migration.md.

Best Practices

  1. Use headphones when testing mic audio to prevent echo/self-interruption
  2. Enable context window compression for sessions longer than 15 minutes
  3. Implement session resumption to handle connection resets gracefully
  4. Use ephemeral tokens for client-side deployments — never expose API keys in browsers
  5. Use send_realtime_input for real-time user input (audio, video, text). Use send_client_content with explicit user/model roles to inject context turns mid-stream
  6. Send audioStreamEnd / audio_stream_end (Hybrid VAD) when the mic is paused or user finishes speaking
  7. Clear audio playback queues on interruption signals (interrupted: true)
  8. Process all parts in each server event — events can contain multiple content parts
  9. Monitor interaction_status (IN_PROGRESS vs IDLE) when using gemini-3.8-live-extended-thinking rather than relying on turn_complete alone

Documentation Lookup

When MCP is Installed (Preferred)

If the search_docs tool (from the Google MCP server) is available, use it as your only documentation source:

  1. Call search_docs with your query
  2. Read the returned documentation
  3. Trust MCP results as source of truth for API details — they are always up-to-date.

[!IMPORTANT] When MCP tools are present, never fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching.

When MCP is NOT Installed (Fallback Only)

If no MCP documentation tools are available, fetch from the official docs index:

llms.txt URL: https://ai.google.dev/gemini-api/docs/llms.txt

This index contains links to all documentation pages in .md.txt format. Use web fetch tools to:

  1. Fetch llms.txt to discover available documentation pages
  2. Fetch specific pages (e.g., https://ai.google.dev/gemini-api/docs/live-session.md.txt)
Key Documentation Pages

[!IMPORTANT] Those are not all the documentation pages. Use the llms.txt index to discover available documentation pages

Supported Languages

The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages.

© google-gemini, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/gemini-live-api-dev of google-gemini/gemini-skills.

  • SKILL.md
  • references/migration.md

Open the folder on GitHubat commit 832c8f9

Compare with similar skills

Gemini Live API Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Live API Dev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Live API Dev this skillgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Gemini Live API DevJetBrains/skills363—~2.6kAutomated safety check: PassNone
Gemini Interactions APIAyuilos/Miffan182—~4.6kAutomated safety check: PassAGPL-3.0
Gemini API DevAyuilos/Miffan1821 repos~1.4kAutomated safety check: PassAGPL-3.0
Azure AI Voicelive Pymicrosoft/skills3.1k6 repos~2.9kAutomated safety check: PassMIT
Gemini API DevJetBrains/skills363—~1.6kAutomated safety check: PassNone

Similar skills

  • Gemini Live API Dev

    JetBrains/skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API.

    363 GitHub stars~2.6k tokensUpdated 3 mo ago
    Backend & APIsAuto-check passed
  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    182 GitHub stars~4.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Gemini API Dev

    Ayuilos/Miffan

    A skill your agent uses when building applications with Gemini API hosted models, including Gemini and Gemma 4, working with multimodal content (text, images, audio, video), implementing function…

    182 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Azure AI Voicelive Py

    microsoft/skills

    Official

    Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

    3.1k GitHub starsUsed in 6 repos~2.9k tokens
    Backend & APIsAuto-check passed
  • Gemini API Dev

    JetBrains/skills

    Official

    A skill your agent uses when building applications with Gemini models, Gemini API, working with multimodal content (text, images, audio, video), implementing function calling, using structured…

    363 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Gemini Interactions API

    JetBrains/skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, streaming responses, background research tasks…

    363 GitHub stars~2.5k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from google-gemini/gemini-skills

  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Gemini Omni Flash API

    google-gemini/gemini-skills

    Official

    A skill your agent uses for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni…

    4.3k GitHub stars~6.4k tokensUpdated today
    Auto-check passed

Questions about Gemini Live API Dev

What does Gemini Live API Dev do?

A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Gemini Live API Dev is an agent skill from google-gemini/gemini-skills, published by the product's own GitHub organization.8 Live.

When should I use Gemini Live API Dev?

Gemini Live API Dev fits situations like: building real-time; bidirectional streaming applications with the Gemini Live API; migrating legacy Live models (2.0/2.5/3; gemini 3.8 Live.

How do I install Gemini Live API Dev in Claude Code?

Run `npx skills add google-gemini/gemini-skills --skill gemini-live-api-dev -a claude-code`. Or copy the skill folder (skills/gemini-live-api-dev in google-gemini/gemini-skills) into .claude/skills/gemini-live-api-dev in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Live API Dev in Codex?

Run `npx skills add google-gemini/gemini-skills --skill gemini-live-api-dev -a codex`. Or copy the skill folder (skills/gemini-live-api-dev in google-gemini/gemini-skills) into .agents/skills/gemini-live-api-dev in your project. Codex loads it when a task matches its description.

Can I use Gemini Live API Dev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-gemini/gemini-skills --skill gemini-live-api-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-live-api-dev, .gemini/skills/gemini-live-api-dev, .github/skills/gemini-live-api-dev and .opencode/skills/gemini-live-api-dev in your project.

What does Gemini Live API Dev need to run?

Going by SKILL.md and its folder, Gemini Live API Dev needs the command-line tools its instructions call (pip and npm). Our summary lists: Python 3; Node.js; A credential in YOUR_API_KEY.

Does Gemini Live API Dev access the network?

SKILL.md names 8 domains. In commands or code: ai.google.dev; the agent is likely to contact it when it follows the instructions. As links in the text: docs.livekit.io, docs.pipecat.ai, docs.fishjam.io, visionagents.ai, voximplant.com, firebase.google.com and colab.research.google.com. This is read from the text; nothing was executed.

Is Gemini Live API Dev safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini Live API Dev use?

Gemini Live API Dev is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Live API Dev use?

About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Live API Dev?

Skills that share tags, products or a category with Gemini Live API Dev: Gemini Live API Dev (JetBrains/skills, 363 stars), Gemini Interactions API (Ayuilos/Miffan, 182 stars), Gemini API Dev (Ayuilos/Miffan, 182 stars) and Azure AI Voicelive Py (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Live API Dev?

google-gemini (a GitHub organization, an official publisher) maintains it in google-gemini/gemini-skills, which has 4,252 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 6, 2026.

Source: google-gemini/gemini-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.