Official agent skill

Gemini Live API

by google in google/skills

Generates a Gemini LiveAPI client service class in the user's chosen programming language.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Gemini Live API

skills CLI
$ npx skills add google/skills --skill gemini-live-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gemini-live-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gemini-live-api .claude/skills/gemini-live-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-live-api
GitHub stars
21k
Token cost
~2.5k tokens
SKILL.md length
1,124 words
Files
4 (incl. references)
Skills in repo
145
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generates a Gemini LiveAPI client service class in the user's chosen programming language.

  • Works in 7 steps: Copy the reference files → Reconcile with the public documentation → Implement the client class → …
  • The user wants to build
  • SKILL.md covers Prerequisites, Reference Files, Steps and Validation Checklist
  • Calls pip

What it does

Gemini Live API is an agent skill from google/skills, published by the product's own GitHub organization. Generates a Gemini LiveAPI client service class in the user's chosen programming language. Use when the user wants to build, scaffold, or integrate a client that connects to the Gemini Enterprise LiveAPI websocket endpoint, handles session setup/resumption, bearer token refresh, and sending/receiving ClientMessage/ServerMessage protos. Don't use for general (non-live, non-bidirectional) Gemini API usage such as one-shot generateContent, embeddings, image/video generation, or fine-tuning — use the gemini-api skill…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/client_server_messages.md` and `references/session_manager.md`).

It sits in AI & LLM Engineering, covering LLM API integration, Realtime and WebSockets and Embeddings. It works with Google Gemini. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • The user wants to build
  • Integrate a client that connects to the Gemini Enterprise LiveAPI websocket endpoint
  • Handles session setup/resumption
  • Bearer token refresh

Example prompts

  • “Use the gemini-live-api skill to generate a Gemini LiveAPI client service class in the user's chosen programming language”
  • “/gemini-live-api”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Copy the reference files
  2. Reconcile with the public documentation
  3. Implement the client class
  4. Write a test file
  5. Generate how_to_run.md
  6. Generate a demo frontend + backend service
  7. Generate how_to_test_with_ui.md

What it can do on your machine

Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Live API loads about 2.5k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 1,124 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:93
    ges, and never instruct the user to run `sudo pip install`.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 1,124 words, ~2,460 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-live-api/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
gemini-live-api
description
Generates a Gemini LiveAPI client service class in the user's chosen programming language. Use when the user wants to build, scaffold, or integrate a client that connects to the Gemini Enterprise LiveAPI websocket endpoint, handles session setup/resumption, bearer token refresh, and sending/receiving `ClientMessage`/`ServerMessage` protos. Don't use for general (non-live, non-bidirectional) Gemini API usage such as one-shot `generateContent`, embeddings, image/video generation, or fine-tuning — use the `gemini-api` skill for those.
metadata.version
1.0.0
metadata.category
AiAndMachineLearning

LiveAPI Service Skill

This skill provides instructions for generating a LiveAPI client service class that connects to the Gemini Enterprise Live API over WebSockets. The generated client handles bidirectional streaming, bearer-token authentication via Application Default Credentials (ADC), transparent session resumption, and ClientMessage / ServerMessage proto exchange.

The skill also produces a demo frontend + backend service so the user can interactively validate the generated client (text, audio, video, transcription, and interrupt handling).

Prerequisites

Before running the generation flow, ensure the following are available on the host:

  • A Google Cloud project with the Vertex AI / Gemini Enterprise Agent Platform APIs enabled.

  • Application Default Credentials configured on the host running the generated client:

    bash
    gcloud auth application-default login
  • A destination output folder supplied by the user (e.g. /tmp/liveapi_out) where the generated code, environment, and demo will be written. Never mutate the host's system Python environment.

  • The user's chosen implementation language (Python is the default and reference language for this skill).

Reference Files

Provided files in references/ (do not treat these as standalone skills — they are loaded on demand):

  • client_server_messages.md: Public reference for the ClientMessage / ServerMessage schemas used by the Live API.
  • client_server_messages.proto: The proto definition generated from client_server_messages.md.
  • session_manager.md: Describes how to correctly handle sessions, buffering, and resumption on disconnection.

Steps

Step 1: Copy the reference files

Copy client_server_messages.md, client_server_messages.proto, and session_manager.md from this skill's references/ folder into the user's destination output folder. These files become the source of truth for the generated client.

Step 2: Reconcile with the public documentation

Examine the public documents linked from client_server_messages.md. If there are any discrepancies between the public documents and the copied client_server_messages.md / client_server_messages.proto, update the copies in the destination folder so the generated client compiles and runs against the current server contract.

Step 3: Implement the client class

Implement a class in the user's chosen language that:

  • Imports the local client_server_messages.proto types (ClientMessage, ServerMessage).
  • Opens a WebSocket connection to the Live API endpoint.
  • Exposes async methods so the user can send and receive data to/from the model.

For languages that require an isolated runtime (e.g. Python), create an isolated environment (e.g. venv) inside the destination folder and generate a bash script (e.g. setup.sh) that recreates the environment and installs dependencies. Never install into the system interpreter or the user's global site-packages, and never instruct the user to run sudo pip install.

Initialization parameters

The user provides the following at construction time:

  • project_id
  • location
  • model_id
  • config: a ClientMessage with the setup field populated.
Authentication

Obtain a bearer token via Application Default Credentials, attach it to the WebSocket connect request as Authorization: Bearer <token>, refresh the token before or upon expiry, and reuse the refreshed token on every reconnection (including go_away and unexpected disconnects). Do not hard-code a long-lived API key as the only auth mechanism.

Public async API

The class MUST expose the following async methods, gated on receipt of a setup_complete ServerMessage before sending:

  • send_realtime_data(data): send realtime input. data is a ClientMessage carrying a realtime_input field.
  • send_client_content(data): send non-realtime, turn-based content that contributes to history. data is a ClientMessage carrying a client_content field.
  • receive(): yield ServerMessage instances parsed from the WebSocket stream.

Do not expose synchronous blocking variants as the primary API surface.

Step 4: Write a test file

Once the client is implemented, generate a test file that initializes the connection and exercises sending text, audio, and video data and receiving the responses. Ask the user for any information required to run the test (project, model, media samples).

Step 5: Generate how_to_run.md

Provide a how_to_run.md in the destination folder that documents the generated class. Include full examples showing how to build ClientMessage payloads for every supported modality, how to send them, and how to receive data from the model.

Step 6: Generate a demo frontend + backend service

Create scripts that deploy the implementation as a service with both a frontend UI and a backend service (any language). The service MUST reuse the ClientMessage / ServerMessage protos from Step 1 for wire traffic. Through the UI the user should be able to:

  • Start a new connection / close the current connection.
  • Select the model to use.
  • Select input sources (audio and/or video from camera or screenshot) and stream them to the model.
  • Send a text message to the model.
  • Hear model audio and see the interleaved model and user transcription / conversation history.

While implementing audio and transcription playback, follow the guidance in Live API best practices.

Show full SKILL.md (395 more words)Show less
Handling the interrupt signal

When a ServerMessage's server_content arrives with interrupted: true, the UI MUST:

  • Ensure played audio and its corresponding transcription remain time-aligned.
  • Immediately stop the currently playing model audio and stop appending to the in-progress transcription bubble.
  • Clear the unplayed audio buffer and any pending unrendered transcription so stale content does not bleed into the next turn.
  • Start new chat bubbles for the next user and model turns.
Handling the transcription finished signal

For streamed input_transcription / output_transcription chunks, append to the currently active bubble while finished is unset, and close that bubble and start a fresh one when finished is observed. Route input_transcription text to user-role bubbles and output_transcription text to model-role bubbles.

Step 7: Generate how_to_test_with_ui.md

Write how_to_test_with_ui.md describing how to launch and use the demo service. It MUST include:

  • The exact shell command(s) or script invocation(s) to start the backend service.
  • The exact shell command(s) or script invocation(s) to start the frontend UI.
  • The host and port (e.g. http://localhost:PORT) the user should open in their browser.
  • How to start a session, select a model, choose input sources (mic, camera, screen), send a text message, and observe model audio and transcription in the UI.

Validation Checklist

Before considering the generation complete, verify each item:

  • client_server_messages.md, client_server_messages.proto, and session_manager.md were copied into the destination folder.
  • The generated client imports the local proto-generated ClientMessage and ServerMessage types.
  • The client connects to the Live API WebSocket at wss://{location}-aiplatform.googleapis.com/ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent (or the wss://aiplatform.googleapis.com/... global variant), and formats the setup model field as projects/{project_id}/locations/{location}/publishers/google/models/{model_id}.
  • Authentication uses ADC-provided bearer tokens sent as Authorization: Bearer <token>, is refreshed before expiry, and reattached on every reconnect.
  • Public async methods send_realtime_data, send_client_content, and receive are present, correctly typed, and gated on setup_complete.
  • Transparent session resumption is enabled (session_resumption.transparent = true), the latest new_handle is tracked, sent-message indexing starts at 1, the buffer is pruned via last_consumed_client_message_index, and buffered messages are replayed on reconnect (including on go_away and WebSocket close codes 1000 / 1006).
  • If using python, an isolated environment (e.g. venv) plus a setup.sh and requirements.txt (or equivalent) exist inside the destination folder; no changes were made to system or user-global Python.
  • how_to_run.md and how_to_test_with_ui.md are present, and the demo UI reuses the same ClientMessage / ServerMessage protos.
  • Interrupt handling and transcription finished handling behave as described above.
  • The client does not target generativelanguage.googleapis.com and does not authenticate via API key in a query string.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/cloud/gemini-live-api of google/skills.

  • SKILL.md
  • references/client_server_messages.md
  • references/client_server_messages.proto
  • references/session_manager.md

Open the folder on GitHubat commit 8a1ac05

Compare with similar skills

Gemini Live API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Live API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Live API this skillgoogle/skills21k—~2.5kAutomated safety check: NotesApache-2.0
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Tao Finetune Cosmos EmbedNVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0
Gemini Interactions APIAyuilos/Miffan182—~4.6kAutomated safety check: PassAGPL-3.0
Gemini Video Understandingeinverne/dotfiles1211 repos~2.6kAutomated safety check: NotesMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0

Similar skills

  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Official

    Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

    3.5k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    182 GitHub stars~4.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

    121 GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check: notes
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.1k GitHub stars~1.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes

More from google/skills

All 145 skills in this repo
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Gemini Live API

What does Gemini Live API do?

Generates a Gemini LiveAPI client service class in the user's chosen programming language. Gemini Live API is an agent skill from google/skills, published by the product's own GitHub organization. Generates a Gemini LiveAPI client service class in the user's chosen programming language.

When should I use Gemini Live API?

Gemini Live API fits situations like: the user wants to build; integrate a client that connects to the Gemini Enterprise LiveAPI websocket endpoint; handles session setup/resumption; bearer token refresh.

How do I install Gemini Live API in Claude Code?

Run `npx skills add google/skills --skill gemini-live-api -a claude-code`. Or copy the skill folder (skills/cloud/gemini-live-api in google/skills) into .claude/skills/gemini-live-api in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Live API in Codex?

Run `npx skills add google/skills --skill gemini-live-api -a codex`. Or copy the skill folder (skills/cloud/gemini-live-api in google/skills) into .agents/skills/gemini-live-api in your project. Codex loads it when a task matches its description.

Can I use Gemini Live API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gemini-live-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-live-api, .gemini/skills/gemini-live-api, .github/skills/gemini-live-api and .opencode/skills/gemini-live-api in your project.

What does Gemini Live API need to run?

Going by SKILL.md and its folder, Gemini Live API needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Gemini Live API access the network?

SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.

Is Gemini Live API safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gemini Live API use?

Gemini Live API is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Live API use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Live API?

Skills that share tags, products or a category with Gemini Live API: Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), Tao Finetune Cosmos Embed (NVIDIA/skills, 3.5k stars), Gemini Interactions API (Ayuilos/Miffan, 182 stars) and Gemini Video Understanding (einverne/dotfiles, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Live API?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.