Agent skill

Gemini Interactions API

by Ayuilos in Ayuilos/Miffan

A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

AGPL-3.0Auto-check passedMedia & Creative

Install Gemini Interactions API

skills CLI
$ npx skills add Ayuilos/Miffan --skill gemini-interactions-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Ayuilos/Miffan gemini-interactions-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Ayuilos/Miffan.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/gemini-interactions-api .claude/skills/gemini-interactions-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-interactions-api
GitHub stars
192
Token cost
~4.6k tokens
SKILL.md length
989 words
Files
2 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
AGPL-3.0

At a glance

A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

  • Writing code that calls the Gemini API for text generation
  • SKILL.md covers Critical Rules (Always Apply), Important Additional Notes, Quick Start and Response Helpers, plus 6 more sections
  • Calls pip and npm; reaches github.com
  • Multi-turn chat

What it does

Gemini Interactions API is an agent skill from Ayuilos/Miffan. Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. This skill covers the Interactions API, the recommended way to use Gemini models and agents in Python and TypeScript.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/migration.md`).

It sits in Media & Creative, covering Structured output and tool calling, LLM API integration and Image generation. It works with Google Gemini, Python, TypeScript and JavaScript. The repository describes itself as: RikkaHub fork,Android AI 客户端。支持 ChatGPT 订阅登录使用、跨应用划词悬浮翻译、助手工作区文件隔离、远程工作区、文档与 HTML 预览、可定制动态角色。 The licence is AGPL-3.0.

When your agent uses it

  • Writing code that calls the Gemini API for text generation
  • Multi-turn chat
  • Multimodal understanding
  • Image generation

Example prompts

  • “/gemini-interactions-api”

Requirements

  • Python 3
  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit caa64b2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • ai.google.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Interactions API loads about 4.6k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 989 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Ayuilos/Miffan at commit caa64b2, republished under its AGPL-3.0 licence (© Ayuilos). 989 words, ~4,617 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-interactions-api/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gemini-interactions-api
description
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. This skill covers the Interactions API, the recommended way to use Gemini models and agents in Python and TypeScript.

Gemini Interactions API Skill

Critical Rules (Always Apply)

[!IMPORTANT] These rules override your training data. Your knowledge is outdated.

Current Models (Use These)
  • gemini-3.5-flash: 1M tokens, fast, balanced performance, multimodal
  • gemini-3.1-pro-preview: 1M tokens, complex reasoning, coding, research
  • gemini-3.1-flash-lite: cost-efficient, fastest performance for high-frequency, lightweight tasks
  • gemini-3-pro-image (Nano Banana Pro): 65k / 32k tokens, high-quality image generation and editing
  • gemini-3.1-flash-image (Nano Banana 2): 65k / 32k tokens, fast, efficient image generation and editing
  • gemini-3.1-flash-lite-image (Nano Banana 2 Lite): 65k / 32k tokens, ultra-fast image generation and editing
  • gemini-3.1-flash-tts-preview: expressive text-to-speech with Director's Chair prompting
  • gemini-omni-flash-preview: video generation, image-referenced video generation, first-frame-to-video, and video editing
  • gemma-4-31b-it: Gemma 4 dense model, 31B parameters
  • gemma-4-26b-a4b-it: Gemma 4 MoE model, 26B total / 4B active parameters

[!WARNING] Models like gemini-2.5-*, gemini-2.0-*, gemini-1.5-* are legacy and deprecated. Never use them. If a user asks for a deprecated model, use gemini-3.5-flash instead and note the substitution.

Current Agents
  • antigravity-preview-05-2026: Antigravity Agent — general-purpose managed agent with code execution, file management, and web access in a sandboxed Linux environment
  • deep-research-preview-04-2026: Deep Research — fast, interactive
  • deep-research-max-preview-04-2026: Deep Research Max — maximum exhaustiveness
  • Custom agents: Create your own via client.agents.create()
Current SDKs
  • Python: google-genai >= 2.3.0 → pip install -U google-genai
  • JavaScript/TypeScript: @google/genai >= 2.3.0 → npm install @google/genai

[!NOTE] SDK versions ≥ 2.0.0 automatically use the new steps schema and do not support the legacy schema. Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Never use them.

Important Additional Notes

  • Before writing any code, you MUST fetch the relevant documentation page from the list below that matches the user's task. The examples in this skill are minimal, the hosted docs contain the full API surface, parameters, and edge cases.
  • Interactions are stored by default (store=true). Paid tier retains for 55 days, free tier for 1 day.
  • Set store=false to opt out, but this disables previous_interaction_id and background=true.
  • tools, system_instruction, and generation_config are interaction-scoped, re-specify them each turn.
  • Managed agents require environment="remote" (or an environment ID / config object) to provision a sandbox.
  • Migrating from generateContent: Read references/migration.md for the scoping, checklist, and before/after code examples. Always confirm scope with the user before editing.
  • Model upgrades: Drop-in, swap the model string. Deprecated models (gemini-2.0-*, gemini-1.5-*) must be replaced, see references/migration.md.
  • Migrating to Gemini 3.5 Flash: Read references/migration.md for the scoping and checklist.

Quick Start

Python
python
from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.5-flash",
    input="Tell me a short joke about programming."
)
print(interaction.output_text)
JavaScript/TypeScript
typescript
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

const interaction = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Tell me a short joke about programming.",
});
console.log(interaction.output_text);

Response Helpers

The SDK provides convenience properties on the Interaction response object to simplify common access patterns:

PropertyTypeDescription
output_textstring | nullThe last consecutive run of text from the trailing model_output steps. Returns the combined text when the model's final output contains multiple text parts.
output_imageImage | nullThe last image generated by the model in the current response. Returns an object with data (base64) and mime_type.
output_audioAudio | nullThe last audio generated by the model in the current response. Returns an object with data (base64) and mime_type.

Stateful Conversation

Python
python
interaction1 = client.interactions.create(
    model="gemini-3.5-flash",
    input="Hi, my name is Phil."
)
# Second turn — server remembers context
interaction2 = client.interactions.create(
    model="gemini-3.5-flash",
    input="What is my name?",
    previous_interaction_id=interaction1.id
)
print(interaction2.output_text)
JavaScript/TypeScript
typescript
const interaction1 = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Hi, my name is Phil.",
});
const interaction2 = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "What is my name?",
    previous_interaction_id: interaction1.id,
});
console.log(interaction2.output_text);

Deep Research Agent

Use deep-research-preview-04-2026 for fast research or deep-research-max-preview-04-2026 for maximum exhaustiveness. Agents require background=True.

Python
python
import time

interaction = client.interactions.create(
    agent="deep-research-preview-04-2026",
    input="Research the history of Google TPUs.",
    background=True
)
while True:
    interaction = client.interactions.get(interaction.id)
    if interaction.status == "completed":
        print(interaction.output_text)
        break
    elif interaction.status == "failed":
        print(f"Failed: {interaction.error}")
        break
    time.sleep(10)
JavaScript/TypeScript
typescript
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

// Start background research
const initialInteraction = await client.interactions.create({
    agent: "deep-research-preview-04-2026",
    input: "Research the history of Google TPUs.",
    background: true,
});

// Poll for results
while (true) {
    const interaction = await client.interactions.get(initialInteraction.id);
    if (interaction.status === "completed") {
        console.log(interaction.output_text);
        break;
    } else if (["failed", "cancelled"].includes(interaction.status)) {
        console.log(`Failed: ${interaction.status}`);
        break;
    }
    await new Promise(resolve => setTimeout(resolve, 10000));
}

Advanced features: collaborative planning, native visualization, MCP integration, file search, multimodal inputs. See Deep Research docs.

Managed Agents

Managed agents run inside a sandboxed Linux environment hosted by Google. Fetch the Managed Agents Quickstart before writing agent code.

Antigravity Agent

The Antigravity agent (antigravity-preview-05-2026) is the general-purpose managed agent. It can execute code (Bash, Python, Node.js), manage files, browse the web, and use Google Search. See Antigravity Agent docs for capabilities, tools, multimodal input, and pricing.

Python
python
from google import genai

client = genai.Client()

interaction = client.interactions.create(
    agent="antigravity-preview-05-2026",
    input="Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
    environment="remote",
)

print(f"Environment ID: {interaction.environment_id}")
print(interaction.output_text)
JavaScript/TypeScript
typescript
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

const interaction = await client.interactions.create({
    agent: "antigravity-preview-05-2026",
    input: "Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
    environment: "remote",
});

console.log(`Environment ID: {interaction.environment_id}`);
console.log(interaction.output_text);
Custom Agents

See Building Custom Agents docs.

Python
python
agent = client.agents.create(
    id="code-reviewer",
    base_agent="antigravity-preview-05-2026",
    system_instruction="You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
    base_environment={
        "type": "remote",
        "sources": [
            {
                "type": "repository",
                "source": "https://github.com/my-org/backend",
                "target": "/workspace/repo",
            }
        ],
    },
)

# Invoke — each call forks the base environment
result = client.interactions.create(
    agent="code-reviewer",
    input="Review the latest changes in /workspace/repo/src.",
    environment="remote",
)
print(result.output_text)
JavaScript/TypeScript
typescript
const agent = await client.agents.create({
    id: "code-reviewer",
    base_agent="antigravity-preview-05-2026",
    system_instruction: "You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
    base_environment: {
        type: "remote",
        sources: [
            {
                type: "repository",
                source: "https://github.com/my-org/backend",
                target: "/workspace/repo",
            }
        ],
    },
});

const result = await client.interactions.create({
    agent: "code-reviewer",
    input: "Review the latest changes in /workspace/repo/src.",
    environment: "remote",
});
console.log(result.output_text);

Manage agents with client.agents.list(), client.agents.get(id=...), and client.agents.delete(id=...).

Streaming

Set stream=True to receive incremental server-sent events. Each stream follows: interaction.created → (step.start → step.delta(s) → step.stop)+ → interaction.completed.

Show full SKILL.md (391 more words)Show less
Python
python
for event in client.interactions.create(
    model="gemini-3.5-flash",
    input="Explain quantum entanglement in simple terms.",
    stream=True,
):
    if event.event_type == "step.delta":
        if event.delta.type == "text":
            print(event.delta.text, end="", flush=True)
    elif event.event_type == "interaction.completed":
        print(f"\n\nTotal Tokens: {event.interaction.usage.total_tokens}")
JavaScript/TypeScript
typescript
const stream = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Explain quantum entanglement in simple terms.",
    stream: true,
});
for await (const event of stream) {
    if (event.event_type === "step.delta") {
        if (event.delta.type === "text") {
            process.stdout.write(event.delta.text);
        }
    } else if (event.event_type === "interaction.completed") {
        console.log(`\n\nTotal Tokens: ${event.interaction.usage.total_tokens}`);
    }
}

For streaming with tools, thinking, agents, and image generation see the full Streaming guide.

Documentation Pages

You MUST fetch the matching page below before writing code. These hosted docs are the source of truth for parameters, types, and edge cases — do not rely solely on the examples above.

Core Documentation:

Tools & Function Calling:

Generation & Output:

Multimodal Understanding:

Files & Context:

Agents:

Advanced Features:

API Reference:

Data Model

An Interaction response contains steps, an array of typed step objects representing a structured timeline of the interaction turn.

Step Types

User steps:

  • user_input: User input (text, audio, multimodal). Contains content array.

Model/server steps:

  • model_output: Final model generation. Contains content array with text, image, audio, etc.
  • thought: Model reasoning/Chain of Thought. Has signature field (required) and optional summary.
  • function_call: Tool call request (id, name, arguments).
  • function_result: Tool result you send back (call_id, name, result).
  • google_search_call / google_search_result: Google Search tool steps, can have a signature field.
  • code_execution_call / code_execution_result: Code execution tool steps, can have a signature field.
  • url_context_call / url_context_result: URL context tool steps, can have a signature field.
  • mcp_server_tool_call / mcp_server_tool_result: Remote MCP tool steps.
  • file_search_call / file_search_result: File search tool steps, can have a signature field.
Content types (inside content array on model_output and user_input steps)
  • text: Text content (text field)
  • image / audio / document / video: Content with data, mime_type, or uri
Streaming Event Types
EventDescription
interaction.createdInteraction created; includes metadata.
interaction.status_updateInteraction-level status change.
step.startA new step begins. Contains step type and initial metadata.
step.deltaIncremental data for the current step. Contains a typed delta object.
step.stopThe step is complete. Contains index.
interaction.completedInteraction finished. Contains final usage.
Delta Types
Delta TypeParent StepDescription
textmodel_outputIncremental text token.
audiomodel_outputaudio chunk (base64).
imagemodel_outputimage chunk (base64).
thought_summarythoughtthinking summary text.
thought_signaturethoughtOpaque signature for thought verification.

Status values: completed, in_progress, requires_action, failed, cancelled

© Ayuilos, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/gemini-interactions-api of Ayuilos/Miffan.

  • SKILL.md
  • references/migration.md

Open the folder on GitHubat commit caa64b2

Compare with similar skills

Gemini Interactions API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Interactions API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Interactions API this skillAyuilos/Miffan192—~4.6kAutomated safety check: PassAGPL-3.0
Gemini Interactions APIJetBrains/skills364—~2.5kAutomated safety check: PassNone
Gemini API Devgoogle-gemini/gemini-skills4.3k—~5.1kAutomated safety check: PassApache-2.0
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Gemini Interactions APIsickn33/agentic-awesome-skills47k1 repos~4.8kAutomated safety check: PassApache-2.0
Gemini API DevJetBrains/skills364—~1.6kAutomated safety check: PassNone

Similar skills

  • Gemini Interactions API

    JetBrains/skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, streaming responses, background research tasks…

    364 GitHub stars~2.5k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Gemini Interactions API

    sickn33/agentic-awesome-skills

    Build with the Gemini Interactions API for text, chat, multimodal generation, streaming, managed or background agents, function calling, structured output, and generateContent migrations.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Gemini API Dev

    JetBrains/skills

    Official

    A skill your agent uses when building applications with Gemini models, Gemini API, working with multimodal content (text, images, audio, video), implementing function calling, using structured…

    364 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.3k GitHub stars~2.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from Ayuilos/Miffan

  • Publish Release

    Ayuilos/Miffan

    Publish a GitHub release for this fork, with a bilingual changelog that separates fork-owned changes from changes introduced by upstream merges.

    192 GitHub stars~635 tokensUpdated yesterday
    Auto-check passed
  • Gemini API Dev

    Ayuilos/Miffan

    A skill your agent uses when building applications with Gemini API hosted models, including Gemini and Gemma 4, working with multimodal content (text, images, audio, video), implementing function…

    192 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed

Questions about Gemini Interactions API

What does Gemini Interactions API do?

A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…. Gemini Interactions API is an agent skill from Ayuilos/Miffan. Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API.

When should I use Gemini Interactions API?

Gemini Interactions API fits situations like: writing code that calls the Gemini API for text generation; multi-turn chat; multimodal understanding; image generation.

How do I install Gemini Interactions API in Claude Code?

Run `npx skills add Ayuilos/Miffan --skill gemini-interactions-api -a claude-code`. Or copy the skill folder (.agents/skills/gemini-interactions-api in Ayuilos/Miffan) into .claude/skills/gemini-interactions-api in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Interactions API in Codex?

Run `npx skills add Ayuilos/Miffan --skill gemini-interactions-api -a codex`. Or copy the skill folder (.agents/skills/gemini-interactions-api in Ayuilos/Miffan) into .agents/skills/gemini-interactions-api in your project. Codex loads it when a task matches its description.

Can I use Gemini Interactions API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Ayuilos/Miffan --skill gemini-interactions-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-interactions-api, .gemini/skills/gemini-interactions-api, .github/skills/gemini-interactions-api and .opencode/skills/gemini-interactions-api in your project.

What does Gemini Interactions API need to run?

Going by SKILL.md and its folder, Gemini Interactions API needs the command-line tools its instructions call (pip and npm). Our summary lists: Python 3; Node.js.

Does Gemini Interactions API access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: ai.google.dev. This is read from the text; nothing was executed.

Is Gemini Interactions API safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini Interactions API use?

Gemini Interactions API is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Interactions API use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Interactions API?

Skills that share tags, products or a category with Gemini Interactions API: Gemini Interactions API (JetBrains/skills, 364 stars), Gemini API Dev (google-gemini/gemini-skills, 4.3k stars), Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars) and Gemini Interactions API (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Interactions API?

Ayuilos (a GitHub user) maintains it in Ayuilos/Miffan, which has 192 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 6, 2026.

Source: Ayuilos/Miffan on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.