Agent skill

Gemini Audio

by einverne in einverne/dotfiles

Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

MITAuto-check: notesMedia & Creative

Install Gemini Audio

skills CLI
$ npx skills add einverne/dotfiles --skill gemini-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install einverne/dotfiles gemini-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-audio .claude/skills/gemini-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-audio
GitHub stars
121
Token cost
~2k tokens
SKILL.md length
551 words
Files
12 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

  • Works in 3 steps: Process environment: export… → Skill directory:… → Project directory: ./.env (project root)
  • Processing audio files
  • SKILL.md covers When to Use This Skill, Prerequisites, Quick Start and Audio Understanding Capabilities, plus 4 more sections
  • Runs Python scripts from its folder; calls python and pip; needs GEMINI_API_KEY

What it does

Gemini Audio is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS. Use when processing audio files, creating transcripts, analyzing speech/music/sounds, or generating natural speech from text.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/README.md`, `references/api-reference.md` and `references/best-practices.md`).

It sits in Media & Creative, covering Transcription, Text to speech and voice and Summarization. It works with Google Gemini. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is MIT.

When your agent uses it

  • Processing audio files
  • Creating transcripts
  • Analyzing speech/music/sounds
  • Generating natural speech from text

Example prompts

  • “/gemini-audio”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Skill directory: .claude/skills/gemini-audio/.env
  3. Project directory: ./.env (project root)

What it can do on your machine

Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ai.google.dev
    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Audio loads about 2k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~18k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:33
    irectory**: `.claude/skills/gemini-audio/.env`
  • NoteMentions a .env fileSKILL.md:34
    3. **Project directory**: `./.env` (project root)
  • NoteMentions a .env fileSKILL.md:38
    Create `.env` file with:
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from einverne/dotfiles at commit c6c0686, republished under its MIT licence (© einverne). 551 words, ~2,010 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-audio/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
gemini-audio
description
Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS. Use when processing audio files, creating transcripts, analyzing speech/music/sounds, or generating natural speech from text.
allowed-tools
Bash, Read, Write, Edit
license
MIT

Gemini Audio API Skill

Process audio with transcription, analysis, and understanding, plus generate natural speech using Google's Gemini API. Supports up to 9.5 hours of audio per request with multiple formats.

When to Use This Skill

Use this skill when you need to:

  • Transcribe audio files to text with timestamps
  • Summarize audio content and extract key points
  • Analyze speech, music, or environmental sounds
  • Generate speech from text with controllable voice and style
  • Process podcasts, interviews, meetings, or any audio content
  • Understand non-speech audio (birdsong, sirens, music)

Prerequisites

API Key Setup

The skill automatically detects your GEMINI_API_KEY in this order:

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Skill directory: .claude/skills/gemini-audio/.env
  3. Project directory: ./.env (project root)

Get your API key: Visit Google AI Studio

Create .env file with:

bash
GEMINI_API_KEY=your_api_key_here
Python Setup

Install required package:

bash
pip install google-genai

Quick Start

Audio Analysis (Transcription, Summarization)
python
from google import genai
import os

# API key auto-detected from environment
client = genai.Client(api_key=os.getenv('GEMINI_API_KEY'))

# Upload audio file
myfile = client.files.upload(file='podcast.mp3')

# Transcribe
response = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=['Generate a transcript of the speech.', myfile]
)
print(response.text)

# Summarize
response = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=['Summarize the key points in 5 bullets.', myfile]
)
print(response.text)
Using Helper Scripts
bash
# Transcribe audio
python .claude/skills/gemini-audio/scripts/transcribe.py audio.mp3

# Summarize audio
python .claude/skills/gemini-audio/scripts/analyze.py audio.mp3 \
  "Summarize key points"

# Analyze specific segment (timestamps in MM:SS format)
python .claude/skills/gemini-audio/scripts/analyze.py audio.mp3 \
  "What is discussed from 02:30 to 05:15?"

# Generate speech
python .claude/skills/gemini-audio/scripts/generate-speech.py \
  "Welcome to our podcast" \
  --output welcome.wav

Audio Understanding Capabilities

Supported Formats
FormatMIME TypeBest Use
WAVaudio/wavUncompressed, highest quality
MP3audio/mp3Compressed, widely compatible
AACaudio/aacCompressed, good quality
FLACaudio/flacLossless compression
OGG Vorbisaudio/oggOpen format
AIFFaudio/aiffApple format
Audio Specifications
  • Maximum length: 9.5 hours per request
  • Multiple files: Unlimited count, combined max 9.5 hours
  • Token rate: 32 tokens/second (1 minute = 1,920 tokens)
  • Processing: Auto-downsampled to 16 Kbps mono
  • File size limits:
    • Inline: 20 MB max total request
    • File API: 2 GB per file, 20 GB project quota
    • Retention: 48 hours auto-delete
Analysis Features
  • Transcription: Full text with punctuation
  • Timestamps: Reference segments (MM:SS format)
  • Multi-speaker: Identify different speakers
  • Non-speech: Analyze music, sounds, ambient audio
  • Languages: Support for multiple languages

Speech Generation (TTS)

Available TTS Models
ModelQualitySpeedCost/1M tokens
gemini-2.5-flash-native-audio-preview-09-2025HighFast$10
gemini-2.5-pro TTS modePremiumSlower$20
Controllable Voice Options
  • Style: Professional, casual, narrative, conversational
  • Pace: Slow, normal, fast
  • Tone: Friendly, serious, enthusiastic
  • Accent: Natural language control
TTS Example
python
response = client.models.generate_content(
    model='gemini-2.5-flash-native-audio-preview-09-2025',
    contents='Generate audio: Welcome to today\'s episode, in a warm, friendly tone.'
)

# Save audio output
with open('output.wav', 'wb') as f:
    f.write(response.audio_data)

Input Methods

python
# Upload and reuse
myfile = client.files.upload(file='large-audio.mp3')

# Use file multiple times
response1 = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=['Transcribe this', myfile]
)

response2 = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=['Summarize this', myfile]
)
Method 2: Inline Data (<20MB)
python
from google.genai import types

with open('small-audio.mp3', 'rb') as f:
    audio_bytes = f.read()

response = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=[
        'Describe this audio',
        types.Part.from_bytes(data=audio_bytes, mime_type='audio/mp3')
    ]
)

Common Use Cases

Transcription
bash
python scripts/transcribe.py meeting.mp3 --include-timestamps
Summary with Key Points
bash
python scripts/analyze.py interview.wav "Extract main topics and key quotes"
Speaker Identification
bash
python scripts/analyze.py discussion.mp3 "Identify speakers and extract dialogue"
Segment Analysis
bash
python scripts/analyze.py podcast.mp3 "Summarize content from 10:30 to 15:45"
Non-Speech Analysis
bash
python scripts/analyze.py ambient.wav "Identify all sounds: voices, music, ambient"

Best Practices

Show full SKILL.md (220 more words)Show less
File Management
  • Use File API for files >20MB or repeated usage
  • Files auto-delete after 48 hours
  • Manage quota (20 GB project limit)
Prompt Engineering
  • Be specific: "Transcribe from 02:30 to 03:29"
  • Use timestamps for segment analysis (MM:SS format)
  • Combine tasks: "Transcribe and summarize"
  • Provide context: "This is a medical interview"
Cost Optimization
  • Use gemini-2.5-flash ($1/1M tokens) for most tasks
  • Upgrade to gemini-2.5-pro ($3/1M tokens) for complex analysis
  • Check token count: 1 min audio = 1,920 tokens
Error Handling
  • Validate file format and size before upload
  • Implement exponential backoff for rate limits
  • Handle 48-hour file expiration

Token Costs & Pricing

Audio Input (32 tokens/second):

  • 1 minute = 1,920 tokens
  • 1 hour = 115,200 tokens
  • 9.5 hours = 1,094,400 tokens

Model Pricing:

  • Gemini 2.5 Flash: $1.00/1M input, $0.10/1M output
  • Gemini 2.5 Pro: $3.00/1M input, $12.00/1M output
  • Gemini 1.5 Flash: $0.70/1M input, $0.175/1M output

TTS Pricing:

  • Flash TTS: $10/1M tokens
  • Pro TTS: $20/1M tokens

Reference Documentation

For detailed information, see:

  • references/api-reference.md - Complete API specifications
  • references/code-examples.md - Comprehensive code examples
  • references/tts-guide.md - Text-to-speech implementation guide
  • references/best-practices.md - Advanced optimization strategies

Scripts Overview

All scripts support 3-step API key detection:

  • transcribe.py: Generate transcripts with optional timestamps
  • analyze.py: General audio analysis with custom prompts
  • generate-speech.py: Text-to-speech generation
  • manage-files.py: Upload, list, and delete audio files

Run any script with --help for detailed usage.

Resources

© einverne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in claude/skills/gemini-audio of einverne/dotfiles.

  • SKILL.md
  • .env.example
  • references/README.md
  • references/api-reference.md
  • references/best-practices.md
  • references/code-examples.md
  • references/tts-guide.md
  • scripts/analyze.py
  • scripts/api_key_helper.py
  • scripts/generate-speech.py
  • scripts/manage-files.py
  • scripts/transcribe.py

Open the folder on GitHubat commit c6c0686

Compare with similar skills

Gemini Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Audio this skilleinverne/dotfiles121—~2kAutomated safety check: NotesMIT
AI MultimodalMicrock/ordinary-claude-skills404—~2.7kAutomated safety check: NotesMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Elevenlabs Transcribeqdhenry/Claude-Command-Suite1.3k—~1.5kAutomated safety check: NotesNone
Video Assemblezenstory-ai/video-recap-skills559—~1.7kAutomated safety check: PassMIT
Gemini Ttsiurysza/module-graph420—~968Automated safety check: PassMIT

Similar skills

  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    404 GitHub stars~2.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • Elevenlabs Transcribe

    qdhenry/Claude-Command-Suite

    Transcribes audio/video files using ElevenLabs Scribe v2 API.

    1.3k GitHub stars~1.5k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    559 GitHub stars~1.7k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Gemini Tts

    iurysza/module-graph

    Generates spoken MP3 audio from text or Markdown with Gemini TTS.

    420 GitHub stars~968 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Transcribe

    comol/ai_rules_1c

    Transcribe video and audio files via Gemini API. An agent skill from comol/ai_rules_1c.

    481 GitHub stars~960 tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes

More from einverne/dotfiles

All 39 skills in this repo
  • DOCX

    einverne/dotfiles

    Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.

    121 GitHub starsUsed in 35 repos~2.5k tokens
    Auto-check: notes
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Auto-check passed
  • Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

    121 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check: notes
  • PPTX

    einverne/dotfiles

    Presentation creation, editing, and analysis. An agent skill from einverne/dotfiles.

    121 GitHub starsUsed in 38 repos~6.4k tokens
    Auto-check: notes
  • Gemini Image Gen

    einverne/dotfiles

    Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.

    121 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Questions about Gemini Audio

What does Gemini Audio do?

Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS. Gemini Audio is an agent skill from einverne/dotfiles.5 hours), plus generate speech with controllable TTS.

When should I use Gemini Audio?

Gemini Audio fits situations like: processing audio files; creating transcripts; analyzing speech/music/sounds; generating natural speech from text.

How do I install Gemini Audio in Claude Code?

Run `npx skills add einverne/dotfiles --skill gemini-audio -a claude-code`. Or copy the skill folder (claude/skills/gemini-audio in einverne/dotfiles) into .claude/skills/gemini-audio in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Audio in Codex?

Run `npx skills add einverne/dotfiles --skill gemini-audio -a codex`. Or copy the skill folder (claude/skills/gemini-audio in einverne/dotfiles) into .agents/skills/gemini-audio in your project. Codex loads it when a task matches its description.

Can I use Gemini Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-audio, .gemini/skills/gemini-audio, .github/skills/gemini-audio and .opencode/skills/gemini-audio in your project.

What does Gemini Audio need to run?

Going by SKILL.md and its folder, Gemini Audio needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Gemini Audio access the network?

SKILL.md names 2 domains. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.

Is Gemini Audio safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Audio use?

Gemini Audio is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Audio use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Audio?

Skills that share tags, products or a category with Gemini Audio: AI Multimodal (Microck/ordinary-claude-skills, 404 stars), Edu Math Video (wy51ai/edulab, 1.4k stars), Elevenlabs Transcribe (qdhenry/Claude-Command-Suite, 1.3k stars) and Video Assemble (zenstory-ai/video-recap-skills, 559 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Audio?

einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.

Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.