Agent skill

Speech Use

by cnemri in cnemri/google-genai-skills

Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs.

MITAuto-check passedMedia & Creative

Install Speech Use

skills CLI
$ npx skills add cnemri/google-genai-skills --skill speech-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cnemri/google-genai-skills speech-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cnemri/google-genai-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speech-use .claude/skills/speech-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-use
GitHub stars
127
Token cost
~616 tokens
SKILL.md length
195 words
Files
5 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs.

  • Works in 3 steps: Generate Speech (TTS) → Create Custom Voice (Voice Cloning) → Transcribe Audio (STT)
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Prerequisites, Usage, Options and References
  • Runs Python scripts from its folder; calls uv; needs GOOGLE_API_KEY and GOOGLE_APPLICATION_CREDENTIALS

What it does

Speech Use is an agent skill from cnemri/google-genai-skills. Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs. Supports Gemini-TTS, Chirp 3, and Instant Custom Voice.

Its SKILL.md is about 620 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/voices.md`, `scripts/create_custom_voice.py` and `scripts/generate_speech.py`).

It sits in Media & Creative, covering Text to speech and voice and Transcription. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Transcription

Example prompts

  • “/speech-use”

Requirements

  • Python 3
  • A credential in GOOGLE_API_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Generate Speech (TTS)
  2. Create Custom Voice (Voice Cloning)
  3. Transcribe Audio (STT)

What it can do on your machine

Read from SKILL.md and the folder at commit 7277476. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY
    • GOOGLE_APPLICATION_CREDENTIALS

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speech Use loads about 616 tokens when it runs, and up to ~937 if it reads all its reference files. Until then it costs about 40 tokens; SKILL.md has 195 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~616
With references · SKILL.md plus every file in references/, read only if the agent opens them
~937

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cnemri/google-genai-skills at commit 7277476, republished under its MIT licence (© cnemri). 195 words, ~616 tokens.

Download SKILL.mdSave it as .claude/skills/speech-use/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
speech-use
description
Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs. Supports Gemini-TTS, Chirp 3, and Instant Custom Voice.

Speech Use

Use this skill to perform Text-to-Speech (TTS), Speech-to-Text (STT), and Voice Cloning operations.

This skill uses portable Python scripts managed by uv.

Prerequisites

  1. Environment Variables:

    • GOOGLE_API_KEY (for TTS via Gemini)
    • GOOGLE_CLOUD_PROJECT (Required for STT and Voice Cloning)
    • GOOGLE_APPLICATION_CREDENTIALS (Recommended for STT/Voice Cloning)
  2. APIs Enabled:

    • Text-to-Speech API (texttospeech.googleapis.com)
    • Speech-to-Text API (speech.googleapis.com)

Usage

1. Generate Speech (TTS)

Generate audio from text using Gemini-TTS.

Standard Voice:

bash
uv run skills/speech-use/scripts/generate_speech.py "Hello world, this is a test." --voice Puck --output hello.wav

Custom Voice (Cloned):

bash
uv run skills/speech-use/scripts/generate_speech.py "This is my custom voice speaking." --voice-cloning-key "YOUR_KEY_HERE" --output custom.wav
2. Create Custom Voice (Voice Cloning)

Generate a voiceCloningKey from a reference audio file and a consent file.

Requirements:

  • reference.wav: 10-30s of clear speech (the voice to clone).
  • consent.wav: The speaker saying: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model."
bash
uv run skills/speech-use/scripts/create_custom_voice.py --reference-audio reference.wav --consent-audio consent.wav

Save the output key to use with generate_speech.py.

3. Transcribe Audio (STT)

Transcribe audio files using Chirp 3.

bash
uv run skills/speech-use/scripts/transcribe_audio.py audio.wav --language en-US --output transcript.txt

Options

generate_speech.py

  • --voice: Prebuilt voice (e.g., Kore, Puck, Fenrir, Aoede).
  • --voice-cloning-key: Key from create_custom_voice.py.
  • --model: Default gemini-2.5-flash-preview-tts.

transcribe_audio.py

  • --model: Default chirp_3.
  • --language: Default auto.
  • --location: Cloud region (default us).

References

Before running scripts, review the reference guides for available voices and options.

  • Voices Guide - 30+ voice options with styles (Puck, Kore, Fenrir, Aoede, etc.)

© cnemri, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/speech-use of cnemri/google-genai-skills.

  • SKILL.md
  • references/voices.md
  • scripts/create_custom_voice.py
  • scripts/generate_speech.py
  • scripts/transcribe_audio.py

Open the folder on GitHubat commit 7277476

Compare with similar skills

Speech Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech Use this skillcnemri/google-genai-skills127—~616Automated safety check: PassMIT
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Videodbaffaan-m/ECC275k3 repos~3.5kAutomated safety check: NotesMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Video Assemblezenstory-ai/video-recap-skills5551 repos~1.7kAutomated safety check: PassMIT
Elevenlabs Transcribeqdhenry/Claude-Command-Suite1.3k—~1.5kAutomated safety check: NotesNone

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    555 GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • Elevenlabs Transcribe

    qdhenry/Claude-Command-Suite

    Transcribes audio/video files using ElevenLabs Scribe v2 API.

    1.3k GitHub stars~1.5k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub starsUsed in 1 repo~2k tokens
    Media & CreativeAuto-check: notes

More from cnemri/google-genai-skills

All 10 skills in this repo
  • Deep Research

    cnemri/google-genai-skills

    Perform autonomous, multi-step research using the Gemini Deep Research Agent (Interactions API).

    127 GitHub stars~613 tokensUpdated 8 mo ago
    Auto-check passed
  • Veo Use

    cnemri/google-genai-skills

    Create and edit videos using Google's Veo 2 and Veo 3 models.

    127 GitHub stars~625 tokensUpdated 8 mo ago
    Auto-check passed
  • Google Developer Knowledge

    cnemri/google-genai-skills

    Search and retrieve Google's developer documentation using the Developer Knowledge API.

    127 GitHub stars~709 tokensUpdated 8 mo ago
    Auto-check passed
  • Nano Banana Use

    cnemri/google-genai-skills

    Generate, edit, and compose images using Gemini Nano Banana models via portable Python scripts.

    127 GitHub stars~587 tokensUpdated 8 mo ago
    Auto-check passed
  • Google Adk Python

    cnemri/google-genai-skills

    Expert guidance on the Google Agent Development Kit (ADK) for Python.

    127 GitHub stars~769 tokensUpdated 8 mo ago
    Auto-check passed
  • Nano Banana Build

    cnemri/google-genai-skills

    Generate and edit high-quality images using Gemini 2.5 Flash Image and Gemini 3 Pro Image (Nano Banana).

    127 GitHub stars~534 tokensUpdated 8 mo ago
    Auto-check passed

Questions about Speech Use

What does Speech Use do?

Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs. Speech Use is an agent skill from cnemri/google-genai-skills. Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs.

When should I use Speech Use?

Speech Use fits situations like: tasks that involve Text to speech and voice; tasks that involve Transcription.

How do I install Speech Use in Claude Code?

Run `npx skills add cnemri/google-genai-skills --skill speech-use -a claude-code`. Or copy the skill folder (skills/speech-use in cnemri/google-genai-skills) into .claude/skills/speech-use in your project. Claude Code loads it when a task matches its description.

How do I install Speech Use in Codex?

Run `npx skills add cnemri/google-genai-skills --skill speech-use -a codex`. Or copy the skill folder (skills/speech-use in cnemri/google-genai-skills) into .agents/skills/speech-use in your project. Codex loads it when a task matches its description.

Can I use Speech Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cnemri/google-genai-skills --skill speech-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-use, .gemini/skills/speech-use, .github/skills/speech-use and .opencode/skills/speech-use in your project.

What does Speech Use need to run?

Going by SKILL.md and its folder, Speech Use needs Python for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named GOOGLE_API_KEY and GOOGLE_APPLICATION_CREDENTIALS. Our summary lists: Python 3; A credential in GOOGLE_API_KEY.

Does Speech Use access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Speech Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Speech Use use?

Speech Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speech Use use?

About 616 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 321 tokens, read only when the agent opens those files.

What are the alternatives to Speech Use?

Skills that share tags, products or a category with Speech Use: HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Videodb (affaan-m/ECC, 275k stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Video Assemble (zenstory-ai/video-recap-skills, 555 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech Use?

cnemri (a GitHub user) maintains it in cnemri/google-genai-skills, which has 127 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on February 6, 2026.

Source: cnemri/google-genai-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.