Agent skill

Speech Build

by cnemri in cnemri/google-genai-skills

Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.

MITAuto-check passedMedia & Creative

Install Speech Build

skills CLI
$ npx skills add cnemri/google-genai-skills --skill speech-build -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cnemri/google-genai-skills speech-build --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cnemri/google-genai-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speech-build .claude/skills/speech-build && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-build
GitHub stars
127
Token cost
~430 tokens
SKILL.md length
78 words
Files
6 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.

  • Works in 2 steps: Generate Speech (Gemini-TTS) → Transcribe Audio (Chirp 3)
  • Tasks that involve Transcription
  • SKILL.md covers Quick Start Setup, Reference Materials and Common Workflows
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Speech Build is an agent skill from cnemri/google-genai-skills. Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Supports Text-to-Speech (Single/Multi-speaker), Instant Custom Voice, and Speech-to-Text (Transcription/Diarization).

Its SKILL.md is about 430 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/prompting.md`, `references/source_code.md` and `references/stt.md`).

It sits in Media & Creative, covering Transcription, Text to speech and voice and Speech recognition and synthesis. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Text to speech and voice
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/speech-build”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Generate Speech (Gemini-TTS)
  2. Transcribe Audio (Chirp 3)

What it can do on your machine

Read from SKILL.md and the folder at commit 7277476. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speech Build loads about 430 tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 78 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~430
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cnemri/google-genai-skills at commit 7277476, republished under its MIT licence (© cnemri). 78 words, ~430 tokens.

Download SKILL.mdSave it as .claude/skills/speech-build/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
speech-build
description
Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Supports Text-to-Speech (Single/Multi-speaker), Instant Custom Voice, and Speech-to-Text (Transcription/Diarization).

Speech Skill (TTS & STT)

Use this skill to implement audio generation and transcription workflows using the google-genai and google-cloud-speech SDKs.

Quick Start Setup

python
from google import genai
from google.genai import types
# For STT: from google.cloud import speech_v2

client = genai.Client()

Reference Materials

Common Workflows

1. Generate Speech (Gemini-TTS)
python
response = client.models.generate_content(
    model="gemini-2.5-flash-preview-tts",
    contents="Hello, world!",
    config=types.GenerateContentConfig(
        response_modalities=["AUDIO"],
        speech_config=types.SpeechConfig(
            voice_config=types.VoiceConfig(
                prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name='Kore')
            )
        )
    )
)
2. Transcribe Audio (Chirp 3)
python
# Requires google-cloud-speech
from google.cloud import speech_v2
# ... (See stt.md for full setup)
response = speech_client.recognize(...)

© cnemri, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/speech-build of cnemri/google-genai-skills.

  • SKILL.md
  • references/prompting.md
  • references/source_code.md
  • references/stt.md
  • references/tts.md
  • references/voices.md

Open the folder on GitHubat commit 7277476

Compare with similar skills

Speech Build next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech Build compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech Build this skillcnemri/google-genai-skills127—~430Automated safety check: PassMIT
Voice AI Engine Developmentaiskillstore/marketplace4305 repos~5.7kAutomated safety check: PassNone
Stepfun Asrdaymade/claude-code-skills1.4k—~3kAutomated safety check: PassMIT
DeepgramAnil-matcha/awesome-muse-connectors1.3k—~857Automated safety check: PassMIT
Groq Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMIT
AudiopodLeoYeAI/openclaw-master-skills2.2k—~5.5kAutomated safety check: PassMIT

Similar skills

  • Voice AI Engine Development

    aiskillstore/marketplace

    Build real-time conversational AI voice engines using async worker pipelines, streaming transcription, LLM agents, and TTS synthesis with interrupt handling and multi-provider support

    430 GitHub starsUsed in 5 repos~5.7k tokens
    Media & CreativeAuto-check passed
  • Stepfun Asr

    daymade/claude-code-skills

    Transcribes Chinese/English audio with StepFun's stepaudio-3-asr-max via its SSE endpoint (not /v1/audio/transcriptions) — one call handles long-form audio with no chunking.

    1.4k GitHub stars~3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Deepgram

    Anil-matcha/awesome-muse-connectors

    Deepgram speech AI: transcribe audio to text and synthesize speech (TTS).

    1.3k GitHub stars~857 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Groq Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses when you need Groq's non-chat endpoints — transcribing or translating audio with Whisper, understanding images with Llama 4 vision, generating speech (TTS), or benchmarking…

    2.8k GitHub stars~1.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Audiopod

    LeoYeAI/openclaw-master-skills

    Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction…

    2.2k GitHub stars~5.5k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Fal Audio

    aiskillstore/marketplace

    Text-to-speech and speech-to-text using fal.ai audio models. An agent skill from aiskillstore/marketplace.

    430 GitHub starsUsed in 5 repos~174 tokens
    Media & CreativeAuto-check passed

More from cnemri/google-genai-skills

All 10 skills in this repo
  • Deep Research

    cnemri/google-genai-skills

    Perform autonomous, multi-step research using the Gemini Deep Research Agent (Interactions API).

    127 GitHub stars~613 tokensUpdated 8 mo ago
    Auto-check passed
  • Speech Use

    cnemri/google-genai-skills

    Generate (TTS), Transcribe (STT), and Clone voices using Google's GenAI and Cloud Speech SDKs.

    127 GitHub stars~616 tokensUpdated 8 mo ago
    Auto-check passed
  • Veo Use

    cnemri/google-genai-skills

    Create and edit videos using Google's Veo 2 and Veo 3 models.

    127 GitHub stars~625 tokensUpdated 8 mo ago
    Auto-check passed
  • Google Developer Knowledge

    cnemri/google-genai-skills

    Search and retrieve Google's developer documentation using the Developer Knowledge API.

    127 GitHub stars~709 tokensUpdated 8 mo ago
    Auto-check passed
  • Nano Banana Use

    cnemri/google-genai-skills

    Generate, edit, and compose images using Gemini Nano Banana models via portable Python scripts.

    127 GitHub stars~587 tokensUpdated 8 mo ago
    Auto-check passed
  • Google Adk Python

    cnemri/google-genai-skills

    Expert guidance on the Google Agent Development Kit (ADK) for Python.

    127 GitHub stars~769 tokensUpdated 8 mo ago
    Auto-check passed

Questions about Speech Build

What does Speech Build do?

Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Speech Build is an agent skill from cnemri/google-genai-skills. Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.

When should I use Speech Build?

Speech Build fits situations like: tasks that involve Transcription; tasks that involve Text to speech and voice; tasks that involve Speech recognition and synthesis.

How do I install Speech Build in Claude Code?

Run `npx skills add cnemri/google-genai-skills --skill speech-build -a claude-code`. Or copy the skill folder (skills/speech-build in cnemri/google-genai-skills) into .claude/skills/speech-build in your project. Claude Code loads it when a task matches its description.

How do I install Speech Build in Codex?

Run `npx skills add cnemri/google-genai-skills --skill speech-build -a codex`. Or copy the skill folder (skills/speech-build in cnemri/google-genai-skills) into .agents/skills/speech-build in your project. Codex loads it when a task matches its description.

Can I use Speech Build in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cnemri/google-genai-skills --skill speech-build -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-build, .gemini/skills/speech-build, .github/skills/speech-build and .opencode/skills/speech-build in your project.

What does Speech Build need to run?

SKILL.md names no scripts, command-line tools or credentials: Speech Build is instructions for the agent only. Our summary lists: Python 3.

Does Speech Build access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Speech Build safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Speech Build use?

Speech Build is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speech Build use?

About 430 tokens (SKILL.md is roughly 1.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Speech Build?

Skills that share tags, products or a category with Speech Build: Voice AI Engine Development (aiskillstore/marketplace, 430 stars), Stepfun Asr (daymade/claude-code-skills, 1.4k stars), Deepgram (Anil-matcha/awesome-muse-connectors, 1.3k stars) and Groq Core Workflow B (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech Build?

cnemri (a GitHub user) maintains it in cnemri/google-genai-skills, which has 127 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on February 6, 2026.

Source: cnemri/google-genai-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.