Agent skill

Google Tts

by sanjay3290 in sanjay3290/ai-skills

Convert documents and text to audio using Google Cloud Text-to-Speech.

Apache-2.0Auto-check passedMedia & Creative

Install Google Tts

skills CLI
$ npx skills add sanjay3290/ai-skills --skill google-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sanjay3290/ai-skills google-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sanjay3290/ai-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/google-tts .claude/skills/google-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
google-tts
GitHub stars
431
Token cost
~883 tokens
SKILL.md length
217 words
Files
4 (incl. scripts)
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert documents and text to audio using Google Cloud Text-to-Speech.

  • Works in 4 steps: If user provides a file path, use… → Default voice: en-US-Neural2-D (male) or… → Generate: python… → …
  • The user wants to: narrate a document
  • SKILL.md covers Setup, Commands, Workflow and Reference
  • Runs Python scripts from its folder; calls python and pip; needs GOOGLE_TTS_API_KEY

What it does

Google Tts is an agent skill from sanjay3290/ai-skills. Convert documents and text to audio using Google Cloud Text-to-Speech. Use this skill when the user wants to: narrate a document, read aloud text, generate audio from a file, convert text to speech, create a recording of documentation or analysis, create a podcast from a document, or use Google TTS/text-to-speech. Trigger phrases: "read this aloud", "narrate this", "create a recording", "text to speech", "TTS", "convert to audio", "audio from document", "listen to this", "generate audio", "google tts", "create a…

Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/extract.py` and `scripts/google_tts.py`).

It sits in Media & Creative, covering Text to speech and voice. It works with Google Cloud and Python. The repository describes itself as: 24 cross-platform agent skills for Claude Code, Cursor, Codex & Gemini CLI — databases, messaging, research, TTS, DevOps, and Google Workspace. The licence is Apache-2.0.

When your agent uses it

  • The user wants to: narrate a document
  • Read aloud text
  • Generate audio from a file
  • Convert text to speech

Example prompts

  • “read this aloud”
  • “narrate this”
  • “create a recording”
  • “/google-tts”

Requirements

  • Python 3
  • A credential in GOOGLE_TTS_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. If user provides a file path, use --file. For generated content, write clean prose to /tmp/tts_input.md first.
  2. Default voice: en-US-Neural2-D (male) or en-US-Neural2-F (female). Use Neural2 for best quality/cost balance.
  3. Generate: python skills/google-tts/scripts/google_tts.py tts --file /tmp/tts_input.md --output ~/Downloads/recording.mp3
  4. Report file location and size. Default output to ~/Downloads/.

What it can do on your machine

Read from SKILL.md and the folder at commit 281d88d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_TTS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Google Tts loads about 883 tokens when it runs. Until then it costs about 135 tokens; SKILL.md has 217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~883

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sanjay3290/ai-skills at commit 281d88d, republished under its Apache-2.0 licence (© sanjay3290). 217 words, ~883 tokens.

Download SKILL.mdSave it as .claude/skills/google-tts/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
google-tts
description
Convert documents and text to audio using Google Cloud Text-to-Speech. Use this skill when the user wants to: narrate a document, read aloud text, generate audio from a file, convert text to speech, create a recording of documentation or analysis, create a podcast from a document, or use Google TTS/text-to-speech. Trigger phrases: "read this aloud", "narrate this", "create a recording", "text to speech", "TTS", "convert to audio", "audio from document", "listen to this", "generate audio", "google tts", "create a podcast".

Google Cloud Text-to-Speech

Converts text and documents into audio using Google Cloud TTS API. Supports Neural2, WaveNet, Studio, and Standard voices across 40+ languages.

Setup

API key via GOOGLE_TTS_API_KEY env var or skills/google-tts/config.json with {"api_key": "..."}. Requires ffmpeg for multi-chunk documents. Optional: pip install PyPDF2 python-docx for PDF/DOCX.

Commands

List Voices
bash
python skills/google-tts/scripts/google_tts.py voices --language en-US --type Neural2
python skills/google-tts/scripts/google_tts.py voices --json
Text-to-Speech
bash
# From text or document (PDF, DOCX, MD, TXT)
python skills/google-tts/scripts/google_tts.py tts --text "Hello world" --output ~/Downloads/hello.mp3
python skills/google-tts/scripts/google_tts.py tts --file /path/to/doc.pdf --output ~/Downloads/narration.mp3

# With voice, rate, pitch, encoding options
python skills/google-tts/scripts/google_tts.py tts --file doc.md --voice en-US-Neural2-F --rate 0.9 --encoding MP3 --output ~/Downloads/out.mp3
Podcast Generation

Takes a JSON script with alternating speakers, synthesizes each with a different voice.

json
[
  {"speaker": "host1", "text": "Welcome to our podcast!"},
  {"speaker": "host2", "text": "Thanks for having me..."}
]
bash
python skills/google-tts/scripts/google_tts.py podcast --script /tmp/script.json --output ~/Downloads/podcast.mp3
python skills/google-tts/scripts/google_tts.py podcast --script /tmp/script.json --voice1 en-US-Neural2-J --voice2 en-US-Neural2-H --rate 0.9 --output ~/Downloads/podcast.mp3

Workflow

Single-Voice Narration
  1. If user provides a file path, use --file. For generated content, write clean prose to /tmp/tts_input.md first.
  2. Default voice: en-US-Neural2-D (male) or en-US-Neural2-F (female). Use Neural2 for best quality/cost balance.
  3. Generate: python skills/google-tts/scripts/google_tts.py tts --file /tmp/tts_input.md --output ~/Downloads/recording.mp3
  4. Report file location and size. Default output to ~/Downloads/.
Podcast from Document
  1. Extract text: python skills/google-tts/scripts/extract.py /path/to/document.pdf
  2. Generate a two-host conversation script as JSON:
    • Natural discussion, not verbatim reading. Host 1 leads, Host 2 reacts/analyzes.
    • Include intro and outro. Vary turn lengths. Keep turns under 4000 chars.
  3. Write script to /tmp/podcast_script.json
  4. Generate: python skills/google-tts/scripts/google_tts.py podcast --script /tmp/podcast_script.json --output ~/Downloads/podcast.mp3
  5. Clean up temp files.

Reference

  • Recommended voice type: Neural2 (~$4/1M chars, high quality)
  • Speaking rate: 0.25-4.0 (0.85-0.95 good for technical content)
  • Pitch: -20.0 to 20.0 semitones
  • Encodings: MP3 (default), LINEAR16 (.wav), OGG_OPUS (.ogg)
  • API limit: 5000 bytes/request. Script auto-chunks at sentence boundaries.

© sanjay3290, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/google-tts of sanjay3290/ai-skills.

  • SKILL.md
  • .gitignore
  • scripts/extract.py
  • scripts/google_tts.py

Open the folder on GitHubat commit 281d88d

Compare with similar skills

Google Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Google Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Google Tts this skillsanjay3290/ai-skills431—~883Automated safety check: PassApache-2.0
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Floe GuardFloe-Labs/floe-guard357—~853Automated safety check: PassMIT
Book Video Factorybytec-ai/book-video-factory322—~1.4kAutomated safety check: NotesNone
Whiteboard Videognipbao/codex-whiteboard-video-skill326—~7.2kAutomated safety check: NotesMIT

Similar skills

  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • Floe Guard

    Floe-Labs/floe-guard

    Know what every AI call really costs — floe-guard meters STT + TTS + LLM + telephony per call (Pipecat, LiveKit — Python & TypeScript), keeps a live ledger of real spend, and hard-stops the next…

    357 GitHub stars~853 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Book Video Factory

    bytec-ai/book-video-factory

    通用的多账号图书短视频生产工作流。用于用户希望建立图书号项目目录、配置账号级片头/声音/BGM/视觉规范,或只提供一本书后依次完成资料研究、口播稿、分镜、图片、配音、字幕、预览与成片导出。适用于新建工作区、批量管理多个账号、继续已有单书任务和检查生产状态;不绑定特定研究、图片、TTS、转录或视频渲染供应商。

    322 GitHub stars~1.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    326 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Aivisspeech Engine Sentry Triage

    Aivis-Project/AivisSpeech-Engine

    AivisSpeech Engine の Sentry issue を調査し、修正すべきエンジン側の不具合と、入力値・ローカル環境・外部サービス由来のノイズを切り分けるためのスキルです。Sentry 側で既知ノイズを永続アーカイブする作業や、voicevoxengine/utility/sentryutility.py と関連テストを更新して既知ノイズを送信前に破棄する作業で使用します。

    181 GitHub stars~545 tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from sanjay3290/ai-skills

All 24 skills in this repo
  • Deep Research

    sanjay3290/ai-skills

    Execute autonomous multi-step research using Google Gemini Deep Research Agent.

    431 GitHub starsUsed in 9 repos~683 tokens
    Auto-check: notes
  • Imagen

    sanjay3290/ai-skills

    Generate images using Google Gemini's image generation capabilities.

    431 GitHub starsUsed in 6 repos~657 tokens
    Auto-check passed
  • Notebooklm

    sanjay3290/ai-skills

    Query and manage Google NotebookLM notebooks with persistent profile auth, source sync, batch/multi queries, and structured exports.

    431 GitHub stars~655 tokensUpdated 28 days ago
    Auto-check passed
  • Whatsapp

    sanjay3290/ai-skills

    Send and receive WhatsApp messages via the unofficial linked-device client pywhats (pip install pywhats) — pair with QR, send text/images, group chat, read receipts, presence/typing, and a…

    431 GitHub stars~1.1k tokensUpdated 28 days ago
    Auto-check passed
  • Postgres

    sanjay3290/ai-skills

    Execute read-only SQL queries against multiple PostgreSQL databases.

    431 GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Atlassian

    sanjay3290/ai-skills

    Manage Jira issues and Confluence wiki pages in Atlassian Cloud.

    431 GitHub stars~1.6k tokensUpdated 28 days ago
    Auto-check passed

Questions about Google Tts

What does Google Tts do?

Convert documents and text to audio using Google Cloud Text-to-Speech. Google Tts is an agent skill from sanjay3290/ai-skills. Convert documents and text to audio using Google Cloud Text-to-Speech.

When should I use Google Tts?

Google Tts fits situations like: the user wants to: narrate a document; read aloud text; generate audio from a file; convert text to speech.

How do I install Google Tts in Claude Code?

Run `npx skills add sanjay3290/ai-skills --skill google-tts -a claude-code`. Or copy the skill folder (skills/google-tts in sanjay3290/ai-skills) into .claude/skills/google-tts in your project. Claude Code loads it when a task matches its description.

How do I install Google Tts in Codex?

Run `npx skills add sanjay3290/ai-skills --skill google-tts -a codex`. Or copy the skill folder (skills/google-tts in sanjay3290/ai-skills) into .agents/skills/google-tts in your project. Codex loads it when a task matches its description.

Can I use Google Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sanjay3290/ai-skills --skill google-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/google-tts, .gemini/skills/google-tts, .github/skills/google-tts and .opencode/skills/google-tts in your project.

What does Google Tts need to run?

Going by SKILL.md and its folder, Google Tts needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GOOGLE_TTS_API_KEY. Our summary lists: Python 3; A credential in GOOGLE_TTS_API_KEY.

Does Google Tts access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Google Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Google Tts use?

Google Tts is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Google Tts use?

About 883 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Google Tts?

Skills that share tags, products or a category with Google Tts: Vox Director (Alisa0808/vox-director, 2.2k stars), Edu Math Video (wy51ai/edulab, 1.4k stars), Floe Guard (Floe-Labs/floe-guard, 357 stars) and Book Video Factory (bytec-ai/book-video-factory, 322 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Google Tts?

sanjay3290 (a GitHub user) maintains it in sanjay3290/ai-skills, which has 431 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on September 10, 2026.

Source: sanjay3290/ai-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.