Agent skill

Gemini Tts

by iurysza in iurysza/module-graph

Generates spoken MP3 audio from text or Markdown with Gemini TTS.

MITAuto-check passedMedia & Creative

Install Gemini Tts

skills CLI
$ npx skills add iurysza/module-graph --skill gemini-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install iurysza/module-graph gemini-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/gemini-tts .claude/skills/gemini-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-tts
GitHub stars
420
Token cost
~968 tokens
SKILL.md length
343 words
Files
5 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Generates spoken MP3 audio from text or Markdown with Gemini TTS.

  • Works in 4 steps: Confirm the command exited successfully. → Confirm the MP3 exists and is non-empty. → Report the exact output path. → …
  • Reading documents aloud
  • SKILL.md covers Default delivery, Setup, Before generation and Discover voices and templates, plus 3 more sections
  • Runs Python scripts from its folder; calls python3; needs GOOGLE_API_KEY and OPENCODE_GOOGLE_API_KEY

What it does

Gemini Tts is an agent skill from iurysza/module-graph. Generates spoken MP3 audio from text or Markdown with Gemini TTS. Use for narration, accessibility, voice previews, or reading documents aloud.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/generate_tts.py`, `scripts/test_generate_tts.py` and `templates.json`). Compatibility notes: Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.

It sits in Media & Creative, covering Text to speech and voice. It works with Google Gemini. The repository describes itself as: A Gradle Plugin for visualizing your project's structure, powered by mermaidjs. The licence is MIT.

When your agent uses it

  • Reading documents aloud
  • Tasks that involve Text to speech and voice

Example prompts

  • “Use the gemini-tts skill to generate spoken MP3 audio from text or Markdown with Gemini TTS”
  • “/gemini-tts”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • A credential in GOOGLE_API_KEY
  • Compatibility (from SKILL.md): Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the command exited successfully.
  2. Confirm the MP3 exists and is non-empty.
  3. Report the exact output path.
  4. If playback was requested, report when no supported player is installed.

What it can do on your machine

Read from SKILL.md and the folder at commit 15b0135. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY
    • OPENCODE_GOOGLE_API_KEY
    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.

    From compatibility in the SKILL.md frontmatter.

Context cost

Gemini Tts loads about 968 tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 343 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~968

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from iurysza/module-graph at commit 15b0135, republished under its MIT licence (© iurysza). 343 words, ~968 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-tts/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
gemini-tts
description
Generates spoken MP3 audio from text or Markdown with Gemini TTS. Use for narration, accessibility, voice previews, or reading documents aloud.
compatibility
Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.
metadata.category
visual-media

Gemini TTS

Generate an MP3 from inline text or a UTF-8 text/Markdown file with the bundled script.

Default delivery

For any unqualified TTS request, use the bundled natural-tech-conference default. The CLI applies it automatically when --template is omitted:

  • voice: Algenib (gravelly, lower pitch)
  • profile: experienced software engineer speaking normally at a small San Francisco technical conference
  • scene: a mid-sized breakout room, explaining the topic plainly to engineering peers
  • delivery: neutral conversational American English, low emotional range, ordinary sentence stress, modest pauses, and no narrator, marketer, keynote, radio-host, or audiobook performance
  • speed: 1.3225 (15% faster than 1.15)

Keep this default unless the user explicitly requests another voice, delivery style, accent, template, or speed.

Setup

Resolve paths relative to this SKILL.md; do not assume a particular install directory.

bash
python3 -m pip install -r <skill-directory>/requirements.txt
export GEMINI_API_KEY='...'

GOOGLE_API_KEY and OPENCODE_GOOGLE_API_KEY are accepted as fallbacks. Set GEMINI_TTS_MODEL to override the default model.

Before generation

Confirm or infer:

  • source text or file
  • output path
  • any explicit override to the default delivery
  • whether playback is wanted

For long input, report the chunk count before making paid API calls. Ask for confirmation when the request is unexpectedly large or the user has not clearly approved generation.

Discover voices and templates

bash
python3 <skill-directory>/scripts/generate_tts.py --list-voices
python3 <skill-directory>/scripts/generate_tts.py --list-templates
python3 <skill-directory>/scripts/generate_tts.py --show-template natural-tech-conference

Bundled templates include the default natural-tech-conference plus mystery-narrator, newscaster, whisper, empathetic, deadpan, promo-hype, and podcast-newsletter.

Generate audio

Using the default delivery:

bash
python3 <skill-directory>/scripts/generate_tts.py \
  --text 'Explain this clearly and naturally.' \
  --output ./narration.mp3

From a file with an explicit alternate template:

bash
python3 <skill-directory>/scripts/generate_tts.py \
  --file ./article.md \
  --template podcast-newsletter \
  --output ./article.mp3 \
  --play

Customize delivery when needed:

bash
python3 <skill-directory>/scripts/generate_tts.py \
  --file ./script.txt \
  --voice Kore \
  --profile 'Calm technical narrator' \
  --scene 'A quiet recording booth' \
  --notes 'Clear diction, measured pace, neutral accent' \
  --speed 1.1 \
  --output ./script.mp3

Explicit CLI flags override an explicit template. An explicit template overrides the catalog default. If the catalog has no default, the hard fallback remains Orus at speed 1.0.

Reliability controls

  • --max-workers N: concurrent chunk requests; default 1
  • --requests-per-minute N: request throttle; default 8; 0 disables it
  • --allow-partial: write an MP3 despite failed chunks; avoid unless the user accepts missing audio

Environment equivalents are GEMINI_TTS_MAX_WORKERS and GEMINI_TTS_RPM.

Verification

After generation:

  1. Confirm the command exited successfully.
  2. Confirm the MP3 exists and is non-empty.
  3. Report the exact output path.
  4. If playback was requested, report when no supported player is installed.

Never print API keys or include them in command examples, logs, or output files.

© iurysza, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in .agents/skills/gemini-tts of iurysza/module-graph.

  • SKILL.md
  • requirements.txt
  • scripts/generate_tts.py
  • scripts/test_generate_tts.py
  • templates.json

Open the folder on GitHubat commit 15b0135

Compare with similar skills

Gemini Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Tts this skilliurysza/module-graph420—~968Automated safety check: PassMIT
Gemini Audioeinverne/dotfiles121—~2kAutomated safety check: NotesMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Video Analyzermikefutia/claude-vision101—~747Automated safety check: NotesNone
Fal AImikeOnBreeze/cc-crossbeam293—~1.9kAutomated safety check: NotesMIT
Fal AI Mediaaffaan-m/ECC276k4 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub stars~2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Video Analyzer

    mikefutia/claude-vision

    Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…

    101 GitHub stars~747 tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Fal AI

    mikeOnBreeze/cc-crossbeam

    This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

    293 GitHub stars~1.9k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes
  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    276k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Gemini

    Anil-matcha/awesome-muse-connectors

    Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.

    1.3k GitHub stars~778 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from iurysza/module-graph

All 16 skills in this repo
  • Audio Transcribe

    iurysza/module-graph

    Transcribes local audio into Markdown with Gemini 3.5 Transcribe, including speaker labels and provider timestamps.

    420 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Chatgpt Imagegen

    iurysza/module-graph

    Generates or edits raster images through OpenAI's Image API.

    420 GitHub stars~726 tokensUpdated 2 days ago
    Auto-check passed
  • Skill Cleaner

    iurysza/module-graph

    Audits installed Agent Skills for duplicates, unused candidates, loaded roots, oversized descriptions, and prompt cost.

    420 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Brainstorming

    iurysza/module-graph

    Explores and validates a feature, component, workflow, or behavior change before implementation.

    420 GitHub stars~723 tokensUpdated 2 days ago
    Auto-check passed
  • Domain Modeling

    iurysza/module-graph

    Maintains project domain language, context maps, diagrams, and architectural decisions under ai-artifacts.

    420 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Reply Bro

    iurysza/module-graph

    Rewrites technical review replies as short, natural Slack or PR messages between engineers.

    420 GitHub stars~501 tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Gemini Tts

What does Gemini Tts do?

Generates spoken MP3 audio from text or Markdown with Gemini TTS. Gemini Tts is an agent skill from iurysza/module-graph. Generates spoken MP3 audio from text or Markdown with Gemini TTS.

When should I use Gemini Tts?

Gemini Tts fits situations like: reading documents aloud; tasks that involve Text to speech and voice.

How do I install Gemini Tts in Claude Code?

Run `npx skills add iurysza/module-graph --skill gemini-tts -a claude-code`. Or copy the skill folder (.agents/skills/gemini-tts in iurysza/module-graph) into .claude/skills/gemini-tts in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Tts in Codex?

Run `npx skills add iurysza/module-graph --skill gemini-tts -a codex`. Or copy the skill folder (.agents/skills/gemini-tts in iurysza/module-graph) into .agents/skills/gemini-tts in your project. Codex loads it when a task matches its description.

Can I use Gemini Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iurysza/module-graph --skill gemini-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-tts, .gemini/skills/gemini-tts, .github/skills/gemini-tts and .opencode/skills/gemini-tts in your project.

What does Gemini Tts need to run?

Going by SKILL.md and its folder, Gemini Tts needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named GOOGLE_API_KEY, OPENCODE_GOOGLE_API_KEY and GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in GOOGLE_API_KEY. Compatibility (from SKILL.md): Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv..

Does Gemini Tts access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gemini Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Tts use?

Gemini Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Tts use?

About 968 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini Tts?

Skills that share tags, products or a category with Gemini Tts: Gemini Audio (einverne/dotfiles, 121 stars), Blog Audio (AgriciDaniel/claude-blog, 2.3k stars), Video Analyzer (mikefutia/claude-vision, 101 stars) and Fal AI (mikeOnBreeze/cc-crossbeam, 293 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Tts?

iurysza (a GitHub user) maintains it in iurysza/module-graph, which has 420 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 8, 2026.

Source: iurysza/module-graph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.