Gemini Audio
einverne/dotfiles
Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.
Generates spoken MP3 audio from text or Markdown with Gemini TTS.
$ npx skills add iurysza/module-graph --skill gemini-tts -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install iurysza/module-graph gemini-tts --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/gemini-tts .claude/skills/gemini-tts && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .claude/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-ttsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add iurysza/module-graph --skill gemini-tts -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install iurysza/module-graph gemini-tts --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/gemini-tts .agents/skills/gemini-tts && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .agents/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iurysza/module-graph --skill gemini-tts -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install iurysza/module-graph gemini-tts --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/gemini-tts .cursor/skills/gemini-tts && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .cursor/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/iurysza/module-graph.git --path .agents/skills/gemini-tts--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add iurysza/module-graph --skill gemini-tts -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install iurysza/module-graph gemini-tts --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/gemini-tts .gemini/skills/gemini-tts && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .gemini/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install iurysza/module-graph gemini-ttsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add iurysza/module-graph --skill gemini-tts -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/gemini-tts .github/skills/gemini-tts && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .github/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iurysza/module-graph --skill gemini-tts -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install iurysza/module-graph gemini-tts --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iurysza/module-graph.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/gemini-tts .opencode/skills/gemini-tts && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gemini-tts" agent skill from https://github.com/iurysza/module-graph/tree/main/.agents/skills/gemini-tts into .opencode/skills/gemini-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gemini-tts", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gemini-ttsGenerates spoken MP3 audio from text or Markdown with Gemini TTS.
Gemini Tts is an agent skill from iurysza/module-graph. Generates spoken MP3 audio from text or Markdown with Gemini TTS. Use for narration, accessibility, voice previews, or reading documents aloud.
Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/generate_tts.py`, `scripts/test_generate_tts.py` and `templates.json`). Compatibility notes: Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.
It sits in Media & Creative, covering Text to speech and voice. It works with Google Gemini. The repository describes itself as: A Gradle Plugin for visualizing your project's structure, powered by mermaidjs. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 15b0135. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GOOGLE_API_KEYOPENCODE_GOOGLE_API_KEYGEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv.
From compatibility in the SKILL.md frontmatter.
Gemini Tts loads about 968 tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 343 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from iurysza/module-graph at commit 15b0135, republished under its MIT licence (© iurysza). 343 words, ~968 tokens.
.claude/skills/gemini-tts/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Generate an MP3 from inline text or a UTF-8 text/Markdown file with the bundled script.
For any unqualified TTS request, use the bundled natural-tech-conference default. The CLI applies it automatically when --template is omitted:
Algenib (gravelly, lower pitch)1.3225 (15% faster than 1.15)Keep this default unless the user explicitly requests another voice, delivery style, accent, template, or speed.
Resolve paths relative to this SKILL.md; do not assume a particular install directory.
python3 -m pip install -r <skill-directory>/requirements.txt
export GEMINI_API_KEY='...'GOOGLE_API_KEY and OPENCODE_GOOGLE_API_KEY are accepted as fallbacks. Set GEMINI_TTS_MODEL to override the default model.
Confirm or infer:
For long input, report the chunk count before making paid API calls. Ask for confirmation when the request is unexpectedly large or the user has not clearly approved generation.
python3 <skill-directory>/scripts/generate_tts.py --list-voices
python3 <skill-directory>/scripts/generate_tts.py --list-templates
python3 <skill-directory>/scripts/generate_tts.py --show-template natural-tech-conferenceBundled templates include the default natural-tech-conference plus mystery-narrator, newscaster, whisper, empathetic, deadpan, promo-hype, and podcast-newsletter.
Using the default delivery:
python3 <skill-directory>/scripts/generate_tts.py \
--text 'Explain this clearly and naturally.' \
--output ./narration.mp3From a file with an explicit alternate template:
python3 <skill-directory>/scripts/generate_tts.py \
--file ./article.md \
--template podcast-newsletter \
--output ./article.mp3 \
--playCustomize delivery when needed:
python3 <skill-directory>/scripts/generate_tts.py \
--file ./script.txt \
--voice Kore \
--profile 'Calm technical narrator' \
--scene 'A quiet recording booth' \
--notes 'Clear diction, measured pace, neutral accent' \
--speed 1.1 \
--output ./script.mp3Explicit CLI flags override an explicit template. An explicit template overrides the catalog default. If the catalog has no default, the hard fallback remains Orus at speed 1.0.
--max-workers N: concurrent chunk requests; default 1--requests-per-minute N: request throttle; default 8; 0 disables it--allow-partial: write an MP3 despite failed chunks; avoid unless the user accepts missing audioEnvironment equivalents are GEMINI_TTS_MAX_WORKERS and GEMINI_TTS_RPM.
After generation:
Never print API keys or include them in command examples, logs, or output files.
© iurysza, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts) in .agents/skills/gemini-tts of iurysza/module-graph.
Open the folder on GitHubat commit 15b0135
Gemini Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gemini Tts this skilliurysza/module-graph | 420 | — | ~968 | Automated safety check: Pass | MIT | |
| Gemini Audioeinverne/dotfiles | 121 | — | ~2k | Automated safety check: Notes | MIT | |
| Blog AudioAgriciDaniel/claude-blog | 2.3k | 1 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Video Analyzermikefutia/claude-vision | 101 | — | ~747 | Automated safety check: Notes | None | |
| Fal AImikeOnBreeze/cc-crossbeam | 293 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Fal AI Mediaaffaan-m/ECC | 276k | 4 repos | ~1.9k | Automated safety check: Pass | MIT |
einverne/dotfiles
Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.
AgriciDaniel/claude-blog
Generate audio narration of blog posts using Google Gemini TTS.
mikefutia/claude-vision
Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…
mikeOnBreeze/cc-crossbeam
This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.
affaan-m/ECC
Unified media generation via fal.ai MCP — image, video, and audio.
Anil-matcha/awesome-muse-connectors
Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing.
iurysza/module-graph
Transcribes local audio into Markdown with Gemini 3.5 Transcribe, including speaker labels and provider timestamps.
iurysza/module-graph
Generates or edits raster images through OpenAI's Image API.
iurysza/module-graph
Audits installed Agent Skills for duplicates, unused candidates, loaded roots, oversized descriptions, and prompt cost.
iurysza/module-graph
Explores and validates a feature, component, workflow, or behavior change before implementation.
iurysza/module-graph
Maintains project domain language, context maps, diagrams, and architectural decisions under ai-artifacts.
iurysza/module-graph
Rewrites technical review replies as short, natural Slack or PR messages between engineers.
Works with
Categories
Generates spoken MP3 audio from text or Markdown with Gemini TTS. Gemini Tts is an agent skill from iurysza/module-graph. Generates spoken MP3 audio from text or Markdown with Gemini TTS.
Gemini Tts fits situations like: reading documents aloud; tasks that involve Text to speech and voice.
Run `npx skills add iurysza/module-graph --skill gemini-tts -a claude-code`. Or copy the skill folder (.agents/skills/gemini-tts in iurysza/module-graph) into .claude/skills/gemini-tts in your project. Claude Code loads it when a task matches its description.
Run `npx skills add iurysza/module-graph --skill gemini-tts -a codex`. Or copy the skill folder (.agents/skills/gemini-tts in iurysza/module-graph) into .agents/skills/gemini-tts in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iurysza/module-graph --skill gemini-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-tts, .gemini/skills/gemini-tts, .github/skills/gemini-tts and .opencode/skills/gemini-tts in your project.
Going by SKILL.md and its folder, Gemini Tts needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named GOOGLE_API_KEY, OPENCODE_GOOGLE_API_KEY and GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in GOOGLE_API_KEY. Compatibility (from SKILL.md): Requires Python 3.10+, google-genai 1.65+, ffmpeg, and a Gemini API key. Optional playback needs afplay, ffplay, or mpv..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gemini Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 968 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gemini Tts: Gemini Audio (einverne/dotfiles, 121 stars), Blog Audio (AgriciDaniel/claude-blog, 2.3k stars), Video Analyzer (mikefutia/claude-vision, 101 stars) and Fal AI (mikeOnBreeze/cc-crossbeam, 293 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
iurysza (a GitHub user) maintains it in iurysza/module-graph, which has 420 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 8, 2026.
Source: iurysza/module-graph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.