Agent skill

AI Text To Speech

by godot-fun in godot-fun/gai

Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).

MITAuto-check passedMedia & Creative

Install AI Text To Speech

skills CLI
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install godot-fun/gai ai-text-to-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .claude/skills/ai-text-to-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-text-to-speech
GitHub stars
184
Token cost
~2k tokens
SKILL.md length
548 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).

  • Works in 5 steps: Python 3.11 + uv → Clone IndexTTS → Install deps with uv (required) → …
  • The user wants TTS
  • SKILL.md covers Rules, Setup (first run), Quick Start and Emotion control (optional), plus 5 more sections
  • Calls uv, python and git; reaches github.com and hf-mirror.com

What it does

AI Text To Speech is an agent skill from godot-fun/gai. Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts). Synthesizes speech from text using a user-provided reference audio for timbre. Use when the user wants TTS, text-to-speech, voice cloning, voice clone, IndexTTS, IndexTTS2, or generating narration/voice lines from a reference WAV.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with Python. The repository describes itself as: A lightweight AI agent and skill workflow framework built with Godot. The licence is MIT.

When your agent uses it

  • The user wants TTS
  • Generating narration/voice lines from a reference WAV

Example prompts

  • “/ai-text-to-speech”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Python 3.11 + uv
  2. Clone IndexTTS
  3. Install deps with uv (required)
  4. Download IndexTTS-2 checkpoints
  5. Register manifest

What it can do on your machine

Read from SKILL.md and the folder at commit a98225b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • python
    • git
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • hf-mirror.com
    • mirrors.aliyun.com

    Also links to:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Text To Speech loads about 2k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 548 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from godot-fun/gai at commit a98225b, republished under its MIT licence (© godot-fun). 548 words, ~2,015 tokens.

Download SKILL.mdSave it as .claude/skills/ai-text-to-speech/SKILL.md (or your agent's skills folder).
name
ai-text-to-speech
description
Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts). Synthesizes speech from text using a user-provided reference audio for timbre. Use when the user wants TTS, text-to-speech, voice cloning, voice clone, IndexTTS, IndexTTS2, or generating narration/voice lines from a reference WAV.

AI Text-to-Speech (IndexTTS2)

Clone a speaker from a reference audio, then synthesize speech from text with IndexTTS2.

Rules

When this skill applies, read and follow skill-dependency-manager — run scripts as documented, install missing tools into .dependency/.

  • Run tts.py at .ai/ai-text-to-speech/tts.py through the index-tts manifest entry (.dependency/index-tts/.venv/). Never use host python, py, python3, or any interpreter outside .dependency/.
  • Do not hand-write IndexTTS Python snippets or uv run webui.py for synthesis — use the bundled script.
  • IndexTTS requires uv for install (pip/conda are unsupported upstream). Python must be >=3.10,<3.12 (use python-3.11).
  • populated: false for index-tts (or missing models) is not a reason to skip. Install / download first, set populated: true, retry the same command.
  • Never overwrite sources. Pass the user's real voice path and text; write only to -o / --output.

Setup (first run)

From project root.

1. Python 3.11 + uv

Ensure python-3.11 and uv are populated under .dependency/ (see skill-dependency-manager). Register:

json
"python-3.11": {
  "populated": true,
  "bin": ".dependency/python-3.11/python.exe"
},
"uv": {
  "populated": true,
  "bin": ".dependency/uv/uv.exe"
}

Use python / uv (no .exe) on Unix.

2. Clone IndexTTS
bash
git clone https://github.com/index-tts/index-tts.git .dependency/index-tts
cd .dependency/index-tts
git lfs install
git lfs pull
3. Install deps with uv (required)
bash
# from .dependency/index-tts
# Windows: skip deepspeed extras if install fails
.dependency/uv/uv.exe sync --extra webui

Slow PyPI (China mirrors):

bash
.dependency/uv/uv.exe sync --extra webui --default-index "https://mirrors.aliyun.com/pypi/simple"

CUDA Toolkit 12.8+ is needed for GPU. CPU works but is slow.

4. Download IndexTTS-2 checkpoints
bash
cd .dependency/index-tts
.dependency/uv/uv.exe tool install "huggingface-hub[cli,hf_xet]"
hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints

Or ModelScope:

bash
.dependency/uv/uv.exe tool install "modelscope"
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpoints

If HuggingFace is slow: HF_ENDPOINT=https://hf-mirror.com (Unix) / $env:HF_ENDPOINT="https://hf-mirror.com" (PowerShell).

5. Register manifest
json
"index-tts": {
  "populated": true,
  "bin": ".dependency/index-tts/.venv/Scripts/python.exe"
}

Use .dependency/index-tts/.venv/bin/python on Unix. Confirm checkpoints/config.yaml exists before synthesizing.

Quick Start

Voice reference + text → WAV (default output: <voice-dir>/ai-text-to-speech/<voice-stem>.wav):

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --text "Hello, welcome to this world."
# → audio/voice/ai-text-to-speech/ref.wav

Explicit output path:

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --text "Hello, this is a test." \
  --output audio/voice/ai-text-to-speech/hello.wav

Directory (writes <voice-stem>.wav inside, e.g. audio/voice/ai-text-to-speech/ref.wav):

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py --voice audio/voice/ref.wav --text "Hello, this is a test." --output audio/voice/ai-text-to-speech

Long script from a UTF-8 text file:

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --text-file script/lines/intro.txt \
  --output audio/voice/ai-text-to-speech/intro.wav

FP16 (faster, less VRAM):

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --text "Testing half-precision inference." \
  --fp16

Emotion control (optional)

ModeFlagsNotes
Emotion reference audio--emotion-audio path.wavSeparate clip for emotion; timbre still from --voice
Emotion weight--emotion-weight 0.6Maps to emo_alpha (0.0–1.0, default 1.0)
Emotion from text--emotion-from-textInfer emotion from synthesis text; prefer --emotion-weight ≈ 0.6
Emotion description--emotion-text "..."Natural-language emotion; implies text emotion mode
Emotion vector--emotion-vector 0,0,0.8,0,0,0,0,08 floats: happy, angry, sad, afraid, disgusted, melancholic, surprised, calm
bash
# Emotion reference audio
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --emotion-audio audio/voice/emo_sad.wav \
  --emotion-weight 0.9 \
  --text "The inn has gone rotten and started auctioning off rooms." \
  --output audio/voice/ai-text-to-speech/sad_line.wav

# Emotion description text
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
  --voice audio/voice/ref.wav \
  --emotion-text "afraid, tense" \
  --emotion-weight 0.6 \
  --text "Hide quickly! He is coming!" \
  --output audio/voice/ai-text-to-speech/afraid_line.wav

Do not combine --emotion-audio, --emotion-vector, and --emotion-text / --emotion-from-text in conflicting ways — pick one emotion source.

Show full SKILL.md (235 more words)Show less

Defaults

OptionDefaultNotes
Output<voice-dir>/ai-text-to-speech/<voice-stem>.wav--output file uses that name; --output directory uses the voice file's stem
Model.dependency/index-tts/checkpointsIndexTTS-2
--fp16offEnable on GPU when VRAM is tight
--emotion-weight1.0Lower (~0.6) for text emotion modes
OverwriteoffPass --force to replace an existing output

Agent workflow

  1. Confirm inputs — need a clear reference voice WAV/MP3 and the text (or --text-file). Ask if either is missing.
  2. Use the user's real paths — do not copy voice files into the repo unless asked.
  3. Trial first — synthesize one short line, play/inspect before long scripts.
  4. Reference audio tips — clean, single-speaker, little noise; a few seconds of clear speech works best.
  5. Missing install — follow Setup; register index-tts in manifest.json; retry the same command.
  6. GPU — prefer CUDA + --fp16 for speed; CPU is acceptable for short tests only.
  7. Revert — delete files under ai-text-to-speech/; sources are never modified.

Troubleshooting

IssueFix
index-tts not populatedClone + uv sync + download checkpoints; update manifest
checkpoints/config.yaml missingRe-run hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints
CUDA / torch errorsInstall CUDA 12.8+; or run on CPU (slow)
OOM / VRAMPass --fp16; shorten text; close other GPU apps
Slow HuggingFaceSet HF_ENDPOINT=https://hf-mirror.com; or use ModelScope
uv sync / DeepSpeed fail on WindowsUse uv sync --extra webui without deepspeed
Unnatural emotionLower --emotion-weight to ~0.6; try a clearer --emotion-audio

CLI

Copy-paste commands: cli/ai-text-to-speech.md

© godot-fun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ai-text-to-speech of godot-fun/gai.

Open the folder on GitHubat commit a98225b

Compare with similar skills

AI Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Text To Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Text To Speech this skillgodot-fun/gai184—~2kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Book Video Factorybytec-ai/book-video-factory321—~1.4kAutomated safety check: NotesNone
Narrator AI CLINarratorAI-Studio/narrator-ai-cli137—~3.7kAutomated safety check: PassMIT
Iflytek Voiceclone Ttsiflytek/iFly-Skills209—~2.1kAutomated safety check: PassApache-2.0
U2 TtsLeoYeAI/openclaw-master-skills2.2k—~4.7kAutomated safety check: NotesMIT

Similar skills

  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Book Video Factory

    bytec-ai/book-video-factory

    通用的多账号图书短视频生产工作流。用于用户希望建立图书号项目目录、配置账号级片头/声音/BGM/视觉规范,或只提供一本书后依次完成资料研究、口播稿、分镜、图片、配音、字幕、预览与成片导出。适用于新建工作区、批量管理多个账号、继续已有单书任务和检查生产状态;不绑定特定研究、图片、TTS、转录或视频渲染供应商。

    321 GitHub stars~1.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Narrator AI CLI

    NarratorAI-Studio/narrator-ai-cli

    Create AI-narrated film/drama commentary videos via CLI. An agent skill from NarratorAI-Studio/narrator-ai-cli.

    137 GitHub stars~3.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Iflytek Voiceclone Tts

    iflytek/iFly-Skills

    A skill your agent uses when user asks to clone a voice, train a custom voice model, or synthesize speech with a cloned voice.

    209 GitHub stars~2.1k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • U2 Tts

    LeoYeAI/openclaw-master-skills

    Text-to-speech conversion using UniSound's TTS WebSocket API for generating high-quality Chinese Mandarin audio from text.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings

More from godot-fun/gai

All 36 skills in this repo
  • Audio Denoise

    godot-fun/gai

    Reduces background noise in a single audio file using FFmpeg afftdn.

    184 GitHub stars~465 tokensUpdated today
    Auto-check passed
  • Audio Fade

    godot-fun/gai

    Applies fade-in and fade-out at the start and end of a single audio file using FFmpeg.

    184 GitHub stars~652 tokensUpdated today
    Auto-check passed
  • Normalizes a single audio file to consistent LUFS loudness with true-peak limiting using FFmpeg.

    184 GitHub stars~684 tokensUpdated today
    Auto-check passed
  • Standardizes a single audio file to 44100 or 48000 Hz and exports 16-bit PCM WAV using FFmpeg.

    184 GitHub stars~638 tokensUpdated today
    Auto-check passed
  • Audio Split

    godot-fun/gai

    Splits a single audio file into two segments (part 1 before the split point, part 2 after) using FFmpeg.

    184 GitHub stars~554 tokensUpdated today
    Auto-check passed
  • Audio To Ogg

    godot-fun/gai

    Converts a single audio file to OGG Vorbis using FFmpeg for Godot-ready compressed assets.

    184 GitHub stars~766 tokensUpdated today
    Auto-check passed

Works with

Questions about AI Text To Speech

What does AI Text To Speech do?

Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts). AI Text To Speech is an agent skill from godot-fun/gai. Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).

When should I use AI Text To Speech?

AI Text To Speech fits situations like: the user wants TTS; generating narration/voice lines from a reference WAV.

How do I install AI Text To Speech in Claude Code?

Run `npx skills add godot-fun/gai --skill ai-text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/ai-text-to-speech in godot-fun/gai) into .claude/skills/ai-text-to-speech in your project. Claude Code loads it when a task matches its description.

How do I install AI Text To Speech in Codex?

Run `npx skills add godot-fun/gai --skill ai-text-to-speech -a codex`. Or copy the skill folder (.agents/skills/ai-text-to-speech in godot-fun/gai) into .agents/skills/ai-text-to-speech in your project. Codex loads it when a task matches its description.

Can I use AI Text To Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add godot-fun/gai --skill ai-text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-text-to-speech, .gemini/skills/ai-text-to-speech, .github/skills/ai-text-to-speech and .opencode/skills/ai-text-to-speech in your project.

What does AI Text To Speech need to run?

Going by SKILL.md and its folder, AI Text To Speech needs the command-line tools its instructions call (uv, python, git and hf). Our summary lists: Python 3.

Does AI Text To Speech access the network?

SKILL.md names 4 domains. In commands or code: github.com, hf-mirror.com and mirrors.aliyun.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is AI Text To Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Text To Speech use?

AI Text To Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Text To Speech use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Text To Speech?

Skills that share tags, products or a category with AI Text To Speech: Edu Math Video (wy51ai/edulab, 1.4k stars), Book Video Factory (bytec-ai/book-video-factory, 321 stars), Narrator AI CLI (NarratorAI-Studio/narrator-ai-cli, 137 stars) and Iflytek Voiceclone Tts (iflytek/iFly-Skills, 209 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Text To Speech?

godot-fun (a GitHub organization) maintains it in godot-fun/gai, which has 184 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 11, 2026.

Source: godot-fun/gai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.