Edu Math Video
wy51ai/edulab
A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…
Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install godot-fun/gai ai-text-to-speech --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .claude/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .claude/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speechType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install godot-fun/gai ai-text-to-speech --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .agents/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .agents/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install godot-fun/gai ai-text-to-speech --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .cursor/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .cursor/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/godot-fun/gai.git --path .agents/skills/ai-text-to-speech--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install godot-fun/gai ai-text-to-speech --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .gemini/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .gemini/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install godot-fun/gai ai-text-to-speechInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .github/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .github/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add godot-fun/gai --skill ai-text-to-speech -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install godot-fun/gai ai-text-to-speech --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ai-text-to-speech .opencode/skills/ai-text-to-speech && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-text-to-speech" agent skill from https://github.com/godot-fun/gai/tree/main/.agents/skills/ai-text-to-speech into .opencode/skills/ai-text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-text-to-speech", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-text-to-speechZero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).
AI Text To Speech is an agent skill from godot-fun/gai. Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts). Synthesizes speech from text using a user-provided reference audio for timbre. Use when the user wants TTS, text-to-speech, voice cloning, voice clone, IndexTTS, IndexTTS2, or generating narration/voice lines from a reference WAV.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Text to speech and voice. It works with Python. The repository describes itself as: A lightweight AI agent and skill workflow framework built with Godot. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a98225b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpythongithfFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comhf-mirror.commirrors.aliyun.comAlso links to:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Text To Speech loads about 2k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 548 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from godot-fun/gai at commit a98225b, republished under its MIT licence (© godot-fun). 548 words, ~2,015 tokens.
.claude/skills/ai-text-to-speech/SKILL.md (or your agent's skills folder).Clone a speaker from a reference audio, then synthesize speech from text with IndexTTS2.
When this skill applies, read and follow skill-dependency-manager — run scripts as documented, install missing tools into .dependency/.
tts.py at .ai/ai-text-to-speech/tts.py through the index-tts manifest entry (.dependency/index-tts/.venv/). Never use host python, py, python3, or any interpreter outside .dependency/.uv run webui.py for synthesis — use the bundled script.uv for install (pip/conda are unsupported upstream). Python must be >=3.10,<3.12 (use python-3.11).populated: false for index-tts (or missing models) is not a reason to skip. Install / download first, set populated: true, retry the same command.-o / --output.From project root.
Ensure python-3.11 and uv are populated under .dependency/ (see skill-dependency-manager). Register:
"python-3.11": {
"populated": true,
"bin": ".dependency/python-3.11/python.exe"
},
"uv": {
"populated": true,
"bin": ".dependency/uv/uv.exe"
}Use python / uv (no .exe) on Unix.
git clone https://github.com/index-tts/index-tts.git .dependency/index-tts
cd .dependency/index-tts
git lfs install
git lfs pull# from .dependency/index-tts
# Windows: skip deepspeed extras if install fails
.dependency/uv/uv.exe sync --extra webuiSlow PyPI (China mirrors):
.dependency/uv/uv.exe sync --extra webui --default-index "https://mirrors.aliyun.com/pypi/simple"CUDA Toolkit 12.8+ is needed for GPU. CPU works but is slow.
cd .dependency/index-tts
.dependency/uv/uv.exe tool install "huggingface-hub[cli,hf_xet]"
hf download IndexTeam/IndexTTS-2 --local-dir=checkpointsOr ModelScope:
.dependency/uv/uv.exe tool install "modelscope"
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpointsIf HuggingFace is slow: HF_ENDPOINT=https://hf-mirror.com (Unix) / $env:HF_ENDPOINT="https://hf-mirror.com" (PowerShell).
"index-tts": {
"populated": true,
"bin": ".dependency/index-tts/.venv/Scripts/python.exe"
}Use .dependency/index-tts/.venv/bin/python on Unix. Confirm checkpoints/config.yaml exists before synthesizing.
Voice reference + text → WAV (default output: <voice-dir>/ai-text-to-speech/<voice-stem>.wav):
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--text "Hello, welcome to this world."
# → audio/voice/ai-text-to-speech/ref.wavExplicit output path:
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--text "Hello, this is a test." \
--output audio/voice/ai-text-to-speech/hello.wavDirectory (writes <voice-stem>.wav inside, e.g. audio/voice/ai-text-to-speech/ref.wav):
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py --voice audio/voice/ref.wav --text "Hello, this is a test." --output audio/voice/ai-text-to-speechLong script from a UTF-8 text file:
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--text-file script/lines/intro.txt \
--output audio/voice/ai-text-to-speech/intro.wavFP16 (faster, less VRAM):
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--text "Testing half-precision inference." \
--fp16| Mode | Flags | Notes |
|---|---|---|
| Emotion reference audio | --emotion-audio path.wav | Separate clip for emotion; timbre still from --voice |
| Emotion weight | --emotion-weight 0.6 | Maps to emo_alpha (0.0–1.0, default 1.0) |
| Emotion from text | --emotion-from-text | Infer emotion from synthesis text; prefer --emotion-weight ≈ 0.6 |
| Emotion description | --emotion-text "..." | Natural-language emotion; implies text emotion mode |
| Emotion vector | --emotion-vector 0,0,0.8,0,0,0,0,0 | 8 floats: happy, angry, sad, afraid, disgusted, melancholic, surprised, calm |
# Emotion reference audio
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--emotion-audio audio/voice/emo_sad.wav \
--emotion-weight 0.9 \
--text "The inn has gone rotten and started auctioning off rooms." \
--output audio/voice/ai-text-to-speech/sad_line.wav
# Emotion description text
.dependency/index-tts/.venv/Scripts/python.exe .ai/ai-text-to-speech/tts.py \
--voice audio/voice/ref.wav \
--emotion-text "afraid, tense" \
--emotion-weight 0.6 \
--text "Hide quickly! He is coming!" \
--output audio/voice/ai-text-to-speech/afraid_line.wavDo not combine --emotion-audio, --emotion-vector, and --emotion-text / --emotion-from-text in conflicting ways — pick one emotion source.
| Option | Default | Notes |
|---|---|---|
| Output | <voice-dir>/ai-text-to-speech/<voice-stem>.wav | --output file uses that name; --output directory uses the voice file's stem |
| Model | .dependency/index-tts/checkpoints | IndexTTS-2 |
--fp16 | off | Enable on GPU when VRAM is tight |
--emotion-weight | 1.0 | Lower (~0.6) for text emotion modes |
| Overwrite | off | Pass --force to replace an existing output |
--text-file). Ask if either is missing.index-tts in manifest.json; retry the same command.--fp16 for speed; CPU is acceptable for short tests only.ai-text-to-speech/; sources are never modified.| Issue | Fix |
|---|---|
index-tts not populated | Clone + uv sync + download checkpoints; update manifest |
checkpoints/config.yaml missing | Re-run hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints |
| CUDA / torch errors | Install CUDA 12.8+; or run on CPU (slow) |
| OOM / VRAM | Pass --fp16; shorten text; close other GPU apps |
| Slow HuggingFace | Set HF_ENDPOINT=https://hf-mirror.com; or use ModelScope |
uv sync / DeepSpeed fail on Windows | Use uv sync --extra webui without deepspeed |
| Unnatural emotion | Lower --emotion-weight to ~0.6; try a clearer --emotion-audio |
Copy-paste commands: cli/ai-text-to-speech.md
© godot-fun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/ai-text-to-speech of godot-fun/gai.
Open the folder on GitHubat commit a98225b
AI Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Text To Speech this skillgodot-fun/gai | 184 | — | ~2k | Automated safety check: Pass | MIT | |
| Edu Math Videowy51ai/edulab | 1.4k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Book Video Factorybytec-ai/book-video-factory | 321 | — | ~1.4k | Automated safety check: Notes | None | |
| Narrator AI CLINarratorAI-Studio/narrator-ai-cli | 137 | — | ~3.7k | Automated safety check: Pass | MIT | |
| Iflytek Voiceclone Ttsiflytek/iFly-Skills | 209 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| U2 TtsLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.7k | Automated safety check: Notes | MIT |
wy51ai/edulab
A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…
bytec-ai/book-video-factory
通用的多账号图书短视频生产工作流。用于用户希望建立图书号项目目录、配置账号级片头/声音/BGM/视觉规范,或只提供一本书后依次完成资料研究、口播稿、分镜、图片、配音、字幕、预览与成片导出。适用于新建工作区、批量管理多个账号、继续已有单书任务和检查生产状态;不绑定特定研究、图片、TTS、转录或视频渲染供应商。
NarratorAI-Studio/narrator-ai-cli
Create AI-narrated film/drama commentary videos via CLI. An agent skill from NarratorAI-Studio/narrator-ai-cli.
iflytek/iFly-Skills
A skill your agent uses when user asks to clone a voice, train a custom voice model, or synthesize speech with a cloned voice.
LeoYeAI/openclaw-master-skills
Text-to-speech conversion using UniSound's TTS WebSocket API for generating high-quality Chinese Mandarin audio from text.
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
godot-fun/gai
Reduces background noise in a single audio file using FFmpeg afftdn.
godot-fun/gai
Applies fade-in and fade-out at the start and end of a single audio file using FFmpeg.
godot-fun/gai
Normalizes a single audio file to consistent LUFS loudness with true-peak limiting using FFmpeg.
godot-fun/gai
Standardizes a single audio file to 44100 or 48000 Hz and exports 16-bit PCM WAV using FFmpeg.
godot-fun/gai
Splits a single audio file into two segments (part 1 before the split point, part 2 after) using FFmpeg.
godot-fun/gai
Converts a single audio file to OGG Vorbis using FFmpeg for Godot-ready compressed assets.
Works with
Categories
Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts). AI Text To Speech is an agent skill from godot-fun/gai. Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).
AI Text To Speech fits situations like: the user wants TTS; generating narration/voice lines from a reference WAV.
Run `npx skills add godot-fun/gai --skill ai-text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/ai-text-to-speech in godot-fun/gai) into .claude/skills/ai-text-to-speech in your project. Claude Code loads it when a task matches its description.
Run `npx skills add godot-fun/gai --skill ai-text-to-speech -a codex`. Or copy the skill folder (.agents/skills/ai-text-to-speech in godot-fun/gai) into .agents/skills/ai-text-to-speech in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add godot-fun/gai --skill ai-text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-text-to-speech, .gemini/skills/ai-text-to-speech, .github/skills/ai-text-to-speech and .opencode/skills/ai-text-to-speech in your project.
Going by SKILL.md and its folder, AI Text To Speech needs the command-line tools its instructions call (uv, python, git and hf). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: github.com, hf-mirror.com and mirrors.aliyun.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AI Text To Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with AI Text To Speech: Edu Math Video (wy51ai/edulab, 1.4k stars), Book Video Factory (bytec-ai/book-video-factory, 321 stars), Narrator AI CLI (NarratorAI-Studio/narrator-ai-cli, 137 stars) and Iflytek Voiceclone Tts (iflytek/iFly-Skills, 209 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
godot-fun (a GitHub organization) maintains it in godot-fun/gai, which has 184 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 11, 2026.
Source: godot-fun/gai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.