Agent skill

Storyboard Tts

by godot-fun in godot-fun/gai

Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech).

MITAuto-check passedMedia & Creative

Install Storyboard Tts

skills CLI
$ npx skills add godot-fun/gai --skill storyboard-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install godot-fun/gai storyboard-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/godot-fun/gai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/storyboard-tts .claude/skills/storyboard-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
storyboard-tts
GitHub stars
183
Token cost
~2.2k tokens
SKILL.md length
664 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech).

  • Works in 7 steps: Batch synthesis only via… → Trial / single-line checks may use… → Parse-only / report-only / subtitle-only… → …
  • The user wants storyboard TTS
  • SKILL.md covers Rules, Inputs, Layout and Quick Start, plus 4 more sections
  • Calls python

What it does

Storyboard Tts is an agent skill from godot-fun/gai. Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech). Batch-writes WAVs under Chinese/ and English/ named by shot id (model loaded once), pads 0.4 s edge silence in place, then a speech-timeline.md and one concatenated SRT per language (Chinese.srt / English.srt). Use when the user wants storyboard TTS, storyboard to speech, narration VO, storyboard-to-speech, bilingual VO export, subtitles, or batch TTS from a storyboard.md.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Comics and storyboards and Transcription. It works with Python. The repository describes itself as: A lightweight AI agent and skill workflow framework built with Godot. The licence is MIT.

When your agent uses it

  • The user wants storyboard TTS
  • Storyboard to speech
  • Storyboard-to-speech
  • Bilingual VO export

Example prompts

  • “Use the storyboard-tts skill to convert storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same…”
  • “/storyboard-tts”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Batch synthesis only via .ai/storyboard-tts/synthesize.py with the index-tts interpreter. Do not hand-write IndexTTS loops, temporary…
  2. Trial / single-line checks may use ai-text-to-speech tts.py, or synthesize.py --limit 1.
  3. Parse-only / report-only / subtitle-only steps use stdlib python (.dependency/python/python).
  4. Never overwrite the storyboard source. Write only under /.
  5. Skip (no VO) / empty lines — no empty WAVs or empty subtitle cues.
  6. Confirm voice reference (and output dir if unclear) before a full batch.
  7. The top-level output directory uses the reference voice filename stem. Subdirectory names stay fixed. Default is //. Do not use -speech`.

What it can do on your machine

Read from SKILL.md and the folder at commit 6394686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Storyboard Tts loads about 2.2k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 664 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from godot-fun/gai at commit 6394686, republished under its MIT licence (© godot-fun). 664 words, ~2,215 tokens.

Download SKILL.mdSave it as .claude/skills/storyboard-tts/SKILL.md (or your agent's skills folder).
name
storyboard-tts
description
Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech). Batch-writes WAVs under Chinese/ and English/ named by shot id (model loaded once), pads 0.4 s edge silence in place, then a speech-timeline.md and one concatenated SRT per language (Chinese.srt / English.srt). Use when the user wants storyboard TTS, storyboard to speech, narration VO, storyboard-to-speech, bilingual VO export, subtitles, or batch TTS from a storyboard.md.

Storyboard TTS

Take a storyboard deliverable and batch-synthesize Chinese + English voice-over with IndexTTS2 (shared setup with ai-text-to-speech).

OutputPath
Chinese VO<storyboard-dir>/<voice-stem>/Chinese/<shot-id>.wav
English VO<storyboard-dir>/<voice-stem>/English/<shot-id>.wav
Duration doc<storyboard-dir>/<voice-stem>/speech-timeline.md
Chinese subs<storyboard-dir>/<voice-stem>/Chinese.srt
English subs<storyboard-dir>/<voice-stem>/English.srt

Shot id from headers (### Shot 01 — … → 01.wav).

Subtitles: one SRT per language. Shots are laid end-to-end on the VO timeline (shot N starts when N−1 ends). Inside a shot, text is split on sentence punctuation (.!?;… and CJK equivalents) into multiple cues; cue lengths share that shot’s WAV duration by non-whitespace character weight. Skip (no VO) / missing audio.

Rules

When this skill applies, read and follow skill-dependency-manager — run scripts as documented, install missing tools into .dependency/.

  1. Batch synthesis only via .ai/storyboard-tts/synthesize.py with the index-tts interpreter. Do not hand-write IndexTTS loops, temporary batch drivers, or N× single tts.py calls for a full storyboard.
  2. Trial / single-line checks may use ai-text-to-speech tts.py, or synthesize.py --limit 1.
  3. Parse-only / report-only / subtitle-only steps use stdlib python (.dependency/python/python).
  4. Never overwrite the storyboard source. Write only under <audio-dir>/.
  5. Skip (no VO) / empty lines — no empty WAVs or empty subtitle cues.
  6. Confirm voice reference (and output dir if unclear) before a full batch.
  7. The top-level output directory uses the reference voice filename stem. Subdirectory names stay fixed. Default <audio-dir> is <storyboard-dir>/<voice-stem>/. Do not use <storyboard-stem>-speech.

Inputs

RequiredNotes
Storyboard .md### Shot NN — title with - **Chinese:** / - **English:**
Reference voiceWAV/MP3 for IndexTTS (--voice, or --voice-zh / --voice-en)
OptionalDefault
Output dir<storyboard-dir>/<voice-stem>/ (reference audio filename, no extension)
Languageboth (--lang chinese / english)
--fp16 / emotion / --devicesame meaning as ai-text-to-speech
--forceoff (skip existing WAVs)
--limit N0 = all jobs (use 1 for trial)
--reportwrite speech-timeline.md + Chinese.srt / English.srt after synth
--no-subtitleswith --report, skip SRT files
Edge padon (0.4 s); --no-pad / --pad-duration / --pad-threshold

Layout

<storyboard-dir>/<voice-stem>/   # e.g. narrator-self-fast/ from narrator-self-fast.wav
  Chinese/
    01.wav
    …
  English/
    01.wav
    …
  shots.json
  speech-timeline.md
  Chinese.srt            # all Chinese cues, continuous timeline
  English.srt            # all English cues, continuous timeline
  _text/                 # only with --write-text

Quick Start

From project root:

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --storyboard path/to/storyboard.md --voice path/to/ref.wav --fp16 --report

Trial run (first line only):

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --storyboard path/to/storyboard.md --voice path/to/ref.wav --fp16 --limit 1

Writes under <storyboard-dir>/<voice-stem>/ (override with --audio-dir only when needed).

This will:

  1. Parse the storyboard → <audio-dir>/shots.json
  2. Load IndexTTS2 once
  3. Write Chinese/<id>.wav and English/<id>.wav (skip existing unless --force)
  4. Pad each WAV in place to 0.4 s leading/trailing silence (--no-pad to skip)
  5. Write speech-timeline.md and Chinese.srt / English.srt when --report
Separate voices / language
bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --storyboard path/to/storyboard.md --voice-zh path/to/zh_ref.wav --voice-en path/to/en_ref.wav --fp16 --report
bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --storyboard path/to/storyboard.md --voice path/to/ref.wav --lang chinese --fp16 --report
Parse, report, or subtitles alone (stdlib python)
bash
.dependency/python/python .ai/storyboard-tts/parse_storyboard.py path/to/storyboard.md -o path/to/<audio-dir>/shots.json
bash
.dependency/python/python .ai/storyboard-tts/duration_report.py --storyboard path/to/storyboard.md --audio-dir path/to/<audio-dir> --shots path/to/<audio-dir>/shots.json -o path/to/<audio-dir>/speech-timeline.md
bash
.dependency/python/python .ai/storyboard-tts/write_subtitles.py --audio-dir path/to/<audio-dir> --shots path/to/<audio-dir>/shots.json

Resume from an existing shots.json:

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --shots path/to/<audio-dir>/shots.json --voice path/to/ref.wav --fp16 --report
Show full SKILL.md (294 more words)Show less

Common Flags (synthesize.py)

FlagNotes
--storyboard / --shotsSource (one required)
--audio-dirOutput root (default: <storyboard-dir>/<voice-stem>/)
--voiceShared speaker ref
--voice-zh / --voice-enPer-language refs
--langboth (default), chinese, english
--limit NFirst N pending jobs only
--forceOverwrite existing WAVs
--reportWrite timeline + SRT subtitles
--report-outCustom timeline path
--no-subtitlesSkip SRT when using --report
--write-textDump lines under _text/
--no-padSkip in-place 0.4 s edge padding (off by default — padding is on)
--pad-durationTarget silence per edge in seconds (default: 0.4)
--pad-thresholdSilence detect threshold in dB (default: -50)
--fp16 / --deviceRuntime
--emotion-* / --random / --verboseSame role as tts.py

On partial failure: script continues remaining jobs, prints Failed jobs: …, exit code 1. Fix install/voice per ai-text-to-speech troubleshooting, re-run (existing OK files are skipped).

Agent Notes

  1. Do not invent narration — use storyboard Chinese/English fields as-is (audio and subtitles).
  2. Prefer one synthesize.py invocation for a full board; model reload cost is the reason.
  3. Chat summary: audio-dir (<storyboard-dir>/<voice-stem>/), counts, path to speech-timeline.md, Chinese.srt / English.srt, Chinese/English total seconds — no full transcripts unless asked. Omit --audio-dir unless overriding.
  4. IndexTTS install lives in ai-text-to-speech; do not duplicate Setup here beyond “populate index-tts if missing”.
  5. Edge padding is built in (default 0.4 s, in place via pad.py). Use --no-pad only when they want raw TTS with no extra silence; --pad-duration if they want a different length.
  6. Loudnorm / OGG / trim remain separate skills after this one.
  7. If audio already exists and only subtitles are needed, run write_subtitles.py alone (stdlib python). Re-running synthesize.py --report without --force still pads existing WAVs, then rewrites the timeline/SRT.

Tests

Stdlib scripts (from repo root):

bash
.dependency/python/python .ai/storyboard-tts/test_parse_storyboard.py
.dependency/python/python .ai/storyboard-tts/test_write_subtitles.py

IndexTTS batch driver (requires populated index-tts; see cli/storyboard-tts.md):

bash
.dependency/index-tts/.venv/Scripts/python.exe .ai/storyboard-tts/synthesize.py --storyboard path/to/storyboard.md --voice .ai/test/audio/han.wav --fp16 --limit 1 --report

© godot-fun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/storyboard-tts of godot-fun/gai.

Open the folder on GitHubat commit 6394686

Compare with similar skills

Storyboard Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Storyboard Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Storyboard Tts this skillgodot-fun/gai183—~2.2kAutomated safety check: PassMIT
Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE2414 repos~5.2kAutomated safety check: PassNone
Web Demo Video SynthesisSven-LI-sankyuu/presentation-skills175—~3.2kAutomated safety check: PassNone
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Whiteboard Videognipbao/codex-whiteboard-video-skill327—~7.2kAutomated safety check: NotesMIT
Content To Videoarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Paper Collage Explainer Generator

    tl2012tl/comfyUI-llama-TE

    For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.

    241 GitHub starsUsed in 4 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Web Demo Video Synthesis

    Sven-LI-sankyuu/presentation-skills

    用于“网页 demo 分段配音 + timeline 驱动录屏 + 后期合成”的 workspace 协作流程:先搭建一个可审计工作目录(cues/timeline/segmentaudio/video/subtitles/final),再由人类 + Codex 迭代维护这些文件,按需只重跑局部步骤,最终合成高质量 MP4。适用于强调可复盘、可编辑、清晰度与字幕安全区可控的场景。

    175 GitHub stars~3.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    327 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: warnings

More from godot-fun/gai

All 36 skills in this repo
  • AI Text To Speech

    godot-fun/gai

    Zero-shot text-to-speech with voice cloning via IndexTTS2 (index-tts).

    184 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audio Denoise

    godot-fun/gai

    Reduces background noise in a single audio file using FFmpeg afftdn.

    184 GitHub stars~465 tokensUpdated today
    Auto-check passed
  • Audio Fade

    godot-fun/gai

    Applies fade-in and fade-out at the start and end of a single audio file using FFmpeg.

    184 GitHub stars~652 tokensUpdated today
    Auto-check passed
  • Normalizes a single audio file to consistent LUFS loudness with true-peak limiting using FFmpeg.

    184 GitHub stars~684 tokensUpdated today
    Auto-check passed
  • Standardizes a single audio file to 44100 or 48000 Hz and exports 16-bit PCM WAV using FFmpeg.

    184 GitHub stars~638 tokensUpdated today
    Auto-check passed
  • Audio Split

    godot-fun/gai

    Splits a single audio file into two segments (part 1 before the split point, part 2 after) using FFmpeg.

    184 GitHub stars~554 tokensUpdated today
    Auto-check passed

Works with

Questions about Storyboard Tts

What does Storyboard Tts do?

Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech). Storyboard Tts is an agent skill from godot-fun/gai. Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech).

When should I use Storyboard Tts?

Storyboard Tts fits situations like: the user wants storyboard TTS; storyboard to speech; storyboard-to-speech; bilingual VO export.

How do I install Storyboard Tts in Claude Code?

Run `npx skills add godot-fun/gai --skill storyboard-tts -a claude-code`. Or copy the skill folder (.agents/skills/storyboard-tts in godot-fun/gai) into .claude/skills/storyboard-tts in your project. Claude Code loads it when a task matches its description.

How do I install Storyboard Tts in Codex?

Run `npx skills add godot-fun/gai --skill storyboard-tts -a codex`. Or copy the skill folder (.agents/skills/storyboard-tts in godot-fun/gai) into .agents/skills/storyboard-tts in your project. Codex loads it when a task matches its description.

Can I use Storyboard Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add godot-fun/gai --skill storyboard-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/storyboard-tts, .gemini/skills/storyboard-tts, .github/skills/storyboard-tts and .opencode/skills/storyboard-tts in your project.

What does Storyboard Tts need to run?

Going by SKILL.md and its folder, Storyboard Tts needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Storyboard Tts access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Storyboard Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Storyboard Tts use?

Storyboard Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Storyboard Tts use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Storyboard Tts?

Skills that share tags, products or a category with Storyboard Tts: Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Web Demo Video Synthesis (Sven-LI-sankyuu/presentation-skills, 175 stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 327 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Storyboard Tts?

godot-fun (a GitHub organization) maintains it in godot-fun/gai, which has 183 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 10, 2026.

Source: godot-fun/gai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.