Agent skill

Add TTS Engine to Voicebox

by jamiepine in jamiepine/voicebox

Walks through adding a new text-to-speech engine to Voicebox end to end: dependency audit, backend, frontend wiring, PyInstaller bundling and frozen-build testing.

MITAuto-check passedDevelopment

Install Add TTS Engine to Voicebox

skills CLI
$ npx skills add jamiepine/voicebox --skill add-tts-engine -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jamiepine/voicebox add-tts-engine --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jamiepine/voicebox.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-tts-engine .claude/skills/add-tts-engine && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-tts-engine
GitHub stars
57k
Token cost
~1.3k tokens
SKILL.md length
561 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Walks through adding a new text-to-speech engine to Voicebox end to end: dependency audit, backend, frontend wiring, PyInstaller bundling and frozen-build testing.

  • Works in 6 steps: Read the guide → Dependency research (Phase 0) → Implement (Phases 1–4) → …
  • Adding support for a new open-source TTS model to Voicebox
  • SKILL.md covers Goal, Reference Doc, Workflow and Key Lessons (from v0.2.3), plus 1 more section
  • Calls git and just

What it does

The skill follows a phased guide kept in the Voicebox repository at `docs/content/docs/developer/tts-engines.mdx`, which the agent must read in full first. Phase 0 is mandatory: clone the model library into a throwaway directory, audit it and its transitive dependencies, and write down findings on PyPI versus non-PyPI packages, PyInstaller directives, runtime data files, native library paths, needed monkey-patches, sample rate and the model download method. Model loading and generation are tested on CPU in a throwaway virtual environment before any code is written.

Implementation comes next: a new backend module that implements the `TTSBackend` protocol and is registered in the backends package, an updated model regex, frontend changes such as the engine type and per-engine languages, then dependency files, CI and Docker, and finally PyInstaller bundling and a test of the frozen build. The stated aim is that you only have to test the final build locally. The skill notes that skipping Phase 0 once turned a release into three patch releases.

When your agent uses it

  • Adding support for a new open-source TTS model to Voicebox
  • Auditing a model library's dependencies before packaging it with PyInstaller
  • Wiring a new engine into the Voicebox frontend and language lists

Example prompts

  • “Add a new text-to-speech engine to Voicebox, starting with the dependency audit.”
  • “Audit the dependencies of this model library before we bundle it with PyInstaller.”
  • “Wire the new engine into the frontend language constants and engine type.”

Requirements

  • A Voicebox checkout that includes its TTS engine guide
  • Python with a throwaway virtual environment for testing the model on CPU
  • PyInstaller, to build and check the frozen binary

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Read the guide
  2. Dependency research (Phase 0)
  3. Implement (Phases 1–4)
  4. PyInstaller bundling (Phase 5)
  5. Verify in dev mode
  6. Use the checklist

What it can do on your machine

Read from SKILL.md and the folder at commit 8af7efe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add TTS Engine to Voicebox loads about 1.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 561 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jamiepine/voicebox at commit 8af7efe, republished under its MIT licence (© jamiepine). 561 words, ~1,340 tokens.

Download SKILL.mdSave it as .claude/skills/add-tts-engine/SKILL.md (or your agent's skills folder).
name
add-tts-engine
description
Use this skill to add a new TTS engine to Voicebox. It walks through dependency research, backend implementation, frontend wiring, PyInstaller bundling, and frozen-build testing. Always start with Phase 0 (dependency audit) before writing any code.

Add TTS Engine

Goal

Integrate a new text-to-speech engine into Voicebox end-to-end: dependency research, backend protocol implementation, frontend UI wiring, PyInstaller bundling, and frozen-build verification. The user should only need to test the final build locally.

Reference Doc

The full phased guide lives at docs/content/docs/developer/tts-engines.mdx. Read this file in its entirety before starting. It contains:

  • Phase 0: Dependency research (mandatory before writing code)
  • Phase 1: Backend implementation (TTSBackend protocol)
  • Phase 2: Route and service integration (usually zero changes)
  • Phase 3: Frontend integration (5 files)
  • Phase 4: Dependencies (requirements.txt, justfile, CI, Docker)
  • Phase 5: PyInstaller bundling (build_binary.py + server.py)
  • Phase 6: Common upstream workarounds
  • Implementation checklist (gate between phases)

Workflow

1. Read the guide
bash
# Read the full TTS engines doc
cat docs/content/docs/developer/tts-engines.mdx

Internalize all phases, especially Phase 0 and Phase 5. The v0.2.3 release was three patch releases because Phase 0 was skipped.

2. Dependency research (Phase 0)

Clone the model library into a temporary directory and audit it. Do NOT skip this.

bash
mkdir /tmp/engine-research && cd /tmp/engine-research
git clone <model-library-url>

Run the grep searches from Phase 0.2 in the guide against the cloned source and its transitive dependencies. Produce a written dependency audit covering:

  1. PyPI vs non-PyPI packages
  2. PyInstaller directives needed (--collect-all, --copy-metadata, --hidden-import)
  3. Runtime data files that must be bundled
  4. Native library paths that need env var overrides in frozen builds
  5. Monkey-patches needed (torch.load, float64, MPS, HF token)
  6. Sample rate
  7. Model download method (from_pretrained vs snapshot_download + from_local)

Test model loading and generation on CPU in the throwaway venv before proceeding.

3. Implement (Phases 1–4)

Follow the guide's phases in order. Key files to modify:

Backend (Phase 1):

  • Create backend/backends/<engine>_backend.py
  • Register in backend/backends/__init__.py (ModelConfig + TTS_ENGINES + factory)
  • Update regex in backend/models.py

Frontend (Phase 3):

  • app/src/lib/api/types.ts — engine union type
  • app/src/lib/constants/languages.ts — ENGINE_LANGUAGES
  • app/src/components/Generation/EngineModelSelector.tsx — ENGINE_OPTIONS, ENGINE_DESCRIPTIONS
  • app/src/lib/hooks/useGenerationForm.ts — Zod schema, model-name mapping
  • app/src/components/ServerSettings/ModelManagement.tsx — MODEL_DESCRIPTIONS

Dependencies (Phase 4):

  • backend/requirements.txt
  • justfile (setup-python, setup-python-release targets)
  • .github/workflows/release.yml
  • Dockerfile (if applicable)
4. PyInstaller bundling (Phase 5)

Register the engine in backend/build_binary.py:

  • --hidden-import for the backend module and model package
  • --collect-all for packages using inspect.getsource, shipping data files, or native libraries
  • --copy-metadata for packages using importlib.metadata

If the engine has native data paths, add os.environ.setdefault() in backend/server.py inside the if getattr(sys, 'frozen', False): block.

Show full SKILL.md (208 more words)Show less
5. Verify in dev mode
bash
just dev

Test the full chain: model download → load → generate → voice cloning.

6. Use the checklist

Walk through the Implementation Checklist at the bottom of tts-engines.mdx. Every item must be checked before handing the build to the user.

Key Lessons (from v0.2.3)

These are the most common failure modes. Phase 0 research catches all of them:

PatternSymptom in Frozen BuildFix
@typechecked / inspect.getsource()"could not get source code"--collect-all <package>
Package ships pretrained model filesFileNotFoundError for .pth.tar, .yaml--collect-all <package>
C library with hardcoded system pathsFileNotFoundError for /usr/share/...--collect-all + env var in server.py
importlib.metadata.version()"No package metadata found"--copy-metadata <package>
torch.load without map_locationCUDA device not available on CPU buildMonkey-patch torch.load
torch.from_numpy on float64 datadtype mismatch RuntimeErrorCast to .float()
token=True in HF download callsAuth failure without stored HF tokenUse snapshot_download(token=None) + from_local()

Notes

  • The route and service layers have zero per-engine dispatch points. main.py requires zero changes.
  • The model config registry in backends/__init__.py handles all dispatch automatically.
  • Use get_torch_device() and model_load_progress() from backends/base.py — don't reimplement device detection or progress tracking.
  • Always test with a clean HuggingFace cache (no pre-downloaded models from dev).
  • Do NOT push or create a release. Hand the build to the user for local testing.

© jamiepine, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-tts-engine of jamiepine/voicebox.

Open the folder on GitHubat commit 8af7efe

Compare with similar skills

Add TTS Engine to Voicebox next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add TTS Engine to Voicebox compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add TTS Engine to Voicebox this skilljamiepine/voicebox57k—~1.3kAutomated safety check: PassMIT
Release NotesAivis-Project/AivisSpeech483—~745Automated safety check: PassLGPL-3.0
Abo Docsjeeftor/audiobook-organizer189—~499Automated safety check: PassMIT
Abo Featurejeeftor/audiobook-organizer189—~446Automated safety check: PassMIT
No Commentstellahq/opensession392—~310Automated safety check: PassMIT
Talking Head Videogooseworks-ai/goose-skills1.2k1 repos~8.4kAutomated safety check: NotesMIT

Similar skills

  • Release Notes

    Aivis-Project/AivisSpeech

    AivisSpeech の新バージョンリリース時に updateInfos のリリースノートドラフトを作成・更新するスキル。「リリースノート」「updateInfos」「アップデート情報」「リリース準備」などのキーワードが出たら使う。updateInfos.draft.json の作成・更新、エンジン側リリースノートとの統合、漏れチェックまでを包括的にサポートする。

    483 GitHub stars~745 tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Abo Docs

    jeeftor/audiobook-organizer

    Update Audiobook Organizer documentation, AGENTS.md, changelog entries, and maintainer-facing workflow notes while keeping repo-local skill references consistent.

    189 GitHub stars~499 tokensUpdated 26 days ago
    DevelopmentAuto-check passed
  • Abo Feature

    jeeftor/audiobook-organizer

    Implement focused Audiobook Organizer features across CLI, organizer core, TUI, web UI, or ABS workflows while following issue, branch, tests, docs, and changelog rules.

    189 GitHub stars~446 tokensUpdated 26 days ago
    DevelopmentAuto-check passed
  • No Comments

    tellahq/opensession

    Review a diff for narration, stale comments, and suppressions that clearer code should replace.

    392 GitHub stars~310 tokensUpdated today
    DevelopmentAuto-check passed
  • Talking Head Video

    gooseworks-ai/goose-skills

    Creates talking head videos from any source material (docs, changelogs, blog posts, notes, transcripts).

    1.2k GitHub starsUsed in 1 repo~8.4k tokens
    DevelopmentAuto-check: notes
  • Dep Updates

    trufflesecurity/trufflehog

    Plan and apply Go dependency updates, including advisory-driven bumps, Trivy/govulncheck validation, and supply-chain review.

    28k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed

More from jamiepine/voicebox

  • Draft Release Notes

    jamiepine/voicebox

    Writes or refreshes the Unreleased section of CHANGELOG.md as a themed narrative built from the commits, PRs and diff since the last version tag.

    57k GitHub stars~941 tokensUpdated today
    Auto-check passed
  • Release Bump

    jamiepine/voicebox

    Ends a release cycle by moving the Unreleased changelog notes under a dated version heading, bumping version files with bumpversion and tagging the commit.

    57k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Pre-Release PR Triage

    jamiepine/voicebox

    Sorts a backlog of open pull requests into must-merge, candidate, superseded and deferred, writes a triage doc and works the merge loop before a release.

    57k GitHub stars~3.1k tokensUpdated today
    Auto-check passed

Questions about Add TTS Engine to Voicebox

What does Add TTS Engine to Voicebox do?

Walks through adding a new text-to-speech engine to Voicebox end to end: dependency audit, backend, frontend wiring, PyInstaller bundling and frozen-build testing. mdx`, which the agent must read in full first. Phase 0 is mandatory: clone the model library into a throwaway directory, audit it and its transitive dependencies, and write down findings on PyPI versus non-PyPI packages, PyInstaller directives, runtime data files, native library paths, needed monkey-patches, sample rate and the model download method.

When should I use Add TTS Engine to Voicebox?

Add TTS Engine to Voicebox fits situations like: adding support for a new open-source TTS model to Voicebox; auditing a model library's dependencies before packaging it with PyInstaller; wiring a new engine into the Voicebox frontend and language lists.

How do I install Add TTS Engine to Voicebox in Claude Code?

Run `npx skills add jamiepine/voicebox --skill add-tts-engine -a claude-code`. Or copy the skill folder (.agents/skills/add-tts-engine in jamiepine/voicebox) into .claude/skills/add-tts-engine in your project. Claude Code loads it when a task matches its description.

How do I install Add TTS Engine to Voicebox in Codex?

Run `npx skills add jamiepine/voicebox --skill add-tts-engine -a codex`. Or copy the skill folder (.agents/skills/add-tts-engine in jamiepine/voicebox) into .agents/skills/add-tts-engine in your project. Codex loads it when a task matches its description.

Can I use Add TTS Engine to Voicebox in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jamiepine/voicebox --skill add-tts-engine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-tts-engine, .gemini/skills/add-tts-engine, .github/skills/add-tts-engine and .opencode/skills/add-tts-engine in your project.

What does Add TTS Engine to Voicebox need to run?

Going by SKILL.md and its folder, Add TTS Engine to Voicebox needs the command-line tools its instructions call (git and just). Our summary lists: A Voicebox checkout that includes its TTS engine guide; Python with a throwaway virtual environment for testing the model on CPU; PyInstaller, to build and check the frozen binary.

Does Add TTS Engine to Voicebox access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add TTS Engine to Voicebox safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add TTS Engine to Voicebox use?

Add TTS Engine to Voicebox is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add TTS Engine to Voicebox use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add TTS Engine to Voicebox?

Skills that share tags, products or a category with Add TTS Engine to Voicebox: Release Notes (Aivis-Project/AivisSpeech, 483 stars), Abo Docs (jeeftor/audiobook-organizer, 189 stars), Abo Feature (jeeftor/audiobook-organizer, 189 stars) and No Comments (tellahq/opensession, 392 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add TTS Engine to Voicebox?

jamiepine (a GitHub user) maintains it in jamiepine/voicebox, which has 56,524 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: jamiepine/voicebox on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.