Agent skill

Piper Tts Training

by sammcj in sammcj/agentic-coding

Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Piper Tts Training

skills CLI
$ npx skills add sammcj/agentic-coding --skill piper-tts-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sammcj/agentic-coding piper-tts-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills_disabled/piper-tts-training .claude/skills/piper-tts-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
piper-tts-training
GitHub stars
162
Token cost
~1.4k tokens
SKILL.md length
461 words
Files
4 (incl. scripts, references)
Skills in repo
64
Repo updated
First seen
Licence
Apache-2.0

At a glance

Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

  • Works in 7 steps: Corpus Preparation → Audio Generation → Quality Validation with Whisper → …
  • Creating new synthetic voices
  • SKILL.md covers Overview, Workflow, Localisation for Australian,… and Dependencies, plus 2 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Piper Tts Training is an agent skill from sammcj/agentic-coding. Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. Use when creating new synthetic voices, fine-tuning existing Piper checkpoints, preparing audio datasets for TTS training, or deploying voice models to devices like Raspberry Pi or Home Assistant. Covers dataset preparation, Whisper-based validation, training configuration, and ONNX export.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/american_spellings.json`, `references/localisation.md` and `scripts/convert_spelling.py`).

It sits in AI & LLM Engineering, covering Text to speech and voice, Fine-tuning and Speech recognition and synthesis. It works with ONNX and Home Assistant. The repository describes itself as: Agentic Coding Rules, Templates etc... The licence is Apache-2.0.

When your agent uses it

  • Creating new synthetic voices
  • Fine-tuning existing Piper checkpoints
  • Preparing audio datasets for TTS training
  • Deploying voice models to devices like Raspberry Pi

Example prompts

  • “/piper-tts-training”

Requirements

  • Python 3
  • Docker

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Corpus Preparation
  2. Audio Generation
  3. Quality Validation with Whisper
  4. Dataset Format (LJSpeech)
  5. Preprocessing
  6. Training
  7. ONNX Export

What it can do on your machine

Read from SKILL.md and the folder at commit 2f25ced. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Piper Tts Training loads about 1.4k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 461 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sammcj/agentic-coding at commit 2f25ced, republished under its Apache-2.0 licence (© sammcj). 461 words, ~1,449 tokens.

Download SKILL.mdSave it as .claude/skills/piper-tts-training/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
piper-tts-training
description
Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. Use when creating new synthetic voices, fine-tuning existing Piper checkpoints, preparing audio datasets for TTS training, or deploying voice models to devices like Raspberry Pi or Home Assistant. Covers dataset preparation, Whisper-based validation, training configuration, and ONNX export.

Piper TTS Voice Training

Train custom text-to-speech voices compatible with Piper's lightweight ONNX runtime.

Overview

Piper produces fast, offline TTS suitable for embedded devices. Training involves:

  1. Corpus preparation (text covering phonetic range)
  2. Audio generation or recording
  3. Quality validation via Whisper transcription
  4. Fine-tuning from existing checkpoint (recommended) or training from scratch
  5. ONNX export for deployment

Fine-tuning vs from-scratch:

  • Fine-tuning: ~1,300 phrases + 1,000 epochs (days on modest GPU)
  • From scratch: ~13,000+ phrases + 2,000+ epochs (weeks/months)

Workflow

1. Corpus Preparation

Gather 1,300-1,500+ phrases covering broad phonetic range:

  • Use piper-recording-studio corpus as base
  • Add domain-specific phrases for your use case
  • Include varied sentence structures and lengths

Critical for non-US English: Ensure corpus uses correct regional spelling. See Localisation.

2. Audio Generation

Generate or record training audio at 22050Hz mono WAV.

If using voice cloning (e.g., Chatterbox TTS):

  • Generate at source sample rate (often 24kHz)
  • Convert to 22050Hz: sox -v 0.95 input.wav -r 22050 -t wav output.wav
  • The -v 0.95 prevents clipping during resampling

Recording requirements:

  • Consistent microphone position and room acoustics
  • Minimal background noise
  • Natural speaking pace (not reading voice)
3. Quality Validation with Whisper

Automate quality checks rather than manual listening:

python
import whisper
from piper_phonemize import phonemize_text

model = whisper.load_model("base")

def validate_sample(audio_path, expected_text):
    result = model.transcribe(audio_path)
    transcribed = result["text"].strip()

    # Compare phonemically to handle spelling/punctuation differences
    expected_phonemes = phonemize_text(expected_text, "en-gb")
    transcribed_phonemes = phonemize_text(transcribed, "en-gb")

    return expected_phonemes == transcribed_phonemes

Retry failed samples up to 3 times. Target 95%+ dataset coverage.

4. Dataset Format (LJSpeech)

Structure your dataset:

dataset/
├── metadata.csv
└── wavs/
    ├── sample_0001.wav
    ├── sample_0002.wav
    └── ...

metadata.csv format: {id}|{text} (pipe-separated, no headers)

sample_0001|The quick brown fox jumps over the lazy dog.
sample_0002|Pack my box with five dozen liquor jugs.
5. Preprocessing

Convert to PyTorch tensors:

bash
python3 -m piper_train.preprocess \
    --language en-gb \
    --input-dir dataset/ \
    --output-dir piper_training_dir/ \
    --dataset-format ljspeech

Use en-gb for Australian/NZ/UK voices (espeak-ng phoneme set).

6. Training

Fine-tuning (recommended):

bash
python3 -m piper_train \
    --dataset-dir piper_training_dir/ \
    --accelerator gpu \
    --devices 1 \
    --batch-size 12 \
    --max_epochs 3000 \
    --resume_from_checkpoint ljspeech-2000.ckpt \
    --checkpoint-epochs 100 \
    --quality high \
    --precision 32

Key parameters:

  • --batch-size: Reduce if VRAM limited (12 works on 8GB)
  • --resume_from_checkpoint: Start from LJSpeech high-quality checkpoint
  • --precision 32: More stable than mixed precision
  • --validation-split 0.0 --num-test-examples 0: Skip validation for small datasets

Monitor with TensorBoard: watch loss_disc_all for convergence.

Show full SKILL.md (186 more words)Show less
7. ONNX Export
bash
python3 -m piper_train.export_onnx checkpoint.ckpt output.onnx.unoptimized
onnxsim output.onnx.unoptimized output.onnx

Create metadata file output.onnx.json from training config.json.

Localisation for Australian, New Zealand and UK English

Piper uses espeak-ng for phonemisation. American pronunciations in training data cause accent drift.

Corpus preparation:

  • Run scripts/convert_spelling.py on corpus text before training
  • Use en-gb or en-au espeak-ng voice for phonemisation
  • Review generated phonemes for Americanisms

Common spelling conversions:

AmericanAustralian/UK
-ize-ise
-or-our
-er-re
-og-ogue
-ense-ence

Phoneme considerations:

  • /r/ linking and intrusion patterns differ
  • Vowel sounds in words like "dance", "bath", "castle"
  • Final -ile pronunciation (hostile, missile)

For complete word lists and phonetic details, see references/localisation.md.

Validation: Use Whisper with language="en" and verify transcriptions match expected regional forms.

Dependencies

Pin versions to avoid API breakage:

pytorch-lightning==1.9.3
torch<2.6.0
piper-phonemize
onnxruntime-gpu
onnxsim

Docker containerisation recommended for reproducibility.

Hardware Requirements

Minimum (fine-tuning):

  • 8GB VRAM GPU (Pascal or newer)
  • 8GB system RAM
  • ~5 days for 1,000 epochs on Tesla P4

From scratch: Multiply time by ~200x.

Troubleshooting

IssueSolution
CUDA OOMReduce batch-size (try 8 or 4)
Checkpoint won't loadCheck pytorch-lightning version matches checkpoint
Garbled outputInsufficient training epochs or dataset too small
Wrong accentCheck espeak-ng language code and corpus spelling

© sammcj, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in Skills_disabled/piper-tts-training of sammcj/agentic-coding.

  • SKILL.md
  • references/american_spellings.json
  • references/localisation.md
  • scripts/convert_spelling.py

Open the folder on GitHubat commit 2f25ced

Compare with similar skills

Piper Tts Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Piper Tts Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Piper Tts Training this skillsammcj/agentic-coding162—~1.4kAutomated safety check: PassApache-2.0
Lilly Community Researchssaaffaakk/Lilly171—~1.4kAutomated safety check: PassMIT
Xybrid Initxybrid-ai/xybrid469—~3kAutomated safety check: PassApache-2.0
Agentstadaspetra/loop2961 repos~2.5kAutomated safety check: PassMIT
Agentselevenlabs/skills482—~6.5kAutomated safety check: PassMIT
Parakeet Sttsundial-org/awesome-openclaw-skills663—~771Automated safety check: PassNone

Similar skills

  • Lilly Community Research

    ssaaffaakk/Lilly

    Lilly community-research skill. An agent skill from ssaaffaakk/Lilly.

    171 GitHub stars~1.4k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Xybrid Init

    xybrid-ai/xybrid

    Generate model metadata for an ML model so it works with xybrid.

    469 GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agents

    elevenlabs/skills

    Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.

    482 GitHub stars~6.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Parakeet Stt

    sundial-org/awesome-openclaw-skills

    Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).

    663 GitHub stars~771 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    xybrid-ai/xybrid

    Test a model end-to-end using the xybrid execution system. An agent skill from xybrid-ai/xybrid.

    469 GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from sammcj/agentic-coding

All 64 skills in this repo
  • Yue2 Music

    sammcj/agentic-coding

    A skill your agent uses when generating songs with YuE2, covering a recording via SheetSage2 audio-to-ABC, editing a score or lyrics with melody preservation, or building a reproducible listening…

    162 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Bento Slides

    sammcj/agentic-coding

    A skill your agent uses when creating or editing Bento (.bento.html) slide decks, including any request for a single-file HTML slide deck.

    162 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Idrive Backup

    sammcj/agentic-coding

    A skill your agent uses whenever the user wants you to manage, discuss or diagnose iDrive Backup configuration on macOS

    162 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check: notes
  • PPTX To Md

    sammcj/agentic-coding

    Convert a PPTX slide deck into per-slide markdown that preserves both the verbatim text and the meaning of embedded screenshots, diagrams and charts in their original layout positions.

    162 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Skill Creator Primer

    sammcj/agentic-coding

    You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

    162 GitHub stars~9.8k tokensUpdated 2 days ago
    Auto-check passed
  • Deferred Task Execution

    sammcj/agentic-coding

    Delays execution of a task until a specified time or after a duration.

    162 GitHub stars~642 tokensUpdated 2 days ago
    Auto-check: notes

Questions about Piper Tts Training

What does Piper Tts Training do?

Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. Piper Tts Training is an agent skill from sammcj/agentic-coding. Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

When should I use Piper Tts Training?

Piper Tts Training fits situations like: creating new synthetic voices; fine-tuning existing Piper checkpoints; preparing audio datasets for TTS training; deploying voice models to devices like Raspberry Pi.

How do I install Piper Tts Training in Claude Code?

Run `npx skills add sammcj/agentic-coding --skill piper-tts-training -a claude-code`. Or copy the skill folder (Skills_disabled/piper-tts-training in sammcj/agentic-coding) into .claude/skills/piper-tts-training in your project. Claude Code loads it when a task matches its description.

How do I install Piper Tts Training in Codex?

Run `npx skills add sammcj/agentic-coding --skill piper-tts-training -a codex`. Or copy the skill folder (Skills_disabled/piper-tts-training in sammcj/agentic-coding) into .agents/skills/piper-tts-training in your project. Codex loads it when a task matches its description.

Can I use Piper Tts Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sammcj/agentic-coding --skill piper-tts-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/piper-tts-training, .gemini/skills/piper-tts-training, .github/skills/piper-tts-training and .opencode/skills/piper-tts-training in your project.

What does Piper Tts Training need to run?

Going by SKILL.md and its folder, Piper Tts Training needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3; Docker.

Does Piper Tts Training access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Piper Tts Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Piper Tts Training use?

Piper Tts Training is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Piper Tts Training use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Piper Tts Training?

Skills that share tags, products or a category with Piper Tts Training: Lilly Community Research (ssaaffaakk/Lilly, 171 stars), Xybrid Init (xybrid-ai/xybrid, 469 stars), Agents (tadaspetra/loop, 296 stars) and Agents (elevenlabs/skills, 482 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Piper Tts Training?

sammcj (a GitHub user) maintains it in sammcj/agentic-coding, which has 162 GitHub stars. The repository holds 64 skills in this directory. The repository was last updated on October 9, 2026.

Source: sammcj/agentic-coding on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.