Agent skill

Kokoro Tts

by calesthio in calesthio/generative-media-skills

Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad.

MITAuto-check passedMedia & Creative

Install Kokoro Tts

skills CLI
$ npx skills add calesthio/generative-media-skills --skill kokoro-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills kokoro-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/text-to-speech/kokoro-tts .claude/skills/kokoro-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kokoro-tts
GitHub stars
197
Token cost
~4.8k tokens
SKILL.md length
2,088 words
Files
2
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad.

  • Works in 3 steps: Python kokoro package (default for… → kokoro-js / ONNX (browser, on-device,… → Kokoro-FastAPI (drop-in…
  • Synthesizing speech offline
  • SKILL.md covers What Kokoro is (documented…, When to use Kokoro vs. when…, Voices, language codes, and… and Text length, chunking, and…, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Kokoro Tts is an agent skill from calesthio/generative-media-skills. Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad. Use when synthesizing speech offline, on-device, in the browser, or on your own server without per-character API cost — for high-volume narration, audiobooks, privacy-constrained pipelines, and prototyping. Covers running Kokoro via the Python kokoro package, ONNX / kokoro-js in the browser, and the OpenAI-compatible Kokoro-FastAPI wrapper; picking voices and language codes; blending voices; chunking…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).

It sits in Media & Creative, covering Text to speech and voice. It works with OpenAI, FastAPI, Python and ONNX. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Synthesizing speech offline
  • On your own server without per-character API cost — for high-volume narration
  • Privacy-constrained pipelines

Example prompts

  • “/kokoro-tts”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Python kokoro package (default for servers/batch)
  2. kokoro-js / ONNX (browser, on-device, Node)
  3. Kokoro-FastAPI (drop-in OpenAI-compatible server)

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co
    • deepwiki.com
    • gist.github.com
    • artificialanalysis.ai
    • pypi.org
    • npmjs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kokoro Tts loads about 4.8k tokens when it runs. Until then it costs about 229 tokens; SKILL.md has 2,088 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~229
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 2,088 words, ~4,825 tokens.

Download SKILL.mdSave it as .claude/skills/kokoro-tts/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
kokoro-tts
description
Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad. Use when synthesizing speech offline, on-device, in the browser, or on your own server without per-character API cost — for high-volume narration, audiobooks, privacy-constrained pipelines, and prototyping. Covers running Kokoro via the Python `kokoro` package, ONNX / kokoro-js in the browser, and the OpenAI-compatible Kokoro-FastAPI wrapper; picking voices and language codes; blending voices; chunking long text; controlling pronunciation via misaki/espeak-ng; and judging when Kokoro is the right tool versus when its lack of voice cloning, narrow emotional range, and weaker non-English quality mean you should reach for a hosted or larger model instead. Not for voice cloning, expressive/emotional character performance, or high-fidelity multilingual work — say so and route elsewhere.

Kokoro TTS (open-weight, self-hosted)

Kokoro is a small, fast, permissively licensed text-to-speech model. Its entire value proposition is that you run it yourself: no API key, no per-character billing, no audio leaving your machine. This skill helps an agent decide whether Kokoro fits a job, run it through the right runtime, and get acceptable output — and, just as importantly, recognize the jobs where Kokoro will disappoint the user and something else is the correct answer.

All version, license, voice-count, ranking, and performance facts below were verified on 2026-07-10 against the sources listed at the end. Treat them as volatile.

What Kokoro is (documented facts)

  • Model. ~82 million parameters. Architecture is StyleTTS 2 (arXiv 2306.07691) with an ISTFTNet decoder (arXiv 2203.02395). The card describes it as "Decoder only: no diffusion, no encoder release." Source: hexgrad/Kokoro-82M model card.
  • License. Apache 2.0, including the weights. v0.19 weights were released in full fp32 on 2024-12-25; v1.0 released 2025-01-27 and is the current default. Because weights are Apache-2.0 you may deploy commercially, redistribute, and fine-tune (subject to attribution). Source: model card.
  • Training data & provenance. "Few hundred hrs" for v1.0, trained exclusively on permissive / non-copyrighted material: public-domain audio, Apache/MIT-licensed content, and synthetic audio generated by closed TTS models, plus <1 hr from Koniwa (CC BY 3.0) and <11 hrs from SIWIS (CC BY 4.0). Reported training cost ≈ $1000 (~1000 A100-80GB GPU-hours). The heavy reliance on synthetic data is the root cause of Kokoro's flat prosody and its uneven non-English quality — keep it in mind. Source: model card.
  • Output. 24 kHz mono audio. Sample rate is fixed at 24000 Hz. Source: model card.
  • Coverage. v1.0 ships 8 languages and 54 voices (American + British English count as one language). Sources: model card, VOICES.md.

When to use Kokoro vs. when not to

Reach for Kokoro when:

  • You need cost-free, high-volume English narration (audiobooks, course/video voiceover, batch document-to-speech, screen readers). Marginal cost is electricity.
  • The pipeline is offline or privacy-constrained — medical, legal, on-device, air-gapped — and no text may be sent to a cloud API.
  • You want in-browser / on-device TTS with no server (kokoro-js + WebGPU/WASM).
  • You are prototyping and want a good-enough voice today without provisioning a paid provider.

Do NOT use Kokoro (and tell the user so) when the job needs:

  • Voice cloning / a specific person's voice. Kokoro has no speaker encoder and no zero-shot cloning. The encoder was deliberately not released. You cannot clone a reference voice. Route to a cloning-capable provider.
  • Emotional or character performance — laughter, crying, shouting, sarcasm, dynamic delivery. Kokoro has no emotion/style tokens and a narrow prosodic range; output is competent but flat. Fine for a neutral narrator, wrong for a video-game character.
  • High-fidelity non-English or many languages. Non-English voices are mostly C/D-graded and trained on little data (see quality grades below); several languages also truncate long text. English is the only tier-1 language.
  • Conversational agents that need real-time bidirectional dialogue with barge-in and personality. Kokoro is a batch/streaming synthesizer, not a dialog voice.

If the request is voice cloning or emotional VO, do not try to fake it with blending or prompt tricks — state the limitation plainly and suggest a cloning/expressive provider.

Voices, language codes, and quality grades

Voice IDs follow [langprefix][gender]_[name], e.g. af_heart = American Female "Heart", bm_george = British Male "George", if_sara = Italian Female "Sara".

Language codes (pass as lang_code in Python; aliases en-us→a, en-gb→b):

codelanguagecodelanguage
aAmerican EnglishiItalian
bBritish EnglishjJapanese
eSpanishpBrazilian Portuguese
fFrenchzMandarin Chinese
hHindi

Source: VOICES.md.

Quality is not uniform — pick by grade, not by name. VOICES.md assigns each voice an "Overall" grade combining a target-quality letter and how much training audio it received (more audio = higher grade). Documented highlights (verified 2026-07-10):

  • Best English female: af_heart (grade A, the card's default), af_bella (A−), af_nicole (B−, breathy/ASMR), bf_emma (B−, British).
  • Solid English: af_aoede, af_kore, af_sarah, am_fenrir, am_michael, am_puck (all C+). bm_fable, bm_george (C, British male).
  • Avoid unless you have a reason: many voices grade C or below; e.g. am_adam (F+), af_jessica / af_river (D). Most non-English voices are C/D, trained on ~minutes of synthetic data.

Production heuristic: default to af_heart (lang a). For a project the user should audition 3–4 A/B-graded voices before committing — grades predict, they do not guarantee, per-sentence quality.

Text length, chunking, and long-form synthesis

Documented limit: Kokoro processes at most 510 phonemized tokens per forward pass (512 with boundary tokens). VOICES.md notes voices "perform best on a goldilocks range of 100–200 tokens," are weak on very short utterances (<10–20 tokens)**, and **rush on long ones (>400). Source: model card / VOICES.md.

Consequences for production:

  • Never feed a whole chapter as one string. Split into sentences/paragraphs and stitch.
  • The Python KPipeline splits automatically; its split_pattern defaults to r'\n+' for English and returns one (graphemes, phonemes, audio) result per chunk, which you concatenate. Source: pipeline.py.
  • Non-English chunking is not fully implemented. Long non-English text can be truncated unless you pre-split it yourself (insert \n at sentence boundaries). This is a common silent-failure trap — verify non-English output length.
  • Heuristic for clean long-form: chunk to roughly 100–250 tokens at sentence boundaries (≈ one to three sentences), synthesize each, and concatenate with a short silence pad. Kokoro-FastAPI's defaults (~175 target / 250 / 450 absolute max tokens) are a reasonable starting point if you build your own splitter.

Pronunciation control (misaki + espeak-ng)

Kokoro does not read graphemes directly — text is converted to phonemes by misaki, hexgrad's G2P library, then fed to the model. English uses misaki's dictionary (spaCy + num2words). Out-of-dictionary words fall back to espeak-ng (EspeakFallback, on by default); espeak-ng is also the backbone for non-English G2P. Install espeak-ng as a system dependency or OOV words degrade to letter-by-letter spelling. Documented example: with fallback, eBook → ˈi bˈʊk; without it, → ˈiː bˈi ˈoʊ ˈoʊ kˈeɪ (spelled out). Source: misaki README.

To fix a mispronounced word (proper noun, brand, acronym, number read wrong):

  1. Inline phoneme override — misaki accepts a markdown-like syntax [word](/phonemes/), e.g. [Misaki](/misˈɑki/) or [Kokoro](/kˈOkəɹO/). Put the IPA/Kokoro phonemes between the slashes; stress marks like ˈ matter.
  2. Phonemize once, reuse — generate phonemes with misaki, hand-correct, and pass phonemes directly to the model so a batch job stays consistent.
  3. Spell it out in text — reword ("A-P-I", "twenty twenty-six") when phonemes are overkill.

Heuristic: always dry-run domain jargon, names, and numbers before a long batch — these are Kokoro's most common error class, and each is a one-line phoneme fix.

Voice blending (mixing)

A Kokoro "voice" is a style vector (voicepack tensor). Blending is a weighted average of two style vectors, which produces a new usable voice. The documented mechanism is a weighted numpy add, style1*(w0/100) + style2*(w1/100), with weights normalized if they don't sum to 100; several tools cap blending at exactly two voices. Source: nazdridoy/kokoro-tts voice-blending docs.

  • In Kokoro-FastAPI, request a blend by combining voice IDs: voice="af_sky+af_bella" (equal), or weighted per that server's syntax. Source: remsky/Kokoro-FastAPI.
  • In native Python you can load two voicepack tensors and average them yourself for full control (any ratio; nothing forces a 2-voice cap if you write the math).

Use blending to: nudge timbre/pitch between two graded voices, or build a house voice that isn't any single shipped one. It does not add emotion, create a new speaker identity from a reference, or rescue a low-grade voice — averaging two C-grade voices yields a C-grade blend.

Runtimes — pick by deployment target

1. Python kokoro package (default for servers/batch)

Best for narration pipelines, audiobooks, and anything on your own box. PyTorch backend; GPU optional. Example (labeled example — adapt paths/voices):

python
# pip install kokoro>=0.9.2 soundfile   ; plus system espeak-ng
from kokoro import KPipeline
import soundfile as sf
import numpy as np

pipeline = KPipeline(lang_code='a')          # 'a' = American English
text = "The quarterly report is ready.\nRevenue rose twelve percent."

chunks = []
for graphemes, phonemes, audio in pipeline(text, voice='af_heart', speed=1.0):
    chunks.append(audio)                     # one result per split (default r'\n+')

sf.write('out.wav', np.concatenate(chunks), 24000)   # 24 kHz mono

Why structured this way: KPipeline does G2P + chunking + inference; iterating yields per-chunk audio you concatenate, which is exactly the long-form pattern above. speed (~0.8–1.3) trades pace for naturalness. For lower-level control, KModel runs a single already-phonemized chunk. Source: model card, Python API.

Show full SKILL.md (808 more words)Show less
2. kokoro-js / ONNX (browser, on-device, Node)

Runs 100% client-side via Transformers.js — no server, nothing uploaded. Model id onnx-community/Kokoro-82M-v1.0-ONNX. Example (labeled example):

js
// npm i kokoro-js
import { KokoroTTS, TextSplitterStream } from "kokoro-js";

const tts = await KokoroTTS.from_pretrained("onnx-community/Kokoro-82M-v1.0-ONNX", {
  dtype: "q8",        // "fp32" | "fp16" | "q8" | "q4" | "q4f16"
  device: "webgpu",   // "wasm" | "webgpu" in-browser, "cpu" in Node ; use fp32 with webgpu
});

const audio = await tts.generate("Hello from the browser.", { voice: "af_heart" });
audio.save("audio.wav");                 // tts.list_voices() lists all IDs

// Streaming: push tokens, get audio incrementally
const splitter = new TextSplitterStream();
const stream = tts.stream(splitter);
(async () => { for await (const { text, phonemes, audio } of stream) audio.save("chunk.wav"); })();

Quantization trade-off: q8/q4 shrink download and speed WASM at some quality cost; fp32 is highest quality and is recommended with WebGPU. Source: kokoro-js README, onnx-community/Kokoro-82M-v1.0-ONNX.

For non-JS ONNX use, kokoro-onnx (Python) runs on onnxruntime (CPU) or onnxruntime-gpu (CUDA) and is what enables Raspberry-Pi / edge deployments.

3. Kokoro-FastAPI (drop-in OpenAI-compatible server)

The fastest way to give an existing app a local TTS backend: a Dockerized wrapper exposing an OpenAI-compatible /v1/audio/speech endpoint, so any client written for OpenAI TTS works by changing the base URL. Supports voice mixing (af_sky+af_bella), MP3/WAV/Opus/FLAC/M4A/PCM, streaming, per-word timestamps, a phoneme endpoint, and CPU/NVIDIA/AMD images. Example (labeled example):

python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8880/v1", api_key="not-needed")

client.audio.speech.create(
    model="kokoro", voice="af_heart", input="Local TTS, OpenAI-shaped API.",
    response_format="mp3",
).stream_to_file("out.mp3")

Source: remsky/Kokoro-FastAPI. Note it is a third-party wrapper (Apache-2.0-licensed model, separate project) — pin a version and verify its endpoint/voice-mixing syntax against its current README, as it evolves.

Hosted endpoints. Several inference platforms (e.g. Replicate, Baseten, and others) host Kokoro if you want the model's economics without self-hosting; those reintroduce a per-use cost and send text off-box, so they undercut the two main reasons to choose Kokoro. Prefer them only for burst capacity or when you can't run the model locally.

Performance expectations (secondary, dated evidence)

RTF (real-time factor) definitions differ between sources — some report audio-seconds-per-compute-second (higher = faster), others the inverse. Read the units.

  • GPU is dozens of times faster than real-time. One benchmark (PyTorch, ~16k chars, chunked ≤510 tokens) reports ~96× RTF on an A10G, ~81× on L4, ~36× on T4; ONNX ran lower (20–37×). Source: Kokoro v1 benchmark gist, retrieved 2026-07-10.
  • CPU is still comfortably faster than real-time on many cores — the same benchmark shows ~5× RTF on a 32-vCPU instance; another reports RTF ≈ 0.45–0.51 (i.e. ~2× real-time) on 4 cores. Source: gist above and a 4-core AMD EPYC run, retrieved 2026-07-10.
  • Footprint: 82M params is tiny; the model loads in well under a GB, and 4 GB RAM suffices for inference (8 GB+ for comfortable batching). Runs on modest hardware and Raspberry-Pi-class devices via ONNX.

Heuristic: for real-time or streaming UX, prefer GPU or a strong multi-core CPU; low-core/edge CPUs work for batch/offline but may fall near or below real-time on long text. Quantized ONNX (q8/q4) helps on constrained CPU/WASM at a quality cost.

Benchmark / ranking standing (dated, mixed evidence)

  • First-party claim (2024-12): the card states Kokoro v0.19 was "#1 ranked" in the TTS Spaces Arena in the weeks around its release. This was a limited-model / single-voice Arena setting — strong signal for its size, not a claim of beating all commercial models. Source: model card.
  • Broader arenas (2026): on wider TTS leaderboards Kokoro sits mid-pack among open-weight models. As of ~2026-03, one aggregated leaderboard placed Kokoro-82M v1.0 ~4th among open-weight models (Elo ≈ 1060), with newer/larger open models (e.g. Step Audio EditX, Elo ≈ 1118) ahead. Secondary source, retrieved 2026-07-10: TTS Arena / Artificial Analysis.

Honest framing for a user: Kokoro is exceptional for 82M parameters and $0 marginal cost, competitive with far larger models on neutral English narration, and clearly behind frontier commercial and larger open models on expressiveness, cloning, and multilingual fidelity. Sell it on economics, privacy, and footprint — not on being the highest-quality voice available.

Output review checklist

Before shipping Kokoro audio, listen for:

  • Mispronounced names / jargon / numbers — the top failure. Fix with a phoneme override.
  • Rushed or clipped delivery on chunks over ~400 tokens — re-split shorter.
  • Truncated non-English long text — verify duration; pre-split with \n.
  • Artifacts on very short lines (one or two words) — pad with context or a trailing period.
  • Wrong-accent voice for the language — match lang_code to the voice's prefix (a/b voices with English text, etc.); mismatches sound off.
  • Flat affect where the script needed emotion — if it reads wrong, the fix is a different model, not more retries.

Safety, licensing, and rights

  • Weights are Apache-2.0 — commercial use, redistribution, and fine-tuning are permitted with attribution. Verify the license of any wrapper (e.g. Kokoro-FastAPI) separately; they are distinct projects. (Verified 2026-07-10.)
  • Synthetic-data provenance: training included synthetic audio from closed TTS models. This is documented and the released data (Koniwa CC BY 3.0, SIWIS CC BY 4.0) is permissively licensed, but if a client has strict provenance requirements, disclose it.
  • No cloning ≠ no misuse risk. Even without cloning, generated speech can be used to impersonate a style or produce misleading audio. Don't generate audio that impersonates a real, identifiable person or is designed to deceive; disclose synthetic voice where the audience could reasonably assume it's human.
  • Voices are model artifacts, not real people — the names (Heart, Emma, George) are labels, not consenting individuals, so there's no per-speaker consent issue; the general synthetic-media disclosure norm still applies.

Sources (verified 2026-07-10)

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/providers/text-to-speech/kokoro-tts of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Kokoro Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kokoro Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kokoro Tts this skillcalesthio/generative-media-skills197—~4.8kAutomated safety check: PassMIT
Aivisspeech Engine Sentry TriageAivis-Project/AivisSpeech-Engine182—~545Automated safety check: PassLGPL-3.0
Video Translatorshang-zhu/violin1.1k—~1kAutomated safety check: NotesMIT
Azure Realtime Podcast Generationmicrosoft/skills3.1k1 repos~947Automated safety check: PassMIT
Webcode Local Windows Tts Installershuyu-labs/WebCode278—~787Automated safety check: PassCustom licence
Agentstadaspetra/loop2961 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Aivisspeech Engine Sentry Triage

    Aivis-Project/AivisSpeech-Engine

    AivisSpeech Engine の Sentry issue を調査し、修正すべきエンジン側の不具合と、入力値・ローカル環境・外部サービス由来のノイズを切り分けるためのスキルです。Sentry 側で既知ノイズを永続アーカイブする作業や、voicevoxengine/utility/sentryutility.py と関連テストを更新して既知ノイズを送信前に破棄する作業で使用します。

    182 GitHub stars~545 tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Translator

    shang-zhu/violin

    Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

    1.1k GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Media & CreativeAuto-check passed
  • A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model…

    278 GitHub stars~787 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Single File To Markdown

    impredicative/podgenai

    Convert a single provided file into a downloadable Markdown file, representing source images and figures as readable text.

    157 GitHub stars~174 tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    197 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    197 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    197 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    197 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    197 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Generated Media QA

    calesthio/generative-media-skills

    Provider-independent quality assurance for AI-generated and AI-assisted media.

    197 GitHub stars~8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Kokoro Tts

What does Kokoro Tts do?

Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad. Kokoro Tts is an agent skill from calesthio/generative-media-skills.0) model by hexgrad.

When should I use Kokoro Tts?

Kokoro Tts fits situations like: synthesizing speech offline; on your own server without per-character API cost — for high-volume narration; privacy-constrained pipelines.

How do I install Kokoro Tts in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill kokoro-tts -a claude-code`. Or copy the skill folder (skills/providers/text-to-speech/kokoro-tts in calesthio/generative-media-skills) into .claude/skills/kokoro-tts in your project. Claude Code loads it when a task matches its description.

How do I install Kokoro Tts in Codex?

Run `npx skills add calesthio/generative-media-skills --skill kokoro-tts -a codex`. Or copy the skill folder (skills/providers/text-to-speech/kokoro-tts in calesthio/generative-media-skills) into .agents/skills/kokoro-tts in your project. Codex loads it when a task matches its description.

Can I use Kokoro Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill kokoro-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kokoro-tts, .gemini/skills/kokoro-tts, .github/skills/kokoro-tts and .opencode/skills/kokoro-tts in your project.

What does Kokoro Tts need to run?

SKILL.md names no scripts, command-line tools or credentials: Kokoro Tts is instructions for the agent only. Our summary lists: Python 3; Node.js.

Does Kokoro Tts access the network?

SKILL.md names 7 domains. As links in the text: github.com, huggingface.co, deepwiki.com, gist.github.com, artificialanalysis.ai, pypi.org and npmjs.com. This is read from the text; nothing was executed.

Is Kokoro Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kokoro Tts use?

Kokoro Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kokoro Tts use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kokoro Tts?

Skills that share tags, products or a category with Kokoro Tts: Aivisspeech Engine Sentry Triage (Aivis-Project/AivisSpeech-Engine, 182 stars), Video Translator (shang-zhu/violin, 1.1k stars), Azure Realtime Podcast Generation (microsoft/skills, 3.1k stars) and Webcode Local Windows Tts Installer (shuyu-labs/WebCode, 278 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kokoro Tts?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.