Agent skill

Venice Audio Speech

by veniceai in veniceai/skills

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

MITAuto-check passedMedia & Creative

Install Venice Audio Speech

skills CLI
$ npx skills add veniceai/skills --skill venice-audio-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-audio-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-audio-speech .claude/skills/venice-audio-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-audio-speech
GitHub stars
144
Token cost
~3.6k tokens
SKILL.md length
1,527 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

  • Tasks that involve Text to speech and voice
  • SKILL.md covers Use when, Minimal request, Request schema and Models, plus 7 more sections
  • Calls curl; reaches api.venice.ai; needs VENICE_API_KEY

What it does

Venice Audio Speech is an agent skill from veniceai/skills. Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output formats (modelspec.supportedformats / defaultformat), streaming, prompt/style control, temperature/topp, speed clamping, and language hints.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with ElevenLabs, MiniMax and Qwen. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/venice-audio-speech”

Requirements

  • A credential in VENICE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Audio Speech loads about 3.6k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 1,527 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 1,527 words, ~3,592 tokens.

Download SKILL.mdSave it as .claude/skills/venice-audio-speech/SKILL.md (or your agent's skills folder).
name
venice-audio-speech
description
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output formats (model_spec.supported_formats / default_format), streaming, prompt/style control, temperature/top_p, speed clamping, and language hints.

Venice TTS (/audio/speech)

POST /api/v1/audio/speech converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's audio.speech.create() works as a drop-in.

MethodPathAuthNotes
POST/api/v1/audio/speechBearer key or x402 (SIWX)JSON body, returns raw audio. Billed per input character.
POST/api/v1/audio/voicesBearer key or x402 (SIWX)multipart/form-data. Clone a voice → vv_… handle. There is no GET /audio/voices.

Use when

  • You want narration, voice replies, or UI audio from text.
  • You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash, Gradium).
  • You want streaming audio returned as it is generated.
  • You need style/emotion control on supported models, or synthesis in a cloned voice.

For music, sound effects and the character-priced ElevenLabs TTS models — v3, v4, v4 Turbo, Multilingual v2 — (async), see venice-audio-music. For transcription (audio → text), see venice-audio-transcription.

Minimal request

bash
curl https://api.venice.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-xai-v1",
    "voice": "eve",
    "input": "Hello, welcome to Venice Voice.",
    "response_format": "mp3",
    "speed": 1.0,
    "streaming": false
  }' --output hello.mp3

Response is the raw audio. Content-Type is audio/mpeg, audio/opus, audio/aac, audio/flac, audio/wav or audio/pcm depending on the format actually produced.

Request schema

The body is strict — unknown fields return 400.

FieldTypeDefaultNotes
inputstring—Required. 1–4096 characters, must contain non-whitespace. Text that sanitizes to nothing speakable (e.g. only markup/emoji) → 400 "Input must contain speakable text".
modelstring—Send it. The OpenAPI schema lists a tts-kokoro default, but that default is never applied: omitting model returns 400 "model is required". Unknown id → 404.
voicestring, ≤ 512the model's default voiceVoices are model-specific; a voice from another model → 400. Also accepts a cloned-voice handle (vv_…) from POST /audio/voices (same model that created it), and — on models with supports_custom_voice_id: true (currently tts-elevenlabs-turbo-v2-5) — a raw provider Voice ID.
response_formatmp3 / opus / aac / flac / wav / pcmthe model's default_formatSupport is per model — read model_spec.supported_formats / default_format. Requesting a format the model doesn't support → 400.
speednumber1.0Schema range 0.25–4.0. Passed to Kokoro unchanged; clamped by xAI (0.7–1.5), ElevenLabs Turbo (0.7–1.2) and MiniMax (0.5–2); ignored by the other models.
streamingboolfalsetrue → chunked audio stream as it's generated. false → buffered file with Content-Length.
languagestring, 2–32 chars—Optional hint; form is model-specific (see below). Gemini Flash drops values outside its locale list and ElevenLabs Turbo drops values longer than 5 chars; Qwen 3, MiniMax and xAI forward the value as given, and a value the model rejects returns 400 "Invalid request parameters: language". Other models ignore it.
promptstring, ≤ 500—Style/emotion instruction. Used by Qwen 3 and Gemini Flash (sent as style instructions); ignored elsewhere.
temperaturenumber, 0–2—Used by Qwen 3, Orpheus, Chatterbox HD, Gemini Flash; ignored elsewhere.
top_pnumber, 0–1—Qwen 3 only; ignored elsewhere.

language by model: Qwen 3 → full names (English, Chinese, …; default auto); xAI → ISO 639-1 (en), passed through as given, so send a valid code (default auto); ElevenLabs Turbo → ISO 639-1 (values longer than 5 chars dropped, shorter ones forwarded); MiniMax → full names (sent as a language boost, unchecked); Gemini Flash → full locale strings such as English (US) or Japanese (Japan) (anything else dropped). Kokoro, Inworld, Chatterbox, Orpheus and Gradium ignore it.

Models

Every id below is in the live GET /models?type=tts list. Prices are model_spec.pricing.input.usd = USD per 1M input characters.

Model IDDefault voiceFormats (default first)PrivacyNotes
tts-kokoroaf_skymp3, opus, aac, flac, wav, pcmprivateMultilingual via voice prefix. speed passed through unclamped.
tts-qwen3-0-6b / tts-qwen3-1-7bVivianmp3privateprompt, temperature, top_p, language.
tts-xai-v1evemp3, wav, pcmanonymized26 voices, ISO language.
tts-inworld-1-5-maxCraigwavanonymizedLow-latency; all voices are English.
tts-chatterbox-hdAurorawavprivatetemperature. Voice cloning (zero-shot).
tts-orpheustarawavprivatetemperature.
tts-elevenlabs-turbo-v2-5Rachelmp3anonymizedAccepts raw ElevenLabs Voice IDs as voice.
tts-minimax-speech-02-hdWiseWomanmp3, pcm, flacanonymizedlanguage. Cloning with this model is not available to regular keys (see below).
tts-gemini-3-1-flashKoremp3, opus, wavanonymizedprompt, temperature, locale language.
tts-gradium-v1Emmawav, pcm, opusanonymizedThe voice picks the language; no language param.

Always inspect GET /models?type=tts before calling: model_spec.voices (authoritative voice list), supported_formats, default_format, supports_custom_voice_id, privacy, pricing, and — on cloning models — voice_cloning. Per-model parameter support (prompt / temperature / top_p / language) is not exposed on /models; use the table above.

A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized models (PRIVATE_TEXT keys are not restricted here).

Voices (from model_spec.voices)

Case-sensitive. Omit voice to get the model's default.

  • Kokoro — <lang><gender>_<name>: a American, b British, z Chinese, f French, h Hindi, i Italian, j Japanese, p Portuguese, e Spanish; f/m gender. Examples: af_sky, af_bella, af_heart, am_adam, am_michael, bf_emma, bm_george, ff_siwis, jf_alpha, zf_xiaoxiao, ef_dora, pm_alex (54 total).
  • Qwen 3 — Vivian, Serena, Ono_Anna, Sohee, Uncle_Fu, Dylan, Eric, Ryan, Aiden
  • xAI — eve, ara, rex, sal, leo, altair, atlas, carina, castor, celeste, cosmo, helios, helix, iris, kepler, lumen, luna, lux, naksh, orion, perseus, rigel, sirius, ursa, zagan, zenith
  • Orpheus — tara, leah, jess, mia, zoe, leo, dan, zac
  • Inworld — Craig, Ashley, Olivia, Sarah, Elizabeth, Priya, Alex, Edward, Theodore, Ronald, Mark, Hades, Luna, Pixie
  • Chatterbox — Aurora, Britney, Siobhan, Vicky, Blade, Carl, Cliff, Richard, Rico
  • ElevenLabs Turbo — Rachel, Aria, Sarah, Laura, Charlotte, Alice, Matilda, Jessica, Lily, Roger, Charlie, George, Callum, River, Liam, Will, Eric, Chris, Brian, Daniel, Bill — or any ElevenLabs Voice ID
  • MiniMax — WiseWoman, FriendlyPerson, InspirationalGirl, CalmWoman, LivelyGirl, LovelyGirl, SweetGirl, ExuberantGirl, DeepVoiceMan, CasualGuy, PatientMan, YoungKnight, DeterminedMan, ImposingManner, ElegantMan
  • Gemini Flash — Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi
  • Gradium — English: Emma, Kent, Eva, Jack. German: Mia, Maximilian. Spanish: Valentina, Sergio. French: Elise, Leo. Portuguese: Alice, Davi

A voice not in the chosen model's list (and not a valid handle / custom Voice ID where allowed) → 400. An ElevenLabs Voice ID that the provider rejects also → 400 (not charged).

Show full SKILL.md (593 more words)Show less

Voice cloning — POST /audio/voices

Clone a voice from an audio sample and get back a handle (vv_…) to pass as voice on /audio/speech. multipart/form-data only; max 25 MB.

bash
curl https://api.venice.ai/api/v1/audio/voices \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -F "model=tts-chatterbox-hd" \
  -F "file=@sample.wav"
json
{ "id": "vv_…", "model": "tts-chatterbox-hd" }
FieldNotes
fileThe voice sample, multipart field file. Validated by extension/MIME and binary signature. Aim for a clean speech recording of at least 5 s (voice_cloning.min_sample_seconds; advisory — Venice does not measure duration).
modelDefaults to tts-chatterbox-hd, the only cloning model open to regular API keys.

tts-chatterbox-hd advertises its cloning contract on /models as model_spec.voice_cloning:

json
{ "mode": "zero_shot", "accepted_formats": ["mp3", "wav", "flac", "mp4"], "min_sample_seconds": 5, "retention_days": 7 }
  • Zero-shot: no voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles stop working 7 days after creation, regardless of use.
  • mp4 covers M4A. Samples in other containers (or with a non-audio signature) → 400 before anything is uploaded.
  • Each successful clone is charged a flat per-clone fee; synthesis is billed separately per character on /audio/speech.

tts-minimax-speech-02-hd also appears in the model enum in the OpenAPI spec, but cloning with it isn't open to regular keys, which get 403 "Voice cloning … is not available on your account". Its model spec on /models carries no voice_cloning object — use that as the signal.

A handle is bound to the model that created it. Pass it with a model that has no cloning support → 400; pairing it with a different cloning model fails.

Streaming

json
{
  "model": "tts-xai-v1",
  "voice": "eve",
  "input": "Hello, this is a long document to narrate. ...",
  "streaming": true,
  "response_format": "mp3"
}

With streaming: true, the body is a chunked (Transfer-Encoding: chunked) audio stream — decode as it arrives. pcm (where supported: Kokoro, xAI, MiniMax, Gradium) is convenient for raw Web Audio playback. If the stream fails before headers are sent you get 500 {"error":"Stream error"}; after that the connection just ends.

OpenAI SDK

ts
import OpenAI from 'openai'
import fs from 'node:fs/promises'

const client = new OpenAI({
  apiKey: process.env.VENICE_API_KEY,
  baseURL: 'https://api.venice.ai/api/v1',
})

const mp3 = await client.audio.speech.create({
  model: 'tts-xai-v1',
  voice: 'eve',
  input: 'Hello from Venice.',
  response_format: 'mp3',
})

await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer()))

Style / emotion (Qwen 3, Gemini Flash)

json
{
  "model": "tts-qwen3-1-7b",
  "voice": "Vivian",
  "input": "We did it!",
  "prompt": "Excited and energetic.",
  "temperature": 0.9,
  "top_p": 0.95
}

tts-gemini-3-1-flash also takes prompt (style instructions) and temperature, but not top_p. For other families, delivery comes from the voice choice itself (e.g. Inworld Hades vs Pixie); prompt / temperature / top_p are silently ignored.

Errors

CodeMeaning
400Missing model, schema error (strict body, input > 4096 / empty / unspeakable), voice not valid for the model, unsupported response_format, bad cloning sample (/audio/voices), handle paired with a non-cloning model.
401Authentication failed.
402Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD
403A PRIVATE_ONLY key calling an anonymized model, region restriction, or a cloning model not open to your account on /audio/voices.
404Unknown model.
413/audio/voices sample over 25 MB.
429Rate limited.
500Inference failure / stream error.
502Temporary upstream TTS failure — {"error":"Speech synthesis failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff.
503Model temporarily offline — retry with jitter.

See venice-errors for body shapes and retry strategy.

Gotchas

  • Always send model — the documented default never applies.
  • input hard cap is 4096 chars. For long content, split on sentence boundaries and concatenate audio client-side.
  • Don't assume mp3: Inworld, Chatterbox, Orpheus and Gradium default to wav, and any response_format outside a model's supported_formats returns 400. Omit response_format or check /models first.
  • speed is only applied by Kokoro (unclamped), xAI, ElevenLabs Turbo and MiniMax (clamped). Keep 0.8–1.3 for natural narration.
  • streaming is a Venice-specific field that isn't in the OpenAI SDK's types; pass it as an extra body field, or call the REST endpoint directly and consume the body.
  • Voice names are case-sensitive (eve ≠ Eve, af_sky ≠ AF_SKY). Note leo (xAI/Orpheus) vs Leo (Gradium, French).
  • Chatterbox cloned handles expire 7 days after creation. Re-clone rather than storing handles long-term.
  • Gradium has no language parameter — pick the voice for the language you want.

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/venice-audio-speech of veniceai/skills.

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Audio Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Audio Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Audio Speech this skillveniceai/skills144—~3.6kAutomated safety check: PassMIT
VoiceoverGTKottman/mortiflix-oss499—~1.8kAutomated safety check: PassAGPL-3.0
Audio Jinglesanqiufong/slides-from-anything1321 repos~1.1kAutomated safety check: PassApache-2.0
Voice Clone Ttsnpc-live/clawfirm156—~1.1kAutomated safety check: PassNone
Web Video PresentationConardLi/garden-skills13k—~3.5kAutomated safety check: PassMIT
Quota Axikunchenguid/quota-axi147—~547Automated safety check: PassMIT

Similar skills

  • Voiceover

    GTKottman/mortiflix-oss

    Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…

    499 GitHub stars~1.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Audio Jingle

    sanqiufong/slides-from-anything

    Audio generation skill — jingles, beds, voiceover, and sound effects.

    132 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Voice Clone Tts

    npc-live/clawfirm

    声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。

    156 GitHub stars~1.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Web Video Presentation

    ConardLi/garden-skills

    Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.

    13k GitHub stars~3.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Quota Axi

    kunchenguid/quota-axi

    Report local Claude, Codex, Cursor, GitHub Copilot, Grok, Kimi, Z.AI, Alibaba, OpenCode Go, Antigravity, Command Code, MiniMax, MiMo, DeepSeek, OpenRouter, ElevenLabs, Devin, Muse, and Higgsfield…

    147 GitHub stars~547 tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

    144 GitHub stars~5.2k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed
  • Transcribe audio files to text via POST /audio/transcriptions.

    144 GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed

Questions about Venice Audio Speech

What does Venice Audio Speech do?

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Venice Audio Speech is an agent skill from veniceai/skills. Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

When should I use Venice Audio Speech?

Venice Audio Speech fits situations like: tasks that involve Text to speech and voice.

How do I install Venice Audio Speech in Claude Code?

Run `npx skills add veniceai/skills --skill venice-audio-speech -a claude-code`. Or copy the skill folder (skills/venice-audio-speech in veniceai/skills) into .claude/skills/venice-audio-speech in your project. Claude Code loads it when a task matches its description.

How do I install Venice Audio Speech in Codex?

Run `npx skills add veniceai/skills --skill venice-audio-speech -a codex`. Or copy the skill folder (skills/venice-audio-speech in veniceai/skills) into .agents/skills/venice-audio-speech in your project. Codex loads it when a task matches its description.

Can I use Venice Audio Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-audio-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-audio-speech, .gemini/skills/venice-audio-speech, .github/skills/venice-audio-speech and .opencode/skills/venice-audio-speech in your project.

What does Venice Audio Speech need to run?

Going by SKILL.md and its folder, Venice Audio Speech needs the command-line tools its instructions call (curl) and credentials named VENICE_API_KEY. Our summary lists: A credential in VENICE_API_KEY.

Does Venice Audio Speech access the network?

SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Venice Audio Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Venice Audio Speech use?

Venice Audio Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Audio Speech use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Audio Speech?

Skills that share tags, products or a category with Venice Audio Speech: Voiceover (GTKottman/mortiflix-oss, 499 stars), Audio Jingle (sanqiufong/slides-from-anything, 132 stars), Voice Clone Tts (npc-live/clawfirm, 156 stars) and Web Video Presentation (ConardLi/garden-skills, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Audio Speech?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.