Voiceover
GTKottman/mortiflix-oss
Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.
$ npx skills add veniceai/skills --skill venice-audio-speech -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install veniceai/skills venice-audio-speech --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-audio-speech .claude/skills/venice-audio-speech && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .claude/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/veniceai/skills/tree/main/skills/venice-audio-speechType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add veniceai/skills --skill venice-audio-speech -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install veniceai/skills venice-audio-speech --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/venice-audio-speech .agents/skills/venice-audio-speech && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .agents/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-audio-speech -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install veniceai/skills venice-audio-speech --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/venice-audio-speech .cursor/skills/venice-audio-speech && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .cursor/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/veniceai/skills.git --path skills/venice-audio-speech--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add veniceai/skills --skill venice-audio-speech -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install veniceai/skills venice-audio-speech --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/venice-audio-speech .gemini/skills/venice-audio-speech && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .gemini/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install veniceai/skills venice-audio-speechInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add veniceai/skills --skill venice-audio-speech -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/venice-audio-speech .github/skills/venice-audio-speech && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .github/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-audio-speech -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install veniceai/skills venice-audio-speech --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/venice-audio-speech .opencode/skills/venice-audio-speech && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech into .opencode/skills/venice-audio-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-speech", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
venice-audio-speechGenerate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.
Venice Audio Speech is an agent skill from veniceai/skills. Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output formats (modelspec.supportedformats / defaultformat), streaming, prompt/style control, temperature/topp, speed clamping, and language hints.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Text to speech and voice. It works with ElevenLabs, MiniMax and Qwen. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.
Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.venice.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VENICE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Venice Audio Speech loads about 3.6k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 1,527 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 1,527 words, ~3,592 tokens.
.claude/skills/venice-audio-speech/SKILL.md (or your agent's skills folder)./audio/speech)POST /api/v1/audio/speech converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's audio.speech.create() works as a drop-in.
| Method | Path | Auth | Notes |
|---|---|---|---|
POST | /api/v1/audio/speech | Bearer key or x402 (SIWX) | JSON body, returns raw audio. Billed per input character. |
POST | /api/v1/audio/voices | Bearer key or x402 (SIWX) | multipart/form-data. Clone a voice → vv_… handle. There is no GET /audio/voices. |
For music, sound effects and the character-priced ElevenLabs TTS models — v3, v4, v4 Turbo, Multilingual v2 — (async), see venice-audio-music. For transcription (audio → text), see venice-audio-transcription.
curl https://api.venice.ai/api/v1/audio/speech \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-xai-v1",
"voice": "eve",
"input": "Hello, welcome to Venice Voice.",
"response_format": "mp3",
"speed": 1.0,
"streaming": false
}' --output hello.mp3Response is the raw audio. Content-Type is audio/mpeg, audio/opus, audio/aac, audio/flac, audio/wav or audio/pcm depending on the format actually produced.
The body is strict — unknown fields return 400.
| Field | Type | Default | Notes |
|---|---|---|---|
input | string | — | Required. 1–4096 characters, must contain non-whitespace. Text that sanitizes to nothing speakable (e.g. only markup/emoji) → 400 "Input must contain speakable text". |
model | string | — | Send it. The OpenAPI schema lists a tts-kokoro default, but that default is never applied: omitting model returns 400 "model is required". Unknown id → 404. |
voice | string, ≤ 512 | the model's default voice | Voices are model-specific; a voice from another model → 400. Also accepts a cloned-voice handle (vv_…) from POST /audio/voices (same model that created it), and — on models with supports_custom_voice_id: true (currently tts-elevenlabs-turbo-v2-5) — a raw provider Voice ID. |
response_format | mp3 / opus / aac / flac / wav / pcm | the model's default_format | Support is per model — read model_spec.supported_formats / default_format. Requesting a format the model doesn't support → 400. |
speed | number | 1.0 | Schema range 0.25–4.0. Passed to Kokoro unchanged; clamped by xAI (0.7–1.5), ElevenLabs Turbo (0.7–1.2) and MiniMax (0.5–2); ignored by the other models. |
streaming | bool | false | true → chunked audio stream as it's generated. false → buffered file with Content-Length. |
language | string, 2–32 chars | — | Optional hint; form is model-specific (see below). Gemini Flash drops values outside its locale list and ElevenLabs Turbo drops values longer than 5 chars; Qwen 3, MiniMax and xAI forward the value as given, and a value the model rejects returns 400 "Invalid request parameters: language". Other models ignore it. |
prompt | string, ≤ 500 | — | Style/emotion instruction. Used by Qwen 3 and Gemini Flash (sent as style instructions); ignored elsewhere. |
temperature | number, 0–2 | — | Used by Qwen 3, Orpheus, Chatterbox HD, Gemini Flash; ignored elsewhere. |
top_p | number, 0–1 | — | Qwen 3 only; ignored elsewhere. |
language by model: Qwen 3 → full names (English, Chinese, …; default auto); xAI → ISO 639-1 (en), passed through as given, so send a valid code (default auto); ElevenLabs Turbo → ISO 639-1 (values longer than 5 chars dropped, shorter ones forwarded); MiniMax → full names (sent as a language boost, unchecked); Gemini Flash → full locale strings such as English (US) or Japanese (Japan) (anything else dropped). Kokoro, Inworld, Chatterbox, Orpheus and Gradium ignore it.
Every id below is in the live GET /models?type=tts list. Prices are model_spec.pricing.input.usd = USD per 1M input characters.
| Model ID | Default voice | Formats (default first) | Privacy | Notes |
|---|---|---|---|---|
tts-kokoro | af_sky | mp3, opus, aac, flac, wav, pcm | private | Multilingual via voice prefix. speed passed through unclamped. |
tts-qwen3-0-6b / tts-qwen3-1-7b | Vivian | mp3 | private | prompt, temperature, top_p, language. |
tts-xai-v1 | eve | mp3, wav, pcm | anonymized | 26 voices, ISO language. |
tts-inworld-1-5-max | Craig | wav | anonymized | Low-latency; all voices are English. |
tts-chatterbox-hd | Aurora | wav | private | temperature. Voice cloning (zero-shot). |
tts-orpheus | tara | wav | private | temperature. |
tts-elevenlabs-turbo-v2-5 | Rachel | mp3 | anonymized | Accepts raw ElevenLabs Voice IDs as voice. |
tts-minimax-speech-02-hd | WiseWoman | mp3, pcm, flac | anonymized | language. Cloning with this model is not available to regular keys (see below). |
tts-gemini-3-1-flash | Kore | mp3, opus, wav | anonymized | prompt, temperature, locale language. |
tts-gradium-v1 | Emma | wav, pcm, opus | anonymized | The voice picks the language; no language param. |
Always inspect GET /models?type=tts before calling: model_spec.voices (authoritative voice list), supported_formats, default_format, supports_custom_voice_id, privacy, pricing, and — on cloning models — voice_cloning. Per-model parameter support (prompt / temperature / top_p / language) is not exposed on /models; use the table above.
A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized models (PRIVATE_TEXT keys are not restricted here).
model_spec.voices)Case-sensitive. Omit voice to get the model's default.
<lang><gender>_<name>: a American, b British, z Chinese, f French, h Hindi, i Italian, j Japanese, p Portuguese, e Spanish; f/m gender. Examples: af_sky, af_bella, af_heart, am_adam, am_michael, bf_emma, bm_george, ff_siwis, jf_alpha, zf_xiaoxiao, ef_dora, pm_alex (54 total).Vivian, Serena, Ono_Anna, Sohee, Uncle_Fu, Dylan, Eric, Ryan, Aideneve, ara, rex, sal, leo, altair, atlas, carina, castor, celeste, cosmo, helios, helix, iris, kepler, lumen, luna, lux, naksh, orion, perseus, rigel, sirius, ursa, zagan, zenithtara, leah, jess, mia, zoe, leo, dan, zacCraig, Ashley, Olivia, Sarah, Elizabeth, Priya, Alex, Edward, Theodore, Ronald, Mark, Hades, Luna, PixieAurora, Britney, Siobhan, Vicky, Blade, Carl, Cliff, Richard, RicoRachel, Aria, Sarah, Laura, Charlotte, Alice, Matilda, Jessica, Lily, Roger, Charlie, George, Callum, River, Liam, Will, Eric, Chris, Brian, Daniel, Bill — or any ElevenLabs Voice IDWiseWoman, FriendlyPerson, InspirationalGirl, CalmWoman, LivelyGirl, LovelyGirl, SweetGirl, ExuberantGirl, DeepVoiceMan, CasualGuy, PatientMan, YoungKnight, DeterminedMan, ImposingManner, ElegantManAchernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, ZubenelgenubiEmma, Kent, Eva, Jack. German: Mia, Maximilian. Spanish: Valentina, Sergio. French: Elise, Leo. Portuguese: Alice, DaviA voice not in the chosen model's list (and not a valid handle / custom Voice ID where allowed) → 400. An ElevenLabs Voice ID that the provider rejects also → 400 (not charged).
POST /audio/voicesClone a voice from an audio sample and get back a handle (vv_…) to pass as voice on /audio/speech. multipart/form-data only; max 25 MB.
curl https://api.venice.ai/api/v1/audio/voices \
-H "Authorization: Bearer $VENICE_API_KEY" \
-F "model=tts-chatterbox-hd" \
-F "file=@sample.wav"{ "id": "vv_…", "model": "tts-chatterbox-hd" }| Field | Notes |
|---|---|
file | The voice sample, multipart field file. Validated by extension/MIME and binary signature. Aim for a clean speech recording of at least 5 s (voice_cloning.min_sample_seconds; advisory — Venice does not measure duration). |
model | Defaults to tts-chatterbox-hd, the only cloning model open to regular API keys. |
tts-chatterbox-hd advertises its cloning contract on /models as model_spec.voice_cloning:
{ "mode": "zero_shot", "accepted_formats": ["mp3", "wav", "flac", "mp4"], "min_sample_seconds": 5, "retention_days": 7 }mp4 covers M4A. Samples in other containers (or with a non-audio signature) → 400 before anything is uploaded./audio/speech.tts-minimax-speech-02-hd also appears in the model enum in the OpenAPI spec, but cloning with it isn't open to regular keys, which get 403 "Voice cloning … is not available on your account". Its model spec on /models carries no voice_cloning object — use that as the signal.
A handle is bound to the model that created it. Pass it with a model that has no cloning support → 400; pairing it with a different cloning model fails.
{
"model": "tts-xai-v1",
"voice": "eve",
"input": "Hello, this is a long document to narrate. ...",
"streaming": true,
"response_format": "mp3"
}With streaming: true, the body is a chunked (Transfer-Encoding: chunked) audio stream — decode as it arrives. pcm (where supported: Kokoro, xAI, MiniMax, Gradium) is convenient for raw Web Audio playback. If the stream fails before headers are sent you get 500 {"error":"Stream error"}; after that the connection just ends.
import OpenAI from 'openai'
import fs from 'node:fs/promises'
const client = new OpenAI({
apiKey: process.env.VENICE_API_KEY,
baseURL: 'https://api.venice.ai/api/v1',
})
const mp3 = await client.audio.speech.create({
model: 'tts-xai-v1',
voice: 'eve',
input: 'Hello from Venice.',
response_format: 'mp3',
})
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer())){
"model": "tts-qwen3-1-7b",
"voice": "Vivian",
"input": "We did it!",
"prompt": "Excited and energetic.",
"temperature": 0.9,
"top_p": 0.95
}tts-gemini-3-1-flash also takes prompt (style instructions) and temperature, but not top_p. For other families, delivery comes from the voice choice itself (e.g. Inworld Hades vs Pixie); prompt / temperature / top_p are silently ignored.
| Code | Meaning |
|---|---|
400 | Missing model, schema error (strict body, input > 4096 / empty / unspeakable), voice not valid for the model, unsupported response_format, bad cloning sample (/audio/voices), handle paired with a non-cloning model. |
401 | Authentication failed. |
402 | Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD |
403 | A PRIVATE_ONLY key calling an anonymized model, region restriction, or a cloning model not open to your account on /audio/voices. |
404 | Unknown model. |
413 | /audio/voices sample over 25 MB. |
429 | Rate limited. |
500 | Inference failure / stream error. |
502 | Temporary upstream TTS failure — {"error":"Speech synthesis failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff. |
503 | Model temporarily offline — retry with jitter. |
See venice-errors for body shapes and retry strategy.
model — the documented default never applies.input hard cap is 4096 chars. For long content, split on sentence boundaries and concatenate audio client-side.mp3: Inworld, Chatterbox, Orpheus and Gradium default to wav, and any response_format outside a model's supported_formats returns 400. Omit response_format or check /models first.speed is only applied by Kokoro (unclamped), xAI, ElevenLabs Turbo and MiniMax (clamped). Keep 0.8–1.3 for natural narration.streaming is a Venice-specific field that isn't in the OpenAI SDK's types; pass it as an extra body field, or call the REST endpoint directly and consume the body.eve ≠ Eve, af_sky ≠ AF_SKY). Note leo (xAI/Orpheus) vs Leo (Gradium, French).language parameter — pick the voice for the language you want.© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/venice-audio-speech of veniceai/skills.
Open the folder on GitHubat commit 5eaeac5
Venice Audio Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Venice Audio Speech this skillveniceai/skills | 144 | — | ~3.6k | Automated safety check: Pass | MIT | |
| VoiceoverGTKottman/mortiflix-oss | 499 | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| Audio Jinglesanqiufong/slides-from-anything | 132 | 1 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Voice Clone Ttsnpc-live/clawfirm | 156 | — | ~1.1k | Automated safety check: Pass | None | |
| Web Video PresentationConardLi/garden-skills | 13k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Quota Axikunchenguid/quota-axi | 147 | — | ~547 | Automated safety check: Pass | MIT |
GTKottman/mortiflix-oss
Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…
sanqiufong/slides-from-anything
Audio generation skill — jingles, beds, voiceover, and sound effects.
npc-live/clawfirm
声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。
ConardLi/garden-skills
Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.
kunchenguid/quota-axi
Report local Claude, Codex, Cursor, GitHub Copilot, Grok, Kimi, Z.AI, Alibaba, OpenCode Go, Antigravity, Command Code, MiniMax, MiMo, DeepSeek, OpenRouter, ElevenLabs, Devin, Muse, and Higgsfield…
tadaspetra/loop
Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.
veniceai/skills
Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.
veniceai/skills
Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.
veniceai/skills
Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.
veniceai/skills
High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.
veniceai/skills
Async music, sound-effect and long-form voice generation via Venice.
veniceai/skills
Transcribe audio files to text via POST /audio/transcriptions.
Works with
Categories
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Venice Audio Speech is an agent skill from veniceai/skills. Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.
Venice Audio Speech fits situations like: tasks that involve Text to speech and voice.
Run `npx skills add veniceai/skills --skill venice-audio-speech -a claude-code`. Or copy the skill folder (skills/venice-audio-speech in veniceai/skills) into .claude/skills/venice-audio-speech in your project. Claude Code loads it when a task matches its description.
Run `npx skills add veniceai/skills --skill venice-audio-speech -a codex`. Or copy the skill folder (skills/venice-audio-speech in veniceai/skills) into .agents/skills/venice-audio-speech in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-audio-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-audio-speech, .gemini/skills/venice-audio-speech, .github/skills/venice-audio-speech and .opencode/skills/venice-audio-speech in your project.
Going by SKILL.md and its folder, Venice Audio Speech needs the command-line tools its instructions call (curl) and credentials named VENICE_API_KEY. Our summary lists: A credential in VENICE_API_KEY.
SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Venice Audio Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Venice Audio Speech: Voiceover (GTKottman/mortiflix-oss, 499 stars), Audio Jingle (sanqiufong/slides-from-anything, 132 stars), Voice Clone Tts (npc-live/clawfirm, 156 stars) and Web Video Presentation (ConardLi/garden-skills, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.
Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.