Agent skill

Venice Audio Music

by veniceai in veniceai/skills

Async music, sound-effect and long-form voice generation via Venice.

MITAuto-check passedMedia & Creative

Install Venice Audio Music

skills CLI
$ npx skills add veniceai/skills --skill venice-audio-music -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-audio-music --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-audio-music .claude/skills/venice-audio-music && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-audio-music
GitHub stars
144
Token cost
~3.1k tokens
SKILL.md length
999 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Async music, sound-effect and long-form voice generation via Venice.

  • Works in 4 steps: POST /audio/quote — price it first → POST /audio/queue — enqueue → POST /audio/retrieve — poll / download → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Use when, Models, Lifecycle and Full loop (TypeScript), plus 3 more sections
  • Calls curl; reaches api.venice.ai; needs VENICE_API_KEY

What it does

Venice Audio Music is an agent skill from veniceai/skills. Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenLabs Voice IDs, language, speed, model capability probing via /models?type=music, pricing shapes, refunds, and polling.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and Music and audio generation. It works with ElevenLabs, MiniMax and x402. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Music and audio generation

Example prompts

  • “/venice-audio-music”

Requirements

  • A credential in VENICE_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. POST /audio/quote — price it first
  2. POST /audio/queue — enqueue
  3. POST /audio/retrieve — poll / download
  4. POST /audio/complete — cleanup

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Audio Music loads about 3.1k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 999 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 999 words, ~3,075 tokens.

Download SKILL.mdSave it as .claude/skills/venice-audio-music/SKILL.md (or your agent's skills folder).
name
venice-audio-music
description
Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenLabs Voice IDs, language, speed, model capability probing via /models?type=music, pricing shapes, refunds, and polling.

Venice Music / Async Audio

Music, sound effects and character-priced voice generation are asynchronous:

MethodPathAuthNotes
POST/api/v1/audio/quoteNone required (key optional)Price in USD.
POST/api/v1/audio/queueBearer key or x402 (SIWX)Charges and enqueues → queue_id. 40 req/min per user.
POST/api/v1/audio/retrieveBearer key or x402 (SIWX)Status JSON or the audio bytes. 120 req/min per user.
POST/api/v1/audio/completeBearer key or x402 (SIWX)Delete the stored media.

For short synchronous text-to-speech use venice-audio-speech. Voice-changer (speech-to-speech) models are refused on these four endpoints with a 400 pointing at /audio/voice-changer/* (callers who can't see the model get 404 instead) — see venice-audio-voice-changer.

Use when

  • You need songs, jingles, score, soundscapes, sound effects, or long narration.
  • The model uses duration-, per-second-, per-job- or character-based pricing and you want a price before submitting.
  • Generation takes long enough that a synchronous call would time out.

Models

Query GET /models?type=music for the current list and each model's model_spec. Representative ids (all in the live list):

KindExamples
Instrumental / songselevenlabs-music, elevenlabs-music-v2-5, lyria-3-pro, sonilo-v1-1-music, stable-audio-25
Songs with lyricsminimax-music-v25, minimax-music-v26, minimax-music-v2 (lyrics required), ace-step-15 (lyrics optional)
Sound effectselevenlabs-sound-effects-v2 (supports loop), sonilo-v1-1-sound-effects, mmaudio-v2-text-to-audio
Voice (text in prompt)elevenlabs-tts-v4, elevenlabs-tts-v4-turbo, elevenlabs-tts-v3, elevenlabs-tts-multilingual-v2, seed-audio-1-0

Lifecycle

1. POST /audio/quote — price it first
bash
curl https://api.venice.ai/api/v1/audio/quote \
  -H "Content-Type: application/json" \
  -d '{ "model": "elevenlabs-music", "duration_seconds": 60 }'

Response: {"quote": 0.69} (USD). No API key needed; sending one lets you price models only your account can see.

FieldNotes
modelRequired.
duration_secondsInteger or numeric string. Only for models that expose duration metadata (min_duration / max_duration / duration_options) — rejected otherwise. Omit to price the model's default_duration.
character_countInteger ≤ prompt_character_limit. Required for models priced by per_thousand_characters.

Unknown fields → 400.

2. POST /audio/queue — enqueue
bash
curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs-music",
    "prompt": "Uplifting indie-folk acoustic track, 120 BPM, major key.",
    "duration_seconds": 60,
    "force_instrumental": true
  }'

Song with lyrics (minimax-music-v25 supports lyrics, the optimizer and instrumental mode, but no duration_seconds):

json
{
  "model": "minimax-music-v25",
  "prompt": "Warm indie-pop ballad, female vocals, acoustic guitar.",
  "lyrics_prompt": "[Verse]\nWalking through the city lights...\n[Chorus]\nWe are the dreamers..."
}

Response: { "model": "...", "queue_id": "...", "status": "QUEUED" }.

The body is strict: every optional field below is rejected with 400 when the model's model_spec says it isn't supported.

FieldNotes
modelRequired.
promptRequired. Between min_prompt_length (default 10) and prompt_character_limit; trimmed. For the voice models this is the text to speak.
lyrics_promptUp to lyrics_character_limit (default 4096). Required when lyrics_required=true; rejected when supports_lyrics=false.
duration_secondsInteger or numeric string. Must be one of duration_options when present, else within min_duration–max_duration. Defaults to default_duration.
force_instrumentalsupports_force_instrumental=true only.
lyrics_optimizerAuto-writes lyrics from prompt. supports_lyrics_optimizer=true only; lyrics_prompt must then be empty.
loopRender a seamless loop (end splices into start). supports_loop=true only — currently elevenlabs-sound-effects-v2.
voiceVoice-enabled models only. One of voices; defaults to default_voice. Models with supports_custom_voice_id=true (the ElevenLabs TTS models) also accept a raw ElevenLabs Voice ID.
language_codeISO 639-1. supports_language_code=true only — no model in the current list sets it.
speedsupports_speed=true only, within min_speed–max_speed.

Model-specific rules also apply: minimax-music-v25 needs a lyrics_prompt of at least 10 chars unless force_instrumental or lyrics_optimizer is true; minimax-music-v26 needs the same unless force_instrumental is true; minimax-music-v2 needs a non-blank lyrics_prompt.

3. POST /audio/retrieve — poll / download
bash
curl https://api.venice.ai/api/v1/audio/retrieve \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"elevenlabs-music","queue_id":"..."}' \
  --output track.mp3
  • Still running: 200 JSON {"status":"PROCESSING","average_execution_time":<ms, P80 estimate>,"execution_duration":<ms since queued>}.
  • Done: 200 with the audio bytes. Content-Type is the audio type; headers x-venice-audio-format, x-venice-inference-time (s), x-venice-model-id, x-venice-model-name, and for Seed Audio also x-venice-audio-duration and x-venice-audio-subtitle.
  • delete_media_on_completion: true deletes the media after this download, so you can skip step 4.

If generation fails (content policy, capacity, provider validation), the charge is refunded (except a DIEM charge from a previous epoch) and the error is returned here.

4. POST /audio/complete — cleanup
bash
curl https://api.venice.ai/api/v1/audio/complete \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"elevenlabs-music","queue_id":"..."}'

Returns {"success": true} once the stored media is deleted (false if the delete didn't go through). Use it after you've saved the bytes, unless you retrieved with delete_media_on_completion: true.

Full loop (TypeScript)

ts
import fs from 'node:fs/promises'

const base = 'https://api.venice.ai/api/v1'
const headers = {
  Authorization: `Bearer ${process.env.VENICE_API_KEY}`,
  'Content-Type': 'application/json',
}

async function generateTrack() {
  // 1. Quote
  const quote = await fetch(`${base}/audio/quote`, {
    method: 'POST', headers,
    body: JSON.stringify({ model: 'elevenlabs-music', duration_seconds: 60 }),
  }).then(r => r.json())
  console.log('price:', quote.quote)

  // 2. Queue
  const { queue_id, model } = await fetch(`${base}/audio/queue`, {
    method: 'POST', headers,
    body: JSON.stringify({
      model: 'elevenlabs-music',
      prompt: 'Uplifting indie-folk acoustic track, 120 BPM.',
      duration_seconds: 60,
      force_instrumental: true,
    }),
  }).then(r => r.json())

  // 3. Poll
  while (true) {
    const res = await fetch(`${base}/audio/retrieve`, {
      method: 'POST', headers,
      body: JSON.stringify({ model, queue_id, delete_media_on_completion: true }),
    })
    if (!res.ok) throw new Error(`retrieve failed: ${res.status} ${await res.text()}`)
    const ct = res.headers.get('content-type') ?? ''
    if (!ct.startsWith('application/json')) {
      await fs.writeFile('track.mp3', Buffer.from(await res.arrayBuffer()))
      break
    }
    const { status } = await res.json()
    if (status !== 'PROCESSING') throw new Error(`unexpected ${status}`)
    await new Promise(r => setTimeout(r, 3000))
  }
  // delete_media_on_completion: true made /audio/complete unnecessary
}
Show full SKILL.md (425 more words)Show less

Capability probing

Each GET /models?type=music entry's model_spec exposes:

  • supports_lyrics, lyrics_required, lyrics_character_limit, supports_lyrics_optimizer
  • supports_force_instrumental, supports_loop, supports_language_code
  • supports_speed, default_speed, min_speed, max_speed
  • voices[], default_voice, supports_custom_voice_id
  • duration_options[], min_duration, max_duration, default_duration
  • min_prompt_length, prompt_character_limit
  • supported_formats, default_format (the output container — informational, not a request field)
  • voice_changer: true marks speech-to-speech models that belong on /audio/voice-changer/*
  • pricing, one of:
    • durations — { "<tier>": { usd, diem, min_seconds, max_seconds } } (e.g. elevenlabs-music, ace-step-15)
    • generation — flat per job (e.g. minimax-music-v25, lyria-3-pro, stable-audio-25)
    • per_second — per generated second (e.g. elevenlabs-sound-effects-v2, sonilo-v1-1-music, seed-audio-1-0)
    • per_thousand_characters — by prompt length (the ElevenLabs TTS models)

Errors

CodeMeaning
400Schema error (strict body), unsupported option for the model, bad duration_seconds, lyrics_optimizer + lyrics_prompt, voice-changer model on these endpoints, a provider-side validation failure reported on retrieve (refunded), or an unknown / foreign queue_id on retrieve/complete ("Request ID is invalid."). Voice errors include details.supported_voices.
401Authentication failed.
402Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD
403A PRIVATE_ONLY key calling an anonymized model, or region restriction.
404Unknown model; or on retrieve, media not found / expired / already deleted.
422Content policy violation (queue or retrieve). May include suggested_prompt. Charge refunded (except a DIEM charge from a previous epoch).
429Rate limited (40/min queue, 120/min retrieve, per user).
500Inference failure.
503Model at capacity — retry later.

See venice-errors for body shapes.

Gotchas

  • Quote before queue. Queue charges up front (credits) or checks your x402 balance against the quote. With an API key, compare the quote to data.balances from GET /api_keys/rate_limits, which works with an INFERENCE key and is already capped at the key's spend limit; the request is charged to the first currency that covers the whole quote (DIEM, then earned credits, then bundled credits, then USD; balances doesn't list earned credits). With a wallet, use /x402/balance/....
  • Sending an unsupported option (lyrics_prompt, voice, speed, language_code, loop, duration_seconds, …) is a 400, not a silent no-op. Build the body from model_spec.
  • Store queue_id and model — every later call needs both.
  • Media is ephemeral. Save the bytes on retrieve; after complete (or delete_media_on_completion) the audio is gone.
  • seed-audio-1-0 takes no duration_seconds: it reserves its 120 s output cap at queue time and settles on the actual length when done.
  • Poll every 2–5 s; use average_execution_time to pick the first delay. Faster polling doesn't speed the job up and eats the 120/min retrieve limit.

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/venice-audio-music of veniceai/skills.

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Audio Music next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Audio Music compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Audio Music this skillveniceai/skills144—~3.1kAutomated safety check: PassMIT
Audio Jinglesanqiufong/slides-from-anything1321 repos~1.1kAutomated safety check: PassApache-2.0
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
ElevenLabs Voiceover Generatordigitalsamba/claude-code-video-toolkit2.2k1 repos~2.7kAutomated safety check: NotesMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Hyperframes Mediachmonitor/chmonitor3011 repos~2.8kAutomated safety check: NotesGPL-3.0

Similar skills

  • Audio Jingle

    sanqiufong/slides-from-anything

    Audio generation skill — jingles, beds, voiceover, and sound effects.

    132 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    301 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Blockrun

    BlockRunAI/blockrun-mcp

    Pay-per-call access to AI models, real-time data, media generation and multi-chain RPC over x402 micropayments (USDC on Base or Solana), or a BlockRun account API key.

    391 GitHub stars~2.7k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

    144 GitHub stars~5.2k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Speech

    veniceai/skills

    Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

    144 GitHub stars~3.6k tokensUpdated 5 days ago
    Auto-check passed
  • Transcribe audio files to text via POST /audio/transcriptions.

    144 GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed

Questions about Venice Audio Music

What does Venice Audio Music do?

Async music, sound-effect and long-form voice generation via Venice. Venice Audio Music is an agent skill from veniceai/skills. Async music, sound-effect and long-form voice generation via Venice.

When should I use Venice Audio Music?

Venice Audio Music fits situations like: tasks that involve Text to speech and voice; tasks that involve Music and audio generation.

How do I install Venice Audio Music in Claude Code?

Run `npx skills add veniceai/skills --skill venice-audio-music -a claude-code`. Or copy the skill folder (skills/venice-audio-music in veniceai/skills) into .claude/skills/venice-audio-music in your project. Claude Code loads it when a task matches its description.

How do I install Venice Audio Music in Codex?

Run `npx skills add veniceai/skills --skill venice-audio-music -a codex`. Or copy the skill folder (skills/venice-audio-music in veniceai/skills) into .agents/skills/venice-audio-music in your project. Codex loads it when a task matches its description.

Can I use Venice Audio Music in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-audio-music -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-audio-music, .gemini/skills/venice-audio-music, .github/skills/venice-audio-music and .opencode/skills/venice-audio-music in your project.

What does Venice Audio Music need to run?

Going by SKILL.md and its folder, Venice Audio Music needs the command-line tools its instructions call (curl) and credentials named VENICE_API_KEY. Our summary lists: A credential in VENICE_API_KEY.

Does Venice Audio Music access the network?

SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Venice Audio Music safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Venice Audio Music use?

Venice Audio Music is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Audio Music use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Audio Music?

Skills that share tags, products or a category with Venice Audio Music: Audio Jingle (sanqiufong/slides-from-anything, 132 stars), Music (tadaspetra/loop, 296 stars), ElevenLabs Voiceover Generator (digitalsamba/claude-code-video-toolkit, 2.2k stars) and Sound Effects (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Audio Music?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.