Agent skill

Venice Audio Transcription

by veniceai in veniceai/skills

Transcribe audio files to text via POST /audio/transcriptions.

MITAuto-check passedMedia & Creative

Install Venice Audio Transcription

skills CLI
$ npx skills add veniceai/skills --skill venice-audio-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-audio-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-audio-transcription .claude/skills/venice-audio-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-audio-transcription
GitHub stars
144
Token cost
~1.7k tokens
SKILL.md length
630 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Transcribe audio files to text via POST /audio/transcriptions.

  • Tasks that involve Transcription
  • SKILL.md covers Use when, Minimal request, Request (multipart/form-data) and Response, plus 5 more sections
  • Calls curl and ffmpeg; reaches api.venice.ai; needs VENICE_API_KEY

What it does

Venice Audio Transcription is an agent skill from veniceai/skills. Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, and per-audio-second pricing. OpenAI-compatible multipart.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. It works with OpenAI. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription

Example prompts

  • “/venice-audio-transcription”

Requirements

  • A credential in VENICE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Audio Transcription loads about 1.7k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 630 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 630 words, ~1,735 tokens.

Download SKILL.mdSave it as .claude/skills/venice-audio-transcription/SKILL.md (or your agent's skills folder).
name
venice-audio-transcription
description
Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, and per-audio-second pricing. OpenAI-compatible multipart.

Venice Transcription (/audio/transcriptions)

POST /api/v1/audio/transcriptions takes an audio file and returns text. It's OpenAI-compatible with multipart/form-data — the OpenAI SDK's audio.transcriptions.create() works unchanged.

MethodPathAuthNotes
POST/api/v1/audio/transcriptionsBearer key or x402 (SIWX)multipart/form-data, file field file, max 25 MB. Billed per second of audio.

Use when

  • You need STT (speech-to-text) for voice notes, meetings, podcasts, short audio.
  • You need word/segment timestamps for subtitles or chapters.
  • You want to pick between Venice-hosted Parakeet, Whisper-family models, ElevenLabs Scribe, or xAI STT.

For video, there is no transcription endpoint any more — POST /video/transcriptions is retired and returns 410. Extract the audio track and send it here, or ask a video-capable chat model via venice-chat.

Minimal request

bash
curl https://api.venice.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -F "file=@./meeting.m4a" \
  -F "model=nvidia/parakeet-tdt-0.6b-v3" \
  -F "response_format=json" \
  -F "timestamps=false"
json
{ "text": "Alright everyone, let's kick off the meeting...", "duration": 184.2 }

With timestamps=true, the JSON also carries a timestamps object (see below).

Request (multipart/form-data)

Only the fields below are read; anything else in the form is ignored.

FieldTypeDefaultNotes
filebinary—Required. Real file part (no base64). Accepted: wav/wave, flac, m4a, aac, mp4, mp3, ogg/oga, webm. Checked by extension/MIME and then by binary signature. Max 25 MB.
modelstring—Send it. The OpenAPI schema lists nvidia/parakeet-tdt-0.6b-v3 as default, but that default is never applied: omitting model returns 400 "model is required".
response_formatjson / textjsonOnly these two. text returns a text/plain body with just the transcript.
timestampsbool (true/false as form string)falseAdds timestamps to the JSON response.
languagestring—ISO 639-1 hint (en, ja, …). Forwarded by Whisper, Wizper, Scribe and xAI STT; ignored by Parakeet (auto-detects).

Response

json
{
  "text": "…",
  "duration": 184.2,
  "timestamps": {
    "word":    [{ "word": "Alright", "start": 0.12, "end": 0.48 }],
    "segment": [{ "text": "Alright everyone…", "start": 0.12, "end": 4.9 }],
    "char":    [{ "char": "A", "start": 0.12, "end": 0.15 }]
  }
}

duration (seconds) and timestamps are optional. Which timestamp arrays appear depends on the model:

ModelTimestamp granularity
openai/whisper-large-v3segment + word
fal-ai/wizpersegment
elevenlabs/scribe-v2word
stt-xai-v1word
nvidia/parakeet-tdt-0.6b-v3may include segment, word and/or char

Models

All five are in the live GET /models?type=asr list. Price is model_spec.pricing.per_audio_second.usd.

Model IDPrivacyNotes
nvidia/parakeet-tdt-0.6b-v3privateVenice-hosted, fast. Ignores language.
openai/whisper-large-v3privateMultilingual; language hint; segment + word timestamps.
fal-ai/wizperprivateWhisper v3 variant; language hint; segment timestamps.
elevenlabs/scribe-v2anonymizedlanguage hint; word timestamps.
stt-xai-v1anonymizedlanguage hint; word timestamps.

A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized ones (PRIVATE_TEXT keys are not restricted here). Failed transcriptions are not charged.

OpenAI SDK

ts
import OpenAI from 'openai'
import fs from 'node:fs'

const client = new OpenAI({
  apiKey: process.env.VENICE_API_KEY,
  baseURL: 'https://api.venice.ai/api/v1',
})

const out = await client.audio.transcriptions.create({
  file: fs.createReadStream('meeting.m4a'),
  model: 'openai/whisper-large-v3',
  response_format: 'json',
  language: 'en',
  // @ts-expect-error — Venice-specific extra, passes through multipart
  timestamps: true,
})

console.log(out.text)

Long files

There's no server-side chunking, and uploads are capped at 25 MB. Split long recordings client-side (on silence, or fixed segments), transcribe each chunk, then concatenate with offset timestamps.

bash
ffmpeg -i long.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
Show full SKILL.md (246 more words)Show less

Errors

CodeMeaning
400Missing model, bad params (e.g. response_format not json/text), no file part (including a JSON body instead of multipart → "No audio file provided"), unsupported extension/MIME, or unrecognized binary signature.
401Authentication failed.
402Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD
403A PRIVATE_ONLY key calling an anonymized model, or region restriction.
404Unknown model.
413File larger than 25 MB ({"code":"PAYLOAD_TOO_LARGE","error":"File exceeds the maximum allowed size of 25 MB."}).
422Upstream provider couldn't process the audio (zero-length, silent, corrupt, unsupported format or language, provider-side refusal). No suggested_prompt.
429Rate limited.
500Inference failure.
502Temporary upstream ASR failure — {"error":"Audio transcription failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff.
503Model temporarily offline — retry with jitter.

See venice-errors for body shapes and retry strategy.

Gotchas

  • Always send model — the documented default never applies.
  • file must be a real multipart file part. JSON + base64 is not supported.
  • There is no verbose_json, srt or vtt. For subtitles, use response_format=json + timestamps=true and render the timings yourself. text drops timestamps entirely.
  • Check which granularity your model returns before building on timestamps.word vs timestamps.segment.
  • A file with a valid extension but a non-audio binary signature is rejected; re-encode to a standard profile (e.g. MP3 44.1 kHz or 16 kHz).
  • On 429, back off; throttle big batches rather than firing everything in parallel.

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/venice-audio-transcription of veniceai/skills.

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Audio Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Audio Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Audio Transcription this skillveniceai/skills144—~1.7kAutomated safety check: PassMIT
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
TranscribeJetBrains/skills3664 repos~776Automated safety check: PassApache-2.0
Openai Whisper APIopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
Local AI Useamd/skills408—~5kAutomated safety check: NotesMIT
Local AI App Integrationamd/skills408—~6kAutomated safety check: PassMIT

Similar skills

  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    408 GitHub stars~5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    408 GitHub stars~6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

    144 GitHub stars~5.2k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Speech

    veniceai/skills

    Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

    144 GitHub stars~3.6k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about Venice Audio Transcription

What does Venice Audio Transcription do?

Transcribe audio files to text via POST /audio/transcriptions. Venice Audio Transcription is an agent skill from veniceai/skills. Transcribe audio files to text via POST /audio/transcriptions.

When should I use Venice Audio Transcription?

Venice Audio Transcription fits situations like: tasks that involve Transcription.

How do I install Venice Audio Transcription in Claude Code?

Run `npx skills add veniceai/skills --skill venice-audio-transcription -a claude-code`. Or copy the skill folder (skills/venice-audio-transcription in veniceai/skills) into .claude/skills/venice-audio-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Venice Audio Transcription in Codex?

Run `npx skills add veniceai/skills --skill venice-audio-transcription -a codex`. Or copy the skill folder (skills/venice-audio-transcription in veniceai/skills) into .agents/skills/venice-audio-transcription in your project. Codex loads it when a task matches its description.

Can I use Venice Audio Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-audio-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-audio-transcription, .gemini/skills/venice-audio-transcription, .github/skills/venice-audio-transcription and .opencode/skills/venice-audio-transcription in your project.

What does Venice Audio Transcription need to run?

Going by SKILL.md and its folder, Venice Audio Transcription needs the command-line tools its instructions call (curl and ffmpeg) and credentials named VENICE_API_KEY. Our summary lists: A credential in VENICE_API_KEY.

Does Venice Audio Transcription access the network?

SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Venice Audio Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Venice Audio Transcription use?

Venice Audio Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Audio Transcription use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Audio Transcription?

Skills that share tags, products or a category with Venice Audio Transcription: 9Router Speech-to-Text (decolua/9router, 31k stars), Transcribe (JetBrains/skills, 366 stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Local AI Use (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Audio Transcription?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.