9Router Speech-to-Text
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
Transcribe audio files to text via POST /audio/transcriptions.
$ npx skills add veniceai/skills --skill venice-audio-transcription -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install veniceai/skills venice-audio-transcription --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-audio-transcription .claude/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .claude/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcriptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add veniceai/skills --skill venice-audio-transcription -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install veniceai/skills venice-audio-transcription --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/venice-audio-transcription .agents/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .agents/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-audio-transcription -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install veniceai/skills venice-audio-transcription --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/venice-audio-transcription .cursor/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .cursor/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/veniceai/skills.git --path skills/venice-audio-transcription--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add veniceai/skills --skill venice-audio-transcription -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install veniceai/skills venice-audio-transcription --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/venice-audio-transcription .gemini/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .gemini/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install veniceai/skills venice-audio-transcriptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add veniceai/skills --skill venice-audio-transcription -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/venice-audio-transcription .github/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .github/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-audio-transcription -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install veniceai/skills venice-audio-transcription --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/venice-audio-transcription .opencode/skills/venice-audio-transcription && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "venice-audio-transcription" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-transcription into .opencode/skills/venice-audio-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-audio-transcription", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
venice-audio-transcriptionTranscribe audio files to text via POST /audio/transcriptions.
Venice Audio Transcription is an agent skill from veniceai/skills. Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, and per-audio-second pricing. OpenAI-compatible multipart.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Transcription. It works with OpenAI. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.
Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.venice.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VENICE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Venice Audio Transcription loads about 1.7k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 630 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 630 words, ~1,735 tokens.
.claude/skills/venice-audio-transcription/SKILL.md (or your agent's skills folder)./audio/transcriptions)POST /api/v1/audio/transcriptions takes an audio file and returns text. It's OpenAI-compatible with multipart/form-data — the OpenAI SDK's audio.transcriptions.create() works unchanged.
| Method | Path | Auth | Notes |
|---|---|---|---|
POST | /api/v1/audio/transcriptions | Bearer key or x402 (SIWX) | multipart/form-data, file field file, max 25 MB. Billed per second of audio. |
For video, there is no transcription endpoint any more — POST /video/transcriptions is retired and returns 410. Extract the audio track and send it here, or ask a video-capable chat model via venice-chat.
curl https://api.venice.ai/api/v1/audio/transcriptions \
-H "Authorization: Bearer $VENICE_API_KEY" \
-F "file=@./meeting.m4a" \
-F "model=nvidia/parakeet-tdt-0.6b-v3" \
-F "response_format=json" \
-F "timestamps=false"{ "text": "Alright everyone, let's kick off the meeting...", "duration": 184.2 }With timestamps=true, the JSON also carries a timestamps object (see below).
multipart/form-data)Only the fields below are read; anything else in the form is ignored.
| Field | Type | Default | Notes |
|---|---|---|---|
file | binary | — | Required. Real file part (no base64). Accepted: wav/wave, flac, m4a, aac, mp4, mp3, ogg/oga, webm. Checked by extension/MIME and then by binary signature. Max 25 MB. |
model | string | — | Send it. The OpenAPI schema lists nvidia/parakeet-tdt-0.6b-v3 as default, but that default is never applied: omitting model returns 400 "model is required". |
response_format | json / text | json | Only these two. text returns a text/plain body with just the transcript. |
timestamps | bool (true/false as form string) | false | Adds timestamps to the JSON response. |
language | string | — | ISO 639-1 hint (en, ja, …). Forwarded by Whisper, Wizper, Scribe and xAI STT; ignored by Parakeet (auto-detects). |
{
"text": "…",
"duration": 184.2,
"timestamps": {
"word": [{ "word": "Alright", "start": 0.12, "end": 0.48 }],
"segment": [{ "text": "Alright everyone…", "start": 0.12, "end": 4.9 }],
"char": [{ "char": "A", "start": 0.12, "end": 0.15 }]
}
}duration (seconds) and timestamps are optional. Which timestamp arrays appear depends on the model:
| Model | Timestamp granularity |
|---|---|
openai/whisper-large-v3 | segment + word |
fal-ai/wizper | segment |
elevenlabs/scribe-v2 | word |
stt-xai-v1 | word |
nvidia/parakeet-tdt-0.6b-v3 | may include segment, word and/or char |
All five are in the live GET /models?type=asr list. Price is model_spec.pricing.per_audio_second.usd.
| Model ID | Privacy | Notes |
|---|---|---|
nvidia/parakeet-tdt-0.6b-v3 | private | Venice-hosted, fast. Ignores language. |
openai/whisper-large-v3 | private | Multilingual; language hint; segment + word timestamps. |
fal-ai/wizper | private | Whisper v3 variant; language hint; segment timestamps. |
elevenlabs/scribe-v2 | anonymized | language hint; word timestamps. |
stt-xai-v1 | anonymized | language hint; word timestamps. |
A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized ones (PRIVATE_TEXT keys are not restricted here). Failed transcriptions are not charged.
import OpenAI from 'openai'
import fs from 'node:fs'
const client = new OpenAI({
apiKey: process.env.VENICE_API_KEY,
baseURL: 'https://api.venice.ai/api/v1',
})
const out = await client.audio.transcriptions.create({
file: fs.createReadStream('meeting.m4a'),
model: 'openai/whisper-large-v3',
response_format: 'json',
language: 'en',
// @ts-expect-error — Venice-specific extra, passes through multipart
timestamps: true,
})
console.log(out.text)There's no server-side chunking, and uploads are capped at 25 MB. Split long recordings client-side (on silence, or fixed segments), transcribe each chunk, then concatenate with offset timestamps.
ffmpeg -i long.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3| Code | Meaning |
|---|---|
400 | Missing model, bad params (e.g. response_format not json/text), no file part (including a JSON body instead of multipart → "No audio file provided"), unsupported extension/MIME, or unrecognized binary signature. |
401 | Authentication failed. |
402 | Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD |
403 | A PRIVATE_ONLY key calling an anonymized model, or region restriction. |
404 | Unknown model. |
413 | File larger than 25 MB ({"code":"PAYLOAD_TOO_LARGE","error":"File exceeds the maximum allowed size of 25 MB."}). |
422 | Upstream provider couldn't process the audio (zero-length, silent, corrupt, unsupported format or language, provider-side refusal). No suggested_prompt. |
429 | Rate limited. |
500 | Inference failure. |
502 | Temporary upstream ASR failure — {"error":"Audio transcription failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff. |
503 | Model temporarily offline — retry with jitter. |
See venice-errors for body shapes and retry strategy.
model — the documented default never applies.file must be a real multipart file part. JSON + base64 is not supported.verbose_json, srt or vtt. For subtitles, use response_format=json + timestamps=true and render the timings yourself. text drops timestamps entirely.timestamps.word vs timestamps.segment.429, back off; throttle big batches rather than firing everything in parallel.© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/venice-audio-transcription of veniceai/skills.
Open the folder on GitHubat commit 5eaeac5
Venice Audio Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Venice Audio Transcription this skillveniceai/skills | 144 | — | ~1.7k | Automated safety check: Pass | MIT | |
| 9Router Speech-to-Textdecolua/9router | 31k | — | ~914 | Automated safety check: Pass | MIT | |
| TranscribeJetBrains/skills | 366 | 4 repos | ~776 | Automated safety check: Pass | Apache-2.0 | |
| Openai Whisper APIopenclaw/openclaw | 392k | 1 repos | ~518 | Automated safety check: Pass | MIT | |
| Local AI Useamd/skills | 408 | — | ~5k | Automated safety check: Notes | MIT | |
| Local AI App Integrationamd/skills | 408 | — | ~6k | Automated safety check: Pass | MIT |
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
JetBrains/skills
Transcribe audio files to text with optional diarization and known-speaker hints.
openclaw/openclaw
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
amd/skills
Integrates local AI capabilities into applications using Embeddable Lemonade.
microsoft/GitHub-Copilot-for-Azure
A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.
veniceai/skills
Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.
veniceai/skills
Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.
veniceai/skills
Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.
veniceai/skills
High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.
veniceai/skills
Async music, sound-effect and long-form voice generation via Venice.
veniceai/skills
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.
Works with
Categories
Transcribe audio files to text via POST /audio/transcriptions. Venice Audio Transcription is an agent skill from veniceai/skills. Transcribe audio files to text via POST /audio/transcriptions.
Venice Audio Transcription fits situations like: tasks that involve Transcription.
Run `npx skills add veniceai/skills --skill venice-audio-transcription -a claude-code`. Or copy the skill folder (skills/venice-audio-transcription in veniceai/skills) into .claude/skills/venice-audio-transcription in your project. Claude Code loads it when a task matches its description.
Run `npx skills add veniceai/skills --skill venice-audio-transcription -a codex`. Or copy the skill folder (skills/venice-audio-transcription in veniceai/skills) into .agents/skills/venice-audio-transcription in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-audio-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-audio-transcription, .gemini/skills/venice-audio-transcription, .github/skills/venice-audio-transcription and .opencode/skills/venice-audio-transcription in your project.
Going by SKILL.md and its folder, Venice Audio Transcription needs the command-line tools its instructions call (curl and ffmpeg) and credentials named VENICE_API_KEY. Our summary lists: A credential in VENICE_API_KEY.
SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Venice Audio Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Venice Audio Transcription: 9Router Speech-to-Text (decolua/9router, 31k stars), Transcribe (JetBrains/skills, 366 stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Local AI Use (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.
Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.