Agent skill

9Router Speech-to-Text

by decolua in decolua/9router

Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

MITAuto-check passedMedia & Creative

Install 9Router Speech-to-Text

skills CLI
$ npx skills add decolua/9router --skill 9router-stt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install decolua/9router 9router-stt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/decolua/9router.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/9router-stt .claude/skills/9router-stt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
9router-stt
GitHub stars
31k
Token cost
~914 tokens
SKILL.md length
224 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

  • Transcribing a recording or voice memo to text
  • SKILL.md covers Discover, Endpoint, Examples and Response shape, plus 1 more section
  • Calls curl and jq; needs NINEROUTER_KEY
  • Generating SRT or VTT subtitles from an audio file

What it does

Requests go to `POST /v1/audio/transcriptions` on a 9Router server, which accepts OpenAI Whisper style multipart form data. The agent first lists speech models at `/v1/models/stt` and checks a model's supported parameters at `/v1/models/info`, then sends the model ID and an audio file (mp3, wav, m4a, webm, ogg or flac) with optional language, prompt, response format and temperature.

Output is JSON text by default; `verbose_json` adds language, duration and timestamped segments, and `srt` or `vtt` return subtitle files. A provider table notes quirks: Groq follows the OpenAI shape, Gemini audio is converted server-side, Deepgram uses token auth, AssemblyAI uploads and polling are handled by the server, and NVIDIA and Hugging Face models are available too. Curl and Node examples are included.

When your agent uses it

  • Transcribing a recording or voice memo to text
  • Generating SRT or VTT subtitles from an audio file
  • Choosing between speech-to-text providers behind one endpoint

Example prompts

  • “Transcribe ./recordings/standup.mp3 with a Groq Whisper model and print the text.”
  • “Make an SRT subtitle file from interview.wav, language English.”
  • “List the speech-to-text models available on my 9Router and pick the Deepgram one.”

Requirements

  • A 9Router server, with `NINEROUTER_URL` set (and `NINEROUTER_KEY` when auth is on)
  • `curl`

What it can do on your machine

Read from SKILL.md and the folder at commit ce4460e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NINEROUTER_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

9Router Speech-to-Text loads about 914 tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 224 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~914

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from decolua/9router at commit ce4460e, republished under its MIT licence (© decolua). 224 words, ~914 tokens.

Download SKILL.mdSave it as .claude/skills/9router-stt/SKILL.md (or your agent's skills folder).
name
9router-stt
description
Speech-to-text via 9Router /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI / NVIDIA / HuggingFace models. Use when the user wants to transcribe audio, convert speech to text, or get subtitles from audio files.

9Router — Speech-to-Text

Requires NINEROUTER_URL (and NINEROUTER_KEY if auth enabled). See https://raw.githubusercontent.com/decolua/9router/refs/heads/master/skills/9router/SKILL.md for setup.

Discover

bash
curl $NINEROUTER_URL/v1/models/stt | jq '.data[].id'
# Per-model params (language, response_format, prompt, temperature support)
curl "$NINEROUTER_URL/v1/models/info?id=openai/whisper-1"

model = STT model ID (e.g. openai/whisper-1, groq/whisper-large-v3, deepgram/nova-3, gemini/gemini-2.5-flash).

Endpoint

POST $NINEROUTER_URL/v1/audio/transcriptions (OpenAI Whisper compatible, multipart/form-data)

FieldRequiredNotes
modelyesfrom /v1/models/stt
fileyesaudio file (mp3, wav, m4a, webm, ogg, flac)
languagenoISO-639-1 (e.g. en, vi)
promptnohint text to guide transcription
response_formatnojson (default) / text / verbose_json / srt / vtt
temperatureno0–1

Examples

bash
curl -X POST "$NINEROUTER_URL/v1/audio/transcriptions" \
  -H "Authorization: Bearer $NINEROUTER_KEY" \
  -F "model=openai/whisper-1" \
  -F "file=@audio.mp3" \
  -F "language=vi"

JS (Node):

js
import { createReadStream } from "node:fs";
const form = new FormData();
form.append("model", "groq/whisper-large-v3-turbo");
form.append("file", new Blob([await (await import("node:fs/promises")).readFile("audio.mp3")]), "audio.mp3");
const r = await fetch(`${process.env.NINEROUTER_URL}/v1/audio/transcriptions`, {
  method: "POST",
  headers: { "Authorization": `Bearer ${process.env.NINEROUTER_KEY}` },
  body: form,
});
const { text } = await r.json();
console.log(text);

Response shape

Default (response_format=json):

json
{ "text": "Xin chào, đây là bản ghi âm." }

verbose_json adds language, duration, segments[] with timestamps. srt / vtt return subtitle text.

Provider quirks

Providermodel formatNotes
openaiwhisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribeNative OpenAI shape
groqwhisper-large-v3, whisper-large-v3-turbo, distil-whisper-large-v3-enFastest; OpenAI shape
geminigemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-liteServer converts to generateContent with audio inline
deepgramnova-3, nova-2, whisper-largeToken auth; server adapts response
assemblyaiuniversal-3-pro, universal-2Async upload+poll handled server-side
nvidianvidia/parakeet-ctc-1.1b-asrNIM endpoint
huggingfaceopenai/whisper-large-v3, openai/whisper-smallHF Inference API
elevenlabsscribe_v1, scribe_v2Whisper-compatible shape; xi-api-key auth, not Bearer. Extra params: timestamps_granularity (word/character/none), tag_audio_events, and speaker labelling via either diarize=true or num_speakers (1–32) — sending both makes the request invalid, so diarize wins and num_speakers is dropped. Blank language is omitted upstream for auto-detect. srt/vtt and verbose_json segments are served from Scribe's own additional_formats render, so segments[] is omitted when the upstream render is unavailable rather than synthesized. Does not accept temperature or prompt — those are not forwarded.

© decolua, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/9router-stt of decolua/9router.

Open the folder on GitHubat commit ce4460e

Compare with similar skills

9Router Speech-to-Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

9Router Speech-to-Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
9Router Speech-to-Text this skilldecolua/9router31k—~914Automated safety check: PassMIT
Keirouter Sttmydisha/keirouter147—~680Automated safety check: PassMIT
Transcribe Anythingswyxio/skills176—~8.5kAutomated safety check: PassMIT
Openai Whisper APIopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Openai Whispercoco-research/coco513—~964Automated safety check: PassCustom licence

Similar skills

  • Keirouter Stt

    mydisha/keirouter

    Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.

    147 GitHub stars~680 tokensUpdated 29 days ago
    Media & CreativeAuto-check passed
  • Transcribe Anything

    swyxio/skills

    Transcribes audio and video files to text using pluggable ASR backends.

    176 GitHub stars~8.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper

    coco-research/coco

    Speech-to-text transcription via OpenAI Whisper. An agent skill from coco-research/coco.

    513 GitHub stars~964 tokensUpdated today
    Media & CreativeAuto-check passed
  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed

More from decolua/9router

All 9 skills in this repo
  • Sets up access to the 9Router AI gateway, an OpenAI-compatible REST endpoint for chat, images, speech, embeddings, web search and web fetch, and indexes its capability skills.

    31k GitHub stars~744 tokensUpdated 2 days ago
    Auto-check passed
  • Sends chat and code-generation requests through a 9Router gateway using OpenAI or Anthropic message formats, with streaming and auto-fallback combos.

    31k GitHub stars~635 tokensUpdated 2 days ago
    Auto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 2 days ago
    Auto-check passed
  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    31k GitHub stars~830 tokensUpdated 2 days ago
    Auto-check passed
  • 9Router Text to Speech

    decolua/9router

    Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers.

    31k GitHub stars~765 tokensUpdated 2 days ago
    Auto-check passed
  • Submits text-to-video or image-to-video jobs to xAI Grok Imagine through 9Router, then polls the job and downloads the finished MP4.

    31k GitHub stars~992 tokensUpdated 2 days ago
    Auto-check passed

Questions about 9Router Speech-to-Text

What does 9Router Speech-to-Text do?

Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others. Requests go to `POST /v1/audio/transcriptions` on a 9Router server, which accepts OpenAI Whisper style multipart form data. The agent first lists speech models at `/v1/models/stt` and checks a model's supported parameters at `/v1/models/info`, then sends the model ID and an audio file (mp3, wav, m4a, webm, ogg or flac) with optional language, prompt, response format and temperature.

When should I use 9Router Speech-to-Text?

9Router Speech-to-Text fits situations like: transcribing a recording or voice memo to text; generating SRT or VTT subtitles from an audio file; choosing between speech-to-text providers behind one endpoint.

How do I install 9Router Speech-to-Text in Claude Code?

Run `npx skills add decolua/9router --skill 9router-stt -a claude-code`. Or copy the skill folder (skills/9router-stt in decolua/9router) into .claude/skills/9router-stt in your project. Claude Code loads it when a task matches its description.

How do I install 9Router Speech-to-Text in Codex?

Run `npx skills add decolua/9router --skill 9router-stt -a codex`. Or copy the skill folder (skills/9router-stt in decolua/9router) into .agents/skills/9router-stt in your project. Codex loads it when a task matches its description.

Can I use 9Router Speech-to-Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add decolua/9router --skill 9router-stt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/9router-stt, .gemini/skills/9router-stt, .github/skills/9router-stt and .opencode/skills/9router-stt in your project.

What does 9Router Speech-to-Text need to run?

Going by SKILL.md and its folder, 9Router Speech-to-Text needs the command-line tools its instructions call (curl and jq) and credentials named NINEROUTER_KEY. Our summary lists: A 9Router server, with `NINEROUTER_URL` set (and `NINEROUTER_KEY` when auth is on); `curl`.

Does 9Router Speech-to-Text access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is 9Router Speech-to-Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does 9Router Speech-to-Text use?

9Router Speech-to-Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 9Router Speech-to-Text use?

About 914 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to 9Router Speech-to-Text?

Skills that share tags, products or a category with 9Router Speech-to-Text: Keirouter Stt (mydisha/keirouter, 147 stars), Transcribe Anything (swyxio/skills, 176 stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Watch Video (coreyhaines31/makerskills, 851 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 9Router Speech-to-Text?

decolua (a GitHub user) maintains it in decolua/9router, which has 30,536 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: decolua/9router on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.