9Router Speech-to-Text
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
Transcribes audio and video files to text using pluggable ASR backends.
$ npx skills add swyxio/skills --skill transcribe-anything -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install swyxio/skills transcribe-anything --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/transcribe-anything .claude/skills/transcribe-anything && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .claude/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/swyxio/skills/tree/main/transcribe-anythingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add swyxio/skills --skill transcribe-anything -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install swyxio/skills transcribe-anything --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/transcribe-anything .agents/skills/transcribe-anything && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .agents/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill transcribe-anything -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install swyxio/skills transcribe-anything --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/transcribe-anything .cursor/skills/transcribe-anything && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .cursor/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/swyxio/skills.git --path transcribe-anything--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add swyxio/skills --skill transcribe-anything -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install swyxio/skills transcribe-anything --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/transcribe-anything .gemini/skills/transcribe-anything && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .gemini/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install swyxio/skills transcribe-anythingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add swyxio/skills --skill transcribe-anything -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/transcribe-anything .github/skills/transcribe-anything && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .github/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill transcribe-anything -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install swyxio/skills transcribe-anything --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/transcribe-anything .opencode/skills/transcribe-anything && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "transcribe-anything" agent skill from https://github.com/swyxio/skills/tree/main/transcribe-anything into .opencode/skills/transcribe-anything/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcribe-anything", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
transcribe-anythingTranscribes audio and video files to text using pluggable ASR backends.
Transcribe Anything is an agent skill from swyxio/skills. Transcribes audio and video files to text using pluggable ASR backends. Default backend is local whisper/MLX Whisper for ASR. Supports pyannote.audio direct diarization, whisperX, insanely-fast-whisper, faster-whisper, whisper.cpp, OpenAI Whisper API, Groq Whisper API, Deepgram, AssemblyAI, Gemini, and Hugging Face models. Handles very long files (1-8+ hours) by preprocessing with ffmpeg: extracts audio from video, converts to optimal ASR format, detects and skips silence, and chunks for API size limits. Supports…
Its SKILL.md is about 8.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `README.md`). Compatibility notes: Requires ffmpeg and at least one ASR backend; MLX Whisper or openai-whisper is the default on macOS. For diarization, prefer managed APIs when configured…
It sits in AI & LLM Engineering, covering Transcription and Speech recognition and synthesis. It works with Whisper, Hugging Face, Deepgram and FFmpeg. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ffmpegwhisperuvpip3curlffprobebrewpython3pippythongityt-dlpFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comapi.openai.comapi.deepgram.comapi.groq.comAlso links to:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENOPENAI_API_KEYGROQ_API_KEYDEEPGRAM_API_KEYASSEMBLYAI_API_KEYGEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires ffmpeg and at least one ASR backend; MLX Whisper or openai-whisper is the default on macOS. For diarization, prefer managed APIs when configured, pyannote.audio on CUDA with Hugging Face access, or NeMo on CUDA without secrets. Never run long diarization on local CPU without explicit opt-in. Cloud backends require their API keys in environment variables.
From compatibility in the SKILL.md frontmatter.
Transcribe Anything loads about 8.5k tokens when it runs. Until then it costs about 210 tokens; SKILL.md has 2,199 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 2,199 words, ~8,516 tokens.
.claude/skills/transcribe-anything/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Transcribes audio and video files to text. Pluggable backends, silence skipping for long files, optional speaker diarization, and multiple output formats.
# ffmpeg — audio extraction, preprocessing, silence detection
brew install ffmpeg
# yt-dlp — downloading video/audio from URLs (optional but recommended)
brew install yt-dlp
# Default ASR backend — OpenAI's whisper CLI
pip3 install --break-system-packages openai-whisper
# Apple Silicon accelerated ASR, preferred when available
uv tool install mlx-whisper# curl_cffi — prevents OAuth errors when downloading private videos
pip3 install --break-system-packages curl_cffi
# faster-whisper — 4x faster than whisper, built-in VAD silence skipping, lower memory
# Best local backend for long files (1hr+)
pip3 install --break-system-packages faster-whisper
# pyannote.audio — preferred local diarization when HF_TOKEN is available
uv venv --python 3.11 .venv-pyannote
uv pip install --python .venv-pyannote/bin/python "pyannote.audio" soundfile
# whisperX — adds precise word-level timestamps; diarization requires HF_TOKEN
# Bundles faster-whisper + pyannote alignment/diarization
pip3 install --break-system-packages whisperxpyannote/whisperX diarization setup (one-time):
export HF_TOKEN=hf_... in your shell profileWithout HF token access, whisperX still works for transcription and word alignment — just no speaker labels. pyannote direct diarization will not work without accepted model terms.
When a transcript already exists, prefer pyannote direct diarization on GPU/cloud GPU over WhisperX diarization. It produces a clean RTTM/exclusive speaker timeline that can be aligned to any ASR output. WhisperX is best when you also need forced word alignment from scratch.
No-secret diarization setup (heavier, but real diarization):
Use NVIDIA NeMo rather than ad hoc speaker embedding clustering when the user asks for diarization and no HF/cloud token is available. NeMo's clustering/MSDD diarizers use VAD + speaker embeddings + clustering; Sortformer is the newer end-to-end diarizer path. Prefer running NeMo on CUDA/cloud GPU. Do not start long NeMo diarization on local CPU unless the user explicitly opts in.
# Prefer Python 3.11 for NeMo audio dependencies.
uv venv --python 3.11 .venv-nemo
source .venv-nemo/bin/activate
uv pip install "nemo_toolkit[asr]"
git clone --depth 1 https://github.com/NVIDIA/NeMo.git /tmp/NeMoAvoid "lightweight diarization" based only on clustering Whisper segments with Resemblyzer/librosa embeddings unless the user explicitly accepts approximate labels. It often collapses to one speaker on long interviews because long ASR segments contain mixed speakers and most embeddings are dominated by the primary talker.
# insanely-fast-whisper — batched GPU inference, 10-20x faster on NVIDIA GPUs
pip3 install --break-system-packages insanely-fast-whisper
# whisper.cpp — C++ native with Metal acceleration on Apple Silicon
# Best option if you want to avoid Python entirely
brew install whisper-cppSet these environment variables if you want to use cloud backends. None are required — local whisper works out of the box.
# OpenAI — best accuracy with gpt-4o-transcribe ($0.006/min)
export OPENAI_API_KEY=sk-...
# Groq — cheapest and fastest cloud option ($0.00004/min with turbo)
export GROQ_API_KEY=gsk_...
# Deepgram — best cloud diarization ($0.0043/min)
export DEEPGRAM_API_KEY=...
# AssemblyAI — cloud diarization + auto-chapters ($0.0062/min)
export ASSEMBLYAI_API_KEY=...
# Gemini — handles 9.5hr files natively, flexible prompting
export GEMINI_API_KEY=...Run this to check what's available:
echo "=== Required ==="
which ffmpeg && echo "ffmpeg: OK" || echo "ffmpeg: MISSING (brew install ffmpeg)"
which whisper && echo "whisper: OK" || echo "whisper: MISSING (pip3 install --break-system-packages openai-whisper)"
echo ""
echo "=== Local Backends ==="
which whisperx && echo "whisperx: OK" || echo "whisperx: not installed"
python3 -c "import faster_whisper" 2>/dev/null && echo "faster-whisper: OK" || echo "faster-whisper: not installed"
python3 -c "import pyannote.audio" 2>/dev/null && echo "pyannote.audio: OK" || echo "pyannote.audio: not installed in system Python"
python3 -c "import nemo.collections.asr" 2>/dev/null && echo "NeMo ASR: OK" || echo "NeMo ASR: not installed"
which mlx_whisper && echo "mlx_whisper: OK" || echo "mlx_whisper: not installed"
which insanely-fast-whisper 2>/dev/null && echo "insanely-fast-whisper: OK" || echo "insanely-fast-whisper: not installed"
which whisper-cpp 2>/dev/null && echo "whisper.cpp: OK" || echo "whisper.cpp: not installed"
echo ""
echo "=== Cloud APIs ==="
[ -n "$OPENAI_API_KEY" ] && echo "OpenAI: configured" || echo "OpenAI: not set"
[ -n "$GROQ_API_KEY" ] && echo "Groq: configured" || echo "Groq: not set"
[ -n "$DEEPGRAM_API_KEY" ] && echo "Deepgram: configured" || echo "Deepgram: not set"
[ -n "$ASSEMBLYAI_API_KEY" ] && echo "AssemblyAI: configured" || echo "AssemblyAI: not set"
[ -n "$GEMINI_API_KEY" ] && echo "Gemini: configured" || echo "Gemini: not set"
echo ""
echo "=== Optional ==="
which yt-dlp && echo "yt-dlp: OK" || echo "yt-dlp: not installed (brew install yt-dlp)"
python3 -c "import curl_cffi" 2>/dev/null && echo "curl_cffi: OK" || echo "curl_cffi: not installed"
[ -n "$HF_TOKEN" ] && echo "HF token: configured (whisperX/pyannote diarization ready if model terms accepted)" || echo "HF token: not set (use NeMo or cloud API for diarization)"Pick the backend based on the user's needs:
| Scenario | Backend | Why |
|---|---|---|
| Default / just works | mlx_whisper or whisper | Fast local ASR on Apple Silicon, good quality |
| Need speaker labels, cloud key available | deepgram or assemblyai | Native diarization and utterances; fastest path |
| Need speaker labels, HF token available and transcript exists | pyannote.audio on CUDA/cloud GPU | Best RTTM/exclusive speaker timeline |
| Need speaker labels, HF token available and word alignment needed | whisperx | Pyannote diarization + word alignment |
| Need speaker labels, no secrets | nemo on CUDA/cloud GPU | Real diarization with public NGC models |
| Very long file, local | faster-whisper | VAD silence skipping, low memory |
| Maximum speed, local GPU | insanely-fast-whisper | Batched inference, 10-20x faster |
| Apple Silicon, no Python | whisper.cpp | Metal acceleration, pure C++ |
| Cheapest cloud, fast | groq | $0.00004/min with turbo model |
| Best cloud accuracy | openai | gpt-4o-transcribe model |
| Cloud with diarization | deepgram or assemblyai | Native speaker labels |
| Flexible Q&A over audio | gemini | Can ask questions, not just transcribe |
If the user doesn't specify, use this priority:
deepgram or assemblyai.HF_TOKEN is set, use pyannote direct diarization on CUDA/cloud GPU after confirming model terms are accepted; align its exclusive RTTM to ASR segments.whisperx --diarize when you need diarization plus forced word-level alignment in one pipeline.mlx_whisper, faster-whisper, then whisper).openai/groq API for transcription-only when a key is set and speed matters.CPU policy: do not automatically run diarization on local CPU for long files. For files over 10 minutes or more than two speakers, ask before using CPU and clearly state expected runtime. Prefer cloud/GPU even if setup takes extra time.
Do not deliver diarization without a QA check that counts speaker labels and samples several speaker changes. If all or nearly all segments are one speaker on a known conversation, mark the diarization attempt failed and switch backend.
Accept any of these input types:
If the input is a URL:
yt-dlp -x --audio-format wav -o "%(title)s.%(ext)s" "{url}"Always preprocess. This step is critical for quality and speed.
# Extract audio from video (or re-encode audio) to ASR-optimal format
ffmpeg -i "{input}" \
-vn \
-ac 1 \
-ar 16000 \
-acodec pcm_s16le \
-af "highpass=f=80,lowpass=f=8000,loudnorm=I=-16:TP=-1.5:LRA=11" \
"{output_stem}_preprocessed.wav"Flags explained:
-vn — strip video-ac 1 — mono (stereo wastes processing time, no ASR benefit)-ar 16000 — 16kHz (what whisper expects internally)-acodec pcm_s16le — 16-bit WAVhighpass=f=80 — remove rumble below speech rangelowpass=f=8000 — remove hiss above speech rangeloudnorm — normalize volume (critical for variable-volume recordings)Check duration after preprocessing:
DURATION=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "{preprocessed_file}" | cut -d. -f1)
echo "Duration: ${DURATION}s ($((DURATION / 3600))h $(((DURATION % 3600) / 60))m)"For files over 30 minutes, apply silence detection and chunking. This is especially important for 1-8 hour recordings.
# Detect silent regions (informational — see what we're working with)
ffmpeg -i "{preprocessed_file}" \
-af silencedetect=noise=-30dB:d=2.0 \
-f null - 2>&1 | grep -c "silence_end"
# Shows number of silence gaps >= 2 secondsSilence threshold guide:
-30dB — clean recordings (studio, podcast)-35dB — moderate background noise-40dB — noisy environmentsLocal backends handle long files natively — no need to chunk. But use VAD to skip silence:
With faster-whisper (built-in VAD):
from faster_whisper import WhisperModel
model = WhisperModel("large-v3", device="cpu", compute_type="int8")
segments, info = model.transcribe(
"preprocessed.wav",
language="en",
word_timestamps=True,
vad_filter=True,
vad_parameters=dict(
min_silence_duration_ms=1000,
speech_pad_ms=400,
threshold=0.5,
),
condition_on_previous_text=False, # prevents hallucination cascades on long files
)With whisper CLI (no built-in VAD — preprocess silence out):
# Remove silences longer than 2s, keeping 0.3s padding
ffmpeg -i "{preprocessed_file}" \
-af "silenceremove=start_periods=1:start_threshold=-30dB:stop_periods=-1:stop_duration=2.0:stop_threshold=-30dB" \
"{output_stem}_trimmed.wav"
# Then transcribe the trimmed file
whisper "{output_stem}_trimmed.wav" --model turbo --language en \
--condition_on_previous_text False \
--word_timestamps True \
--output_format json \
--output_dir ./Important for long files: Always use --condition_on_previous_text False with whisper on files over 30 minutes. Without this, a single hallucination can cascade and corrupt hours of transcript (whisper repeats the same phrase endlessly).
Cloud APIs (OpenAI, Groq) have a 25MB limit. Compress first, then chunk if needed.
# Compress to opus (smallest format for speech) — 1 hour ≈ 14MB
ffmpeg -i "{preprocessed_file}" -ac 1 -ar 16000 -c:a libopus -b:a 32k "{output_stem}.ogg"
# Check file size
SIZE_MB=$(du -m "{output_stem}.ogg" | cut -f1)
echo "File size: ${SIZE_MB}MB"If the compressed file is under 25MB, send it directly. Otherwise, chunk on silence boundaries:
# Split into ~20-minute chunks on silence boundaries
# (under 25MB each at opus 32kbps)
ffmpeg -i "{output_stem}.ogg" \
-f segment \
-segment_time 1200 \
-c copy \
"{output_stem}_chunk_%03d.ogg"For each chunk, track the start offset for timestamp correction:
# Get duration of each chunk for timestamp reassembly
for f in {output_stem}_chunk_*.ogg; do
dur=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$f")
echo "$f: ${dur}s"
donewhisper "{input_file}" \
--model turbo \
--language en \
--output_format json \
--output_dir "{output_dir}" \
--word_timestamps True \
--condition_on_previous_text False \
--fp16 FalseModel selection for Apple Silicon (CPU — no CUDA):
turbo — best balance of speed and quality (recommended default)large-v3 — highest quality, 2-3x slower than turbomedium.en — faster, English-only, good for clear speechsmall.en — fast, acceptable quality for clean recordingsbase.en — fastest, use only for quick previewsNote: --fp16 False is required on CPU (Apple Silicon without MLX). Whisper defaults to fp16 which only works on CUDA.
whisperx "{input_file}" \
--model large-v3 \
--language en \
--diarize \
--min_speakers 2 \
--max_speakers 6 \
--hf_token "{HF_TOKEN}" \
--compute_type int8 \
--output_dir "{output_dir}" \
--output_format jsonIf no HF token is available, whisperX still works for transcription and word alignment, just without diarization:
whisperx "{input_file}" \
--model large-v3 \
--language en \
--compute_type int8 \
--output_dir "{output_dir}" \
--output_format jsonUse this when the transcript already exists or when ASR and diarization should be decoupled. It outputs RTTM speaker turns. For readable transcript assignment, prefer the exclusive_speaker_diarization output because it guarantees at most one speaker at a time.
import os
import torch
from pyannote.audio import Pipeline
audio = "{preprocessed_wav}" # mono 16 kHz WAV
num_speakers = 5 # set when known; improves clustering
token = os.environ["HF_TOKEN"]
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1",
token=token,
)
if not torch.cuda.is_available():
raise RuntimeError(
"No CUDA GPU available. Do not run long diarization on CPU unless the user explicitly opts in."
)
device = "cuda"
pipeline.to(torch.device(device))
output = pipeline(audio, num_speakers=num_speakers)
diarization = output.speaker_diarization
exclusive = output.exclusive_speaker_diarization
with open("diarization.rttm", "w") as f:
diarization.write_rttm(f)
with open("diarization-exclusive.rttm", "w") as f:
exclusive.write_rttm(f)Operational notes:
HF_TOKEN from the environment or a secret manager.num_speakers when known. For meetings/interviews, ask the user for speaker count and names before diarization.Use this when the user requests diarization and there is no HF_TOKEN, Deepgram key, or AssemblyAI key. This produces RTTM speaker turns that must be aligned back onto the ASR transcript.
AUDIO="{input_file}"
WORK="{output_dir}/nemo-diarization"
NEMO_REPO="${NEMO_REPO:-/tmp/NeMo}"
mkdir -p "$WORK/input" "$WORK/output"
ffmpeg -y -i "$AUDIO" -ac 1 -ar 16000 "$WORK/input/audio.16k.wav"
DURATION=$(ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 "$WORK/input/audio.16k.wav")
python - <<PY
import json
manifest = {
"audio_filepath": "$WORK/input/audio.16k.wav",
"offset": 0,
"duration": float("$DURATION"),
"label": "infer",
"text": "-",
"num_speakers": 2,
"rttm_filepath": None,
"uem_filepath": None,
}
open("$WORK/input/manifest.json", "w").write(json.dumps(manifest) + "\\n")
PY
test -d "$NEMO_REPO" || git clone --depth 1 https://github.com/NVIDIA/NeMo.git "$NEMO_REPO"
uv venv --python 3.11 "$WORK/.venv"
source "$WORK/.venv/bin/activate"
uv pip install "nemo_toolkit[asr]"
python "$NEMO_REPO/examples/speaker_tasks/diarization/clustering_diarizer/offline_diar_infer.py" \
diarizer.manifest_filepath="$WORK/input/manifest.json" \
diarizer.out_dir="$WORK/output" \
diarizer.speaker_embeddings.model_path=titanet_large \
diarizer.vad.model_path=vad_multilingual_marblenet \
diarizer.speaker_embeddings.parameters.save_embeddings=False \
diarizer.clustering.parameters.oracle_num_speakers=TrueIf speaker count is unknown, use num_speakers: null and oracle_num_speakers=False, but prefer a known count for interviews. For two-person interviews, setting num_speakers: 2 usually prevents single-speaker collapse.
Expected RTTM:
find "$WORK/output" -name '*.rttm' -printAlign RTTM to Whisper/faster-whisper segments by choosing the speaker with the largest time overlap for each segment. If a Whisper segment spans multiple RTTM speakers, split it only if word timestamps are available; otherwise assign by majority overlap and keep the raw RTTM for audit.
Prefer Sortformer when installed examples support it and the audio is short enough for the available hardware. Sortformer is NeMo's newer end-to-end diarizer that predicts speaker labels directly from audio. On CPU it may be slower than clustering diarization; on CUDA it is a better candidate for high-quality diarization.
Check for available scripts:
find "$NEMO_REPO/examples" -iname '*sortformer*' -o -iname '*diar*infer*.py'If Sortformer is available in the checked-out NeMo version, run the provided inference script with an output RTTM path, then use the same RTTM-to-ASR alignment step.
from faster_whisper import WhisperModel
model = WhisperModel("large-v3", device="cpu", compute_type="int8")
segments, info = model.transcribe(
"{input_file}",
language="en",
beam_size=5,
word_timestamps=True,
vad_filter=True,
vad_parameters=dict(min_silence_duration_ms=1000),
condition_on_previous_text=False,
)
for segment in segments:
print(f"[{segment.start:.2f} -> {segment.end:.2f}] {segment.text}")insanely-fast-whisper \
--file-name "{input_file}" \
--model-name openai/whisper-large-v3-turbo \
--task transcribe \
--language en \
--batch-size 24 \
--timestamp word \
--transcript-path "{output_stem}.json"# Download model if needed
whisper-cpp-download-model large-v3
# Transcribe with Metal GPU acceleration
whisper-cpp \
-m ~/.local/share/whisper-cpp/ggml-large-v3.bin \
-f "{preprocessed_wav}" \
-l en \
-t 8 \
--output-json \
--print-progressNote: whisper.cpp requires WAV input (not mp3/ogg). Always preprocess to WAV first.
curl -s https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file="@{input_file}" \
-F model="gpt-4o-transcribe" \
-F language="en" \
-F response_format="verbose_json" \
-F 'timestamp_granularities[]=word' \
-F 'timestamp_granularities[]=segment' \
> "{output_stem}_openai.json"For multiple chunks, loop and offset timestamps:
OFFSET=0
for chunk in {output_stem}_chunk_*.ogg; do
curl -s https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file="@$chunk" \
-F model="gpt-4o-transcribe" \
-F language="en" \
-F response_format="verbose_json" \
-F 'timestamp_granularities[]=segment' \
> "${chunk%.ogg}_transcript.json"
DUR=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$chunk")
OFFSET=$(echo "$OFFSET + $DUR" | bc)
doneSame OpenAI-compatible format, different base URL:
curl -s https://api.groq.com/openai/v1/audio/transcriptions \
-H "Authorization: Bearer $GROQ_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file="@{input_file}" \
-F model="whisper-large-v3-turbo" \
-F language="en" \
-F response_format="verbose_json" \
-F 'timestamp_granularities[]=word' \
-F 'timestamp_granularities[]=segment' \
> "{output_stem}_groq.json"curl -s -X POST "https://api.deepgram.com/v1/listen?model=nova-3&smart_format=true&diarize=true&language=en&utterances=true" \
-H "Authorization: Token $DEEPGRAM_API_KEY" \
-H "Content-Type: audio/wav" \
--data-binary "@{input_file}" \
> "{output_stem}_deepgram.json"import assemblyai as aai
aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"]
config = aai.TranscriptionConfig(
speaker_labels=True,
language_code="en",
auto_chapters=True,
word_boost=["custom", "vocabulary", "terms"],
)
transcript = aai.Transcriber().transcribe("{input_file}", config=config)
for utterance in transcript.utterances:
print(f"Speaker {utterance.speaker}: {utterance.text}")import google.generativeai as genai
genai.configure(api_key=os.environ["GEMINI_API_KEY"])
model = genai.GenerativeModel("gemini-2.5-flash")
audio = genai.upload_file("{input_file}")
response = model.generate_content([
audio,
"""Transcribe this audio verbatim. Format as markdown with:
- Timestamps every ~30 seconds as ### headers (e.g., ### [00:01:30])
- Speaker labels if you can distinguish voices (Speaker 1, Speaker 2, etc.)
- Paragraph breaks at natural topic shifts
Do not summarize or omit anything. Transcribe every word spoken."""
])
print(response.text)Note: Gemini handles up to ~9.5 hours natively (no chunking needed) but timestamps are approximate and output is unstructured text shaped by your prompt.
The user's default preference is markdown-formatted plain text. Convert the raw backend output to this format.
# Transcript: {filename}
**Date transcribed:** {date}
**Duration:** {duration}
**Backend:** {backend}
**Model:** {model}
***
## [00:00:00]
{text of first segment or group of segments...}
## [00:05:23]
{text continues with periodic timestamp headers...}
***
*Transcribed with {backend} ({model})*Formatting rules:
## [HH:MM:SS] timestamp headers every 2-5 minutes (not every segment — that's too noisy)**Speaker 1:** text...*** horizontal rules at major topic shifts or long pauses (>10s)# Transcript: {filename}
**Speakers:** 3 detected
**Duration:** 1h 23m
***
## [00:00:00]
**Speaker 1:** Welcome everyone to today's session. We're going to be talking about...
**Speaker 2:** Thanks for having me. I'm excited to share...
## [00:05:12]
**Speaker 1:** Let's dive into the first topic...Before delivering speaker-labeled output, run a simple validation pass:
python - "{diarized_json}" <<'PY'
import json, collections, sys
path = sys.argv[1]
data = json.load(open(path))
segments = data.get("segments", data if isinstance(data, list) else [])
counts = collections.Counter(s.get("speaker", "UNKNOWN") for s in segments)
total = sum(counts.values())
print(counts)
if len(counts) < 2:
raise SystemExit("FAILED: diarization found fewer than 2 speakers")
speaker, n = counts.most_common(1)[0]
if total and n / total > 0.95:
raise SystemExit(f"FAILED: {speaker} owns {n/total:.1%} of segments; likely speaker collapse")
PYAlso inspect at least the first 10 minutes and two later sections manually. A good diarization pass should show plausible alternation at questions, short confirmations, and interruptions. If it fails:
num_speakers and oracle_num_speakers=True.If the user requests a different format:
SRT subtitles:
whisper "{input}" --model turbo --output_format srt --output_dir ./VTT subtitles:
whisper "{input}" --model turbo --output_format vtt --output_dir ./JSON with word timestamps:
whisper "{input}" --model turbo --output_format json --word_timestamps True --output_dir ./Plain text (no timestamps):
whisper "{input}" --model turbo --output_format txt --output_dir ./TSV (tab-separated, for spreadsheets):
whisper "{input}" --model turbo --output_format tsv --output_dir ./Whisper supports an initial_prompt that biases the model toward specific terminology:
whisper "{input}" --model turbo --language en \
--initial_prompt "This conversation discusses Kubernetes, GitLab CI/CD, Terraform, and Infrastructure as Code. Names mentioned: Sid Sijbrandij, David Thompson."For cloud APIs:
# OpenAI - use the prompt parameter
curl ... -F prompt="Technical terms: LLM, RAG, vector database, embeddings. Names: Sid Sijbrandij."
# AssemblyAI - use word_boost
config = aai.TranscriptionConfig(
word_boost=["Kubernetes", "GitLab", "Sijbrandij", "Terraform"],
boost_param="high",
)
# Deepgram - use keywords
curl ... "https://api.deepgram.com/v1/listen?keywords=Kubernetes:2&keywords=GitLab:2"When to use custom vocabulary:
Music segments cause hallucinations. Trim them:
# Skip first 30s (intro music) and last 30s (outro)
TOTAL=$(ffprobe -v error -show_entries format=duration -of csv=p=0 input.mp3 | cut -d. -f1)
END=$((TOTAL - 30))
ffmpeg -i input.mp3 -ss 30 -to $END -ac 1 -ar 16000 -acodec pcm_s16le trimmed.wavIf each speaker has their own audio track:
# Extract each track
ffmpeg -i recording.mkv -map 0:a:0 -ac 1 -ar 16000 speaker1.wav
ffmpeg -i recording.mkv -map 0:a:1 -ac 1 -ar 16000 speaker2.wav
# Transcribe each separately (no diarization needed)
whisper speaker1.wav --model turbo --output_format json
whisper speaker2.wav --model turbo --output_format json
# Then interleave by timestamps in the markdown output# Split channels
ffmpeg -i stereo.wav -af "pan=mono|c0=FL" -ar 16000 left.wav
ffmpeg -i stereo.wav -af "pan=mono|c0=FR" -ar 16000 right.wav
# Transcribe each channel as a separate speaker# Download audio only
yt-dlp -x --audio-format wav -o "%(title)s.%(ext)s" "{url}"
# Then run the standard preprocessing + transcription pipelineThe hallucination cascade problem. Fix: use --condition_on_previous_text False. If already set, the input likely has long silence or music — preprocess with silence removal.
Treat this as a failed diarization, not a usable transcript. Common causes:
Fix order:
Deepgram or AssemblyAI) if a key is available.whisperx --diarize if HF_TOKEN is available and pyannote model terms are accepted.num_speakers set for the interview, preferably on CUDA/cloud GPU.Whisper's Python implementation doesn't use Metal/MPS well. Options:
whisper.cpp for Metal GPU acceleration--model turbo instead of large-v3 (2-3x faster, minimal quality loss)--model medium.en for English-only (4-5x faster)faster-whisper with compute_type="int8" (halves memory)medium or small)Modern macOS Python (Homebrew) requires:
pip3 install --break-system-packages {package}Whisper's native word timestamps are approximate. For precise word-level timing, use whisperX which adds forced phoneme alignment.
--language to let whisper auto-detectlarge-v3 (best multilingual model)--initial_prompt with example text in the target language| Backend | Install | Diarization | VAD | Speed (1hr file, Apple Silicon) | Cost |
|---|---|---|---|---|---|
| whisper | pip install openai-whisper | No | No | ~20-40 min (turbo) | Free |
| whisperx | pip install whisperx | Yes, requires HF token/model terms | Yes | ~15-30 min | Free |
| pyannote.audio | uv pip install "pyannote.audio" | Yes, requires HF token/model terms | Yes | GPU/cloud recommended | Free |
| NeMo clustering/MSDD | uv pip install "nemo_toolkit[asr]" | Yes, no secret | Yes | GPU/cloud recommended | Free |
| NeMo Sortformer | nemo_toolkit[asr] + supported examples | Yes, no secret | Model-dependent | GPU recommended | Free |
| faster-whisper | pip install faster-whisper | No | Yes (Silero) | ~10-20 min | Free |
| insanely-fast-whisper | pip install insanely-fast-whisper | Experimental | No | ~5-10 min (GPU) | Free |
| whisper.cpp | brew install whisper-cpp | Basic | No | ~10-15 min (Metal) | Free |
| OpenAI API | API key | No | N/A | ~1-2 min | $0.006/min |
| Groq API | API key | No | N/A | ~seconds | $0.00004/min |
| Deepgram | API key | Yes (native) | N/A | ~1-2 min | $0.0043/min |
| AssemblyAI | API key | Yes (native) | N/A | ~2-5 min | $0.0062/min |
| Gemini | API key | Prompted | N/A | ~1-3 min | Token-based |
© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in transcribe-anything of swyxio/skills.
Open the folder on GitHubat commit 038ef34
Transcribe Anything next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Transcribe Anything this skillswyxio/skills | 176 | — | ~8.5k | Automated safety check: Pass | MIT | |
| 9Router Speech-to-Textdecolua/9router | 31k | — | ~914 | Automated safety check: Pass | MIT | |
| Use Local Whispersbusso/claudeclaw | 194 | — | ~1.3k | Automated safety check: Notes | MIT | |
| Watch Videocoreyhaines31/makerskills | 851 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Keirouter Sttmydisha/keirouter | 147 | — | ~680 | Automated safety check: Pass | MIT | |
| Bggg Tiktok Readvideobinggandata/bggg-skills | 605 | — | ~1.6k | Automated safety check: Pass | MIT |
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
sbusso/claudeclaw
A skill your agent uses when the user wants local voice transcription instead of OpenAI Whisper API.
coreyhaines31/makerskills
When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.
mydisha/keirouter
Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.
binggandata/bggg-skills
把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。
jeremylongshore/tons-of-skills-marketplace
Deep dive into migrating to Deepgram from other transcription providers.
swyxio/skills
Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.
swyxio/skills
Design, implement, audit, or refresh protected username and handle namespaces for public products.
swyxio/skills
Fully automated new Mac setup for fullstack web developers and AI engineers.
swyxio/skills
Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…
swyxio/skills
Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.
swyxio/skills
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.
Works with
Categories
Transcribes audio and video files to text using pluggable ASR backends. Transcribe Anything is an agent skill from swyxio/skills. Transcribes audio and video files to text using pluggable ASR backends.
Transcribe Anything fits situations like: someone says transcribe this; convert to text; get the transcript; transcribe this video/audio/podcast/recording.
Run `npx skills add swyxio/skills --skill transcribe-anything -a claude-code`. Or copy the skill folder (transcribe-anything in swyxio/skills) into .claude/skills/transcribe-anything in your project. Claude Code loads it when a task matches its description.
Run `npx skills add swyxio/skills --skill transcribe-anything -a codex`. Or copy the skill folder (transcribe-anything in swyxio/skills) into .agents/skills/transcribe-anything in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill transcribe-anything -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcribe-anything, .gemini/skills/transcribe-anything, .github/skills/transcribe-anything and .opencode/skills/transcribe-anything in your project.
Going by SKILL.md and its folder, Transcribe Anything needs the command-line tools its instructions call (ffmpeg, whisper, uv, pip3, curl and ffprobe) and credentials named HF_TOKEN, OPENAI_API_KEY, GROQ_API_KEY and DEEPGRAM_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in GROQ_API_KEY. Compatibility (from SKILL.md): Requires ffmpeg and at least one ASR backend; MLX Whisper or openai-whisper is the default on macOS. For diarization, prefer managed APIs when configured, pyannote.audio on CUDA with Hugging Face access, or NeMo on CUDA without secrets. Never run long diarization on local CPU without explicit opt-in. Cloud backends require their API keys in environment variables. .
SKILL.md names 5 domains. In commands or code: github.com, api.openai.com, api.deepgram.com and api.groq.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Transcribe Anything is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.5k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Transcribe Anything: 9Router Speech-to-Text (decolua/9router, 31k stars), Use Local Whisper (sbusso/claudeclaw, 194 stars), Watch Video (coreyhaines31/makerskills, 851 stars) and Keirouter Stt (mydisha/keirouter, 147 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.
Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.