Openai Whisper API
trpc-group/trpc-agent-go
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/18-multimodal/whisper .claude/skills/whisper && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .claude/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisperType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/18-multimodal/whisper .agents/skills/whisper && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .agents/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/18-multimodal/whisper .cursor/skills/whisper && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .cursor/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 18-multimodal/whisper--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/18-multimodal/whisper .gemini/skills/whisper && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .gemini/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisperInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/18-multimodal/whisper .github/skills/whisper && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .github/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/18-multimodal/whisper .opencode/skills/whisper && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "whisper" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/whisper into .opencode/skills/whisper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
whisperTranscribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
This skill covers OpenAI's Whisper speech recognition model for speech-to-text, podcast or video transcription, meeting notes, noisy audio and multilingual processing. It supports 99 languages, transcription, translation to English and language identification, and it notes the model was trained on 680,000 hours of audio.
Install with pip install -U openai-whisper on Python 3.8 to 3.11. A table compares the six sizes, from tiny at 39M parameters to large at 1550M, with English-only and multilingual availability, relative speed and VRAM needs, and it suggests turbo for a balance of speed and quality and base for prototyping. Python examples show transcribe with automatic or fixed language, transcribe versus translate tasks, an initial prompt to give context, word-level timestamps and temperature fallback.
It also shows command-line usage, batch processing over a list of audio files, and a pointer to faster-whisper for streaming. Alternatives named are AssemblyAI for managed APIs and speaker diarization, Deepgram for real-time streaming and Google Speech-to-Text for cloud use.
10 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
whisperpipffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comarxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Whisper Speech Recognition loads about 1.9k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 326 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
# Ubuntu: sudo apt install ffmpegAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 326 words, ~1,859 tokens.
.claude/skills/whisper/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.OpenAI's multilingual speech recognition model.
Use when:
Metrics:
Use alternatives instead:
# Requires Python 3.8-3.11
pip install -U openai-whisper
# Requires ffmpeg
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg
# Windows: choco install ffmpegimport whisper
# Load model
model = whisper.load_model("base")
# Transcribe
result = model.transcribe("audio.mp3")
# Print text
print(result["text"])
# Access segments
for segment in result["segments"]:
print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")# Available models
models = ["tiny", "base", "small", "medium", "large", "turbo"]
# Load specific model
model = whisper.load_model("turbo") # Fastest, good quality| Model | Parameters | English-only | Multilingual | Speed | VRAM |
|---|---|---|---|---|---|
| tiny | 39M | ✓ | ✓ | ~32x | ~1 GB |
| base | 74M | ✓ | ✓ | ~16x | ~1 GB |
| small | 244M | ✓ | ✓ | ~6x | ~2 GB |
| medium | 769M | ✓ | ✓ | ~2x | ~5 GB |
| large | 1550M | ✗ | ✓ | 1x | ~10 GB |
| turbo | 809M | ✗ | ✓ | ~8x | ~6 GB |
Recommendation: Use turbo for best speed/quality, base for prototyping
# Auto-detect language
result = model.transcribe("audio.mp3")
# Specify language (faster)
result = model.transcribe("audio.mp3", language="en")
# Supported: en, es, fr, de, it, pt, ru, ja, ko, zh, and 89 more# Transcription (default)
result = model.transcribe("audio.mp3", task="transcribe")
# Translation to English
result = model.transcribe("spanish.mp3", task="translate")
# Input: Spanish audio → Output: English text# Improve accuracy with context
result = model.transcribe(
"audio.mp3",
initial_prompt="This is a technical podcast about machine learning and AI."
)
# Helps with:
# - Technical terms
# - Proper nouns
# - Domain-specific vocabulary# Word-level timestamps
result = model.transcribe("audio.mp3", word_timestamps=True)
for segment in result["segments"]:
for word in segment["words"]:
print(f"{word['word']} ({word['start']:.2f}s - {word['end']:.2f}s)")# Retry with different temperatures if confidence low
result = model.transcribe(
"audio.mp3",
temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0)
)# Basic transcription
whisper audio.mp3
# Specify model
whisper audio.mp3 --model turbo
# Output formats
whisper audio.mp3 --output_format txt # Plain text
whisper audio.mp3 --output_format srt # Subtitles
whisper audio.mp3 --output_format vtt # WebVTT
whisper audio.mp3 --output_format json # JSON with timestamps
# Language
whisper audio.mp3 --language Spanish
# Translation
whisper spanish.mp3 --task translateimport os
audio_files = ["file1.mp3", "file2.mp3", "file3.mp3"]
for audio_file in audio_files:
print(f"Transcribing {audio_file}...")
result = model.transcribe(audio_file)
# Save to file
output_file = audio_file.replace(".mp3", ".txt")
with open(output_file, "w") as f:
f.write(result["text"])# For streaming audio, use faster-whisper
# pip install faster-whisper
from faster_whisper import WhisperModel
model = WhisperModel("base", device="cuda", compute_type="float16")
# Transcribe with streaming
segments, info = model.transcribe("audio.mp3", beam_size=5)
for segment in segments:
print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")import whisper
# Automatically uses GPU if available
model = whisper.load_model("turbo")
# Force CPU
model = whisper.load_model("turbo", device="cpu")
# Force GPU
model = whisper.load_model("turbo", device="cuda")
# 10-20× faster on GPU# Generate SRT subtitles
whisper video.mp4 --output_format srt --language English
# Output: video.srtfrom langchain.document_loaders import WhisperTranscriptionLoader
loader = WhisperTranscriptionLoader(file_path="audio.mp3")
docs = loader.load()
# Use transcription in RAG
from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings
vectorstore = Chroma.from_documents(docs, OpenAIEmbeddings())# Use ffmpeg to extract audio
ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav
# Then transcribe
whisper audio.wav| Model | Real-time factor (CPU) | Real-time factor (GPU) |
|---|---|---|
| tiny | ~0.32 | ~0.01 |
| base | ~0.16 | ~0.01 |
| turbo | ~0.08 | ~0.01 |
| large | ~1.0 | ~0.05 |
Real-time factor: 0.1 = 10× faster than real-time
Top-supported languages:
Full list: 99 languages total
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in 18-multimodal/whisper of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 7 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Whisper Speech Recognition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Whisper Speech Recognition this skillOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.9k | Automated safety check: Notes | MIT | |
| Openai Whisper APItrpc-group/trpc-agent-go | 1.9k | 12 repos | ~288 | Automated safety check: Pass | Apache-2.0 | |
| 9Router Speech-to-Textdecolua/9router | 31k | — | ~914 | Automated safety check: Pass | MIT | |
| Openai Whisper APIopenclaw/openclaw | 392k | 1 repos | ~518 | Automated safety check: Pass | MIT | |
| Openai Whispercoco-research/coco | 513 | — | ~964 | Automated safety check: Pass | Custom licence | |
| Keirouter Sttmydisha/keirouter | 147 | — | ~680 | Automated safety check: Pass | MIT |
trpc-group/trpc-agent-go
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
openclaw/openclaw
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
coco-research/coco
Speech-to-text transcription via OpenAI Whisper. An agent skill from coco-research/coco.
mydisha/keirouter
Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.
deepgram/deepgram-python-sdk
Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.
Categories
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI. This skill covers OpenAI's Whisper speech recognition model for speech-to-text, podcast or video transcription, meeting notes, noisy audio and multilingual processing. It supports 99 languages, transcription, translation to English and language identification, and it notes the model was trained on 680,000 hours of audio.
Whisper Speech Recognition fits situations like: transcribing a podcast or recorded meeting to text; translating non-English speech into English text; getting word-level timestamps for subtitles; batch transcribing a folder of audio files.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a claude-code`. Or copy the skill folder (18-multimodal/whisper in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/whisper in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a codex`. Or copy the skill folder (18-multimodal/whisper in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/whisper in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper, .gemini/skills/whisper, .github/skills/whisper and .opencode/skills/whisper in your project.
Going by SKILL.md and its folder, Whisper Speech Recognition needs the command-line tools its instructions call (whisper, pip and ffmpeg). Our summary lists: Python 3.8 to 3.11 with the openai-whisper package; Enough VRAM for the model size, from about 1 GB for tiny to about 10 GB for large.
SKILL.md names 2 domains. As links in the text: github.com and arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Whisper Speech Recognition is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Whisper Speech Recognition: Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Openai Whisper (coco-research/coco, 513 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.