Agent skill

Whisper Speech Recognition

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

MITAuto-check: notesAI & LLM Engineering

Install Whisper Speech Recognition

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs whisper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/18-multimodal/whisper .claude/skills/whisper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
whisper
GitHub stars
13k
Used in
7 other repos
Token cost
~1.9k tokens
SKILL.md length
326 words
Files
2 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

  • Works in 10 steps: Use turbo model - Best speed/quality for… → Specify language - Faster than auto-detect → Add initial prompt - Improves technical… → …
  • Transcribing a podcast or recorded meeting to text
  • SKILL.md covers When to use Whisper, Quick start, Model sizes and Transcription options, plus 10 more sections
  • Calls whisper, pip and ffmpeg

What it does

This skill covers OpenAI's Whisper speech recognition model for speech-to-text, podcast or video transcription, meeting notes, noisy audio and multilingual processing. It supports 99 languages, transcription, translation to English and language identification, and it notes the model was trained on 680,000 hours of audio.

Install with pip install -U openai-whisper on Python 3.8 to 3.11. A table compares the six sizes, from tiny at 39M parameters to large at 1550M, with English-only and multilingual availability, relative speed and VRAM needs, and it suggests turbo for a balance of speed and quality and base for prototyping. Python examples show transcribe with automatic or fixed language, transcribe versus translate tasks, an initial prompt to give context, word-level timestamps and temperature fallback.

It also shows command-line usage, batch processing over a list of audio files, and a pointer to faster-whisper for streaming. Alternatives named are AssemblyAI for managed APIs and speaker diarization, Deepgram for real-time streaming and Google Speech-to-Text for cloud use.

When your agent uses it

  • Transcribing a podcast or recorded meeting to text
  • Translating non-English speech into English text
  • Getting word-level timestamps for subtitles
  • Batch transcribing a folder of audio files

Example prompts

  • “Transcribe ./audio/standup.mp3 with Whisper and save the text.”
  • “Translate this Spanish interview recording into English using the turbo model.”
  • “Transcribe every mp3 in ./podcast/episodes and include word timestamps.”

Requirements

  • Python 3.8 to 3.11 with the openai-whisper package
  • Enough VRAM for the model size, from about 1 GB for tiny to about 10 GB for large

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Use turbo model - Best speed/quality for English
  2. Specify language - Faster than auto-detect
  3. Add initial prompt - Improves technical terms
  4. Use GPU - 10-20× faster
  5. Batch process - More efficient
  6. Convert to WAV - Better compatibility
  7. Split long audio - <30 min chunks
  8. Check language support - Quality varies by language
  9. Use faster-whisper - 4× faster than openai-whisper
  10. Monitor VRAM - Scale model size to hardware

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • whisper
    • pip
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Whisper Speech Recognition loads about 1.9k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 326 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:46
    # Ubuntu: sudo apt install ffmpeg

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 326 words, ~1,859 tokens.

Download SKILL.mdSave it as .claude/skills/whisper/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
whisper
description
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Whisper, Speech Recognition, ASR, Multimodal, Multilingual, OpenAI, Speech-To-Text, Transcription, Translation, Audio Processing
dependencies
openai-whisper, transformers, torch

Whisper - Robust Speech Recognition

OpenAI's multilingual speech recognition model.

When to use Whisper

Use when:

  • Speech-to-text transcription (99 languages)
  • Podcast/video transcription
  • Meeting notes automation
  • Translation to English
  • Noisy audio transcription
  • Multilingual audio processing

Metrics:

  • 72,900+ GitHub stars
  • 99 languages supported
  • Trained on 680,000 hours of audio
  • MIT License

Use alternatives instead:

  • AssemblyAI: Managed API, speaker diarization
  • Deepgram: Real-time streaming ASR
  • Google Speech-to-Text: Cloud-based

Quick start

Installation
bash
# Requires Python 3.8-3.11
pip install -U openai-whisper

# Requires ffmpeg
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg
# Windows: choco install ffmpeg
Basic transcription
python
import whisper

# Load model
model = whisper.load_model("base")

# Transcribe
result = model.transcribe("audio.mp3")

# Print text
print(result["text"])

# Access segments
for segment in result["segments"]:
    print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")

Model sizes

python
# Available models
models = ["tiny", "base", "small", "medium", "large", "turbo"]

# Load specific model
model = whisper.load_model("turbo")  # Fastest, good quality
ModelParametersEnglish-onlyMultilingualSpeedVRAM
tiny39M✓✓~32x~1 GB
base74M✓✓~16x~1 GB
small244M✓✓~6x~2 GB
medium769M✓✓~2x~5 GB
large1550M✗✓1x~10 GB
turbo809M✗✓~8x~6 GB

Recommendation: Use turbo for best speed/quality, base for prototyping

Transcription options

Language specification
python
# Auto-detect language
result = model.transcribe("audio.mp3")

# Specify language (faster)
result = model.transcribe("audio.mp3", language="en")

# Supported: en, es, fr, de, it, pt, ru, ja, ko, zh, and 89 more
Task selection
python
# Transcription (default)
result = model.transcribe("audio.mp3", task="transcribe")

# Translation to English
result = model.transcribe("spanish.mp3", task="translate")
# Input: Spanish audio → Output: English text
Initial prompt
python
# Improve accuracy with context
result = model.transcribe(
    "audio.mp3",
    initial_prompt="This is a technical podcast about machine learning and AI."
)

# Helps with:
# - Technical terms
# - Proper nouns
# - Domain-specific vocabulary
Timestamps
python
# Word-level timestamps
result = model.transcribe("audio.mp3", word_timestamps=True)

for segment in result["segments"]:
    for word in segment["words"]:
        print(f"{word['word']} ({word['start']:.2f}s - {word['end']:.2f}s)")
Temperature fallback
python
# Retry with different temperatures if confidence low
result = model.transcribe(
    "audio.mp3",
    temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0)
)

Command line usage

bash
# Basic transcription
whisper audio.mp3

# Specify model
whisper audio.mp3 --model turbo

# Output formats
whisper audio.mp3 --output_format txt     # Plain text
whisper audio.mp3 --output_format srt     # Subtitles
whisper audio.mp3 --output_format vtt     # WebVTT
whisper audio.mp3 --output_format json    # JSON with timestamps

# Language
whisper audio.mp3 --language Spanish

# Translation
whisper spanish.mp3 --task translate

Batch processing

python
import os

audio_files = ["file1.mp3", "file2.mp3", "file3.mp3"]

for audio_file in audio_files:
    print(f"Transcribing {audio_file}...")
    result = model.transcribe(audio_file)

    # Save to file
    output_file = audio_file.replace(".mp3", ".txt")
    with open(output_file, "w") as f:
        f.write(result["text"])

Real-time transcription

python
# For streaming audio, use faster-whisper
# pip install faster-whisper

from faster_whisper import WhisperModel

model = WhisperModel("base", device="cuda", compute_type="float16")

# Transcribe with streaming
segments, info = model.transcribe("audio.mp3", beam_size=5)

for segment in segments:
    print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")

GPU acceleration

python
import whisper

# Automatically uses GPU if available
model = whisper.load_model("turbo")

# Force CPU
model = whisper.load_model("turbo", device="cpu")

# Force GPU
model = whisper.load_model("turbo", device="cuda")

# 10-20× faster on GPU

Integration with other tools

Subtitle generation
bash
# Generate SRT subtitles
whisper video.mp4 --output_format srt --language English

# Output: video.srt
With LangChain
python
from langchain.document_loaders import WhisperTranscriptionLoader

loader = WhisperTranscriptionLoader(file_path="audio.mp3")
docs = loader.load()

# Use transcription in RAG
from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings

vectorstore = Chroma.from_documents(docs, OpenAIEmbeddings())
Extract audio from video
bash
# Use ffmpeg to extract audio
ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav

# Then transcribe
whisper audio.wav

Best practices

  1. Use turbo model - Best speed/quality for English
  2. Specify language - Faster than auto-detect
  3. Add initial prompt - Improves technical terms
  4. Use GPU - 10-20× faster
  5. Batch process - More efficient
  6. Convert to WAV - Better compatibility
  7. Split long audio - <30 min chunks
  8. Check language support - Quality varies by language
  9. Use faster-whisper - 4× faster than openai-whisper
  10. Monitor VRAM - Scale model size to hardware

Performance

ModelReal-time factor (CPU)Real-time factor (GPU)
tiny~0.32~0.01
base~0.16~0.01
turbo~0.08~0.01
large~1.0~0.05

Real-time factor: 0.1 = 10× faster than real-time

Language support

Top-supported languages:

  • English (en)
  • Spanish (es)
  • French (fr)
  • German (de)
  • Italian (it)
  • Portuguese (pt)
  • Russian (ru)
  • Japanese (ja)
  • Korean (ko)
  • Chinese (zh)

Full list: 99 languages total

Limitations

  1. Hallucinations - May repeat or invent text
  2. Long-form accuracy - Degrades on >30 min audio
  3. Speaker identification - No diarization
  4. Accents - Quality varies
  5. Background noise - Can affect accuracy
  6. Real-time latency - Not suitable for live captioning

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in 18-multimodal/whisper of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/languages.md

Open the folder on GitHubat commit 773a529

Used in 7 other repositories

We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 7 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Whisper Speech Recognition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Whisper Speech Recognition compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Whisper Speech Recognition this skillOrchestra-Research/AI-Research-SKILLs13k7 repos~1.9kAutomated safety check: NotesMIT
Openai Whisper APItrpc-group/trpc-agent-go1.9k12 repos~288Automated safety check: PassApache-2.0
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Openai Whisper APIopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
Openai Whispercoco-research/coco513—~964Automated safety check: PassCustom licence
Keirouter Sttmydisha/keirouter147—~680Automated safety check: PassMIT

Similar skills

  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed
  • Openai Whisper

    coco-research/coco

    Speech-to-text transcription via OpenAI Whisper. An agent skill from coco-research/coco.

    513 GitHub stars~964 tokensUpdated today
    Media & CreativeAuto-check passed
  • Keirouter Stt

    mydisha/keirouter

    Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.

    147 GitHub stars~680 tokensUpdated 29 days ago
    Media & CreativeAuto-check passed
  • Deepgram Audio Intelligence for Python

    deepgram/deepgram-python-sdk

    Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.

    469 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    Auto-check passed

Questions about Whisper Speech Recognition

What does Whisper Speech Recognition do?

Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI. This skill covers OpenAI's Whisper speech recognition model for speech-to-text, podcast or video transcription, meeting notes, noisy audio and multilingual processing. It supports 99 languages, transcription, translation to English and language identification, and it notes the model was trained on 680,000 hours of audio.

When should I use Whisper Speech Recognition?

Whisper Speech Recognition fits situations like: transcribing a podcast or recorded meeting to text; translating non-English speech into English text; getting word-level timestamps for subtitles; batch transcribing a folder of audio files.

How do I install Whisper Speech Recognition in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a claude-code`. Or copy the skill folder (18-multimodal/whisper in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/whisper in your project. Claude Code loads it when a task matches its description.

How do I install Whisper Speech Recognition in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a codex`. Or copy the skill folder (18-multimodal/whisper in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/whisper in your project. Codex loads it when a task matches its description.

Can I use Whisper Speech Recognition in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill whisper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper, .gemini/skills/whisper, .github/skills/whisper and .opencode/skills/whisper in your project.

What does Whisper Speech Recognition need to run?

Going by SKILL.md and its folder, Whisper Speech Recognition needs the command-line tools its instructions call (whisper, pip and ffmpeg). Our summary lists: Python 3.8 to 3.11 with the openai-whisper package; Enough VRAM for the model size, from about 1 GB for tiny to about 10 GB for large.

Does Whisper Speech Recognition access the network?

SKILL.md names 2 domains. As links in the text: github.com and arxiv.org. This is read from the text; nothing was executed.

Is Whisper Speech Recognition safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Whisper Speech Recognition use?

Whisper Speech Recognition is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Whisper Speech Recognition use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Whisper Speech Recognition?

Skills that share tags, products or a category with Whisper Speech Recognition: Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Openai Whisper (coco-research/coco, 513 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Whisper Speech Recognition?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.