Agent skill

Whisper

by AlexAI-MCP in AlexAI-MCP/hermes-CCC

OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

MITAuto-check passedMedia & Creative

Install Whisper

skills CLI
$ npx skills add AlexAI-MCP/hermes-CCC --skill whisper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexAI-MCP/hermes-CCC whisper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/whisper .claude/skills/whisper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
whisper
GitHub stars
135
Token cost
~1.9k tokens
SKILL.md length
656 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

  • Tasks that involve Transcription
  • SKILL.md covers Purpose, Install, Command Line Usage and Python API, plus 18 more sections
  • Calls whisper and pip
  • Tasks that involve Speech recognition and synthesis

What it does

Whisper is an agent skill from AlexAI-MCP/hermes-CCC. OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with Whisper. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/whisper”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • whisper
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Whisper loads about 1.9k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 656 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 656 words, ~1,859 tokens.

Download SKILL.mdSave it as .claude/skills/whisper/SKILL.md (or your agent's skills folder).
name
whisper
description
OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.
version
1.0.0
author
hermes-CCC (ported from Hermes Agent by NousResearch)
license
MIT

Whisper

Purpose

  • Use this skill for local speech-to-text transcription and subtitle generation.
  • Prefer it when privacy, offline processing, or batch audio workflows matter.
  • Whisper works well for meetings, podcasts, interviews, voice notes, and extracted video audio.
  • It supports multilingual transcription and language detection.

Install

  • Standard Whisper package:
bash
pip install openai-whisper
  • Faster inference alternative:
bash
pip install faster-whisper
  • faster-whisper is often the better default for production batch jobs because it is typically faster at similar accuracy.

Command Line Usage

  • Basic CLI transcription:
bash
whisper audio.mp3 --model medium --language en
  • Use GPU explicitly:
bash
whisper audio.mp3 --model medium --language en --device cuda
  • Generate subtitles in SRT format:
bash
whisper audio.mp3 --model medium --language en --output_format srt
  • The CLI is a good fit for one-off transcription, shell scripts, and media preprocessing jobs.

Python API

python
import whisper

model = whisper.load_model("medium")
result = model.transcribe("audio.mp3")

print(result["text"])
print(result["language"])
print(result["segments"][:2])
  • Standard load pattern:
  • import whisper
  • model = whisper.load_model('medium')

What transcribe() Returns

  • text: the full transcript

  • segments: timestamped segment-level outputs

  • language: detected or selected language code

  • Example shape:

python
{
    "text": "Full transcript text",
    "language": "en",
    "segments": [
        {"id": 0, "start": 0.0, "end": 4.5, "text": "Hello everyone"},
    ],
}

Model Sizes

  • tiny

  • base

  • small

  • medium

  • large-v3

  • The tradeoff is simple:

  • smaller models are faster and cheaper

  • larger models are slower but more accurate

Model Selection Guidance

  • tiny: fast experiments and low-resource CPU runs
  • base: simple automation on clean audio
  • small: balanced for lightweight production tasks
  • medium: common quality default for serious transcription
  • large-v3: best accuracy when latency and VRAM are acceptable

Language Selection

  • Set the language when you know it:
bash
whisper audio.mp3 --model medium --language en
  • Explicit language hints usually improve speed and stability.
  • If the language is unknown, let the model detect it.

Language Detection

  • Whisper can estimate the spoken language from audio features.
  • Example pattern:
python
import whisper

model = whisper.load_model("medium")
audio = whisper.load_audio("audio.mp3")
audio = whisper.pad_or_trim(audio)
mel = whisper.log_mel_spectrogram(audio).to(model.device)

_, probs = model.detect_language(mel)
language = max(probs, key=probs.get)
print(language)
  • model.detect_language(audio) is the key workflow concept, though in practice you pass the processed spectrogram tensor.

Subtitle Generation

  • Generate .srt subtitles from the CLI:
bash
whisper audio.mp3 --model medium --language en --output_format srt
  • Subtitle outputs are useful for:
  • video captions
  • podcast transcripts
  • lecture indexing
  • searchable archives

Batch Processing Multiple Files

  • Simple shell loop:
bash
Get-ChildItem *.mp3 | ForEach-Object {
  whisper $_.FullName --model medium --language en --output_format srt
}
  • Python batch example:
python
from pathlib import Path

import whisper

model = whisper.load_model("medium")

for path in Path("audio").glob("*.mp3"):
    result = model.transcribe(str(path))
    out_path = path.with_suffix(".txt")
    out_path.write_text(result["text"], encoding="utf-8")
  • Batch processing is a common pattern for meeting folders, call archives, and downloaded media collections.

GPU vs CPU

  • GPU example from the CLI:
bash
whisper audio.mp3 --model medium --device cuda
  • GPU is strongly preferred for:

  • medium

  • large-v3

  • multi-file batch jobs

  • CPU is acceptable for:

  • tiny

  • base

  • occasional short clips

Common Use Cases

  • YouTube audio transcription after extracting audio from video
  • meeting notes from Zoom or Teams recordings
  • podcast transcription for search and republishing
  • lecture indexing and subtitle generation
  • multilingual voice note transcription

faster-whisper

  • Install with:
bash
pip install faster-whisper
  • It is often around 4x faster while maintaining comparable accuracy.
  • It is a strong choice for MLOps pipelines where throughput matters.

faster-whisper Example

python
from faster_whisper import WhisperModel

model = WhisperModel("medium", device="cuda", compute_type="float16")
segments, info = model.transcribe("audio.mp3", beam_size=5)

print(info.language, info.language_probability)

for segment in segments:
    print(f"[{segment.start:.2f} -> {segment.end:.2f}] {segment.text}")
  • faster-whisper is particularly useful for server-side or queue-based transcription systems.
Show full SKILL.md (272 more words)Show less

Output Management

  • Save plain text for downstream NLP.
  • Save SRT for subtitles and media players.
  • Save structured segments for search, speaker labeling pipelines, or analytics.

Segment-Level Workflows

  • Segment timestamps make it easy to:
  • jump to precise moments in media
  • chunk transcripts for embeddings
  • build searchable meeting or lecture interfaces
  • align captions with edited video

Quality Tips

  • Use the cleanest source audio available.
  • Downmix weird multi-channel recordings if channels are corrupted or imbalanced.
  • Remove long leading silence when possible.
  • Pick a larger model for noisy audio, accents, or technical vocabulary.

Common Failure Modes

  • Slow runtime:

  • switch to faster-whisper

  • move from CPU to GPU

  • use a smaller model

  • Bad language choice:

  • set --language en or another known language explicitly

  • inspect the detected language before large batch runs

  • Poor transcript quality:

  • upgrade from base or small to medium or large-v3

  • improve the source audio

  • split very long recordings into manageable chunks

  • Memory issues:

  • use a smaller model

  • run on GPU with enough VRAM

  • use faster-whisper with a suitable compute type

  • Start with medium for general English transcription.
  • Move to large-v3 when quality matters more than speed.
  • Use faster-whisper for production batch jobs.
  • Save both transcript text and structured segments.

When To Use This Skill

  • You need local transcription rather than a hosted ASR API.
  • You want subtitle files for recorded media.
  • You need multilingual speech recognition with no cloud dependency.
  • You are building meeting, media, or archival transcription pipelines.

Quick Reference

  • Install: pip install openai-whisper
  • Faster option: pip install faster-whisper
  • CLI: whisper audio.mp3 --model medium --language en
  • GPU: --device cuda
  • SRT subtitles: --output_format srt
  • Python load: model = whisper.load_model("medium")
  • Result fields: text, segments, language

© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/whisper of AlexAI-MCP/hermes-CCC.

Open the folder on GitHubat commit 8107e89

Compare with similar skills

Whisper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Whisper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Whisper this skillAlexAI-MCP/hermes-CCC135—~1.9kAutomated safety check: PassMIT
Openai Whisper APICoWork-OS/CoWork-OS477—~411Automated safety check: PassMIT
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Openai Whisper APIopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Wjs Transcribing Audiojianshuo/claude-skills131—~4.4kAutomated safety check: NotesMIT

Similar skills

  • Openai Whisper API

    CoWork-OS/CoWork-OS

    Transcribe audio via OpenAI Whisper, Atlas Cloud, or MuAPI speech-to-text APIs.

    477 GitHub stars~411 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Wjs Transcribing Audio

    jianshuo/claude-skills

    A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.

    131 GitHub stars~4.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Whisper Transcription

    benchflow-ai/skillsbench

    Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from AlexAI-MCP/hermes-CCC

All 44 skills in this repo
  • GitHub Code Review

    AlexAI-MCP/hermes-CCC

    Review GitHub pull requests with a findings-first engineering mindset.

    135 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • GitHub PR Workflow

    AlexAI-MCP/hermes-CCC

    Run a disciplined GitHub pull request workflow from branch creation through merge.

    135 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Memory

    AlexAI-MCP/hermes-CCC

    Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Route

    AlexAI-MCP/hermes-CCC

    Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Skill

    AlexAI-MCP/hermes-CCC

    Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Traj

    AlexAI-MCP/hermes-CCC

    Capture Claude Code interaction trajectories in training-friendly formats.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Whisper

What does Whisper do?

OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation. Whisper is an agent skill from AlexAI-MCP/hermes-CCC. OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

When should I use Whisper?

Whisper fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Whisper in Claude Code?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill whisper -a claude-code`. Or copy the skill folder (skills/whisper in AlexAI-MCP/hermes-CCC) into .claude/skills/whisper in your project. Claude Code loads it when a task matches its description.

How do I install Whisper in Codex?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill whisper -a codex`. Or copy the skill folder (skills/whisper in AlexAI-MCP/hermes-CCC) into .agents/skills/whisper in your project. Codex loads it when a task matches its description.

Can I use Whisper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill whisper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper, .gemini/skills/whisper, .github/skills/whisper and .opencode/skills/whisper in your project.

What does Whisper need to run?

Going by SKILL.md and its folder, Whisper needs the command-line tools its instructions call (whisper and pip). Our summary lists: Python 3.

Does Whisper access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Whisper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Whisper use?

Whisper is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Whisper use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Whisper?

Skills that share tags, products or a category with Whisper: Openai Whisper API (CoWork-OS/CoWork-OS, 477 stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Openai Whisper API (openclaw/openclaw, 392k stars) and Watch Video (coreyhaines31/makerskills, 851 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Whisper?

AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.

Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.