Agent skill

Whisper Transcription

by benchflow-ai in benchflow-ai/skillsbench

Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.

Apache-2.0Auto-check passedMedia & Creative

Install Whisper Transcription

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .claude/skills/whisper-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
whisper-transcription
GitHub stars
1.8k
Token cost
~1.1k tokens
SKILL.md length
101 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.

  • You need speech-to-text with accurate timing information for each word
  • SKILL.md covers Installation, Model Selection, Basic Usage with Word Timestamps and Detecting Specific Words, plus 3 more sections
  • Calls pip and ffmpeg
  • Tasks that involve Transcription

What it does

Whisper Transcription is an agent skill from benchflow-ai/skillsbench. Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Use when you need speech-to-text with accurate timing information for each word.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with Whisper. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • You need speech-to-text with accurate timing information for each word
  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/whisper-transcription”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Whisper Transcription loads about 1.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 101 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 101 words, ~1,136 tokens.

Download SKILL.mdSave it as .claude/skills/whisper-transcription/SKILL.md (or your agent's skills folder).
name
whisper-transcription
description
Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Use when you need speech-to-text with accurate timing information for each word.

Whisper Transcription

OpenAI Whisper provides accurate speech-to-text with word-level timestamps.

Installation

bash
pip install openai-whisper

Model Selection

Use the tiny model for fast transcription - it's sufficient for most tasks and runs much faster:

ModelSizeSpeedAccuracy
tiny39 MBFastestGood for clear speech
base74 MBFastBetter accuracy
small244 MBMediumHigh accuracy

Recommendation: Start with tiny - it handles clear interview/podcast audio well.

Basic Usage with Word Timestamps

python
import whisper
import json

def transcribe_with_timestamps(audio_path, output_path):
    """
    Transcribe audio and get word-level timestamps.

    Args:
        audio_path: Path to audio/video file
        output_path: Path to save JSON output
    """
    # Use tiny model for speed
    model = whisper.load_model("tiny")

    # Transcribe with word timestamps
    result = model.transcribe(
        audio_path,
        word_timestamps=True,
        language="en"  # Specify language for better accuracy
    )

    # Extract words with timestamps
    words = []
    for segment in result["segments"]:
        if "words" in segment:
            for word_info in segment["words"]:
                words.append({
                    "word": word_info["word"].strip(),
                    "start": word_info["start"],
                    "end": word_info["end"]
                })

    with open(output_path, "w") as f:
        json.dump(words, f, indent=2)

    return words

Detecting Specific Words

python
def find_words(transcription, target_words):
    """
    Find specific words in transcription with their timestamps.

    Args:
        transcription: List of word dicts with 'word', 'start', 'end'
        target_words: Set of words to find (lowercase)

    Returns:
        List of matches with word and timestamp
    """
    matches = []
    target_lower = {w.lower() for w in target_words}

    for item in transcription:
        word = item["word"].lower().strip()
        # Remove punctuation for matching
        clean_word = ''.join(c for c in word if c.isalnum())

        if clean_word in target_lower:
            matches.append({
                "word": clean_word,
                "timestamp": item["start"]
            })

    return matches

Complete Example: Find Filler Words

python
import whisper
import json

# Filler words to detect
FILLER_WORDS = {
    "um", "uh", "hum", "hmm", "mhm",
    "like", "so", "well", "yeah", "okay",
    "basically", "actually", "literally"
}

def detect_fillers(audio_path, output_path):
    # Load tiny model (fast!)
    model = whisper.load_model("tiny")

    # Transcribe
    result = model.transcribe(audio_path, word_timestamps=True, language="en")

    # Find fillers
    fillers = []
    for segment in result["segments"]:
        for word_info in segment.get("words", []):
            word = word_info["word"].lower().strip()
            clean = ''.join(c for c in word if c.isalnum())

            if clean in FILLER_WORDS:
                fillers.append({
                    "word": clean,
                    "timestamp": round(word_info["start"], 2)
                })

    with open(output_path, "w") as f:
        json.dump(fillers, f, indent=2)

    return fillers

# Usage
detect_fillers("/root/input.mp4", "/root/annotations.json")

Audio Extraction (if needed)

Whisper can process video files directly, but for cleaner results:

bash
# Extract audio as 16kHz mono WAV
ffmpeg -i input.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 audio.wav

Multi-Word Phrases

For detecting phrases like "you know" or "I mean":

python
def find_phrases(transcription, phrases):
    """Find multi-word phrases in transcription."""
    matches = []
    words = [w["word"].lower().strip() for w in transcription]

    for phrase in phrases:
        phrase_words = phrase.lower().split()
        phrase_len = len(phrase_words)

        for i in range(len(words) - phrase_len + 1):
            if words[i:i+phrase_len] == phrase_words:
                matches.append({
                    "word": phrase,
                    "timestamp": transcription[i]["start"]
                })

    return matches

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Whisper Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Whisper Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Whisper Transcription this skillbenchflow-ai/skillsbench1.8k—~1.1kAutomated safety check: PassApache-2.0
Wjs Transcribing Audiojianshuo/claude-skills131—~4.4kAutomated safety check: NotesMIT
WhisperAlexAI-MCP/hermes-CCC135—~1.9kAutomated safety check: PassMIT
Faster Whispersundial-org/awesome-openclaw-skills663—~3kAutomated safety check: PassNone
Openai Whisper APICoWork-OS/CoWork-OS477—~411Automated safety check: PassMIT
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT

Similar skills

  • Wjs Transcribing Audio

    jianshuo/claude-skills

    A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.

    131 GitHub stars~4.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Whisper

    AlexAI-MCP/hermes-CCC

    OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed
  • Faster Whisper

    sundial-org/awesome-openclaw-skills

    Local speech-to-text using faster-whisper. An agent skill from sundial-org/awesome-openclaw-skills.

    663 GitHub stars~3k tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    CoWork-OS/CoWork-OS

    Transcribe audio via OpenAI Whisper, Atlas Cloud, or MuAPI speech-to-text APIs.

    477 GitHub stars~411 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Whisper Transcription

What does Whisper Transcription do?

Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Whisper Transcription is an agent skill from benchflow-ai/skillsbench. Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.

When should I use Whisper Transcription?

Whisper Transcription fits situations like: you need speech-to-text with accurate timing information for each word; tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Whisper Transcription in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a claude-code`. Or copy the skill folder (tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription in benchflow-ai/skillsbench) into .claude/skills/whisper-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Whisper Transcription in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a codex`. Or copy the skill folder (tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription in benchflow-ai/skillsbench) into .agents/skills/whisper-transcription in your project. Codex loads it when a task matches its description.

Can I use Whisper Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper-transcription, .gemini/skills/whisper-transcription, .github/skills/whisper-transcription and .opencode/skills/whisper-transcription in your project.

What does Whisper Transcription need to run?

Going by SKILL.md and its folder, Whisper Transcription needs the command-line tools its instructions call (pip and ffmpeg). Our summary lists: Python 3.

Does Whisper Transcription access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Whisper Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Whisper Transcription use?

Whisper Transcription is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Whisper Transcription use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Whisper Transcription?

Skills that share tags, products or a category with Whisper Transcription: Wjs Transcribing Audio (jianshuo/claude-skills, 131 stars), Whisper (AlexAI-MCP/hermes-CCC, 135 stars), Faster Whisper (sundial-org/awesome-openclaw-skills, 663 stars) and Openai Whisper API (CoWork-OS/CoWork-OS, 477 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Whisper Transcription?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.