Wjs Transcribing Audio
jianshuo/claude-skills
A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.
Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .claude/skills/whisper-transcription && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .claude/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcriptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .agents/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .agents/skills/whisper-transcription && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .agents/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .cursor/skills/whisper-transcription && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .cursor/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benchflow-ai/skillsbench.git --path tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .gemini/skills/whisper-transcription && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .gemini/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benchflow-ai/skillsbench whisper-transcriptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .github/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .github/skills/whisper-transcription && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .github/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benchflow-ai/skillsbench whisper-transcription --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription .opencode/skills/whisper-transcription && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "whisper-transcription" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription into .opencode/skills/whisper-transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "whisper-transcription", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
whisper-transcriptionTranscribe audio/video to text with word-level timestamps using OpenAI Whisper.
Whisper Transcription is an agent skill from benchflow-ai/skillsbench. Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Use when you need speech-to-text with accurate timing information for each word.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with Whisper. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Whisper Transcription loads about 1.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 101 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 101 words, ~1,136 tokens.
.claude/skills/whisper-transcription/SKILL.md (or your agent's skills folder).OpenAI Whisper provides accurate speech-to-text with word-level timestamps.
pip install openai-whisperUse the tiny model for fast transcription - it's sufficient for most tasks and runs much faster:
| Model | Size | Speed | Accuracy |
|---|---|---|---|
| tiny | 39 MB | Fastest | Good for clear speech |
| base | 74 MB | Fast | Better accuracy |
| small | 244 MB | Medium | High accuracy |
Recommendation: Start with tiny - it handles clear interview/podcast audio well.
import whisper
import json
def transcribe_with_timestamps(audio_path, output_path):
"""
Transcribe audio and get word-level timestamps.
Args:
audio_path: Path to audio/video file
output_path: Path to save JSON output
"""
# Use tiny model for speed
model = whisper.load_model("tiny")
# Transcribe with word timestamps
result = model.transcribe(
audio_path,
word_timestamps=True,
language="en" # Specify language for better accuracy
)
# Extract words with timestamps
words = []
for segment in result["segments"]:
if "words" in segment:
for word_info in segment["words"]:
words.append({
"word": word_info["word"].strip(),
"start": word_info["start"],
"end": word_info["end"]
})
with open(output_path, "w") as f:
json.dump(words, f, indent=2)
return wordsdef find_words(transcription, target_words):
"""
Find specific words in transcription with their timestamps.
Args:
transcription: List of word dicts with 'word', 'start', 'end'
target_words: Set of words to find (lowercase)
Returns:
List of matches with word and timestamp
"""
matches = []
target_lower = {w.lower() for w in target_words}
for item in transcription:
word = item["word"].lower().strip()
# Remove punctuation for matching
clean_word = ''.join(c for c in word if c.isalnum())
if clean_word in target_lower:
matches.append({
"word": clean_word,
"timestamp": item["start"]
})
return matchesimport whisper
import json
# Filler words to detect
FILLER_WORDS = {
"um", "uh", "hum", "hmm", "mhm",
"like", "so", "well", "yeah", "okay",
"basically", "actually", "literally"
}
def detect_fillers(audio_path, output_path):
# Load tiny model (fast!)
model = whisper.load_model("tiny")
# Transcribe
result = model.transcribe(audio_path, word_timestamps=True, language="en")
# Find fillers
fillers = []
for segment in result["segments"]:
for word_info in segment.get("words", []):
word = word_info["word"].lower().strip()
clean = ''.join(c for c in word if c.isalnum())
if clean in FILLER_WORDS:
fillers.append({
"word": clean,
"timestamp": round(word_info["start"], 2)
})
with open(output_path, "w") as f:
json.dump(fillers, f, indent=2)
return fillers
# Usage
detect_fillers("/root/input.mp4", "/root/annotations.json")Whisper can process video files directly, but for cleaner results:
# Extract audio as 16kHz mono WAV
ffmpeg -i input.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 audio.wavFor detecting phrases like "you know" or "I mean":
def find_phrases(transcription, phrases):
"""Find multi-word phrases in transcription."""
matches = []
words = [w["word"].lower().strip() for w in transcription]
for phrase in phrases:
phrase_words = phrase.lower().split()
phrase_len = len(phrase_words)
for i in range(len(words) - phrase_len + 1):
if words[i:i+phrase_len] == phrase_words:
matches.append({
"word": phrase,
"timestamp": transcription[i]["start"]
})
return matches© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription of benchflow-ai/skillsbench.
Open the folder on GitHubat commit 9a1f4dd
Whisper Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Whisper Transcription this skillbenchflow-ai/skillsbench | 1.8k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Wjs Transcribing Audiojianshuo/claude-skills | 131 | — | ~4.4k | Automated safety check: Notes | MIT | |
| WhisperAlexAI-MCP/hermes-CCC | 135 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Faster Whispersundial-org/awesome-openclaw-skills | 663 | — | ~3k | Automated safety check: Pass | None | |
| Openai Whisper APICoWork-OS/CoWork-OS | 477 | — | ~411 | Automated safety check: Pass | MIT | |
| 9Router Speech-to-Textdecolua/9router | 31k | — | ~914 | Automated safety check: Pass | MIT |
jianshuo/claude-skills
A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.
AlexAI-MCP/hermes-CCC
OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.
sundial-org/awesome-openclaw-skills
Local speech-to-text using faster-whisper. An agent skill from sundial-org/awesome-openclaw-skills.
CoWork-OS/CoWork-OS
Transcribe audio via OpenAI Whisper, Atlas Cloud, or MuAPI speech-to-text APIs.
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
openclaw/openclaw
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
benchflow-ai/skillsbench
This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
benchflow-ai/skillsbench
AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.
benchflow-ai/skillsbench
Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.
benchflow-ai/skillsbench
Build deterministic, verifiable data visualizations with D3.js (v6).
benchflow-ai/skillsbench
DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.
Works with
Categories
Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Whisper Transcription is an agent skill from benchflow-ai/skillsbench. Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.
Whisper Transcription fits situations like: you need speech-to-text with accurate timing information for each word; tasks that involve Transcription; tasks that involve Speech recognition and synthesis.
Run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a claude-code`. Or copy the skill folder (tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription in benchflow-ai/skillsbench) into .claude/skills/whisper-transcription in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a codex`. Or copy the skill folder (tasks-extra/video-filler-word-remover/environment/skills/whisper-transcription in benchflow-ai/skillsbench) into .agents/skills/whisper-transcription in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill whisper-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper-transcription, .gemini/skills/whisper-transcription, .github/skills/whisper-transcription and .opencode/skills/whisper-transcription in your project.
Going by SKILL.md and its folder, Whisper Transcription needs the command-line tools its instructions call (pip and ffmpeg). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Whisper Transcription is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Whisper Transcription: Wjs Transcribing Audio (jianshuo/claude-skills, 131 stars), Whisper (AlexAI-MCP/hermes-CCC, 135 stars), Faster Whisper (sundial-org/awesome-openclaw-skills, 663 stars) and Openai Whisper API (CoWork-OS/CoWork-OS, 477 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.
Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.