Agent skill

Whisper Transcription

by guia-matthieu in guia-matthieu/clawfu-skills

Transcribe audio and video files to text using OpenAI Whisper.

MITAuto-check: notesMedia & Creative

Install Whisper Transcription

skills CLI
$ npx skills add guia-matthieu/clawfu-skills --skill whisper-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guia-matthieu/clawfu-skills whisper-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guia-matthieu/clawfu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/automation/whisper-transcription .claude/skills/whisper-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
whisper-transcription
GitHub stars
150
Token cost
~1.2k tokens
SKILL.md length
321 words
Files
3 (incl. scripts)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Transcribe audio and video files to text using OpenAI Whisper.

  • Works in 4 steps: GPU acceleration - 10x faster with CUDA… → Audio extraction - Script auto-extracts… → Chunking - Long files auto-split for… → …
  • : converting podcasts to blog posts
  • SKILL.md covers When to Use This Skill, What Claude Does vs What You…, Dependencies and Commands, plus 7 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Whisper Transcription is an agent skill from guia-matthieu/clawfu-skills. Transcribe audio and video files to text using OpenAI Whisper. Use when: converting podcasts to blog posts; creating video subtitles; extracting quotes from interviews; repurposing video content to text; building searchable audio archives

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/main.py`).

It sits in Media & Creative, covering Transcription. It works with Whisper and YouTube. The repository describes itself as: 172 expert marketing skills for AI agents — ClawFu MCP Server. The licence is MIT.

When your agent uses it

  • : converting podcasts to blog posts
  • Creating video subtitles
  • Extracting quotes from interviews
  • Repurposing video content to text

Example prompts

  • “/whisper-transcription”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. GPU acceleration - 10x faster with CUDA GPU
  2. Audio extraction - Script auto-extracts audio from video
  3. Chunking - Long files auto-split for memory efficiency
  4. Language detection - Automatic, or specify with --language

What it can do on your machine

Read from SKILL.md and the folder at commit 4108f5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Whisper Transcription loads about 1.2k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:40
    # Ubuntu: sudo apt install ffmpeg

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from guia-matthieu/clawfu-skills at commit 4108f5c, republished under its MIT licence (© guia-matthieu). 321 words, ~1,180 tokens.

Download SKILL.mdSave it as .claude/skills/whisper-transcription/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
whisper-transcription
description
Transcribe audio and video files to text using OpenAI Whisper. Use when: converting podcasts to blog posts; creating video subtitles; extracting quotes from interviews; repurposing video content to text; building searchable audio archives
license
MIT
metadata.author
ClawFu
metadata.version
1.0.0
metadata.mcp-server
@clawfu/mcp-skills

Whisper Transcription

Transcribe any audio or video to text using OpenAI's Whisper model - the same technology powering ChatGPT voice features.

When to Use This Skill

  • Podcast repurposing - Convert episodes to blog posts, show notes, social snippets
  • Video subtitles - Generate SRT/VTT files for YouTube, social media
  • Interview extraction - Pull quotes and insights from recorded calls
  • Content audit - Make audio/video libraries searchable
  • Translation - Transcribe and translate foreign language content

What Claude Does vs What You Decide

Claude DoesYou Decide
Structures production workflowFinal creative direction
Suggests technical approachesEquipment and tool choices
Creates templates and checklistsQuality standards
Identifies best practicesBrand/voice decisions
Generates script outlinesFinal script approval

Dependencies

bash
pip install openai-whisper torch ffmpeg-python click
# Also requires ffmpeg installed on system
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg

Commands

Transcribe Single File
bash
python scripts/main.py transcribe audio.mp3 --model medium --output transcript.txt
python scripts/main.py transcribe video.mp4 --format srt --output subtitles.srt
Batch Transcription
bash
python scripts/main.py batch ./recordings/ --format txt --output ./transcripts/
Transcribe + Translate
bash
python scripts/main.py translate foreign-audio.mp3 --to en
Extract Timestamps
bash
python scripts/main.py timestamps podcast.mp3 --format json

Examples

Example 1: Podcast to Blog Post
bash
# Transcribe 1-hour podcast
python scripts/main.py transcribe episode-42.mp3 --model medium

# Output: episode-42.txt (full transcript with timestamps)
# Processing time: ~5 min for 1 hour audio on M1 Mac
Example 2: YouTube Subtitles
bash
# Generate SRT for video upload
python scripts/main.py transcribe marketing-video.mp4 --format srt

# Output: marketing-video.srt
# Upload directly to YouTube/Vimeo
Example 3: Batch Process Interview Library
bash
# Transcribe all recordings in folder
python scripts/main.py batch ./customer-interviews/ --model small --format txt

# Output: ./customer-interviews/*.txt (one per audio file)

Model Selection Guide

ModelSpeedAccuracyVRAMBest For
tinyFastest~70%1GBQuick drafts, short clips
baseFast~80%1GBSocial media clips
smallMedium~85%2GBPodcasts, interviews
mediumSlow~90%5GBProfessional transcripts
largeSlowest~95%10GBCritical accuracy needs

Recommendation: Start with small for most marketing content. Use medium for client deliverables.

Output Formats

FormatExtensionUse Case
txt.txtBlog posts, analysis
srt.srtVideo subtitles (YouTube)
vtt.vttWeb video subtitles
json.jsonProgrammatic access
tsv.tsvSpreadsheet analysis

Performance Tips

  1. GPU acceleration - 10x faster with CUDA GPU
  2. Audio extraction - Script auto-extracts audio from video
  3. Chunking - Long files auto-split for memory efficiency
  4. Language detection - Automatic, or specify with --language

Skill Boundaries

What This Skill Does Well
  • Structuring audio production workflows
  • Providing technical guidance
  • Creating quality checklists
  • Suggesting creative approaches
What This Skill Cannot Do
  • Replace audio engineering expertise
  • Make subjective creative decisions
  • Access or edit audio files directly
  • Guarantee commercial success

Skill Metadata

  • Mode: cyborg
yaml
category: automation
subcategory: audio-processing
dependencies: [openai-whisper, torch, ffmpeg-python]
difficulty: beginner
time_saved: 10+ hours/week

© guia-matthieu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/automation/whisper-transcription of guia-matthieu/clawfu-skills.

  • SKILL.md
  • scripts/main.py
  • scripts/requirements.txt

Open the folder on GitHubat commit 4108f5c

Compare with similar skills

Whisper Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Whisper Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Whisper Transcription this skillguia-matthieu/clawfu-skills150—~1.2kAutomated safety check: NotesMIT
Voice Memo SyncLeoYeAI/openclaw-master-skills2.2k—~5.1kAutomated safety check: PassMIT
Video To Subtitle Summaryimlewc/video-to-subtitle-summary-skill218—~4.6kAutomated safety check: NotesMIT
Youtube PublishAndonywang123/Epost197—~3.4kAutomated safety check: WarnNone
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Gemini Yt Video Transcriptsundial-org/awesome-openclaw-skills663—~293Automated safety check: PassNone

Similar skills

  • Voice Memo Sync

    LeoYeAI/openclaw-master-skills

    Sync, transcribe, and intelligently organize voice memos, audio/video files, and URLs.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Video To Subtitle Summary

    imlewc/video-to-subtitle-summary-skill

    A skill your agent uses when user provides a short video platform URL or local video/audio file and wants subtitles/AI summary, or when user asks to list their own AI Douyin historical tasks.

    218 GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Youtube Publish

    Andonywang123/Epost

    Prepare an English YouTube release with local Chinese-to-English translation, subtitles and cover localization, then use a deterministic script connected to dedicated Chrome and YouTube Studio to…

    197 GitHub stars~3.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Gemini Yt Video Transcript

    sundial-org/awesome-openclaw-skills

    Create a verbatim transcript for a YouTube URL using Google Gemini (speaker labels, paragraph breaks; no time codes).

    663 GitHub stars~293 tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~2.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from guia-matthieu/clawfu-skills

All 31 skills in this repo
  • Ab Test Stats

    guia-matthieu/clawfu-skills

    Calculate A/B test statistical significance. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~1k tokensUpdated 10 days ago
    Auto-check passed
  • Cohort Analysis

    guia-matthieu/clawfu-skills

    Analyze user retention by cohort. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~864 tokensUpdated 10 days ago
    Auto-check passed
  • Competitive Analysis

    guia-matthieu/clawfu-skills

    Analyze your competitive landscape using Porter's Five Forces and modern frameworks—understand industry dynamics, identify strategic opportunities, and position your business for sustainable…

    150 GitHub stars~5.2k tokensUpdated 10 days ago
    Auto-check passed
  • Competitor Monitor

    guia-matthieu/clawfu-skills

    Monitor competitor websites for changes. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~815 tokensUpdated 10 days ago
    Auto-check passed
  • Content Repurposer

    guia-matthieu/clawfu-skills

    Transform long-form content into multiple short-form pieces.

    150 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Deliverability Checker

    guia-matthieu/clawfu-skills

    Check email deliverability and DNS configuration. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~1.1k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Whisper Transcription

What does Whisper Transcription do?

Transcribe audio and video files to text using OpenAI Whisper. Whisper Transcription is an agent skill from guia-matthieu/clawfu-skills. Transcribe audio and video files to text using OpenAI Whisper.

When should I use Whisper Transcription?

Whisper Transcription fits situations like: : converting podcasts to blog posts; creating video subtitles; extracting quotes from interviews; repurposing video content to text.

How do I install Whisper Transcription in Claude Code?

Run `npx skills add guia-matthieu/clawfu-skills --skill whisper-transcription -a claude-code`. Or copy the skill folder (skills/automation/whisper-transcription in guia-matthieu/clawfu-skills) into .claude/skills/whisper-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Whisper Transcription in Codex?

Run `npx skills add guia-matthieu/clawfu-skills --skill whisper-transcription -a codex`. Or copy the skill folder (skills/automation/whisper-transcription in guia-matthieu/clawfu-skills) into .agents/skills/whisper-transcription in your project. Codex loads it when a task matches its description.

Can I use Whisper Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guia-matthieu/clawfu-skills --skill whisper-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/whisper-transcription, .gemini/skills/whisper-transcription, .github/skills/whisper-transcription and .opencode/skills/whisper-transcription in your project.

What does Whisper Transcription need to run?

Going by SKILL.md and its folder, Whisper Transcription needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Whisper Transcription access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Whisper Transcription safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Whisper Transcription use?

Whisper Transcription is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Whisper Transcription use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Whisper Transcription?

Skills that share tags, products or a category with Whisper Transcription: Voice Memo Sync (LeoYeAI/openclaw-master-skills, 2.2k stars), Video To Subtitle Summary (imlewc/video-to-subtitle-summary-skill, 218 stars), Youtube Publish (Andonywang123/Epost, 197 stars) and Watch Video (coreyhaines31/makerskills, 851 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Whisper Transcription?

guia-matthieu (a GitHub user) maintains it in guia-matthieu/clawfu-skills, which has 150 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 1, 2026.

Source: guia-matthieu/clawfu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.