Agent skill

Use Local Whisper

by sbusso in sbusso/claudeclaw

A skill your agent uses when the user wants local voice transcription instead of OpenAI Whisper API.

MITAuto-check: notesAI & LLM Engineering

Install Use Local Whisper

skills CLI
$ npx skills add sbusso/claudeclaw --skill use-local-whisper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sbusso/claudeclaw use-local-whisper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sbusso/claudeclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/use-local-whisper .claude/skills/use-local-whisper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
use-local-whisper
GitHub stars
194
Token cost
~1.3k tokens
SKILL.md length
445 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants local voice transcription instead of OpenAI Whisper API.

  • Works in 3 steps: Pre-flight → Apply Code Changes → Verify
  • The user wants local voice transcription instead of OpenAI Whisper API
  • SKILL.md covers Prerequisites, Phase 1: Pre-flight, Phase 2: Apply Code Changes and Phase 3: Verify, plus 2 more sections
  • Calls git, brew and ffmpeg; reaches huggingface.co and github.com

What it does

Use Local Whisper is an agent skill from sbusso/claudeclaw. Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Transcription. It works with Whisper, WhatsApp, Homebrew and FFmpeg. The repository describes itself as: Use Claude to orchestrate agents like OpenClaw. The licence is MIT.

When your agent uses it

  • The user wants local voice transcription instead of OpenAI Whisper API
  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Transcription

Example prompts

  • “/use-local-whisper”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Pre-flight
  2. Apply Code Changes
  3. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 1395af4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • brew
    • ffmpeg
    • npm
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Use Local Whisper loads about 1.3k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:135
    Environment variables (optional, set in `.env`):

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sbusso/claudeclaw at commit 1395af4, republished under its MIT licence (© sbusso). 445 words, ~1,275 tokens.

Download SKILL.mdSave it as .claude/skills/use-local-whisper/SKILL.md (or your agent's skills folder).
name
use-local-whisper
description
Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

Use Local Whisper

Switches voice transcription from OpenAI's Whisper API to local whisper.cpp. Runs entirely on-device — no API key, no network, no cost.

Channel support: Currently WhatsApp only. The transcription module (src/transcription.ts) uses Baileys types for audio download. Other channels (Telegram, Discord, etc.) would need their own audio-download logic before this skill can serve them.

Note: The Homebrew package is whisper-cpp, but the CLI binary it installs is whisper-cli.

Prerequisites

  • voice-transcription skill must be applied first (WhatsApp channel)
  • macOS with Apple Silicon (M1+) recommended
  • whisper-cpp installed: brew install whisper-cpp (provides the whisper-cli binary)
  • ffmpeg installed: brew install ffmpeg
  • A GGML model file downloaded to data/models/

Phase 1: Pre-flight

Check if already applied

Check if src/transcription.ts already uses whisper-cli:

bash
grep 'whisper-cli' src/transcription.ts && echo "Already applied" || echo "Not applied"

If already applied, skip to Phase 3 (Verify).

Check dependencies are installed
bash
whisper-cli --help >/dev/null 2>&1 && echo "WHISPER_OK" || echo "WHISPER_MISSING"
ffmpeg -version >/dev/null 2>&1 && echo "FFMPEG_OK" || echo "FFMPEG_MISSING"

If missing, install via Homebrew:

bash
brew install whisper-cpp ffmpeg
Check for model file
bash
ls data/models/ggml-*.bin 2>/dev/null || echo "NO_MODEL"

If no model exists, download the base model (148MB, good balance of speed and accuracy):

bash
mkdir -p data/models
curl -L -o data/models/ggml-base.bin "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.bin"

For better accuracy at the cost of speed, use ggml-small.bin (466MB) or ggml-medium.bin (1.5GB).

Phase 2: Apply Code Changes

Ensure WhatsApp fork remote
bash
git remote -v

If whatsapp is missing, add it:

bash
git remote add whatsapp https://github.com/qwibitai/claudeclaw-whatsapp.git
Merge the skill branch
bash
git fetch whatsapp skill/local-whisper
git merge whatsapp/skill/local-whisper || {
  git checkout --theirs package-lock.json
  git add package-lock.json
  git merge --continue
}

This modifies src/transcription.ts to use the whisper-cli binary instead of the OpenAI API.

Validate
bash
npm run build

Phase 3: Verify

Ensure launchd PATH includes Homebrew

The ClaudeClaw launchd service runs with a restricted PATH. whisper-cli and ffmpeg are in /opt/homebrew/bin/ (Apple Silicon) or /usr/local/bin/ (Intel), which may not be in the plist's PATH.

Service name: Derived from the directory name: com.claudeclaw.<dirname> (macOS) / claudeclaw-<dirname> (Linux). For example, if cwd is my-assistant, the service is com.claudeclaw.my-assistant. Determine the correct service name before running service commands below.

Check the current PATH:

bash
grep -A1 'PATH' ~/Library/LaunchAgents/com.claudeclaw.plist

If /opt/homebrew/bin is missing, add it to the <string> value inside the PATH key in the plist. Then reload:

bash
launchctl unload ~/Library/LaunchAgents/com.claudeclaw.plist
launchctl load ~/Library/LaunchAgents/com.claudeclaw.plist
Show full SKILL.md (154 more words)Show less
Build and restart
bash
npm run build
launchctl kickstart -k gui/$(id -u)/com.claudeclaw
Test

Send a voice note in any registered group. The agent should receive it as [Voice: <transcript>].

Check logs
bash
tail -f logs/claudeclaw.log | grep -i -E "voice|transcri|whisper"

Look for:

  • Transcribed voice message — successful transcription
  • whisper.cpp transcription failed — check model path, ffmpeg, or PATH

Configuration

Environment variables (optional, set in .env):

VariableDefaultDescription
WHISPER_BINwhisper-cliPath to whisper.cpp binary
WHISPER_MODELdata/models/ggml-base.binPath to GGML model file

Troubleshooting

"whisper.cpp transcription failed": Ensure both whisper-cli and ffmpeg are in PATH. The launchd service uses a restricted PATH — see Phase 3 above. Test manually:

bash
ffmpeg -f lavfi -i anullsrc=r=16000:cl=mono -t 1 -f wav /tmp/test.wav -y
whisper-cli -m data/models/ggml-base.bin -f /tmp/test.wav --no-timestamps -nt

Transcription works in dev but not as service: The launchd plist PATH likely doesn't include /opt/homebrew/bin. See "Ensure launchd PATH includes Homebrew" in Phase 3.

Slow transcription: The base model processes ~30s of audio in <1s on M1+. If slower, check CPU usage — another process may be competing.

Wrong language: whisper.cpp auto-detects language. To force a language, you can set WHISPER_LANG and modify src/transcription.ts to pass -l $WHISPER_LANG.

© sbusso, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/use-local-whisper of sbusso/claudeclaw.

Open the folder on GitHubat commit 1395af4

Compare with similar skills

Use Local Whisper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Use Local Whisper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Use Local Whisper this skillsbusso/claudeclaw194—~1.3kAutomated safety check: NotesMIT
Transcribe Anythingswyxio/skills176—~8.5kAutomated safety check: PassMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Bggg Tiktok Readvideobinggandata/bggg-skills605—~1.6kAutomated safety check: PassMIT
Openai Whisper APItrpc-group/trpc-agent-go1.9k12 repos~288Automated safety check: PassApache-2.0
Lecture To Notesysyecust/lecture-to-notes273—~14kAutomated safety check: NotesCustom licence

Similar skills

  • Transcribe Anything

    swyxio/skills

    Transcribes audio and video files to text using pluggable ASR backends.

    176 GitHub stars~8.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Bggg Tiktok Readvideo

    binggandata/bggg-skills

    把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。

    605 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed
  • Lecture To Notes

    ysyecust/lecture-to-notes

    A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified…

    273 GitHub stars~14k tokensUpdated 7 days ago
    Documents & OfficeAuto-check: notes
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes

More from sbusso/claudeclaw

All 23 skills in this repo
  • Debug

    sbusso/claudeclaw

    Debug container agent issues. An agent skill from sbusso/claudeclaw.

    194 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes
  • X Integration

    sbusso/claudeclaw

    X (Twitter) integration for ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Gmail

    sbusso/claudeclaw

    Add Gmail integration to ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Qmd

    sbusso/claudeclaw

    Add QMD (Query Markup Documents) as an advanced memory search backend.

    194 GitHub stars~629 tokensUpdated 1 mo ago
    Auto-check passed
  • Add Telegram

    sbusso/claudeclaw

    Add Telegram as a channel. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Telegram Swarm

    sbusso/claudeclaw

    Add Agent Swarm (Teams) support to Telegram. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3.7k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Use Local Whisper

What does Use Local Whisper do?

A skill your agent uses when the user wants local voice transcription instead of OpenAI Whisper API. Use Local Whisper is an agent skill from sbusso/claudeclaw. Use when the user wants local voice transcription instead of OpenAI Whisper API.

When should I use Use Local Whisper?

Use Local Whisper fits situations like: the user wants local voice transcription instead of OpenAI Whisper API; tasks that involve Speech recognition and synthesis; tasks that involve Transcription.

How do I install Use Local Whisper in Claude Code?

Run `npx skills add sbusso/claudeclaw --skill use-local-whisper -a claude-code`. Or copy the skill folder (skills/use-local-whisper in sbusso/claudeclaw) into .claude/skills/use-local-whisper in your project. Claude Code loads it when a task matches its description.

How do I install Use Local Whisper in Codex?

Run `npx skills add sbusso/claudeclaw --skill use-local-whisper -a codex`. Or copy the skill folder (skills/use-local-whisper in sbusso/claudeclaw) into .agents/skills/use-local-whisper in your project. Codex loads it when a task matches its description.

Can I use Use Local Whisper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sbusso/claudeclaw --skill use-local-whisper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/use-local-whisper, .gemini/skills/use-local-whisper, .github/skills/use-local-whisper and .opencode/skills/use-local-whisper in your project.

What does Use Local Whisper need to run?

Going by SKILL.md and its folder, Use Local Whisper needs the command-line tools its instructions call (git, brew, ffmpeg, npm and curl).

Does Use Local Whisper access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Use Local Whisper safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Use Local Whisper use?

Use Local Whisper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Use Local Whisper use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Use Local Whisper?

Skills that share tags, products or a category with Use Local Whisper: Transcribe Anything (swyxio/skills, 176 stars), Watch Video (coreyhaines31/makerskills, 851 stars), Bggg Tiktok Readvideo (binggandata/bggg-skills, 605 stars) and Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Use Local Whisper?

sbusso (a GitHub user) maintains it in sbusso/claudeclaw, which has 194 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 12, 2026.

Source: sbusso/claudeclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.