Agent skill

Add Voice Transcription

by sbusso in sbusso/claudeclaw

Add voice message transcription to ClaudeClaw using OpenAI's Whisper API.

MITAuto-check: notesMedia & Creative

Install Add Voice Transcription

skills CLI
$ npx skills add sbusso/claudeclaw --skill add-voice-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sbusso/claudeclaw add-voice-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sbusso/claudeclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/add-voice-transcription .claude/skills/add-voice-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-voice-transcription
GitHub stars
194
Token cost
~1.1k tokens
SKILL.md length
484 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Add voice message transcription to ClaudeClaw using OpenAI's Whisper API.

  • Works in 4 steps: Pre-flight → Apply Code Changes → Configure → …
  • Tasks that involve Transcription
  • SKILL.md covers Phase 1: Pre-flight, Phase 2: Apply Code Changes, Phase 3: Configure and Phase 4: Verify, plus 1 more section
  • Calls git, npm and npx; reaches api.openai.com and github.com; needs OPENAI_API_KEY

What it does

Add Voice Transcription is an agent skill from sbusso/claudeclaw. Add voice message transcription to ClaudeClaw using OpenAI's Whisper API. Automatically transcribes WhatsApp voice notes so the agent can read and respond to them.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. It works with OpenAI, WhatsApp and Whisper. The repository describes itself as: Use Claude to orchestrate agents like OpenClaw. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription

Example prompts

  • “/add-voice-transcription”

Requirements

  • Node.js
  • A credential in OPENAI_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Pre-flight
  2. Apply Code Changes
  3. Configure
  4. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 1395af4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npm
    • npx
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com
    • github.com

    Also links to:

    • platform.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Voice Transcription loads about 1.1k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 484 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:89
    Add to `.env`:
  • NoteMentions a .env fileSKILL.md:98
    mkdir -p data/env && cp .env data/env/env
  • NoteMentions a .env fileSKILL.md:101
    ds environment from `data/env/env`, not `.env` directly.
  • NoteMentions a .env fileSKILL.md:129
    NAI_API_KEY not set` — key missing from `.env`
  • NoteMentions a .env fileSKILL.md:137
    1. Check `OPENAI_API_KEY` is set in `.env` AND synced to `data/env/env`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sbusso/claudeclaw at commit 1395af4, republished under its MIT licence (© sbusso). 484 words, ~1,137 tokens.

Download SKILL.mdSave it as .claude/skills/add-voice-transcription/SKILL.md (or your agent's skills folder).
name
add-voice-transcription
description
Add voice message transcription to ClaudeClaw using OpenAI's Whisper API. Automatically transcribes WhatsApp voice notes so the agent can read and respond to them.

Add Voice Transcription

This skill adds automatic voice message transcription to ClaudeClaw's WhatsApp channel using OpenAI's Whisper API. When a voice note arrives, it is downloaded, transcribed, and delivered to the agent as [Voice: <transcript>].

Phase 1: Pre-flight

Check if already applied

Check if src/transcription.ts exists. If it does, skip to Phase 3 (Configure). The code changes are already in place.

Ask the user

Use AskUserQuestion to collect information:

AskUserQuestion: Do you have an OpenAI API key for Whisper transcription?

If yes, collect it now. If no, direct them to create one at https://platform.openai.com/api-keys.

Phase 2: Apply Code Changes

Prerequisite: WhatsApp must be installed first (skill/whatsapp merged). This skill modifies WhatsApp channel files.

Ensure WhatsApp fork remote
bash
git remote -v

If whatsapp is missing, add it:

bash
git remote add whatsapp https://github.com/qwibitai/claudeclaw-whatsapp.git
Merge the skill branch
bash
git fetch whatsapp skill/voice-transcription
git merge whatsapp/skill/voice-transcription || {
  git checkout --theirs package-lock.json
  git add package-lock.json
  git merge --continue
}

This merges in:

  • src/transcription.ts (voice transcription module using OpenAI Whisper)
  • Voice handling in src/channels/whatsapp.ts (isVoiceMessage check, transcribeAudioMessage call)
  • Transcription tests in src/channels/whatsapp.test.ts
  • openai npm dependency in package.json
  • OPENAI_API_KEY in .env.example

If the merge reports conflicts, resolve them by reading the conflicted files and understanding the intent of both sides.

Validate code changes
bash
npm install --legacy-peer-deps
npm run build
npx vitest run src/channels/whatsapp.test.ts

All tests must pass and build must be clean before proceeding.

Phase 3: Configure

Get OpenAI API key (if needed)

If the user doesn't have an API key:

I need you to create an OpenAI API key:

  1. Go to https://platform.openai.com/api-keys
  2. Click "Create new secret key"
  3. Give it a name (e.g., "ClaudeClaw Transcription")
  4. Copy the key (starts with sk-)

Cost: $0.006 per minute of audio ($0.003 per typical 30-second voice note)

Wait for the user to provide the key.

Add to environment

Add to .env:

bash
OPENAI_API_KEY=<their-key>

Sync to container environment:

bash
mkdir -p data/env && cp .env data/env/env

The container reads environment from data/env/env, not .env directly.

Service name: Derived from the directory name: com.claudeclaw.<dirname> (macOS) / claudeclaw-<dirname> (Linux). For example, if cwd is my-assistant, the service is com.claudeclaw.my-assistant. Determine the correct service name before running service commands below.

Show full SKILL.md (173 more words)Show less
Build and restart
bash
npm run build
launchctl kickstart -k gui/$(id -u)/com.claudeclaw  # macOS
# Linux: systemctl --user restart claudeclaw

Phase 4: Verify

Test with a voice note

Tell the user:

Send a voice note in any registered WhatsApp chat. The agent should receive it as [Voice: <transcript>] and respond to its content.

Check logs if needed
bash
tail -f logs/claudeclaw.log | grep -i voice

Look for:

  • Transcribed voice message — successful transcription with character count
  • OPENAI_API_KEY not set — key missing from .env
  • OpenAI transcription failed — API error (check key validity, billing)
  • Failed to download audio message — media download issue

Troubleshooting

Voice notes show "[Voice Message - transcription unavailable]"
  1. Check OPENAI_API_KEY is set in .env AND synced to data/env/env
  2. Verify key works: curl -s https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY" | head -c 200
  3. Check OpenAI billing — Whisper requires a funded account
Voice notes show "[Voice Message - transcription failed]"

Check logs for the specific error. Common causes:

Agent doesn't respond to voice notes

Verify the chat is registered and the agent is running. Voice transcription only runs for registered groups.

© sbusso, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/add-voice-transcription of sbusso/claudeclaw.

Open the folder on GitHubat commit 1395af4

Compare with similar skills

Add Voice Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Voice Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Voice Transcription this skillsbusso/claudeclaw194—~1.1kAutomated safety check: NotesMIT
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Openai Whisper APIopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
Keirouter Sttmydisha/keirouter147—~680Automated safety check: PassMIT
Openai Whispercoco-research/coco513—~964Automated safety check: PassCustom licence
Openai Whisper APItrpc-group/trpc-agent-go1.9k12 repos~288Automated safety check: PassApache-2.0

Similar skills

  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 1 repo~518 tokens
    Media & CreativeAuto-check passed
  • Keirouter Stt

    mydisha/keirouter

    Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.

    147 GitHub stars~680 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Openai Whisper

    coco-research/coco

    Speech-to-text transcription via OpenAI Whisper. An agent skill from coco-research/coco.

    513 GitHub stars~964 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes

More from sbusso/claudeclaw

All 23 skills in this repo
  • Debug

    sbusso/claudeclaw

    Debug container agent issues. An agent skill from sbusso/claudeclaw.

    194 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes
  • X Integration

    sbusso/claudeclaw

    X (Twitter) integration for ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Gmail

    sbusso/claudeclaw

    Add Gmail integration to ClaudeClaw. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Qmd

    sbusso/claudeclaw

    Add QMD (Query Markup Documents) as an advanced memory search backend.

    194 GitHub stars~629 tokensUpdated 1 mo ago
    Auto-check passed
  • Add Telegram

    sbusso/claudeclaw

    Add Telegram as a channel. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes
  • Add Telegram Swarm

    sbusso/claudeclaw

    Add Agent Swarm (Teams) support to Telegram. An agent skill from sbusso/claudeclaw.

    194 GitHub stars~3.7k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Add Voice Transcription

What does Add Voice Transcription do?

Add voice message transcription to ClaudeClaw using OpenAI's Whisper API. Add Voice Transcription is an agent skill from sbusso/claudeclaw. Add voice message transcription to ClaudeClaw using OpenAI's Whisper API.

When should I use Add Voice Transcription?

Add Voice Transcription fits situations like: tasks that involve Transcription.

How do I install Add Voice Transcription in Claude Code?

Run `npx skills add sbusso/claudeclaw --skill add-voice-transcription -a claude-code`. Or copy the skill folder (skills/add-voice-transcription in sbusso/claudeclaw) into .claude/skills/add-voice-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Add Voice Transcription in Codex?

Run `npx skills add sbusso/claudeclaw --skill add-voice-transcription -a codex`. Or copy the skill folder (skills/add-voice-transcription in sbusso/claudeclaw) into .agents/skills/add-voice-transcription in your project. Codex loads it when a task matches its description.

Can I use Add Voice Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sbusso/claudeclaw --skill add-voice-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-voice-transcription, .gemini/skills/add-voice-transcription, .github/skills/add-voice-transcription and .opencode/skills/add-voice-transcription in your project.

What does Add Voice Transcription need to run?

Going by SKILL.md and its folder, Add Voice Transcription needs the command-line tools its instructions call (git, npm, npx and curl) and credentials named OPENAI_API_KEY. Our summary lists: Node.js; A credential in OPENAI_API_KEY.

Does Add Voice Transcription access the network?

SKILL.md names 3 domains. In commands or code: api.openai.com and github.com; the agent is likely to contact these when it follows the instructions. As links in the text: platform.openai.com. This is read from the text; nothing was executed.

Is Add Voice Transcription safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Add Voice Transcription use?

Add Voice Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Voice Transcription use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Voice Transcription?

Skills that share tags, products or a category with Add Voice Transcription: 9Router Speech-to-Text (decolua/9router, 31k stars), Openai Whisper API (openclaw/openclaw, 392k stars), Keirouter Stt (mydisha/keirouter, 147 stars) and Openai Whisper (coco-research/coco, 513 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Voice Transcription?

sbusso (a GitHub user) maintains it in sbusso/claudeclaw, which has 194 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 12, 2026.

Source: sbusso/claudeclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.