Agent skill

Audio Transcription

by mitsuhiko in mitsuhiko/agent-stuff

Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

Apache-2.0Auto-check passedMedia & Creative

Install Audio Transcription

skills CLI
$ npx skills add mitsuhiko/agent-stuff --skill audio-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mitsuhiko/agent-stuff audio-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mitsuhiko/agent-stuff.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audio-transcription .claude/skills/audio-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-transcription
GitHub stars
3.2k
Token cost
~1k tokens
SKILL.md length
318 words
Files
3
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

  • Works in 5 steps: Preserve temporary inputs immediately.… → Use cached local models, not cloud APIs.… → Force language when known. For Armin's… → …
  • Tasks that involve Transcription
  • SKILL.md covers Core rules, Fast path, Cached models and Manual command template, plus 1 more section
  • Runs Python scripts from its folder; calls uvx

What it does

Audio Transcription is an agent skill from mitsuhiko/agent-stuff. Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `precache-models.py` and `transcribe-audio.py`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. The repository describes itself as: These are commands I use with agents, mostly Claude. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/audio-transcription”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Preserve temporary inputs immediately. Voice Memo share-sheet paths under ~/Library/Containers/com.apple.VoiceMemos/Data/tmp/.com.apple.uik…
  2. Use cached local models, not cloud APIs. Prefer MLX Whisper via uvx --from mlx-whisper mlx_whisper; Hugging Face models must be cached in…
  3. Force language when known. For Armin's own dictations this is usually English with an Austrian/German accent, even when the filename is…
  4. For bad audio, run a hallucination-resistant pass. Use --condition-on-previous-text False, --word-timestamps True, and…
  5. Deliver a cleaned best-effort transcript. Compare model output with timestamps/JSON, remove obvious Whisper loops, and mark uncertain…

What it can do on your machine

Read from SKILL.md and the folder at commit 0865c84. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uvx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio Transcription loads about 1k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 318 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mitsuhiko/agent-stuff at commit 0865c84, republished under its Apache-2.0 licence (© mitsuhiko). 318 words, ~1,004 tokens.

Download SKILL.mdSave it as .claude/skills/audio-transcription/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
audio-transcription
description
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

Use this skill whenever the user asks to transcribe an audio/video file, a Voice Memos export, dictation, lecture, meeting recording, or "bad audio".

Core rules

  1. Preserve temporary inputs immediately. Voice Memo share-sheet paths under ~/Library/Containers/com.apple.VoiceMemos/Data/tmp/.com.apple.uikit.itemprovider... can disappear. Before probing or experimenting, copy the file to stable /private/tmp/audio-transcription-inputs/.
  2. Use cached local models, not cloud APIs. Prefer MLX Whisper via uvx --from mlx-whisper mlx_whisper; Hugging Face models must be cached in ~/.cache/huggingface/hub/.
  3. Force language when known. For Armin's own dictations this is usually English with an Austrian/German accent, even when the filename is German. Do not infer language from filename alone.
  4. For bad audio, run a hallucination-resistant pass. Use --condition-on-previous-text False, --word-timestamps True, and --hallucination-silence-threshold 2.
  5. Deliver a cleaned best-effort transcript. Compare model output with timestamps/JSON, remove obvious Whisper loops, and mark uncertain spans as [unclear] rather than inventing words.

Fast path

Run from this skill directory:

bash
cd /Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription
./transcribe-audio.py "/path/to/audio.m4a" --language en --quality balanced

The script:

  • stages a stable copy of the input under /private/tmp/audio-transcription-inputs/
  • ensures the selected model is cached (downloads only if missing)
  • writes txt, srt, vtt, tsv, and json to /private/tmp/audio-transcriptions/<name>-<timestamp>/
  • detects obvious hallucination loops and, in balanced mode, reruns with the full model if needed

Useful variants:

bash
# Quick draft, fastest cached model
./transcribe-audio.py audio.m4a --language en --quality fast

# Bad/important audio, slower full model
./transcribe-audio.py audio.m4a --language en --quality best \
  --prompt "Armin Ronacher dictating about AI, data centers, Vienna, Donauinsel, shareholder value."

# Auto language detection when language is genuinely unknown
./transcribe-audio.py audio.m4a --language auto --quality balanced

Cached models

Default model IDs:

  • Fast/balanced: mlx-community/whisper-large-v3-turbo
  • Best fallback: mlx-community/whisper-large-v3-mlx

Pre-cache / refresh both models:

bash
cd /Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription
./precache-models.py

Verify cache manually:

bash
find ~/.cache/huggingface/hub -maxdepth 1 -type d -name 'models--mlx-community--whisper-large-v3*' -print

If a model is already cached, mlx_whisper should say Fetching 4 files: 100% almost instantly.

Manual command template

If the helper script is not suitable, use this command directly:

bash
mkdir -p /private/tmp/audio-transcriptions/manual
uvx --from mlx-whisper mlx_whisper "/stable/copy/of/audio.m4a" \
  --model mlx-community/whisper-large-v3-turbo \
  --language en \
  --condition-on-previous-text False \
  --word-timestamps True \
  --hallucination-silence-threshold 2 \
  --output-format all \
  --output-dir /private/tmp/audio-transcriptions/manual \
  --output-name transcript \
  --verbose False

For especially rough audio, replace the model with mlx-community/whisper-large-v3-mlx.

Quality checks

Inspect the generated .txt first, then the .srt/.json around suspicious areas.

Red flags that require rerun or cleanup:

  • repeated phrases for many lines (eg. in nature loops)
  • many zero-duration segments
  • avg_logprob is NaN or compression ratios are very high in JSON
  • text contradicts obvious context words supplied in the prompt

When finalizing, lightly punctuate and paragraph the transcript, but do not over-edit uncertain content.

© mitsuhiko, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/audio-transcription of mitsuhiko/agent-stuff.

  • SKILL.md
  • precache-models.py
  • transcribe-audio.py

Open the folder on GitHubat commit 0865c84

Compare with similar skills

Audio Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Transcription this skillmitsuhiko/agent-stuff3.2k—~1kAutomated safety check: PassApache-2.0
Speech Recognitiondpearson2699/swift-ios-skills1.2k—~3.7kAutomated safety check: PassCustom licence
Stepfun Asrdaymade/claude-code-skills1.4k—~3kAutomated safety check: PassMIT
Sttmikeyobrien/rho373—~149Automated safety check: PassMIT
Groq Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMIT
WhisperAlexAI-MCP/hermes-CCC135—~1.9kAutomated safety check: PassMIT

Similar skills

  • Speech Recognition

    dpearson2699/swift-ios-skills

    Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills.

    1.2k GitHub stars~3.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Stepfun Asr

    daymade/claude-code-skills

    Transcribes Chinese/English audio with StepFun's stepaudio-3-asr-max via its SSE endpoint (not /v1/audio/transcriptions) — one call handles long-form audio with no chunking.

    1.4k GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Stt

    mikeyobrien/rho

    Speech-to-text — transcribe voice to text using device microphone.

    373 GitHub stars~149 tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Groq Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses when you need Groq's non-chat endpoints — transcribing or translating audio with Whisper, understanding images with Llama 4 vision, generating speech (TTS), or benchmarking…

    2.8k GitHub stars~1.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Whisper

    AlexAI-MCP/hermes-CCC

    OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed
  • Muapi AI Clipping

    hashgraph-online/awesome-codex-plugins

    Turn a long video into N viral-ready short clips with a single managed API call.

    1.3k GitHub stars~1.7k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from mitsuhiko/agent-stuff

All 16 skills in this repo
  • Anachb

    mitsuhiko/agent-stuff

    Austrian public transport (VOR AnachB) for all of Austria. An agent skill from mitsuhiko/agent-stuff.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Librarian

    mitsuhiko/agent-stuff

    Cache and refresh remote git repositories under ~/.cache/checkouts/<host/<org/<repo so future references can reuse a local copy.

    3.2k GitHub starsUsed in 1 repo~519 tokens
    Auto-check passed
  • Web Browser

    mitsuhiko/agent-stuff

    Automate and interact with web pages through Chrome or Chromium using the Chrome DevTools Protocol (CDP): navigate, click, fill forms, inspect content, take screenshots, and debug console or network…

    3.2k GitHub stars~1.5k tokensUpdated 13 days ago
    Auto-check: warnings
  • Google Workspace

    mitsuhiko/agent-stuff

    Access Google Workspace APIs (Drive, Docs, Calendar, Gmail, Sheets, Slides, Chat, People) via local helper scripts without MCP.

    3.2k GitHub stars~919 tokensUpdated 13 days ago
    Auto-check passed
  • Oebb Scotty

    mitsuhiko/agent-stuff

    Austrian rail travel planner (ÖBB Scotty). An agent skill from mitsuhiko/agent-stuff.

    3.2k GitHub stars~2.4k tokensUpdated 13 days ago
    Auto-check passed
  • Sentry

    mitsuhiko/agent-stuff

    Fetch and analyze Sentry issues, events, transactions, and logs.

    3.2k GitHub stars~1.8k tokensUpdated 13 days ago
    Auto-check passed

Questions about Audio Transcription

What does Audio Transcription do?

Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio. Audio Transcription is an agent skill from mitsuhiko/agent-stuff. Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

When should I use Audio Transcription?

Audio Transcription fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Audio Transcription in Claude Code?

Run `npx skills add mitsuhiko/agent-stuff --skill audio-transcription -a claude-code`. Or copy the skill folder (skills/audio-transcription in mitsuhiko/agent-stuff) into .claude/skills/audio-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Audio Transcription in Codex?

Run `npx skills add mitsuhiko/agent-stuff --skill audio-transcription -a codex`. Or copy the skill folder (skills/audio-transcription in mitsuhiko/agent-stuff) into .agents/skills/audio-transcription in your project. Codex loads it when a task matches its description.

Can I use Audio Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mitsuhiko/agent-stuff --skill audio-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-transcription, .gemini/skills/audio-transcription, .github/skills/audio-transcription and .opencode/skills/audio-transcription in your project.

What does Audio Transcription need to run?

Going by SKILL.md and its folder, Audio Transcription needs Python for the scripts in its folder and the command-line tools its instructions call (uvx). Our summary lists: Python 3.

Does Audio Transcription access the network?

SKILL.md contains no URLs. Its commands use uvx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Audio Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio Transcription use?

Audio Transcription is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Transcription use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio Transcription?

Skills that share tags, products or a category with Audio Transcription: Speech Recognition (dpearson2699/swift-ios-skills, 1.2k stars), Stepfun Asr (daymade/claude-code-skills, 1.4k stars), Stt (mikeyobrien/rho, 373 stars) and Groq Core Workflow B (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Transcription?

mitsuhiko (a GitHub user) maintains it in mitsuhiko/agent-stuff, which has 3,181 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 27, 2026.

Source: mitsuhiko/agent-stuff on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.