Agent skill

Local Audio Transcriber

by chujianyun in chujianyun/skills

本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。

Custom licenceAuto-check passedAI & LLM Engineering

Install Local Audio Transcriber

skills CLI
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chujianyun/skills local-audio-transcriber --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/local-audio-transcriber .claude/skills/local-audio-transcriber && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
local-audio-transcriber
GitHub stars
742
Token cost
~1k tokens
SKILL.md length
178 words
Files
3 (incl. scripts)
Skills in repo
35
Repo updated
First seen
Licence
Custom licence

At a glance

本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。

  • Works in 7 steps: 确认用户提供的是可访问的本地音频/视频文件路径或附件。 → 先判断机器类型和可用引擎 → Apple Silicon(M1/M2/M3/M4)优先安装并使用… → …
  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers 适用边界, 工作流程, 常用命令 and 参数取舍, plus 2 more sections
  • Runs Python scripts from its folder; calls python3 and python

What it does

Local Audio Transcriber is an agent skill from chujianyun/skills. 本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `agents/openai.yaml` and `scripts/transcribe.py`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis. It works with Whisper and Python. The repository describes itself as: WuMing's Claude Skills.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/local-audio-transcriber”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. 确认用户提供的是可访问的本地音频/视频文件路径或附件。
  2. 先判断机器类型和可用引擎
  3. Apple Silicon(M1/M2/M3/M4)优先安装并使用 MLX,本地调用 Apple GPU/统一内存
  4. 非 Apple Silicon、CUDA 机器或 MLX 不可用时,再使用 faster-whisper
  5. 运行转写脚本,中文录音优先指定 --language zh;不确定语言时省略语言参数
  6. 向用户直接发送转写文本。文本很长时,优先交付本地文件路径,并贴出开头和关键说明。
  7. 默认只保存 .md 和 .srt 文件;不要生成 txt、json 或 vtt。

What it can do on your machine

Read from SKILL.md and the folder at commit 13b27aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Local Audio Transcriber loads about 1k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 178 words (~1,005 tokens).

name
local-audio-transcriber

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (scripts) in skills/local-audio-transcriber of chujianyun/skills.

  • SKILL.md
  • agents/openai.yaml
  • scripts/transcribe.py

Open the folder on GitHubat commit 13b27aa

Compare with similar skills

Local Audio Transcriber next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Local Audio Transcriber compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Local Audio Transcriber this skillchujianyun/skills742—~1kAutomated safety check: PassCustom licence
Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs13k7 repos~1.9kAutomated safety check: NotesMIT
Deepgram Audio Intelligence for Pythondeepgram/deepgram-python-sdk469—~2.3kAutomated safety check: PassMIT
Agentstadaspetra/loop2961 repos~2.5kAutomated safety check: PassMIT
Deepgram Flux Conversational STTdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Local Asrysyecust/lecture-to-notes273—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Deepgram Audio Intelligence for Python

    deepgram/deepgram-python-sdk

    Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.

    469 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Deepgram Flux Conversational STT

    deepgram/deepgram-python-sdk

    Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.

    469 GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Local Asr

    ysyecust/lecture-to-notes

    把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell…

    273 GitHub stars~1.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Bili Note

    mingchen666/Reviva

    Turn Bilibili videos that are already registered and parsed in MindSpace, or Bilibili opus/article posts, into evidence-linked Markdown learning notes.

    244 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from chujianyun/skills

All 35 skills in this repo
  • Baize

    chujianyun/skills

    将公开或用户有权访问的 Wiki 完整转换为面向 Agent 的离线知识 Skill,并按原 Wiki 层级保存 Markdown、生成检索索引和逐文档哈希清单,支持无变化不落盘的手动或自动增量更新。当用户要求把 Wiki、文档站或帮助中心做成 Skill、同步已有 Wiki Skill、保持文档目录树或设置 Wiki Skill 自动更新时使用;不用于只摘要单篇网页或绕过登录、付费墙和访问控制。

    742 GitHub stars~905 tokensUpdated yesterday
    Auto-check passed
  • 为照片、身份证、护照、学位证、毕业证、资格证、营业执照等证件或证书扫描件及 PDF 添加本地文字水印。用户提出照片加版权水印、身份证或学历证件添加“仅限某用途”水印、资质文件批量加水印、生成水印预览,或需要保持头像、二维码、印章等区域可辨认时使用。支持 JPEG、PNG、WebP 和 PDF,默认先预览、保留原件并清除图片元数据。不用于去除水印、伪造或篡改证件内容,也不用于仅设计 Logo…

    742 GitHub stars~820 tokensUpdated yesterday
    Auto-check passed
  • GitHub Code Interpreter

    chujianyun/skills

    GitHub 源码解读助手。适用于用户提供 GitHub 仓库链接,并希望解读源码、理解原理、分析架构、生成学习报告或快速上手文档时使用。会在 working 目录下生成源码解读和快速上手两份文档。默认先交付初稿,不自动复查;如果用户明确同意,再安排后续复查。不适用于仅克隆仓库或只要一句简介的场景。

    742 GitHub stars~611 tokensUpdated yesterday
    Auto-check passed
  • Llama Index Wiki

    chujianyun/skills

    LlamaIndex 官方用户文档离线知识库,用于检索并回答 LlamaIndex Python 框架的安装、RAG、数据加载、索引、检索与查询、Agent、Workflow、模型、Embedding、向量库、评估、可观测性、部署、LlamaCloud 和 LlamaParse 等问题,也可生成有文档依据的示例代码与排障建议。当用户提到…

    742 GitHub stars~480 tokensUpdated yesterday
    Auto-check passed
  • Paper Interpreter

    chujianyun/skills

    论文解读助手。适用于用户发送 arXiv 论文链接,并希望下载论文、解读论文、生成读书笔记、做论文拆解或输出详细报告时使用。会在工作目录创建论文文件夹、下载 PDF 与 TeX Source(如有)、生成中文 Markdown 报告。默认先交付初稿,不自动复查;如果用户明确同意,再安排后续复查。不适用于只要简短推荐语的情况。

    742 GitHub stars~810 tokensUpdated yesterday
    Auto-check passed
  • Qianwenai Wiki

    chujianyun/skills

    千问AI平台(Qianwen AI Platform / DashScope)官方文档离线知识库,用于检索并回答模型选择、API Key、OpenAI 兼容接口、DashScope SDK、文本与多模态生成、图像/视频/语音、Realtime API、Embedding、Reranking、Function Calling、MCP、批量调用、计费、Token Plan、API/SDK/CLI…

    742 GitHub stars~717 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Local Audio Transcriber

What does Local Audio Transcriber do?

本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。. Local Audio Transcriber is an agent skill from chujianyun/skills.

When should I use Local Audio Transcriber?

Local Audio Transcriber fits situations like: tasks that involve Speech recognition and synthesis.

How do I install Local Audio Transcriber in Claude Code?

Run `npx skills add chujianyun/skills --skill local-audio-transcriber -a claude-code`. Or copy the skill folder (skills/local-audio-transcriber in chujianyun/skills) into .claude/skills/local-audio-transcriber in your project. Claude Code loads it when a task matches its description.

How do I install Local Audio Transcriber in Codex?

Run `npx skills add chujianyun/skills --skill local-audio-transcriber -a codex`. Or copy the skill folder (skills/local-audio-transcriber in chujianyun/skills) into .agents/skills/local-audio-transcriber in your project. Codex loads it when a task matches its description.

Can I use Local Audio Transcriber in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chujianyun/skills --skill local-audio-transcriber -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/local-audio-transcriber, .gemini/skills/local-audio-transcriber, .github/skills/local-audio-transcriber and .opencode/skills/local-audio-transcriber in your project.

What does Local Audio Transcriber need to run?

Going by SKILL.md and its folder, Local Audio Transcriber needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Local Audio Transcriber access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Local Audio Transcriber safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Local Audio Transcriber use?

Local Audio Transcriber has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Local Audio Transcriber use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Local Audio Transcriber?

Skills that share tags, products or a category with Local Audio Transcriber: Whisper Speech Recognition (Orchestra-Research/AI-Research-SKILLs, 13k stars), Deepgram Audio Intelligence for Python (deepgram/deepgram-python-sdk, 469 stars), Agents (tadaspetra/loop, 296 stars) and Deepgram Flux Conversational STT (deepgram/deepgram-python-sdk, 469 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Local Audio Transcriber?

chujianyun (a GitHub user) maintains it in chujianyun/skills, which has 742 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.

Source: chujianyun/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.