Whisper Speech Recognition
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install chujianyun/skills local-audio-transcriber --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/local-audio-transcriber .claude/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .claude/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriberType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install chujianyun/skills local-audio-transcriber --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/local-audio-transcriber .agents/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .agents/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install chujianyun/skills local-audio-transcriber --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/local-audio-transcriber .cursor/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .cursor/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/chujianyun/skills.git --path skills/local-audio-transcriber--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install chujianyun/skills local-audio-transcriber --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/local-audio-transcriber .gemini/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .gemini/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install chujianyun/skills local-audio-transcriberInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/local-audio-transcriber .github/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .github/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chujianyun/skills --skill local-audio-transcriber -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install chujianyun/skills local-audio-transcriber --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chujianyun/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/local-audio-transcriber .opencode/skills/local-audio-transcriber && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "local-audio-transcriber" agent skill from https://github.com/chujianyun/skills/tree/main/skills/local-audio-transcriber into .opencode/skills/local-audio-transcriber/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "local-audio-transcriber", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
local-audio-transcriber本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。
Local Audio Transcriber is an agent skill from chujianyun/skills. 本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `agents/openai.yaml` and `scripts/transcribe.py`).
It sits in AI & LLM Engineering, covering Speech recognition and synthesis. It works with Whisper and Python. The repository describes itself as: WuMing's Claude Skills.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 13b27aa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Local Audio Transcriber loads about 1k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 178 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 178 words (~1,005 tokens).
SKILL.md and 2 other files (scripts) in skills/local-audio-transcriber of chujianyun/skills.
Open the folder on GitHubat commit 13b27aa
Local Audio Transcriber next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Local Audio Transcriber this skillchujianyun/skills | 742 | — | ~1k | Automated safety check: Pass | Custom licence | |
| Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.9k | Automated safety check: Notes | MIT | |
| Deepgram Audio Intelligence for Pythondeepgram/deepgram-python-sdk | 469 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Agentstadaspetra/loop | 296 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Deepgram Flux Conversational STTdeepgram/deepgram-python-sdk | 469 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Local Asrysyecust/lecture-to-notes | 273 | — | ~1.6k | Automated safety check: Pass | Custom licence |
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
deepgram/deepgram-python-sdk
Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.
tadaspetra/loop
Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.
deepgram/deepgram-python-sdk
Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.
ysyecust/lecture-to-notes
把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell…
mingchen666/Reviva
Turn Bilibili videos that are already registered and parsed in MindSpace, or Bilibili opus/article posts, into evidence-linked Markdown learning notes.
chujianyun/skills
将公开或用户有权访问的 Wiki 完整转换为面向 Agent 的离线知识 Skill,并按原 Wiki 层级保存 Markdown、生成检索索引和逐文档哈希清单,支持无变化不落盘的手动或自动增量更新。当用户要求把 Wiki、文档站或帮助中心做成 Skill、同步已有 Wiki Skill、保持文档目录树或设置 Wiki Skill 自动更新时使用;不用于只摘要单篇网页或绕过登录、付费墙和访问控制。
chujianyun/skills
为照片、身份证、护照、学位证、毕业证、资格证、营业执照等证件或证书扫描件及 PDF 添加本地文字水印。用户提出照片加版权水印、身份证或学历证件添加“仅限某用途”水印、资质文件批量加水印、生成水印预览,或需要保持头像、二维码、印章等区域可辨认时使用。支持 JPEG、PNG、WebP 和 PDF,默认先预览、保留原件并清除图片元数据。不用于去除水印、伪造或篡改证件内容,也不用于仅设计 Logo…
chujianyun/skills
GitHub 源码解读助手。适用于用户提供 GitHub 仓库链接,并希望解读源码、理解原理、分析架构、生成学习报告或快速上手文档时使用。会在 working 目录下生成源码解读和快速上手两份文档。默认先交付初稿,不自动复查;如果用户明确同意,再安排后续复查。不适用于仅克隆仓库或只要一句简介的场景。
chujianyun/skills
LlamaIndex 官方用户文档离线知识库,用于检索并回答 LlamaIndex Python 框架的安装、RAG、数据加载、索引、检索与查询、Agent、Workflow、模型、Embedding、向量库、评估、可观测性、部署、LlamaCloud 和 LlamaParse 等问题,也可生成有文档依据的示例代码与排障建议。当用户提到…
chujianyun/skills
论文解读助手。适用于用户发送 arXiv 论文链接,并希望下载论文、解读论文、生成读书笔记、做论文拆解或输出详细报告时使用。会在工作目录创建论文文件夹、下载 PDF 与 TeX Source(如有)、生成中文 Markdown 报告。默认先交付初稿,不自动复查;如果用户明确同意,再安排后续复查。不适用于只要简短推荐语的情况。
chujianyun/skills
千问AI平台(Qianwen AI Platform / DashScope)官方文档离线知识库,用于检索并回答模型选择、API Key、OpenAI 兼容接口、DashScope SDK、文本与多模态生成、图像/视频/语音、Realtime API、Embedding、Reranking、Function Calling、MCP、批量调用、计费、Token Plan、API/SDK/CLI…
Categories
本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。. Local Audio Transcriber is an agent skill from chujianyun/skills.
Local Audio Transcriber fits situations like: tasks that involve Speech recognition and synthesis.
Run `npx skills add chujianyun/skills --skill local-audio-transcriber -a claude-code`. Or copy the skill folder (skills/local-audio-transcriber in chujianyun/skills) into .claude/skills/local-audio-transcriber in your project. Claude Code loads it when a task matches its description.
Run `npx skills add chujianyun/skills --skill local-audio-transcriber -a codex`. Or copy the skill folder (skills/local-audio-transcriber in chujianyun/skills) into .agents/skills/local-audio-transcriber in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chujianyun/skills --skill local-audio-transcriber -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/local-audio-transcriber, .gemini/skills/local-audio-transcriber, .github/skills/local-audio-transcriber and .opencode/skills/local-audio-transcriber in your project.
Going by SKILL.md and its folder, Local Audio Transcriber needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Local Audio Transcriber has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.
About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Local Audio Transcriber: Whisper Speech Recognition (Orchestra-Research/AI-Research-SKILLs, 13k stars), Deepgram Audio Intelligence for Python (deepgram/deepgram-python-sdk, 469 stars), Agents (tadaspetra/loop, 296 stars) and Deepgram Flux Conversational STT (deepgram/deepgram-python-sdk, 469 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
chujianyun (a GitHub user) maintains it in chujianyun/skills, which has 742 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.
Source: chujianyun/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.