Agent skill

Murmur

by xiaopengde in xiaopengde/murmur

把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用…

MITAuto-check: notesAI & LLM Engineering

Install Murmur

skills CLI
$ npx skills add xiaopengde/murmur --skill murmur -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xiaopengde/murmur murmur --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
murmur
GitHub stars
109
Token cost
~2.9k tokens
SKILL.md length
766 words
Files
27 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用…

  • Works in 7 steps: 触发条件 → 端到端 5 步流程 → 目录约定 → …
  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers 0. 触发条件, 1. 端到端 5 步流程, 2. 目录约定 and 3. 各 agent 平台的小差异, plus 3 more sections
  • Runs Python, Shell and PowerShell scripts from its folder; calls python, bash and uvx; reaches pypi.tuna.tsinghua.edu.cn and hf-mirror.com

What it does

Murmur is an agent skill from xiaopengde/murmur. 把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用 whisper-ctranslate2)。零云端、零订阅、隐私不出本机。适用:替代飞书妙计/通义听悟/Otter.ai 这类付费转录服务、需要在 VS Code 或 Word 里直接拿到可读稿、在 AI agent(Claude Code/Copilot/Codex/Cursor)里端到端跑通。不适用:实时转录、强噪声多人重叠会议、需要严格说话人分离的场景。

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 30 other files, including scripts (for example `README.md`, `docs/install-mac.md` and `docs/install-windows.md`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Word documents. It works with Microsoft Word, Linux, Visual Studio Code and macOS. The repository describes itself as: 🎙 Local zero-cost audio→markdown/docx pipeline. Whisper + AI cleanup. Replaces Otter.ai/飞书妙计. Agent-ready (Claude Code/Codex/Copilot/Cursor). macOS/Windows/Linux. The licence is MIT.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Word documents

Example prompts

  • “/murmur”

Requirements

  • Python 3
  • A Bash shell
  • PowerShell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 触发条件
  2. 端到端 5 步流程
  3. 目录约定
  4. 各 agent 平台的小差异
  5. 故障排查速查
  6. 一次完整跑通的最简命令序列
  7. 这个 skill 的设计哲学

What it can do on your machine

Read from SKILL.md and the folder at commit eef63ab. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python, Shell and PowerShell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • bash
    • uvx
    • ffmpeg
    • brew
    • winget

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.tuna.tsinghua.edu.cn
    • hf-mirror.com

    Also links to:

    • agentskills.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Murmur loads about 2.9k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 766 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:271
    - 如果用户在 macOS 且首次安装,可能需要 sudo 提示

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from xiaopengde/murmur at commit eef63ab, republished under its MIT licence (© xiaopengde). 766 words, ~2,940 tokens.

Download SKILL.mdSave it as .claude/skills/murmur/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.
name
murmur
description
把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用 whisper-ctranslate2)。零云端、零订阅、隐私不出本机。适用:替代飞书妙计/通义听悟/Otter.ai 这类付费转录服务、需要在 VS Code 或 Word 里直接拿到可读稿、在 AI agent(Claude Code/Copilot/Codex/Cursor)里端到端跑通。不适用:实时转录、强噪声多人重叠会议、需要严格说话人分离的场景。

Murmur — 本地零成本音频转录与清洗工作流

仓库:https://github.com/xiaopengde/murmur 适用 agent:GitHub Copilot Agent / Claude Code / OpenAI Codex / Cursor / 任意遵循 agentskills.io 的 agent


0. 触发条件

满足以下任一触发该 skill:

  • 用户说"帮我转录 / 把这段录音转成文字 / 出逐字稿 / 出会议纪要 / 飞书妙计太贵"
  • 工作目录里出现 .m4a / .mp3 / .wav / .mp4 / .webm / .flac / .ogg 文件
  • 用户明确说"用 Murmur"

1. 端到端 5 步流程

步骤 A — 环境检查(必跑,不要跳)

假设用户是新机器,刚装完 Claude Code / Codex / Copilot,没装其他任何东西。所以先检查环境:

bash
# macOS / Linux
bash scripts/doctor.sh

# Windows (PowerShell)
powershell -ExecutionPolicy Bypass -File scripts/doctor.ps1

doctor 脚本会分两块输出:

  1. 核心依赖状态:ffmpeg / uvx / pandoc / python3 / 平台和芯片 / 模型缓存(ModelScope / HuggingFace)
  2. Murmur onboarding 状态:默认输出格式、默认离线模型是否已经由用户明确选择

如果核心依赖有 ❌,跑对应的 install 脚本:

bash
# macOS
bash scripts/install-mac.sh

# Windows (需要管理员 PowerShell)
powershell -ExecutionPolicy Bypass -File scripts/install-windows.ps1

安装完再跑一次 doctor。如果要在自动化里强校验,使用 strict 模式:

bash
bash scripts/doctor.sh --strict
powershell -ExecutionPolicy Bypass -File scripts/doctor.ps1 -Strict

--strict / -Strict 只有在核心依赖齐全且 onboarding 已完成时才返回 0;依赖缺失或 onboarding 未完成都会返回非 0。doctor 如果提示 onboarding 未完成,不要说“可以开始转录”,下一步必须跑 python scripts/transcribe.py --onboarding。

端到端验证(首次安装强烈推荐,或排查问题时;会使用临时配置,不污染用户默认值):

bash
bash scripts/doctor.sh --smoke                                    # macOS / Linux
powershell -ExecutionPolicy Bypass -File scripts/doctor.ps1 -Smoke  # Windows

会自动生成 2 秒测试音频跑完整 pipeline。大陆 Apple Silicon 会优先用 ModelScope large-v3-turbo 4bit(首次约 464MB),其他环境用 tiny 模型(首次约 75MB);通过后说明 ffmpeg → uvx → mlx/whisper → 文件输出全链路工作。失败时会保留临时目录方便排查。

步骤 B — 新用户 / 首次 onboarding(硬门禁)

无论用户是否已经给了音频文件,都必须先跑:

bash
python scripts/transcribe.py --onboarding

读取 JSON:

  • 如果 needs_onboarding=false:可以进入步骤 C。
  • 如果 needs_onboarding=true:agent 必须使用 AskQuestion / 候选框让用户选择,不能自己决定,不能直接执行 JSON 里的 example。JSON 会明确包含:
    • must_ask_user: true
    • do_not_choose_for_user: true
    • do_not_run_example_without_user_choice: true

必须向用户询问两个候选项:

  1. 默认输出格式:md 或 docx
  2. 默认离线模型:large-v3-turbo / large-v3 / medium / small

用户选完后,运行:

bash
python scripts/transcribe.py --init-defaults --format <md|docx> --set-default-model <model>

只有这个命令成功后,才允许继续转录。--format / --model 是单次覆盖参数,不能绕过首次 onboarding;未完成 onboarding 时,转录主流程会直接退出并要求先完成 onboarding。

如果用户后来想改默认:

bash
python scripts/transcribe.py --set-default md       # 或 docx
python scripts/transcribe.py --set-default-model medium
步骤 C — 跑转录并转述进度
bash
python scripts/transcribe.py <音频文件> [--lang zh] [--output-dir .] [--model medium] [--cn]

脚本内部会:

  1. 用 ffmpeg 把任意输入转成 16kHz 单声道 WAV(关键——直接喂 m4a 会触发 Whisper 幻觉循环)
  2. 检测平台:
    • Apple Silicon Mac → uvx --from mlx-whisper mlx_whisper(GPU 加速,最快)
    • Windows / Linux / Intel Mac → uvx whisper-ctranslate2(CTranslate2 后端,CPU 也很快)
  3. 输出 转录原稿.txt + 字幕.srt 到目标目录
  4. 清理临时 WAV
⛔️ 进度转述硬约束(agent 必读,最常翻车的一步)

反模式(你大概率会犯,越自信的模型越容易犯):把这条转录命令当成「一个会返回结果的调用」,用一次性、长时间的前台阻塞等待(例如 Waiting 12m for shell / 一个 timeout=30min 的同步调用)等它跑完,期间不回头看终端。你可能觉得「下载和推理中间没我什么事,等结束再汇报最高效」——这是错的。脚本每 30 秒就在终端打一行心跳(📥 模型下载/准备中... 已用时 x:xx / ⏳ 转录推理中... 已用时 x:xx),但只要你一次性阻塞到结束,这些心跳全被你埋掉,用户面对的就是 十几分钟纯静默的终端,完全不知道是在下载、在推理、还是已经卡死。说一句「我会盯着进度」不等于你真的在盯。

必须这样做(不可省略、不可优化掉):

  1. 禁止用单次「阻塞等到命令结束」的方式跑转录。要么把它放后台跑,要么用短 timeout(如每次最多等 30–45 秒)然后反复回来读终端。
  2. 每约 30 秒必须主动 poll 一次终端输出。
  3. 每次 poll 后,只要看到新的心跳或阶段变化(📥 下载/准备 → ⏳ 转录推理 → ✅ 完成),立刻用一句前台消息转述给用户(例如「模型还在下载,约 40%」「已进入转录推理,已用时 2:00」「转录完成,正在整理输出」)。
  4. 从命令开始到结束,中间不允许出现一次超过 ~45 秒的静默。如果你发现自己正打算「等它跑完再说」,停下——那就是上面的反模式。

首次没有缓存时,下载/准备完成后才算进入推理阶段;这两个阶段都要分别转述,不要把「还在下载」说成「在推理」。

🇨🇳 大陆网络:transcribe.py 会按时区/语言自动判断是否在大陆。命中后默认使用 --model-source auto:

  • 对已验证映射,优先从 ModelScope 下载模型到 ~/.cache/murmur/models/,然后把本地目录交给现有 mlx-whisper / whisper-ctranslate2 推理。
  • 对没有 ModelScope 映射的模型,继续给 whisper 子进程注入 HF_ENDPOINT=https://hf-mirror.com(HuggingFace 镜像)和 UV_INDEX_URL=https://pypi.tuna.tsinghua.edu.cn/simple(uv 拉依赖走清华)。

⚠️ 走 ModelScope 时注意 VPN:ModelScope 是国内 CDN,大陆直连最快。如果用户开了全局 VPN / 代理,流量会绕到海外再回国内,反而拖慢下载。看到走 ModelScope 时,agent 应提醒用户:「正在从国内源 ModelScope 下载,如开着全局 VPN 建议临时关掉直连更快」。(这点和走 HuggingFace 时相反——HF 路线开 VPN 才快。)

可手动指定模型源:

bash
python scripts/transcribe.py 录音.m4a --model-source modelscope  # 强制优先 ModelScope
python scripts/transcribe.py 录音.m4a --model-source hf          # 强制原 HuggingFace/引擎默认源

当前已验证映射:

  • Apple Silicon / mlx-whisper:large-v3-turbo → mlx-community/whisper-large-v3-turbo-4bit(ModelScope,约 464MB;下载后自动适配 model.safetensors → weights.safetensors)
  • Windows / Linux / Intel Mac / whisper-ctranslate2:large-v3-turbo → mobiuslabsgmbh/faster-whisper-large-v3-turbo(ModelScope,约 1.62GB,CTranslate2 格式;需要按目标平台 smoke test)

用户已经手动设过的同名环境变量不会被覆盖。显式 --cn / --no-cn 强制单次开关。

持久化偏好(避免每次都加 --cn):

bash
python scripts/transcribe.py --set-default-cn on    # 以后每次自动启用
python scripts/transcribe.py --set-default-cn off   # 以后每次走官方源
python scripts/transcribe.py --set-default-cn auto  # 恢复按时区/语言自动判断(默认)

也可以直接 bash scripts/install-mac.sh --cn(或 -CN for Windows),安装脚本会在结束时把偏好写进配置。

换更小的模型(CPU 慢机器常用):

bash
python scripts/transcribe.py 录音.m4a --model medium                   # 单次
python scripts/transcribe.py --set-default-model medium                # 永久(写入 config)
python scripts/transcribe.py --set-default-model ""                    # 清空恢复内置默认 large-v3-turbo

支持 tiny / base / small / medium / large-v2 / large-v3 / large-v3-turbo 短名,会按引擎自动映射:

  • Apple Silicon / mlx-whisper 使用显式 HuggingFace repo 映射(例如 large-v3-turbo → mlx-community/whisper-large-v3-turbo,不是 ...-turbo-mlx)
  • whisper-ctranslate2 透传短名(如 large-v3-turbo / medium)

也支持透传完整 HF repo 名给高级用户。转录失败时,脚本会打印 resolved model,并尽量区分网络问题与 repo 不存在 / 私有 / 映射错误。

预期耗时:

  • M2/M3:音频时长 × 0.3-0.5
  • Windows / Linux CPU:音频时长 × 1-2(首次会更慢,模型加载约 30s)
  • 首次跑会先下载模型到 ~/.cache/murmur/models/(ModelScope 路线)或 ~/.cache/huggingface/hub/(原 HuggingFace 路线;Win 是 %USERPROFILE%\.cache\huggingface\hub\),日志显示 📥 模型下载/准备中;缓存就绪后才显示 ⏳ 转录推理中,之后秒级冷启动

⚠️ 关键约定:脚本里已经默认关掉了 condition-on-previous-text,因为这是 No.1 大坑(不关会输出"X 点 X 点 X 点……"或"谢谢观看"成段重复)。不要修改这个默认值。

步骤 D — LLM 清洗成可读 Markdown

读 转录原稿.txt,按以下步骤执行(不要跳过任何一步):

1. 组装 prompt(两种方式选一)

  • 自动(推荐):
    bash
    python scripts/clean.py 转录原稿.txt [--scene interview|meeting|podcast]
    # 输出组装好的 prompt,复制到 LLM 对话框即可
  • 手动:打开 docs/prompts/clean-transcript.md,把 ## === PROMPT 开始 === 到 ## === PROMPT 结束 === 之间的内容原封不动复制出来(不要总结、不要省略、不要用下面的简化版代替)。然后在末尾的占位符处贴入 转录原稿.txt 的全部内容。

2. 补场景描述(如已知,加在 prompt 最前面)

  • 面试录音:这是一段中文面试录音,面试官代称"面试官",应试者代称"我"。
  • 会议录音:这是一段工作会议录音,已知参与者:[姓名/代号]。
  • 不确定则不加(LLM 会从上下文推断)

3. 发给 LLM 执行(GPT-4o / Claude / Gemini / 国内大模型均可)

  • 原稿超过 1.5 万字:先让 LLM 列出章节结构,再逐章节批次清洗,最后拼成一个文件
  • 说话人 ≥ 3 人且混乱:先听前 30 秒录音,告诉 LLM 每个人的角色/声音特征

4. 对照「输出前自检清单」核查

拿到 LLM 输出后,过一遍 docs/prompts/clean-transcript.md 末尾的「输出前自检清单」(7 项)。有不合格项让 LLM 补改后再保存。

5. 保存为 逐字稿-清洗版.md,放在和音频同目录。

6. 主动把成果递到用户眼前(不要只报路径)

清洗稿是用户唯一真正关心的交付物,落盘后必须:

  • 在编辑器里主动打开 逐字稿-清洗版.md(Cursor / VS Code:用打开文件的工具直接打开它,让用户一抬眼就看到稿子;Claude Code / Codex CLI 等无法直接开文件时,在回复里贴出开头一段做即时预览)。
  • 明确告诉用户文件的完整保存路径。
  • 不要只甩一句「文件已保存在 xxx 目录」就完事——那样用户还得自己去翻文件树,体验差一口气。
步骤 E — 如果默认是 docx,转 docx
bash
python scripts/md2docx.py 逐字稿-清洗版.md

会在同目录输出 逐字稿-清洗版.docx,用 pandoc 实现,跨平台一致。

如果用户配置了默认 docx,不要问"要不要转 docx",直接转就完事——这是设默认的意义。转好后同样主动打开/告知 docx 的完整路径,别让用户自己去找。

Show full SKILL.md (307 more words)Show less
步骤 F(可选)— 复盘纪要

仅当用户明确说"复盘 / 纪要 / 总结 / retro"时做。读 docs/prompts/retrospective.md,按里面的模板生成 复盘纪要.md。

默认只交付步骤 D 的清洗稿,不要主动给复盘——大多数人只要可读逐字稿。


2. 目录约定

一次会议/面试 = 一个子目录:

<场景名-YYYY-MM-DD>/
├── 录音.m4a                ← 用户的原始音频
├── 字幕.srt                ← transcribe.py 输出,保留
├── 转录原稿.txt            ← transcribe.py 输出,保留(清洗依据)
├── 逐字稿-清洗版.md        ← LLM 清洗产物,主要交付物
├── 逐字稿-清洗版.docx      ← 默认 docx 时附加产物
└── 复盘纪要.md             ← 可选

如果用户音频文件本来就在某个目录里,就在那个目录就地输出;不要无中生有创建子目录除非用户要求。


3. 各 agent 平台的小差异

通用红线:所有平台都必须遵守步骤 C 的「进度转述硬约束」——不要用一次性长阻塞等转录跑完。下面是各平台具体怎么做到「后台跑 / 短轮询 + 每 30s 转述」。

GitHub Copilot Agent (VS Code)
  • 不要用 mode='sync' + 一个 30min 长 timeout 一次性等完——那样心跳全被埋掉。改用后台运行(isBackground=true)或 mode='sync' 配短 timeout(30–45s)然后反复再读同一终端的输出
  • 看进度请读取该终端已有输出,不要 send_to_terminal 任何命令到正在转录的持久 zsh——会 Ctrl+C 掉进程
  • 每约 30s 读一次终端,把新心跳转述给用户
Claude Code
  • 不要用一个超长 timeout 的 Bash 调用阻塞到结束。把转录放后台(命令尾 & 或后台模式),再用 BashOutput 每约 30s 轮询一次输出并转述心跳
  • 如果用户在 macOS 且首次安装,可能需要 sudo 提示
OpenAI Codex CLI
  • 默认权限可能不够;安装步骤需要用户手动 sudo;转录命令本身不需要特权
  • 不要一次性阻塞等待。用短超时反复 poll 终端,每约 30s 转述一次进度心跳
Cursor
  • 跟 VS Code Copilot 行为一致:后台跑或短 timeout + 反复读终端,严禁单次长阻塞等待

4. 故障排查速查

症状原因解法
转录文本反复 "X 点 X 点 X 点…" 或某句话整段重复condition-on-previous-text 未关用本仓库的 transcribe.py 不会有这个问题;如果手动改过命令,加回 --condition-on-previous-text False
全程 "谢谢观看" 成段重复音频开头有静音 + 没做 ffmpeg 预处理用本仓库的 transcribe.py 自动处理;手动跑时记得先 ffmpeg -ar 16000 -ac 1
速度极慢用成了 openai-whisper PyPI 版(纯 CPU + Python)确认走的是 mlx-whisper(Mac AS)或 whisper-ctranslate2(其他)
模型下载卡住HuggingFace 网络问题优先加 --model-source modelscope 或 --cn 走 ModelScope 已验证模型;常用国内的话直接 --set-default-cn on
uvx 首次拉 mlx-whisper / whisper-ctranslate2 卡住PyPI 访问慢同样加 --cn,会同时注入 UV_INDEX_URL=https://pypi.tuna.tsinghua.edu.cn/simple
CPU 机器转录慢、显存不够模型太大换小模型:--model medium 单次,或 --set-default-model medium 永久
brew install / winget install 卡在下载国内访问 Homebrew bottle / GitHub Releases 慢重跑安装脚本时加 CN flag:<br>Mac: bash scripts/install-mac.sh --cn(启用 USTC 镜像)<br>Win: powershell -ExecutionPolicy Bypass -File scripts\install-windows.ps1 -CN(启用 Scoop/PyPI 兜底)<br>脚本默认会按时区/语言自动判断,加 flag 是强制启用
Mac 上 install-mac.sh 报 "command not found: brew"Homebrew 没装让用户先装 Homebrew(脚本会给提示)
Windows 上 install-windows.ps1 报权限错误PowerShell 没用管理员模式右键 PowerShell → 以管理员身份运行
Windows 上 winget 找不到旧版 Windows 10 没装 winget让用户从 Microsoft Store 装 "App Installer"
英文术语全错多语言混读触发了语言切换transcribe.py 默认 --language zh,对中文为主的录音最稳;纯英文用 --lang en
说话人混在一起Whisper 不带 diarization清洗阶段靠上下文推断;多人混乱场景用 **说话人 1/2/3** 占位
转录到一半进程被杀同一持久终端被 send_to_terminal 干扰 / 电脑休眠见 §3 "GitHub Copilot Agent" 注意事项

完整故障排查见 docs/troubleshooting.md。


5. 一次完整跑通的最简命令序列

假设用户给你一个 面试.m4a,目标默认 docx:

bash
# 1) 环境检查(macOS)
bash scripts/doctor.sh

# 2) 缺啥装啥
bash scripts/install-mac.sh

# 3) 转录(首次会问 md/docx,让用户回答)
python scripts/transcribe.py 面试.m4a

# 4) agent 你来读 转录原稿.txt,按 docs/prompts/clean-transcript.md
#    清洗输出 逐字稿-清洗版.md(这一步是 LLM 自己做,不调脚本)

# 5) 默认 docx → 转
python scripts/md2docx.py 逐字稿-清洗版.md

# 6) 用户说要复盘?再来一步
#    读 docs/prompts/retrospective.md,输出 复盘纪要.md

整个过程只有一次用户交互(首次 onboarding 的格式 + 模型选择),之后都是无人值守。


6. 这个 skill 的设计哲学

写给后续维护者 / fork 这个 skill 的 agent:

  1. 新机器零假设——doctor 和 install 脚本必须能在刚装完 OS 的 Mac/Win 上跑通
  2. 不替用户做关键选择——格式默认值这种长期影响的设置必须问,但只问一次
  3. LLM 清洗 prompt 和代码解耦——prompts 在 docs/prompts/,agent 直接读,方便用户改
  4. 跨平台分支必须显式——不要假设 Apple Silicon 和 Windows 走同一条命令;transcribe.py 里的平台判断是核心逻辑
  5. 错误恢复优先——任何关键命令失败都要给可执行的下一步建议,不要只报错

© xiaopengde, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 26 other files (scripts) in the repository root of xiaopengde/murmur.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • docs/install-mac.md
  • docs/install-windows.md
  • docs/prompts/clean-transcript.md
  • docs/prompts/retrospective.md
  • docs/troubleshooting.md
  • examples/sample-clean-output.md
  • scripts/clean.py
  • scripts/cn_env.py
  • scripts/config.py
  • scripts/doctor.ps1
  • scripts/doctor.sh
  • scripts/install-linux.sh
  • scripts/install-mac.sh
  • … and 10 more

Open the folder on GitHubat commit eef63ab

Compare with similar skills

Murmur next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Murmur compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Murmur this skillxiaopengde/murmur109—~2.9kAutomated safety check: NotesMIT
Tdoc DOCXLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: NotesMIT
Video EditingLeoYeAI/openclaw-master-skills2.2k—~2.7kAutomated safety check: NotesMIT
Local Asrysyecust/lecture-to-notes269—~1.6kAutomated safety check: PassCustom licence
Local RAGnigo81/nigo-skills133—~1.4kAutomated safety check: PassMIT
Vlog Auto Editznyupup/ai-video-editing-skill144—~6.8kAutomated safety check: PassMIT

Similar skills

  • Tdoc DOCX

    LeoYeAI/openclaw-master-skills

    Word 文档全能处理技能 | Complete Word Document Processing Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Video Editing

    LeoYeAI/openclaw-master-skills

    Automated video editing skill for talk/vlog/standup videos. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~2.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Local Asr

    ysyecust/lecture-to-notes

    把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell…

    269 GitHub stars~1.6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Local RAG

    nigo81/nigo-skills

    本地向量知识库,支持按项目管理文档(docx/doc/pdf/md),语义检索。默认用硅基流动免费 API,零模型安装即可使用。支持多项目隔离、中文制度文档专用切片、Embedding+Rerank 两阶段检索。触发词:知识库、向量检索、RAG、制度检索、文档入库、语义搜索、local…

    133 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vlog Auto Edit

    znyupup/ai-video-editing-skill

    AI Agent自动剪辑旅行Vlog的完整工作流。从原始素材到成品视频,系统级只需ffmpeg,其余在Python venv内完成。by nyx研究所 (GitHub @znyupup · B站/小红书 @nyx研究所)

    144 GitHub stars~6.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Docling

    zhuzhaoyun/Molio

    PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown.

    431 GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed

Questions about Murmur

What does Murmur do?

把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用…. Murmur is an agent skill from xiaopengde/murmur.

When should I use Murmur?

Murmur fits situations like: tasks that involve Speech recognition and synthesis; tasks that involve Word documents.

How do I install Murmur in Claude Code?

Run `npx skills add xiaopengde/murmur --skill murmur -a claude-code`. Or copy the skill folder (the xiaopengde/murmur repository) into .claude/skills/murmur in your project. Claude Code loads it when a task matches its description.

How do I install Murmur in Codex?

Run `npx skills add xiaopengde/murmur --skill murmur -a codex`. Or copy the skill folder (the xiaopengde/murmur repository) into .agents/skills/murmur in your project. Codex loads it when a task matches its description.

Can I use Murmur in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiaopengde/murmur --skill murmur -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/murmur, .gemini/skills/murmur, .github/skills/murmur and .opencode/skills/murmur in your project.

What does Murmur need to run?

Going by SKILL.md and its folder, Murmur needs Python, a shell and PowerShell for the scripts in its folder and the command-line tools its instructions call (python, bash, uvx, ffmpeg, brew and winget). Our summary lists: Python 3; A Bash shell; PowerShell.

Does Murmur access the network?

SKILL.md names 3 domains. In commands or code: pypi.tuna.tsinghua.edu.cn and hf-mirror.com; the agent is likely to contact these when it follows the instructions. As links in the text: agentskills.io. This is read from the text; nothing was executed.

Is Murmur safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Murmur use?

Murmur is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Murmur use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Murmur?

Skills that share tags, products or a category with Murmur: Tdoc DOCX (LeoYeAI/openclaw-master-skills, 2.2k stars), Video Editing (LeoYeAI/openclaw-master-skills, 2.2k stars), Local Asr (ysyecust/lecture-to-notes, 269 stars) and Local RAG (nigo81/nigo-skills, 133 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Murmur?

xiaopengde (a GitHub user) maintains it in xiaopengde/murmur, which has 109 GitHub stars. The repository was last updated on May 31, 2026.

Source: xiaopengde/murmur on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.