Agent skill

Voice Persona

by davepoon in davepoon/buildwithclaude

让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl.

MITAuto-check passedMedia & Creative

Install Voice Persona

skills CLI
$ npx skills add davepoon/buildwithclaude --skill voice-persona -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davepoon/buildwithclaude voice-persona --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/all-skills/skills/voice-persona .claude/skills/voice-persona && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-persona
GitHub stars
3.6k
Token cost
~947 tokens
SKILL.md length
221 words
Files
2 (incl. scripts)
Skills in repo
247
Repo updated
First seen
Licence
MIT

At a glance

让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl.

  • Tasks that involve Text to speech and voice
  • SKILL.md covers 与同类差异(2026-09 实测调研), 快速验证 / Smoke Test(30 秒,免 key), 使用前提(依赖) and 用法, plus 4 more sections
  • Runs Python scripts from its folder; calls python3 and pip; needs LLM_API_KEY
  • Tasks that involve Messaging and chat bots

What it does

Voice Persona is an agent skill from davepoon/buildwithclaude. 让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl. WeChat silk) and reply in persona voices.

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/voice_persona.py`).

It sits in Media & Creative, covering Text to speech and voice, Messaging and chat bots and Transcription. It works with WeChat and Telegram. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Messaging and chat bots
  • Tasks that involve Transcription

Example prompts

  • “/voice-persona”

Requirements

  • Python 3
  • A credential in LLM_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 616deb5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Persona loads about 947 tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 221 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~947

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davepoon/buildwithclaude at commit 616deb5, republished under its MIT licence (© davepoon). 221 words, ~947 tokens.

Download SKILL.mdSave it as .claude/skills/voice-persona/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
voice-persona
description
让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl. WeChat silk) and reply in persona voices.
category
communication
license
MIT

Voice-Persona · 语音对话 + 人格音色

把 Agent 变成"能听懂语音、会用人格声音回话"的机器人——首个 IM 语音双工闭环技能。

  • 🎙 听懂语音:任意语音文件转文字,支持微信 .silk 格式(独门——大多数开源方案解不了微信语音)
  • 🔊 回语音气泡:人格回复 → Edge TTS 多音色 → silk 编码回发微信语音(pilk 双向,全网首个打通)
  • 🗣 人格回话:内置 6 个人格(元气少女/温柔知心/沉稳大叔/新闻播报/阳光伙伴/随性好友),各配专属中文音色 + 口吻
  • 💰 零 API 成本:faster-whisper(本地)+ Edge TTS(免费)+ pilk(开源)全免费,可选 LLM 增强口吻

Voice chat for agents: transcribe voice (incl. WeChat silk) → reply with persona voices → back to WeChat voice bubbles. First IM two-way voice skill.

与同类差异(2026-09 实测调研)

方案问题
jozhn/wechat-voice-decode-skill只做单向解码转写,无 TTS/人格/语音回复,绑死 openclaw 路径
zhayujie/chatgpt-on-wechat (46k★)语音依赖云端 ASR/TTS,个人微信通道封号风险,无 silk 原生闭环
whisperbot / Telegram 转写 bot 族全部单向"语音→文字",无多音色人格语音回复
SillyTavern (33k★)角色音色强大但只在自家 UI,不接微信/Telegram 语音消息
微软云 voice skills(Azure/OpenAI TTS)单项能力 + 付费云 API,无 IM 场景

别人做"听写"或"朗读",voice-persona 做"在微信/Telegram 里用带人格的嗓音说话"。

快速验证 / Smoke Test(30 秒,免 key)

从本技能目录运行(本仓库已自带脚本,无需额外 clone):

bash
cd plugins/all-skills/skills/voice-persona   # 或直接进入 voice-persona 技能目录
python3 scripts/voice_persona.py demo
# 🎙 语音 → 文字 → [元气少女/沉稳大叔/新闻播报] 三音色回复音频
# ✅ 全链路通过:语音输入 → 转写 → 人格回复 → 语音输出

使用前提(依赖)

bash
pip install faster-whisper pilk edge-tts    # 首次使用安装
# 可选:--llm 增强口吻需要 LLM_API_KEY / OpenAI 兼容 URL(与本集合其他技能一致)

用法

所有命令从技能目录内以相对路径调用脚本(无需安装到系统 PATH):

bash
# 1. 语音 → 文字(微信语音直接传 .silk 文件即可)
python3 scripts/voice_persona.py stt wechat_voice.silk
python3 scripts/voice_persona.py stt meeting.m4a --model small

# 2. 文字 → 人格回复 → 语音 mp3
python3 scripts/voice_persona.py speak "明天记得交报告" --persona yunjian --out reply.mp3

# 3. 只取人格化文本(接你自己的 TTS/IM)
python3 scripts/voice_persona.py chat "明天记得交报告" --persona xiaoyi

# 4. 列出人格
python3 scripts/voice_persona.py list

# 5. 音频 → 微信 silk 语音(可回发微信语音气泡)
python3 scripts/voice_persona.py to_silk reply.mp3 --out reply.silk

# 6. 一条命令双向闭环:人格语音 → 微信格式
python3 scripts/voice_persona.py speak "明天早上十点开会" --persona xiaoyi --out r.mp3
python3 scripts/voice_persona.py to_silk r.mp3          # → r.silk(#!SILK_V3)
在 Hermes 中安装(可选)

本仓库自带脚本可直接运行;若使用 Hermes Agent,也可从技能源仓库安装以自动接入:

bash
hermes skills install jiawood2006/hermes-skills/skills/voice-persona

人格库(可扩展)

在 voice_persona.py 顶部 PERSONAS / STYLE_WORDS 增加即可:

key人格Edge TTS 音色口吻
xiaoyi元气少女 · 小伊zh-CN-XiaoyiNeural嘿嘿、啦/呀
xiaoxuan温柔知心 · 晓萱zh-CN-XiaoxuanNeural别担心、慢慢来
yunjian沉稳大叔 · 云健zh-CN-YunjianNeural直说、结论先行
yunyang新闻播报 · 云扬zh-CN-YunyangNeural播报、条理
yunxi阳光伙伴 · 云希zh-CN-YunxiNeural加油、没问题
xiaochen随性好友 · 晓辰zh-CN-XiaochenNeural口语化、像朋友

Agent 接入(Hermes 等框架)

STT 接入:把平台收到的语音消息文件交给 stt 子命令 → 拿到文字进对话流。 Hermes 的 stt.provider: local_command + HERMES_LOCAL_STT_COMMAND 环境变量可直接把微信语音自动接入(指向本技能 scripts/ 下的脚本绝对路径):

HERMES_LOCAL_STT_COMMAND=<python> <...>/skills/voice-persona/scripts/voice_persona.py stt {input_path} --model {model} --output_dir {output_dir} --language {language}

TTS 接入:人格化文本 → speak 出 mp3 → 平台发送语音。

已知陷阱

  • 微信 silk 必须用 pilk 解码:pilk.decode() 输出的是无 RIFF 头的裸 PCM,whisper 读不了——要用 pilk.silk_to_wav()(输出标准 wav,已实测)
  • 首次运行 faster-whisper 会下载模型(small ~500MB),等待 1-3 分钟属正常
  • Edge TTS 需联网(微软免费接口);断网时 speak 会失败,stt/chat 不受影响
  • 转写质量:Intel Mac CPU int8 下 small 模型中文够用;嘈杂/方言建议 base/small 或重发

验证记录(真实产出)

  • 微信 .silk 语音实测转写成功("API这一块我会让海沃云来对接…"39字符/17字符两段)
  • demo 全链路:say 合成中文 → stt 转写 → 3 人格回复 → 3 个 mp3(27-31KB)✅

© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in plugins/all-skills/skills/voice-persona of davepoon/buildwithclaude.

  • SKILL.md
  • scripts/voice_persona.py

Open the folder on GitHubat commit 616deb5

Compare with similar skills

Voice Persona next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Persona compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Persona this skilldavepoon/buildwithclaude3.6k—~947Automated safety check: PassMIT
AI Video Scriptzrt-ai-lab/opencode-skills287—~1.7kAutomated safety check: PassNone
FeedgrabiBigQiang/feedgrab614—~2kAutomated safety check: PassMIT
Lingzaoatian-create/lingzao-skill2961 repos~8.8kAutomated safety check: PassMIT
Ffmpeg Best Practicelinyqh/speclip-skills110—~2.3kAutomated safety check: PassNone
Voice Memoletta-ai/lettabot327—~487Automated safety check: PassApache-2.0

Similar skills

  • AI Video Script

    zrt-ai-lab/opencode-skills

    A skill your agent uses when a request asks for a Chinese-first AI video script with shot plans, image prompts, narration, subtitles, or handoff contracts for scene generation, image generation…

    287 GitHub stars~1.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Feedgrab

    iBigQiang/feedgrab

    Universal content grabber — fetch any URL and return structured Markdown.

    614 GitHub stars~2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Lingzao

    atian-create/lingzao-skill

    Use Lingzao creator-content tools for Xiaohongshu/XHS, Douyin, and WeChat official-account public content.

    296 GitHub starsUsed in 1 repo~8.8k tokens
    Media & CreativeAuto-check passed
  • Ffmpeg Best Practice

    linyqh/speclip-skills

    Lean FFmpeg playbook for reliable video compression, WeChat-compatible MP4 export, clip stitching, audio mixing, subtitle handling, and quick fallback decisions.

    110 GitHub stars~2.3k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed
  • Voice Memo

    letta-ai/lettabot

    Reply with voice memos using text-to-speech. An agent skill from letta-ai/lettabot.

    327 GitHub stars~487 tokensUpdated 4 mo ago
    Media & CreativeAuto-check passed
  • Video Downloader

    kangarooking/kangarooking-skills

    Download or open videos and recover platform captions, audio transcripts, keyframes, screen text, visual facts, and editing observations as a plain multimodaltranscript.md.

    662 GitHub stars~8.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from davepoon/buildwithclaude

All 247 skills in this repo
  • iOS Hig Design Guide

    davepoon/buildwithclaude

    Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.

    3.6k GitHub stars~735 tokensUpdated yesterday
    Auto-check passed
  • Video Downloader

    davepoon/buildwithclaude

    Download YouTube videos with customizable quality and format options.

    3.6k GitHub starsUsed in 1 repo~871 tokens
    Auto-check passed
  • Qwen Vision

    davepoon/buildwithclaude

    A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

    3.6k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Atlas Cloud Media

    davepoon/buildwithclaude

    Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.

    3.6k GitHub stars~852 tokensUpdated yesterday
    Auto-check passed
  • Browser Extension Launch

    davepoon/buildwithclaude

    面向没有编程经验的用户,把想法做成可试用的浏览器插件,并完成检查、商店材料、审核提交和上线验证;也用于继续已有插件、排错和发布新版。用户说“帮我做个插件”“把插件上架”“继续我的插件”时使用。普通网站开发、仅查询插件知识不触发。

    3.6k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Slack Gif Creator

    davepoon/buildwithclaude

    Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.

    3.6k GitHub starsUsed in 12 repos~4.3k tokens
    Auto-check passed

Works with

Questions about Voice Persona

What does Voice Persona do?

让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl. Voice Persona is an agent skill from davepoon/buildwithclaude. 让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl.

When should I use Voice Persona?

Voice Persona fits situations like: tasks that involve Text to speech and voice; tasks that involve Messaging and chat bots; tasks that involve Transcription.

How do I install Voice Persona in Claude Code?

Run `npx skills add davepoon/buildwithclaude --skill voice-persona -a claude-code`. Or copy the skill folder (plugins/all-skills/skills/voice-persona in davepoon/buildwithclaude) into .claude/skills/voice-persona in your project. Claude Code loads it when a task matches its description.

How do I install Voice Persona in Codex?

Run `npx skills add davepoon/buildwithclaude --skill voice-persona -a codex`. Or copy the skill folder (plugins/all-skills/skills/voice-persona in davepoon/buildwithclaude) into .agents/skills/voice-persona in your project. Codex loads it when a task matches its description.

Can I use Voice Persona in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill voice-persona -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-persona, .gemini/skills/voice-persona, .github/skills/voice-persona and .opencode/skills/voice-persona in your project.

What does Voice Persona need to run?

Going by SKILL.md and its folder, Voice Persona needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named LLM_API_KEY. Our summary lists: Python 3; A credential in LLM_API_KEY.

Does Voice Persona access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Voice Persona safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Voice Persona use?

Voice Persona is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice Persona use?

About 947 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Persona?

Skills that share tags, products or a category with Voice Persona: AI Video Script (zrt-ai-lab/opencode-skills, 287 stars), Feedgrab (iBigQiang/feedgrab, 614 stars), Lingzao (atian-create/lingzao-skill, 296 stars) and Ffmpeg Best Practice (linyqh/speclip-skills, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Persona?

davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,610 GitHub stars. The repository holds 247 skills in this directory. The repository was last updated on October 9, 2026.

Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.