Agent skill

Wjs Voicedrop Reading Aloud

by jianshuo in jianshuo/claude-skills

A skill your agent uses when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices.

MITAuto-check: notesMedia & Creative

Install Wjs Voicedrop Reading Aloud

skills CLI
$ npx skills add jianshuo/claude-skills --skill wjs-voicedrop-reading-aloud -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jianshuo/claude-skills wjs-voicedrop-reading-aloud --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/wjs-voicedrop-reading-aloud .claude/skills/wjs-voicedrop-reading-aloud && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wjs-voicedrop-reading-aloud
GitHub stars
131
Token cost
~894 tokens
SKILL.md length
180 words
Files
1
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices.

  • Works in 5 steps: 取内容并认真通读 → 编排重写成朗读脚本 → 加语音指令(克制) → …
  • The user wants text turned into an audiobook / read aloud — they give a passage
  • SKILL.md covers 铁律, 工作流, 音色表(已实测可用,全部支持语音指令) and 常见坑
  • Calls python3

What it does

Wjs Voicedrop Reading Aloud is an agent skill from jianshuo/claude-skills. Use when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices. Triggers — "做成有声书", "朗读出来", "读给我听", "念出来", "read aloud", "/wjs-voicedrop-reading-aloud".

Its SKILL.md is about 890 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: 13 Claude Code skills for video production (transcribe / translate / dub / multicam / subtitles / reframe) + WeChat publishing. Compatible with Claude Code, OpenAI Codex CLI… The licence is MIT.

When your agent uses it

  • The user wants text turned into an audiobook / read aloud — they give a passage
  • A VoiceDrop article and want an mp3 narration with expressive voices
  • /wjs-voicedrop-reading-aloud

Example prompts

  • “read aloud”
  • “/wjs-voicedrop-reading-aloud”
  • “/wjs-voicedrop-reading-aloud”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 取内容并认真通读
  2. 编排重写成朗读脚本
  3. 加语音指令(克制)
  4. 朗读脚本格式(tts.py --script)
  5. 合成与验收

What it can do on your machine

Read from SKILL.md and the folder at commit b2690f5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • volcengine.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wjs Voicedrop Reading Aloud loads about 894 tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 180 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~894

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:10
    ode/volcano-tts/tts.py`(先 `source ~/code/.env`)。
  • NoteMentions a .env fileSKILL.md:73
    source ~/code/.env

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jianshuo/claude-skills at commit b2690f5, republished under its MIT licence (© jianshuo). 180 words, ~894 tokens.

Download SKILL.mdSave it as .claude/skills/wjs-voicedrop-reading-aloud/SKILL.md (or your agent's skills folder).
name
wjs-voicedrop-reading-aloud
description
Use when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices. Triggers — "做成有声书", "朗读出来", "读给我听", "念出来", "read aloud", "/wjs-voicedrop-reading-aloud".

wjs-voicedrop-reading-aloud

文字 → 有声书 mp3。不是照字面念,而是先编排:不同性质的内容换不同声音,关键转折处加语音指令,然后用火山引擎豆包 seed-tts-2.0 合成。

合成工具:~/code/volcano-tts/tts.py(先 source ~/code/.env)。

铁律

  1. 绝不把带【】/[ ] 标注的文本直接喂裸 API —— 裸 API 会把标注原样念出来(实测实锤)。永远走 tts.py,它把标注切段转成 context_texts 语音指令。
  2. 只能用 2.0 音色(_uranus_bigtts 后缀) —— moon/mars/tob 音色在 seed-tts-2.0 资源下报错 resource ID is mismatched。
  3. 合成前把编排好的朗读脚本给用户看一眼(除非用户说直接出)——编排是再创作,声音分配和指令值得确认。headless/自动化场景跳过此步。

工作流

1. 取内容并认真通读
  • 纯文本/文件:直接读。
  • URL:WebFetch;SPA 页面(如 docs.volcengine.com)用 browse skill 渲染后取正文。
  • VoiceDrop 文章:voicedrop MCP 的 read_article。

通读时标记出:正文叙述 / 直接引用(引号、blockquote)/ 大白话吐槽与内心 OS / 数据、列表、表格 / 标题与小节。

2. 编排重写成朗读脚本

这是核心步骤,是重写不是转录:

  • 去掉一切视觉残留:markdown 符号、链接、图片说明、脚注编号。
  • 表格、列表、数据改写成口语句子(「三个原因:第一…」)。
  • 标题不逐字念,化进过渡句,或用停顿+换气带过。
  • 太书面的长句改口语,但保留作者的用词风格。
  • 按内容性质分配音色(见音色表):正文一个主声贯穿;引用换引用声;吐槽/大白话换插话声。声音切换是给听众的「格式信号」,等价于视觉上的引用块。
  • 不要频繁换声:一般 2~3 个声音封顶,切换只发生在内容性质真正变化处。
3. 加语音指令(克制)

在句前加 [心理活动、细腻表情、肢体动作等描述],如 [放慢,一字一顿,点出要害]。

  • 只在需要的地方加:情绪转折、节奏变化、重音、引用的口吻模仿。平铺直叙的段落一个不加,靠全局指令兜底。
  • 经验密度:每 3~5 句最多一处;一段平静的叙述可以整段没有。
  • 每处标注就是一次切段(一次 API 调用+拼接点),切太碎会让语流变散。
  • 标注写成对朗读者说的表演提示(心理活动/表情/动作皆可),不要写成对听众的说明。
  • 指令要戏剧化、情绪化才有效(实测):模型对情绪/音色类指令跟随很强(哭腔、耳语、亢奋大喊、像法官宣判、+50% 时长级别的变化),对含蓄舞台提示(「语气一沉」「带一丝惋惜」)和机械精确指令(「停顿一秒」「放慢一倍」)跟随很弱。写法上宁可夸张:「请把声音压到接近耳语,凑近话筒,像说破一个秘密」远强于「压低声音」。
4. 朗读脚本格式(tts.py --script)
# 注释行
@voice narrator zh_male_yuanboxiaoshu_uranus_bigtts
@voice quote zh_male_yizhipiannan_uranus_bigtts
@voice casual zh_male_fanjuanqingnian_uranus_bigtts

@narrator
[语气平静从容,像老朋友聊天]先讲一个真事。……他公开断言:

@quote
[带着当年的笃定与体面]股价已经站上了一个永久的高原。

@narrator
几天后,市场开始了最惨烈的下跌。

@casual
[像随口吐槽]这人判断力真差。
5. 合成与验收
bash
source ~/code/.env
python3 ~/code/volcano-tts/tts.py -f script.txt --script -o out.mp3 \
  -i "这是一段有声书朗读,自然口语化,像讲故事,不要播音腔"
  • -i 全局指令必带,定整体基调;--speech-rate、--subtitle(字级时间戳)按需。
  • 验收:afinfo out.mp3 看时长是否与字数匹配(中文约 4~5 字/秒);如首次改动过工具或有疑虑,用 ~/.claude/skills/wjs-transcribing-audio/scripts/volc_asr_stream.py 抽查一段,确认标注没被念出来。
  • 用 SendUserFile 把 mp3 发给用户。

音色表(已实测可用,全部支持语音指令)

用途voice_type名称
旁白主声(男,默认)zh_male_yuanboxiaoshu_uranus_bigtts渊博小叔
旁白主声(女,备选)zh_female_zhixingnv_uranus_bigtts知性女声
旁白(对话感/播客感)zh_male_shenyeboke_uranus_bigtts深夜播客
解说腔zh_male_cixingjieshuonan_uranus_bigtts磁性解说男声
引用/名人语录/译文zh_male_yizhipiannan_uranus_bigtts译制片男
大白话/吐槽/内心 OSzh_female_shuangkuaisisi_uranus_bigtts爽快思思(用户定:大白话用女声,与男声旁白区分)
大白话(男声备选)zh_male_fanjuanqingnian_uranus_bigtts反卷青年
通用女声zh_female_vv_uranus_bigttsVivi

更多 2.0 音色:https://www.volcengine.com/docs/6561/1257544 (只认 _uranus_bigtts 后缀)。

常见坑

  • 指令悄无声息不生效(听起来全一样) → 三个已实锤的坑,tts.py 均已修复:① context_texts/section_id 必须嵌在 additions(JSON 字符串)里,放 req_params 顶层被静默忽略;② context_texts 传多条时只有第一条生效——段内指令必须独占一条,不能「全局+段内」并列(全局排前面会把段内全顶掉);③ 转述包装(「按照这个指示朗读:X」)会稀释效果,标注原文直发最强。怀疑时用「极慢哭腔」指令 A/B 对比时长(应 +25% 以上),基线合成是确定性的(同句同参时长完全一致),差异小于 5% 即未生效。
  • 标注被念出来 → 忘了走 tts.py 的标注解析,或用了它不认的括号格式(支持 【】 与 [ ])。
  • resource ID is mismatched with speaker related resource → 用了非 2.0 音色。
  • 输出中间语气断裂感明显 → 标注切段太碎,合并标注、减少切换。
  • 听到「星号」「井号」→ 编排步骤没洗干净 markdown(tts.py 未开 markdown 过滤,靠编排时清除)。

© jianshuo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in wjs-voicedrop-reading-aloud of jianshuo/claude-skills.

Open the folder on GitHubat commit b2690f5

Compare with similar skills

Wjs Voicedrop Reading Aloud next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wjs Voicedrop Reading Aloud compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wjs Voicedrop Reading Aloud this skilljianshuo/claude-skills131—~894Automated safety check: NotesMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from jianshuo/claude-skills

All 38 skills in this repo
  • Wjs Segmenting Video

    jianshuo/claude-skills

    A skill your agent uses when the user has a long-form video (interview / lecture / podcast / conversation) and a transcript SRT, and wants to extract 3–6 stand-alone topical short clips from it.

    131 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Uploading Video

    jianshuo/claude-skills

    Upload one or many videos to YouTube. An agent skill from jianshuo/claude-skills.

    131 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Converting Wp To Hugo

    jianshuo/claude-skills

    A skill your agent uses when migrating a WordPress site to a Hugo static site on GitHub Pages from a WXR export (.xml) plus the wp-content/uploads folder — preserving /archives/<id/ URLs, localizing…

    131 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Burning Subtitles

    jianshuo/claude-skills

    A skill your agent uses when the user has a video + an SRT and wants the subtitles either burned into the pixels (libass, always-visible) or soft-muxed as a togglable track.

    131 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Cleaning Spam

    jianshuo/claude-skills

    A skill your agent uses when the user complains about spam on his X/Twitter posts — 同城面付 / 寻固炮 / 线下上门 / 免费破处 这类引流号在他推文下刷的 emoji 垃圾回复 — and wants them removed.

    131 GitHub stars~532 tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Creating Video Book

    jianshuo/claude-skills

    A skill your agent uses when the user wants a book turned into YouTube chapter videos — 每章用 VoiceDrop 读书的有声书 mp3 做音轨,配 GPT Image 2 画面和中心思想大字,输出 1920×1080 横屏视频发 YouTube。Triggers — "把这本书做成视频"…

    131 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Wjs Voicedrop Reading Aloud

What does Wjs Voicedrop Reading Aloud do?

A skill your agent uses when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices. Wjs Voicedrop Reading Aloud is an agent skill from jianshuo/claude-skills. Use when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices.

When should I use Wjs Voicedrop Reading Aloud?

Wjs Voicedrop Reading Aloud fits situations like: the user wants text turned into an audiobook / read aloud — they give a passage; A VoiceDrop article and want an mp3 narration with expressive voices; /wjs-voicedrop-reading-aloud.

How do I install Wjs Voicedrop Reading Aloud in Claude Code?

Run `npx skills add jianshuo/claude-skills --skill wjs-voicedrop-reading-aloud -a claude-code`. Or copy the skill folder (wjs-voicedrop-reading-aloud in jianshuo/claude-skills) into .claude/skills/wjs-voicedrop-reading-aloud in your project. Claude Code loads it when a task matches its description.

How do I install Wjs Voicedrop Reading Aloud in Codex?

Run `npx skills add jianshuo/claude-skills --skill wjs-voicedrop-reading-aloud -a codex`. Or copy the skill folder (wjs-voicedrop-reading-aloud in jianshuo/claude-skills) into .agents/skills/wjs-voicedrop-reading-aloud in your project. Codex loads it when a task matches its description.

Can I use Wjs Voicedrop Reading Aloud in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jianshuo/claude-skills --skill wjs-voicedrop-reading-aloud -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wjs-voicedrop-reading-aloud, .gemini/skills/wjs-voicedrop-reading-aloud, .github/skills/wjs-voicedrop-reading-aloud and .opencode/skills/wjs-voicedrop-reading-aloud in your project.

What does Wjs Voicedrop Reading Aloud need to run?

Going by SKILL.md and its folder, Wjs Voicedrop Reading Aloud needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Wjs Voicedrop Reading Aloud access the network?

SKILL.md names 1 domain. As links in the text: volcengine.com. This is read from the text; nothing was executed.

Is Wjs Voicedrop Reading Aloud safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Wjs Voicedrop Reading Aloud use?

Wjs Voicedrop Reading Aloud is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wjs Voicedrop Reading Aloud use?

About 894 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wjs Voicedrop Reading Aloud?

Skills that share tags, products or a category with Wjs Voicedrop Reading Aloud: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wjs Voicedrop Reading Aloud?

jianshuo (a GitHub user) maintains it in jianshuo/claude-skills, which has 131 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on August 20, 2026.

Source: jianshuo/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.