Agent skill

Multi Voice Dubbing

by ZJU-REAL in ZJU-REAL/Easel

多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕. An agent skill from ZJU-REAL/Easel.

Apache-2.0Auto-check: notesMedia & Creative

Install Multi Voice Dubbing

skills CLI
$ npx skills add ZJU-REAL/Easel --skill multi-voice-dubbing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJU-REAL/Easel multi-voice-dubbing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJU-REAL/Easel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openclaw/multi-voice-dubbing .claude/skills/multi-voice-dubbing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-voice-dubbing
GitHub stars
3.4k
Token cost
~1.2k tokens
SKILL.md length
292 words
Files
2
Skills in repo
114
Repo updated
First seen
Licence
Apache-2.0

At a glance

多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕. An agent skill from ZJU-REAL/Easel.

  • Works in 4 steps: 选角(cast.json)——读选角指南定音色 → 逐行对白(lines.json) → 合成多声线音轨 + 字幕 → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 配音质量分层(重要:治「像 AI 平读」), 谁会用到, 输入 and 执行步骤, plus 4 more sections
  • Calls python; reaches api.siliconflow.cn; needs VOICE_API_KEY and GEMINI_API_KEY

What it does

Multi Voice Dubbing is an agent skill from ZJU-REAL/Easel. 多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕。 当用户说“多角色/双人/剧本/对话配音、多人对白、不同角色不同声音、有声剧配音”时使用。 单一公共音色用 tts-voiceover;克隆本人音色用 voice-clone;整部短剧制作由 short-drama 编排。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EASEL-META.md`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “多角色/双人/剧本/对话配音、多人对白、不同角色不同声音、有声剧配音”
  • “/multi-voice-dubbing”

Requirements

  • Python 3
  • A credential in VOICE_API_KEY
  • A credential in GEMINI_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 选角(cast.json)——读选角指南定音色
  2. 逐行对白(lines.json)
  3. 合成多声线音轨 + 字幕
  4. 用到视频里

What it can do on your machine

Read from SKILL.md and the folder at commit 278f420. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.siliconflow.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOICE_API_KEY
    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multi Voice Dubbing loads about 1.2k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 292 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:31
    conFlow**(云端 CosyVoice2、无需 GPU、中文最稳)配置落 `.env`:
  • NoteMentions a .env fileSKILL.md:78
    > 需在 `.env` 配对应 provider 的 key(`voice_clone.py check --provider minimax`);缺 key 该角色**自动回退 edge**(平)并告警。

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZJU-REAL/Easel at commit 278f420, republished under its Apache-2.0 licence (© ZJU-REAL). 292 words, ~1,198 tokens.

Download SKILL.mdSave it as .claude/skills/multi-voice-dubbing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
multi-voice-dubbing
description
多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕。 当用户说“多角色/双人/剧本/对话配音、多人对白、不同角色不同声音、有声剧配音”时使用。 单一公共音色用 tts-voiceover;克隆本人音色用 voice-clone;整部短剧制作由 short-drama 编排。
layer
produce

多角色 / 对话配音(Multi-voice Dubbing)

把「多人对话 / 多角色脚本」合成为多声线音轨——每个角色一个符合其人设的声音, 不再是「全程一个声音」。核心引擎 skills/shared/scripts/multivoice.py: 逐行调 tts.py(edge-tts 免费多音色)或 voice_clone.py(云端表现力 provider / 克隆音色)合成, 把每行 emotion 喂进 provider 的真情感通道,ffmpeg 拼成一轨 + 生成对齐的说话人字幕。

创意(谁说什么、什么情绪)由你 LLM 产出;音色映射在 cast。确定性 IO(逐行合成/拼接/字幕)走引擎。 产物 voice.mp3 可直接当 narration 喂 auto-short-video/assemble.py 或加进任意视频;voice.srt 是带角色名的字幕。

配音质量分层(重要:治「像 AI 平读」)

edge-tts 没有情感引擎,只能变速变调,再怎么调也像机器平读。要"像人"必须用有情感通道的云 provider(用户自备 key,无需 GPU):

引擎(cast 里 engine)质量需要情感机制
edge(默认,免费兜底)⚠️ 平、机器感,仅供草稿无 key、外网仅 rate/pitch/volume 微调
clone+openai-compatible→SiliconFlow CosyVoice2(推荐)好、中文自然VOICE_API_KEY(便宜/新用户赠额)内联 <|endofprompt|> 指令
clone+gemini好、真免费GEMINI_API_KEY(Google 免费层,国内需代理)自然语言前缀
clone+minimax / dashscope好、有情绪各家 keyemotion 枚举 / instruct

推荐 SiliconFlow(云端 CosyVoice2、无需 GPU、中文最稳)配置落 .env:

bash
VOICE_PROVIDER=openai-compatible
VOICE_BASE_URL=https://api.siliconflow.cn/v1
VOICE_API_KEY=<你的key>
VOICE_MODEL=FunAudioLLM/CosyVoice2-0.5B
VOICE_INSTRUCT_MODE=inline

cast 里角色:--engine clone --provider openai-compatible --voice-id FunAudioLLM/CosyVoice2-0.5B:alex(8 音色 alex/anna/benjamin/…)。

  • 逐行 emotion 自动驱动演绎:lines.json 每行的 emotion(愤怒/崩溃大哭/冷笑/温柔…)→ 引擎按 provider 转成对应情感参数。写具体越贴戏越好。
  • 想要「像人」→ 至少给主角/关键角色配一个云 provider(engine=clone);配了 key 才有情绪,没 key 自动回退 edge(平)并告警。
  • provider 配置见 voice_clone.py 头部(各家 env);voice_clone.py check --provider <名> 离线校验 key 是否齐。

谁会用到

短剧对白(short-drama 已内部委派)、论文双人问答讲解(paper-explainer:主讲+提问者)、 访谈/播客脚本、有声剧、任何「多个说话人」的口播。单人整段口播用 tts-voiceover 即可。

输入

字段必填说明
cast.json是选角表:每个说话人 → 音色(edge 音色 + pitch/rate,或克隆音色)。含「旁白/主讲」条目
lines.json是逐行对白:有序 [{speaker, text, emotion}](speaker 用 cast 里的名字;emotion 如 冷/怒/紧张/温柔,自动匹配语气)
输出路径否voice.mp3(默认与调用方约定);voice.srt 同名

执行步骤

1. 选角(cast.json)——读选角指南定音色

先读 skills/shared/references/voice-casting.md(音色→人物原型对照 + 同性别区分 + 旁白独立 + 情绪韵律)。 为每个说话人定一个符合其性别/年龄/气质/身份的音色,旁白/主讲单列且与所有角色不同:

bash
python skills/shared/scripts/multivoice.py cast init  --cast <路径>/cast.json
python skills/shared/scripts/multivoice.py cast add   --cast <路径>/cast.json \
    --name 林策 --role male_lead --voice zh-CN-YunxiNeural --rate=-5% --pitch=-3Hz --note "冷峻男主"   # 负值参数用等号
python skills/shared/scripts/multivoice.py cast add   --cast <路径>/cast.json \
    --name 苏晚 --role female_lead --voice zh-CN-XiaoxiaoNeural --note "温婉女主"
python skills/shared/scripts/multivoice.py cast check --cast <路径>/cast.json    # 校验:音色有效/旁白独立/无撞音色

cast.json 也可直接写(就是 JSON)。想让某角色像真人有情绪(治 AI 平读)→ 该角色用云 provider:

bash
python skills/shared/scripts/multivoice.py cast add --cast <路径>/cast.json \
    --name 霸总 --role male_lead --engine clone --provider minimax --voice-id <你的音色id> --note "克隆/表现力音色"

需在 .env 配对应 provider 的 key(voice_clone.py check --provider minimax);缺 key 该角色自动回退 edge(平)并告警。

2. 逐行对白(lines.json)

把脚本 / 对话稿拆成有序逐行:

json
{"lines":[
  {"speaker":"主讲","text":"这篇论文解决了一个关键问题。","emotion":"平"},
  {"speaker":"提问","text":"等等,为什么现有方法不行?","emotion":"惊"},
  {"speaker":"主讲","text":"因为它们忽略了时序依赖。","emotion":"坚定"}
]}

speaker 必须与 cast 里的名字一致(不一致会回退旁白音色并告警);emotion 可选。

3. 合成多声线音轨 + 字幕
bash
python skills/shared/scripts/multivoice.py dub \
    --cast <路径>/cast.json --lines <路径>/lines.json -o <路径>/voice.mp3

产出 voice.mp3(每角色独立声线、情绪自动调韵律)+ voice.srt(带角色名,时序按逐行实测时长对齐)。 dub 会打印用了几种声线——确认 ≥2 种(多人对话不该只有一个声音)。

4. 用到视频里
  • voice.mp3 作 narration + voice.srt 作字幕,喂 auto-short-video/scripts/assemble.py(图/视频合成)。
  • 或与 BGM 混音(audio-mix)、加进已有视频(video-editing)。

降级

  • 缺外网/edge-tts 不通 → 无法合成,如实告知(多声线依赖 edge-tts)。
  • cast 指定克隆音色但缺 key/失败 → 该角色自动回退 edge 免费音色并告警,不阻断。
  • 说话人不在 cast → 回退旁白音色并告警(建议补进 cast)。

Profile 感知

  • 有 Profile:style.md 融进音色气质选择(角色气质→音色);preferences.md 红线过滤台词。
  • 无 Profile:按 voice-casting.md 默认对照选音色。

规则

  1. 多人对话必须多声线:每个说话人独立音色、旁白/主讲单列——绝不允许全程一个声音(cast check + dub 声线数把关)。
  2. 要像人就上云 provider:edge 只配草稿;成品/关键角色用 engine=clone+provider(有真情感通道),无需 GPU。
  3. 音色贴人物:按 voice-casting.md 的原型对照 + 同性别用 pitch/rate 区分。
  4. 情绪标注要具体:lines 每行标 emotion(愤怒/崩溃大哭/冷笑/温柔…),会喂进 provider 情感通道驱动演绎。
  5. 确定性留档:cast.json / lines.json 落文件,改台词/换音色重跑 dub 即可,不必重来。

参考来源

见 EASEL-META.md。多声线配音(cast + 逐行 lines + 按角色音色逐行合成 + ffmpeg 拼接 + 说话人字幕)为 Easel 自研; 底层封装 edge-tts(tts.py)与云端克隆(voice_clone.py)。

© ZJU-REAL, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/openclaw/multi-voice-dubbing of ZJU-REAL/Easel.

  • SKILL.md
  • EASEL-META.md

Open the folder on GitHubat commit 278f420

Compare with similar skills

Multi Voice Dubbing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Voice Dubbing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Voice Dubbing this skillZJU-REAL/Easel3.4k—~1.2kAutomated safety check: NotesApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from ZJU-REAL/Easel

All 114 skills in this repo
  • Gzh Design

    ZJU-REAL/Easel

    微信公众号文章排版引擎:把 Markdown / Word(.docx) / PDF / 纯文本转成可直接粘贴进公众号编辑器的 HTML,自动章节编号、关键词标记、引言卡、目录、代码块、图片/GIF、作者签名;主题从 references/theme-index.md…

    3.4k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • 微信公众号文章自动创作与发布工具。给定参考文章、文字或文档,自动搜索整理全网相关信息、生成图文并茂的公众号文章,并发布到微信公众号草稿箱。特别强调反 AI 检测写作。

    3.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Card Design

    ZJU-REAL/Easel

    社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。

    3.4k GitHub stars~657 tokensUpdated today
    Auto-check passed
  • Ecom Details Image

    ZJU-REAL/Easel

    生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Infographic

    ZJU-REAL/Easel

    将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

    3.4k GitHub stars~643 tokensUpdated today
    Auto-check passed
  • Novel Writer

    ZJU-REAL/Easel

    长篇小说/网文连载创作:从世界观、人设和三级大纲写到逐章正文,并用文件化状态维护伏笔、前情和跨章一致性. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Multi Voice Dubbing

What does Multi Voice Dubbing do?

多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕. An agent skill from ZJU-REAL/Easel. Multi Voice Dubbing is an agent skill from ZJU-REAL/Easel.

When should I use Multi Voice Dubbing?

Multi Voice Dubbing fits situations like: tasks that involve Text to speech and voice.

How do I install Multi Voice Dubbing in Claude Code?

Run `npx skills add ZJU-REAL/Easel --skill multi-voice-dubbing -a claude-code`. Or copy the skill folder (skills/openclaw/multi-voice-dubbing in ZJU-REAL/Easel) into .claude/skills/multi-voice-dubbing in your project. Claude Code loads it when a task matches its description.

How do I install Multi Voice Dubbing in Codex?

Run `npx skills add ZJU-REAL/Easel --skill multi-voice-dubbing -a codex`. Or copy the skill folder (skills/openclaw/multi-voice-dubbing in ZJU-REAL/Easel) into .agents/skills/multi-voice-dubbing in your project. Codex loads it when a task matches its description.

Can I use Multi Voice Dubbing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJU-REAL/Easel --skill multi-voice-dubbing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-voice-dubbing, .gemini/skills/multi-voice-dubbing, .github/skills/multi-voice-dubbing and .opencode/skills/multi-voice-dubbing in your project.

What does Multi Voice Dubbing need to run?

Going by SKILL.md and its folder, Multi Voice Dubbing needs the command-line tools its instructions call (python) and credentials named VOICE_API_KEY and GEMINI_API_KEY. Our summary lists: Python 3; A credential in VOICE_API_KEY; A credential in GEMINI_API_KEY.

Does Multi Voice Dubbing access the network?

SKILL.md names 1 domain. In commands or code: api.siliconflow.cn; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Multi Voice Dubbing safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Multi Voice Dubbing use?

Multi Voice Dubbing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multi Voice Dubbing use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multi Voice Dubbing?

Skills that share tags, products or a category with Multi Voice Dubbing: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Voice Dubbing?

ZJU-REAL (a GitHub organization) maintains it in ZJU-REAL/Easel, which has 3,376 GitHub stars. The repository holds 114 skills in this directory. The repository was last updated on October 9, 2026.

Source: ZJU-REAL/Easel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.