Agent skill

Tts Voiceover

by ZJU-REAL in ZJU-REAL/Easel

文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。配了 VOICEPROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI…

Apache-2.0Auto-check: notesMedia & Creative

Install Tts Voiceover

skills CLI
$ npx skills add ZJU-REAL/Easel --skill tts-voiceover -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJU-REAL/Easel tts-voiceover --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJU-REAL/Easel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openclaw/tts-voiceover .claude/skills/tts-voiceover && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tts-voiceover
GitHub stars
3.4k
Token cost
~896 tokens
SKILL.md length
178 words
Files
2
Skills in repo
114
Repo updated
First seen
Licence
Apache-2.0

At a glance

文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。配了 VOICEPROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI…

  • Works in 3 steps: 挑音色(可选) → 合成配音 speak → 后处理(可选,复用已有共享脚本)
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 输入, 输出, 前置:外网代理 and 执行步骤, plus 3 more sections
  • Calls python; needs VOICE_API_KEY

What it does

Tts Voiceover is an agent skill from ZJU-REAL/Easel. 文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。配了 VOICEPROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI 配音”“口播语音”“旁白”“朗读”“把这段文字读出来”“生成语音”“语音合成”时使用。

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EASEL-META.md`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “把这段文字读出来”
  • “/tts-voiceover”

Requirements

  • Python 3
  • A credential in VOICE_API_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 挑音色(可选)
  2. 合成配音 speak
  3. 后处理(可选,复用已有共享脚本)

What it can do on your machine

Read from SKILL.md and the folder at commit 278f420. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tts Voiceover loads about 896 tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~896

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:9
    ` 到 `AGENTS.md` 末尾给出的 Easel 项目根,确认当前目录有 `.env` 和 `skills/shared/scripts/`。云 TTS 配置只能用项目根的 `model_registry.py configured
  • NoteMentions a .env fileSKILL.md:12
    **默认闭源优先**——配了 `.env` 的 `VOICE_PROVIDER`(+ VOICE_API_KEY) 就走闭源云 TTS(voice_clone,

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZJU-REAL/Easel at commit 278f420, republished under its Apache-2.0 licence (© ZJU-REAL). 178 words, ~896 tokens.

Download SKILL.mdSave it as .claude/skills/tts-voiceover/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
tts-voiceover
description
文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。**配了 VOICE_PROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底**(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI 配音”“口播语音”“旁白”“朗读”“把这段文字读出来”“生成语音”“语音合成”时使用。
layer
produce

文字转语音配音(TTS Voiceover)

配置检查路径铁律:先 cd 到 AGENTS.md 末尾给出的 Easel 项目根,确认当前目录有 .env 和 skills/shared/scripts/。云 TTS 配置只能用项目根的 model_registry.py configured --group voice --env-file .env 和 voice_clone.py check ... --env-file .env 判断;不得在 workspace 跑 ./shared/scripts/...,也不得用 env / printenv 推断 Key/URL 缺失。

把文案 / 脚本合成为 AI 语音(口播、旁白、朗读)。共享脚本 skills/shared/scripts/tts.py speak: 默认闭源优先——配了 .env 的 VOICE_PROVIDER(+ VOICE_API_KEY) 就走闭源云 TTS(voice_clone, 按句合成+拼接+分句 SRT,有情感、像真人),没 key 才退 edge(AI 味、生硬,仅兜底)。 --engine closed/edge 可强制;闭源音色用 --voice 传 voice-id(如 FunAudioLLM/CosyVoice2-0.5B:alex), 旁白默认 alex,可用 VOICE_NARRATOR_VOICE_ID 覆盖。合成后语音可交 audio_ops.py/video_ops.py 混音或加到视频。

输入

字段必填说明
text / file是待配音的文本,或文本文件路径(长文本推荐 --file)
voice否音色,默认 zh-CN-XiaoxiaoNeural(晓晓)
rate/volume/pitch否语速 / 音量 / 音调微调
output否默认 outputs/主题名/{name}.mp3

输出

  • 配音音频文件(mp3,可选 wav/m4a),放入 outputs/主题名/
  • 可选同步输出 SRT 字幕(--subtitle),供视频烧字幕用
  • 打印实际执行的 edge-tts 命令 + 输出文件时长/大小/音色

前置:外网代理

edge-tts 调微软在线服务,必须能访问外网。内网环境先设代理:

bash
export https_proxy=http://<代理host>:<端口> http_proxy=http://<代理host>:<端口>

脚本会自动读环境变量代理并透传给 edge-tts(也可用 --proxy 覆盖)。

执行步骤

脚本路径(相对项目根):skills/shared/scripts/tts.py。每个子命令支持 -h。

0. 挑音色(可选)
bash
python skills/shared/scripts/tts.py voices          # 常用中文音色 + 简介
python skills/shared/scripts/tts.py voices --all    # 拉全量 zh- 音色(需外网)
1. 合成配音 speak
bash
# 最简:一句话 → mp3
python skills/shared/scripts/tts.py speak --text "欢迎来到本期内容" \
  -o outputs/主题名/intro.mp3

# 长文本从文件读 + 换音色 + 加速 10%
python skills/shared/scripts/tts.py speak --file script.txt \
  -o outputs/主题名/narration.mp3 --voice zh-CN-YunxiNeural --rate +10%

# 同步出 SRT 字幕(视频烧字幕用)
python skills/shared/scripts/tts.py speak --file script.txt \
  -o outputs/主题名/vo.mp3 --subtitle outputs/主题名/vo.srt

# 输出 wav(需 ffmpeg,便于后续无损处理)
python skills/shared/scripts/tts.py speak --text "……" \
  -o outputs/主题名/vo.wav --format wav

参数:--rate +10%(语速)、--volume +20%(音量)、--pitch +2Hz(音调)。

2. 后处理(可选,复用已有共享脚本)

配音出来后按需接下游脚本,无需在本 SKILL 重造能力:

bash
# ① 配音 + BGM 混音(原声 1.0 / BGM 0.3)→ 用 audio_ops concat / video_ops bgm
python skills/shared/scripts/video_ops.py bgm -i vo.mp3 -o vo_bgm.mp3 \
  --music bgm.mp3 --voice-volume 1.0 --music-volume 0.3

# ② 配音音量归一化到社媒响度(-14 LUFS)
python skills/shared/scripts/audio_ops.py normalize vo.mp3 -o vo_norm.mp3

# ③ 把配音作为旁白加到视频
python skills/shared/scripts/video_ops.py bgm -i clip.mp4 -o clip_vo.mp4 \
  --music vo.mp3 --voice-volume 0.4 --music-volume 1.0

常用中文音色

音色特点
zh-CN-XiaoxiaoNeural晓晓 · 女声,温暖亲和,通用首选(默认)
zh-CN-XiaoyiNeural晓伊 · 女声,活泼年轻,口播/种草
zh-CN-YunxiNeural云希 · 男声,清朗自然,旁白/解说
zh-CN-YunyangNeural云扬 · 男声,专业沉稳,新闻/播报
zh-CN-YunjianNeural云健 · 男声,浑厚有力,激情内容

粤语用 zh-HK-HiuMaanNeural(曉曼),台式用 zh-TW-HsiaoChenNeural(曉臻)。

规则

  1. 绝不覆盖原始素材 — 只写新文件到 outputs/主题名/。
  2. 长文本走 --file — 避免命令行过长 / 换行转义问题。
  3. 先设代理 — edge-tts 需外网,网络失败脚本会给明确提示。
  4. 不重造能力 — 混音/归一化/加视频旁白复用 audio_ops.py / video_ops.py。
  5. 无 Profile 也能用 — 无画像时用默认音色晓晓。

Profile 感知

有 Profile 时可读取账号偏好音色 / 语速 / 平台调性(如口播偏活泼晓伊、 知识类偏沉稳云扬)作为默认参数;无 Profile 退到通用默认(晓晓、正常语速)。

© ZJU-REAL, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/openclaw/tts-voiceover of ZJU-REAL/Easel.

  • SKILL.md
  • EASEL-META.md

Open the folder on GitHubat commit 278f420

Compare with similar skills

Tts Voiceover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tts Voiceover compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tts Voiceover this skillZJU-REAL/Easel3.4k—~896Automated safety check: NotesApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from ZJU-REAL/Easel

All 114 skills in this repo
  • Gzh Design

    ZJU-REAL/Easel

    微信公众号文章排版引擎:把 Markdown / Word(.docx) / PDF / 纯文本转成可直接粘贴进公众号编辑器的 HTML,自动章节编号、关键词标记、引言卡、目录、代码块、图片/GIF、作者签名;主题从 references/theme-index.md…

    3.4k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • 微信公众号文章自动创作与发布工具。给定参考文章、文字或文档,自动搜索整理全网相关信息、生成图文并茂的公众号文章,并发布到微信公众号草稿箱。特别强调反 AI 检测写作。

    3.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Card Design

    ZJU-REAL/Easel

    社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。

    3.4k GitHub stars~657 tokensUpdated today
    Auto-check passed
  • Ecom Details Image

    ZJU-REAL/Easel

    生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Infographic

    ZJU-REAL/Easel

    将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

    3.4k GitHub stars~643 tokensUpdated today
    Auto-check passed
  • Novel Writer

    ZJU-REAL/Easel

    长篇小说/网文连载创作:从世界观、人设和三级大纲写到逐章正文,并用文件化状态维护伏笔、前情和跨章一致性. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Tts Voiceover

What does Tts Voiceover do?

文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。配了 VOICEPROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI…. Tts Voiceover is an agent skill from ZJU-REAL/Easel.

When should I use Tts Voiceover?

Tts Voiceover fits situations like: tasks that involve Text to speech and voice.

How do I install Tts Voiceover in Claude Code?

Run `npx skills add ZJU-REAL/Easel --skill tts-voiceover -a claude-code`. Or copy the skill folder (skills/openclaw/tts-voiceover in ZJU-REAL/Easel) into .claude/skills/tts-voiceover in your project. Claude Code loads it when a task matches its description.

How do I install Tts Voiceover in Codex?

Run `npx skills add ZJU-REAL/Easel --skill tts-voiceover -a codex`. Or copy the skill folder (skills/openclaw/tts-voiceover in ZJU-REAL/Easel) into .agents/skills/tts-voiceover in your project. Codex loads it when a task matches its description.

Can I use Tts Voiceover in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJU-REAL/Easel --skill tts-voiceover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tts-voiceover, .gemini/skills/tts-voiceover, .github/skills/tts-voiceover and .opencode/skills/tts-voiceover in your project.

What does Tts Voiceover need to run?

Going by SKILL.md and its folder, Tts Voiceover needs the command-line tools its instructions call (python) and credentials named VOICE_API_KEY. Our summary lists: Python 3; A credential in VOICE_API_KEY.

Does Tts Voiceover access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tts Voiceover safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tts Voiceover use?

Tts Voiceover is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tts Voiceover use?

About 896 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tts Voiceover?

Skills that share tags, products or a category with Tts Voiceover: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tts Voiceover?

ZJU-REAL (a GitHub organization) maintains it in ZJU-REAL/Easel, which has 3,376 GitHub stars. The repository holds 114 skills in this directory. The repository was last updated on October 9, 2026.

Source: ZJU-REAL/Easel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.