Agent skill

Voice Clone

by ZJU-REAL in ZJU-REAL/Easel

上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。

Apache-2.0Auto-check: notesMedia & Creative

Install Voice Clone

skills CLI
$ npx skills add ZJU-REAL/Easel --skill voice-clone -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJU-REAL/Easel voice-clone --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJU-REAL/Easel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openclaw/voice-clone .claude/skills/voice-clone && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-clone
GitHub stars
3.4k
Token cost
~715 tokens
SKILL.md length
175 words
Files
1
Skills in repo
114
Repo updated
First seen
Licence
Apache-2.0

At a glance

上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。

  • Works in 2 steps: 登记音色(enroll,得到 voice_id) → 合成(clone)
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 前置:配置 API key, 输入, 输出(outputs/主题名/) and 执行步骤, plus 3 more sections
  • Calls python; needs DASHSCOPE_API_KEY and MINIMAX_API_KEY

What it does

Voice Clone is an agent skill from ZJU-REAL/Easel. 上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。

Its SKILL.md is about 720 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with MiniMax and OpenAI. The repository describes itself as: An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”
  • “/voice-clone”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY
  • A credential in MINIMAX_API_KEY

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. 登记音色(enroll,得到 voice_id)
  2. 合成(clone)

What it can do on your machine

Read from SKILL.md and the folder at commit 278f420. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY
    • MINIMAX_API_KEY
    • FISH_API_KEY
    • VOICE_API_KEY
    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Clone loads about 715 tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 175 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~715

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:17
    ` 到 `AGENTS.md` 末尾给出的 Easel 项目根,确认当前目录有 `.env` 和 `skills/shared/scripts/`,再运行注册表、`check`、`enroll` 或 `clone`。不得改用 workspa
  • NoteMentions a .env fileSKILL.md:19
    选 provider 并在 `.env` 填 key,再 `check` 离线校验:
  • NoteMentions a .env fileSKILL.md:21
    e.py check --provider minimax --env-file .env
  • NoteMentions a .env fileSKILL.md:24
    | provider | 服务 | .env 需配 |
  • NoteMentions a .env fileSKILL.md:34
    y.py configured --group voice --env-file .env`。只有一个可用时显式选择;多个可用且用户没点名时,列出 provider/模型询问本次使用哪个,不按 `VOICE_PROVIDER` 擅自选择。

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZJU-REAL/Easel at commit 278f420, republished under its Apache-2.0 licence (© ZJU-REAL). 175 words, ~715 tokens.

Download SKILL.mdSave it as .claude/skills/voice-clone/SKILL.md (or your agent's skills folder).
name
voice-clone
description
上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。
layer
produce

声音克隆配音

用本人语音样本克隆音色,再合成任意文案。走云端 provider(用户自备 key),本地无需 GPU。 全部走 skills/shared/scripts/voice_clone.py。

不想克隆、用现成公共音色见 tts-voiceover(edge-tts,免费无需 key); AI 生成音乐/BGM 见 ai-music;合成后与 BGM 混音见 audio-mix。

前置:配置 API key

配置检查路径铁律:先 cd 到 AGENTS.md 末尾给出的 Easel 项目根,确认当前目录有 .env 和 skills/shared/scripts/,再运行注册表、check、enroll 或 clone。不得改用 workspace 的 ./shared/scripts/...,也不得以 env / printenv 没显示变量为由判断 VOICE_BASE_URL/Key 缺失。check 支持时显式传 --env-file .env。

选 provider 并在 .env 填 key,再 check 离线校验:

bash
python skills/shared/scripts/voice_clone.py check --provider minimax --env-file .env
provider服务.env 需配
dashscope阿里 CosyVoice 声音复刻DASHSCOPE_API_KEY(可选 DASHSCOPE_TTS_MODEL/DASHSCOPE_BASE_URL)
minimaxMiniMax 语音克隆MINIMAX_API_KEY、MINIMAX_GROUP_ID(可选 MINIMAX_MODEL)
fish-audioFish AudioFISH_API_KEY(可选 FISH_BASE_URL)
openai-compatibleOpenAI 兼容 /audio/speechVOICE_API_KEY、VOICE_BASE_URL(预置 voice,非零样本克隆)
geminiGoogle Gemini TTSGEMINI_API_KEY(可选 GEMINI_TTS_MODEL/GEMINI_VOICE/GEMINI_BASE_URL)

⚠️ 各 provider 依公开 API 文档实现,端点/模型名可用 env 覆盖以适配实际参数。

执行前先跑 model_registry.py configured --group voice --env-file .env。只有一个可用时显式选择;多个可用且用户没点名时,列出 provider/模型询问本次使用哪个,不按 VOICE_PROVIDER 擅自选择。

输入

字段必填说明
语音样本克隆时必填本人清晰无噪的语音(一般 10s-1min,具体看 provider 要求)
文案合成时必填要用克隆音色说出来的文字

输出(outputs/主题名/)

  • 合成的语音音频(mp3/wav)

执行步骤

脚本路径(相对项目根):skills/shared/scripts/voice_clone.py(各子命令支持 -h)。

1. 登记音色(enroll,得到 voice_id)
bash
# minimax:上传样本文件
python skills/shared/scripts/voice_clone.py enroll --provider minimax \
  --sample me.mp3 --name my_voice
# dashscope:用公网可访问的样本 URL
python skills/shared/scripts/voice_clone.py enroll --provider dashscope \
  --sample-url https://.../me.wav --name myv

(fish-audio 用已有 model_id 或内联参考音频,openai-compatible 用预置 voice 名,无需 enroll。)

2. 合成(clone)
bash
python skills/shared/scripts/voice_clone.py clone --provider minimax \
  --voice-id my_voice --text "大家好,欢迎来到我的频道" --speed 1.0 \
  -o outputs/主题名/vo.mp3

fish-audio 也可直接给参考音频:--sample ref.mp3 --sample-text "参考音频的文字"。

合规红线(必须遵守)

  1. 只能克隆你有权使用的声音(本人,或已获明确授权的人)。
  2. 不得克隆他人/名人声音用于误导、诈骗、冒充、伪造。
  3. 合成内容不得用于虚假信息或侵权用途。 —— 越线不做,并向用户说明。

规则

  1. 样本质量决定克隆效果:清晰、无背景噪、语气自然、时长足够。
  2. 先 check 确认 key,再 enroll,再 clone。
  3. 合成语音可接 audio-mix 加 BGM、接 auto-subtitle 出字幕、接视频。
  4. 产物统一进 outputs/主题名/。

参考来源

主流声音克隆云服务:阿里 CosyVoice(声音复刻)、MiniMax(语音克隆 + T2A)、Fish Audio、 OpenAI 兼容 TTS。本地零样本克隆(GPT-SoVITS/CosyVoice 本地)需 GPU,故走云端 provider。

© ZJU-REAL, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/openclaw/voice-clone of ZJU-REAL/Easel.

Open the folder on GitHubat commit 278f420

Compare with similar skills

Voice Clone next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Clone compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Clone this skillZJU-REAL/Easel3.4k—~715Automated safety check: NotesApache-2.0
Voice Clone Ttsnpc-live/clawfirm156—~1.1kAutomated safety check: PassNone
KrillinAI CLI Operatorkrillinai/OpenCreator13k—~869Automated safety check: PassApache-2.0
Web Video PresentationConardLi/garden-skills13k—~3.5kAutomated safety check: PassMIT
Video Translatorshang-zhu/violin1.1k—~1kAutomated safety check: NotesMIT
9Router Text to Speechdecolua/9router31k—~765Automated safety check: PassMIT

Similar skills

  • Voice Clone Tts

    npc-live/clawfirm

    声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。

    156 GitHub stars~1.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • KrillinAI CLI Operator

    krillinai/OpenCreator

    Routes agents to the right KrillinAI command for subtitles, dubbing, video rendering, covers and speech, and explains how to read its JSON and manifest output.

    13k GitHub stars~869 tokensUpdated today
    Media & CreativeAuto-check passed
  • Web Video Presentation

    ConardLi/garden-skills

    Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.

    13k GitHub stars~3.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Video Translator

    shang-zhu/violin

    Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

    1.1k GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • 9Router Text to Speech

    decolua/9router

    Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers.

    31k GitHub stars~765 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes

More from ZJU-REAL/Easel

All 114 skills in this repo
  • Gzh Design

    ZJU-REAL/Easel

    微信公众号文章排版引擎:把 Markdown / Word(.docx) / PDF / 纯文本转成可直接粘贴进公众号编辑器的 HTML,自动章节编号、关键词标记、引言卡、目录、代码块、图片/GIF、作者签名;主题从 references/theme-index.md…

    3.4k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • 微信公众号文章自动创作与发布工具。给定参考文章、文字或文档,自动搜索整理全网相关信息、生成图文并茂的公众号文章,并发布到微信公众号草稿箱。特别强调反 AI 检测写作。

    3.4k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Card Design

    ZJU-REAL/Easel

    社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。

    3.4k GitHub stars~657 tokensUpdated yesterday
    Auto-check passed
  • Ecom Details Image

    ZJU-REAL/Easel

    生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Infographic

    ZJU-REAL/Easel

    将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

    3.4k GitHub stars~643 tokensUpdated yesterday
    Auto-check passed
  • Novel Writer

    ZJU-REAL/Easel

    长篇小说/网文连载创作:从世界观、人设和三级大纲写到逐章正文,并用文件化状态维护伏笔、前情和跨章一致性. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Voice Clone

What does Voice Clone do?

上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。. Voice Clone is an agent skill from ZJU-REAL/Easel.

When should I use Voice Clone?

Voice Clone fits situations like: tasks that involve Text to speech and voice.

How do I install Voice Clone in Claude Code?

Run `npx skills add ZJU-REAL/Easel --skill voice-clone -a claude-code`. Or copy the skill folder (skills/openclaw/voice-clone in ZJU-REAL/Easel) into .claude/skills/voice-clone in your project. Claude Code loads it when a task matches its description.

How do I install Voice Clone in Codex?

Run `npx skills add ZJU-REAL/Easel --skill voice-clone -a codex`. Or copy the skill folder (skills/openclaw/voice-clone in ZJU-REAL/Easel) into .agents/skills/voice-clone in your project. Codex loads it when a task matches its description.

Can I use Voice Clone in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJU-REAL/Easel --skill voice-clone -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-clone, .gemini/skills/voice-clone, .github/skills/voice-clone and .opencode/skills/voice-clone in your project.

What does Voice Clone need to run?

Going by SKILL.md and its folder, Voice Clone needs the command-line tools its instructions call (python) and credentials named DASHSCOPE_API_KEY, MINIMAX_API_KEY, FISH_API_KEY and VOICE_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY; A credential in MINIMAX_API_KEY.

Does Voice Clone access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voice Clone safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Voice Clone use?

Voice Clone is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice Clone use?

About 715 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Clone?

Skills that share tags, products or a category with Voice Clone: Voice Clone Tts (npc-live/clawfirm, 156 stars), KrillinAI CLI Operator (krillinai/OpenCreator, 13k stars), Web Video Presentation (ConardLi/garden-skills, 13k stars) and Video Translator (shang-zhu/violin, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Clone?

ZJU-REAL (a GitHub organization) maintains it in ZJU-REAL/Easel, which has 3,376 GitHub stars. The repository holds 114 skills in this directory. The repository was last updated on October 9, 2026.

Source: ZJU-REAL/Easel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.