Agent skill

Doubao Text-to-Speech Generator

by xvirobotics in xvirobotics/metabot

Generates Chinese speech audio from text using Volcengine's Doubao TTS API, for narration, voiceovers or multi-voice podcasts.

MITAuto-check: notesMedia & Creative

Install Doubao Text-to-Speech Generator

skills CLI
$ npx skills add xvirobotics/metabot --skill doubao-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xvirobotics/metabot doubao-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xvirobotics/metabot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/doubao-tts .claude/skills/doubao-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doubao-tts
GitHub stars
994
Token cost
~1.1k tokens
SKILL.md length
220 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Generates Chinese speech audio from text using Volcengine's Doubao TTS API, for narration, voiceovers or multi-voice podcasts.

  • Works in 4 steps: Write the script — Create the podcast… → Generate audio — Use bin/doubao-tts -f… → Copy to outputs — cp podcast.mp3… → …
  • Generating narration or a voiceover from written text
  • SKILL.md covers When to Use, Quick Usage, Available Voices (已验证可用) and API Details, plus 3 more sections
  • Calls ffmpeg, curl and python3; reaches openspeech.bytedance.com; needs VOLCENGINE_TTS_ACCESS_KEY

What it does

This skill wraps Volcengine's Doubao TTS API behind a local CLI tool, supporting both a real-time path for short text under 300 characters and an asynchronous long-form path for up to 100K characters, where text is submitted and then polled for a result that returns an audio URL valid for an hour.

It ships a list of verified Chinese female and male voice IDs with short descriptions, and reads its Volcengine app ID, access key, and optional resource ID from environment variables already configured in the host app. For a podcast, it has you write the script first, generate audio per speaker with the matching voice ID for a multi-voice result, then stitch the segments together with ffmpeg.

Output is copied to a per-chat outputs folder so it can be delivered back to the user; voice families beyond the Chinese set listed need a different resource ID enabled in the Volcengine console first.

When your agent uses it

  • Generating narration or a voiceover from written text
  • Turning a written script into a multi-voice podcast episode
  • Reading a block of text aloud as an audio file

Example prompts

  • “Turn this script into a podcast episode with two Chinese voices.”
  • “Generate a voiceover for this product demo in a male voice.”
  • “Read this paragraph aloud and save it as an MP3.”

Requirements

  • Volcengine Doubao TTS API credentials
  • ffmpeg (for multi-voice podcasts)

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Write the script — Create the podcast script as markdown or plain text
  2. Generate audio — Use bin/doubao-tts -f script.txt -v zh_male_aojiaobazong_moon_bigtts -o podcast.mp3
  3. Copy to outputs — cp podcast.mp3 /tmp/metabot-outputs// to send to user
  4. For multi-voice podcasts, generate each speaker's segments separately, then concatenate with ffmpeg

What it can do on your machine

Read from SKILL.md and the folder at commit 916637e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffmpeg
    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openspeech.bytedance.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOLCENGINE_TTS_ACCESS_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doubao Text-to-Speech Generator loads about 1.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 220 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:60
    Variables (already configured in MetaBot .env)

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xvirobotics/metabot at commit 916637e, republished under its MIT licence (© xvirobotics). 220 words, ~1,057 tokens.

Download SKILL.mdSave it as .claude/skills/doubao-tts/SKILL.md (or your agent's skills folder).
name
doubao-tts
description
Generate high-quality speech audio using Doubao (豆包/Volcengine) TTS API. Use this skill when the user asks to generate audio, podcasts, voiceovers, or text-to-speech output.

Doubao TTS — 豆包语音合成

Generate high-quality speech audio from text using Volcengine's Doubao TTS API. Supports short-form (real-time) and long-form (async, up to 100K characters) synthesis.

When to Use

  • User asks to generate audio, podcasts, voiceovers, or narration
  • User wants text-to-speech for any content
  • User asks to "read this aloud" or "make an audio version"

Quick Usage

Use the doubao-tts CLI tool (installed at bin/doubao-tts):

bash
# Short text (real-time, < 300 chars)
bin/doubao-tts "你好世界" -o output.mp3

# Long text from file (async mode, up to 100K chars)
bin/doubao-tts -f article.txt -o podcast.mp3

# Pipe content
echo "Hello world" | bin/doubao-tts -o hello.mp3

# Choose voice
bin/doubao-tts "你好" -v zh_male_aojiaobazong_moon_bigtts -o output.mp3

# Adjust speed/volume/pitch
bin/doubao-tts "你好" --speed 1.2 --volume 1.5 -o output.mp3

Available Voices (已验证可用)

Chinese Female
Voice IDDescription
zh_female_sajiaonvyou_moon_bigtts撒娇女友 (default)
zh_female_gaolengyujie_moon_bigtts高冷御姐
zh_female_tianmeixiaoyuan_moon_bigtts甜美校园
zh_female_yuanqinvyou_moon_bigtts元气女友
zh_female_wanwanxiaohe_moon_bigtts弯弯小何
zh_female_linjianvhai_moon_bigtts邻家女孩
Chinese Male
Voice IDDescription
zh_male_aojiaobazong_moon_bigtts傲娇霸总
zh_male_jingqiangkanye_moon_bigtts京腔侃爷
zh_male_wennuanahu_moon_bigtts温暖阿虎
zh_male_yangguangqingnian_moon_bigtts阳光青年

Note: 其他音色 (BV系列, mars后缀) 需要不同的 resource ID。如需更多音色,请在火山引擎控制台开通对应资源。

API Details

Environment Variables (already configured in MetaBot .env)
VOLCENGINE_TTS_APPID=<app_id>
VOLCENGINE_TTS_ACCESS_KEY=<access_key>
VOLCENGINE_TTS_RESOURCE_ID=volc.service_type.10029  (optional)
Short-form API (real-time, < 300 chars)
  • Endpoint: https://openspeech.bytedance.com/api/v3/tts/unidirectional
  • Response: chunked JSON with base64 audio in data field
  • Latency: < 1 second
Long-form API (async, up to 100K chars)
  • Submit: POST https://openspeech.bytedance.com/api/v1/tts_async/submit
  • Query: GET https://openspeech.bytedance.com/api/v1/tts_async/query?appid=X&task_id=Y
  • Response: audio_url (valid for 1 hour)
  • Latency: seconds to minutes depending on text length

Workflow for Podcasts

  1. Write the script — Create the podcast script as markdown or plain text
  2. Generate audio — Use bin/doubao-tts -f script.txt -v zh_male_aojiaobazong_moon_bigtts -o podcast.mp3
  3. Copy to outputs — cp podcast.mp3 /tmp/metabot-outputs/<chatId>/ to send to user
  4. For multi-voice podcasts, generate each speaker's segments separately, then concatenate with ffmpeg

Multi-Voice Podcast Example

bash
# Generate segments for different speakers
bin/doubao-tts -f host_lines.txt -v zh_male_aojiaobazong_moon_bigtts -o host.mp3
bin/doubao-tts -f guest_lines.txt -v zh_female_gaolengyujie_moon_bigtts -o guest.mp3

# Concatenate (requires ffmpeg)
echo "file 'host.mp3'" > list.txt
echo "file 'guest.mp3'" >> list.txt
ffmpeg -f concat -safe 0 -i list.txt -c copy podcast.mp3

Raw curl (if CLI not available)

bash
# Short-form
curl -X POST "https://openspeech.bytedance.com/api/v3/tts/unidirectional" \
  -H "Content-Type: application/json" \
  -H "X-Api-App-Id: $VOLCENGINE_TTS_APPID" \
  -H "X-Api-Access-Key: $VOLCENGINE_TTS_ACCESS_KEY" \
  -H "X-Api-Resource-Id: volc.service_type.10029" \
  -H "X-Api-Request-Id: $(uuidgen)" \
  -d '{
    "req_params": {
      "text": "你好世界",
      "speaker": "zh_female_sajiaonvyou_moon_bigtts",
      "audio_params": {"format": "mp3", "sample_rate": 24000}
    }
  }' | python3 -c "
import sys, json, base64
chunks = []
for line in sys.stdin:
    line = line.strip()
    if not line: continue
    try:
        d = json.loads(line)
        if d.get('data'): chunks.append(base64.b64decode(d['data']))
    except: pass
sys.stdout.buffer.write(b''.join(chunks))
" > output.mp3

© xvirobotics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/doubao-tts of xvirobotics/metabot.

Open the folder on GitHubat commit 916637e

Compare with similar skills

Doubao Text-to-Speech Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doubao Text-to-Speech Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doubao Text-to-Speech Generator this skillxvirobotics/metabot994—~1.1kAutomated safety check: NotesMIT
Podcast Generationbytedance/deer-flow84k2 repos~2.2kAutomated safety check: PassMIT
Video Podcast Makerdtsola/xiaoyaosearch1k—~3.4kAutomated safety check: PassMIT
Podcastzarazhangrui/personalized-podcast438—~2.3kAutomated safety check: NotesNone
Video Podcast Maker NanoAgents365-ai/video-podcast-maker1.7k—~3.6kAutomated safety check: PassMIT
Azure Realtime Podcast Generationmicrosoft/skills3.1k1 repos~947Automated safety check: PassMIT

Similar skills

  • Podcast Generation

    bytedance/deer-flow

    Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script.

    84k GitHub starsUsed in 2 repos~2.2k tokens
    Media & CreativeAuto-check passed
  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Podcast

    zarazhangrui/personalized-podcast

    Generate a podcast episode from content you provide. An agent skill from zarazhangrui/personalized-podcast.

    438 GitHub stars~2.3k tokensUpdated 6 mo ago
    Media & CreativeAuto-check: notes
  • Video Podcast Maker Nano

    Agents365-ai/video-podcast-maker

    Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…

    1.7k GitHub stars~3.6k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Media & CreativeAuto-check passed
  • Z Qwen Audio Studio

    tjxj/z-skills

    A skill your agent uses when creating complete generated audio with qwen-audio-3.1-tts-next, including podcasts, radio drama, advertisements, multiple speakers, reference voices, ambience, sound…

    548 GitHub stars~716 tokensUpdated 19 days ago
    Media & CreativeAuto-check passed

More from xvirobotics/metabot

  • MetaBot Agent Teams CLI

    xvirobotics/metabot

    Documents the metabot teams command surface for creating durable teams, spawning teammates, dispatching tasks and inspecting their runs across engine sessions.

    994 GitHub stars~693 tokensUpdated 25 days ago
    Auto-check passed
  • MetaBot Scheduler

    xvirobotics/metabot

    Schedules persistent cron and one-shot tasks on the MetaBot server that survive Claude session restarts and can be managed by other bots.

    994 GitHub stars~1k tokensUpdated 25 days ago
    Auto-check passed
  • MetaBot CLI

    xvirobotics/metabot

    Documents the unified `metabot` CLI for personal memory, the skill hub, durable agent messaging, the agent registry, T5T status and scheduling.

    994 GitHub stars~573 tokensUpdated 25 days ago
    Auto-check passed
  • Agent Team And Skill Builder

    xvirobotics/metabot

    Decides whether a request wants a full agent team, one agent, or one skill, then follows the matching flow file to generate it.

    994 GitHub stars~628 tokensUpdated 25 days ago
    Auto-check: notes

Questions about Doubao Text-to-Speech Generator

What does Doubao Text-to-Speech Generator do?

Generates Chinese speech audio from text using Volcengine's Doubao TTS API, for narration, voiceovers or multi-voice podcasts. This skill wraps Volcengine's Doubao TTS API behind a local CLI tool, supporting both a real-time path for short text under 300 characters and an asynchronous long-form path for up to 100K characters, where text is submitted and then polled for a result that returns an audio URL valid for an hour.

When should I use Doubao Text-to-Speech Generator?

Doubao Text-to-Speech Generator fits situations like: generating narration or a voiceover from written text; turning a written script into a multi-voice podcast episode; reading a block of text aloud as an audio file.

How do I install Doubao Text-to-Speech Generator in Claude Code?

Run `npx skills add xvirobotics/metabot --skill doubao-tts -a claude-code`. Or copy the skill folder (.claude/skills/doubao-tts in xvirobotics/metabot) into .claude/skills/doubao-tts in your project. Claude Code loads it when a task matches its description.

How do I install Doubao Text-to-Speech Generator in Codex?

Run `npx skills add xvirobotics/metabot --skill doubao-tts -a codex`. Or copy the skill folder (.claude/skills/doubao-tts in xvirobotics/metabot) into .agents/skills/doubao-tts in your project. Codex loads it when a task matches its description.

Can I use Doubao Text-to-Speech Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xvirobotics/metabot --skill doubao-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doubao-tts, .gemini/skills/doubao-tts, .github/skills/doubao-tts and .opencode/skills/doubao-tts in your project.

What does Doubao Text-to-Speech Generator need to run?

Going by SKILL.md and its folder, Doubao Text-to-Speech Generator needs the command-line tools its instructions call (ffmpeg, curl and python3) and credentials named VOLCENGINE_TTS_ACCESS_KEY. Our summary lists: Volcengine Doubao TTS API credentials; ffmpeg (for multi-voice podcasts).

Does Doubao Text-to-Speech Generator access the network?

SKILL.md names 1 domain. In commands or code: openspeech.bytedance.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Doubao Text-to-Speech Generator safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Doubao Text-to-Speech Generator use?

Doubao Text-to-Speech Generator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doubao Text-to-Speech Generator use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doubao Text-to-Speech Generator?

Skills that share tags, products or a category with Doubao Text-to-Speech Generator: Podcast Generation (bytedance/deer-flow, 84k stars), Video Podcast Maker (dtsola/xiaoyaosearch, 1k stars), Podcast (zarazhangrui/personalized-podcast, 438 stars) and Video Podcast Maker Nano (Agents365-ai/video-podcast-maker, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doubao Text-to-Speech Generator?

xvirobotics (a GitHub organization) maintains it in xvirobotics/metabot, which has 994 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 15, 2026.

Source: xvirobotics/metabot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.