Agent skill

Doubao Tts

by OpenMinis in OpenMinis/MinisSkills

Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files.

MITAuto-check passedMedia & Creative

Install Doubao Tts

skills CLI
$ npx skills add OpenMinis/MinisSkills --skill doubao-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenMinis/MinisSkills doubao-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenMinis/MinisSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/doubao-tts .claude/skills/doubao-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doubao-tts
GitHub stars
446
Token cost
~1.8k tokens
SKILL.md length
546 words
Files
2 (incl. scripts)
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files.

  • Works in 4 steps: Log in to the Volcengine Console → Go to Doubao Voice → Text-to-Speech… → Create an application or use an existing… → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Get an API Key (Recommended,…, Environment Variables, Usage and API Description, plus 4 more sections
  • Runs Python scripts from its folder; calls uv; reaches openspeech.bytedance.com; needs DOUBAO_TTS_API_KEY and DOUBAO_TTS_TOKEN

What it does

Doubao Tts is an agent skill from OpenMinis/MinisSkills. Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files. When a user mentions "Doubao TTS," "Doubao Text-to-Speech," "doubao tts," "Volcengine TTS," "volcengine tts," "speech synthesis," "text-to-speech," "TTS," "generate audio," "read text aloud," or any scenario requiring the Doubao/Volcengine Text-to-Speech API, this skill must be triggered.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/tts.py`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Skills collection for Minis. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “Doubao TTS,”
  • “Doubao Text-to-Speech,”
  • “doubao tts,”
  • “/doubao-tts”

Requirements

  • Python 3
  • A credential in DOUBAO_TTS_API_KEY
  • A credential in DOUBAO_TTS_TOKEN

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Log in to the Volcengine Console
  2. Go to Doubao Voice → Text-to-Speech Large Model → Application Management
  3. Create an application or use an existing one
  4. On the API Key Management page, get the API Key → corresponding to DOUBAO_TTS_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit ae8c5db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openspeech.bytedance.com

    Also links to:

    • console.volcengine.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DOUBAO_TTS_API_KEY
    • DOUBAO_TTS_TOKEN
    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doubao Tts loads about 1.8k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from OpenMinis/MinisSkills at commit ae8c5db, republished under its MIT licence (© OpenMinis). 546 words, ~1,791 tokens.

Download SKILL.mdSave it as .claude/skills/doubao-tts/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
doubao-tts
description
Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files. When a user mentions "Doubao TTS," "Doubao Text-to-Speech," "doubao tts," "Volcengine TTS," "volcengine tts," "speech synthesis," "text-to-speech," "TTS," "generate audio," "read text aloud," or any scenario requiring the Doubao/Volcengine Text-to-Speech API, this skill must be triggered.
version
2.1.0

Doubao TTS Skill (V3)

Use the Volcengine Doubao Text-to-Speech V3 HTTP SSE unidirectional streaming API to convert text into audio files.

  1. Log in to the Volcengine Console
  2. Go to Doubao Voice → Text-to-Speech Large Model → Application Management
  3. Create an application or use an existing one
  4. On the API Key Management page, get the API Key → corresponding to DOUBAO_TTS_API_KEY

If you have not enabled the service yet, enable it first on the Text-to-Speech Large Model page.

Legacy Console (AppID + Token)

At the bottom of the application details page in the legacy console, you can find:

  • APP ID → DOUBAO_TTS_APPID
  • Access Token → DOUBAO_TTS_TOKEN

Environment Variables

Variable NameDescriptionRecommended
DOUBAO_TTS_API_KEYAPI Key (new console, X-Api-Key)✅
DOUBAO_TTS_APPIDAppID (legacy console, X-Api-App-Id)
DOUBAO_TTS_TOKENAccess Token (legacy console, X-Api-Access-Key)
DOUBAO_TTS_RESOURCE_IDResource ID; leave blank to use the default seed-tts-2.0

Check whether they are configured:

sh
[ -n "$DOUBAO_TTS_API_KEY" ] && echo "API_KEY: set" || echo "API_KEY: not set"
[ -n "$DOUBAO_TTS_APPID" ] && echo "APPID: set" || echo "APPID: not set"
[ -n "$DOUBAO_TTS_TOKEN" ] && echo "TOKEN: set" || echo "TOKEN: not set"

If they are not configured, tell the user to set them (prefer API Key): Set DOUBAO_TTS_API_KEY | Set DOUBAO_TTS_RESOURCE_ID

Legacy console (AppID + Token): Set DOUBAO_TTS_APPID | Set DOUBAO_TTS_TOKEN

Usage

Call the script: /var/minis/skills/doubao-tts/scripts/tts.py

sh
# Basic usage
uv run --script --cache-dir /root/.cache/uv \
  /var/minis/skills/doubao-tts/scripts/tts.py \
  --text "Hello, and welcome to Doubao Text-to-Speech." \
  --output /var/minis/workspace/output.mp3

# Specify voice and speech rate
uv run --script --cache-dir /root/.cache/uv \
  /var/minis/skills/doubao-tts/scripts/tts.py \
  --text "The weather is really nice today." \
  --speaker zh_female_cancan_uranus_bigtts \
  --speech-rate 10 \
  --output /var/minis/workspace/output.mp3

# English
uv run --script --cache-dir /root/.cache/uv \
  /var/minis/skills/doubao-tts/scripts/tts.py \
  --text "Hello! Nice to meet you." \
  --speaker en_female_dacey_uranus_bigtts \
  --output /var/minis/workspace/output.mp3

API Description

  • Endpoint: https://openspeech.bytedance.com/api/v3/tts/unidirectional/sse (SSE streaming)
  • Authentication (choose one):
    • New console: Header X-Api-Key (API Key)
    • Legacy console: Header X-Api-App-Id + X-Api-Access-Key (AppID + Token)
  • Resource ID: Specifies the model version to call (see the table below)
  • Usage response: The script includes X-Control-Require-Usage-Tokens-Return: text_words by default, which returns the number of billable characters (text_words) when synthesis ends
Resource IDDescription
seed-tts-1.0Doubao Text-to-Speech Model 1.0 character version (default, compatible with all BV*_streaming voices)
seed-tts-1.0-concurrDoubao Text-to-Speech Model 1.0 concurrency version
seed-tts-2.0Doubao Text-to-Speech Model 2.0 (supports only 2.0 voices)

Parameters

ParameterDescription
--textText to synthesize (required)
--outputOutput file path (required)
--api-keyAPI Key (new console, takes precedence over APPID/TOKEN)
--appidAppID (legacy console)
--tokenAccess Token (legacy console)
--speakerVoice, default zh_female_shuangkuaisisi_uranus_bigtts (Shuangkuai Sisi 2.0)
--encodingFormat: mp3/pcm/ogg_opus, default mp3
--speech-rateSpeech rate [-50, 100], where 0 is the default and 100 is 2x speed
--loudnessVolume [-50, 100], where 0 is the default
--sample-rateSample rate, default 24000
--emotionEmotion, such as happy/sad/angry/narrator
--emotion-scaleEmotion intensity [1, 5] (used with --emotion)
--resource-idResource ID (overrides the environment variable)
--jsonOutput result in JSON format
Show full SKILL.md (183 more words)Show less

Quick Reference for Common Voices

speakerNameScenario
zh_female_shuangkuaisisi_uranus_bigttsShuangkuai Sisi 2.0 ⭐ DefaultGeneral
zh_female_cancan_uranus_bigttsZhixing Cancan 2.0Role-playing
zh_female_tianmeixiaoyuan_uranus_bigttsTianmei Xiaoyuan 2.0General
zh_female_vv_uranus_bigttsVivi 2.0General, Chinese/Japanese/Indonesian/Mexican Spanish, Sichuan/Shaanxi/Northeastern dialects
zh_female_xiaohe_uranus_bigttsXiaohe 2.0General
zh_male_m191_uranus_bigttsYunzhou 2.0General
zh_male_taocheng_uranus_bigttsXiaotian 2.0General
zh_female_kefunvsheng_uranus_bigttsNuanyang Female Voice 2.0Customer service
en_female_dacey_uranus_bigttsDaceyMultilingual (English)
en_male_tim_uranus_bigttsTimMultilingual (English)
Doubao Text-to-Speech Model 1.0 (seed-tts-1.0, requires changing --resource-id)
speakerNameScenario
BV700_streamingCancanGeneral, supports 22 emotions
BV001_streamingGeneral Female VoiceGeneral
BV002_streamingGeneral Male VoiceGeneral
BV701_streamingQingcangAudiobook
BV503_streamingEnergetic Female Voice-ArianaEnglish

⚠️ 1.0 and 2.0 voices cannot be mixed. seed-tts-2.0 supports only voices ending in *_uranus_bigtts.

Common Emotion Values

pleased (pleased) / sorry (sorry) / happy (happy) / sad (sad) / angry (angry) / scare (scared) / surprise (surprised) / hate (disgust) / tear (tearful voice) / narrator (narrator) / storytelling (storytelling)

Complete Workflow

  1. Check whether environment variables are configured (first DOUBAO_TTS_API_KEY, then DOUBAO_TTS_APPID + DOUBAO_TTS_TOKEN)
  2. Call the tts.py script to generate an audio file in /var/minis/workspace/
  3. Return it to the user as a minis://workspace/xxx.mp3 link, which they can click to play directly.

© OpenMinis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in doubao-tts of OpenMinis/MinisSkills.

  • SKILL.md
  • scripts/tts.py

Open the folder on GitHubat commit ae8c5db

Compare with similar skills

Doubao Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doubao Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doubao Tts this skillOpenMinis/MinisSkills446—~1.8kAutomated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from OpenMinis/MinisSkills

All 49 skills in this repo
  • Android UI Automation

    OpenMinis/MinisSkills

    Automate Android apps that have no public API or web version by driving the UI layer through the Accessibility Service.

    446 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Codex On Ish

    OpenMinis/MinisSkills

    Install and run OpenAI Codex CLI inside the Minis/iSH Alpine sandbox on iOS, where rustls TLS and async sockets are broken.

    446 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Evidence Chain Builder

    OpenMinis/MinisSkills

    Score a claim against the evidence behind it. An agent skill from OpenMinis/MinisSkills.

    446 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Hyperframes CLI

    OpenMinis/MinisSkills

    HyperFrames CLI and Minis rendering. An agent skill from OpenMinis/MinisSkills.

    446 GitHub stars~3.5k tokensUpdated 2 days ago
    Auto-check passed
  • Whenpeak

    OpenMinis/MinisSkills

    Predict when a person's brain works best from their sleep, using the WhenPeak performance-intelligence API, and turn it into concrete scheduling advice.

    446 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Openstreetmap Marker

    OpenMinis/MinisSkills

    Generate a mobile-first, immersive, Amap(Gaode)-style custom landmark marker HTML map for travel itinerary planning, place showcasing, location sharing, etc.

    446 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Doubao Tts

What does Doubao Tts do?

Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files. Doubao Tts is an agent skill from OpenMinis/MinisSkills. Use Doubao Text-to-Speech (Volcengine TTS) to convert text into audio files.

When should I use Doubao Tts?

Doubao Tts fits situations like: tasks that involve Text to speech and voice.

How do I install Doubao Tts in Claude Code?

Run `npx skills add OpenMinis/MinisSkills --skill doubao-tts -a claude-code`. Or copy the skill folder (doubao-tts in OpenMinis/MinisSkills) into .claude/skills/doubao-tts in your project. Claude Code loads it when a task matches its description.

How do I install Doubao Tts in Codex?

Run `npx skills add OpenMinis/MinisSkills --skill doubao-tts -a codex`. Or copy the skill folder (doubao-tts in OpenMinis/MinisSkills) into .agents/skills/doubao-tts in your project. Codex loads it when a task matches its description.

Can I use Doubao Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenMinis/MinisSkills --skill doubao-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doubao-tts, .gemini/skills/doubao-tts, .github/skills/doubao-tts and .opencode/skills/doubao-tts in your project.

What does Doubao Tts need to run?

Going by SKILL.md and its folder, Doubao Tts needs Python for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named DOUBAO_TTS_API_KEY, DOUBAO_TTS_TOKEN and API_KEY. Our summary lists: Python 3; A credential in DOUBAO_TTS_API_KEY; A credential in DOUBAO_TTS_TOKEN.

Does Doubao Tts access the network?

SKILL.md names 2 domains. In commands or code: openspeech.bytedance.com; the agent is likely to contact it when it follows the instructions. As links in the text: console.volcengine.com. This is read from the text; nothing was executed.

Is Doubao Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Doubao Tts use?

Doubao Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doubao Tts use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doubao Tts?

Skills that share tags, products or a category with Doubao Tts: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doubao Tts?

OpenMinis (a GitHub organization) maintains it in OpenMinis/MinisSkills, which has 446 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 7, 2026.

Source: OpenMinis/MinisSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.