Agent skill

Byteplus Seed Speech Tts

by calesthio in calesthio/generative-media-skills

Production guidance for international BytePlus Seed Speech text-to-speech.

MITAuto-check passedMedia & Creative

Install Byteplus Seed Speech Tts

skills CLI
$ npx skills add calesthio/generative-media-skills --skill byteplus-seed-speech-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills byteplus-seed-speech-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/text-to-speech/byteplus-seed-speech-tts .claude/skills/byteplus-seed-speech-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
byteplus-seed-speech-tts
GitHub stars
197
Token cost
~3.4k tokens
SKILL.md length
1,496 words
Files
2
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Production guidance for international BytePlus Seed Speech text-to-speech.

  • Selecting TTS 1.0 versus 2.0
  • SKILL.md covers Verification and evidence, Boundary and taxonomy, Select model family and protocol and Authentication and headers, plus 10 more sections
  • Reaches voice.ap-southeast-1.bytepluses.com; needs BYTEPLUS_API_KEY
  • Unidirectional streaming

What it does

Byteplus Seed Speech Tts is an agent skill from calesthio/generative-media-skills. Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling, and output QA. Do not use for mainland-China Volcengine Doubao Speech endpoints.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Selecting TTS 1.0 versus 2.0
  • Unidirectional streaming
  • Current voices and languages
  • Prompt/prosody controls

Example prompts

  • “/byteplus-seed-speech-tts”

Requirements

  • A credential in BYTEPLUS_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are http and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • voice.ap-southeast-1.bytepluses.com

    Also links to:

    • docs.byteplus.com
    • byteplus.com
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BYTEPLUS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Byteplus Seed Speech Tts loads about 3.4k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,496 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,496 words, ~3,392 tokens.

Download SKILL.mdSave it as .claude/skills/byteplus-seed-speech-tts/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
byteplus-seed-speech-tts
description
Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling, and output QA. Do not use for mainland-China Volcengine Doubao Speech endpoints.

BytePlus Seed Speech TTS

Use this skill for the international BytePlus Seed Speech service. The current documented API host is in the ap-southeast-1.bytepluses.com domain. Do not transfer credentials, endpoints, prices, voice IDs, or commercial assumptions from mainland-China Volcengine Doubao Speech.

Verification and evidence

  • Documented fact: behavior stated in official BytePlus documentation or its applicable legal pages.
  • Production heuristic: a practical workflow that must be tested with the selected account, voice, and protocol.
  • Empirical observation: output measured from an actual dated request.

Provider facts were verified 2026-07-12. Voice inventory, language compatibility, pricing, free quota, concurrency, endpoints, retention, and protocol behavior are volatile. Re-check the official voice list, API page, console, pricing, order form, and DPA immediately before production.

Boundary and taxonomy

BytePlus calls the international product Seed Speech. Current synthesis resource families include:

  • seed-tts-1.0: TTS 1.0;
  • seed-tts-2.0: TTS 2.0;
  • seed-icl-1.0: authorized Voice Replication 1.0 synthesis;
  • seed-icl-2.0: authorized Voice Replication 2.0 synthesis.

Seed-TTS research papers describe model research and are not API contracts. Do not expose research variants as hosted model IDs.

This leaf covers prebuilt TTS and synthesis with a voice already enrolled and authorized. Voice enrollment itself requires a separate consent, data, and lifecycle review. It does not cover ASR, speech-to-speech agents, interpretation, podcasts, or mainland Volcengine APIs.

Select model family and protocol

TTS 1.0

Use when a chosen voice or required feature is documented for TTS 1.0, especially where the selected protocol documents timestamp or voice-specific emotion support. Do not assume TTS 1.0 and 2.0 voices share identifiers or controls.

TTS 2.0

BytePlus documents broader multilingual and contextual/instruction control. context_texts can give natural-language performance direction; only the first list value is documented as effective. section_id can connect prior synthesis context, with current documentation describing up to 30 rounds or ten minutes.

TTS 2.0 does not support SSML in the documented V3 bidirectional or unidirectional paths. Never send SSML as though it were a supported control.

Bidirectional WebSocket

Use when text arrives incrementally and audio must stream while preserving session context:

text
wss://voice.ap-southeast-1.bytepluses.com/api/v3/tts/bidirection

Lifecycle: establish WebSocket, StartConnection, wait for ConnectionStarted, StartSession, wait for SessionStarted, send one or more TaskRequest events, consume sentence/audio events, FinishSession, wait for SessionFinished, then close. Event codes currently include sentence start 350, sentence end 351, and TTS response 352.

Unidirectional HTTP streaming

Use when the complete text is known and audio should stream back over HTTP:

text
POST https://voice.ap-southeast-1.bytepluses.com/api/v3/tts/unidirectional

Read newline/streamed JSON responses, base64-decode each audio payload, and concatenate in arrival order. Finish on provider success code 20000000.

No public BytePlus async long-text workflow was established during verification. Do not redirect international jobs to Volcengine async endpoints.

Authentication and headers

New integrations use a console-generated X-Api-Key. The selected resource ID is mandatory. Current HTTP documentation also requires the fixed X-Api-App-Key shown by the provider; verify it from the current page rather than assuming it is a secret or reusing a stale value.

Example header shape:

http
X-Api-Key: ${BYTEPLUS_API_KEY}
X-Api-Resource-Id: seed-tts-2.0
X-Api-App-Key: <current documented fixed value>
X-Api-Request-Id: <uuid>
Content-Type: application/json

The WebSocket path uses an X-Api-Connect-Id for troubleshooting. Preserve the returned X-Tt-Logid and client request/connection ID. Never log credentials or sensitive source text.

Legacy App ID/access-key headers remain documented for compatibility, with legacy resource IDs. Do not mix a new API key with guessed legacy fields or copy Volcengine auth.

Voice and language selection

Treat support as a tuple:

text
voice ID + resource family + protocol + language + control

Load the official voice list at execution time. Verify that the account is entitled to the voice and that the chosen voice supports the target language, protocol, emotion, timestamps, and other requested controls. Do not invent IDs or infer capability from an ID suffix.

The TTS 2.0 overview currently lists English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian, Brazilian Portuguese, Italian, and Korean. Protocol/voice pages may expose additional voice-specific language codes. The selected voice page is authoritative for a production call.

Request controls

Current V3 request fields include text, speaker, audio format/sample rate, speech rate, loudness, voice-specific emotion, and an escaped JSON additions string.

Documented ranges include:

  • speech_rate: -50 to 100, where -50 represents 0.5x and 100 represents 2x;
  • loudness_rate: -50 to 100;
  • emotion_scale: 1 to 5 for voices supporting emotion, with non-linear perceived change;
  • formats: mp3, ogg_opus, and pcm;
  • sample-rate values listed by the current protocol page from 8 kHz through 48 kHz.

Streaming WAV can produce repeated headers; prefer PCM when an unframed stream is needed, or use MP3/Opus with verified concatenation/decoding.

additions controls include language policy, Markdown/emoji handling, end silence, formula normalization, TTS 2.0 context, and optional cache behavior. Cache hits are documented as not returning timestamps.

Normalize numbers, dates, currencies, abbreviations, URLs, formulas, Markdown, emoji, and punctuation before estimating cost and synthesis. Save the exact normalized source used for the request.

Timestamps and subtitles

Current V3 parameter tables say enable_timestamp works only for TTS 1.0 and ICL 1.0. Cache responses omit timestamps. Treat any timing observed from another model/protocol combination as an empirical account-level result, not a documented capability.

Therefore:

  • never promise TTS 2.0 timing without a dated integration test of the exact voice/protocol;
  • verify timestamps are monotonic, bounded by decoded duration, and aligned after text normalization;
  • human-review names, numbers, punctuation, and line grouping;
  • retain a fallback transcription/alignment workflow;
  • do not construct captions from estimated character duration.

Billing and capacity

Official public pricing checked 2026-07-12 lists TTS 2.0 packages and pay-as-you-go character billing, a 20,000-character activation trial, and separately purchased concurrency. It defines spaces, punctuation, symbols, and carriage returns as billable characters.

Because prices, packages, and terminology can change, create a dated cost record from the current pricing page/console. Count normalized request text using the provider's current billing rule. Confirm whether the account limit is expressed as concurrency or QPS and test queue behavior before batch production.

Do not state that a plan or trial is guaranteed outside the applicable account and date.

Show full SKILL.md (556 more words)Show less

The applicable BytePlus agreement and DPA govern processing. The customer is responsible for notices, lawful basis/consent, and authorized instructions. BytePlus's current enterprise AI FAQ qualifies training claims: do not simplify it to “data is never used for training” without the corporate-customer and express-authorization context.

For any replicated voice:

  • obtain explicit, purpose-specific authority from the speaker/rights holder;
  • record permitted use, audience, territory, term, revocation, and deletion procedure;
  • verify account, voice ID, and language are authorized;
  • avoid public-figure or deceptive impersonation;
  • minimize enrollment/source audio and restrict access;
  • resolve retention, residency, deletion, and model-improvement terms in the current contract/order form.

A successful enrollment or voiceprint check is not legal consent.

Error handling

Handle provider/application codes and preserve X-Tt-Logid.

  • Invalid text, parameters, unsupported controls, or unauthorized voice: repair; do not retry unchanged.
  • Concurrency/quota: queue, back off with jitter, or adjust purchased capacity.
  • General server/network failures: bounded retry when idempotency/charge behavior is understood.
  • Ambiguous partial stream: reject incomplete audio unless the workflow explicitly supports resume/reassembly.

Do not print the API key, legacy token, sensitive text, or full response bodies containing customer data.

Production QA

Before a full batch, generate short samples and approve voice, language, pronunciation, rate, emotion, and context behavior.

Check:

  • decoded codec, sample rate, channel count, duration, and corruption;
  • missing/repeated/truncated words and stream ordering;
  • names, acronyms, numbers, currencies, formulas, and multilingual switching;
  • clipping, true peak, loudness consistency, noise, clicks, sibilance, and silence;
  • performance direction and voice identity;
  • back-transcription against normalized text;
  • timestamp/caption alignment where used;
  • request/resource/voice/log IDs, source revision, consent record, price evidence, and output hash.

Provider output can vary across calls. Do not promise bit-identical speech from the same request unless measured and contractually supported.

Example 1: TTS 2.0 HTTP narration

This is a complete example, not a mandatory formula.

Intent: low-latency Chinese product narration with calm performance direction.

Preflight the current voice list and choose an authorized TTS 2.0 Chinese voice. Use seed-tts-2.0, MP3 at 24 kHz, exact normalized script, and one context_texts instruction: “Speak naturally, calmly, and warmly; avoid exaggeration.” Submit to the unidirectional HTTP endpoint, consume every audio chunk in order, and stop on 20000000.

The payload shape is:

json
{
  "user": {"id": "production-job-2841"},
  "req_params": {
    "text": "欢迎回来。项目已经准备就绪,我们现在可以开始。",
    "speaker": "<authorized-current-voice-id>",
    "audio_params": {
      "format": "mp3",
      "sample_rate": 24000,
      "speech_rate": 0,
      "loudness_rate": 0
    },
    "additions": "{\"explicit_language\":\"zh-cn\",\"disable_markdown_filter\":true,\"context_texts\":[\"请用自然、沉稳、亲切的语气表达,不要夸张。\"]}"
  }
}

QA decoded audio, duration, script completeness, pronunciation, performance, and log ID. Do not request timestamps unless this exact voice/protocol has passed a dated test.

Example 2: incremental assistant response

This is a complete example, not a mandatory formula.

Intent: stream an authorized English assistant voice while response text arrives incrementally.

Use the bidirectional WebSocket lifecycle. Start connection/session with an authorized voice/resource pair, send complete semantic chunks as TaskRequest events, consume sentence/audio events, and finish only after the final chunk. Keep one performance instruction and context policy across the session. Buffer enough text to avoid unnatural fragments without delaying the first useful audio excessively.

QA chunk boundaries, missing/repeated text, ordering, cancellation, reconnection, rate limits, and final transcript/audio parity. If the selected voice does not support bidirectional operation, switch protocol or voice after explicit approval; do not silently substitute.

Sources

Official sources verified 2026-07-12:

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/providers/text-to-speech/byteplus-seed-speech-tts of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Byteplus Seed Speech Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Byteplus Seed Speech Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Byteplus Seed Speech Tts this skillcalesthio/generative-media-skills197—~3.4kAutomated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    197 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    197 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    197 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    197 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    197 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Generated Media QA

    calesthio/generative-media-skills

    Provider-independent quality assurance for AI-generated and AI-assisted media.

    197 GitHub stars~8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Byteplus Seed Speech Tts

What does Byteplus Seed Speech Tts do?

Production guidance for international BytePlus Seed Speech text-to-speech. Byteplus Seed Speech Tts is an agent skill from calesthio/generative-media-skills. Production guidance for international BytePlus Seed Speech text-to-speech.

When should I use Byteplus Seed Speech Tts?

Byteplus Seed Speech Tts fits situations like: selecting TTS 1.0 versus 2.0; unidirectional streaming; current voices and languages; prompt/prosody controls.

How do I install Byteplus Seed Speech Tts in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill byteplus-seed-speech-tts -a claude-code`. Or copy the skill folder (skills/providers/text-to-speech/byteplus-seed-speech-tts in calesthio/generative-media-skills) into .claude/skills/byteplus-seed-speech-tts in your project. Claude Code loads it when a task matches its description.

How do I install Byteplus Seed Speech Tts in Codex?

Run `npx skills add calesthio/generative-media-skills --skill byteplus-seed-speech-tts -a codex`. Or copy the skill folder (skills/providers/text-to-speech/byteplus-seed-speech-tts in calesthio/generative-media-skills) into .agents/skills/byteplus-seed-speech-tts in your project. Codex loads it when a task matches its description.

Can I use Byteplus Seed Speech Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill byteplus-seed-speech-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/byteplus-seed-speech-tts, .gemini/skills/byteplus-seed-speech-tts, .github/skills/byteplus-seed-speech-tts and .opencode/skills/byteplus-seed-speech-tts in your project.

What does Byteplus Seed Speech Tts need to run?

Going by SKILL.md and its folder, Byteplus Seed Speech Tts needs credentials named BYTEPLUS_API_KEY. Our summary lists: A credential in BYTEPLUS_API_KEY.

Does Byteplus Seed Speech Tts access the network?

SKILL.md names 4 domains. In commands or code: voice.ap-southeast-1.bytepluses.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.byteplus.com, byteplus.com and arxiv.org. This is read from the text; nothing was executed.

Is Byteplus Seed Speech Tts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Byteplus Seed Speech Tts use?

Byteplus Seed Speech Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Byteplus Seed Speech Tts use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Byteplus Seed Speech Tts?

Skills that share tags, products or a category with Byteplus Seed Speech Tts: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Byteplus Seed Speech Tts?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.