MoneyPrinterTurbo Video Generator
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
Production guidance for mainland-China Volcengine Doubao Speech text-to-speech.
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-tts --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .claude/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .claude/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-ttsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-tts --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .agents/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .agents/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-tts --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .cursor/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .cursor/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/providers/text-to-speech/volcengine-doubao-speech-tts--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-tts --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .gemini/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .gemini/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-ttsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .github/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .github/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills volcengine-doubao-speech-tts --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/providers/text-to-speech/volcengine-doubao-speech-tts .opencode/skills/volcengine-doubao-speech-tts && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "volcengine-doubao-speech-tts" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/volcengine-doubao-speech-tts into .opencode/skills/volcengine-doubao-speech-tts/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "volcengine-doubao-speech-tts", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
volcengine-doubao-speech-ttsProduction guidance for mainland-China Volcengine Doubao Speech text-to-speech.
Volcengine Doubao Speech Tts is an agent skill from calesthio/generative-media-skills. Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats, expiring results, quota/error handling, authorized cloned voices, privacy, and audio QA. Do not use for international BytePlus endpoints.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).
It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are http and json).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
openspeech.bytedance.comAlso links to:
docs.volcengine.comarxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VOLC_ACCESS_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Volcengine Doubao Speech Tts loads about 2.9k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,224 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,224 words, ~2,871 tokens.
.claude/skills/volcengine-doubao-speech-tts/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this skill for mainland-China Volcengine Doubao Speech. Do not transfer BytePlus API hosts, keys, pricing, legal terms, or assumptions into this workflow. The services share model-family language but are separate regional products.
Facts were verified 2026-07-12. Voice lists, protocol compatibility, pricing, activation, quotas, retention, regional availability, and legal terms are volatile. Re-check the Chinese official docs and authenticated console.
Current synthesis resource families include:
seed-tts-1.0: TTS 1.0 on V3;seed-tts-2.0: TTS 2.0 on V3;seed-icl-1.0: authorized cloned-voice 1.0 synthesis;seed-icl-2.0: authorized cloned-voice 2.0 synthesis.Legacy volc.service_type.10029 and the V1 API belong to older TTS 1.0 contracts. Do not call a TTS 2.0 voice through V1 or substitute legacy model selectors for a V3 resource ID.
This leaf covers synthesis and long-text production. Voice enrollment is included only as a consent/authorization boundary; ASR, speech agents, interpretation, podcasts, speech-to-speech, and international BytePlus are out of scope.
Use when text arrives incrementally:
wss://openspeech.bytedance.com/api/v3/tts/bidirectionFollow the documented connection/session/task lifecycle. Not every voice supports every protocol; verify the current voice list and account entitlement before locking bidirectional delivery.
Use when the complete text is known but audio should stream. Official Volcengine documentation includes V3 WebSocket, HTTP chunked, and SSE surfaces. Select the exact documented endpoint and response parser at implementation time; do not assume transport schemas are interchangeable.
Use for batch narration/audiobooks and inputs within the current documented limit:
POST https://openspeech.bytedance.com/api/v3/tts/submit
POST https://openspeech.bytedance.com/api/v3/tts/queryCurrent docs state a maximum of 100,000 characters. Submit returns a task_id; query statuses are 1 running, 2 success, and 3 failure. On success, download immediately. Server audio is documented as retained seven days; each returned URL is valid for one hour and can be refreshed by querying again.
Submit and query consume the purchased concurrency pool. Use bounded polling with jitter and persist the provider task ID.
Volcengine V3 streaming and async pages use different documented header combinations. Follow the exact selected page.
Current V3 new-console integrations use X-Api-Key on relevant streaming surfaces. The async long-text API documents:
X-Api-App-Id: ${VOLC_APP_ID}
X-Api-Access-Key: ${VOLC_ACCESS_KEY}
X-Api-Resource-Id: seed-tts-2.0
X-Api-Request-Id: <uuid>Legacy V1 uses the unusual Authorization: Bearer;${token} form. Do not copy that to V3 without current documentation.
Preserve request IDs and X-Tt-Logid where returned. Never log secrets, tokens, enrollment audio, or sensitive source text.
Treat support as:
speaker + resource ID + protocol + language + control + account entitlementLoad the current official voice list. Do not freeze counts, infer support from a suffix, or invent an ID. Some multilingual/current voices may be limited to unidirectional operation.
Documented controls include voice-specific emotion, non-linear emotion_scale 1-5, speech_rate -50 to 100, loudness_rate -50 to 100, sample rate/format, end silence, language filtering, Markdown/emoji/formula handling, and optional AIGC watermark settings.
TTS 2.0 uses contextual/instruction controls and does not support SSML on the documented V3 paths. SSML support belongs to compatible TTS 1.0 paths only; verify exact tags and limits before use.
Mixing voices or carrying controls across versions is not universal. The current voice/protocol page is authoritative.
Timestamp documentation varies by version/protocol. Volcengine async results can return sentence/word timing structures, but TTS 2.0/ICL 2.0 support and normalization behavior must be tested with the exact voice.
Set a stable caller-side unique_id and save it with task_id, source hash, normalized text, resource/speaker, and attempt number.
After an ambiguous timeout, query known task state before resubmitting. Blind resubmission can duplicate work or charges. On success, download before URL expiry, compute a checksum, verify decoding/duration, and move the asset into controlled storage. Result URLs are not durable delivery locations.
Public RMB price amounts could not be confirmed from the unauthenticated documentation available during verification; the authenticated console and order terms are authoritative. Async long text is documented as using the corresponding short-text TTS/voice-replication pricing rather than a separate premium.
Before paid use, record:
Do not promise a public free tier, region set, or fixed rate without current account evidence.
Voice and associated identity can be sensitive personal data. Public product docs do not replace the applicable contract, privacy terms, and legal review.
For cloned voices:
No technical enrollment or similarity check proves legal authorization. Do not claim unverified residency, retention, training exclusion, or regulatory compliance.
Current APIs use success code 20000000; preserve detailed provider code/message/log ID.
3: surface failure details; do not conceal or endlessly poll.Create a pronunciation set for names, brands, acronyms, numbers, dates, currencies, formulas, dialect, and domain terms. Generate approval samples before long jobs.
Verify:
This is a complete example, not a mandatory formula.
Intent: synthesize a mainland-China audiobook chapter with an authorized TTS 2.0 voice.
Preflight current voice compatibility, normalize the chapter, hash it, calculate current billed characters, and secure price/concurrency approval. Submit with async-required headers, stable unique_id, namespace: "BidirectionalTTS", seed-tts-2.0, MP3 24 kHz, and the exact authorized current speaker ID.
{
"user": {"uid": "audiobook-pipeline"},
"unique_id": "chapter-0047-v3-9b2f8c4a",
"namespace": "BidirectionalTTS",
"req_params": {
"text": "<normalized-chapter-text>",
"speaker": "<authorized-current-voice-id>",
"audio_params": {
"format": "mp3",
"sample_rate": 24000,
"speech_rate": 0,
"loudness_rate": 0,
"enable_timestamp": true
},
"additions": "{\"disable_markdown_filter\":true,\"explicit_language\":\"zh-cn\"}"
}
}Poll by task_id; on status 2, download immediately, store timing payload, checksum, decoded duration, and URL expiry. Human-review chapter and captions. If this exact TTS 2.0 voice does not return usable timing, align separately.
This is a complete example, not a mandatory formula.
Intent: stream Mandarin assistant speech as text arrives.
Select a current voice explicitly documented for bidirectional V3 and the matching resource. Use PCM or a tested streaming codec, stable connection/session IDs, and a bounded semantic text-buffer policy. Follow StartConnection/StartSession/TaskRequest/FinishSession events and cancel cleanly when upstream text is abandoned.
QA first-audio behavior, chunk ordering, sentence boundaries, reconnect/cancel, missing/repeated text, entitlement, concurrency, and final transcript parity. If the desired voice is unidirectional-only, request user approval for a voice or protocol change.
Official sources verified 2026-07-12:
© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/providers/text-to-speech/volcengine-doubao-speech-tts of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Volcengine Doubao Speech Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Volcengine Doubao Speech Tts this skillcalesthio/generative-media-skills | 197 | — | ~2.9k | Automated safety check: Pass | MIT | |
| MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo | 130k | — | ~2.1k | Automated safety check: Warn | MIT | |
| HyperFrames Media Useheygen-com/hyperframes | 60k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Openspec OnboardSAP/e-mobility-charging-stations-simulator | 227 | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Blog AudioAgriciDaniel/claude-blog | 2.3k | 1 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Musictadaspetra/loop | 296 | 2 repos | ~827 | Automated safety check: Pass | MIT |
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
SAP/e-mobility-charging-stations-simulator
Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.
AgriciDaniel/claude-blog
Generate audio narration of blog posts using Google Gemini TTS.
tadaspetra/loop
Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.
hoquanghai/Auto-Create-Video
Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
Categories
Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Volcengine Doubao Speech Tts is an agent skill from calesthio/generative-media-skills. Production guidance for mainland-China Volcengine Doubao Speech text-to-speech.
Volcengine Doubao Speech Tts fits situations like: TTS 1.0/2.0 selection; V3 bidirectional; unidirectional streaming; asynchronous long-text submit/query jobs.
Run `npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a claude-code`. Or copy the skill folder (skills/providers/text-to-speech/volcengine-doubao-speech-tts in calesthio/generative-media-skills) into .claude/skills/volcengine-doubao-speech-tts in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a codex`. Or copy the skill folder (skills/providers/text-to-speech/volcengine-doubao-speech-tts in calesthio/generative-media-skills) into .agents/skills/volcengine-doubao-speech-tts in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill volcengine-doubao-speech-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/volcengine-doubao-speech-tts, .gemini/skills/volcengine-doubao-speech-tts, .github/skills/volcengine-doubao-speech-tts and .opencode/skills/volcengine-doubao-speech-tts in your project.
Going by SKILL.md and its folder, Volcengine Doubao Speech Tts needs credentials named VOLC_ACCESS_KEY. Our summary lists: A credential in VOLC_ACCESS_KEY.
SKILL.md names 3 domains. In commands or code: openspeech.bytedance.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.volcengine.com and arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Volcengine Doubao Speech Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Volcengine Doubao Speech Tts: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.