Topic · Media & Creative
Best text to speech and voice skills, page 9
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 385 | 385.Fish Audio API Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice. | waker240/ | 201 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | 14 days ago |
| 386 | 386.Google Tts Convert documents and text to audio using Google Cloud Text-to-Speech. | sanjay3290/ | 432 | — | ~883 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 387 | 387.Flow Next QA Live-app real-user QA pass derived from the spec. An agent skill from gmickel/flow-next. | gmickel/ | 709 | — | ~1.2k | Automated safety check: Notes | MIT | today |
| 388 | 388.Oma Video Create short, explainer, or recorded-demo videos through the OMA video CLI. | first-fluke/ | 1.3k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 389 | 389.Aer Paper Body A skill your agent uses when drafting or revising the body sections of an AER, AER:Insights, or AEJ manuscript — institutional background, data, empirical strategy, results, mechanisms, and… | brycewang-stanford/ | 4.6k | 1 repo | ~3.3k | Automated safety check: Pass | Unknown | 5 days ago |
| 390 | 390.AI Music AI 音乐 / BGM 生成:给短视频、社媒内容生成原创背景音乐 / 配乐 / 纯音乐。通过可插拔 provider(阿里 DashScope / Suno 类第三方 API)文生音乐,异步提交→轮询→下载,产物可再裁剪/归一化或加到视频。当用户说“AI 音乐”“AI 配乐”“生成 BGM”“背景音乐”“原创音乐”“AI 作曲”“纯音乐”“给视频配乐”“做首曲子”时使用。与… | ZJU-REAL/ | 3.4k | — | ~896 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 391 | 多角色对话配音:按 cast 和逐行对白为不同角色分配音色与情绪,合成多声线音轨和带角色名字幕. An agent skill from ZJU-REAL/Easel. | ZJU-REAL/ | 3.4k | — | ~1.2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 392 | 392.Tts Voiceover 文字转语音配音:把文案/脚本合成为 AI 语音口播、旁白、朗读音频。配了 VOICEPROVIDER 默认走闭源云 TTS(CosyVoice2 等,有情感、像真人),edge 仅无 key 时兜底(edge 偏机械/AI 味);同步输出分句 SRT 字幕、mp3/wav/m4a。合成后可与 BGM 混音或加到视频作旁白。当用户说“配音”“文字转语音”“TTS”“AI… | ZJU-REAL/ | 3.4k | — | ~896 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 393 | 393.Voice Clone 上传本人语音样本克隆专属音色,再用它合成口播、旁白或带货语音。当用户说“声音克隆、克隆/复刻我的声音、用我的声音配音、定制专属音色”时使用,需要用户自备云端 provider 凭证。使用公共现成音色时改用 tts-voiceover。 | ZJU-REAL/ | 3.4k | — | ~715 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 394 | 394.Unified LLM API Call model APIs through @prismshadow/mmsp (MMSP) — streaming text generation, image generation, speech synthesis, embeddings and the supported-model registry with one client. | Prism-Shadow/ | 2.5k | — | ~6.7k | Automated safety check: Pass | Apache-2.0 | today |
| 395 | Install and use the official Multilingual Voiceover package, pinned by digest, for paid hosted work on the Beatra service. | sickn33/ | 47k | 1 repo | ~3.2k | Automated safety check: Pass | MIT-0 | yesterday |
| 396 | Install and use the official AI Podcast Voiceover package, pinned by digest, for paid hosted work on the Beatra service. | sickn33/ | 47k | 1 repo | ~3.1k | Automated safety check: Pass | MIT-0 | yesterday |
| 397 | Before accepting an agent's 'done / shipped / fixed' claim, verify it against ground truth (git ancestry + the commit's own diff) using the DOS kernel's dos verify and dos commit-audit — never the… | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 398 | 398.Md2video Audio Convert Markdown documents into narrated MP4 videos with synchronized visuals and voice narration. | sickn33/ | 47k | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 399 | Install and use the official AI Voice Cloning package, pinned by digest, for paid hosted work on the Beatra service. | sickn33/ | 47k | 1 repo | ~3.1k | Automated safety check: Pass | MIT-0 | yesterday |
| 400 | Install and use the official AI Voiceover Generator package, pinned by digest, for paid hosted work on the Beatra service. | sickn33/ | 47k | 1 repo | ~3.2k | Automated safety check: Pass | MIT-0 | yesterday |
| 401 | End-to-end video production orchestration for the content-creation workspace. | Pluviobyte/ | 1.6k | — | ~7.4k | Automated safety check: Notes | Unknown | 19 days ago |
| 402 | A skill your agent uses when user asks to translate videos, dub video content, or localize videos into other languages. | iflytek/ | 209 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 403 | 403.Tts Text-to-speech synthesis with ElevenLabs and system voices. An agent skill from alsk1992/CloddsBot. | alsk1992/ | 3k | — | ~1.2k | Automated safety check: Pass | MIT | 8 days ago |
| 404 | 404.Voice Voice recognition, wake words, and voice-controlled trading. An agent skill from alsk1992/CloddsBot. | alsk1992/ | 3k | — | ~1.1k | Automated safety check: Pass | MIT | 8 days ago |
| 405 | 405.Abo Issue Verify Verify that an Audiobook Organizer issue is done by checking acceptance criteria, code changes, tests, docs, changelog, and ABS matrix obligations. | jeeftor/ | 190 | — | ~499 | Automated safety check: Pass | MIT | 1 mo ago |
| 406 | 406.Comment Judge LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep). | fmflurry/ | 171 | — | ~2.5k | Automated safety check: Pass | MIT | 3 days ago |
| 407 | A skill your agent uses when creating cloned voices with Alibaba Cloud Model Studio CosyVoice customization models, especially cosyvoice-v3.5-plus or cosyvoice-v3.5-flash, from reference audio and… | cinience/ | 397 | — | ~904 | Automated safety check: Pass | MIT | 2 mo ago |
| 408 | A skill your agent uses when designing custom voices with Alibaba Cloud Model Studio CosyVoice customization models, especially cosyvoice-v3.5-plus or cosyvoice-v3.5-flash, from a voice prompt plus… | cinience/ | 397 | — | ~847 | Automated safety check: Pass | MIT | 2 mo ago |
| 409 | 409.Aliyun Qwen Tts A skill your agent uses when generating human-like speech audio with Model Studio DashScope Qwen TTS models (qwen3-tts-flash, qwen3-tts-instruct-flash). | cinience/ | 397 | — | ~951 | Automated safety check: Pass | MIT | 2 mo ago |
| 410 | A skill your agent uses when real-time speech synthesis is needed with Alibaba Cloud Model Studio Qwen TTS Realtime models. | cinience/ | 397 | — | ~788 | Automated safety check: Pass | MIT | 2 mo ago |
| 411 | A skill your agent uses when cloning voices with Alibaba Cloud Model Studio Qwen TTS VC models. | cinience/ | 397 | — | ~661 | Automated safety check: Pass | MIT | 2 mo ago |
| 412 | A skill your agent uses when designing custom voices with Alibaba Cloud Model Studio Qwen TTS VD models. | cinience/ | 397 | — | ~675 | Automated safety check: Pass | MIT | 2 mo ago |
| 413 | A skill your agent uses when replacing lip sync in existing videos with Alibaba Cloud Model Studio VideoRetalk (videoretalk). | cinience/ | 397 | — | ~735 | Automated safety check: Pass | MIT | 2 mo ago |
| 414 | Add or revise scene-bound narration or speech-linked captions in the active iPolloWork video using actual voice settings, measured audio and alignment. | Devin-AXIS/ | 6.8k | — | ~282 | Automated safety check: Pass | Unknown | today |
| 415 | Offline experimental phone-workflow helper that compares supplied prosodic features and suggests bounded TTS parameter changes. | CALLE-AI/ | 107 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 416 | 416.Speech Build Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. | cnemri/ | 127 | — | ~430 | Automated safety check: Pass | MIT | 8 mo ago |
| 417 | Teaches how to generate audio without assets in fluttersoloud using loadWaveform() with the 9 WaveForm oscillators, runtime setWaveform() tweaks, speechText() (built-in robotic TTS that creates and… | alnitak/ | 425 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 418 | 418.Cohub Generate Generate or transform images, video, speech, and music with Cohub multimodal models via cohub generate. | netaart/ | 572 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 419 | A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence. | microsoft/ | 255 | 1 repo | ~852 | Automated safety check: Pass | MIT | yesterday |
| 420 | A skill your agent uses when the user needs a Chinese narration phrase timeline from real audio: phrase-timeline.json, 逐句或逐短语时间标注, Whisper word-boundary alignment, timed captions, or… | ChenShuo2004/ | 197 | — | ~500 | Automated safety check: Pass | MIT | 2 days ago |
| 421 | 421.Offline Voice How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and… | notivn/ | 127 | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 422 | 422.Qwen3 Tts Local 内容创作者与影片与视频编辑在制作视频配音、有声书或无网环境朗读时,请用此技能一键生成多语言、多音色的本地高质量音频。基于Edge-TTS引擎,零配置完全离线,高效产出专业级语音,彻底告别网络限制。 | anbeime/ | 7.8k | — | ~586 | Automated safety check: Pass | No licence | yesterday |
| 423 | 影片与视频编辑、内容创作者在制作视频配音或有声书时,当需要克隆音色、生成情感化配音或流式实时语音合成请用此技能。支持1.7B高质量与0.6B快速双模型,一键实现多语言方言配音,让语音创作更高效更自然。 | anbeime/ | 7.8k | — | ~757 | Automated safety check: Pass | No licence | yesterday |
| 424 | Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. | aAAaqwq/ | 105 | 5 repos | ~913 | Automated safety check: Pass | MIT | 2 days ago |
| 425 | Command-line interface for MiniMax AI — chat (MiniMax-M3, MiniMax-M2.7) and speech-2.x TTS via the MiniMax API. | HKUDS/ | 52k | — | ~920 | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 426 | 426.iOS Hig A skill your agent uses when designing iOS interfaces, implementing accessibility (VoiceOver, Dynamic Type), handling dark mode, ensuring adequate touch targets, providing animation/haptic feedback… | johnrogers/ | 231 | — | ~705 | Automated safety check: Pass | MIT | 8 mo ago |
| 427 | 427.Deepgram Voice Select and tune a Deepgram TTS voice - curated voice list, full Aura voice catalog via API key, and tuning parameters | vellum-ai/ | 1.4k | — | ~2.6k | Automated safety check: Pass | MIT | yesterday |
| 428 | 428.Elevenlabs Voice Select and tune an ElevenLabs TTS voice - curated voice list, custom/cloned voices via API key, and tuning parameters | vellum-ai/ | 1.4k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 429 | 429.Voice Setup Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting | vellum-ai/ | 1.4k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 430 | 430.Pitch Deck Build a pitch deck as one self-contained HTML file (keyboard and click navigation, staggered slide-entry animations, a persistent brand mark, cited numbers with real links), write its timed… | ooiyeefei/ | 495 | — | ~920 | Automated safety check: Pass | MIT | 2 mo ago |
| 431 | 431.Tts Turn supplied text into spoken audio, single or multi-speaker. | win4r/ | 343 | 1 repo | ~2.7k | Automated safety check: Pass | No licence | 12 days ago |
| 432 | A skill your agent uses when user asks to clone a voice, train a custom voice model, or synthesize speech with a cloned voice. | iflytek/ | 209 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
Explore related skills
Category
More topics in Media & Creative
- Video production890
- Image generation724
- AI video generation667
- Transcription569
- Motion graphics387
- Design review and critique240
- Logo and visual identity226
- Comics and storyboards219
- Image editing193
- Social media graphics190
- Video scripts and shorts180
- Infographics159
- Music and audio generation150
- Podcasting122
- Generative and creative coding60