Topic · Media & Creative
Best text to speech and voice skills, page 11
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 481 | Render a 'mosaic-grid-reveal' video from a config — a real-DOM FULL-BLEED N×N mosaic of real product tiles that pops in one tile at a time (scatter order, ease-out-back overshoot), the grid clears… | gooseworks-ai/ | 1.2k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 482 | 482.Video Polish Takes an existing screen recording or demo video and adds professional zoom/pan effects synchronized to the narration. | gooseworks-ai/ | 1.2k | 1 repo | ~3.5k | Automated safety check: Notes | MIT | yesterday |
| 483 | Dung khi mot CA NHAN muon len hinh bang AI thay vi tu quay — pipeline avatar AI: 3 tier cong cu, 4 workflow gom avatar don, dich da ngon ngu, san xuat hang loat va hybrid nguoi that cong AI; nhan… | minhnv0807/ | 609 | — | ~4.8k | Automated safety check: Pass | MIT | 28 days ago |
| 484 | Dung khi mot CA NHAN can AM THANH bang AI — clone giong noi, lam podcast, audiobook, voiceover cho video: 3 use case gom voiceover ngan cho TikTok va Reels, podcast 30-60 phut, audiobook; quy trinh… | minhnv0807/ | 609 | — | ~3.9k | Automated safety check: Pass | MIT | 28 days ago |
| 485 | 485.Voice Persona 让 Agent 变成能语音对话的机器人:语音文件转文字(支持微信 silk 格式)+ 多音色人格回复(Edge TTS 免费中文音色),全本地零 API 成本。Voice persona chat: transcribe voice messages (incl. | davepoon/ | 3.6k | — | ~947 | Automated safety check: Pass | MIT | yesterday |
| 486 | 486.API Scripts Call nodetool.scripts from a code action: create a voice script, add speakers with voices and spoken lines, voice the lines as audio takes, and assemble the takes into a voiceover timeline with word… | nodetool-ai/ | 560 | — | ~1k | Automated safety check: Pass | AGPL-3.0 | today |
| 487 | 487.Script Video Direct a NodeTool video whose timing follows written voiceover, such as an explainer or narrated b-roll. | nodetool-ai/ | 560 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 488 | Demos working software: first proves every claimed behavior by exercising the real artifact, then replays the proof as a slow, voice-narrated walkthrough. | lexler/ | 239 | — | ~715 | Automated safety check: Pass | Apache-2.0 | today |
| 489 | 489.Hyperframes Create HyperFrames HTML video compositions — animations, title cards, overlays, captions, GSAP timelines, registry blocks/components, voiceovers, audio-reactive visuals, and scene transitions. | OpenMinis/ | 446 | — | ~6.2k | Automated safety check: Pass | MIT | 2 days ago |
| 490 | Long-form and episodic video production pipeline (1-10 min, YouTube/web). | pawbytes/ | 113 | — | ~2.6k | Automated safety check: Pass | MIT | 7 days ago |
| 491 | 491.Hyperframes CLI Use the HyperFrames CLI development loop: init, add, catalog, capture, lint, check, snapshot, compare, grade-compare, preview, play, present, beats, keyframes, single or batch render, publish… | aiskillstore/ | 433 | 1 repo | ~2.7k | Automated safety check: Pass | No licence | today |
| 492 | Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills. | pexoai/ | 804 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 493 | 493.Voiceover Produce narrated voice-overs: choose the right voice per video style and language, generate segments, hand them to video tools. | OtoDock/ | 190 | — | ~688 | Automated safety check: Pass | Unknown | 2 days ago |
| 494 | Thin orchestrator for the end-to-end video localization pipeline. | jianshuo/ | 131 | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 495 | A skill your agent uses when the user wants to teach / learn an English word as a video — turn a single English word into a self-contained HLS "supercut" lesson built from the mira video base. | jianshuo/ | 131 | — | ~809 | Automated safety check: Pass | MIT | 1 mo ago |
| 496 | A skill your agent uses when the user wants text turned into an audiobook / read aloud — they give a passage, a file, a URL, or a VoiceDrop article and want an mp3 narration with expressive voices. | jianshuo/ | 131 | — | ~894 | Automated safety check: Notes | MIT | 1 mo ago |
| 497 | A skill your agent uses when producing a video ad with Scenario from a product shot or brand assets: a mobile 9:16 or landscape 16:9 commercial, a TikTok, Reels, Shorts, YouTube, or CTV spot, or… | scenario-labs/ | 946 | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 498 | 498.Fcpx Assistant Final Cut Pro X (FCPX) assistant — auto video production, TTS voiceover, media management, batch export | AI 自动成片、TTS 配音、素材管理、批量导出. | LeoYeAI/ | 2.2k | — | ~3.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 499 | 499.Minimax Studio Create voice, music, and video with MiniMax AI models. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~5.9k | Automated safety check: Notes | MIT | 2 mo ago |
| 500 | 500.Flow Next Prose Use while drafting any substantial reply, report, review walkthrough, or summary for the user - read the artifact prose contract before writing, not after, and draft under its rules. | gmickel/ | 709 | — | ~405 | Automated safety check: Pass | MIT | today |
| 501 | Video post-production rules: audio mastering, color, captions, platform export. | AnastasiyaW/ | 154 | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 502 | Test desktop apps with NVDA, JAWS, Narrator, VoiceOver and UIA tooling. | Community-Access/ | 423 | — | ~1.5k | Automated safety check: Pass | MIT | 16 days ago |
| 503 | 503.No Comments Review a diff for narration, stale comments, and suppressions that clearer code should replace. | tellahq/ | 394 | — | ~310 | Automated safety check: Pass | MIT | today |
| 504 | 504.Tts Text-to-speech — make the device speak text aloud. An agent skill from mikeyobrien/rho. | mikeyobrien/ | 373 | — | ~161 | Automated safety check: Pass | MIT | 9 days ago |
| 505 | 505.Tts Text-to-speech on macOS -- make the device speak text aloud. | mikeyobrien/ | 373 | — | ~177 | Automated safety check: Pass | MIT | 9 days ago |
| 506 | 506.Audio Studio Compose music, lyrics, soundscapes and voiceovers, or implement browser audio feedback. | WrongStack/ | 371 | — | ~910 | Automated safety check: Pass | MIT | today |
| 507 | Create exported TypeScript animations with Motion Canvas scenes, signals, generators and narration cues. | WrongStack/ | 371 | — | ~957 | Automated safety check: Pass | MIT | today |
| 508 | 508.Gtts Google Text-to-Speech (gTTS) for converting text to audio. An agent skill from benchflow-ai/skillsbench. | benchflow-ai/ | 1.8k | — | ~845 | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 509 | 509.Openai Tts OpenAI Text-to-Speech API for high-quality speech synthesis. | benchflow-ai/ | 1.8k | — | ~946 | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 510 | 510.Voice Clone Tts 声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。 | npc-live/ | 156 | — | ~1.1k | Automated safety check: Pass | No licence | 3 mo ago |
| 511 | 511.AI Task Hub AI task hub for image analysis, background removal, speech-to-text, text-to-speech, markdown conversion, and async execute/poll/presentation orchestration. | LeoYeAI/ | 2.2k | — | ~1.2k | Automated safety check: Pass | MIT | 2 mo ago |
| 512 | Convert written medical content into podcast or video scripts optimized for audio delivery. | LeoYeAI/ | 2.2k | — | ~4.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 513 | 每日名言語音任務。產生「語音 + 封面圖靜態影片 +(選配)HeyGen 數位人影片」並發送給主人. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~3.5k | Automated safety check: Pass | MIT | 2 mo ago |
| 514 | 514.Cadence Video 用 cadence 框架 做"代码渲染、音画同步"的视频:两条并列入口(脚本+TTS 实测时长 / 歌曲分析对齐)写同一份 project.json,可混合;带网页端时间线编辑器(拖拽剪辑、改字重生成语音、场景与帧特效标注)和无头 Chrome+ffmpeg… | pa001024/ | 137 | — | ~703 | Automated safety check: Pass | MIT | today |
| 515 | 515.Cartesia Cartesia text-to-speech: synthesize speech audio and browse voices. | Anil-matcha/ | 1.3k | — | ~773 | Automated safety check: Pass | MIT | 4 days ago |
| 516 | 516.Deepgram Deepgram speech AI: transcribe audio to text and synthesize speech (TTS). | Anil-matcha/ | 1.3k | — | ~857 | Automated safety check: Pass | MIT | 4 days ago |
| 517 | 517.Elevenlabs ElevenLabs: check subscription usage, list voices, and generate text-to-speech audio. | Anil-matcha/ | 1.3k | — | ~735 | Automated safety check: Pass | MIT | 4 days ago |
| 518 | 518.Gemini Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing. | Anil-matcha/ | 1.3k | — | ~778 | Automated safety check: Pass | MIT | 4 days ago |
| 519 | 519.Hume AI Hume AI Octave TTS and EVI speech-to-speech configs: synthesize speech and manage EVI configs. | Anil-matcha/ | 1.3k | — | ~746 | Automated safety check: Pass | MIT | 4 days ago |
| 520 | 520.Playht PlayHT text-to-speech: synthesize speech, list voices, and manage instant voice clones. | Anil-matcha/ | 1.3k | — | ~802 | Automated safety check: Pass | MIT | 4 days ago |
| 521 | 521.Invoking Gemini Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. | oaustegard/ | 150 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 522 | Creates talking head videos from any source material (docs, changelogs, blog posts, notes, transcripts). | gooseworks-ai/ | 1.2k | 1 repo | ~8.4k | Automated safety check: Notes | MIT | yesterday |
| 523 | Apply production-ready Deepgram SDK patterns for TypeScript and Python. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 524 | Configure CI/CD pipelines for ElevenLabs with mocked unit tests and gated integration tests. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 525 | Diagnose and fix ElevenLabs API errors by HTTP status code. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 526 | Implement ElevenLabs text-to-speech and voice cloning workflows. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 527 | Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 528 | Optimize ElevenLabs costs through model selection, character-efficient patterns, caching, and usage monitoring with budget alerts. | jeremylongshore/ | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | today |
Explore related skills
More topics in Media & Creative
- Video production890
- Image generation724
- AI video generation667
- Transcription569
- Motion graphics387
- Design review and critique240
- Logo and visual identity226
- Comics and storyboards219
- Image editing193
- Social media graphics190
- Video scripts and shorts180
- Infographics159
- Music and audio generation150
- Podcasting122
- Generative and creative coding60