Topic · Media & Creative
Best text to speech and voice skills, page 7
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 289 | 289.Ig Reel Write an Instagram Reel from a raw idea - hook options off 26 formulas, the spoken script, the on-screen text, and a timed beat sheet - in the user's own voice and scored before they shoot it. | Jakeschincariol/ | 697 | — | ~1.6k | Automated safety check: Pass | MIT | 26 days ago |
| 290 | 290.Launch Video A skill your agent uses when the user wants a full launch video, hero video, or 20–60s product film combining product proof, brand, narration, music, and authored motion. | ucsandman/ | 251 | — | ~980 | Automated safety check: Pass | MIT | 1 mo ago |
| 291 | 291.AI Video The model-agnostic AI-video router and brief — the counterpart to image-prompt. | social-media-skills/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 8 days ago |
| 292 | 292.Auto Dubbing Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language. | sonilo-ai/ | 115 | — | ~4.5k | Automated safety check: Notes | MIT | today |
| 293 | Translate and dub videos from one language to another, replacing the original audio with TTS while keeping the video intact. | NoizAI/ | 526 | — | ~1.3k | Automated safety check: Notes | No licence | 11 days ago |
| 294 | 294.Tts Speak or render text to speech locally. An agent skill from codewhale-hq/Codewhale. | codewhale-hq/ | 41k | — | ~279 | Automated safety check: Notes | MIT | today |
| 295 | Async music, sound-effect and long-form voice generation via Venice. | veniceai/ | 143 | — | ~3.1k | Automated safety check: Pass | MIT | 3 days ago |
| 296 | Build call automation workflows with Azure Communication Services Call Automation Java SDK. | microsoft/ | 3.1k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | today |
| 297 | 297.Edge Tts Text-to-speech conversion using node-edge-tts npm package for generating audio from text. | sundial-org/ | 663 | 1 repo | ~1.8k | Automated safety check: Pass | No licence | 7 mo ago |
| 298 | 298.AI Media A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with… | ericrisco/ | 174 | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 299 | 299.Abo Abs Tests Work on Audiobook Organizer Audiobookshelf harness validation, ABS E2E tests, matrix updates, reset contracts, and ABS-facing behavior verification. | jeeftor/ | 190 | — | ~460 | Automated safety check: Pass | MIT | 29 days ago |
| 300 | Translate narration, ideas, or knowledge points into editorial halftone paper-collage metaphors, still prompts, storyboards, or stop-motion clips. | vllm-project/ | 7.1k | — | ~976 | Automated safety check: Pass | Apache-2.0 | today |
| 301 | 301.Agents Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills. | elevenlabs/ | 482 | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 302 | 302.Listenhub Tts 使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。 | smallnest/ | 290 | — | ~1.5k | Automated safety check: Notes | MIT | 26 days ago |
| 303 | Build multi-step AI content creation pipelines combining image, video, audio, and text. | NeverSight/ | 216 | 1 repo | ~1.8k | Automated safety check: Pass | No licence | today |
| 304 | 304.AI Voiceover The AI narration / voiceover mini-skill (ElevenLabs-led). An agent skill from social-media-skills/skills. | social-media-skills/ | 128 | — | ~1.1k | Automated safety check: Pass | MIT | 8 days ago |
| 305 | 305.Proofread Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before… | sonilo-ai/ | 115 | — | ~4.4k | Automated safety check: Notes | MIT | today |
| 306 | 306.Clean Audio Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video… | hassancs91/ | 325 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 307 | Wires xAI's Grok voice synthesis into an app you already have, so assistant replies get a play button, optional automatic playback, or narration of any text. | cursor/ | 10k | — | ~3.6k | Automated safety check: Pass | No licence | yesterday |
| 308 | Wires Grok speech-to-speech into an app's own microphone and audio playback over a realtime WebSocket, replacing an STT-LLM-TTS cascade or OpenAI Realtime. | cursor/ | 10k | — | ~1.7k | Automated safety check: Pass | No licence | yesterday |
| 309 | 309.Growth Log Write growth log entries that extract reusable patterns from completed work — root cause, transferable rule, and a recognizable signal — instead of diary-style event narration, with a 4-8 sentence… | affaan-m/ | 276k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 310 | 310.AI Video Script A skill your agent uses when a request asks for a Chinese-first AI video script with shot plans, image prompts, narration, subtitles, or handoff contracts for scene generation, image generation… | zrt-ai-lab/ | 287 | — | ~1.7k | Automated safety check: Pass | No licence | 2 mo ago |
| 311 | 311.Honey Review Review a diff for what Honey would cut — over-engineering (speculative generality, hand-rolled stdlib, single-caller abstractions) and over-verbosity (dead code, narration, redundant comments). | Green-PT/ | 313 | — | ~478 | Automated safety check: Pass | MIT | 1 mo ago |
| 312 | Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. | veniceai/ | 143 | — | ~3.6k | Automated safety check: Pass | MIT | 3 days ago |
| 313 | 313.Video Director Canonical entrypoint for every AI chat request that asks to make, generate, plan, or edit a video, including promotional films, ads, short videos, product videos, reference-image videos… | Jamailar/ | 1.8k | — | ~9.6k | Automated safety check: Pass | Unknown | today |
| 314 | Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/. | GAIK-project/ | 100 | — | ~2.4k | Automated safety check: Notes | MIT | today |
| 315 | A skill your agent uses when directing FLUX 3 audio, dialogue, or voiceover. | black-forest-labs/ | 127 | — | ~584 | Automated safety check: Pass | MIT | 1 mo ago |
| 316 | 生成自然真实的双人访谈播客,使用共享TTS模块支持3种引擎(Edge TTS / IndexTTS2 / MiniMax)和情感控制 | huangserva/ | 167 | — | ~941 | Automated safety check: Pass | No licence | 7 days ago |
| 317 | 用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。 | bozhouDev/ | 150 | — | ~667 | Automated safety check: Notes | MIT | 2 mo ago |
| 318 | 318.Abo Docs Update Audiobook Organizer documentation, AGENTS.md, changelog entries, and maintainer-facing workflow notes while keeping repo-local skill references consistent. | jeeftor/ | 190 | — | ~499 | Automated safety check: Pass | MIT | 29 days ago |
| 319 | 319.SaaS Video Make a SaaS demo or explainer video with Pexo. An agent skill from pexoai/pexo-skills. | pexoai/ | 802 | — | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 320 | 320.Gm Craft The Art of Game Mastering — narration, NPC, pacing, and improvisation wisdom that makes a session feel magical. | Sstobo/ | 155 | — | ~2k | Automated safety check: Warn | Unknown | 1 mo ago |
| 321 | 321.Setup API Key Guides users through setting up an ElevenLabs API key for REST API and SDK workflows. | elevenlabs/ | 482 | — | ~954 | Automated safety check: Notes | MIT | today |
| 322 | 322.Quota Axi Report local Claude, Codex, Cursor, GitHub Copilot, Grok, Kimi, Z.AI, Alibaba, OpenCode Go, Antigravity, Command Code, MiniMax, MiMo, DeepSeek, OpenRouter, ElevenLabs, Devin, Muse, and Higgsfield… | kunchenguid/ | 146 | — | ~547 | Automated safety check: Pass | MIT | today |
| 323 | 323.Voice Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip. | guaardvark/ | 257 | — | ~798 | Automated safety check: Pass | MIT | today |
| 324 | Create AI-powered podcasts with text-to-speech, music, and audio editing. | NeverSight/ | 216 | 1 repo | ~2k | Automated safety check: Pass | No licence | today |
| 325 | 325.Marketing Studio A skill your agent uses when generating any brand video, animation, image, or audio asset (logo reveal, social clip, product demo, launch video, OG image, README GIF, music, voiceover) for any… | ucsandman/ | 251 | — | ~1k | Automated safety check: Pass | MIT | 1 mo ago |
| 326 | 用火山引擎 Podcast AI 模型生成中文双人对话播客。当用户要把文章、报告、话题文本转成播客音频、生成对话式音频内容时使用,需要环境具备火山引擎 APPID 和 ACCESSKEY。支持 mp3/oggopus/pcm/aac、语速调节、自定义音色、断点续传。不用于:单人朗读式 TTS(用普通语音合成)、英文播客(模型主要优化中文)、播客文稿本身的撰写(先用写作类 skill… | staruhub/ | 727 | — | ~475 | Automated safety check: Pass | MIT | 1 mo ago |
| 327 | 327.Voice Realism Rewrite AI video / text-to-speech prompts so the generated VOICE sounds like a real human, not a robot. | nestyme/ | 151 | — | ~2.3k | Automated safety check: Pass | No licence | 9 days ago |
| 328 | 328.Motion Graphics A skill your agent uses when the user wants a short, design-led motion graphic where motion is the message: kinetic typography, stat or number count-up, chart/data-viz hit, logo sting, brand lockup… | chmonitor/ | 299 | 2 repos | ~3.4k | Automated safety check: Pass | GPL-3.0 | 4 days ago |
| 329 | 329.Narration 做一条由旁白带着走的成片(资讯快评、知识讲解、产品介绍、步骤教程),或给已有画面配一段多句的旁白时用:先按实测语速算出稿子该有多少字、按句分行写稿,再选一个全片不变的音色、用同一套语气说明一次合成,量时长、查接缝、只重配不合格的句子,最后把画面剪到旁白上。不用于把视频配成另一种语言(翻译配音),也不用于只念一句话。 | JimLiu/ | 533 | — | ~698 | Automated safety check: Pass | Unknown | today |
| 330 | 330.Ugc Plan and run UGC-style creator ads and social proof videos with genmedia. | fal-ai-community/ | 250 | — | ~1.5k | Automated safety check: Pass | No licence | 10 days ago |
| 331 | Convert written documents to narrated video scripts with TTS audio and word-level timing. | jwynia/ | 169 | — | ~3.5k | Automated safety check: Pass | MIT | 7 mo ago |
| 332 | 332.Scholar Ling Design and analyze studies in sociolinguistics, language variation, acoustic phonetics, discourse analysis, language contact, and computational linguistics. | joshzyj/ | 168 | — | ~6.7k | Automated safety check: Pass | Unknown | 21 days ago |
| 333 | 333.Lesson Make a narrated lesson video about lpm from the real desktop app, recorded on a pristine data directory. | gug007/ | 152 | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 334 | Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). | boraoztunc/ | 397 | 2 repos | ~3.7k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 335 | 335.Elevenlabs Tts This skill converts text to high-quality audio files using ElevenLabs API. | glebis/ | 390 | — | ~792 | Automated safety check: Notes | MIT | yesterday |
| 336 | 336.Fish Audio Generate AI audio using Fish Audio models. An agent skill from refly-ai/refly-skills. | refly-ai/ | 204 | — | ~538 | Automated safety check: Pass | No licence | 2 mo ago |
Explore related skills
Category
More topics in Media & Creative
- Video production892
- Image generation723
- AI video generation666
- Transcription577
- Motion graphics384
- Design review and critique239
- Logo and visual identity229
- Comics and storyboards220
- Image editing193
- Social media graphics187
- Video scripts and shorts186
- Infographics158
- Music and audio generation150
- Podcasting124
- Generative and creative coding60