Topic · Media & Creative
Best text to speech and voice skills, page 6
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | 241.Hyperframes CLI HyperFrames CLI and Minis rendering. An agent skill from OpenMinis/MinisSkills. | OpenMinis/ | 444 | — | ~3.5k | Automated safety check: Pass | MIT | yesterday |
| 242 | 242.Runpod Cloud GPU processing via RunPod serverless. An agent skill from digitalsamba/claude-code-video-toolkit. | digitalsamba/ | 2.2k | — | ~2.1k | Automated safety check: Notes | MIT | 3 days ago |
| 243 | 243.Suggest Sfx Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor. | hassancs91/ | 325 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 244 | 244.Fal AI Media Unified media generation via fal.ai MCP — image, video, and audio. | affaan-m/ | 276k | 4 repos | ~1.9k | Automated safety check: Pass | MIT | 4 days ago |
| 245 | Prepare Fish official-API narration, OpenAI word timestamps, and reviewed user-supplied local music/SFX for HyperFrames. | waker240/ | 199 | — | ~1.2k | Automated safety check: Notes | Apache-2.0 | 12 days ago |
| 246 | 当用户要把产品事实、口播、数字人、产品界面和 CTA 制作成可验收的数字人产品介绍视频时使用:统一预检 ChatCut、FFmpeg、ComfyUI、Fish/TTS 与 Remotion,按 plan、sample、batch 三种模式编排,先完成可审批样片再批量或出成片。用于有明确产品包的横版/竖版产品视频流水线;不要用于通用自动剪辑、电商短视频复刻、纯视频选题策划或未经确认的批量生成与发布。 | ChenShuo2004/ | 194 | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 247 | 247.Voice Changer Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API. | elevenlabs/ | 482 | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 248 | The AI music + sound-design skill for social -- original/licensed audio beds and sound design for Reels/TikToks/Shorts/videos. | social-media-skills/ | 128 | — | ~1.9k | Automated safety check: Pass | MIT | 7 days ago |
| 249 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. | deepgram/ | 276 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 250 | 250.Audio Ducking Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps. | sonilo-ai/ | 115 | — | ~1.8k | Automated safety check: Notes | MIT | today |
| 251 | Builds a short spoken morning news briefing from a few web searches and turns it into one audio clip, on request or on a schedule. | THU-SAGE/ | 303 | — | ~533 | Automated safety check: Pass | MIT | 4 mo ago |
| 252 | Agent-callable ElevenLabs tools — generate spoken audio from text, create sound effects and multi-speaker dialogue, re-voice and clean up audio, transcribe audio and video, design synthetic voices… | zapier/ | 176 | — | ~3.7k | Automated safety check: Pass | Elastic-2.0 | 1 mo ago |
| 253 | Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API. | diegosouzapw/ | 74k | — | ~5.5k | Automated safety check: Pass | MIT | today |
| 254 | 254.Kling Official Official Kling direct API guidance for OpenMontage providers. | calesthio/ | 66k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | 5 days ago |
| 255 | Synthetic terminal-style screen recording guidance for Remotion TerminalScene. | calesthio/ | 66k | — | ~2.5k | Automated safety check: Pass | MIT | 5 days ago |
| 256 | 256.Agentvibes 🎤 AgentVibes Voice Management - Manage your text-to-speech voices across multiple providers (Piper TTS, Piper, macOS Say). | paulpreibisch/ | 155 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 22 days ago |
| 257 | Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead… | nodetool-ai/ | 560 | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | today |
| 258 | A skill your agent uses when writing or editing Pixi'VN story content — defining labels (scenes) with newLabel, writing dialogue steps, adding player choices with… | DRincs-Productions/ | 149 | — | ~4.9k | Automated safety check: Pass | LGPL-2.1 | 8 days ago |
| 259 | 259.Cosyvoice Ssml 内部技能,仅用于数字人 / VideoRetalk / 资产库角色口播视频链路中的 CosyVoice TTS 子步骤。普通 AI 聊天、普通音频、普通短视频、产品视频或广告视频请求不得直接调用本技能;这些视频请求必须先走 video-director。只有 video-director 已确认这是数字人口播,且需要把已批准的台词合成为角色驱动音频时,才可激活本技能。 | Jamailar/ | 1.8k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 260 | 260.Demo Video A skill your agent uses when asked to make a demo, training, how-to or tutorial video with sound or voice-over showing an HMIS function or configuration (e.g. | hmislk/ | 236 | — | ~3.9k | Automated safety check: Pass | GPL-3.0 | today |
| 261 | 261.Audio Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio. | autonomous-ai/ | 381 | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 262 | 用于“网页 demo 分段配音 + timeline 驱动录屏 + 后期合成”的 workspace 协作流程:先搭建一个可审计工作目录(cues/timeline/segmentaudio/video/subtitles/final),再由人类 + Codex 迭代维护这些文件,按需只重跑局部步骤,最终合成高质量 MP4。适用于强调可复盘、可编辑、清晰度与字幕安全区可控的场景。 | Sven-LI-sankyuu/ | 175 | — | ~3.2k | Automated safety check: Pass | No licence | 2 mo ago |
| 263 | 在创作者明确确认后,执行短剧项目的图片、视频、TTS/配音或时间线音乐生产任务,并把结果与精简运行记录落回项目。用户说“生成这张图/这段视频/这句配音/这段配乐”“开始跑图/跑视频/合成语音/生成音乐”“把已确认提示词送去生产”,或要求批量执行已确认媒体任务时使用;不负责创作提示词、镜头、台词、歌词或声音身份,也绝不把预览、继续、预算说明或既有接受状态当作本次付费生产确认。 | zenstory-ai/ | 2.6k | — | ~1.9k | Automated safety check: Pass | MIT | 6 days ago |
| 264 | 264.Abo PR Create Commit, push, and create Audiobook Organizer pull requests into protected master after verification and PR body preparation. | jeeftor/ | 190 | — | ~540 | Automated safety check: Pass | MIT | 28 days ago |
| 265 | 265.Scene Splitter Splits a plain English story into a numbered list of SCENES — each scene being one moment that gets exactly one illustration AND one narration clip downstream. | hassancs91/ | 101 | — | ~2.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 266 | 把一篇技术长文/论文解读自动做成章节式解说视频(1080p, 5-8 分钟)。双主题:warm(奶油底+珊瑚红+cozy-handdrawn 透明插图,亲和感)和 midnight(深蓝黑底+琥珀金+宋体标题+executive-tech 插图,AI 科技感),storyboard 一个 theme 字段切换。每章三种 layout 混排:illustration(左文右图+Ken… | wwwzhouhui/ | 283 | — | ~2.2k | Automated safety check: Pass | No licence | 3 days ago |
| 267 | Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations. | glebis/ | 390 | — | ~2.3k | Automated safety check: Notes | MIT | yesterday |
| 268 | 268.Voice Isolator Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API. | elevenlabs/ | 482 | — | ~923 | Automated safety check: Pass | MIT | today |
| 269 | 269.Audio Track A skill your agent uses when the user wants music, voiceover, narration, or a soundtrack added to a video asset, OR wants standalone generated audio for any purpose (e.g. | ucsandman/ | 251 | — | ~1.1k | Automated safety check: Notes | MIT | 1 mo ago |
| 270 | 270.Hig Inputs Apple HIG guidance for input methods and interaction patterns: gestures, Apple Pencil, keyboards, game controllers, pointers, Digital Crown, eye tracking, focus system, remotes, spatial… | raintree-technology/ | 143 | 5 repos | ~1.5k | Automated safety check: Pass | MIT | 28 days ago |
| 271 | 271.To Narration A skill your agent uses when writing or revising spoken narration from complete, lineage-valid video beats, including scripts that must remain open to later visual direction. | sugarforever/ | 167 | — | ~819 | Automated safety check: Pass | MIT | today |
| 272 | 272.Elevenlabs Convert documents and text to audio using ElevenLabs text-to-speech. | sanjay3290/ | 431 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 28 days ago |
| 273 | 273.Speech Engine Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. | elevenlabs/ | 482 | — | ~2.5k | Automated safety check: Warn | MIT | today |
| 274 | 274.Keirouter Tts Text-to-speech via KeiRouter /v1/audio/speech using OpenAI / ElevenLabs / Deepgram / Edge TTS / Google TTS / Inworld voices. | mydisha/ | 147 | — | ~599 | Automated safety check: Pass | MIT | 28 days ago |
| 275 | Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays… | cosmicstack-labs/ | 476 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 276 | Convert a single provided file into a downloadable Markdown file, representing source images and figures as readable text. | impredicative/ | 157 | — | ~174 | Automated safety check: Pass | LGPL-3.0 | 4 days ago |
| 277 | 277.Story Create high-quality MulmoScript through structured multi-phase creative process | receptron/ | 475 | — | ~3.5k | Automated safety check: Notes | No licence | today |
| 278 | Expert in building voice AI applications - from real-time voice agents to voice-enabled apps. | davila7/ | 32k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | today |
| 279 | Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media. | madebyaris/ | 126 | — | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 280 | Use the September 2026 image, video, speech and Avatar V adapters with explicit model/host contracts. | calesthio/ | 66k | — | ~1.4k | Automated safety check: Pass | AGPL-3.0 | 5 days ago |
| 281 | 281.Arkcli Deploy arkcli +deploy:普通创建推理接入点(Endpoint)的统一首选入口。用户说『创建/新建/create 一个 endpoint/接入点』或『部署/上线/deploy 某模型』时优先走这里;但脚本化 / CI / 无护栏 / 原始 raw CRUD 创建是唯一例外,必须改走 arkcli-infer-endpoint,不能由本 skill… | volcengine/ | 140 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 282 | 282.Voxclaw Give your agent a voice. An agent skill from malpern/VoxClaw. | malpern/ | 208 | — | ~1.9k | Automated safety check: Pass | No licence | 2 mo ago |
| 283 | Multi-stage workflow for producing videos set in the Three-Body universe, from lore research and storyboards to voice, sound and music generation and final compositing. | anbeime/ | 7.7k | — | ~1.6k | Automated safety check: Pass | No licence | yesterday |
| 284 | Generates media for intelligent textbooks - slide decks and presentations (MARP web decks in docs/slides/ or PowerPoint .pptx lecture downloads), illustrated stories and graphic novels… | dmccreary/ | 105 | — | ~1.7k | Automated safety check: Pass | CC-BY-NC-4.0 | yesterday |
| 285 | 285.Line Editing This skill should be used when the user asks to "line edit", "edit my prose", "polish this chapter", "tighten the prose", "improve the sentences", "copyedit", "proofread", "proof pass", "check… | danjdewhurst/ | 283 | 1 repo | ~3.1k | Automated safety check: Notes | MIT | today |
| 286 | 286.Story Narrator Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech). | hassancs91/ | 101 | — | ~2.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 287 | 把单个学科概念(中文/英文都行,如"声现象""杠杆原理""光合作用")自动做成动态教学内容。两种产出都由 HyperFrames 渲染、共用同一个 index.html(mode 变量切换):① 配音教学视频 — 分镜 + Minimax / Edge TTS 中文配音 + 字幕 + 完整 MP4(1080p, ~90s, 发视频号/给孩子看);② 无声循环动图 — 同一内容的紧凑无声版… | wwwzhouhui/ | 283 | — | ~1.5k | Automated safety check: Pass | No licence | 3 days ago |
| 288 | Fills an uploaded PowerPoint into an existing classroom stage with its layout intact, then checks each page, writes narration and actions, and generates speech. | THU-MAIC/ | 40k | — | ~2.5k | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in Media & Creative
- Video production892
- Image generation723
- AI video generation666
- Transcription577
- Motion graphics384
- Design review and critique239
- Logo and visual identity229
- Comics and storyboards220
- Image editing193
- Social media graphics187
- Video scripts and shorts186
- Infographics158
- Music and audio generation150
- Podcasting124
- Generative and creative coding60