Topic · Media & Creative
Best text to speech and voice skills, page 5
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | 音频素材生成与获取。批量 Edge TTS 旁白生成(支持 storyboard pacing 字段驱动语速)、BGM/SFX 检索(BGM 节奏匹配 BPM 规则)、音频时长提取。包含 Edge voice 配置、速度调整规则、durations.json 格式规范(含 audiovisualrelation 说明)和关键的音频时序规则。 | bilibili/ | 129 | — | ~2.5k | Automated safety check: Pass | Unknown | 5 mo ago |
| 194 | 194.Pneuma Clipcraft AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills. | pandazki/ | 161 | — | ~7.5k | Automated safety check: Notes | MIT | today |
| 195 | Generate accessibility infrastructure for VoiceOver, Dynamic Type, and accessibility features. | rshankras/ | 785 | — | ~1.4k | Automated safety check: Notes | MIT | 2 mo ago |
| 196 | Intelligently compress and rewrite documents into TTS-friendly scripts. | huangserva/ | 167 | — | ~936 | Automated safety check: Notes | No licence | 7 days ago |
| 197 | arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli… | volcengine/ | 140 | — | ~743 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 198 | 198.Voiceover Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved… | GTKottman/ | 476 | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | today |
| 199 | 199.Teachany K-12 interactive courseware creation. An agent skill from infometa/workbuddyskills. | infometa/ | 346 | 1 repo | ~1.4k | Automated safety check: Notes | No licence | today |
| 200 | End-to-end viral tech reel production for Instagram Reels and TikTok using 2026 trend grammar — retention-first pacing, punch-ins, 3D cinematic AI-generated shots, motion design graphics, proof… | tornikebolokadze1-cyber/ | 143 | — | ~3k | Automated safety check: Pass | CC0-1.0 | yesterday |
| 201 | Plan and generate one continuous product video shot from one reference image, with a coherent camera move and product identity checks. | T8mars/ | 615 | — | ~255 | Automated safety check: Pass | MIT | yesterday |
| 202 | Internal skill that rewrites and re-renders sound prompts in a draft PeonPing pack to follow a reroll caption, then records the change in a log. | PeonPing/ | 5.1k | — | ~910 | Automated safety check: Pass | MIT | 2 days ago |
| 203 | 203.Voice Compose Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generatevoice.js CLI. | Utopai-Research/ | 355 | — | ~631 | Automated safety check: Pass | Unknown | 9 days ago |
| 204 | 204.App Builder App Builder — generate complete, runnable fullstack WebUI applications (FastAPI backend + pure HTML/CSS/JS frontend) around on-device AI Model Packs (OCR, TTS, ASR, Super-Resolution, etc.). | qualcomm/ | 247 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 205 | 205.Vibe Scene Author, repair, render, and inspect VibeFrame scene projects built from STORYBOARD.md and DESIGN.md. | vericontext/ | 175 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 206 | 把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”… | cacity/ | 168 | — | ~2.3k | Automated safety check: Notes | MIT | 7 days ago |
| 207 | 207.Diy Yt Creator A skill your agent uses when the user wants to create a new YouTube Short using one of the templates in this repo's templates/ folder. | coleam00/ | 148 | — | ~844 | Automated safety check: Pass | No licence | 4 mo ago |
| 208 | 208.Stage Plan The "ingest + plan" half of end-to-end video orchestration — ingest the user's material from evidence, then decompose intent into ONE cross-modal EDL (plan.json: edit/generate/compose/provided… | Orkas-AI/ | 498 | — | ~4k | Automated safety check: Pass | MIT | 17 days ago |
| 209 | 209.Unit Video 把一个阅读单元的「备读」做成「领读视频 / 备读视频 / 讲解视频」: 作者第一人称、课件式讲解、16:9 横屏。当前仅支持《道德经》。从 books/道德经/chNN/NN.md 出发, 经「口播稿 → TTS 时间轴 → 分镜 atoms/steps → 渐进渲染」产出 video/道德经/chNN/NN/领读视频.mp4。用户要求把《道德经》某章、某单元或备读内容做成视频、… | plustar35/ | 146 | — | ~1.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 210 | A skill your agent uses when building a finished product ad from FLUX 3 - shot design, voiceover, action-to-word sync, evidence-gated copy, deterministic assembly, and QC gates that catch clipped… | black-forest-labs/ | 127 | — | ~9.8k | Automated safety check: Pass | MIT | 29 days ago |
| 211 | 211.Voice To Video 口播文字稿一键成片:TTS 配音(edge-tts,带逐句/逐词时间戳)→ HTML 动画合成(场景由时间戳驱动,画面跟着声音走)→ Playwright 逐帧确定性渲染合成 MP4。Use whenever the user wants to 把口播稿/文字稿/文案/文章做成视频、配音视频、解说视频、知识类短视频、口播视频、文字转视频、TTS video、voice-over… | wwwzhouhui/ | 283 | — | ~1.1k | Automated safety check: Pass | No licence | 3 days ago |
| 212 | 212.Music Generate music using ElevenLabs Music API. An agent skill from bozhouDev/video-skills-toolkit. | bozhouDev/ | 150 | — | ~3.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 213 | Turns a document into a narrated roadshow video through ten staged roles, from document analysis and slide planning to audio, subtitles and final composition. | anbeime/ | 7.7k | — | ~1.4k | Automated safety check: Notes | No licence | yesterday |
| 214 | 214.Tutorial Videos Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals)… | matthiasn/ | 1.2k | — | ~2.2k | Automated safety check: Notes | GPL-3.0 | today |
| 215 | Orchestrates repo pipeline from a single URL: branch Zhihu vs generic page → spider → DeepSeek caption when needed → Edge TTS → Remotion render to output/video/video.mp4. | szhshp/ | 291 | — | ~1.2k | Automated safety check: Notes | No licence | 1 mo ago |
| 216 | 216.Qwen Edit AI image editing prompting patterns for Qwen-Image-Edit. An agent skill from digitalsamba/claude-code-video-toolkit. | digitalsamba/ | 2.2k | — | ~711 | Automated safety check: Pass | MIT | 3 days ago |
| 217 | 217.Make Tsx Step 2 of the AI Video Editor pipeline — build the visual beats (Remotion TSX shots) over a project's master cut and bake a composited preview. | hassancs91/ | 325 | — | ~1.9k | Automated safety check: Notes | MIT | 1 mo ago |
| 218 | 218.Tts A skill your agent uses whenever the user wants to convert text into speech, generate audio from text, or produce voiceovers. | NoizAI/ | 526 | — | ~1.9k | Automated safety check: Pass | No licence | 11 days ago |
| 219 | 219.Text To Speech Convert text to speech using ElevenLabs voice AI. An agent skill from elevenlabs/skills. | elevenlabs/ | 482 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 220 | Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. | sammcj/ | 162 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 221 | 221.Video Voiceover 把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。 | zenstory-ai/ | 559 | — | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 222 | Generate neural narration audio using Azure AI Speech (REST text-to-speech). | calesthio/ | 66k | — | ~1.4k | Automated safety check: Pass | MIT | 5 days ago |
| 223 | 223.Dashscope DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). | calesthio/ | 66k | — | ~1.5k | Automated safety check: Notes | AGPL-3.0 | 5 days ago |
| 224 | 224.Doubao Tts Generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. | calesthio/ | 66k | — | ~988 | Automated safety check: Notes | AGPL-3.0 | 5 days ago |
| 225 | 225.Fish Audio Tts Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via referenceid. | calesthio/ | 66k | — | ~1.4k | Automated safety check: Notes | AGPL-3.0 | 5 days ago |
| 226 | 226.Gemini Omni Generate and conversationally edit short videos with Google Gemini Omni Flash (gemini-omni-flash-preview). | calesthio/ | 66k | — | ~2.1k | Automated safety check: Notes | AGPL-3.0 | 5 days ago |
| 227 | 227.Setup API Key Guides users through setting up an ElevenLabs API key for ElevenLabs MCP tools. | calesthio/ | 66k | — | ~717 | Automated safety check: Notes | MIT | 5 days ago |
| 228 | 228.Text To Speech Generate speech audio from text using HeyGen's Starfish TTS model. | calesthio/ | 66k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | 5 days ago |
| 229 | 229.Video Translate Translate and dub existing videos into multiple languages using HeyGen. | calesthio/ | 66k | — | ~2.7k | Automated safety check: Pass | AGPL-3.0 | 5 days ago |
| 230 | Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper. | deepgram/ | 469 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 231 | 231.Motion Canvas Production pipeline for Motion Canvas — TypeScript-based programmatic vector animation with real-time preview. | dracohu2025-cloud/ | 227 | — | ~1.5k | Automated safety check: Pass | No licence | 22 days ago |
| 232 | 232.World Narration 根据已冻结规则事实生成静态房间景物正文时使用。保留场景与空间事实,同时允许自然合理的文学补白,不创造可操作玩法. An agent skill from oiuv/mud. | oiuv/ | 179 | — | ~417 | Automated safety check: Pass | MIT | today |
| 233 | ElevenLabs の新モデル追加時に使用。provider2agent.ts のモデルリスト更新とテストスクリプト更新を行う。 | receptron/ | 475 | — | ~613 | Automated safety check: Notes | No licence | today |
| 234 | Make Manim teaching animations with narration and subtitles when motion helps a concept. | madhvantyagi/ | 337 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 235 | Inspect Audiobook Organizer GitHub issues, comments, acceptance criteria, linked PRs, and next steps. | jeeftor/ | 190 | — | ~417 | Automated safety check: Pass | MIT | 28 days ago |
| 236 | 236.Sherpa Onnx Tts Local text-to-speech via sherpa-onnx (offline, no cloud). An agent skill from trpc-group/trpc-agent-go. | trpc-group/ | 1.9k | 9 repos | ~850 | Automated safety check: Pass | Apache-2.0 | today |
| 237 | A skill your agent uses when user asks to synthesize speech, convert text to audio, or read text aloud. | iflytek/ | 209 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 238 | 238.Computer Use Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser. | katipally/ | 339 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 239 | 239.Hyperframes Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. | scott-fryxell/ | 125 | — | ~3.2k | Automated safety check: Pass | MIT | 11 days ago |
| 240 | 240.Repo Conventions NeuroLink's review standards — the critical rules to enforce, what NOT to comment on, the security bar, hot paths. | juspay/ | 144 | — | ~1.3k | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in Media & Creative
- Video production892
- Image generation723
- AI video generation666
- Transcription577
- Motion graphics384
- Design review and critique239
- Logo and visual identity229
- Comics and storyboards220
- Image editing193
- Social media graphics187
- Video scripts and shorts186
- Infographics158
- Music and audio generation150
- Podcasting124
- Generative and creative coding60