Search
Text to speech and voice
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model… | shuyu-labs/ | 278 | — | ~787 | Automated safety check: Pass | Unknown | 3 mo ago |
| 98 | Design, evaluate, implement, or review image and video generation connectors for Timeline Studio. | MartinDelophy/ | 905 | — | ~989 | Automated safety check: Pass | MIT | yesterday |
| 99 | Build or resume a reusable 12-page personalized English picture book in the Moonlit watercolor house style, using Codex built-in image generation for a per-book character sheet plus illustrations… | lincwang123-bot/ | 174 | — | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 100 | 100.Narrator AI CLI Create AI-narrated film/drama commentary videos via CLI. An agent skill from NarratorAI-Studio/narrator-ai-cli. | NarratorAI-Studio/ | 137 | — | ~3.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 101 | Lilly community-research skill. An agent skill from ssaaffaakk/Lilly. | ssaaffaakk/ | 171 | — | ~1.4k | Automated safety check: Pass | MIT | 4 days ago |
| 102 | Turn a ~5-second voiceover line, opinion sentence, or abstract concept into a premium editorial halftone paper-collage B-roll clip — optionally with a fitted voiceover. | MegaTroll222/ | 216 | — | ~5.2k | Automated safety check: Pass | MIT | 2 mo ago |
| 103 | 103.Whiteboard Video 手绘白板风"边画边讲"讲解视频出片 skill(Excalidraw 风格逐笔动画 + 火山引擎配音 + 烧录字幕 + 品牌水印与片尾卡 + 横竖两张封面 + 各平台发布文案)。当用户说"做一期白板视频 / 边画边讲 / 手绘讲解视频 / 用 excalidraw 做视频 / whiteboard video",或要在本仓库里新建一期、改场景、换贴纸或真实… | trustfuture/ | 382 | — | ~1.8k | Automated safety check: Notes | MIT | 17 days ago |
| 104 | Chinese-language skill that produces a 25-second vertical video of a digital shopping-guide avatar for e-commerce, chaining AI image, voice and video generation. | anbeime/ | 7.8k | — | ~1k | Automated safety check: Pass | No licence | yesterday |
| 105 | Prefer the hosting AI Agent to write narration into input.txt from crawl JSON (no external LLM). | szhshp/ | 291 | — | ~1.5k | Automated safety check: Notes | No licence | yesterday |
| 106 | Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands. | modelstudioai/ | 542 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 107 | 107.Dotty Av Test Run local black-box voice tests against the physical Dotty robot by playing a TTS prompt through the workstation speakers while the C920 records video and room audio. | BrettKinny/ | 113 | — | ~866 | Automated safety check: Notes | MIT | 5 days ago |
| 108 | 108.Video Cut 把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。 | zenstory-ai/ | 561 | — | ~1.6k | Automated safety check: Pass | MIT | 7 days ago |
| 109 | Synthesize speech from text with Qwen TTS models. An agent skill from QianWen-AI/qianwen-ai. | QianWen-AI/ | 105 | — | ~4.2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 110 | A skill your agent uses when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI… | JetBrains/ | 366 | 3 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 111 | 111.Release Notes AivisSpeech の新バージョンリリース時に updateInfos のリリースノートドラフトを作成・更新するスキル。「リリースノート」「updateInfos」「アップデート情報」「リリース準備」などのキーワードが出たら使う。updateInfos.draft.json の作成・更新、エンジン側リリースノートとの統合、漏れチェックまでを包括的にサポートする。 | Aivis-Project/ | 483 | — | ~745 | Automated safety check: Pass | LGPL-3.0 | 3 mo ago |
| 112 | 112.Edge Tts Text-to-speech conversion using uvx edge-tts for generating audio from text. | aahl/ | 162 | 1 repo | ~1.1k | Automated safety check: Pass | MIT | 26 days ago |
| 113 | Renders landscape videos with the KrillinAI CLI, either the original footage with bilingual subtitles or a dubbed video with target-language subtitles. | krillinai/ | 13k | — | ~422 | Automated safety check: Pass | Apache-2.0 | today |
| 114 | Build conversational AI voice agents on the ElevenLabs platform. | jezweb/ | 1.1k | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 115 | 115.Remotion Molio's builtin skill for MAKING a video from any source — wiki notes, articles, scripts, product info, or a brief — and rendering it to MP4. | zhuzhaoyun/ | 433 | — | ~4k | Automated safety check: Pass | Unknown | yesterday |
| 116 | 116.Neurolink Guide Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink. | juspay/ | 148 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 117 | 117.Connector Setup A skill your agent uses when the user wants to add, connect, sign in to or fix an MCP connector (a server that gives you tools for an app such as GitHub, Linear or Notion). | katipally/ | 342 | — | ~860 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 118 | Use native WebMCP tools to edit the project open in Timeline Studio. | MartinDelophy/ | 905 | — | ~2.6k | Automated safety check: Pass | MIT | yesterday |
| 119 | 119.Erm Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings. | dougcalobrisi/ | 112 | — | ~1.3k | Automated safety check: Notes | MIT | 1 mo ago |
| 120 | 120.Movie Voice Tts 使用千问 Audio 3.0 TTS Plus、MiniMax Speech 2.8 HD 和 Fish Audio S2.1 Pro 生成或克隆电影角色旁白,包括直接复用用户选定的 Fish 公共音色编号并保存原生时间戳;供应商没有原生时间戳时,使用火山引擎语音识别恢复中文字幕或字级时间戳。制作中英文第一人称电影旁白、三模型克隆试音、供应商对比、整篇配音、字幕定时或音文对齐时使用。 | straighttttt/ | 135 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | 2 mo ago |
| 121 | 121.Tts Voiceover Generate voiceover audio using local VoxCPM2 TTS model with voice cloning. | liancheng-zcy/ | 106 | — | ~569 | Automated safety check: Pass | MIT | 4 mo ago |
| 122 | Generate faceless content video prompts for Seedance 2.0 on Higgsfield. | rediumvex/ | 409 | — | ~4.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 123 | 123.Local AI Use Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API. | amd/ | 408 | — | ~5k | Automated safety check: Notes | MIT | 2 days ago |
| 124 | Generates a HeyGen digital-human video layer using a fixed avatar profile, local narration and an approved circular avatar placement. | Pluviobyte/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Unknown | 20 days ago |
| 125 | 125.Stage Edit Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained… | Orkas-AI/ | 499 | — | ~2.4k | Automated safety check: Pass | MIT | 19 days ago |
| 126 | 126.Fal AI This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API. | mikeOnBreeze/ | 293 | — | ~1.9k | Automated safety check: Notes | MIT | 7 mo ago |
| 127 | 127.Xybrid Init Generate model metadata for an ML model so it works with xybrid. | xybrid-ai/ | 469 | — | ~3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 128 | 128.Sayit Live, low-latency spoken narration for hands-free agent sessions through the sayit TTS command. | callebtc/ | 160 | — | ~2.9k | Automated safety check: Pass | MIT | 18 days ago |
| 129 | 129.Vox Director Compile a topic, brief, article, or research artifact into a provider-neutral Flovart ProductionSpec for a VOX-inspired editorial paper-collage explainer. | avabbbb/ | 139 | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 130 | 130.Sound Fx A skill your agent uses whenever the user wants to generate sound effects, ambient audio, or short audio clips from a text description. | NoizAI/ | 526 | — | ~1.4k | Automated safety check: Pass | No licence | 13 days ago |
| 131 | 131.Gemini Tts Generates spoken MP3 audio from text or Markdown with Gemini TTS. | iurysza/ | 420 | — | ~968 | Automated safety check: Pass | MIT | 3 days ago |
| 132 | Creates professional 50-60s vertical explainer videos (9:16 format, 1080x1920 @ 30fps) for any given topic using Remotion, styled with modern aesthetics and narrated by natural voiceover using… | Cuongyd196/ | 171 | — | ~2.2k | Automated safety check: Notes | No licence | 1 mo ago |
| 133 | 133.Video Reference 按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。 | zenstory-ai/ | 561 | — | ~1.3k | Automated safety check: Pass | MIT | 7 days ago |
| 134 | 134.Tabz Browser Browser automation via 70 tabz MCP tools. An agent skill from GGPrompts/TabzChrome. | GGPrompts/ | 147 | — | ~730 | Automated safety check: Pass | MIT | 13 days ago |
| 135 | 135.Vocab Chat Create a vocabulary learning chat MulmoScript with messenger-style animated UI (voiceover approach). | receptron/ | 475 | — | ~2.5k | Automated safety check: Notes | No licence | today |
| 136 | Install AgentVibes TTS voice system for BMAD agents. An agent skill from paulpreibisch/AgentVibes. | paulpreibisch/ | 155 | — | ~557 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 137 | 137.Summarize Call Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants). | reysu/ | 270 | — | ~3.8k | Automated safety check: Notes | MIT | 1 mo ago |
| 138 | 138.Blockrun Pay-per-call access to AI models, real-time data, media generation and multi-chain RPC over x402 micropayments (USDC on Base or Solana), or a BlockRun account API key. | BlockRunAI/ | 391 | — | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 139 | 139.Keirouter Entry point for KeiRouter — local/remote AI gateway with OpenAI-compatible REST for chat, image, TTS, embeddings, web search, web fetch. | mydisha/ | 147 | — | ~995 | Automated safety check: Pass | MIT | 1 mo ago |
| 140 | Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration. | ConardLi/ | 13k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 141 | 141.Add Tts Model Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 142 | 142.Hyperframes Create HTML-based video compositions, animated title cards, social overlays, captioned talking-head videos, audio-reactive visuals, and shader transitions using HyperFrames. | johnson7788/ | 327 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 143 | Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API. | bozhouDev/ | 150 | — | ~1.9k | Automated safety check: Notes | MIT | 2 mo ago |
| 144 | 144.Hyperframes CLI HyperFrames CLI tool — hyperframes init, lint, inspect, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark. | coleam00/ | 149 | — | ~1.6k | Automated safety check: Pass | No licence | 5 mo ago |