Search
Text to speech and voice
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | 241.Computer Use Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser. | katipally/ | 342 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 242 | A skill your agent uses when user asks to synthesize speech, convert text to audio, or read text aloud. | iflytek/ | 209 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 243 | 243.Runpod Cloud GPU processing via RunPod serverless. An agent skill from digitalsamba/claude-code-video-toolkit. | digitalsamba/ | 2.2k | — | ~2.1k | Automated safety check: Notes | MIT | yesterday |
| 244 | 244.Repo Conventions NeuroLink's review standards — the critical rules to enforce, what NOT to comment on, the security bar, hot paths. | juspay/ | 148 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 245 | 245.Hyperframes Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. | scott-fryxell/ | 125 | — | ~3.2k | Automated safety check: Pass | MIT | 13 days ago |
| 246 | 246.Hyperframes CLI HyperFrames CLI and Minis rendering. An agent skill from OpenMinis/MinisSkills. | OpenMinis/ | 446 | — | ~3.5k | Automated safety check: Pass | MIT | 3 days ago |
| 247 | 247.Suggest Sfx Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor. | hassancs91/ | 328 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 248 | Prepare Fish official-API narration, OpenAI word timestamps, and reviewed user-supplied local music/SFX for HyperFrames. | waker240/ | 201 | — | ~1.2k | Automated safety check: Notes | Apache-2.0 | 15 days ago |
| 249 | 249.Fal AI Media Unified media generation via fal.ai MCP — image, video, and audio. | affaan-m/ | 277k | 4 repos | ~1.9k | Automated safety check: Pass | MIT | today |
| 250 | 当用户要把产品事实、口播、数字人、产品界面和 CTA 制作成可验收的数字人产品介绍视频时使用:统一预检 ChatCut、FFmpeg、ComfyUI、Fish/TTS 与 Remotion,按 plan、sample、batch 三种模式编排,先完成可审批样片再批量或出成片。用于有明确产品包的横版/竖版产品视频流水线;不要用于通用自动剪辑、电商短视频复刻、纯视频选题策划或未经确认的批量生成与发布。 | ChenShuo2004/ | 197 | — | ~1.3k | Automated safety check: Pass | MIT | 3 days ago |
| 251 | The AI music + sound-design skill for social -- original/licensed audio beds and sound design for Reels/TikToks/Shorts/videos. | social-media-skills/ | 134 | — | ~1.9k | Automated safety check: Pass | MIT | 10 days ago |
| 252 | 252.Voice Changer Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API. | elevenlabs/ | 482 | — | ~2.9k | Automated safety check: Pass | MIT | 2 days ago |
| 253 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. | deepgram/ | 276 | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 254 | 254.Audio Ducking Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps. | sonilo-ai/ | 115 | — | ~1.8k | Automated safety check: Notes | MIT | 2 days ago |
| 255 | Builds a short spoken morning news briefing from a few web searches and turns it into one audio clip, on request or on a schedule. | THU-SAGE/ | 303 | — | ~533 | Automated safety check: Pass | MIT | 4 mo ago |
| 256 | Agent-callable ElevenLabs tools — generate spoken audio from text, create sound effects and multi-speaker dialogue, re-voice and clean up audio, transcribe audio and video, design synthetic voices… | zapier/ | 177 | — | ~3.7k | Automated safety check: Pass | Elastic-2.0 | 1 mo ago |
| 257 | Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API. | diegosouzapw/ | 75k | — | ~5.5k | Automated safety check: Pass | MIT | today |
| 258 | 258.Kling Official Official Kling direct API guidance for OpenMontage providers. | calesthio/ | 66k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | 8 days ago |
| 259 | Synthetic terminal-style screen recording guidance for Remotion TerminalScene. | calesthio/ | 66k | — | ~2.5k | Automated safety check: Pass | MIT | 8 days ago |
| 260 | 260.Agentvibes 🎤 AgentVibes Voice Management - Manage your text-to-speech voices across multiple providers (Piper TTS, Piper, macOS Say). | paulpreibisch/ | 155 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 261 | Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead… | nodetool-ai/ | 560 | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | today |
| 262 | 262.Cosyvoice Ssml 内部技能,仅用于数字人 / VideoRetalk / 资产库角色口播视频链路中的 CosyVoice TTS 子步骤。普通 AI 聊天、普通音频、普通短视频、产品视频或广告视频请求不得直接调用本技能;这些视频请求必须先走 video-director。只有 video-director 已确认这是数字人口播,且需要把已批准的台词合成为角色驱动音频时,才可激活本技能。 | Jamailar/ | 1.8k | — | ~4.9k | Automated safety check: Pass | Unknown | 2 days ago |
| 263 | 263.Audio Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio. | autonomous-ai/ | 407 | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 264 | A skill your agent uses when writing or editing Pixi'VN story content — defining labels (scenes) with newLabel, writing dialogue steps, adding player choices with… | DRincs-Productions/ | 149 | — | ~4.9k | Automated safety check: Pass | LGPL-2.1 | 10 days ago |
| 265 | 265.Demo Video A skill your agent uses when asked to make a demo, training, how-to or tutorial video with sound or voice-over showing an HMIS function or configuration (e.g. | hmislk/ | 236 | — | ~3.9k | Automated safety check: Pass | GPL-3.0 | today |
| 266 | 在创作者明确确认后,执行短剧项目的图片、视频、TTS/配音或时间线音乐生产任务,并把结果与精简运行记录落回项目。用户说“生成这张图/这段视频/这句配音/这段配乐”“开始跑图/跑视频/合成语音/生成音乐”“把已确认提示词送去生产”,或要求批量执行已确认媒体任务时使用;不负责创作提示词、镜头、台词、歌词或声音身份,也绝不把预览、继续、预算说明或既有接受状态当作本次付费生产确认。 | zenstory-ai/ | 2.7k | — | ~1.9k | Automated safety check: Pass | MIT | 8 days ago |
| 267 | 用于“网页 demo 分段配音 + timeline 驱动录屏 + 后期合成”的 workspace 协作流程:先搭建一个可审计工作目录(cues/timeline/segmentaudio/video/subtitles/final),再由人类 + Codex 迭代维护这些文件,按需只重跑局部步骤,最终合成高质量 MP4。适用于强调可复盘、可编辑、清晰度与字幕安全区可控的场景。 | Sven-LI-sankyuu/ | 175 | — | ~3.2k | Automated safety check: Pass | No licence | 2 mo ago |
| 268 | 268.Abo PR Create Commit, push, and create Audiobook Organizer pull requests into protected master after verification and PR body preparation. | jeeftor/ | 190 | — | ~540 | Automated safety check: Pass | MIT | 1 mo ago |
| 269 | 269.Scene Splitter Splits a plain English story into a numbered list of SCENES — each scene being one moment that gets exactly one illustration AND one narration clip downstream. | hassancs91/ | 102 | — | ~2.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 270 | 把一篇技术长文/论文解读自动做成章节式解说视频(1080p, 5-8 分钟)。双主题:warm(奶油底+珊瑚红+cozy-handdrawn 透明插图,亲和感)和 midnight(深蓝黑底+琥珀金+宋体标题+executive-tech 插图,AI 科技感),storyboard 一个 theme 字段切换。每章三种 layout 混排:illustration(左文右图+Ken… | wwwzhouhui/ | 283 | — | ~2.2k | Automated safety check: Pass | No licence | 5 days ago |
| 271 | Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations. | glebis/ | 391 | — | ~2.3k | Automated safety check: Notes | MIT | 3 days ago |
| 272 | 272.Voice Isolator Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API. | elevenlabs/ | 482 | — | ~923 | Automated safety check: Pass | MIT | 2 days ago |
| 273 | 273.Audio Track A skill your agent uses when the user wants music, voiceover, narration, or a soundtrack added to a video asset, OR wants standalone generated audio for any purpose (e.g. | ucsandman/ | 251 | — | ~1.1k | Automated safety check: Notes | MIT | 1 mo ago |
| 274 | 274.Hig Inputs Apple HIG guidance for input methods and interaction patterns: gestures, Apple Pencil, keyboards, game controllers, pointers, Digital Crown, eye tracking, focus system, remotes, spatial… | raintree-technology/ | 144 | 5 repos | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 275 | 275.Elevenlabs Convert documents and text to audio using ElevenLabs text-to-speech. | sanjay3290/ | 432 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 276 | 276.Speech Engine Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. | elevenlabs/ | 482 | — | ~2.5k | Automated safety check: Warn | MIT | 2 days ago |
| 277 | 277.Keirouter Tts Text-to-speech via KeiRouter /v1/audio/speech using OpenAI / ElevenLabs / Deepgram / Edge TTS / Google TTS / Inworld voices. | mydisha/ | 147 | — | ~599 | Automated safety check: Pass | MIT | 1 mo ago |
| 278 | Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays… | cosmicstack-labs/ | 476 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 279 | Convert a single provided file into a downloadable Markdown file, representing source images and figures as readable text. | impredicative/ | 157 | — | ~174 | Automated safety check: Pass | LGPL-3.0 | 6 days ago |
| 280 | 280.Story Create high-quality MulmoScript through structured multi-phase creative process | receptron/ | 475 | — | ~3.5k | Automated safety check: Notes | No licence | today |
| 281 | Expert in building voice AI applications - from real-time voice agents to voice-enabled apps. | davila7/ | 33k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | today |
| 282 | Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media. | madebyaris/ | 126 | — | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 283 | Use the September 2026 image, video, speech and Avatar V adapters with explicit model/host contracts. | calesthio/ | 66k | — | ~1.4k | Automated safety check: Pass | AGPL-3.0 | 8 days ago |
| 284 | 284.Arkcli Deploy arkcli +deploy:普通创建推理接入点(Endpoint)的统一首选入口。用户说『创建/新建/create 一个 endpoint/接入点』或『部署/上线/deploy 某模型』时优先走这里;但脚本化 / CI / 无护栏 / 原始 raw CRUD 创建是唯一例外,必须改走 arkcli-infer-endpoint,不能由本 skill… | volcengine/ | 140 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 285 | 285.Voxclaw Give your agent a voice. An agent skill from malpern/VoxClaw. | malpern/ | 208 | — | ~1.9k | Automated safety check: Pass | No licence | 2 mo ago |
| 286 | Multi-stage workflow for producing videos set in the Three-Body universe, from lore research and storyboards to voice, sound and music generation and final compositing. | anbeime/ | 7.8k | — | ~1.6k | Automated safety check: Pass | No licence | yesterday |
| 287 | Generates media for intelligent textbooks - slide decks and presentations (MARP web decks in docs/slides/ or PowerPoint .pptx lecture downloads), illustrated stories and graphic novels… | dmccreary/ | 105 | — | ~1.7k | Automated safety check: Pass | CC-BY-NC-4.0 | yesterday |
| 288 | 288.Ig Reel Write an Instagram Reel from a raw idea - hook options off 26 formulas, the spoken script, the on-screen text, and a timed beat sheet - in the user's own voice and scored before they shoot it. | Jakeschincariol/ | 760 | — | ~1.6k | Automated safety check: Pass | MIT | 28 days ago |