Topic · Media & Creative
Best text to speech and voice skills, page 10
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 433 | 433.Higgsfield Audio A skill your agent uses when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in… | OSideMedia/ | 713 | — | ~11k | Automated safety check: Pass | MIT | 13 days ago |
| 434 | 434.Voice TTS speech + mic/speaker mute for privacy. An agent skill from autonomous-ai/Physical-AI-Operating-System. | autonomous-ai/ | 407 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 435 | 435.Abo Tests Select, write, and run Audiobook Organizer verification for Go CLI, organizer, TUI, server/app, web build, documentation, and release hygiene changes. | jeeftor/ | 190 | — | ~537 | Automated safety check: Pass | MIT | 1 mo ago |
| 436 | 436.Tts Integration Add or modify text-to-speech providers in assistant-api with transport-aware behavior (WS/SSE/SDK/HTTP), correct packet lifecycle, and UI/provider wiring. | rapidaai/ | 745 | — | ~895 | Automated safety check: Pass | Unknown | 4 days ago |
| 437 | 437.Arkcli Onboard arkcli 接入向导(workflow):把某个模型接入到自己的应用/服务的端到端引导 —— 从'我想用某模型'到拿到可调用的 Endpoint(+ 可选示例代码)。当用户说'我想在我的 app/服务里用豆包/某模型''怎么把方舟模型接进来''帮我接入 XX 模型''想正式用上某模型'这类不含 deploy/部署关键词、但本质是正式接入的意图时触发。已明确说'部署/创建… | volcengine/ | 140 | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 438 | 438.Ppt Tts Script 将 PPT/PPTX/PDF 演示文档转换为拟人化逐字稿(Markdown + JSON 双格式输出),并将逐字稿回写到 PPTX 演讲者备注中。使用五阶段流水线:素材提取、大纲生成、逐字稿撰写(通过 Gemini CLI)、合并输出、备注回写。当用户需要将演示文档转为演讲稿、朗读稿、解说词时使用此技能。触发短语包括但不限于:「PPT 转语音」「生成逐字稿」「演讲稿生成」「帮我把 PPT… | ninehills/ | 280 | — | ~1.5k | Automated safety check: Pass | No licence | 3 mo ago |
| 439 | 439.Fish Audio Generate expressive audio clips using Fish Audio S2 TTS with bracket emotion tags. | vellum-ai/ | 1.4k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |
| 440 | Build real-time conversational AI voice engines using async worker pipelines, streaming transcription, LLM agents, and TTS synthesis with interrupt handling and multi-provider support | aiskillstore/ | 433 | 5 repos | ~5.7k | Automated safety check: Pass | No licence | yesterday |
| 441 | 441.Stepfun Asr Transcribes Chinese/English audio with StepFun's stepaudio-3-asr-max via its SSE endpoint (not /v1/audio/transcriptions) — one call handles long-form audio with no chunking. | daymade/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | today |
| 442 | 442.Stepfun Tts Generates Chinese/Japanese speech with StepFun's Contextual TTS — default stepaudio-2.5-tts, stepaudio-3-tts for whisper/inline-() prosody. | daymade/ | 1.4k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 443 | 443.Fal Lip Sync Create talking head videos and lip sync audio to video via fal.ai. | nexu-io/ | 100k | — | ~308 | Automated safety check: Pass | Apache-2.0 | today |
| 444 | 444.Speech Generate spoken audio from text using OpenAI's API with built-in voices. | nexu-io/ | 100k | — | ~291 | Automated safety check: Pass | Apache-2.0 | today |
| 445 | Text-to-speech models, voices, formats, and streaming via Venice.ai. | nexu-io/ | 100k | — | ~296 | Automated safety check: Pass | Apache-2.0 | today |
| 446 | Agent-callable HeyGen tools — generate AI avatar videos, translate and lip-sync videos, synthesize speech, and clone or browse voices and avatars. | zapier/ | 177 | — | ~4.4k | Automated safety check: Pass | Elastic-2.0 | 1 mo ago |
| 447 | 447.Magic Moment Make a "magic moment" video — turn a creator's talking-head recording into a vertical video preserving the source narration where their Muse story replays through artifacts and brief exchanges… | win4r/ | 343 | 1 repo | ~1.9k | Automated safety check: Pass | No licence | 13 days ago |
| 448 | A skill your agent uses when a PERSONAL brand needs AUDIO — voice cloning with ElevenLabs, Murf, or PlayHT, podcast production, audiobooks, and voiceover: short voiceover for TikTok and Reels, a 30… | minhnv0807/ | 609 | — | ~4k | Automated safety check: Pass | MIT | 28 days ago |
| 449 | 449.Opener Variator Rewrite subsection openers so they stop reading like a generated table-of-contents: remove \"overview/narration\" stems and reduce repeated opener cadences across H3s. | WILLOSCAR/ | 513 | — | ~1.3k | Automated safety check: Pass | No licence | 5 days ago |
| 450 | 450.Senpi Why Answer "what makes Senpi different?" / "why Senpi?" / "Senpi vs other trading apps, bots or AI chatbots?" — the positioning answer, led by value, not a feature dump. | Senpi-ai/ | 134 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 451 | 451.Txt2mp3 将文本内容转换为 MP3 语音文件,可选择直接朗读。当用户提到文字转语音、TTS、文本朗读、AI 配音时使用. An agent skill from CoderWanFeng/python-office. | CoderWanFeng/ | 1.4k | — | ~214 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 452 | 452.Demo Video A skill your agent uses when the user asks to create a demo video, product walkthrough, feature showcase, animated presentation, marketing video, or GIF from screenshots or scene descriptions. | alirezarezvani/ | 28k | — | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 453 | A skill your agent uses when setting up a new or existing project on @drincs/pixi-vn, wiring the main.ts entry point, or calling the top-level Game API (Game.init, Game.start, Game.onEnd… | DRincs-Productions/ | 149 | — | ~6.6k | Automated safety check: Pass | LGPL-2.1 | 9 days ago |
| 454 | 454.Slides Video Produce slides-driven narration videos (口播视频) where each slide maps 1:1 to one voiceover section. | sugarforever/ | 137 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 455 | A skill your agent uses when the user says 'teach me', 'explain as you go', 'mentor mode', 'walk me through', 'help me learn', 'explain why', 'learning mode', or wants real-time plain language… | cwinvestments/ | 423 | — | ~1.6k | Automated safety check: Pass | MIT | 14 days ago |
| 456 | 456.Comment Cleanup Delete and tighten code comments in source files after they are written. | haacked/ | 134 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 457 | 把中文口播文案做成像素风 + 坐标系隐喻 + 打字机字幕的动画解说视频(含 edge-tts 配音):先问 5 项输入、写分镜表与 3 张关键帧等确认,再写 spec.json、配音、逐帧渲染 1080p mp4。用户给文案并要像素风 / 注意力赌博那种风格、或确认是像素解说系列续集时使用;独立单片和其他风格走 $cs-code-video,暗夜星空知识片走 $cs-knowledge-film。 | ChenShuo2004/ | 197 | — | ~2.2k | Automated safety check: Pass | MIT | 3 days ago |
| 458 | 458.Abo Web UI Build, debug, and verify the current Audiobook Organizer local browser UI in web/, cmd/web.go, cmd/gui.go, internal/server, and internal/app. | jeeftor/ | 190 | — | ~502 | Automated safety check: Pass | MIT | 1 mo ago |
| 459 | 459.Tts Text-to-speech on Linux -- make the device speak text aloud. | mikeyobrien/ | 373 | — | ~405 | Automated safety check: Notes | MIT | 10 days ago |
| 460 | 460.LLM Eval Harness Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. | daymade/ | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | today |
| 461 | Stage 1 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 462 | Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted. | NVIDIA/ | 3.6k | — | ~3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 463 | 463.Audio Jingle Audio generation skill — jingles, beds, voiceover, and sound effects. | sanqiufong/ | 132 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 464 | 464.Stage Generate AI-generated footage plus execution of bounded semantic video-edit segments already approved by EDIT/AUTO. | Orkas-AI/ | 499 | — | ~1.7k | Automated safety check: Pass | MIT | 18 days ago |
| 465 | 465.Storyboard Turns user materials into a shot-by-shot storyboard with bilingual narration (Chinese + English VO, both required) and AI video prompts per shot. | godot-fun/ | 183 | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 466 | Muxes per-shot storyboard video with Chinese and English voice-over by retiming video to match VO duration (setpts), writing Video-Chinese/ and Video-English/ (matched by shot id 01, 02, …). | godot-fun/ | 183 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 467 | Turns a storyboard shot (video prompt + narration) into a self-contained fullscreen HTML CSS animation for the browser. | godot-fun/ | 183 | — | ~4.3k | Automated safety check: Pass | MIT | today |
| 468 | 468.Storyboard Tts Converts storyboard markdown (bilingual Chinese + English narration per shot) into speech audio via IndexTTS2 (same stack as ai-text-to-speech). | godot-fun/ | 183 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 469 | 469.Guided Demo A pattern for adding a self-narrating (typed out) guided walkthrough to any HTML/web application. | sammcj/ | 162 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 470 | 470.Media Production Create and process finished videos with Remotion, Motion Canvas, Manim, FFmpeg or an available AI video provider. | WrongStack/ | 371 | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 471 | 471.Hig Technologies Apple HIG guidance for Apple technology integrations: Siri, Apple Pay, HealthKit, HomeKit, ARKit, machine learning, generative AI, iCloud, Sign in with Apple, SharePlay, CarPlay, Game Center, in-app… | raintree-technology/ | 143 | — | ~1.9k | Automated safety check: Pass | MIT | 29 days ago |
| 472 | 472.To Video A skill your agent uses when approved video-planning artifacts must be handed to HyperFrames for production, with optional ListenHub narration. | sugarforever/ | 168 | — | ~743 | Automated safety check: Pass | MIT | yesterday |
| 473 | A skill your agent uses when the user wants a 王建硕-style WeChat article (article.md) turned into a narrated short MP4 video — TTS voiceover via 火山引擎 Volcano TTS, HyperFrames CSS/GSAP animation per… | jianshuo/ | 131 | — | ~5.5k | Automated safety check: Notes | MIT | 1 mo ago |
| 474 | A skill your agent uses when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. | jianshuo/ | 131 | — | ~5.3k | Automated safety check: Notes | MIT | 1 mo ago |
| 475 | Gates whether a code comment should exist and forces the ones that stay to explain why, not what. | PostHog/ | 40k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 476 | Generate a voiceover (VO) clip via ElevenLabs text-to-speech, ROUTED THROUGH THE elevenlabs-proxy so it bills the Ads agent. | gooseworks-ai/ | 1.2k | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 477 | Render a 'brand identity reveal' video from a config — a single poster frame in a real, softly-lit space (real wall, soft-focus plant in the corner, dappled leaf shadow, illuminated poster) whose… | gooseworks-ai/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 478 | Assemble a cosmic-mythology-voiceover reel from a config — one spoken voiceover carries the whole narrative while N curated stills in one chosen look are weighted beat-synced across the delivered VO… | gooseworks-ai/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 479 | Render a 'hand-swipe flavor listicle' video from a config — the brand's REAL product cutouts each float on a flat flavor-matched color field under a persistent title and slide RIGHT-TO-LEFT one to… | gooseworks-ai/ | 1.2k | — | ~2.9k | Automated safety check: Pass | MIT | 2 days ago |
| 480 | Render an 'Instagram-Live social-proof gallery' video from a config — ~5 real brand product stills each framed as an Instagram-LIVE card (IG gradient-ring avatar, username, verified check, red LIVE… | gooseworks-ai/ | 1.2k | — | ~1k | Automated safety check: Pass | MIT | 2 days ago |
Explore related skills
Category
More topics in Media & Creative
- Video production890
- Image generation724
- AI video generation667
- Transcription569
- Motion graphics387
- Design review and critique240
- Logo and visual identity226
- Comics and storyboards219
- Image editing193
- Social media graphics190
- Video scripts and shorts180
- Infographics159
- Music and audio generation150
- Podcasting122
- Generative and creative coding60