Topic · Media & Creative
Best text to speech and voice skills for Claude Code, Codex and other agents.
- skills
- 629
- official
- 14
Text to speech and voice skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades. | heygen-com/ | 58k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Walks through adding a new text-to-speech engine to Voicebox end to end: dependency audit, backend, frontend wiring, PyInstaller bundling and frozen-build testing. | jamiepine/ | 57k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 3 | Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music. | harry0703/ | 129k | — | ~2.1k | Automated safety check: Warn | MIT | yesterday |
| 4 | A skill your agent uses when making a demo, tutorial, or feature video of Nuclear. | nukeop/ | 19k | — | ~892 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 5 | Automates JianYing Pro video editing through a Python wrapper, JyWrapper, covering drafts, media import, subtitles, screen recording, voiceover and export. | luoluoluo22/ | 3.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 26 days ago |
| 6 | Routes agents to the right KrillinAI command for subtitles, dubbing, video rendering, covers and speech, and explains how to read its JSON and manifest output. | krillinai/ | 13k | — | ~869 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 7 | Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script. | bytedance/ | 83k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 8 | Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all… | Alisa0808/ | 2.2k | — | ~5.6k | Automated safety check: Pass | MIT | 2 days ago |
| 9 | A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice… | google-gemini/ | 4.3k | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat… | Agents365-ai/ | 1.7k | — | ~4.9k | Automated safety check: Pass | MIT | 6 days ago |
| 11 | 11.Music Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop. | tadaspetra/ | 296 | 3 repos | ~827 | Automated safety check: Pass | MIT | 5 days ago |
| 12 | 12.Hyperframes Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. | boraoztunc/ | 393 | 9 repos | ~7.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 13 | Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work. | SAP/ | 227 | 25 repos | ~3.5k | Automated safety check: Pass | MIT | yesterday |
| 14 | TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK. | trypostit/ | 676 | 2 repos | ~3.5k | Automated safety check: Pass | MIT | today |
| 15 | 15.Blog Audio Generate audio narration of blog posts using Google Gemini TTS. | AgriciDaniel/ | 2.3k | 1 repo | ~2.2k | Automated safety check: Notes | MIT | 5 days ago |
| 16 | 16.Videodb Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable… | affaan-m/ | 274k | 3 repos | ~3.5k | Automated safety check: Notes | MIT | 2 days ago |
| 17 | AI 电影/短剧解说视频自动生成(AI 解说大师 CLI Skill)。当用户需要创建电影解说视频、短剧解说、影视二创、AI 配音旁白视频、film commentary、video narration、drama dubbing、movie narration 时触发。内置电影素材库、BGM、多语种配音、解说模板。通过 narrator-ai-cli 命令行实现:搜片→选模板→选… | NarratorAI-Studio/ | 3k | — | ~4.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 18 | Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop. | tadaspetra/ | 296 | 3 repos | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 19 | Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references. | dtsola/ | 1k | — | ~3.4k | Automated safety check: Pass | MIT | 5 days ago |
| 20 | 终极口播视频 skill:中文口播稿 + 成品配音 → CPU 字级时间戳 → SHOTBOOK 层矩阵分镜 → Remotion 电影感成片(横屏默认/竖屏)。当用户要"做口播视频"、"解说/科普视频"、"把文案变成视频"、"给配音配画面动效"时使用。默认使用成品配音,可选 Fish Audio 从稿子合成配音与时间戳;数字人生成技术不在本 skill… | Vincentwei1021/ | 1.4k | — | ~8.7k | Automated safety check: Notes | Unknown | 6 days ago |
| 21 | Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion. | Agents365-ai/ | 1.7k | — | ~4.5k | Automated safety check: Pass | MIT | 6 days ago |
| 22 | Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor. | chmonitor/ | 298 | 5 repos | ~1.3k | Automated safety check: Pass | GPL-3.0 | 2 days ago |
| 23 | Validates a KrillinAI multi-stage output plan with the pipeline command's dry run, then maps it to the individual stage commands that do the real work. | krillinai/ | 13k | — | ~488 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 24 | A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot… | wy51ai/ | 1.4k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 9 days ago |
| 25 | Joins Google Meet, Teams or Zoom video calls as an AI bot with voice and visual presence through the AgentCall service, in audio, text-to-speech or webpage modes. | pattern-ai-labs/ | 165 | 1 repo | ~25k | Automated safety check: Pass | MIT | 22 days ago |
| 26 | Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt. | hoquanghai/ | 319 | 1 repo | ~3.7k | Automated safety check: Pass | MIT | 5 mo ago |
| 27 | Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports. | NousResearch/ | 252k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 28 | 给一个主题,产出一条黑底 MG 风格(幕底可选星点或点阵波)、有配音字幕章节进度条的科普讲解视频(中文或英文;Remotion 代码动画;时长由用户定,常用 3–5 分钟)。内含可编译模板、图元库、配音/分镜/渲染工具、风格与动效规范、多 agent 分工协议与 QC 判据,以及一条完整样片(《RAG 与知识库》)作为质量标尺。Turn any topic into a narrated… | Vincentwei1021/ | 2.3k | — | ~2.7k | Automated safety check: Pass | Unknown | 19 days ago |
| 29 | Build conversational AI voice agents on the ElevenLabs platform. | jezweb/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 30 | Dub a video into another language and generate subtitles using the default Together + Cartesia stack. | shang-zhu/ | 1.1k | — | ~1k | Automated safety check: Notes | MIT | 1 mo ago |
| 31 | Real time video translation / dubbing skill. An agent skill from InsiderX-Pro/video-translator. | InsiderX-Pro/ | 928 | — | ~635 | Automated safety check: Pass | MIT | 3 mo ago |
| 32 | Guides an approval-gated video editing pipeline from raw footage to an audited DaVinci Resolve timeline, with scripting, optional TTS and a blueprint at each stage. | liuluhaixiu/ | 480 | — | ~2.8k | Automated safety check: Notes | MIT | 3 mo ago |
| 33 | Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style. | digitalsamba/ | 2.2k | 1 repo | ~2.7k | Automated safety check: Notes | MIT | 2 days ago |
| 34 | Renders a source video as a portrait video with the KrillinAI CLI, with short bilingual subtitles or a dubbed audio track, then checks the result. | krillinai/ | 13k | — | ~527 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 35 | 35.Xiaoai Tts Control Xiaoai speaker via OpenXiaoAI Voice API for high-quality TTS playback. | coderzc/ | 386 | — | ~749 | Automated safety check: Pass | MIT | 13 days ago |
| 36 | 36.Floe Guard Know what every AI call really costs — floe-guard meters STT + TTS + LLM + telephony per call (Pipecat, LiveKit — Python & TypeScript), keeps a live ledger of real spend, and hard-stops the next… | Floe-Labs/ | 357 | — | ~853 | Automated safety check: Pass | MIT | 2 days ago |
| 37 | 把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在… | itwanger/ | 18k | — | ~1.1k | Automated safety check: Pass | No licence | today |
| 38 | Generate sound effects from text descriptions using ElevenLabs. | tadaspetra/ | 296 | 2 repos | ~1.1k | Automated safety check: Pass | MIT | 5 days ago |
| 39 | Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg… | limin112/ | 359 | — | ~2.4k | Automated safety check: Pass | No licence | 14 days ago |
| 40 | Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound… | chmonitor/ | 298 | 1 repo | ~2.8k | Automated safety check: Notes | GPL-3.0 | 2 days ago |
| 41 | 通用的多账号图书短视频生产工作流。用于用户希望建立图书号项目目录、配置账号级片头/声音/BGM/视觉规范,或只提供一本书后依次完成资料研究、口播稿、分镜、图片、配音、字幕、预览与成片导出。适用于新建工作区、批量管理多个账号、继续已有单书任务和检查生产状态;不绑定特定研究、图片、TTS、转录或视频渲染供应商。 | bytec-ai/ | 321 | — | ~1.4k | Automated safety check: Notes | No licence | 2 mo ago |
| 42 | Guide a creator through a real couple's custom wedding video, from a shareable story intake card and Kimi writing pack through narration, external GPT image prompts, image-to-video packs, music and… | aaronyi97/ | 288 | — | ~1k | Automated safety check: Pass | MIT | 27 days ago |
| 43 | HyperFrames CLI tool — hyperframes init, lint, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark. | nateherkai/ | 1.2k | 3 repos | ~1.2k | Automated safety check: Pass | Unknown | 9 days ago |
| 44 | Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers. | decolua/ | 30k | — | ~765 | Automated safety check: Pass | MIT | 6 days ago |
| 45 | AivisSpeech エディタの Sentry issue を調査し、修正すべき Electron / Vue 側の不具合と、ローカル環境・ブラウザ実装・外部通信由来のノイズを切り分けるためのスキルです。Sentry 側で既知ノイズを整理する作業や、src/domain/sentryEventFilter.ts と関連テストを更新して既知ノイズを送信前に破棄する作業で使用します。 | Aivis-Project/ | 483 | — | ~636 | Automated safety check: Pass | LGPL-3.0 | 2 mo ago |
| 46 | 46.Video Script 对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选… | zenstory-ai/ | 460 | — | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 47 | Analyze images, video, speech, motion, products, and websites; route local vision, audio, depth, tracking, matting, identity, and restoration models; auto-edit, replicate, enhance, caption, voice… | MartinDelophy/ | 891 | — | ~7k | Automated safety check: Pass | MIT | today |
| 48 | Diagnose and revise defensive academic writing while preserving claim ceilings, evidence status, scope conditions, rival explanations, and conceptual hierarchy. | lensback940701/ | 255 | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
Questions, answered from the data.
What is the best text to speech and voice skill?
HyperFrames Media Use from heygen-com/hyperframes ranks first of the 629 text to speech and voice skills listed here, with the highest score: its repository has 58k GitHub stars, 2 other GitHub owners carry a copy, its SKILL.md loads about 2.1k tokens and it passes the automated safety check with no findings. Next come Add TTS Engine to Voicebox and MoneyPrinterTurbo Video Generator.
Which text to speech and voice skills are official?
14 of the 629 text to speech and voice skills are official, published by the vendor's own GitHub organization: Gemini API Dev, Openspec Onboard, Speech, Azure Realtime Podcast Generation, Elevenlabs and 9 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in Media & Creative
- Video production887
- Image generation726
- AI video generation635
- Transcription543
- Motion graphics377
- Design review and critique238
- Logo and visual identity229
- Comics and storyboards225
- Image editing192
- Social media graphics191
- Video scripts and shorts185
- Infographics160
- Music and audio generation159
- Podcasting127
- Generative and creative coding61