Topic · AI & LLM Engineering
Best speech recognition and synthesis skills for Claude Code, Codex and other agents.
- skills
- 272
- official
- 18
Speech recognition and synthesis skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model. | bradautomates/ | 18k | — | ~4.3k | Automated safety check: Notes | MIT | 12 days ago |
| 2 | Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI. | Orchestra-Research/ | 13k | 8 repos | ~1.9k | Automated safety check: Notes | MIT | 3 mo ago |
| 3 | 3.Triage Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang. | TalAter/ | 6.8k | — | ~810 | Automated safety check: Notes | MIT | today |
| 4 | 逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct… | mcncarl/ | 4.3k | — | ~780 | Automated safety check: Pass | Unknown | 3 days ago |
| 5 | 5.Dy Note DyNote: systematically and efficiently extract raw Douyin/DY video data and analyze videos, comments, accounts, hashtags, and short-video scenes into evidence-graded learning notes, summaries… | Rimagination/ | 172 | 1 repo | ~4.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | 钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。 | DingTalk-Real-AI/ | 3.2k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 7 | Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop. | tadaspetra/ | 296 | 3 repos | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 8 | Joins Google Meet, Teams or Zoom video calls as an AI bot with voice and visual presence through the AgentCall service, in audio, text-to-speech or webpage modes. | pattern-ai-labs/ | 165 | 1 repo | ~25k | Automated safety check: Pass | MIT | 21 days ago |
| 9 | Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video. | HUANGCHIHHUNGLeo/ | 2.2k | — | ~639 | Automated safety check: Pass | MIT | 6 days ago |
| 10 | Build conversational AI voice agents on the ElevenLabs platform. | jezweb/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 11 | Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others. | decolua/ | 30k | — | ~745 | Automated safety check: Pass | MIT | 6 days ago |
| 12 | 逸尘自用的互联网研究总入口。用于跨平台且跨阶段、用户尚未确定工具,或明确要求对公司、产品、人物、技术、行业和领域做横纵分析、发展史加现状对比或有来源约束的系统深度研究;先生成有截止日期和证据闸门的计划,再把搜索发现、候选核验、有限归档、按需转写和证据综合路由到… | mcncarl/ | 4.3k | — | ~1.9k | Automated safety check: Pass | Unknown | 3 days ago |
| 13 | Install and use crv (claude-real-video) — a tool that lets any AI agent watch videos by extracting scene-aware keyframes, deduplicating them, and transcribing audio. | HUANGCHIHHUNGLeo/ | 2.2k | — | ~2k | Automated safety check: Notes | MIT | 6 days ago |
| 14 | Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass. | AI-Builder-Club/ | 1.3k | — | ~850 | Automated safety check: Pass | No licence | 21 days ago |
| 15 | Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers. | YaoApp/ | 8.1k | — | ~416 | Automated safety check: Pass | Unknown | 2 days ago |
| 16 | A skill your agent uses when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast. | Varnan-Tech/ | 672 | 1 repo | ~1.9k | Automated safety check: Notes | MIT | 1 mo ago |
| 17 | Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS. | ysyecust/ | 269 | — | ~783 | Automated safety check: Pass | Unknown | 4 days ago |
| 18 | Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants. | KIRVO-REPORTING/ | 105 | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 19 | AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill. | speechlab0210/ | 105 | — | ~4.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 20 | Analyze, summarize, and extract insights from DeLive transcription sessions. | XimilalaXiang/ | 281 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 21 | Add or modify end-of-speech integrations in assistant-api with strict separation from VAD internals. | rapidaai/ | 744 | — | ~877 | Automated safety check: Pass | Unknown | today |
| 22 | Build complete, working voice agent applications using the gradbot framework. | gradium-ai/ | 126 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 23 | Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python. | deepgram/ | 469 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 24 | 24.Agents Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop. | tadaspetra/ | 296 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 4 days ago |
| 25 | Guide the user through installing, configuring, and launching voxtype — local on-device voice dictation (speech-to-text that types wherever the cursor is). | pchalasani/ | 2k | — | ~857 | Automated safety check: Notes | MIT | 2 days ago |
| 26 | Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0. | bozhouDev/ | 149 | — | ~1.8k | Automated safety check: Notes | MIT | 2 mo ago |
| 27 | Lilly community-research skill. An agent skill from ssaaffaakk/Lilly. | ssaaffaakk/ | 171 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 28 | 把课堂视频(本地或 B 站/YouTube)、文字稿、课件三者(任意组合)整理成一份详细的中文 Markdown 课堂笔记,输出按课程标题命名的 {titlename}.md(首行为 文档标题)+ 相对路径图片。Markdown 工作流,与上游 lecture-to-notes 的 LaTeX/PDF 输出并行存在;上游 skill 完全不动。触发词:markdown 笔记、md 笔记、视频转… | ysyecust/ | 269 | — | ~3.9k | Automated safety check: Pass | Unknown | 4 days ago |
| 29 | Explain and validate local setup paths for this repository with Docker and without Docker. | rapidaai/ | 744 | — | ~595 | Automated safety check: Pass | Unknown | today |
| 30 | YouTube 영상에서 쇼츠 클립 3개를 자동 제작합니다. An agent skill from uxjoseph/content-marketing-team. | uxjoseph/ | 107 | — | ~665 | Automated safety check: Pass | No licence | 9 mo ago |
| 31 | 31.Test Model Test a model end-to-end using the xybrid execution system. An agent skill from xybrid-ai/xybrid. | xybrid-ai/ | 465 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands. | modelstudioai/ | 541 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 33 | Run local black-box voice tests against the physical Dotty robot by playing a TTS prompt through the workstation speakers while the C920 records video and room audio. | BrettKinny/ | 113 | — | ~866 | Automated safety check: Notes | MIT | yesterday |
| 34 | End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and… | GoogleCloudPlatform/ | 106 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | 35.Murmur 把一段中文(或任意 Whisper 支持语言)的会议/面试录音用本地 Whisper large-v3 转成文本,再清洗成带说话人标签、修过 ASR 错字、分好章节的 markdown 文档(可选再转成 docx)。跨平台(macOS Apple Silicon 用 mlx-whisper,Windows/Linux/Intel Mac 用… | xiaopengde/ | 109 | — | ~2.9k | Automated safety check: Notes | MIT | 4 mo ago |
| 36 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize… | deepgram/ | 276 | — | ~1.5k | Automated safety check: Pass | MIT | 5 days ago |
| 37 | 37.Video Cut 把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。 | zenstory-ai/ | 553 | — | ~1.6k | Automated safety check: Pass | MIT | 3 days ago |
| 38 | Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection. | deepgram/ | 469 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 39 | 39.Deps Bump Safely update Txtify dependencies or resolve Dependabot alerts. | lkmeta/ | 135 | — | ~585 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 40 | 40.Autoshorts Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a… | Upload-Post/ | 150 | — | ~5.3k | Automated safety check: Notes | MIT | 5 mo ago |
| 41 | Build voice AI agents with LiveKit Cloud and the Agents SDK. | allgpt-co/ | 488 | — | ~3.3k | Automated safety check: Notes | MIT | yesterday |
| 42 | 42.Watch Watch a video (URL or local path) like an editor. An agent skill from taoufik123-collab/claude-watch. | taoufik123-collab/ | 930 | — | ~5.8k | Automated safety check: Warn | MIT | 2 mo ago |
| 43 | AI Agent自动剪辑旅行Vlog的完整工作流。从原始素材到成品视频,系统级只需ffmpeg,其余在Python venv内完成。by nyx研究所 (GitHub @znyupup · B站/小红书 @nyx研究所) | znyupup/ | 144 | — | ~6.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 44 | Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images. | linzzzzzz/ | 569 | — | ~2.8k | Automated safety check: Warn | MIT | 1 mo ago |
| 45 | 45.Intro Video Build Remotion intro, reel, and brand-film videos with remocn and ASR-timed captions. | TheOrcDev/ | 106 | — | ~1.6k | Automated safety check: Pass | No licence | 9 days ago |
| 46 | Build a code-grounded implementation plan before coding. An agent skill from rapidaai/voice-ai. | rapidaai/ | 744 | — | ~721 | Automated safety check: Pass | Unknown | today |
| 47 | 47.Local AI Use Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API. | amd/ | 395 | — | ~5k | Automated safety check: Notes | MIT | today |
| 48 | 48.Xybrid Init Generate model metadata for an ML model so it works with xybrid. | xybrid-ai/ | 465 | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
Questions, answered from the data.
What is the best speech recognition and synthesis skill?
Watch Video Q&A from bradautomates/claude-video ranks first of the 272 speech recognition and synthesis skills listed here, with the highest score: its repository has 18k GitHub stars, its SKILL.md loads about 4.3k tokens and it has informational notes only in the automated safety check. Next come Whisper Speech Recognition and Triage.
Which speech recognition and synthesis skills are official?
18 of the 272 speech recognition and synthesis skills are official, published by the vendor's own GitHub organization: Transformers.js, Azure AI Voicelive Py, Azure AI Voicelive TS, Azure Speech To Text REST Py, Azure AI Openai Dotnet and 13 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23