Topic · AI & LLM Engineering
Best speech recognition and synthesis skills, page 2
Speech recognition and synthesis skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Local AI Use Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API. | amd/ | 398 | — | ~5k | Automated safety check: Notes | MIT | today |
| 50 | 50.Xybrid Init Generate model metadata for an ML model so it works with xybrid. | xybrid-ai/ | 466 | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 51 | 51.Local Asr 把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell… | ysyecust/ | 270 | — | ~1.6k | Automated safety check: Pass | Unknown | 5 days ago |
| 52 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription. | deepgram/ | 276 | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 53 | 53.Bili Note Turn Bilibili videos that are already registered and parsed in MindSpace, or Bilibili opus/article posts, into evidence-linked Markdown learning notes. | mingchen666/ | 237 | — | ~1.1k | Automated safety check: Pass | No licence | 17 days ago |
| 54 | Guides Python code that calls the Deepgram Management APIs to administer projects, keys, members, usage, billing and stored Voice Agent configurations. | deepgram/ | 469 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 55 | 55.Model Scout A skill your agent uses when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search… | RisorseArtificiali/ | 117 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 56 | Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio. | buildfastwithai/ | 785 | — | ~1.7k | Automated safety check: Pass | MIT | 16 days ago |
| 57 | Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants). | reysu/ | 270 | — | ~3.8k | Automated safety check: Notes | MIT | 1 mo ago |
| 58 | Transcribe audio via OpenAI Audio Transcriptions API (Whisper). | trpc-group/ | 1.8k | 13 repos | ~288 | Automated safety check: Pass | Apache-2.0 | today |
| 59 | 59.Audio2txt 从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。 | CoderWanFeng/ | 1.4k | — | ~230 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 60 | Integrate Egma into a repository containing voice agent code. | egma-ai/ | 144 | — | ~500 | Automated safety check: Pass | Unknown | 8 days ago |
| 61 | Add or modify VAD providers and tuning in assistant-api with strict separation from EOS internals. | rapidaai/ | 744 | — | ~758 | Automated safety check: Pass | Unknown | 2 days ago |
| 62 | 62.Video Cut 把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。 | zenstory-ai/ | 555 | — | ~1.6k | Automated safety check: Pass | MIT | 4 days ago |
| 63 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Management APIs for projects, API keys, members, invites, requests, usage, billing, models… | deepgram/ | 276 | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 64 | Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. | livekit-examples/ | 264 | 1 repo | ~2.4k | Automated safety check: Notes | MIT | 2 days ago |
| 65 | Covers basic transcription with the Deepgram Python SDK's listen.v1 endpoint, for one-shot REST transcription of a file or URL and live WebSocket streaming with interim results. | deepgram/ | 469 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 66 | 66.Gaik Toolkit GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit. | GAIK-project/ | 100 | — | ~5.7k | Automated safety check: Pass | MIT | yesterday |
| 67 | A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified… | ysyecust/ | 270 | — | ~14k | Automated safety check: Notes | Unknown | 5 days ago |
| 68 | 68.Verify Verify a Txtify change end-to-end. An agent skill from lkmeta/txtify. | lkmeta/ | 135 | — | ~583 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 69 | Convert a long, messy, repetitive, speech-recognition-heavy task dump into a precise executable to-do list. | FAIRY123456789/ | 103 | — | ~1.1k | Automated safety check: Pass | MIT | 8 days ago |
| 70 | Connect voice agents to WhatsApp Calling through Kapso. An agent skill from gokapso/agent-skills. | gokapso/ | 175 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 71 | 71.Watch Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path). | mathiaschu/ | 141 | — | ~4k | Automated safety check: Warn | MIT | 4 mo ago |
| 72 | A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file. | NoizAI/ | 526 | — | ~917 | Automated safety check: Pass | No licence | 10 days ago |
| 73 | Convert a HuggingFace ASR fine-tune into a sherpa-onnx external model, publish it, and add it to the Anti-Vocale community catalog. | RisorseArtificiali/ | 117 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 74 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription. | deepgram/ | 276 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 75 | 75.Hri A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action. | terrense/ | 111 | — | ~268 | Automated safety check: Pass | MIT | 1 mo ago |
| 76 | Write or edit tests for a voice agent that can run on the Egma platform. | egma-ai/ | 144 | — | ~2.7k | Automated safety check: Pass | Unknown | 8 days ago |
| 77 | arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli… | volcengine/ | 140 | — | ~743 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 78 | 用于把 AI 生成的视频、本地素材、口播素材或产品短片剪成可发布到 TikTok 的竖屏成片. An agent skill from binggandata/bggg-skills. | binggandata/ | 603 | — | ~954 | Automated safety check: Pass | MIT | 1 mo ago |
| 79 | Explain and validate local setup paths for this repo with Docker and without Docker. | rapidaai/ | 744 | — | ~683 | Automated safety check: Pass | Unknown | 2 days ago |
| 80 | 80.App Builder App Builder — generate complete, runnable fullstack WebUI applications (FastAPI backend + pure HTML/CSS/JS frontend) around on-device AI Model Packs (OCR, TTS, ASR, Super-Resolution, etc.). | qualcomm/ | 246 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 81 | Local speech-to-text with the Whisper CLI (no API key). An agent skill from huangruiteng/CS-Notes. | huangruiteng/ | 4k | 19 repos | ~228 | Automated safety check: Pass | MIT | 2 days ago |
| 82 | Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. | sammcj/ | 162 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 83 | 83.Cut Video Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses. | louisedesadeleer/ | 108 | — | ~8.1k | Automated safety check: Warn | MIT | 3 mo ago |
| 84 | 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. | zenstory-ai/ | 555 | — | ~1.1k | Automated safety check: Pass | MIT | 4 days ago |
| 85 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text Intelligence / Read (/v1/read) for sentiment, summarization, topic detection, and intent… | deepgram/ | 276 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 86 | 86.Dashscope DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). | calesthio/ | 65k | — | ~1.5k | Automated safety check: Notes | AGPL-3.0 | 5 days ago |
| 87 | Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper. | deepgram/ | 469 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 88 | 88.Agents SDK Build AI agents on Cloudflare Workers using the Agents SDK. An agent skill from hodgef/apiker. | hodgef/ | 127 | 3 repos | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 89 | 89.Ax Audio This skill helps an LLM generate correct audio code with @ax-llm/ax. | dosco/ | 107 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 90 | Build a code-grounded implementation plan before coding. An agent skill from rapidaai/voice-ai. | rapidaai/ | 744 | — | ~727 | Automated safety check: Pass | Unknown | 2 days ago |
| 91 | 91.Docling PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown. | zhuzhaoyun/ | 432 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 92 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. | deepgram/ | 276 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 93 | 播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills. | chubbyguan/ | 1.2k | — | ~1.1k | Automated safety check: Notes | MIT | 2 days ago |
| 94 | Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling. | deepgram/ | 469 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 95 | Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models. | mydisha/ | 147 | — | ~680 | Automated safety check: Pass | MIT | 27 days ago |
| 96 | 96.Vibe To Spec Convert messy voice input, imperfect speech recognition, half-formed product ideas, and iterative corrections into an implementation-ready software specification for Codex, Claude Code, Copilot… | FAIRY123456789/ | 103 | — | ~997 | Automated safety check: Pass | MIT | 8 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23