Search
Speech recognition and synthesis
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Xybrid Init Generate model metadata for an ML model so it works with xybrid. | xybrid-ai/ | 469 | — | ~3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 50 | 50.Local Asr 把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell… | ysyecust/ | 273 | — | ~1.6k | Automated safety check: Pass | Unknown | 8 days ago |
| 51 | 51.Bili Note Turn Bilibili videos that are already registered and parsed in MindSpace, or Bilibili opus/article posts, into evidence-linked Markdown learning notes. | mingchen666/ | 244 | — | ~1.1k | Automated safety check: Pass | No licence | yesterday |
| 52 | 按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。 | zenstory-ai/ | 561 | — | ~1.3k | Automated safety check: Pass | MIT | 7 days ago |
| 53 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription. | deepgram/ | 276 | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 54 | Guides Python code that calls the Deepgram Management APIs to administer projects, keys, members, usage, billing and stored Voice Agent configurations. | deepgram/ | 469 | — | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 55 | 55.Model Scout A skill your agent uses when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search… | RisorseArtificiali/ | 118 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 56 | Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio. | buildfastwithai/ | 785 | — | ~1.7k | Automated safety check: Pass | MIT | 19 days ago |
| 57 | Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants). | reysu/ | 270 | — | ~3.8k | Automated safety check: Notes | MIT | 1 mo ago |
| 58 | 58.Audio2txt 从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。 | CoderWanFeng/ | 1.4k | — | ~230 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 59 | Integrate Egma into a repository containing voice agent code. | egma-ai/ | 144 | — | ~500 | Automated safety check: Pass | Unknown | 11 days ago |
| 60 | Transcribe audio via OpenAI Audio Transcriptions API (Whisper). | trpc-group/ | 1.9k | 12 repos | ~288 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 61 | Add or modify VAD providers and tuning in assistant-api with strict separation from EOS internals. | rapidaai/ | 745 | — | ~758 | Automated safety check: Pass | Unknown | 4 days ago |
| 62 | A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file. | NoizAI/ | 526 | — | ~917 | Automated safety check: Pass | No licence | 13 days ago |
| 63 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Management APIs for projects, API keys, members, invites, requests, usage, billing, models… | deepgram/ | 276 | — | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 64 | Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. | livekit-examples/ | 264 | 1 repo | ~2.4k | Automated safety check: Notes | MIT | yesterday |
| 65 | Covers basic transcription with the Deepgram Python SDK's listen.v1 endpoint, for one-shot REST transcription of a file or URL and live WebSocket streaming with interim results. | deepgram/ | 469 | — | ~2.9k | Automated safety check: Pass | MIT | 2 days ago |
| 66 | 66.Gaik Toolkit GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit. | GAIK-project/ | 100 | — | ~5.7k | Automated safety check: Pass | MIT | 2 days ago |
| 67 | A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified… | ysyecust/ | 273 | — | ~14k | Automated safety check: Notes | Unknown | 8 days ago |
| 68 | 68.Verify Verify a Txtify change end-to-end. An agent skill from lkmeta/txtify. | lkmeta/ | 135 | — | ~583 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 69 | Convert a long, messy, repetitive, speech-recognition-heavy task dump into a precise executable to-do list. | FAIRY123456789/ | 103 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 70 | Connect voice agents to WhatsApp Calling through Kapso. An agent skill from gokapso/agent-skills. | gokapso/ | 177 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 71 | 71.Watch Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path). | mathiaschu/ | 142 | — | ~4k | Automated safety check: Warn | MIT | 4 mo ago |
| 72 | Convert a HuggingFace ASR fine-tune into a sherpa-onnx external model, publish it, and add it to the Anti-Vocale community catalog. | RisorseArtificiali/ | 118 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 73 | 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. | zenstory-ai/ | 561 | — | ~1.1k | Automated safety check: Pass | MIT | 7 days ago |
| 74 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription. | deepgram/ | 276 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 75 | 75.Hri A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action. | terrense/ | 111 | — | ~268 | Automated safety check: Pass | MIT | 1 mo ago |
| 76 | 播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills. | chubbyguan/ | 1.2k | — | ~1.1k | Automated safety check: Notes | MIT | 3 days ago |
| 77 | Write or edit tests for a voice agent that can run on the Egma platform. | egma-ai/ | 144 | — | ~2.7k | Automated safety check: Pass | Unknown | 11 days ago |
| 78 | arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli… | volcengine/ | 140 | — | ~743 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 79 | 用于把 AI 生成的视频、本地素材、口播素材或产品短片剪成可发布到 TikTok 的竖屏成片. An agent skill from binggandata/bggg-skills. | binggandata/ | 605 | — | ~954 | Automated safety check: Pass | MIT | 1 mo ago |
| 80 | Explain and validate local setup paths for this repo with Docker and without Docker. | rapidaai/ | 745 | — | ~683 | Automated safety check: Pass | Unknown | 4 days ago |
| 81 | 81.App Builder App Builder — generate complete, runnable fullstack WebUI applications (FastAPI backend + pure HTML/CSS/JS frontend) around on-device AI Model Packs (OCR, TTS, ASR, Super-Resolution, etc.). | qualcomm/ | 247 | — | ~2.6k | Automated safety check: Pass | Unknown | yesterday |
| 82 | Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches. | sammcj/ | 162 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 83 | 83.Cut Video Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses. | louisedesadeleer/ | 108 | — | ~8.1k | Automated safety check: Warn | MIT | 3 mo ago |
| 84 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text Intelligence / Read (/v1/read) for sentiment, summarization, topic detection, and intent… | deepgram/ | 276 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 85 | Local speech-to-text with the Whisper CLI (no API key). An agent skill from huangruiteng/CS-Notes. | huangruiteng/ | 4k | 18 repos | ~228 | Automated safety check: Pass | MIT | 2 days ago |
| 86 | 86.Dashscope DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). | calesthio/ | 66k | — | ~1.5k | Automated safety check: Notes | AGPL-3.0 | 7 days ago |
| 87 | Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper. | deepgram/ | 469 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 88 | Build AI applications with OpenAI Agents SDK - text agents, voice agents, multi-agent handoffs, tools with Zod schemas, guardrails, and streaming. | coco-research/ | 531 | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 89 | 89.Agents SDK Build AI agents on Cloudflare Workers using the Agents SDK. An agent skill from hodgef/apiker. | hodgef/ | 127 | 3 repos | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 90 | A skill your agent uses when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast. | Varnan-Tech/ | 674 | — | ~1.9k | Automated safety check: Notes | MIT | yesterday |
| 91 | 91.Ax Audio This skill helps an LLM generate correct audio code with @ax-llm/ax. | dosco/ | 107 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 92 | Build a code-grounded implementation plan before coding. An agent skill from rapidaai/voice-ai. | rapidaai/ | 745 | — | ~727 | Automated safety check: Pass | Unknown | 4 days ago |
| 93 | 93.Docling PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown. | zhuzhaoyun/ | 433 | — | ~2.6k | Automated safety check: Pass | Unknown | yesterday |
| 94 | A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. | deepgram/ | 276 | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 95 | Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling. | deepgram/ | 469 | — | ~3.6k | Automated safety check: Pass | MIT | 2 days ago |
| 96 | Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models. | mydisha/ | 147 | — | ~680 | Automated safety check: Pass | MIT | 1 mo ago |