Audio Transcription
mitsuhiko/agent-stuff
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.
“对大量语音转写稿进行校对、整理、分段处理,支持断点续传和恢复”
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/long-audio-transcript-processor .claude/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .claude/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/long-audio-transcript-processor .agents/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .agents/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/long-audio-transcript-processor .cursor/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .cursor/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cafe3310/public-agent-skills.git --path skills/long-audio-transcript-processor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/long-audio-transcript-processor .gemini/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .gemini/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/long-audio-transcript-processor .github/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .github/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/long-audio-transcript-processor .opencode/skills/long-audio-transcript-processor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "long-audio-transcript-processor" agent skill from https://github.com/cafe3310/public-agent-skills/tree/main/skills/long-audio-transcript-processor into .opencode/skills/long-audio-transcript-processor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "long-audio-transcript-processor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
long-audio-transcript-processorLong Audio Transcript Processor is a skill in cafe3310/public-agent-skills (255 stars). Its SKILL.md is about 1k tokens, with 6 other files in the folder (scripts, assets). Licence: Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c45501. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Long Audio Transcript Processor loads about 1k tokens when it runs. Until then it costs about 16 tokens; SKILL.md has 212 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from cafe3310/public-agent-skills at commit 6c45501, republished under its Apache-2.0 licence (© cafe3310). 212 words, ~1,041 tokens.
.claude/skills/long-audio-transcript-processor/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.此技能旨在通过文件系统辅助,安全、有序地处理超长语音转写文本。它通过分段处理、上下文维护(术语表、主题记录)和状态追踪,确保处理过程的可持续性和高质量。 此技能最好使用最高性能的模型(而不是写代码用的快速模型)以确保最佳质量。
当用户提供一个或多个长篇语音转写文件,并要求进行:
首先,必须初始化工作区。询问用户是否已准备好源文件。
运行初始化脚本:
python3 .gemini/skills/long-audio-transcript-processor/scripts/setup_workspace.py "path/to/file1.txt" "path/to/file2.txt" ...(注意:请根据实际技能安装路径调整脚本路径,通常是 .gemini/skills/...)
初始化后,工作区结构如下:
语音转写处理_YYYY-MM-DD-HH-MM/
├── 0-工作日志.md # 进度追踪与计划
├── 1-原始文件/ # 存放用户提供的原始语音文本
├── 2-要求和信息/ # 存放活动背景、发言人等信息(用户补充)
├── 3-校对和术语表.md # 动态更新的术语库和错误模式
├── 4-分段主题.md # 记录已处理分段的主题脉络
└── 5-最终输出/ # 存放校对完成的分段文件关键操作:
2-要求和信息/ 目录下,也在该目录下创建 Markdown 文档记录用户的说明。这是保证后续处理准确性的基石。在进入循环前,总是先读取以下文件以加载上下文(确保跨分段的信息一致性):
0-工作日志.md (检查进度)2-要求和信息/ 下的所有背景和要求文档3-校对和术语表.md (加载最新积累的术语和校对规则)4-分段主题.md (加载已有上下文主题)5-最终输出/ 下的文件 -- 列出文件名即可步骤:
0-工作日志.md 中找到第一个未完成([ ])的分段。sed 命令从 1-原始文件/ 中提取对应行范围,并重定向写入到 5-最终输出/ 下的对应文件中。sed -n '开始行,结束行p' "1-原始文件/文件名.txt" > "5-最终输出/文件名_开始行-结束行.txt"read_file 读取上一步生成的 5-最终输出/ 下的文件内容。(...)。write_file 将校对后的完整文本写回 5-最终输出/ 的对应文件(覆盖掉刚才的底稿)。3-校对和术语表.md。仅追加,用行号段落区分不同分段的内容。4-分段主题.md。仅追加,用行号段落区分不同分段的内容。0-工作日志.md 中标记分段为 [x]。3-校对和术语表.md 并修正 5-最终输出 中的对应文件。如果对话中断,不要 重新初始化。
直接执行 分段处理循环 的“在进入循环前”步骤:通过读取 2-要求和信息/ 和 3-校对和术语表.md 完整找回记忆。
然后继续下一个未完成的分段。
除了标准转写外,用户可能要求并行生成其他产物(如 Q&A 问答库、摘要、待办事项)。
2-要求和信息/ 下创建任务说明文档(如 额外任务_问答积累.md)。5-最终输出/ 下,文件命名应清晰(如 问题和回答-主题.md)。在所有分段处理完成后,进入此阶段。
成果汇总与提醒:
订正处理标准:
{🔴 原始内容}:标记被修改或删除的原始文本。{🟢 新内容}:标记新增或修正后的文本。{‼️ 需特别留意,可能出错}:标记模型认为存在矛盾、风险或不确定的地方,或者用户特别强调的注意点。3-校对和术语表.md 和 5-最终输出/ 中的对应文件,确保下一次处理或合并时使用最新数据的正确版本。0-工作日志.md: 核心状态文件。必须 保持最新。3-校对和术语表.md: 动态更新的知识库。发现新术语或特定错误模式时,务必更新此文件,以保证后续分段处理的一致性。该文件的编辑也遵循只追加原则。4-分段主题.md: 帮助 LLM 保持对长文本整体脉络的理解。3-校对和术语表, 4-分段主题.md 和任何额外输出文件,严格遵循 只追加、不修改的原则,避免在 loop 过程中,覆盖已有内容。2-要求和信息/ 中,并通知后续处理环节。同行 (测试)、访客 (学生))。禁止更改时间戳行,识别的角色必须添加在发言行行首。发言人 2 01:30:00
这是一个验证。发言人 2 01:30:00
(测试负责人-老张):这是一个验证。© cafe3310, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, assets) in skills/long-audio-transcript-processor of cafe3310/public-agent-skills.
Open the folder on GitHubat commit 6c45501
Long Audio Transcript Processor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Long Audio Transcript Processor this skillcafe3310/public-agent-skills | 255 | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Audio Transcriptionmitsuhiko/agent-stuff | 3.2k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Yao Audio TranscriptionYaoApp/yao | 8.1k | — | ~416 | Automated safety check: Pass | Custom licence | |
| Baoyu Youtube TranscriptJimLiu/baoyu-skills | 26k | 1 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Transcription0xsline/OpenChatCut | 2.2k | 1 repos | ~1.1k | Automated safety check: Pass | AGPL-3.0 | |
| Youtube Transcriptbrowser-act/skills | 6.1k | — | ~2.1k | Automated safety check: Pass | MIT |
mitsuhiko/agent-stuff
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.
YaoApp/yao
Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers.
JimLiu/baoyu-skills
Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.
0xsline/OpenChatCut
A skill your agent uses when a video/audio task needs OpenChatCut transcription, captions, subtitles, subtitle styling, transcript search, transcript readiness checks, or enabling captions…
browser-act/skills
YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…
steipete/agent-scripts
yt-dlp downloads: video, audio, subtitles, transcripts, clips, playlists.
cafe3310/public-agent-skills
A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…
cafe3310/public-agent-skills
A specialized skill for embedding and extracting resilient watermarks in text by manipulating sentence lengths and using Fountain Codes.
cafe3310/public-agent-skills
从 Obsidian 知识库中扫描指定时间范围内未完成事件,生成/更新未完成事件整理文档. An agent skill from cafe3310/public-agent-skills.
cafe3310/public-agent-skills
一个全面、自主的深度研究框架。当用户请求对复杂主题、市场调研、技术格局进行深入的多维度调查,或需要大量网页浏览、数据合成和结构化报告的任何任务时,使用此技能。它协调子代理(subagents)并使用基于文件系统的状态管理来防止上下文膨胀。
cafe3310/public-agent-skills
将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档. An agent skill from cafe3310/public-agent-skills.
cafe3310/public-agent-skills
通过 markdown.new API 将网页、整站或搜索结果转换为干净的 Markdown. An agent skill from cafe3310/public-agent-skills.
Run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a claude-code`. Or copy the skill folder (skills/long-audio-transcript-processor in cafe3310/public-agent-skills) into .claude/skills/long-audio-transcript-processor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a codex`. Or copy the skill folder (skills/long-audio-transcript-processor in cafe3310/public-agent-skills) into .agents/skills/long-audio-transcript-processor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/long-audio-transcript-processor, .gemini/skills/long-audio-transcript-processor, .github/skills/long-audio-transcript-processor and .opencode/skills/long-audio-transcript-processor in your project.
Going by SKILL.md and its folder, Long Audio Transcript Processor needs Python for the scripts in its folder and the command-line tools its instructions call (python3).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Long Audio Transcript Processor is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Long Audio Transcript Processor: Audio Transcription (mitsuhiko/agent-stuff, 3.2k stars), Yao Audio Transcription (YaoApp/yao, 8.1k stars), Baoyu Youtube Transcript (JimLiu/baoyu-skills, 26k stars) and Transcription (0xsline/OpenChatCut, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cafe3310 (a GitHub user) maintains it in cafe3310/public-agent-skills, which has 255 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on June 26, 2026.
Source: cafe3310/public-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.