Agent skill

Long Audio Transcript Processor

by cafe3310 in cafe3310/public-agent-skills

“对大量语音转写稿进行校对、整理、分段处理,支持断点续传和恢复”

— description from SKILL.md by cafe3310
Apache-2.0Auto-check passed

Install Long Audio Transcript Processor

skills CLI
$ npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cafe3310/public-agent-skills long-audio-transcript-processor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/long-audio-transcript-processor .claude/skills/long-audio-transcript-processor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
long-audio-transcript-processor
GitHub stars
255
Token cost
~1k tokens
SKILL.md length
212 words
Files
7 (incl. scripts, assets)
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

  • Works in 5 steps: 准备阶段 (Initialization) → 分段处理循环 (Processing Loop) → 中断恢复 (Resuming) → …
  • SKILL.md covers 使用时机, 工作流, 关键文件说明 and 注意事项
  • Runs Python scripts from its folder; calls python3

About this skill

Long Audio Transcript Processor is a skill in cafe3310/public-agent-skills (255 stars). Its SKILL.md is about 1k tokens, with 6 other files in the folder (scripts, assets). Licence: Apache-2.0.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 准备阶段 (Initialization)
  2. 分段处理循环 (Processing Loop)
  3. 中断恢复 (Resuming)
  4. 衍生任务 (Derivative Tasks)
  5. 调整与整合 (Adjustment and Integration)

What it can do on your machine

Read from SKILL.md and the folder at commit 6c45501. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Long Audio Transcript Processor loads about 1k tokens when it runs. Until then it costs about 16 tokens; SKILL.md has 212 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cafe3310/public-agent-skills at commit 6c45501, republished under its Apache-2.0 licence (© cafe3310). 212 words, ~1,041 tokens.

Download SKILL.mdSave it as .claude/skills/long-audio-transcript-processor/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
long-audio-transcript-processor
description
对大量语音转写稿进行校对、整理、分段处理,支持断点续传和恢复
license
Apache-2.0
author
github/cafe3310
depends_on_binary
python3, sed

语音转写处理工作流 (Long Audio Transcript Processor)

此技能旨在通过文件系统辅助,安全、有序地处理超长语音转写文本。它通过分段处理、上下文维护(术语表、主题记录)和状态追踪,确保处理过程的可持续性和高质量。 此技能最好使用最高性能的模型(而不是写代码用的快速模型)以确保最佳质量。

使用时机

当用户提供一个或多个长篇语音转写文件,并要求进行:

  1. 校对与修正:修复识别错误、词语重复、中断、错漏等,校对前后行数一致。
  2. 整理与格式化:区分发言人,整理段落。
  3. 长文本处理:文本过长,无法一次性放入 LLM 上下文。
  4. 需要中断与恢复:工作可能跨越多次对话。

工作流

1. 准备阶段 (Initialization)

首先,必须初始化工作区。询问用户是否已准备好源文件。

运行初始化脚本:

bash
python3 .gemini/skills/long-audio-transcript-processor/scripts/setup_workspace.py "path/to/file1.txt" "path/to/file2.txt" ...

(注意:请根据实际技能安装路径调整脚本路径,通常是 .gemini/skills/...)

初始化后,工作区结构如下:

text
语音转写处理_YYYY-MM-DD-HH-MM/
├── 0-工作日志.md           # 进度追踪与计划
├── 1-原始文件/             # 存放用户提供的原始语音文本
├── 2-要求和信息/           # 存放活动背景、发言人等信息(用户补充)
├── 3-校对和术语表.md       # 动态更新的术语库和错误模式
├── 4-分段主题.md           # 记录已处理分段的主题脉络
└── 5-最终输出/             # 存放校对完成的分段文件

关键操作:

  1. 检查生成的目录。
  2. 同步背景知识:确保将用户提供的任何背景文档(如 MRD、PRD、会议背景、相关邮件等)文档复制到 2-要求和信息/ 目录下,也在该目录下创建 Markdown 文档记录用户的说明。这是保证后续处理准确性的基石。
2. 分段处理循环 (Processing Loop)

在进入循环前,总是先读取以下文件以加载上下文(确保跨分段的信息一致性):

  • 0-工作日志.md (检查进度)
  • 2-要求和信息/ 下的所有背景和要求文档
  • 3-校对和术语表.md (加载最新积累的术语和校对规则)
  • 4-分段主题.md (加载已有上下文主题)
  • 5-最终输出/ 下的文件 -- 列出文件名即可

步骤:

  1. 确定任务:从 0-工作日志.md 中找到第一个未完成([ ])的分段。
  2. 提取分段并初始化目标文件:
    • 使用 sed 命令从 1-原始文件/ 中提取对应行范围,并重定向写入到 5-最终输出/ 下的对应文件中。
    • 命令示例:sed -n '开始行,结束行p' "1-原始文件/文件名.txt" > "5-最终输出/文件名_开始行-结束行.txt"
    • 目的:确保所有原始行都已先进入目标文件,作为后续改写的底稿,防止内容遗漏。
  3. 执行校对与订正:
    • 读取:使用 read_file 读取上一步生成的 5-最终输出/ 下的文件内容。
    • 参考:术语表、背景信息、前序分段主题。
    • 改写校对:
      • 对读取到的内容进行逐行/逐段校对。
      • 修正错别字、重复、同音错误,减少冗余的语气词。
      • 标记不确定内容为 (...)。
      • 保持一致性:必须确保校对后的行数与原始行数完全一致。如果需要合并发言,请在原位保留行占位。
    • 写回:使用 write_file 将校对后的完整文本写回 5-最终输出/ 的对应文件(覆盖掉刚才的底稿)。
  4. 更新辅助文档:
    • 更新术语表。如果在本段发现了新的专有名词、特定错误模式或确认了某个模糊术语的正确写法,追加 到 3-校对和术语表.md。仅追加,用行号段落区分不同分段的内容。
    • 更新主题。将本段核心主题 追加 到 4-分段主题.md。仅追加,用行号段落区分不同分段的内容。
  5. 更新状态:
    • 更新日志:在 0-工作日志.md 中标记分段为 [x]。
  6. 反馈与沉淀:
    • 告知用户本段已完成,简述关键修改、新发现的术语或不确定点。
    • 询问是否有新增术语或错误修正。
    • 如有反馈,更新 3-校对和术语表.md 并修正 5-最终输出 中的对应文件。
3. 中断恢复 (Resuming)

如果对话中断,不要 重新初始化。 直接执行 分段处理循环 的“在进入循环前”步骤:通过读取 2-要求和信息/ 和 3-校对和术语表.md 完整找回记忆。 然后继续下一个未完成的分段。

4. 衍生任务 (Derivative Tasks)

除了标准转写外,用户可能要求并行生成其他产物(如 Q&A 问答库、摘要、待办事项)。

  • 定义:在 2-要求和信息/ 下创建任务说明文档(如 额外任务_问答积累.md)。
  • 执行:在每次“输出结果”后,检查衍生任务要求,从当前分段中提取相关信息。
  • 存储:衍生文件存放在 5-最终输出/ 下,文件命名应清晰(如 问题和回答-主题.md)。
  • 原则:衍生文件的更新同样遵循**“只追加”**,用行号段落区分不同分段的内容。
5. 调整与整合 (Adjustment and Integration)

在所有分段处理完成后,进入此阶段。

  1. 成果汇总与提醒:

    • 列出所有产出的核心文件(分段文本、术语表、摘要、衍生文档)。
    • 提醒用户:
      • a) 可以提出任何额外的整理、合并或格式调整要求(例如:合并为一个文档,提取特定角色的发言)。
      • b) 可以将内容发送给相关人员(如 PM、技术负责人)进行审阅和订正。
  2. 订正处理标准:

    • 当用户反馈订正意见时,或者用户转发他人的订正内容时,必须使用以下 Emoji 标记修改痕迹,以便清晰追踪变动:
      • {🔴 原始内容}:标记被修改或删除的原始文本。
      • {🟢 新内容}:标记新增或修正后的文本。
      • {‼️ 需特别留意,可能出错}:标记模型认为存在矛盾、风险或不确定的地方,或者用户特别强调的注意点。
    • 更新动作:根据订正内容,同步更新 3-校对和术语表.md 和 5-最终输出/ 中的对应文件,确保下一次处理或合并时使用最新数据的正确版本。

关键文件说明

  • 0-工作日志.md: 核心状态文件。必须 保持最新。
  • 3-校对和术语表.md: 动态更新的知识库。发现新术语或特定错误模式时,务必更新此文件,以保证后续分段处理的一致性。该文件的编辑也遵循只追加原则。
  • 4-分段主题.md: 帮助 LLM 保持对长文本整体脉络的理解。

注意事项

  • 文件操作:所有生成的文件应严格存放在初始化的工作目录中。
  • 行号对应:读取原始文件时,务必确保行号范围准确。
  • 只追加原则:对于 3-校对和术语表, 4-分段主题.md 和任何额外输出文件,严格遵循 只追加、不修改的原则,避免在 loop 过程中,覆盖已有内容。
  • 第一时间记录额外要求:如果用户在过程中提出了额外要求,尤其是需要调整处理方式或补充信息,应立即记录在 2-要求和信息/ 中,并通知后续处理环节。
  • 保持行数一致:生成校对后输出时,禁止更改时间戳行。如有多次发言合并,可以用 (空) 标记被合并的发言。这是为了方便对原始文稿和校对后结果进行 diff 检查。
  • 说话人识别:语音转写的说话人一般正确,除非用户指出需要做识别。识别时,根据上下文(提问 vs 回答、专家 vs 小白、行业背景差异)区分真实角色,在输出中显式标记角色(如 同行 (测试)、访客 (学生))。禁止更改时间戳行,识别的角色必须添加在发言行行首。
发言人 2  01:30:00
这是一个验证。
发言人 2  01:30:00
(测试负责人-老张):这是一个验证。
  • 背景驱动:禁止在不读取背景文档的情况下开始处理分段。

© cafe3310, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, assets) in skills/long-audio-transcript-processor of cafe3310/public-agent-skills.

  • SKILL.md
  • assets/derivative_task_template.md
  • assets/proofreading_glossary_template.md
  • assets/segment_topics_template.md
  • assets/speaker_info_template.md
  • assets/work_log_template.md
  • scripts/setup_workspace.py

Open the folder on GitHubat commit 6c45501

Compare with similar skills

Long Audio Transcript Processor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Long Audio Transcript Processor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Long Audio Transcript Processor this skillcafe3310/public-agent-skills255—~1kAutomated safety check: PassApache-2.0
Audio Transcriptionmitsuhiko/agent-stuff3.2k—~1kAutomated safety check: PassApache-2.0
Yao Audio TranscriptionYaoApp/yao8.1k—~416Automated safety check: PassCustom licence
Baoyu Youtube TranscriptJimLiu/baoyu-skills26k1 repos~2.4kAutomated safety check: PassMIT
Transcription0xsline/OpenChatCut2.2k1 repos~1.1kAutomated safety check: PassAGPL-3.0
Youtube Transcriptbrowser-act/skills6.1k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Audio Transcription

    mitsuhiko/agent-stuff

    Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

    3.2k GitHub stars~1k tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers.

    8.1k GitHub stars~416 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Baoyu Youtube Transcript

    JimLiu/baoyu-skills

    Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.

    26k GitHub starsUsed in 1 repo~2.4k tokens
    Media & CreativeAuto-check passed
  • Transcription

    0xsline/OpenChatCut

    A skill your agent uses when a video/audio task needs OpenChatCut transcription, captions, subtitles, subtitle styling, transcript search, transcript readiness checks, or enabling captions…

    2.2k GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Youtube Transcript

    browser-act/skills

    YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…

    6.1k GitHub stars~2.1k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • Video Transcript Downloader

    steipete/agent-scripts

    yt-dlp downloads: video, audio, subtitles, transcripts, clips, playlists.

    7.3k GitHub stars~558 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from cafe3310/public-agent-skills

All 29 skills in this repo
  • Impeccable

    cafe3310/public-agent-skills

    A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…

    255 GitHub stars~4.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Text Watermark Fountain

    cafe3310/public-agent-skills

    A specialized skill for embedding and extracting resilient watermarks in text by manipulating sentence lengths and using Fountain Codes.

    255 GitHub stars~921 tokensUpdated 3 mo ago
    Auto-check passed
  • Obsidian Todo Collector

    cafe3310/public-agent-skills

    从 Obsidian 知识库中扫描指定时间范围内未完成事件,生成/更新未完成事件整理文档. An agent skill from cafe3310/public-agent-skills.

    255 GitHub stars~575 tokensUpdated 3 mo ago
    Auto-check passed
  • Deep Research

    cafe3310/public-agent-skills

    一个全面、自主的深度研究框架。当用户请求对复杂主题、市场调研、技术格局进行深入的多维度调查,或需要大量网页浏览、数据合成和结构化报告的任何任务时,使用此技能。它协调子代理(subagents)并使用基于文件系统的状态管理来防止上下文膨胀。

    255 GitHub stars~996 tokensUpdated 3 mo ago
    Auto-check passed
  • Long Audio To Obsidian

    cafe3310/public-agent-skills

    将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档. An agent skill from cafe3310/public-agent-skills.

    255 GitHub stars~772 tokensUpdated 3 mo ago
    Auto-check passed
  • Markdown New

    cafe3310/public-agent-skills

    通过 markdown.new API 将网页、整站或搜索结果转换为干净的 Markdown. An agent skill from cafe3310/public-agent-skills.

    255 GitHub stars~344 tokensUpdated 3 mo ago
    Auto-check passed

Questions about Long Audio Transcript Processor

How do I install Long Audio Transcript Processor in Claude Code?

Run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a claude-code`. Or copy the skill folder (skills/long-audio-transcript-processor in cafe3310/public-agent-skills) into .claude/skills/long-audio-transcript-processor in your project. Claude Code loads it when a task matches its description.

How do I install Long Audio Transcript Processor in Codex?

Run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a codex`. Or copy the skill folder (skills/long-audio-transcript-processor in cafe3310/public-agent-skills) into .agents/skills/long-audio-transcript-processor in your project. Codex loads it when a task matches its description.

Can I use Long Audio Transcript Processor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cafe3310/public-agent-skills --skill long-audio-transcript-processor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/long-audio-transcript-processor, .gemini/skills/long-audio-transcript-processor, .github/skills/long-audio-transcript-processor and .opencode/skills/long-audio-transcript-processor in your project.

What does Long Audio Transcript Processor need to run?

Going by SKILL.md and its folder, Long Audio Transcript Processor needs Python for the scripts in its folder and the command-line tools its instructions call (python3).

Does Long Audio Transcript Processor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Long Audio Transcript Processor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Long Audio Transcript Processor use?

Long Audio Transcript Processor is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Long Audio Transcript Processor use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Long Audio Transcript Processor?

Skills that share tags, products or a category with Long Audio Transcript Processor: Audio Transcription (mitsuhiko/agent-stuff, 3.2k stars), Yao Audio Transcription (YaoApp/yao, 8.1k stars), Baoyu Youtube Transcript (JimLiu/baoyu-skills, 26k stars) and Transcription (0xsline/OpenChatCut, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Long Audio Transcript Processor?

cafe3310 (a GitHub user) maintains it in cafe3310/public-agent-skills, which has 255 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on June 26, 2026.

Source: cafe3310/public-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.