Voxtype Install
pchalasani/claude-code-tools
Guide the user through installing, configuring, and launching voxtype — local on-device voice dictation (speech-to-text that types wherever the cursor is).
转录稿纠错与轻度优化。本技能应在用户需要按用户词典纠正 ASR 转录稿同音字与英文专有名称漂移时使用。不要用于:重写为课程章节、报告、总结,或完全空白的素材创作。
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cat-xierluo/legal-skills transcription-corrector --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/transcription-corrector .claude/skills/transcription-corrector && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .claude/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-correctorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cat-xierluo/legal-skills transcription-corrector --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/transcription-corrector .agents/skills/transcription-corrector && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .agents/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cat-xierluo/legal-skills transcription-corrector --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/transcription-corrector .cursor/skills/transcription-corrector && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .cursor/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cat-xierluo/legal-skills.git --path skills/transcription-corrector--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cat-xierluo/legal-skills transcription-corrector --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/transcription-corrector .gemini/skills/transcription-corrector && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .gemini/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cat-xierluo/legal-skills transcription-correctorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/transcription-corrector .github/skills/transcription-corrector && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .github/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cat-xierluo/legal-skills transcription-corrector --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cat-xierluo/legal-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/transcription-corrector .opencode/skills/transcription-corrector && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "transcription-corrector" agent skill from https://github.com/cat-xierluo/legal-skills/tree/main/skills/transcription-corrector into .opencode/skills/transcription-corrector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription-corrector", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
transcription-corrector转录稿纠错与轻度优化。本技能应在用户需要按用户词典纠正 ASR 转录稿同音字与英文专有名称漂移时使用。不要用于:重写为课程章节、报告、总结,或完全空白的素材创作。
Transcription Corrector is an agent skill from cat-xierluo/legal-skills. 转录稿纠错与轻度优化。本技能应在用户需要按用户词典纠正 ASR 转录稿同音字与英文专有名称漂移时使用。不要用于:重写为课程章节、报告、总结,或完全空白的素材创作。
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `CHANGELOG.md`, `config/user_dictionary.example.yaml` and `references/config-decoupling.md`).
It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit db2c58c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, env and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Transcription Corrector loads about 4.8k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 26 tokens; SKILL.md has 1,390 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cat-xierluo/legal-skills at commit db2c58c, republished under its MIT licence (© cat-xierluo). 1,390 words, ~4,849 tokens.
.claude/skills/transcription-corrector/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.首次使用流程见 references/first_use.md。工作流中的判断准则见 references/correction_patterns.md。
把 raw 转录稿里"听得对但写不对"的同音/形近/拼写错误按用户词典统一改对,并按需做轻度润色。不做:改写、总结、删减、改事实、生成课程章节。
| 模式 | 触发 | 输出 |
|---|---|---|
| 纠错模式(默认) | 拿到 raw 转录稿,要先做"校对" | archive/.../{原文件}_corrected.md + 校对对照日志 |
| 纠错 + 轻度润色 | 还要顺手把发言人合并/标点整理 | 在纠错基础上再润色,输出到 archive/.../{原文件}_polished.md |
与下游处理工作的边界:
version + terms 列表)已成为本 skill 公开的接口约定;其他下游处理工作可以按需读取同一份词典文件。工作流按 5 个 Phase 组织,每个 Phase 含 1-3 个 Step。默认开启 Phase 1-3,Phase 4 默认关闭(需要用户明确开启)。
Phase 1 准备
Step 1.1 读取配置(词典)
Step 1.2 TXT → MD 转换(如需要)
Phase 2 识别
Step 2.1 识别文件结构
Step 2.2 通读识别高置信误转写
Phase 3 修正(默认开启)
Step 3.1 词典统一替换(改字)
Step 3.2 基础空白清理(改空白)
Step 3.3 口语词精简(删字)
Phase 4 润色(默认关闭)
Step 4.1 发言人合并
Step 4.2 标点规范
Step 4.3 段落切分
Phase 5 输出
Step 5.1 主归档(必出)
Step 5.2 源文件目录镜像(必出)为什么这样分组:
转录稿修正的本质是"对每一处误转写做出独立判断"。这个判断的输入是上下文,输出是一处 Edit,三者必须由同一个 agent 在同一个上下文里完成。
为什么按"每处 Edit"展开
terms: ["WorkBuddy"] 告诉 agent "正确的写法是 WorkBuddy",但不告诉 agent "这处原文是否需要改"。由此推出自然的执行流:
Read 把源文件读进上下文(必须通读全文)Edit 把每一处的决策落盘Write 写出 correction_log.md / meta.json / 源目录镜像为什么这个流是自洽的
skill 各项润色开关由 config/config.env 控制。加载顺序:
key=value 覆盖config/config.env(本 skill 自带)config/config.env.example(默认值的"母本",缺失 step 2 时用)配置示例(config/config.env.example):
# Phase 4 润色开关(默认全开,按需关闭)
POLISH_MERGE_SPEAKERS=true # Step 4.1
POLISH_PUNCTUATION=true # Step 4.2
POLISH_PARAGRAPH_SPLIT=true # Step 4.3
POLISH_TOPIC_SECTION=true # Step 4.4
POLISH_SUMMARY=true # Step 4.5
# 输出策略
SOURCE_MIRROR=true # 源文件目录镜像
SOURCE_MIRROR_CONFLICT=timestamp # 冲突加时间戳
# 日志详细度
LOG_VERBOSITY=verbose # verbose | compact缺失的键统一按本节表格的"默认值"列取值(不在文件中硬编码,每个 Step 自己的 § 段注明默认值)。
词典文件查找顺序(找到第一个存在的即用):
dictionary 路径config/user_dictionary.yaml(本 skill 自带)config/user_dictionary.example.yaml 为 config/user_dictionary.yaml词典 YAML 格式:
version: 1
terms:
- "WorkBuddy"
- "Claude Code"
- "Codex"当输入是 .txt 文件时,先做格式转换再进入 Phase 2。
触发条件:输入文件扩展名是 .txt。
转换规则(保守——只做格式包装,不改内容):
发言人 1 00:33:09):^(发言人|说话人|Speaker|S)\s*[\d一二三四五六七八九十]+[\.、]?\s*(\d{1,2}:\d{2}(?::\d{2})?)$**发言人 1 00:33:09**---、===)等全部保留###/####- xxx 改成 markdown 列表——保持原样输出:
{原文件 basename}_converted.md,保留在 archive 子目录.txt 文件保持不动示例:
输入(raw.txt):
发言人 1 00:33:09
我先说一下比较好的部分,就是整个最终的报告篇幅是比较充分的。
然后大家在一开始不知道怎么样一个报告是比较完善的,也知道怎么样去借鉴他山之石。
从这个阿尔法 gpt里面的这报告相当于蒸馏他的一个一个想法啊。
这个这个在前期一开始搭基建的时候确实是很有必要的一个措施。
发言人 1 00:33:57
呃然后。我注意到咱们在最终生成法律服务意见书的时候...输出(raw_converted.md):
**发言人 1 00:33:09**
我先说一下比较好的部分,就是整个最终的报告篇幅是比较充分的。
然后大家在一开始不知道怎么样一个报告是比较完善的,也知道怎么样去借鉴他山之石。
从这个阿尔法 gpt里面的这报告相当于蒸馏他的一个一个想法啊。
这个这个在前期一开始搭基建的时候确实是很有必要的一个措施。
**发言人 1 00:33:57**
呃然后。我注意到咱们在最终生成法律服务意见书的时候...不识别的格式(保持原文):
通读 raw 稿,识别组织方式:
**发言人 1 00:33:09**)— 多数 ASR 输出识别目的:决定 Phase 4 润色时是否合并同发言人连续发言。不要改写分块结构,保留原貌。
必须完整通读全文,以 Agent 视角扫描。不列静态错误表——ASR 错误无穷无尽(大小写漂移、空格、漏词、错位、同音字、形近字、断词重连、音节吞并…),任何预设列表都会漏判。
核心原则:
仅在上下文明确指向某个词典术语时,才将近似误转写替换为该术语。 模糊或低置信场景保持原文,不做猜测式替换。
启动扫描的高频易错类别(不是错误表,是判断启动点):
形式漂移识别(参考但不穷举):
禁止:
Phase 3 三个 Step 改的是三类不同对象——Step 3.1 改字(语义)、Step 3.2 改空白(格式)、Step 3.3 删字(语言字符)。三类操作各自独立、可单独审计、单独计日志。
对 Step 2.2 识别出的高置信误转写:
{原文} → {改后},并标注行号和上下文片段为什么独立于 Phase 4 润色:基础空白清理是"格式标准化",不改变原文语义、不删减任何信息;与 Phase 4 发言人合并 / 标点规范 / 段落切分那种"风格调整"性质不同,所以默认开启(用户不需要为此做开关决策)。
清理动作清单(仅做以下五类,不做任何超出此范围的事):
| 类别 | 动作 | 触发条件 |
|---|---|---|
| 行首/行尾空白 | 删除 | 任意行 |
| 连续空行 | 压缩为单个空行 | 任意位置 |
| 全角/半角空格 | 统一为半角空格 | 出现在英文单词之间或英文标点旁 |
| 专有名称内部断裂空格 | 合并 | 上下文明确指向词典中的专有名称(如"Work Buddy" → "WorkBuddy") |
| 中文数字 + 空格 + 级/章/节 | 合并 | 模式:[一二三四五六七八九十百千]+ +级/章/节 合并为 [一二...]+级/章/节(如"二 级" → "二级") |
标点周围空白规则:
,。、;:?!""''「」() 前后不留空格(若出现则删除), . ; : ? ! " ' ( ) 后空一格(中文段落中保留)安全执行(避免越段吞并):本步的"中文标点前后多余空格"清理有一个结构性陷阱——speaker 块的边界(**说话人 N HH:MM:SS** + 内容 + 空行 + 下一块标题)依赖"内容末标点 → 换行 → 空行"这一序列。如果清理正则用 \s+ 把标点之后的换行+空行也吃成单个空格,会跨块吞并,speaker 标题、H2 标题、speaker 块都会被合并成单行。执行时按行处理、禁止跨行匹配 \s+,仅做"行内"动作(行尾空格、同句内的 X, Y → X,Y)。如必须做跨段清理,先单独诊断再决定。
严格不做(避免越界):
**发言人 1 00:33:09**)与 Step 3.1 的关系:基础空白清理在 Step 3.1 词典替换之后执行;遇到"Work Buddy" → "WorkBuddy" 这类专有名称合并时,优先走 Step 3.1(词典替换路径),Step 3.2 仅做 Step 3.1 不覆盖的"无词典目标的纯格式"清理。
校对日志记录:所有 Step 3.2 的清理动作也写到 correction_log.md,在"基础空白清理"小节集中记录(与 Step 3.1 高置信替换区分开),便于审计。
为什么独立于 Phase 4 润色:与 Step 3.2 基础空白清理性质相同——属于"读起来对"的范畴,不改变原意(口语词不携带信息);与 Phase 4 发言人合并 / 标点规范 / 段落切分那种"风格调整"性质不同,所以默认开启(用户不需要为此做开关决策)。
白名单(可删的纯填充词)——AI 通读时识别"作为填充词 vs 作为实义",仅删独立出现且无语义负担的:
| 词 | 典型位置 | 删除条件 |
|---|---|---|
| 呃 | 段中独立 / 句首 | 认知停顿,不构成回答 |
| 啊 | 段中独立 / 句末 | 感叹停顿 |
| 哦 | 段中独立 | 过渡停顿(不作"明白了"回应) |
| 哎 | 段中独立 | 过渡停顿 |
| 那个 | 独立使用 | 纯指代填充(不指代具体名词) |
| 就是说 | 独立使用 | 纯连接填充 |
| 然后呢 | 段末 | 纯过渡填充 |
| 对吧 | 段末 | 纯反问填充 |
| 你知道吗 | 段中 | 纯引导填充 |
| 怎么说呢 | 段中 | 纯思考填充 |
| 这样子 | 独立使用 | 纯总结填充 |
| 其实呢 | 段首 | 纯转折填充 |
| 但是呢 | 段首 | 纯转折填充 |
保留规则(不删):
段首 vs 段尾 vs 段中 决策树:
| 位置 | 判定标准 | 处理(v2 激进版起) |
|---|---|---|
| 段首 | 紧跟 **发言人 X HH:MM:SS** 之后第一个非空行的第一个词 | 删除(v2 起激进) |
| 段中独立 | 句中位置、不是紧邻关键名词/动词 | 删除 |
| 段中并列停顿 啊 | 两个名词/动名词之间 + 顿号/逗号 | 保留(删除会破坏并列结构) |
| 段尾 | 一段内容末尾的最后一个词 | 保留(标记节奏收束) |
详见 references/correction_patterns.md §11 段中并列停顿 啊 的识别模式。
判断原则:
严格不做(避免越界):
校对日志记录:所有 Step 3.3 的精简动作写到 correction_log.md 的"口语词精简"小节,格式:
| 行号 | 原文片段 | 被删词 | 理由 |
|------|----------|--------|------|
| L207 | "呃 这个怎么说呢,我们今天来" | "呃" | 段首认知停顿,后接"怎么说呢"+完整内容,整段保留——本行不删(认知停顿后接完整内容的"呃"按 §3.3 保留规则例外)|
| L312 | "对,所以呢,我们" | "所以呢" | "所以呢" 不在白名单,保留 |
| L451 | "然后呢我们看一下" | "然后呢" | 段首过渡填充,可删 |与 Step 3.2 的关系:Step 3.2 动空白字符,Step 3.3 动语言字符。两者都是"读起来对"操作,都默认开启,都在 correction_log.md 单独小节集中记录。
必检(防止漏跑整节):跑完必须 grep 自检(详见 references/correction_patterns.md §12):
correction_log.md 必须包含 ## 口语词精简 小节correction_log.md 必须包含 ## 高置信替换 / ## 基础空白清理 / ## 口语词精简 / ## 未处理 四个小节meta.json 必含 summary / process_time / dictionary_source / polish_enabled 字段Phase 4 下的所有 Step 由 config/config.env 配置文件中的开关决定是否启用。每个 Step 默认值的来源:
config.env:用 config.env.example 的值(即默认值)config.env.example:用本节速查表的默认值| Step | 类别 | 动作 | 边界 | 配置开关 | 默认值 |
|---|---|---|---|---|---|
| Step 4.1 | 发言人合并 | 同一发言人连续多段、可合并为一段时,合并 | 仅去重分块,不删任何原话 | POLISH_MERGE_SPEAKERS | true |
| Step 4.2 | 标点规范 | 中英文标点统一、句末补全 | 不改语气词、不删"呃/啊"等自然语流标记 | POLISH_PUNCTUATION | true |
| Step 4.3 | 段落切分 | 极长段(>500字)按语义边界切 | 不重组、不重排、不总结 | POLISH_PARAGRAPH_SPLIT | true |
| Step 4.4 | 主题分章 + H2 标题 | 在明确的话题切换点插入 ## 一、标题 行(带中文大写序号) | 仅在切换点插入,不重组内容、不改写、不总结 | POLISH_TOPIC_SECTION | true |
| Step 4.5 | 会议纪要 / 总结 | 每章"本章摘要"+"关键洞察"就地嵌入对应 H2 标题下;文首主标题 + 文末关键词各一行 | 不引入新事实、不重切 H2、不评价、不另设顶部聚合块 | POLISH_SUMMARY | true |
触发条件:话题发生明确切换(不同讨论对象 / 不同议程 / 不同主体焦点)时,插入 ## 一、标题 行;同一话题内部的子切换不插入。
识别原则:
锚点定位流程(避免行号漂移):H2 锚点必须从 converted.md 上的结构标记(**说话人 N HH:MM:SS** 或杨卫薪行)的实际行号取,不得从上游 raw 文件的早期阅读印象猜行号——raw 与 converted 之间存在 metadata 块位移(Phase 1.2 包装前几行),且肉眼从 raw 数行号与脚本扫出来的真实行号常常差几行,把 H2 插错章节。先写 5 行扫描脚本打印 converted.md 的全部 speaker 行真实行号,再据此填 H2 锚点表;插完按"## 标题行 + 后 5 行"抽检,确认落在正确章节。
标题文本从何而来:
编号规则:
一、、二、、三、...、 顿号分隔(中文正式排版惯例):## 一、工具选择与日常使用corrected.md 中序号唯一,不重复严格不做:
校对日志记录:所有 Step 4.4 的插入写到 correction_log.md 的"主题分章"小节(与 Step 4.1/4.2/4.3 区分),格式:
| 行号 | 标题文本 | 上下文锚点(前后 50 字)| 切换类型 |
|------|----------|------------------------|----------|
| L487 | ## 律所管理与本地部署 | "...他就不会愿意去做这个事事情。" | 硬切换 |与 Phase 4 其他 Step 的关系:Step 4.1-4.3 是"风格调整",Step 4.4 是"结构标注"——前者改分块形式,后者只插入锚点,不重排。
为什么本 skill 做这个而不下游做:原 §1.2 早期版本说"下游做章节重组时再处理",但下游 skill(如 course-generator)做的是"基于转录稿生成课程",与"忠实标注原话结构"是不同任务。H2 主题分章是"忠实标注",放在本 skill 更合身。
POLISH_SUMMARY 开关)触发条件:
POLISH_SUMMARY=true,或用户在调用时显式声明("生成纪要" / "总结" / "AI 洞察")核心原则(v1.4 修订:摘要就地嵌入,不聚合到顶部):
<!-- AI-SUMMARY --> markers,不产生"所有章节摘要汇聚在一起"的独立区块输出位置:
corrected.md:① 文首 H1 下加一行主标题;② 每个 H2 标题下就地嵌入本章摘要 + 关键洞察;③ 文末加一行关键词_summary.md,不写 markers就地嵌入格式(每个 H2 之下):
## 一、章节名(Step 4.4 标题,原样)
**本章摘要**:50-100 字概括本章节核心内容(引用原话中的具体事实/数字/专有名,不引入新事实,不评价)
**关键洞察**:
- 洞察 1(15-30 字,横向观察本章节)
- 洞察 2
**发言人 1 00:02:09**
(该章原话正文,全部保留,紧跟在摘要 / 洞察之后)
...文首主标题 / 文末关键词:
**主标题**:5-15 字概括全文核心议题**关键词**:词1, 词2, ...(5-8 个,逗号分隔)为什么改为就地嵌入(v1.4):
生成原则:
严格不做:
## AI 摘要 聚合块、不写 <!-- AI-SUMMARY --> markers与 Step 4.4 的边界:
| Step 4.4 | Step 4.5 | |
|---|---|---|
| 输出文件 | corrected.md(修改) | corrected.md(修改,就地嵌入) |
| 操作 | 在原话章节前插入 H2 标记 | 每个 H2 标题下、原话前嵌入摘要 / 洞察;文首主标题、文末关键词各一行 |
| 保留原话 | 全部保留 | 原话全部保留 + 摘要中引用片段 |
| 引入新信息 | 不允许 | 不允许 |
幂等与验证:
^## 二级标题之下紧跟一个 **本章摘要** 行;② 全文无 <!-- AI-SUMMARY 字样;③ 主标题 / 关键词各一行存在校对日志记录:Step 4.5 的产出写到 correction_log.md 的"会议纪要"小节,格式:
| 字段 | 值 |
|------|-----|
| 主标题(文首单行)| 主标题文本 |
| 章节数 | 8(与 Step 4.4 H2 一致)|
| 总洞察数 | X 条 |
| 摘要字数总计 | Y 字 |
| 关键词(文末单行)| 词1, 词2, ... |
| 嵌入方式 | 就地嵌入每个 H2 之下;无顶部聚合块;无 markers |Phase 4 整体边界(适用于 Step 4.1-4.4):
输出采用 双写 策略:主归档在 skill 内部,源文件目录同步一份易访问副本。原始文件保持不动。
输出位置:{skill 目录}/archive/{YYYYMMDD_HHMMSS}_{原文件 basename}/
文件:
{原文件}_corrected.md — 纠错后的完整文本(保留原始分块结构){原文件}_polished.md — 仅在用户开启润色时生成correction_log.md — 校对对照日志meta.json — 处理时间、词典来源、是否润色、原始文件路径输出位置:原始文件所在目录(与原文件同目录)。
文件:
{原文件 basename}_corrected.md — 与 5.1 内容完全一致,仅是易访问副本_corrected 后缀,扩展名改为 .md(原文件是 .md 时,输出仍为 .md;原文件是 .txt 时,输出为 .md)为什么双写:
_corrected.md(polished 副本与 log 不镜像——它们的主归档位置就在 skill 内部)冲突处理:
_corrected.md → 用时间戳后缀区分(_corrected_20260608_153022.md),避免覆盖历史校对结果meta.json 记录 source_mirror: "skipped"原始文件保持不动——双写策略只生成"新文件",绝不就地修改原文件。
当某次跑出的 _corrected.md 经用户确认无问题、视为该转录稿的"稳定版"时,在该 archive 子目录下加 STABLE.md 标记文件:
# Stable Version Marker
**v3 (2026-06-08 12:00) 标记为该转录稿的稳定版本。**
## 为什么 v3 是 stable
- v1.0.0 (2026-06-07 14:30) — 7 处,仅 Step 3
- v1.0.2 (2026-06-08 00:41) — 42 处,Step 3 + 3.5,漏跑 Step 3.6
- v2 (2026-06-08 11:00) — 170 处,补跑 Step 3.6(84 处)
- v3 (2026-06-08 12:00) — 226 处,激进版段首删除(+52 处)
## 锁定本版本
- 后续不再对 `260606_AI技能创新大赛_corrected.md`(v3)做改动
- 源文件目录三个版本并列保留供对比
- 如发现新的误转写需要补判,新建 `archive/<新时间戳>_*` 目录重跑,不就地修改 v3为什么需要 STABLE.md:
_corrected_v3_20260608_120000.md)按时间戳命名,本身是"稳定"暗示,但 archive 内部 260606_AI技能创新大赛_corrected.md 没 STABLE 标记会被误认成"最新版"什么时候不写 STABLE.md:
archive 目录会保留每次跑的全量产物(主归档 + correction_log + meta.json),不主动清理历史版本。但当 archive 下子目录数 ≥ 5 时,建议在 STABLE.md / DECISIONS.md 备注保留策略:
_corrected.md / correction_log.md / meta.json(记录完整迭代轨迹)analysis_report.md 等中间产物可手动归档_corrected.md(用户可能要从源文件目录镜像之外的路径找回终版)gating the "漏跑章节"型 bug 的 grep 必检已写在 Step 3.3 末尾;这里加的是"跑过但跑错"的物理结构回归——Step 3.2 越段吞并(speaker 块、H2 标题被跨段正则粘连)与 Step 4.4 行号漂移(H2 插错章节)这两类失败模式已分别写入 Step 3.2「安全执行」段与 Step 4.4「锚点定位流程」段,本步只做落地抽检,把"是否真发生了这两类回归"显式锁住。
抽检项(任一不过即视为本批未完成,不得打完成凭据):
_corrected.md 中 ^\*\*说话人 \d+ \d+:\d+(?::\d+)?\*\* 与 ^杨卫薪 \d+:\d+(?::\d+)? 的 speaker 标记行总数,应等于或近似于 converted.md 中的总数(差 ≤ 1 因末尾 speaker 块合并属正常)。如果总数大幅缩小(如 96 块 → 60 块),说明 Step 3.2 越段吞并了内容。_corrected.md 抽取每个 speaker 块的第 1 行非空文本,长度应 ≥ 5 个字符;如果变短(如「可以听到。」只有 4 字),说明上一块被吞并。^## (一|二|三|四|五|六|七|八|九|十)、 之下的紧跟 5 行内应出现 speaker 块首行(**说话人** / 杨卫薪);找不到即为章节定位错位。_corrected.md 倒数行 ≤ 5 行内应有 ^\*\*关键词\*\*:...,开头 metadata 块之后第一个 H2 之前应有 ^\*\*主标题\*\*:...。执行建议:把以上写成 5-10 行脚本(脚本文件无需持久化),跑完打印四行 PASS/FAIL;任一 FAIL 就对照 Step 3.2 / Step 4.4 的边界段排查,不要靠目测放过。
为什么本步独立于已有 grep 必检:Step 3.3 末尾的 grep 检查的是"log 章节齐不齐"——它查日志的存在,不查产物的物理形态。本步查的是产物本身有没有结构性破坏。两者覆盖正交失败模式,缺一不可。
patch / write_file 输入错必修(Type 1 失败模式):以上抽检是结构层"跑过但跑错"。还有一类输入字面错——patch 工具的 old_string / new_string 输入时手误(如把 排障 误输 排排),旧字串可能在文中多处命中或被部分匹配,diff 工具会显示"成功替换",但实际写入了错别字。所有 patch 调用的 diff 输出必须逐字过一遍再贴回——尤其是长文件、多字、含相似字词(如 排排/排障/排序/排除)的场景;write_file 整文件覆写时,对成品做"复制 5 行片段回读"核对写入内容与预期一致。两类失败都属"未核对的成功"。
# 校对对照日志
- **原始文件**:/path/to/raw.md
- **处理时间**:2026-06-07 14:30:22
- **词典来源**:`config/user_dictionary.yaml`(含 18 项术语)
- **润色开关**:关闭
## 高置信替换
| 原文 | 改后 | 出现次数 | 上下文示例 |
|------|------|----------|------------|
| workbody | WorkBuddy | 3 | "通过第二版呢进行测试..." |
| work body | WorkBuddy | 2 | "我们把这个实践中办案..." |
| claude-code | Claude Code | 1 | "...重新调回那个 claude-code 里面..." |
## 未处理(保留原文)
- "智能件" — 上下文不足以判断是否为"智能体"
- "代理" — 同上
## 润色变更(如开启)
- 合并发言人 1 连续 3 段为 1 段
- ...见 references/config-decoupling.md。
© cat-xierluo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (references) in skills/transcription-corrector of cat-xierluo/legal-skills.
Open the folder on GitHubat commit db2c58c
Transcription Corrector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Transcription Corrector this skillcat-xierluo/legal-skills | 720 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Voxtype Installpchalasani/claude-code-tools | 2k | — | ~857 | Automated safety check: Notes | MIT | |
| Speech To TextNoizAI/skills | 526 | — | ~917 | Automated safety check: Pass | None | |
| Podcast Transcript FetcherVarnan-Tech/opendirectory | 674 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Audio Transcriptionmitsuhiko/agent-stuff | 3.2k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Speech Recognitiondpearson2699/swift-ios-skills | 1.2k | — | ~3.7k | Automated safety check: Pass | Custom licence |
pchalasani/claude-code-tools
Guide the user through installing, configuring, and launching voxtype — local on-device voice dictation (speech-to-text that types wherever the cursor is).
NoizAI/skills
A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file.
Varnan-Tech/opendirectory
A skill your agent uses when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast.
mitsuhiko/agent-stuff
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.
dpearson2699/swift-ios-skills
Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills.
cnemri/google-genai-skills
Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.
cat-xierluo/legal-skills
Converts a lawyer's ordinary complaint or a described case into the Supreme People's Court's elements-style Word template, with layout checks on the result.
cat-xierluo/legal-skills
Analyzes raw lecture transcripts for verbal tics, pacing, time use and promise follow-through, with optional slide-by-slide comparison and cross-session tracking.
cat-xierluo/legal-skills
Detects and rewrites machine-sounding patterns in the body text of Chinese articles while keeping the author's facts, headings and legal terms intact.
cat-xierluo/legal-skills
Finds GitHub projects mentioned in articles or screenshots and stars them, tracks updates to your starred repos, and builds an HTML dashboard to browse them.
cat-xierluo/legal-skills
Sets up or incrementally updates AGENTS.md and CLAUDE.md for legal professionals, with a minimal safety baseline and a check that a new session loads and follows the rules.
cat-xierluo/legal-skills
Chinese-language skill that organizes a case file into a multi-role mock trial with judge, parties and clerk, producing a transcript, issue review and a to-strengthen list.
Categories
转录稿纠错与轻度优化。本技能应在用户需要按用户词典纠正 ASR 转录稿同音字与英文专有名称漂移时使用。不要用于:重写为课程章节、报告、总结,或完全空白的素材创作。. Transcription Corrector is an agent skill from cat-xierluo/legal-skills.
Transcription Corrector fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.
Run `npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a claude-code`. Or copy the skill folder (skills/transcription-corrector in cat-xierluo/legal-skills) into .claude/skills/transcription-corrector in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a codex`. Or copy the skill folder (skills/transcription-corrector in cat-xierluo/legal-skills) into .agents/skills/transcription-corrector in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cat-xierluo/legal-skills --skill transcription-corrector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcription-corrector, .gemini/skills/transcription-corrector, .github/skills/transcription-corrector and .opencode/skills/transcription-corrector in your project.
SKILL.md names no scripts, command-line tools or credentials: Transcription Corrector is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Transcription Corrector is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Transcription Corrector: Voxtype Install (pchalasani/claude-code-tools, 2k stars), Speech To Text (NoizAI/skills, 526 stars), Podcast Transcript Fetcher (Varnan-Tech/opendirectory, 674 stars) and Audio Transcription (mitsuhiko/agent-stuff, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cat-xierluo (a GitHub user) maintains it in cat-xierluo/legal-skills, which has 720 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 10, 2026.
Source: cat-xierluo/legal-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.