Agent skill

Voice Optimization

by kangarooking in kangarooking/system-prompt-skills

当系统提示面向语音交互场景(语音助手、语音搜索、有声回答、电话客服 AI)时调用。适用于需要将文本输出优化为口语表达的系统提示设计。不适用于纯文本聊天界面,不适用于语音合成(TTS)技术选型,不适用于图像/视频多模态场景。

MITAuto-check passedAI & LLM Engineering

Install Voice Optimization

skills CLI
$ npx skills add kangarooking/system-prompt-skills --skill voice-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kangarooking/system-prompt-skills voice-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kangarooking/system-prompt-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/voice-optimization .claude/skills/voice-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-optimization
GitHub stars
205
Token cost
~642 tokens
SKILL.md length
144 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

当系统提示面向语音交互场景(语音助手、语音搜索、有声回答、电话客服 AI)时调用。适用于需要将文本输出优化为口语表达的系统提示设计。不适用于纯文本聊天界面,不适用于语音合成(TTS)技术选型,不适用于图像/视频多模态场景。

  • Works in 6 steps: 简洁优先原则:语音场景中用户注意力窗口极短,回答必须开门见山,禁止"好的,让我来回… → 口语化转换:将书面语转换为自然口语——使用短句、主动语态、日常词汇,避免从句嵌套和… → 格式降级:去除 Markdown… → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers R — 原文 (Reading), I — 方法论骨架 (Interpretation), A1 — 案例分析 (Past Application) and A2 — 触发场景 (Future Trigger) ★, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Voice Optimization is an agent skill from kangarooking/system-prompt-skills. 当系统提示面向语音交互场景(语音助手、语音搜索、有声回答、电话客服 AI)时调用。适用于需要将文本输出优化为口语表达的系统提示设计。不适用于纯文本聊天界面,不适用于语音合成(TTS)技术选型,不适用于图像/视频多模态场景。

Its SKILL.md is about 640 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Text to speech and voice. It works with Perplexity. The repository describes itself as: 从 165 个顶级 AI 产品系统提示词中蒸馏出的 15 个可执行 Agent skill. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/voice-optimization”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. 简洁优先原则:语音场景中用户注意力窗口极短,回答必须开门见山,禁止"好的,让我来回答您的问题"等前言式表达。
  2. 口语化转换:将书面语转换为自然口语——使用短句、主动语态、日常词汇,避免从句嵌套和术语堆砌。
  3. 格式降级:去除 Markdown 表格、代码块、嵌套列表等视觉格式,改用自然语言描述或简短列举。
  4. 音频安全约束:禁止识别特定说话人身份、禁止模仿真实人物声线特征、禁止唱歌或哼唱旋律。
  5. 语言限制声明:明确支持的语种范围,超出范围时引导用户调整设置而非强行处理。
  6. 长度自适应:根据问题复杂度动态调整回答长度——简单问题一两句话,复杂问题控制在合理时长内。

What it can do on your machine

Read from SKILL.md and the folder at commit 252cd52. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Optimization loads about 642 tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 144 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~642

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kangarooking/system-prompt-skills at commit 252cd52, republished under its MIT licence (© kangarooking). 144 words, ~642 tokens.

Download SKILL.mdSave it as .claude/skills/voice-optimization/SKILL.md (or your agent's skills folder).
name
voice-optimization
description
当系统提示面向语音交互场景(语音助手、语音搜索、有声回答、电话客服 AI)时调用。适用于需要将文本输出优化为口语表达的系统提示设计。不适用于纯文本聊天界面,不适用于语音合成(TTS)技术选型,不适用于图像/视频多模态场景。
tags
语音, 口语优化, 简洁输出, 语音安全, 对话设计
related_skills
mobile-adaptation, citation-system

语音场景优化

R — 原文 (Reading)

Perplexity Voice 要求"请快速说话"、仅支持英语、禁止说话人识别、禁止唱歌哼唱、禁止模仿;Claude Mobile 强调"始终先给答案、无前言"、列表在小屏幕上更易扫描;Sesame AI Maya 专为语音优化的对话模式。核心模式:简洁优先、口语化适配、音频安全约束、语言限制、去除视觉格式。

I — 方法论骨架 (Interpretation)

  1. 简洁优先原则:语音场景中用户注意力窗口极短,回答必须开门见山,禁止"好的,让我来回答您的问题"等前言式表达。
  2. 口语化转换:将书面语转换为自然口语——使用短句、主动语态、日常词汇,避免从句嵌套和术语堆砌。
  3. 格式降级:去除 Markdown 表格、代码块、嵌套列表等视觉格式,改用自然语言描述或简短列举。
  4. 音频安全约束:禁止识别特定说话人身份、禁止模仿真实人物声线特征、禁止唱歌或哼唱旋律。
  5. 语言限制声明:明确支持的语种范围,超出范围时引导用户调整设置而非强行处理。
  6. 长度自适应:根据问题复杂度动态调整回答长度——简单问题一两句话,复杂问题控制在合理时长内。

A1 — 案例分析 (Past Application)

案例: Perplexity Voice 的音频安全边界
  • 问题: 语音交互中用户可能要求模仿名人声音、识别通话对象、或要求 AI 唱歌,这些行为涉及隐私、版权和安全风险。
  • 设计模式的使用: Perplexity Voice 在系统提示中设置明确禁区——不对语音输入进行说话人识别("No speaker identification from voice"),不执行唱歌或哼唱请求,不进行人物模仿。同时限定仅支持英语,超出能力范围时引导用户修改设置。
  • 结论: 语音场景有独特的安全边界(声纹、模仿、演唱),这些在纯文本场景中不存在,需要专项防护。
案例: Claude Mobile 的回答优先策略
  • 问题: 移动端语音回答中,用户听到冗长的开场白会快速失去耐心,尤其在驾驶、行走等场景下。
  • 设计模式的使用: Claude Mobile 明确指令"Always lead with answer. No preamble.",将答案前置,解释后置。对于不同复杂度的问题设定长度层级——简单问题 1-2 句,操作指南用短列表,实质性问题 2-3 段。
  • 结论: 语音场景对延迟感知极度敏感,去除前言可显著提升用户满意度和信息获取效率。

A2 — 触发场景 (Future Trigger) ★

用户在什么情境下需要?
  1. 设计语音助手(如智能音箱、车载助手)的系统提示
  2. 为现有文本聊天机器人添加语音交互模式
  3. 构建电话客服 AI 的对话脚本
  4. 优化播客生成或有声内容合成中的口语表达
语言信号
  • "语音场景下的输出优化"
  • "需要口语化回答"
  • "用户通过语音提问"
  • "回答会被朗读出来"
  • "如何让 AI 说话更自然"
与相邻 skill 的区分
  • 与 mobile-adaptation 区别:移动适配关注屏幕尺寸约束,语音优化关注听觉通道约束;但两者都强调简洁优先,常联合使用
  • 与 citation-system 区别:引用系统在语音场景中需要特殊处理(无法使用视觉标记),但语音优化不涉及引用格式设计本身

E — 可执行步骤 (Execution)

  1. 步骤 1:设定回答长度层级 - 完成标准:为问题复杂度定义 3-4 个层级(简单/操作/中等/复杂),每个层级规定最大句数或预估朗读时长,并在系统提示中以示例说明。
  2. 步骤 2:编写前言禁令与答案前置规则 - 完成标准:明确声明"禁止在回答开头添加确认性前言",提供正确和错误的示例对比(如错误:"好的,让我为您解答...",正确:直接给出答案)。
  3. 步骤 3:定义格式降级规则 - 完成标准:列出需降级的视觉格式(表格→自然语言描述、嵌套列表→扁平列举、代码块→口语化步骤说明),并给出每种降级的示例。
  4. 步骤 4:设定音频安全边界 - 完成标准:明确禁止的行为清单(说话人识别、声线模仿、唱歌哼唱、人物扮演),定义超出能力范围时的标准回退话术。
  5. 步骤 5:添加口语化转换指南 - 完成标准:列出书面语到口语的转换规则(从句→短句、被动→主动、术语→日常词汇),提供 3 个以上转换示例。

B — 边界 (Boundary) ★

不要在以下情况使用
  • 纯文本聊天界面,用户通过键盘输入和屏幕阅读
  • 语音合成(TTS)引擎的技术选型或参数调优
  • 音频信号处理(降噪、回声消除等)
  • 多模态场景中语音仅为辅助通道(如视频会议中的字幕场景)
常见失败模式
  • 照搬文本输出:直接将文本聊天回复用于语音场景,导致冗长前言、视觉格式(表格、代码块)被朗读出来,用户体验极差
  • 忽视音频独有风险:仅优化表达方式但未设置说话人识别、模仿等音频特有的安全边界
  • 过度简化:将所有回答压缩为一句话,丢失必要信息和上下文,应按复杂度分级而非一刀切
  • 忽略语言限制:未声明支持的语种范围,导致多语言场景下输出混乱或质量下降

© kangarooking, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in voice-optimization of kangarooking/system-prompt-skills.

Open the folder on GitHubat commit 252cd52

Compare with similar skills

Voice Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Optimization this skillkangarooking/system-prompt-skills205—~642Automated safety check: PassMIT
Xsaimoeru-ai/airi50k1 repos~1.3kAutomated safety check: PassMIT
Hriterrense/ros2-multimodal-robot-collab111—~268Automated safety check: PassMIT
Video Understandingzenstory-ai/video-recap-skills555—~1.1kAutomated safety check: PassMIT
Unified LLM APIPrism-Shadow/penguin-harness2.5k—~6.7kAutomated safety check: PassApache-2.0
Dailymajiayu000/claude-skill-registry6664 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Xsai

    moeru-ai/airi

    A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…

    50k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Hri

    terrense/ros2-multimodal-robot-collab

    A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action.

    111 GitHub stars~268 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Video Understanding

    zenstory-ai/video-recap-skills

    把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

    555 GitHub stars~1.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Unified LLM API

    Prism-Shadow/penguin-harness

    Call model APIs through @prismshadow/mmsp (MMSP) — streaming text generation, image generation, speech synthesis, embeddings and the supported-model registry with one client.

    2.5k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Daily

    majiayu000/claude-skill-registry

    Documentation and capabilities reference for Daily. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 4 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Render Cinematic Music Video

    majiayu000/claude-skill-registry

    Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and…

    666 GitHub starsUsed in 1 repo~1k tokens
    AI & LLM EngineeringAuto-check passed

More from kangarooking/system-prompt-skills

All 15 skills in this repo
  • Persona Design

    kangarooking/system-prompt-skills

    当需要为 AI 产品定义核心身份、角色声明和能力边界时调用此 skill。典型场景包括:设计新 AI 产品的 system prompt 首段、为不同场景创建差异化角色(如教学助手 vs 编程代理)、重新定义 AI 与用户的关系框架。

    205 GitHub starsUsed in 1 repo~956 tokens
    Auto-check passed
  • Tool Specification

    kangarooking/system-prompt-skills

    当需要为 AI 定义工具接口、设计调用规范、实现工具发现与编排机制时调用此 skill。典型场景包括:设计 AI agent 的工具集、定义 JSON Schema/XML/TypeScript 格式的工具描述、实现工具权限控制与并行调度、设计子代理委托架构。

    205 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Memory System

    kangarooking/system-prompt-skills

    当需要为 AI 设计记忆存储、检索、应用和更新机制时调用此 skill。典型场景包括:设计持久化记忆架构(用户偏好、历史上下文、项目知识)、定义记忆的创建/读取/更新/删除生命周期、实现静默记忆应用(不在回复中透露记忆内容)、管理敏感记忆边界。

    205 GitHub stars~1.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Personality System

    kangarooking/system-prompt-skills

    当需要在基础身份之上叠加可切换的人格风格层时调用此 skill。典型场景包括:为同一产品提供多种人格选项(如 GPT-5.1 的 friendly/professional/quirky 模式)、设计人格切换机制、防止人格泄露到用户内容中。

    205 GitHub stars~1k tokensUpdated 5 mo ago
    Auto-check passed
  • Conversation Flow

    kangarooking/system-prompt-skills

    当系统提示词需要定义 AI 如何分类用户意图、路由到不同处理流程、决定澄清策略和自主度级别时调用此 Skill。适用于多任务型 AI 助手、客服机器人、编程工具、研究助手等需要结构化对话管理的场景。不适用于:纯问答型系统(无任务执行)、单轮交互(无对话状态)、简单的 prompt 模板(无路由逻辑)。当需求仅涉及"输出什么格式"而非"如何决定输出什么"时,应该用…

    205 GitHub starsUsed in 1 repo~788 tokens
    Auto-check passed
  • Safety Guardrails

    kangarooking/system-prompt-skills

    当需要为 AI 系统设计多层安全防线、内容过滤策略和伦理边界时调用此 skill。典型场景包括:设计拒绝策略与升级机制、防御 prompt 注入攻击、实现领域特定安全规则(教育、医疗、金融等)、定义 AI 的价值观锚点。

    205 GitHub stars~1.2k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Voice Optimization

What does Voice Optimization do?

当系统提示面向语音交互场景(语音助手、语音搜索、有声回答、电话客服 AI)时调用。适用于需要将文本输出优化为口语表达的系统提示设计。不适用于纯文本聊天界面,不适用于语音合成(TTS)技术选型,不适用于图像/视频多模态场景。. Voice Optimization is an agent skill from kangarooking/system-prompt-skills.

When should I use Voice Optimization?

Voice Optimization fits situations like: tasks that involve Text to speech and voice.

How do I install Voice Optimization in Claude Code?

Run `npx skills add kangarooking/system-prompt-skills --skill voice-optimization -a claude-code`. Or copy the skill folder (voice-optimization in kangarooking/system-prompt-skills) into .claude/skills/voice-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Voice Optimization in Codex?

Run `npx skills add kangarooking/system-prompt-skills --skill voice-optimization -a codex`. Or copy the skill folder (voice-optimization in kangarooking/system-prompt-skills) into .agents/skills/voice-optimization in your project. Codex loads it when a task matches its description.

Can I use Voice Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kangarooking/system-prompt-skills --skill voice-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-optimization, .gemini/skills/voice-optimization, .github/skills/voice-optimization and .opencode/skills/voice-optimization in your project.

What does Voice Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Voice Optimization is instructions for the agent only.

Does Voice Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voice Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice Optimization use?

Voice Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice Optimization use?

About 642 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Optimization?

Skills that share tags, products or a category with Voice Optimization: Xsai (moeru-ai/airi, 50k stars), Hri (terrense/ros2-multimodal-robot-collab, 111 stars), Video Understanding (zenstory-ai/video-recap-skills, 555 stars) and Unified LLM API (Prism-Shadow/penguin-harness, 2.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Optimization?

kangarooking (a GitHub user) maintains it in kangarooking/system-prompt-skills, which has 205 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on May 4, 2026.

Source: kangarooking/system-prompt-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.