TTS speech + mic/speaker mute for privacy. An agent skill from autonomous-ai/Physical-AI-Operating-System.

Apache-2.0Auto-check passedMedia & Creative

Install Voice

skills CLI
$ npx skills add autonomous-ai/Physical-AI-Operating-System --skill voice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/Physical-AI-Operating-System voice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/Physical-AI-Operating-System.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/voice .claude/skills/voice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice
GitHub stars
381
Token cost
~2.5k tokens
SKILL.md length
1,313 words
Files
2
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

TTS speech + mic/speaker mute for privacy. An agent skill from autonomous-ai/Physical-AI-Operating-System.

  • Works in 4 steps: Determine if you need explicit speech… → Optionally check if TTS is busy: GET… → If tts_speaking is true, wait or skip → …
  • Silence requests
  • SKILL.md covers Quick Start, Brief speech that completes…, Workflow — explicit additional… and Examples, plus 8 more sections
  • Calls curl

What it does

Voice is an agent skill from autonomous-ai/Physical-AI-Operating-System. TTS speech + mic/speaker mute for privacy. MUST trigger on meetings, calls, privacy, silence requests. "meeting"/"call"/"private" = mic+speaker mute. "be quiet"/"silent" = speaker mute only. Always call HW markers — never just text.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: The open-source operating system for physical AI. The licence is Apache-2.0.

When your agent uses it

  • Silence requests
  • Tasks that involve Text to speech and voice

Example prompts

  • “meeting”
  • “private”
  • “be quiet”
  • “/voice”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Determine if you need explicit speech beyond your normal reply
  2. Optionally check if TTS is busy: GET /voice/status
  3. If tts_speaking is true, wait or skip
  4. Call POST /voice/speak with plain text

What it can do on your machine

Read from SKILL.md and the folder at commit f1b9ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice loads about 2.5k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,313 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/Physical-AI-Operating-System at commit f1b9ebe, republished under its Apache-2.0 licence (© autonomous-ai). 1,313 words, ~2,466 tokens.

Download SKILL.mdSave it as .claude/skills/voice/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
voice
description
TTS speech + mic/speaker mute for privacy. MUST trigger on meetings, calls, privacy, silence requests. "meeting"/"call"/"private" = mic+speaker mute. "be quiet"/"silent" = speaker mute only. Always call HW markers — never just text.

Voice — Speak Through Speaker

Quick Start

Choose the relevant path first:

  • Mic/speaker mute, unmute, or privacy request: apply Ambient Audio Guard, then the applicable mute/unmute or Meeting Mode section below. Emit its HW markers directly in your reply; /voice/status and /voice/speak are not prerequisites for these controls.
  • Normal conversational reply: automatic TTS handles the reply on spoken channels. No explicit speech call is needed.
  • Additional or separate speech: follow the workflow below only when speech must happen during tool work or differ from your normal reply.

Brief speech that completes the task

  • Simple command or automatic reaction: normally one short sentence. Include the actual outcome, required question, or essential next step; do not add a second acknowledgment just to sound conversational.
  • Mute/privacy controls: keep the confirmation AND how to unmute. Enrollment: keep consent/name questions and the success or failure confirmation. Guard alerts: keep the hazard and required action; returning-user summaries may need several short sentences. Never remove these to meet a word target.
  • Direct questions: answer first, then only the detail needed. Requested stories, explanations, exact readbacks, and important safety guidance may be longer.
  • All analysis belongs in the provider's native thinking channel, not ordinary text or an explicit speech API payload. Do not recap reasoning after thinking ends; if no native channel is available, omit analysis. Required tool calls need no spoken introduction. Do not make extra calls solely to narrate progress.
  • Explicit early speech is for a useful user-facing cue needed before an action or during a wait, not for route/skill/cooldown commentary. Say it once; do not repeat it in the final reply. Automatic TTS is sufficient for routine events.

Workflow — explicit additional or separate speech

  1. Determine if you need explicit speech beyond your normal reply:
    • Normal conversational reply -> do NOT call this skill, TTS is automatic
    • Need to speak while also performing tool calls -> use POST /voice/speak
    • Need to speak different text than your chat reply -> use POST /voice/speak
    • Reacting to a sensing event before reply is finalized -> use POST /voice/speak
  2. Optionally check if TTS is busy: GET /voice/status
  3. If tts_speaking is true, wait or skip
  4. Call POST /voice/speak with plain text

Examples

Input: Normal conversational reply Output: Do NOT call this skill. Just reply normally — your text is automatically spoken.

Input: You need to greet the user while also activating a scene Output: Call POST /voice/speak with {"text": "Good morning!"} in parallel with the Scene API call.

Input: You want to say something different from your chat reply Output: Call POST /voice/speak with the spoken text. Then provide your chat reply separately.

Input: User says "say something" / "tell me a joke" Output: Do NOT call this skill. Just reply normally with the joke — automatic TTS handles it.

Tools

Use Bash with curl to call the HTTP API at http://127.0.0.1:5001.

Speak text
bash
curl -s -X POST http://127.0.0.1:5001/voice/speak \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello, this is a test."}'

Text max 2000 characters. Returns immediately; audio plays in background.

Check voice status
bash
curl -s http://127.0.0.1:5001/voice/status

Response:

json
{
  "voice_available": true,
  "voice_listening": false,
  "tts_available": true,
  "tts_speaking": false
}

Error Handling

  • If tts_speaking is true, the speaker is busy. Wait briefly or skip the explicit speech.
  • If voice_available or tts_available is false, inform the user: "Voice output is currently unavailable."
  • If the API is unreachable, fall back to chat-only reply. Speech is non-critical.

Rules

  • Normal replies = automatic TTS. Do NOT call /voice/speak for every response.
  • Use /voice/speak explicitly only when:
    • You need to say something while ALSO performing tool calls (speech in parallel)
    • You want to speak a different text than your chat reply
    • You are reacting to a sensing event and want to speak before your reply is finalized
  • Keep spoken text plain and short — 1-3 sentences. No markdown, no emoji, no formatting. Plain natural speech only.
    • Exception: reading a draft back for approval. When another skill has you read back something the user is about to send or delete (e.g. connectors before sending, sharing, or deleting), speak it in full. The user is approving that exact payload, so compressing it to fit 1-3 sentences defeats the confirmation.
  • Match the user's language — if they speak Vietnamese, speak Vietnamese.
  • Text max 2000 characters.
  • For volume control, use the Audio skill, not this skill.

Ambient Audio Guard (read FIRST)

If the user's message contains the literal token [ambient] (typically alongside a [user] priority marker, e.g. [user] [ambient] ...), it is overheard passive audio — NOT directed at the device. Do NOT trigger mute markers from a single bare word like "call", "meeting", "private", or a clipped fragment.

Mute markers [HW:/voice/mute:{}] and [HW:/speaker/mute:{}] may fire on ambient audio ONLY when the transcript contains a clear, complete intent:

  • "I'm on a call" / "I have a meeting" / "I need privacy" / "stop listening"
  • Or directly addresses it by name with a mute request

When ambient is ambiguous, reply naturally or stay quiet — DO NOT mute. Voice commands (no [ambient] token in the message) follow the normal trigger tables below.

Show full SKILL.md (517 more words)Show less

Mic Mute/Unmute (Privacy)

Users can mute the mic for privacy (meetings, calls). Use HW markers — no curl needed.

Mute mic
[HW:/voice/mute:{}]

Stops all listening — STT, wake word, sound detection. The device becomes fully deaf. Unmute via physical button, web toggle, or Telegram command.

Trigger phrases (MANDATORY — must call HW marker, not just reply with text)

Any phrase about privacy, meetings, calls, not wanting to be heard, or asking the device to stop listening MUST trigger [HW:/voice/mute:{}]. Do NOT just acknowledge — you MUST include the HW marker.

User saysAction
"don't listen" / "stop listening" / "mute" / "mute mic"[HW:/voice/mute:{}] — MUST call
"I'm in a meeting" / "I have a meeting" / "I need a private meeting" / "meeting"[HW:/voice/mute:{}] — MUST call
"I'm on a call" / "I have a call" / "phone call"[HW:/voice/mute:{}] — MUST call
"privacy" / "private" / "give me privacy" / "need privacy"[HW:/voice/mute:{}] — MUST call
"don't hear me" / "in a meeting" / "mute mic" / "stop hearing"[HW:/voice/mute:{}] — MUST call
Examples

Input: "I have a meeting now" Output: [HW:/voice/mute:{}] OK, I'll stop listening. Press the button when you need me.

Input: "Stop listening" Output: [HW:/voice/mute:{}] Got it, mic off. Press my button to unmute.

Input: "I need a private meeting" Output: [HW:/voice/mute:{}] Got it, going silent. Press the button when you're done.

Input: "I'm on a call" Output: [HW:/voice/mute:{}] Muting now. Press the button to unmute when you're done.

Unmute mic
[HW:/voice/unmute:{}]

Use when a Telegram or web chat user asks to unmute remotely. Voice unmute is not possible (the device is deaf when muted). Physical button also unmutes.

User says (via Telegram/web)Action
"unmute" / "start listening" / "listen again" / "mic on"[HW:/voice/unmute:{}] — only works from Telegram/web, not voice

Speaker Mute/Unmute (Silent Mode)

Suppress all audio output — TTS, music, backchannel. The device stays silent but still listens.

Ambient guard above also applies here — bare fragments like "quiet" or "silence" in [ambient] audio do NOT trigger speaker mute.

Mute speaker
[HW:/speaker/mute:{}]
Unmute speaker
[HW:/speaker/unmute:{}]

Mic still works when speaker is muted — user can unmute via voice command.

CRITICAL: "unmute" ≠ "mute". Read the EXACT word. Do NOT call mute when user says unmute.

User saysAction
"be quiet" / "silent mode" / "don't talk" / "hush" / "silence"[HW:/speaker/mute:{}] — MUST call
"unmute" / "you can talk" / "unmute speaker" / "speak again" / "talk again"[HW:/speaker/unmute:{}] — MUST call (UN-mute, not mute!)
Examples

Input: "Be quiet" Output: [HW:/speaker/mute:{}] Going silent. Just say "you can talk" when you want me back.

Input: "You can talk now" Output: [HW:/speaker/unmute:{}] I'm back!

Meeting Mode (mic + speaker mute)

When user mentions a meeting or call and wants full silence (not just speaker), mute BOTH mic and speaker. Ambient guard above applies — explicit intent required, not bare fragments.

Input: "I'm in a meeting" Output: [HW:/voice/mute:{}][HW:/speaker/mute:{}] Meeting mode — fully silent. Press the button when you're done.

Rules

  • Mic mute is the last thing the device hears via voice — after mic mute, only physical button, web toggle, or Telegram can unmute
  • Voice unmute for mic is impossible (the device is deaf) — tell user to press the button
  • Speaker mute: user can still voice-unmute (mic still on)
  • TTS still works when only mic is muted — the device can speak but not hear
  • Always confirm mute with how to unmute

© autonomous-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/voice of autonomous-ai/Physical-AI-Operating-System.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit f1b9ebe

Compare with similar skills

Voice next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice this skillautonomous-ai/Physical-AI-Operating-System381—~2.5kAutomated safety check: PassApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Videodbaffaan-m/ECC275k3 repos~3.5kAutomated safety check: NotesMIT
Musictadaspetra/loop2963 repos~827Automated safety check: PassMIT
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 3 repos~827 tokens
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes

More from autonomous-ai/Physical-AI-Operating-System

All 28 skills in this repo
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Claude Code Buddy

    autonomous-ai/Physical-AI-Operating-System

    Push Claude Code activity to the user's device (e.g. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Connectors

    autonomous-ai/Physical-AI-Operating-System

    Discover and use linked third-party services (Gmail, Google Calendar, Google Drive, Notion, Figma, Asana, Linear, GitHub, Ahrefs, Facebook Fan Page and others).

    381 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Audio

    autonomous-ai/Physical-AI-Operating-System

    Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

    381 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Voice

What does Voice do?

TTS speech + mic/speaker mute for privacy. An agent skill from autonomous-ai/Physical-AI-Operating-System. Voice is an agent skill from autonomous-ai/Physical-AI-Operating-System. TTS speech + mic/speaker mute for privacy.

When should I use Voice?

Voice fits situations like: silence requests; tasks that involve Text to speech and voice.

How do I install Voice in Claude Code?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill voice -a claude-code`. Or copy the skill folder (skills/voice in autonomous-ai/Physical-AI-Operating-System) into .claude/skills/voice in your project. Claude Code loads it when a task matches its description.

How do I install Voice in Codex?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill voice -a codex`. Or copy the skill folder (skills/voice in autonomous-ai/Physical-AI-Operating-System) into .agents/skills/voice in your project. Codex loads it when a task matches its description.

Can I use Voice in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill voice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice, .gemini/skills/voice, .github/skills/voice and .opencode/skills/voice in your project.

What does Voice need to run?

Going by SKILL.md and its folder, Voice needs the command-line tools its instructions call (curl).

Does Voice access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Voice safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice use?

Voice is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice?

Skills that share tags, products or a category with Voice: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Videodb (affaan-m/ECC, 275k stars) and Music (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/Physical-AI-Operating-System, which has 381 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/Physical-AI-Operating-System on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.