Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

Apache-2.0Auto-check passedMedia & Creative

Install Audio

skills CLI
$ npx skills add autonomous-ai/Physical-AI-Operating-System --skill audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/Physical-AI-Operating-System audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/Physical-AI-Operating-System.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audio .claude/skills/audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio
GitHub stars
381
Token cost
~1k tokens
SKILL.md length
458 words
Files
2
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

  • Works in 4 steps: Determine the user's audio hardware need → For explicit volume/tone/recording… → Execute the appropriate API call → …
  • TTS/speech (that is the Voice skill)
  • SKILL.md covers Quick Start, Workflow, Examples and Tools, plus 3 more sections
  • Calls curl

What it does

Audio is an agent skill from autonomous-ai/Physical-AI-Operating-System. Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio. Do NOT use for TTS/speech (that is the Voice skill).

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: The open-source operating system for physical AI. The licence is Apache-2.0.

When your agent uses it

  • TTS/speech (that is the Voice skill)
  • Tasks that involve Text to speech and voice

Example prompts

  • “/audio”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Determine the user's audio hardware need
  2. For explicit volume/tone/recording requests, call the requested endpoint directly and handle its response. Use GET /audio for device…
  3. Execute the appropriate API call
  4. Confirm the action to the user

What it can do on your machine

Read from SKILL.md and the folder at commit f1b9ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio loads about 1k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 458 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/Physical-AI-Operating-System at commit f1b9ebe, republished under its Apache-2.0 licence (© autonomous-ai). 458 words, ~1,010 tokens.

Download SKILL.mdSave it as .claude/skills/audio/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
audio
description
Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio. Do NOT use for TTS/speech (that is the Voice skill).

Audio Control

Quick Start

Control the device's speaker and microphone hardware directly. Use this for volume adjustments, test tones, and raw audio recording. This is LOW-LEVEL hardware control only.

Workflow

  1. Determine the user's audio hardware need:
    • Volume adjustment -> use POST /audio/volume
    • Check current volume -> use GET /audio/volume
    • Diagnostics / test -> use POST /audio/play-tone
    • Raw recording -> use POST /audio/record
  2. For explicit volume/tone/recording requests, call the requested endpoint directly and handle its response. Use GET /audio for device diagnostics, not routine preflight. Relative volume changes still need GET /audio/volume.
  3. Execute the appropriate API call
  4. Confirm the action to the user

Examples

Input: "Louder please" / "Turn it up" Output: Check current volume with GET /audio/volume, then increase by ~15 with POST /audio/volume. Confirm: "Volume set to 85%."

Input: "Set volume to 50%" Output: Call POST /audio/volume with {"volume": 50}. Confirm: "Volume set to 50%."

Input: "Mute" / "Too loud" Output: Call POST /audio/volume with {"volume": 0}. Confirm: "Muted."

Input: "I can't hear you" Output: Check current volume with GET /audio/volume, then increase it. Confirm with the new level.

Input: "Say something" / "Tell me a joke" Output: Do NOT use this skill. Just reply normally — your voice pipeline handles TTS automatically.

Tools

Use Bash with curl to call the HTTP API at http://127.0.0.1:5001.

Check audio devices
bash
curl -s http://127.0.0.1:5001/audio

Response:

json
{
  "output_device": 0,
  "input_device": 1,
  "available": true
}
Set volume
bash
curl -s -X POST http://127.0.0.1:5001/audio/volume \
  -H "Content-Type: application/json" \
  -d '{"volume": 70}'

Volume range: 0 (mute) to 100 (max).

Get current volume
bash
curl -s http://127.0.0.1:5001/audio/volume

Response: {"control": "Speaker", "volume": 70}

Play test tone
bash
curl -s -X POST "http://127.0.0.1:5001/audio/play-tone?frequency=440&duration_ms=500"

Plays a sine wave. Use for audio testing only. Keep it short (< 1 second).

Show full SKILL.md (209 more words)Show less
Record audio
bash
record_dir=$(mktemp -d /tmp/autonomous-recording.XXXXXX) &&
curl --fail --silent --show-error -X POST \
  "http://127.0.0.1:5001/audio/record?duration_ms=3000" \
  --output "$record_dir/recording.wav" &&
printf 'Recording saved: %s\n' "$record_dir/recording.wav"

Records from the microphone and returns WAV bytes; save them to a unique file rather than printing binary data to the model. Adapt the duration to the user's request. Report the path only after the request succeeds; a failed request may leave an empty or partial file and is not a recording result. Do not re-record automatically after an uncertain outcome.

Error Handling

  • If GET /audio returns "available": false, inform the user: "The speaker/microphone is not connected right now."
  • If the API is unreachable, inform the user: "I couldn't access the audio hardware. The service may be unavailable."
  • If the user requests a volume outside 0-100, clamp to the valid range.

Rules

  • Default volume is usually 70%. Adjust based on user preference.
  • Audio = volume knob, raw mic recording, test beeps. No AI speech processing.
  • Voice = AI-powered speech (TTS/STT). Uses Audio hardware underneath but is a separate skill.
  • When the user says "I can't hear you" or "too loud", adjust volume via this skill.
  • Do NOT use this skill for TTS or speech output — that is handled by the Voice skill and the automatic voice pipeline.

Output Template

[Audio] {action} — {details}

Examples:

  • [Audio] Volume set — 70%
  • [Audio] Volume set — muted (0%)
  • [Audio] Test tone played — 440Hz, 500ms
  • [Audio] Recording captured — 3000ms

© autonomous-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/audio of autonomous-ai/Physical-AI-Operating-System.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit f1b9ebe

Compare with similar skills

Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio this skillautonomous-ai/Physical-AI-Operating-System381—~1kAutomated safety check: PassApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Videodbaffaan-m/ECC275k3 repos~3.5kAutomated safety check: NotesMIT
Musictadaspetra/loop2963 repos~827Automated safety check: PassMIT
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 3 repos~827 tokens
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes

More from autonomous-ai/Physical-AI-Operating-System

All 28 skills in this repo
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Claude Code Buddy

    autonomous-ai/Physical-AI-Operating-System

    Push Claude Code activity to the user's device (e.g. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Connectors

    autonomous-ai/Physical-AI-Operating-System

    Discover and use linked third-party services (Gmail, Google Calendar, Google Drive, Notion, Figma, Asana, Linear, GitHub, Ahrefs, Facebook Fan Page and others).

    381 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Camera

    autonomous-ai/Physical-AI-Operating-System

    Camera control — snapshot, stream, and privacy toggle. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~3k tokensUpdated today
    Auto-check passed

Questions about Audio

What does Audio do?

Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio. Audio is an agent skill from autonomous-ai/Physical-AI-Operating-System. Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

When should I use Audio?

Audio fits situations like: TTS/speech (that is the Voice skill); tasks that involve Text to speech and voice.

How do I install Audio in Claude Code?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill audio -a claude-code`. Or copy the skill folder (skills/audio in autonomous-ai/Physical-AI-Operating-System) into .claude/skills/audio in your project. Claude Code loads it when a task matches its description.

How do I install Audio in Codex?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill audio -a codex`. Or copy the skill folder (skills/audio in autonomous-ai/Physical-AI-Operating-System) into .agents/skills/audio in your project. Codex loads it when a task matches its description.

Can I use Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio, .gemini/skills/audio, .github/skills/audio and .opencode/skills/audio in your project.

What does Audio need to run?

Going by SKILL.md and its folder, Audio needs the command-line tools its instructions call (curl).

Does Audio access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Audio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio use?

Audio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio?

Skills that share tags, products or a category with Audio: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Videodb (affaan-m/ECC, 275k stars) and Music (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/Physical-AI-Operating-System, which has 381 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/Physical-AI-Operating-System on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.