Route voice input from the realtime voice agent pipeline based on message tags.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Input Branching

skills CLI
$ npx skills add autonomous-ai/Physical-AI-Operating-System --skill input-branching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/Physical-AI-Operating-System input-branching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/Physical-AI-Operating-System.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/input-branching .claude/skills/input-branching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
input-branching
GitHub stars
381
Token cost
~1.2k tokens
SKILL.md length
560 words
Files
2
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

Route voice input from the realtime voice agent pipeline based on message tags.

  • Works in 6 steps: [voice-instruction] is the primary… → [HANDLED] → always NO_REPLY, and never… → Log context from [HANDLED] silently —… → …
  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers Message Formats, Rules, Spoken delivery and Examples
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Input Branching is an agent skill from autonomous-ai/Physical-AI-Operating-System. Route voice input from the realtime voice agent pipeline based on message tags.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis. The repository describes itself as: The open-source operating system for physical AI. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/input-branching”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. [voice-instruction] is the primary input. When present, use it over [transcript].
  2. [HANDLED] → always NO_REPLY, and never act on it. No exceptions. Even if the reply seems wrong, incomplete, or claims something was done…
  3. Log context from [HANDLED] silently — memory/mood/wellbeing only. If the exchange reveals mood, intent, or information worth tracking…
  4. Never echo tags. [voice-instruction], [transcript], [HANDLED], [REPLY] are routing metadata, not user-facing text.
  5. No prefix = normal voice event. Process the message as-is.
  6. Compound instructions: execute EVERY clause. A [voice-instruction] often carries several actions plus a question ("Rotate to the right…

What it can do on your machine

Read from SKILL.md and the folder at commit f1b9ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Input Branching loads about 1.2k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 560 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/Physical-AI-Operating-System at commit f1b9ebe, republished under its Apache-2.0 licence (© autonomous-ai). 560 words, ~1,219 tokens.

Download SKILL.mdSave it as .claude/skills/input-branching/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
input-branching
description
Route voice input from the realtime voice agent pipeline based on message tags.

Input Branching

Voice input passes through a realtime voice agent before reaching you. The agent handles chit-chat directly and delegates everything else. Message tags tell you which path was taken.

Message Formats

Delegated with instruction
[voice-instruction] Play jazz music on Spotify
[transcript] play some jazz please on spotify
  • [voice-instruction] — the realtime agent's clean summary of what the user wants. Use this as the primary input.
  • [transcript] — raw STT transcript. Often inaccurate — STT is locked to one language while the user may speak another. Treat as noisy supplementary context only, never as the source of truth.
Delegated without instruction (fallback)
turn on the lights

No tags. Process as a normal voice event.

Handled (history entry)
[HANDLED] Hey, how's it going?
[REPLY] I'm doing great! How about you?
  • [HANDLED] — the user's original message, already answered via TTS.
  • [REPLY] — what the realtime agent said. The user already heard it.

This is a history entry only. The conversation already happened. You are being notified, not asked to respond.

Rules

  1. [voice-instruction] is the primary input. When present, use it over [transcript].
  2. [HANDLED] → always NO_REPLY, and never act on it. No exceptions. Even if the reply seems wrong, incomplete, or claims something was done that wasn't — the user already heard it. Do not correct, echo, paraphrase, add to, finish or redo it.
  3. Log context from [HANDLED] silently — memory/mood/wellbeing only. If the exchange reveals mood, intent, or information worth tracking (fatigue, stress, preferences), update memory/mood/wellbeing. Never record names or identities from it. No device, camera, face, voice-enrollment or messaging calls, even when a skill you loaded earlier (e.g. face-enroll) is mid-flow. A pending task resumes only on a live (non-[HANDLED]) message, including one marked [realtime-handoff].
  4. Never echo tags. [voice-instruction], [transcript], [HANDLED], [REPLY] are routing metadata, not user-facing text. Interpret routing, transcript differences, and speaker metadata silently. Assistant reply text is spoken aloud: start with any required HW markers and the user-facing answer, without narrating which input or skill you chose.
  5. No prefix = normal voice event. Process the message as-is.
  6. Compound instructions: execute EVERY clause. A [voice-instruction] often carries several actions plus a question ("Rotate to the right, hold that position, then describe what you see"). Handle each clause, in order, with the skill it belongs to — and never confirm a clause you did not actually perform. Answering the question half while silently skipping the movement (or vice versa) is the worst outcome: the user hears "done" for something that never happened.
Show full SKILL.md (179 more words)Show less

Spoken delivery

Keep routing and tool planning in the provider's native thinking channel; do not copy a reasoning summary into ordinary text before/after tools or the final reply. If unavailable, omit analysis. For a simple voice command, give one short confirmation after the required action, including a failure or essential next step when applicable. For a question, lead with the answer and use only the detail needed; longer explanations, stories and requested readbacks remain valid. Brevity must not drop clauses from compound instructions, required consent, or important safety information. Use available context and avoid redundant skill reads or repeated checks merely to refine phrasing.

Examples

Delegated — use instruction, ignore noisy transcript:

[voice-instruction] Set brightness to 50%
[transcript] set the brightness to like half

→ Process "Set brightness to 50%". Route to LED skill.

Delegated — transcript in wrong language (user spoke Vietnamese, STT is English):

[voice-instruction] Turn off the light and play some relaxing music
[transcript] ton of delay and play some relate music

→ Process the [voice-instruction]. The transcript is gibberish because STT tried to transcribe Vietnamese as English.

Delegated — transcript garbled by accent/noise:

[voice-instruction] What is the weather like tomorrow
[transcript] what is the wetter like to morrow

→ Process the [voice-instruction]. The transcript has misheard words.

Delegated without instruction:

play some jazz music

→ Process as-is. Route to music skill.

Handled — silent:

[HANDLED] What time is it?
[REPLY] It's 3:15 PM.

→ NO_REPLY.

Handled — log mood:

[HANDLED] I'm so tired today
[REPLY] That sounds rough. Maybe take a short break?

→ Log fatigue via mood/wellbeing. NO_REPLY.

© autonomous-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/input-branching of autonomous-ai/Physical-AI-Operating-System.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit f1b9ebe

Compare with similar skills

Input Branching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Input Branching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Input Branching this skillautonomous-ai/Physical-AI-Operating-System381—~1.2kAutomated safety check: PassApache-2.0
TriageTalAter/annyang6.8k1 repos~810Automated safety check: NotesMIT
Yichen Asrmcncarl/yichen-skills4.3k—~780Automated safety check: PassCustom licence
Dingtalk MinutesDingTalk-Real-AI/dingtalk-workspace-cli3.2k—~2.3kAutomated safety check: PassApache-2.0
Youtube FetcherJimmySadek/youtube-fetcher-to-markdown485—~3.1kAutomated safety check: PassMIT
Yichen Web Researchmcncarl/yichen-skills4.3k—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Triage

    TalAter/annyang

    Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.

    6.8k GitHub starsUsed in 1 repo~810 tokens
    AI & LLM EngineeringAuto-check: notes
  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.3k GitHub stars~780 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Dingtalk Minutes

    DingTalk-Real-AI/dingtalk-workspace-cli

    钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。

    3.2k GitHub stars~2.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…

    485 GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Yichen Web Research

    mcncarl/yichen-skills

    逸尘自用的互联网研究总入口。用于跨平台且跨阶段、用户尚未确定工具,或明确要求对公司、产品、人物、技术、行业和领域做横纵分析、发展史加现状对比或有来源约束的系统深度研究;先生成有截止日期和证据闸门的计划,再把搜索发现、候选核验、有限归档、按需转写和证据综合路由到…

    4.3k GitHub stars~1.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Volcengine Asr

    ysyecust/lecture-to-notes

    Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS.

    270 GitHub stars~783 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from autonomous-ai/Physical-AI-Operating-System

All 28 skills in this repo
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Claude Code Buddy

    autonomous-ai/Physical-AI-Operating-System

    Push Claude Code activity to the user's device (e.g. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Connectors

    autonomous-ai/Physical-AI-Operating-System

    Discover and use linked third-party services (Gmail, Google Calendar, Google Drive, Notion, Figma, Asana, Linear, GitHub, Ahrefs, Facebook Fan Page and others).

    381 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Audio

    autonomous-ai/Physical-AI-Operating-System

    Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

    381 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Input Branching

What does Input Branching do?

Route voice input from the realtime voice agent pipeline based on message tags. Input Branching is an agent skill from autonomous-ai/Physical-AI-Operating-System. Route voice input from the realtime voice agent pipeline based on message tags.

When should I use Input Branching?

Input Branching fits situations like: tasks that involve Speech recognition and synthesis.

How do I install Input Branching in Claude Code?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill input-branching -a claude-code`. Or copy the skill folder (skills/input-branching in autonomous-ai/Physical-AI-Operating-System) into .claude/skills/input-branching in your project. Claude Code loads it when a task matches its description.

How do I install Input Branching in Codex?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill input-branching -a codex`. Or copy the skill folder (skills/input-branching in autonomous-ai/Physical-AI-Operating-System) into .agents/skills/input-branching in your project. Codex loads it when a task matches its description.

Can I use Input Branching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill input-branching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/input-branching, .gemini/skills/input-branching, .github/skills/input-branching and .opencode/skills/input-branching in your project.

What does Input Branching need to run?

SKILL.md names no scripts, command-line tools or credentials: Input Branching is instructions for the agent only.

Does Input Branching access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Input Branching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Input Branching use?

Input Branching is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Input Branching use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Input Branching?

Skills that share tags, products or a category with Input Branching: Triage (TalAter/annyang, 6.8k stars), Yichen Asr (mcncarl/yichen-skills, 4.3k stars), Dingtalk Minutes (DingTalk-Real-AI/dingtalk-workspace-cli, 3.2k stars) and Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Input Branching?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/Physical-AI-Operating-System, which has 381 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/Physical-AI-Operating-System on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.