Camera control — snapshot, stream, and privacy toggle. An agent skill from autonomous-ai/Physical-AI-Operating-System.

Apache-2.0Auto-check passed

Install Camera

skills CLI
$ npx skills add autonomous-ai/Physical-AI-Operating-System --skill camera -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/Physical-AI-Operating-System camera --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/Physical-AI-Operating-System.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/camera .claude/skills/camera && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
camera
GitHub stars
381
Token cost
~3k tokens
SKILL.md length
1,668 words
Files
2
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

Camera control — snapshot, stream, and privacy toggle. An agent skill from autonomous-ai/Physical-AI-Operating-System.

  • Works in 2 steps: Run the Capture Protocol command (the… → Respond from the returned description,…
  • What do you see
  • SKILL.md covers Quick Start, Already-captured frame (reuse,…, Capture Protocol and Never describe the view…, plus 8 more sections
  • Calls curl

What it does

Camera is an agent skill from autonomous-ai/Physical-AI-Operating-System. Camera control — snapshot, stream, and privacy toggle. Trigger on "what do you see", "look at this", "take a photo", "don't look", "stop looking", "stop watching", "stop staring", "camera off", "camera on", "give me privacy". MUST call [HW:/camera/disable:{}] or [HW:/camera/enable:{}] when toggling — never just reply with text.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

The repository describes itself as: The open-source operating system for physical AI. The licence is Apache-2.0.

When your agent uses it

  • What do you see
  • Give me privacy

Example prompts

  • “what do you see”
  • “look at this”
  • “take a photo”
  • “/camera”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Run the Capture Protocol command (the /camera check and the look are one shell call). CAMERA_OFF / CAMERA_UNAVAILABLE → answer from that…
  2. Respond from the returned description, or inspect the returned path with an image tool when the server provides only a path (see Capture…

What it can do on your machine

Read from SKILL.md and the folder at commit f1b9ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Camera loads about 3k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,668 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/Physical-AI-Operating-System at commit f1b9ebe, republished under its Apache-2.0 licence (© autonomous-ai). 1,668 words, ~3,020 tokens.

Download SKILL.mdSave it as .claude/skills/camera/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
camera
description
Camera control — snapshot, stream, and privacy toggle. Trigger on "what do you see", "look at this", "take a photo", "don't look", "stop looking", "stop watching", "stop staring", "camera off", "camera on", "give me privacy". MUST call [HW:/camera/disable:{}] or [HW:/camera/enable:{}] when toggling — never just reply with text.

Camera

Quick Start

Accesses the device's built-in camera at http://127.0.0.1:5001 to take snapshots or check the environment. Only use when the user explicitly asks you to look at something.

Already-captured frame (reuse, don't re-snapshot)

If the incoming turn contains a line like:

[vision-image] <absolute-path-to-a.jpg> (a photo was JUST captured ...)

a photo was already taken for this exact request by the realtime voice layer (it captured the frame, then handed the turn to you — e.g. it timed out mid-answer), and the OS layer delivers it with this very message — either as an [image description] line (when the main model is text-only, a vision model has already analyzed the photo for you) or as an attached image. Answer the visual question from that description/attachment. Do NOT call /api/vision/look or /camera/snapshot again — re-snapshotting wastes time and may capture a different moment than what the user asked about. Do NOT read the path with a file tool — it is there for traceability only, and on text-only models a file-read image is silently dropped.

Use the capture protocol below when no current image or description was supplied for the request.

Capture Protocol

One tool call. It checks that the camera can actually deliver a frame, and only then takes the photo, sizes it, and gives you back what is in it:

bash
c=$(curl -s http://127.0.0.1:5001/camera); case "$c" in
  *'"disabled":true'*) echo "CAMERA_OFF" ;;
  *'"has_frame":false'*) echo "CAMERA_UNAVAILABLE" ;;
  *) curl -sX POST http://127.0.0.1:5000/api/vision/look \
       -H 'Content-Type: application/json' \
       -d '{"question":"<what the user asked>"}' ;;
esac

Read the output:

  • CAMERA_OFF → the user turned the camera off (privacy). Say so in one sentence and that they can say "camera on" to turn it back on. Stop. Do not look, do not enable it yourself.
  • CAMERA_UNAVAILABLE → the hardware is not delivering frames (not connected or not detected). Say the camera is not working right now. Stop.
  • {"data":{"description":"...","path":"..."}} → answer from description. It is the only thing in this turn that actually saw the frame. No description field (the main model can read images itself) → open path with your file/image tool.
  • Any other error → tell the user you couldn't see it this time. Do not guess.

Keeping the check inside the same shell command costs no extra tool round: a separate GET /camera turn would add one more model call (~5s) for every look. The server handles servo freeze, frame wait, and image sizing. No preparatory aim or sleep is needed unless the user explicitly requested a movement (see Move first, then snapshot).

For raw-frame export rather than a visual answer, use GET http://127.0.0.1:5001/camera/snapshot?save=true&width=768&quality=75 (HAL, port 5001). It writes a file and returns its path, without a description; the path alone is not visual evidence.

Never describe the view without an image

If you are about to say what you see, this turn MUST contain either a [vision-image] line or a /api/vision/look call whose answer you actually looked at, or a /api/vision/look description. Describing the room from memory, from an earlier turn's photo, or from a plausible guess ("same view — the desk, your screen…") is a fabrication, even when the guess happens to be close. No image → say you'll take a look and take one; never invent.

Move first, then snapshot

When the request combines a movement and a visual question ("turn right, hold it there, and tell me what you see"), fire the servo calls with curl during the turn (POST /servo/aim, POST /servo/hold), then call /api/vision/look. [HW:...] markers are executed only after your reply is composed, so a marker-based aim would move the device after the photo — you would describe the old view.

Workflow

  1. Run the Capture Protocol command (the /camera check and the look are one shell call). CAMERA_OFF / CAMERA_UNAVAILABLE → answer from that word and stop.
  2. Respond from the returned description, or inspect the returned path with an image tool when the server provides only a path (see Capture Protocol).

You also receive camera snapshots automatically as part of sensing events ([sensing:*] messages with images). You do not need the camera API for those — just look at the attached image.

Examples

Input: "What do you see right now?" Output: POST /api/vision/look → say: "I can see your desk with a laptop and a coffee mug. Looks like a productive setup!"

Input: "Is anyone in the room?" Output: POST /api/vision/look → say: "I can see one person sitting at the desk."

Input: "Take a photo" or "Send me a photo" Output: POST /api/vision/look → say what description reports.

Input: (sensing event with image already attached) Output: Do NOT call the camera API. Just look at the attached image and react.

Tools

Bash with curl — http://127.0.0.1:5000 for /api/vision/look, http://127.0.0.1:5001 for HAL camera control.

Look at the scene
bash
curl -sX POST http://127.0.0.1:5000/api/vision/look -H 'Content-Type: application/json' -d '{"question":"..."}'

Returns {"data":{"description":"...","path":"..."}}. See Capture Protocol.

Live stream
bash
curl -s http://127.0.0.1:5001/camera/stream

Returns an MJPEG stream (multipart/x-mixed-replace). Only use when continuous video is needed. Prefer snapshot for one-time checks.

Camera On/Off (Privacy Control)

Users can toggle the camera via voice or chat. Use HW markers — no curl needed.

Disable camera
[HW:/camera/disable:{}]

The user wants privacy. Camera stays off until the user explicitly re-enables it (voice or web toggle).

Enable camera
[HW:/camera/enable:{}]

Enabling is not seeing. The marker only flips the privacy switch; you have not captured a frame, and the hardware may not even be delivering one (a camera that is unplugged still "enables" fine). Never say "I can see you again", "there you are" or describe anything after an enable — say the camera is on and stop, or, if the user wants to be seen, run the Capture Protocol in the same turn and answer from what it returns.

Trigger phrases (MANDATORY — must call HW marker, not just reply with text)

Any phrase meaning "stop looking" or "camera off" MUST trigger [HW:/camera/disable:{}]. Any phrase meaning "look at me" or "camera on" MUST trigger [HW:/camera/enable:{}]. Do NOT just acknowledge — you MUST include the HW marker.

User saysAction
"don't look" / "stop looking" / "stop watching" / "privacy mode" / "camera off" / "don't watch me" / "give me privacy" / "stop staring"[HW:/camera/disable:{}] — MUST call
"look at me" / "camera on" / "you can look now" / "start watching"[HW:/camera/enable:{}] — MUST call
Show full SKILL.md (706 more words)Show less
"Look at ..." is ambiguous — route by what follows

The verb alone does NOT mean "turn the camera on". Only phrases about the device's own camera state belong in the table above.

User saysMeaningRoute to
"look at me" / "camera on" / "you can look now"turn the camera back on[HW:/camera/enable:{}]
"look at this" / "look at what I'm holding" / "what is this"a visual question about an object/api/vision/look (Workflow above)
"look at the desk / table / wall"a fixed locationservo-control /servo/aim
"look at the cup and follow it"a movable object to trackservo-tracking /servo/track
"find my keyboard" / "where is my cup" / "look around for X" / "where are you"something that may be OUT of view — a search, not a lookservo-control /servo/search (curl, in-turn)

"Find my X" is a search, not a look. One frame from wherever the head happens to point answers "what do you see", not "where is it". /servo/search sweeps the room, centres on the object, and returns the frame with a box on it — do not answer a find-request with a snapshot.

"Look at this" is a visual question, not a privacy toggle. The user is holding something up to be identified. Replying "Got it, camera on" answers a question they did not ask.

Examples

Input: "Look at this" / "Look at what I'm holding" Output: Capture Protocol command → say what the object is. Do NOT call [HW:/camera/enable:{}] — if the command prints CAMERA_OFF, say the camera is off and stop.

Input: "Don't watch me" Output: [HW:/camera/disable:{}] Got it, camera off. Just say "look at me" when you want me to see again.

Input: "Stop watching me" Output: [HW:/camera/disable:{}] I'll look away. Let me know when you want me back.

Input: "Look at me" / "Camera on" Output: [HW:/camera/enable:{}] Camera back on! — nothing more: no "I can see you", no description. You have not looked yet.

Camera off or unavailable (IMPORTANT)

A camera that is off (privacy) or has no frame (hardware) is a complete answer by itself: say it and stop. Do not enable the camera on the user's behalf, do not call /api/vision/look or /camera/snapshot anyway, do not ask them to enable it and then look — every extra step is a model round the user waits for. The Capture Protocol command already makes this decision for you from GET /camera (disabled, has_frame).

Error Handling

  • If capture fails, report the returned error without describing an unseen frame. /api/vision/look reports capture/description failures as errors; a raw /camera/snapshot request can return 503 when the camera is unavailable.
  • One failed /api/vision/look is final for this turn. Do NOT "try once more" and do NOT fall back to GET /camera/snapshot — it is the same capture path and fails the same way, costing another tool round. An error mentioning "not delivering frames" / "not connected or not detected" means the camera hardware is absent: tell the user the camera is not connected, and stop.
  • If the API is unreachable, inform the user that the camera is temporarily unavailable.
  • Never spend a separate tool round on GET /camera — the Capture Protocol command checks it in the same shell call as the look. Skip the call entirely when a current image/description was already supplied.
  • If a sensing event already included an image, do not call the camera API again.

Rules

  • Visual questions use /api/vision/look — the server handles capture and model-compatible image evidence.
  • Raw-frame export uses /camera/snapshot?save=true&width=768&quality=75 — read the returned path; never invent filenames or treat a path as a description.
  • Image delivery is handled automatically by the system — do not manually send images via tools.
  • Never use the camera proactively without the user's request — respect privacy.
  • Never disable/enable camera on your own — only toggle when the user explicitly asks or when a system trigger requires it (guard mode, scene change).
  • Don't repeatedly snapshot without reason.
  • Don't call the camera API when a sensing event already included an image.
  • Prefer /api/vision/look for one-time visual questions; use raw snapshot for frame export and /camera/stream only for continuous video.
  • When describing what you see, be specific and helpful.
  • If camera is unavailable, inform the user clearly and move on.

Output Template

After a visual tool call, answer from its description or the image you inspected. For a privacy toggle, include the corresponding marker:

[HW:/camera/disable:{}] Camera off.

© autonomous-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/camera of autonomous-ai/Physical-AI-Operating-System.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit f1b9ebe

Compare with similar skills

Camera next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Camera compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Camera this skillautonomous-ai/Physical-AI-Operating-System381—~3kAutomated safety check: PassApache-2.0
Privacy Policythedaviddias/Front-End-Checklist74k—~617Automated safety check: PassMIT
Privacy Policyphuryn/pm-skills27k—~2.7kAutomated safety check: PassMIT
Data Privacy Controlssickn33/agentic-awesome-skills47k1 repos~3.4kAutomated safety check: PassMIT
Performing Privacy Impact Assessmentmukul975/Anthropic-Cybersecurity-Skills34k—~2.6kAutomated safety check: PassApache-2.0
Privacy By Designsickn33/agentic-awesome-skills47k2 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Privacy Policy

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing whether a website has a visible, accessible privacy policy link, particularly in the footer navigation.

    74k GitHub stars~617 tokensUpdated 2 days ago
    Legal & ComplianceAuto-check passed
  • Privacy Policy

    phuryn/pm-skills

    Draft a detailed privacy policy covering data types, jurisdiction, GDPR and compliance considerations, and clauses needing legal review.

    27k GitHub stars~2.7k tokensUpdated 23 days ago
    Legal & ComplianceAuto-check passed
  • Data Privacy Controls

    sickn33/agentic-awesome-skills

    Data privacy control register: data category, lawful basis, retention period, access roles, encryption and consent requirement per module.

    47k GitHub starsUsed in 1 repo~3.4k tokens
    Legal & ComplianceAuto-check passed
  • Performing Privacy Impact Assessment

    mukul975/Anthropic-Cybersecurity-Skills

    Automates the Privacy Impact Assessment (PIA) workflow including data flow mapping, privacy risk scoring matrices, GDPR Article 35 DPIA and CCPA/CPRA alignment checks, data inventory cataloging, and…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Privacy By Design

    sickn33/agentic-awesome-skills

    A skill your agent uses when building apps that collect user data.

    47k GitHub starsUsed in 2 repos~1.8k tokens
    Legal & ComplianceAuto-check passed
  • Monitor Stream

    ruvnet/ruflo

    Stream live swarm events using the Monitor tool for real-time observability

    74k GitHub stars~188 tokensUpdated today
    DevOps & CloudAuto-check passed

More from autonomous-ai/Physical-AI-Operating-System

All 28 skills in this repo
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Claude Code Buddy

    autonomous-ai/Physical-AI-Operating-System

    Push Claude Code activity to the user's device (e.g. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Connectors

    autonomous-ai/Physical-AI-Operating-System

    Discover and use linked third-party services (Gmail, Google Calendar, Google Drive, Notion, Figma, Asana, Linear, GitHub, Ahrefs, Facebook Fan Page and others).

    381 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Audio

    autonomous-ai/Physical-AI-Operating-System

    Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

    381 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Camera

What does Camera do?

Camera control — snapshot, stream, and privacy toggle. An agent skill from autonomous-ai/Physical-AI-Operating-System. Camera is an agent skill from autonomous-ai/Physical-AI-Operating-System. Camera control — snapshot, stream, and privacy toggle.

When should I use Camera?

Camera fits situations like: what do you see; give me privacy.

How do I install Camera in Claude Code?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill camera -a claude-code`. Or copy the skill folder (skills/camera in autonomous-ai/Physical-AI-Operating-System) into .claude/skills/camera in your project. Claude Code loads it when a task matches its description.

How do I install Camera in Codex?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill camera -a codex`. Or copy the skill folder (skills/camera in autonomous-ai/Physical-AI-Operating-System) into .agents/skills/camera in your project. Codex loads it when a task matches its description.

Can I use Camera in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill camera -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/camera, .gemini/skills/camera, .github/skills/camera and .opencode/skills/camera in your project.

What does Camera need to run?

Going by SKILL.md and its folder, Camera needs the command-line tools its instructions call (curl).

Does Camera access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Camera safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Camera use?

Camera is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Camera use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Camera?

Skills that share tags, products or a category with Camera: Privacy Policy (thedaviddias/Front-End-Checklist, 74k stars), Privacy Policy (phuryn/pm-skills, 27k stars), Data Privacy Controls (sickn33/agentic-awesome-skills, 47k stars) and Performing Privacy Impact Assessment (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Camera?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/Physical-AI-Operating-System, which has 381 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/Physical-AI-Operating-System on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.