Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face).

Apache-2.0Auto-check passed

Install Servo Tracking

skills CLI
$ npx skills add autonomous-ai/Physical-AI-Operating-System --skill servo-tracking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/Physical-AI-Operating-System servo-tracking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/Physical-AI-Operating-System.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/servo-tracking .claude/skills/servo-tracking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
servo-tracking
GitHub stars
381
Token cost
~1.3k tokens
SKILL.md length
580 words
Files
2
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face).

  • Works in 4 steps: User names an object to track. → If the user also hints at a direction… → Prefix reply with… → …
  • Fixed locations (desk
  • SKILL.md covers Quick Start, Workflow, Examples and How to Control Tracking, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Servo Tracking is an agent skill from autonomous-ai/Physical-AI-Operating-System. Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face). NEVER use for furniture or fixed locations (desk, table, wall, floor, ceiling, door, window, workspace, room) — those are directions; use servo-control /servo/aim. NEVER use for direction words (left, right, up, down, center). If the user combines a direction AND an object ("look at the desk and follow the cup", "point at the table and track the…

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`).

The repository describes itself as: The open-source operating system for physical AI. The licence is Apache-2.0.

When your agent uses it

  • Fixed locations (desk
  • Room) — those are directions
  • Use servo-control /servo/aim
  • Direction words (left

Example prompts

  • “look at the desk and follow the cup”
  • “point at the table and track the phone”
  • “look at”
  • “/servo-tracking”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. User names an object to track.
  2. If the user also hints at a direction ("look at desk and follow cup", "point at the table and watch the phone"), fire /servo/aim FIRST so…
  3. Prefix reply with [HW:/servo/track:{"target":["",""]}] — the device detects and follows.
  4. To stop, prefix with [HW:/servo/track/stop:{}].

What it can do on your machine

Read from SKILL.md and the folder at commit f1b9ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Servo Tracking loads about 1.3k tokens when it runs. Until then it costs about 182 tokens; SKILL.md has 580 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~182
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/Physical-AI-Operating-System at commit f1b9ebe, republished under its Apache-2.0 licence (© autonomous-ai). 580 words, ~1,310 tokens.

Download SKILL.mdSave it as .claude/skills/servo-tracking/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
servo-tracking
description
Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face). NEVER use for furniture or fixed locations (desk, table, wall, floor, ceiling, door, window, workspace, room) — those are directions; use servo-control /servo/aim. NEVER use for direction words (left, right, up, down, center). If the user combines a direction AND an object ("look at the desk and follow the cup", "point at the table and track the phone"), fire TWO markers in order — servo-control /servo/aim first, then /servo/track here. The verb "look at" is ambiguous — if what follows is furniture/location, route to servo-control instead.

Vision Tracking

Quick Start

Tracks and follows any object by name. YOLOWorld detects the object in the camera frame, TrackerVit follows it in real-time with servo movement.

Workflow

  1. User names an object to track.
  2. If the user also hints at a direction ("look at desk and follow cup", "point at the table and watch the phone"), fire /servo/aim FIRST so the camera is pointing at the right region before YOLO detection runs. Then fire /servo/track. The aim should complete in ~2s before the track call takes over.
  3. Prefix reply with [HW:/servo/track:{"target":["<label1>","<label2>"]}] — the device detects and follows.
    • target accepts a list of candidate labels. Pass 2-4 synonyms/variants to maximise the chance YOLOWorld finds the object on the first try.
    • A single string also works ({"target":"cup"}) for backward compatibility, but the list form is preferred when the object could reasonably have multiple names.
  4. To stop, prefix with [HW:/servo/track/stop:{}].

Examples

Input: "Look at the desk and follow the cup" Output: [HW:/servo/aim:{"direction":"desk"}][HW:/servo/track:{"target":["cup","mug","coffee cup"]}] Looking at the desk and locking onto the cup.

Input: "Point at the table and track my phone" Output: [HW:/servo/aim:{"direction":"desk"}][HW:/servo/track:{"target":["phone","smartphone","mobile phone"]}] Aimed at the table, tracking your phone.

Input: "Follow the cup" Output: [HW:/servo/track:{"target":["cup","mug","coffee cup"]}] OK, following the cup!

Input: "Look at the bottle" Output: [HW:/servo/track:{"target":["bottle","water bottle","plastic bottle"]}] Watching the bottle.

Input: "Track that person" Output: [HW:/servo/track:{"target":["person","man","woman"]}] Following them now.

Input: "Watch my phone" Output: [HW:/servo/track:{"target":["phone","smartphone","mobile phone"]}] Got it, tracking your phone.

Input: "Follow the teddy bear" Output: [HW:/servo/track:{"target":["teddy bear","stuffed animal","plush toy"]}] Tracking the teddy bear!

Input: "Stop following" / "Stop tracking" Output: [HW:/servo/track/stop:{}] Stopped tracking.

Input: "What can you track?" Output: I can track most common objects — cups, bottles, phones, laptops, books, people, bags, and more. Just tell me what to follow!

How to Control Tracking

No exec/curl needed. Inline markers at start of reply:

[HW:/servo/track:{"target":["cup","mug","coffee cup"]}] Following the cup.
[HW:/servo/track:{"target":["person"]}] Tracking you now.
[HW:/servo/track/stop:{}] Stopped tracking.
Target names

Prefer a list of 2-4 candidate labels in English. YOLOWorld evaluates all candidates and picks the highest-confidence detection across the set — synonyms increase the chance of a successful first-try detection when the user's wording doesn't exactly match COCO/training vocabulary.

Common objects: person, cup, bottle, glass, phone, laptop, keyboard, mouse, book, pen, notebook, bag, chair, monitor, remote control, plate, bowl, plant, vase, clock, lamp, speaker, headphones, watch, glasses, hat, shoe, toy, ball, teddy bear.

Any label works (open-vocabulary detection). Don't include too many unrelated items in the list — that risks matching a nearby but wrong object.

Show full SKILL.md (185 more words)Show less
How it works internally
  1. Camera captures current frame
  2. YOLOWorld API evaluates all candidate labels (~1-2s)
  3. Highest-confidence detection becomes the initial bbox
  4. TrackerVit locks on and follows in a move-then-freeze loop (~7 FPS)
  5. Servo base_yaw + base_pitch nudges to keep object centered
  6. Auto-stops when object is lost, out of range, or after 5 minutes

Error Handling

  • If the object is not found: "I can't see a {target} right now. Try pointing me toward it, or try a different name."
  • If tracking stops unexpectedly: "I lost the {target}. Want me to try again?"
  • If servo is not available: "Servo is not available right now."

Rules

  • Only one object can be tracked at a time. Starting a new track stops the previous one.
  • Tracking auto-stops when the object leaves the frame, gets occluded, or after 5 minutes.
  • When tracking stops, the servo holds its last position (no snap back to idle).
  • Do NOT use this for emotional reactions — use the Emotion skill instead.
  • Prefer target as a list of 2-4 English synonyms for better first-try detection. Avoid packing unrelated labels into the list.

© autonomous-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/servo-tracking of autonomous-ai/Physical-AI-Operating-System.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit f1b9ebe

Compare with similar skills

Servo Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Servo Tracking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Servo Tracking this skillautonomous-ai/Physical-AI-Operating-System381—~1.3kAutomated safety check: PassApache-2.0
Fal Visionnexu-io/open-design100k—~295Automated safety check: PassApache-2.0
Senior Computer Visiondavila7/claude-code-templates32k3 repos~1.4kAutomated safety check: PassMIT
Horizon Trackruvnet/ruflo74k—~744Automated safety check: NotesMIT
Senior Computer Visionalirezarezvani/claude-skills28k2 repos~3.2kAutomated safety check: PassMIT
Vision Sftwshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • Fal Vision

    nexu-io/open-design

    Analyze images — segment objects, detect, run OCR, describe, and answer visual questions via fal.ai vision models.

    100k GitHub stars~295 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    davila7/claude-code-templates

    World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems.

    32k GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Horizon Track

    ruvnet/ruflo

    Track long-horizon objectives across multiple sessions with milestone checkpoints, progress persistence, and drift detection

    74k GitHub stars~744 tokensUpdated today
    Product & Project ManagementAuto-check: notes
  • Senior Computer Vision

    alirezarezvani/claude-skills

    Computer vision engineering skill for object detection, image segmentation, and visual AI systems.

    28k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Vision Sft

    wshobson/agents

    Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from autonomous-ai/Physical-AI-Operating-System

All 28 skills in this repo
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Claude Code Buddy

    autonomous-ai/Physical-AI-Operating-System

    Push Claude Code activity to the user's device (e.g. An agent skill from autonomous-ai/Physical-AI-Operating-System.

    381 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Connectors

    autonomous-ai/Physical-AI-Operating-System

    Discover and use linked third-party services (Gmail, Google Calendar, Google Drive, Notion, Figma, Asana, Linear, GitHub, Ahrefs, Facebook Fan Page and others).

    381 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Audio

    autonomous-ai/Physical-AI-Operating-System

    Low-level speaker and microphone hardware control — adjust volume, play test tones, record raw audio.

    381 GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Servo Tracking

What does Servo Tracking do?

Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face). Servo Tracking is an agent skill from autonomous-ai/Physical-AI-Operating-System. Use ONLY to follow/track/watch a movable physical OBJECT vision can recognize (cup, bottle, phone, hand, person, pen, book, remote, toy, keys, pet, a specific face).

When should I use Servo Tracking?

Servo Tracking fits situations like: fixed locations (desk; room) — those are directions; use servo-control /servo/aim; direction words (left.

How do I install Servo Tracking in Claude Code?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill servo-tracking -a claude-code`. Or copy the skill folder (skills/servo-tracking in autonomous-ai/Physical-AI-Operating-System) into .claude/skills/servo-tracking in your project. Claude Code loads it when a task matches its description.

How do I install Servo Tracking in Codex?

Run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill servo-tracking -a codex`. Or copy the skill folder (skills/servo-tracking in autonomous-ai/Physical-AI-Operating-System) into .agents/skills/servo-tracking in your project. Codex loads it when a task matches its description.

Can I use Servo Tracking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/Physical-AI-Operating-System --skill servo-tracking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/servo-tracking, .gemini/skills/servo-tracking, .github/skills/servo-tracking and .opencode/skills/servo-tracking in your project.

What does Servo Tracking need to run?

SKILL.md names no scripts, command-line tools or credentials: Servo Tracking is instructions for the agent only.

Does Servo Tracking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Servo Tracking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Servo Tracking use?

Servo Tracking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Servo Tracking use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Servo Tracking?

Skills that share tags, products or a category with Servo Tracking: Fal Vision (nexu-io/open-design, 100k stars), Senior Computer Vision (davila7/claude-code-templates, 32k stars), Horizon Track (ruvnet/ruflo, 74k stars) and Senior Computer Vision (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Servo Tracking?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/Physical-AI-Operating-System, which has 381 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/Physical-AI-Operating-System on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.