Agent skill

Phone Harness

by ShawnPana in ShawnPana/phone-harness

Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen.

MITAuto-check passedMobile

Install Phone Harness

skills CLI
$ npx skills add ShawnPana/phone-harness --skill phone-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ShawnPana/phone-harness phone-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phone-harness
GitHub stars
3.2k
Token cost
~7.7k tokens
SKILL.md length
4,426 words
Files
47
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen.

  • Works in 5 steps: Name what should change before you act -… → Do one thing, then check that one thing:… → Once a sequence is proven, batch it: a… → …
  • Tasks that involve Mobile testing and debugging
  • SKILL.md covers Which phone?, Working method (every phone), Cloud phones and Cloud iPhone, plus 3 more sections
  • Runs Python scripts from its folder; calls adb, uv and git

What it does

Phone Harness is an agent skill from ShawnPana/phone-harness. Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen.

Its SKILL.md is about 7.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 51 other files (for example `.github/FUNDING.yml`, `.github/PULL_REQUEST_TEMPLATE.md` and `.github/workflows/ci.yml`).

It sits in Mobile, covering Mobile testing and debugging. It works with Android. The repository describes itself as: let your agent control your phone. The licence is MIT.

When your agent uses it

  • Tasks that involve Mobile testing and debugging

Example prompts

  • “s phone - an iPhone through the Mac”
  • “/phone-harness”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Name what should change before you act - a title, a row, a field's
  2. Do one thing, then check that one thing: wait_for_text("Got it"),
  3. Once a sequence is proven, batch it: a whole sub-task in one script is far
  4. When a check fails, isolate: re-run that one action, look at the screen,
  5. Keep what you learn in the agent workspace's agent_helpers.py.

What it can do on your machine

Read from SKILL.md and the folder at commit 873a36e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • adb
    • uv
    • git
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, git and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phone Harness loads about 7.7k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 4,426 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ShawnPana/phone-harness at commit 873a36e, republished under its MIT licence (© ShawnPana). 4,426 words, ~7,677 tokens.

Download SKILL.mdSave it as .claude/skills/phone-harness/SKILL.md (or your agent's skills folder). This skill also uses 46 other files; get the full folder from GitHub.
name
phone-harness
description
Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen.

phone-harness

Direct control of a phone from Python scripts. The same helpers drive every kind of phone below; what differs is how each one sees and touches the screen, and that is what the per-phone sections below are for. Read Which phone, the shared Working method, and then only the section for the phone in front of you.

Which phone?

PhoneHow it is reachedEyesHandsRead
Cloud Android - rented from Phone Harness Cloud, the user's own saved phonephone-harness cloud start connects it; nothing to exportaccessibility tree (exact), Vision OCR fallback on a Macadb inputWorking method, then Cloud phones and Android
Cloud iPhone - a second iPhone hosted by Phone Harness Cloud, not the one in the user's pocket; an iOS entry in the account catalogphone-harness cloud start with its id, label, or ios when that is the only oneserver-side OCR (source: "pixels"), from any OSops over HTTPSWorking method, then Cloud iPhone
Android on the desk - USB or paired Wi-Fithe harness finds itaccessibility tree (exact), Vision OCR fallback on a Macadb inputWorking method, then Android
Personal iPhone - the user's own phone, through the Mac's iPhone Mirroring windowthe harness opens Mirroring and presses Connect; the user locks the phonescreenshots + Vision OCRHID-level CGEvents into the windowWorking method, then iPhone

"My phone" means the personal one. The personal iPhone and the cloud iPhone are two different devices with different apps and logins: the first is in the user's pocket and mirrored onto the Mac, the second lives in a data centre and is reached over HTTPS. When the user says "my phone", "my iPhone", or names an app they use, drive the personal iPhone. Use a cloud phone only when the user names it ("the cloud phone", cloud start) or the task is clearly about it. If the personal phone cannot be reached, say so and ask; never fall through to a cloud session that happens to be running, even one the user started - it is theirs, and it is a different phone.

phone-harness config shows the default platform (ios on a Mac, android elsewhere) and every other setting. phone-harness cloud shows whether a cloud phone is attached; while one is, scripts drive it regardless of the config default. PHONE_HARNESS_PLATFORM=android overrides per call. On a cloud iPhone any explicit platform selects a local backend; see Cloud iPhone.

When not to use any of them: if the task is doable on the Mac or the web - a website, an API, an app with a web equivalent - do it there and leave the phone alone. Use a phone when the task genuinely needs one: phone-only apps, things tied to a phone number or 2FA, checking how something looks on a phone.

Working method (every phone)

bash
phone-harness <<'PY'
# task: report the OS version and model name from Settings
# step: open Settings, find About, read the screen
open_app("Settings")
print([o["text"] for o in ocr()][:10])
PY
  • Invoke as phone-harness with a heredoc. Helpers are pre-imported. Start every script with # task: (the user's request in one sentence, identical across the scripts of one request) and # step: (what this script does).
  • Tell the user what you are doing as you go. A phone task is many short scripts and the user sees none of them: one line before each script (what you are about to do), one line after (what you saw). Never run two scripts in a row in silence. The # task:/# step: comments are not this - the user cannot see them.
  • Read with ocr(), not by eyeballing screenshots. Every visible string comes back with a tap-ready centre: [{text, confidence, source, x, y, w, h}]. Filter in Python before printing. tap_text("Weather") taps by label and, on failure, raises with what IS visible - read the exception before retrying. screenshot() costs more but shows what OCR cannot: icons, images, state.
  • Act, verify, adapt. There is no DOM to assert against and no return value that means "it worked", so:
    1. Name what should change before you act - a title, a row, a field's contents. Most phone failures are silent no-ops; if you cannot name the expected change you cannot tell success from one.
    2. Do one thing, then check that one thing: wait_for_text("Got it"), wait_for_app("com.android.chrome") or wait_stable(), then one ocr(). Do not wait(n) for a screen to load.
    3. Once a sequence is proven, batch it: a whole sub-task in one script is far faster than a call per turn. Keep one cheap check at the end.
    4. When a check fails, isolate: re-run that one action, look at the screen, form one guess, test it. Do not re-run the batch hoping it lands.
    5. Keep what you learn in the agent workspace's agent_helpers.py. A checkout uses agent-workspace/. A pip or uv install uses ~/.phone-harness/agent-workspace, or ~/.local/share/phone-harness/agent-workspace when ~/.phone-harness is a git checkout, so an upgrade does not wipe it. PH_AGENT_WORKSPACE overrides that path. The directory is created from the defaults only when it is missing. An existing workspace is never overwritten. To get a fresh default, delete the workspace folder; it is recreated on the next run.
  • The harness reports, you decide. Helpers return observations, never a verdict on your intent. Diff two ocr() sets, watch one label, count rows, poll until something appears - whatever matches what you asked for.
  • Navigation: home(), back(), app_switcher(), open_app("Notes"), type_text("..."), press("enter"), long_press(x, y), tap(x, y).
  • scroll says what you want to SEE; swipe says which way the finger goes. English disagrees the same way - "scroll down the page" and "swipe up for the next video" describe the same motion. scroll("down") reveals what is further down; swipe("up") is a thumb flick, the phrasing everyone uses for "next". scroll, scroll_screen, scroll_until and scroll_collect take the content direction; only swipe takes finger motion. "left"/"right" work on both. Which one moves what differs per phone - see each section.
  • Walking a list: scroll_until(done) stops when your predicate on the visible boxes is met; scroll_collect(extract, key=...) walks and de-dupes, returning {items, stop, scrolls} with stop of 'reached-end' or 'max-scrolls'. Both end on YOUR check, so an extractor that misses rows ends the walk early. scroll_screen() is the single step and returns {before, after, boxes} for you to judge. at= aims the gesture at an inner list or strip; only the scroll view under that point moves.
  • type_text needs a focused field on every phone. Tap the field, wait for the keyboard, then type; verify with ocr(), because unfocused text goes elsewhere or nowhere.
  • Consent. This is the user's real phone and real accounts. Stop and ask before anything outward-facing or hard to reverse: sending, posting, following, purchasing, deleting, changing settings. Never type a PIN, a password or a 2FA code; on a cloud phone, hand the controls over instead (phone-harness cloud open, below).
  • The harness reconnects what it can; the rest is the user's. On a personal iPhone, start the task with ensure_mirroring(): it opens iPhone Mirroring if it is closed, presses the window's Connect once, and brings the window to the front so the user can watch. Pairing, unlocking, tapping Allow, locking an iPhone that says "iPhone in Use" or "Timed Out", approving a cloud sign-in in the browser: when the harness still says the phone is not reachable, relay its message and ask - never loop-poll. Retry once after the user says it is done; that retry presses Connect again.

Cloud phones

There are two things. A phone is a row in the account catalog (GET /me -> available): a label, a platform, an id, and whether it is a temporary phone. The service keeps no default phone; a preference is the chooser's. A session is one of those phones rented right now. cloud ls prints one row per catalog phone and writes any live session on that row (* is the session this process targets). A session lines up with the row that has the same device_id or profile_id; a session with neither lines up with the temporary row. A cloud iPhone session does not echo that id: its device is the word "iPhone", and its profile (iphone-...) matches no catalog row. The catalog row's state is running while a session holds it. A session whose ids match no row counts as unmatched. When exactly one row of that platform and kind is running and exactly one unmatched session has that same platform and kind, the session is written on that row. A session that still matches nothing is listed on its own.

cloud start with no name starts this machine's cloud.phone setting (config set cloud.phone "Desk phone") when one is set; otherwise it sends no platform or kind and gets the service's rule for an unnamed start (GET /me -> unnamed): a temporary Android. A name is a catalog id or label, or ios / android when only one entry has that platform. Several matches print the catalog and stop (a terminal may ask which one). The CLI turns the chosen row into platform and kind, plus device_id for an iPhone or profile_id for the saved Android. The temporary row sends neither id.

What the helpers do next follows the session's link, not the platform name. An adb link is the Android section. A control link ({url, token}) is Cloud iPhone.

bash
phone-harness cloud                          # signed in? which session is attached? time left?
phone-harness cloud ls                       # catalog, live sessions inlined; --json for the rows
phone-harness cloud start                    # cloud.phone if set, else a temporary Android; opens the live view
phone-harness cloud start ios --for 15m      # one iPhone, fifteen minutes
phone-harness cloud start "Temporary phone" --for 90s
phone-harness --session SID <<'PY'           # this process only; or PHONE_HARNESS_SESSION
print(screen_info())
PY
phone-harness cloud stop                     # ends billing; returns at once
phone-harness cloud phone                    # the saved Android's stored state
phone-harness cloud watch                    # reopen the read-only live view (--print to share)
phone-harness cloud open                     # the dashboard: the user takes the controls
  • Pin a session when more than one can be up. --session SID or PHONE_HARNESS_SESSION selects one phone for this process, and every cloud command and helper honours it. With neither, a person at a terminal gets the attached session: the last cloud start, or cloud use SID. cloud use rewrites that shared fallback. Parallel agents should pass --session and leave cloud use alone, or they will steal the fallback from each other.
  • It bills by the minute while it is up. Start it once and keep it for the whole conversation. Stop it when the user is done, and say that you did. --for is the length: 90s, 15m / 15min, or a bare number of minutes. Default 15, cap 30 (cloud.minutes, cloud.max_minutes).
  • Stop before the deadline; a session cannot be extended. One that simply runs out keeps a saved Android's data but not its running state. The harness warns on stderr under two minutes; phone-harness cloud shows the time left. cloud stop returns at once. Stopping an adb phone that has a profile saves it; cloud start of that phone waits out a save still in progress. A control link keeps its data on the phone, so there is nothing to wait for.
  • Already running. Starting a phone this machine already holds for the same catalog choice reuses that session. If the service says it is held elsewhere (409 profile_running), the message names the SID and says to attach with --session SID. Do that; do not start a second one. 409 device_unavailable means the phone is granted but not on its host. 409 session_limit names how many sessions are in use. 503 phones_busy includes Retry-After. 403 account_blocked means rentals are blocked; a 403 about credit or balance means add credit. A 400 that names no such phone reprints the catalog.
  • Not signed in means the user runs phone-harness cloud login and approves it in a browser. Relay that; you cannot do it for them.
  • Watching versus controlling. cloud start opens the live view so the user can watch. cloud watch reopens it. For a password, a 2FA prompt, or the user taking over, run phone-harness cloud open and wait until they say they are done.
  • cloud phone is the saved Android only (GET /me/profile). cloud phone reset --yes wipes that phone's stored state: apps, logins, and data. The CLI has no iPhone reset; that is the operator's.
  • A cloud session that ends does not hand you another phone. If the session attached to this script expires, the error says session <sid> expired. If the service loses it, the error says the cloud phone failed and is unreachable. Either way the script stops. It does not drive the personal phone, and the CLI does not rent a replacement. Start the same catalog phone again with a longer duration and continue the task: phone-harness cloud start <same phone> --for <longer>. <same phone> is that phone's id or label (or ios / android when only one entry has that platform). <longer> is a duration longer than the session that just ended (90s, 15m, or a number of minutes). A --session pin whose session expired is the same error, naming that sid. A missing cloud state file still means no cloud phone was attached, and the personal phone is used as today. An unreadable state file is an error, not a cue to switch phones. cloud stop clears the failure; a new cloud start does too.
  • The adb address the CLI shows is not a secret; the unlock code is. Do not look for it, print it, or ask the user for it.

Cloud iPhone

Start it from the catalog: phone-harness cloud start ios when that is unambiguous, or the entry's id or label (cloud ls). The create body is platform: "ios", kind: "device", and device_id set to that entry's id. A ready session carries control: {url, token, expires_at} and no adb. From then on the helpers drive it over HTTPS. This machine needs no adb, no shell, and no Mac. It is a phone granted to the user's account and kept between sessions - not the personal iPhone mirrored on the Mac, which is a different device with its own apps and logins. cloud stop ends billing and returns at once; the phone keeps its own data. There is no Android-style disk save, and nothing on this phone for cloud phone to wipe.

An Android session can stay up beside it. Target one with --session SID. cloud use SID only changes the terminal fallback.

Length, billing, ls, watch, and stop are the Cloud phones commands above. An iPhone session reports billing_mode: disabled. Run scripts as plain phone-harness (with --session when you must pin one). PHONE_HARNESS_PLATFORM=... selects a local backend and skips the attached cloud phone.

The mirroring section describes a Mac window. This phone is the HTTPS path.

The client posts {"op", "kw"} to <control url>/op and reads the host's list from GET <control url>/ops (remote.py). supports(op) is that list. send of an op the host did not list raises Unsupported before the POST (this cloud phone cannot '...'). The op names and the shapes below are the client vocabulary in transport.py. Result shapes are what this repo documents or decodes.

OpArgumentsResult this client uses
screen.capturepath optional(png_path, bounds). On the wire the host returns png_b64 and bounds; the client writes the PNG. Timeout 60s.
screen.bounds{x, y, w, h, id} or None
screen.requirebounds, or raise
screen.textmin_confidence[{text, confidence, source, x, y, w, h}]. source is "pixels" or "tree".
screen.text_pixelsmin_confidencesame, forced through pixel OCR
input.tapx, yhost result, passed through
input.pressx, y, durationhost result. Long press.
input.dragx1, y1, x2, y2, duration, stepshost result
input.scrollx, y, dy, dx, stepshost result. Content-space deltas.
input.keyscombohost result
input.texts, delay, keystrokeshost result. Timeout max(60, 0.2 * len(s) + 30) seconds.
nav.home / nav.back / nav.recentshost result, or Unsupported when the host says unsupported
apps.launchname, freshapp id. Timeout 90s.
apps.currentapp id or None
apps.listinclude_system[app id]. Timeout 90s.
session.statebackend-defined string; 'ready' when usable
session.requirebounds, or raise telling the user what to do
session.refocusNone here. The client does not call the host.
session.detaila string. Local.
focus.probean opaque snapshot. Local: (True,).
focus.diffbefore, after{raised, stole_focus}. Local: both false.
tree[node] where the device has one
rawbackend-specifichost result

Trust GET /ops when the host's list is shorter than this table.

  • Coordinates are screen.bounds. tap(x, y), find_window(), send("screen.require"), and OCR centres all use that space: {x, y, w, h, id}. A point read off the PNG goes through tap_image_point() once screen_info() returns img_px. Otherwise tap an OCR centre.
  • screen_info(), image_point(), and tap_image_point() measure the PNG so a point you read off a screenshot can become a tap. If screen_info() raises ModuleNotFoundError: No module named 'Quartz', that helper imported the macOS Vision module to read the PNG's pixel size. ocr() sends screen.text to the host and still works - keep using tap_text(). If screen_info() returns img_px, those three helpers work on any OS. PR #108 ("screen_info() no longer needs macOS") is the change that reads the PNG header instead of importing Quartz; until screen_info() returns img_px off a Mac, treat the Quartz error as those three helpers only.
  • type_text() types into the focused field. Tap the field first. Text sent before the field has focus goes nowhere, so check with ocr() afterwards.
  • Act, then look. One action, then ocr() or screenshot(), then the next. Do not wait(n) for a screen to load. The client allows screen.capture 60s, apps.launch and apps.list 90s, and input.text the timeout in the table.
  • Errors the client raises. HTTP 400 with unsupported: true raises Unsupported (retrying will not help). HTTP 404 raises RuntimeError: the session is gone or the grant expired. Any other HTTP status raises RuntimeError with the host's error string. Create-time device_unavailable, profile_running, and phones_busy happen on cloud start, before any op, and are described under Cloud phones.
  • ui(), back(), and current_app() raise Unsupported when the host does not list tree, nav.back, and apps.current. home() is the Home button. shell() is Unsupported.
  • Confirm open_app. It sends apps.launch and can return with the app still not in front. Check with ocr() or a screenshot. If it missed, home() and tap the icon. tap_icon() in agent-workspace/agent_helpers.py aims 35 points above the label for the mirroring window, so measure this phone's icon yourself.
Show full SKILL.md (1,523 more words)Show less

Android

The phone is reached over adb - a USB phone if plugged in, else the paired Wi-Fi phone, else the attached cloud Android - so there is nothing to select. phone-harness android shows known phones and what is attached.

bash
PHONE_HARNESS_PLATFORM=android phone-harness <<'PY'
open_app("chrome"); wait_for_app("com.android.chrome")
tap_ui("Got it")                    # exact label from the accessibility tree
PY
  • Coordinates are device pixels; the screenshot is 1:1 with tap(x, y).
  • ocr() is the accessibility tree (source: "tree") - exact, no misreads. Prefer ui() / find_nodes() / tap_ui(): they also see elements with no visible text (icons with a content-description, fields by resource-id like tap_ui("url_bar")).
  • Some screens never give up their tree - a playing video, some Settings pages. On a Mac, ocr() then reads the screenshot with Vision instead (source: "pixels", fuzzier, still tap-ready; the first read costs ~12s while uiautomator gives up, later reads are fast). Elsewhere it raises saying so; take a screenshot() and look at it instead of retrying. ui() / tap_ui() need the real tree and keep raising on such screens.
  • adb is the native language here, and it is first-class. shell(cmd) runs adb shell cmd on whichever phone the harness chose: shell("input tap 360 640"), shell("input keyevent KEYCODE_BACK"), shell("am start -n pkg/.Activity"), shell("dumpsys notification --noredact"), shell("pm list packages -3"). The input helpers are one-line wrappers over the same commands - use whichever you think in. What the harness adds: finding and reconnecting the phone, and the screen as a short list instead of XML.
  • scroll and swipe are different gestures here. scroll moves the content and stops: no momentum, the same distance every time, so use it (and scroll_until / scroll_collect) to walk a list without skipping rows. swipe is a flick and coasts past whatever was next - right for "next video" or changing pages, wrong for reading a list.
  • open_app("TikTok") takes the name a person would say, a package id, or a fragment of one, and returns the package it launched; when nothing matches, the error lists what is installed. back(), current_app(), list_apps() exist.
  • press() takes single keys only ("enter", "back", "tab"); chords raise Unsupported. type_text types ASCII: adb cannot type emoji or accented letters.
  • Verify cheaply, then read. adb reports nothing about outcomes - a tap on empty space "succeeds". After an action: wait_for_app(...) (~0.1s a poll) or wait_for_text(...) (returns the box or None), then ui()/ocr() once. The tree costs ~2-3s a call on a slow phone and a screenshot ~0.5s, so batching a proven sub-task is worth a lot; a batch of unverified steps fails silently and tells you nothing about which one broke.
  • No focus to keep: nothing on the Mac has to be frontmost, and interruption(before, after) always reports nothing disturbed.
  • A desk phone locks itself after its screen timeout. connection_state() reports locked; taps and ocr() refuse with the same message. Ask the user to unlock - never type a PIN. screenshot() still works locked, so you can show them what you see. For a task longer than a minute, ask the user, then phone-harness android awake --bg keeps it awake without changing any phone setting; phone-harness android rest ends that. Cloud phones do not lock.
  • Connecting a desk phone is the user's job (USB debugging + Allow, or Wireless debugging + phone-harness android pair CODE); on no-device the error names the missing step - relay it, don't retry-loop.
  • A saved primary that is unreachable is a failure, not a different phone. If phone-harness android use has named a primary and that phone cannot be reached, the error says the primary is unreachable. Do not expect the harness to connect some other paired or mDNS phone instead. When no primary is configured, a USB phone that is plugged in is still used.

iPhone (iPhone Mirroring)

The Mac's iPhone Mirroring app renders the phone as a window; the harness captures that window and OCRs it with Vision for eyes, and posts HID-level events into it for hands. All coordinates are global macOS screen points.

  • Start every task with ensure_mirroring(). It opens iPhone Mirroring if it is not running, brings the window to the front so the user sees what you see, presses the interstitial's Connect once through accessibility, and waits for the live stream. A phone that is in use shows up within about three seconds and it raises then, with the window's own words, rather than sitting out a timeout. It is the cheapest first line a script can have: a phone that cannot be driven fails there, not five taps in. Call it once in the first script of a task, not before every action. After that, the default build works the phone without taking the user's focus: capture is by window id and taps and keystrokes are event records delivered straight to the app. Scrolling is the exception - macOS routes a scroll to whichever window sits under the pointer, so a scroll raises the mirroring window for the length of the gesture and hands focus straight back. Expect a brief flicker on scrolls and nothing on anything else. PHONE_HARNESS_BACKGROUND=0 forces the classic path, which focuses before every action.
  • Icons without labels: screenshot(), view the image, and use tap_image_point(x, y, image_size=...) with coordinates measured in the screenshot. Do not pass screenshot pixel coordinates to tap(): it expects global screen points. If using tap(), convert with image_point() from the current screen_info(); never estimate the window offset.
  • Use scroll for anything scrollable. On macOS 26 a vertical touch-drag is dropped, so swipe("up")/swipe("down") move nothing in a list or a feed - measured on Settings and on TikTok. Horizontal still works, so swipe("left") / swipe("right") remain the way to flip Home Screen pages and carousels, which a scroll cannot do. (Breaking change: scroll used to take finger motion too, so the old scroll("up") is today's scroll("down"). swipe is unchanged.)
  • open_app("Notes") goes through Spotlight.
  • Home-Screen labels are not tap targets. tap_text("Weather") hits the label and nothing happens; the icon is ~35 points above it. Use tap_icon("Weather") (agent helper) on the Home Screen; tap_text works for in-app buttons and list rows.
  • type_text pastes; it does not type. The keystroke path runs through iOS autocorrect, which rewrites words as they land ("Thu" becomes "thru"). Pass keystrokes=True for fields that need real key events. The typed text stays on the Mac clipboard afterwards. If a tap will not take focus, press("tab") moves between fields.
  • What the harness cannot clear is physical. Connect only works while the iPhone is locked; "iPhone in Use" and "Timed Out ... due to iPhone use" mean it was not. ensure_mirroring() presses Connect once and then raises quoting the screen (connection_state() is ready / blocked / no-window / not-running). STOP and relay it, ask the user to lock the phone, and retry once when they say so. The app reconnects by itself the moment the phone is locked, so the retry usually finds it live; if not, it presses Connect again. Do not press it yourself in a loop and do not tap the window (an interstitial is a Mac view; taps meant for the phone go nowhere). Unlocking the physical phone pauses the session ("iPhone in Use") - the same rule applies mid-task. Taps, scrolls, swipes, typing, key presses, open_app, home, and screenshot raise that same message and do not press Connect. Ask the user to lock the phone, then call ensure_mirroring() and continue from the step that raised. The Mac login prompt that sometimes appears in the window is never typed into.
  • Unfocused input is swallowed silently - for events you post yourself. The helpers are immune in the background build, but raw CGEvents and the PHONE_HARNESS_BACKGROUND=0 path need the window frontmost: activate() before posting, and re-activate if a click steals focus. The failure looks exactly like "scrolling is broken" or "the list already ended" - when a gesture changes nothing on screen, check focus before inventing another theory.
  • The window is a video stream. macOS accessibility sees nothing inside it; AppleScript click at fails silently. Only HID-level CGEvents work.
  • The window moves. Never cache coordinates across calls; ocr() and swipe() re-query bounds every time.
  • Mouse taps map to touches 1:1, but there is no multi-touch: no pinch, no two-finger gestures.
  • Raw Quartz is the escape hatch: import Quartz in your script for anything the helpers don't cover - but raw CGEvents don't ride the helpers' delivery path, and where they land is its own question per event type. Check what actually happened on screen rather than assuming the event arrived.

For task-specific edits, use the agent workspace's agent_helpers.py: agent-workspace/ in a checkout, ~/.phone-harness/agent-workspace for a pip or uv install (or ~/.local/share/phone-harness/agent-workspace when ~/.phone-harness is a git checkout). PH_AGENT_WORKSPACE overrides that. Installs seed that directory from the packaged defaults when it does not exist yet; an existing workspace is left as it is. To get a fresh default, delete the workspace folder; it is recreated on the next run. For setup or permission problems, read install.md.

Updates

phone-harness update brings this install up to date. A uv tool runs uv tool upgrade phone-harness. A git checkout runs git pull --ff-only in that checkout. Anything else prints pip install -U phone-harness and does not run it. The CLI never updates itself. Once a day, on a terminal, it may print one stderr line when PyPI has a newer version:

phone-harness X.Y.Z is available (you have A.B.C): run phone-harness update

Relay that line to the user. PHONE_HARNESS_NO_UPDATE_CHECK=1 turns the notice off. A network problem never fails the command.

© ShawnPana, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 46 other files in the repository root of ShawnPana/phone-harness.

  • SKILL.md
  • .github/FUNDING.yml
  • .github/PULL_REQUEST_TEMPLATE.md
  • .github/workflows/ci.yml
  • .github/workflows/publish.yml
  • .gitignore
  • LICENSE
  • README.md
  • agent-workspace/agent_helpers.py
  • install.md
  • integrations/hermes/README.md
  • integrations/hermes/__init__.py
  • integrations/hermes/plugin.yaml
  • onboarding.md
  • phone-harness
  • pyproject.toml
  • … and 31 more

Open the folder on GitHubat commit 873a36e

Compare with similar skills

Phone Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phone Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phone Harness this skillShawnPana/phone-harness3.2k—~7.7kAutomated safety check: PassMIT
Mobile QAtloncorp/tlon-apps107—~2.4kAutomated safety check: PassMIT
Androidyang1ming/android-harness176—~259Automated safety check: PassMIT
Debug Receiverstimusus/Shuttle2229—~1.9kAutomated safety check: PassApache-2.0
Debug Bridgegetknit/knit133—~2.3kAutomated safety check: PassGPL-3.0
Capturing Screenshots And Screenrecordskydoves/android-testing-skills334—~3.7kAutomated safety check: PassApache-2.0

Similar skills

  • Mobile QA

    tloncorp/tlon-apps

    Run a mobile QA checklist on a physical Android device over adb for tlon-apps, then triage what fails into fixes.

    107 GitHub stars~2.4k tokensUpdated yesterday
    MobileAuto-check passed
  • Android

    yang1ming/android-harness

    Direct Android device control through ADB. An agent skill from yang1ming/android-harness.

    176 GitHub stars~259 tokensUpdated 2 mo ago
    MobileAuto-check passed
  • Debug Receivers

    timusus/Shuttle2

    Drive the S2 debug build's playback and queue over ADB broadcasts — play the whole library, play/pause, skip, seek, remove a queue item, toggle shuffle/repeat, dump playback state as JSON, reimport…

    229 GitHub stars~1.9k tokensUpdated today
    MobileAuto-check passed
  • Debug Bridge

    getknit/knit

    Drive and verify Knit on a device or emulator through the headless debug bridge (am broadcast to app.getknit.knit.debug.<ACTION, replies as JSON) — send a message on one phone and confirm it landed…

    133 GitHub stars~2.3k tokensUpdated yesterday
    MobileAuto-check passed
  • Capturing Screenshots And Screenrecord

    skydoves/android-testing-skills

    A skill your agent uses to capture visual artefacts from a device for test failures, golden image generation, QA repro, and demo videos.

    334 GitHub stars~3.7k tokensUpdated 4 mo ago
    MobileAuto-check passed
  • Android Emulator Skill

    Moustachauve/WLED-Android

    Production-ready scripts for Android app testing, building, and automation.

    169 GitHub starsUsed in 1 repo~911 tokens
    MobileAuto-check passed

Works with

Categories

Questions about Phone Harness

What does Phone Harness do?

Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen. Phone Harness is an agent skill from ShawnPana/phone-harness. Control the user's phone - an iPhone through the Mac's iPhone Mirroring window, an Android over adb, a rented cloud Android, or a cloud iPhone over HTTPS: open apps, tap, type, swipe, read the screen.

When should I use Phone Harness?

Phone Harness fits situations like: tasks that involve Mobile testing and debugging.

How do I install Phone Harness in Claude Code?

Run `npx skills add ShawnPana/phone-harness --skill phone-harness -a claude-code`. Or copy the skill folder (the ShawnPana/phone-harness repository) into .claude/skills/phone-harness in your project. Claude Code loads it when a task matches its description.

How do I install Phone Harness in Codex?

Run `npx skills add ShawnPana/phone-harness --skill phone-harness -a codex`. Or copy the skill folder (the ShawnPana/phone-harness repository) into .agents/skills/phone-harness in your project. Codex loads it when a task matches its description.

Can I use Phone Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ShawnPana/phone-harness --skill phone-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phone-harness, .gemini/skills/phone-harness, .github/skills/phone-harness and .opencode/skills/phone-harness in your project.

What does Phone Harness need to run?

Going by SKILL.md and its folder, Phone Harness needs Python for the scripts in its folder and the command-line tools its instructions call (adb, uv, git and pip). Our summary lists: Python 3.

Does Phone Harness access the network?

SKILL.md contains no URLs. Its commands use uv, git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Phone Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phone Harness use?

Phone Harness is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phone Harness use?

About 7.7k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phone Harness?

Skills that share tags, products or a category with Phone Harness: Mobile QA (tloncorp/tlon-apps, 107 stars), Android (yang1ming/android-harness, 176 stars), Debug Receivers (timusus/Shuttle2, 229 stars) and Debug Bridge (getknit/knit, 133 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phone Harness?

ShawnPana (a GitHub user) maintains it in ShawnPana/phone-harness, which has 3,223 GitHub stars. The repository was last updated on October 11, 2026.

Source: ShawnPana/phone-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.