Agent skill

Atl Browser

by LeoYeAI in LeoYeAI/openclaw-master-skills

Mobile browser and native app automation via ATL (iOS Simulator).

MITAuto-check passedProductivity & Automation

Install Atl Browser

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill atl-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills atl-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/atl-mobile .claude/skills/atl-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
atl-browser
GitHub stars
2.2k
Token cost
~6.5k tokens
SKILL.md length
1,119 words
Files
3 (incl. scripts)
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Mobile browser and native app automation via ATL (iOS Simulator).

  • Works in 7 steps: Start Simulator → Build & Install AtlBrowser → Verify Server → …
  • Tasks that involve App automation through connectors
  • SKILL.md covers 🔀 Two Servers: Browser & Native, 📱 Native App Automation (Port…, 🔄 Native App Workflow Example and 💡 Core Insight: Vision-Free…, plus 5 more sections
  • Runs Shell scripts from its folder; calls curl, xcrun and jq; reaches store.com and apple.com

What it does

Atl Browser is an agent skill from LeoYeAI/openclaw-master-skills. Mobile browser and native app automation via ATL (iOS Simulator). Navigate, click, screenshot, and automate web and native app tasks on iPhone/iPad simulators.

Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `_meta.json` and `scripts/setup.sh`).

It sits in Productivity & Automation, covering App automation through connectors. It works with iOS. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Tasks that involve App automation through connectors

Example prompts

  • “/atl-browser”

Requirements

  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Start Simulator
  2. Build & Install AtlBrowser
  3. Verify Server
  4. Clean UI Before Acting
  5. Verify State After Actions
  6. Use Viewport Coordinates for Taps
  7. Screenshot is Your Debugging Superpower

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • curl
    • xcrun
    • jq
    • xcodebuild

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • store.com
    • apple.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Atl Browser loads about 6.5k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 1,119 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,119 words, ~6,532 tokens.

Download SKILL.mdSave it as .claude/skills/atl-browser/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
atl-browser
description
Mobile browser and native app automation via ATL (iOS Simulator). Navigate, click, screenshot, and automate web and native app tasks on iPhone/iPad simulators.

ATL — Agent Touch Layer

The automation layer between AI agents and iOS

ATL provides HTTP-based automation for iOS Simulator — both browser (mobile Safari) and native apps. Think Playwright, but for mobile.

🔀 Two Servers: Browser & Native

ATL uses two separate servers for browser and native app automation:

ServerPortUse CaseKey Commands
Browser9222Web automation in mobile Safarigoto, markElements, clickMark, evaluate
Native9223iOS app automation (Settings, Contacts, any app)openApp, snapshot, tapRef, find
┌─────────────────────────────────────────────────────────────┐
│  BROWSER SERVER (9222)     │     NATIVE SERVER (9223)      │
│  (mobile Safari/WebView)   │     (iOS apps via XCTest)     │
│                            │                                │
│  markElements + clickMark  │     snapshot + tapRef         │
│  CSS selectors             │     accessibility tree        │
│  DOM evaluation            │     element references        │
│  tap, swipe, screenshot    │     tap, swipe, screenshot    │
└─────────────────────────────────────────────────────────────┘

Why two ports? Native app automation requires XCTest APIs (XCUIApplication, XCUIElement) which are only available in UI Test bundles. The native server runs as a UI Test that exposes an HTTP API.

Starting the Servers
bash
# Browser server (starts automatically with AtlBrowser app)
xcrun simctl launch booted com.atl.browser
curl http://localhost:9222/ping  # → {"status":"ok"}

# Native server (run as UI Test)
cd ~/Atl/core/AtlBrowser
xcodebuild test -workspace AtlBrowser.xcworkspace \
  -scheme AtlBrowser \
  -destination 'id=<SIMULATOR_UDID>' \
  -only-testing:AtlBrowserUITests/NativeServer/testNativeServer &
  
# Wait for it to start, then:
curl http://localhost:9223/ping  # → {"status":"ok","mode":"native"}
Quick Port Reference
TaskPortExample
Browse websites9222curl localhost:9222/command -d '{"method":"goto",...}'
Open native app9223curl localhost:9223/command -d '{"method":"openApp",...}'
Screenshot (browser)9222curl localhost:9222/command -d '{"method":"screenshot"}'
Screenshot (native)9223curl localhost:9223/command -d '{"method":"screenshot"}'

📱 Native App Automation (Port 9223)

Native automation uses port 9223 and automates any iOS app using the accessibility tree — no DOM, no JavaScript, just direct element interaction.

Opening & Closing Apps
bash
# Open an app by bundle ID
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"openApp","params":{"bundleId":"com.apple.Preferences"}}'
# → {"success":true,"result":{"bundleId":"com.apple.Preferences","mode":"native","state":"running"}}

# Check current app state
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"appState"}'
# → {"success":true,"result":{"mode":"native","bundleId":"com.apple.Preferences","state":"running"}}

# Close current app
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"closeApp"}'
# → {"success":true,"result":{"closed":true}}
Common Bundle IDs
AppBundle ID
Settingscom.apple.Preferences
Contactscom.apple.MobileAddressBook
Calculatorcom.apple.calculator
Calendarcom.apple.mobilecal
Photoscom.apple.mobileslideshow
Notescom.apple.mobilenotes
Reminderscom.apple.reminders
Clockcom.apple.mobiletimer
Mapscom.apple.Maps
Safaricom.apple.mobilesafari
The snapshot Command

snapshot returns the accessibility tree — all visible elements with their properties and tap-able references.

bash
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"snapshot","params":{"interactiveOnly":true}}' | jq '.result'

Example output:

json
{
  "count": 12,
  "elements": [
    {
      "ref": "e0",
      "type": "cell",
      "label": "Wi-Fi",
      "value": "MyNetwork",
      "identifier": "",
      "x": 0,
      "y": 142,
      "width": 393,
      "height": 44,
      "isHittable": true,
      "isEnabled": true
    },
    {
      "ref": "e1",
      "type": "cell",
      "label": "Bluetooth",
      "value": "On",
      "identifier": "",
      "x": 0,
      "y": 186,
      "width": 393,
      "height": 44,
      "isHittable": true,
      "isEnabled": true
    },
    {
      "ref": "e2",
      "type": "button",
      "label": "Back",
      "value": null,
      "identifier": "Back",
      "x": 0,
      "y": 44,
      "width": 80,
      "height": 44,
      "isHittable": true,
      "isEnabled": true
    }
  ]
}

Parameters:

  • interactiveOnly (bool, default: false) — Only return hittable elements
  • maxDepth (int, optional) — Limit tree traversal depth
The tapRef Command

Tap an element by its reference from the last snapshot:

bash
# Take snapshot first
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"snapshot","params":{"interactiveOnly":true}}'

# Tap element e0 (Wi-Fi cell from example above)
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"tapRef","params":{"ref":"e0"}}'
# → {"success":true}
The find Command

Find and interact with elements by text — no need to parse snapshot manually:

bash
# Find and tap "Wi-Fi"
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"find","params":{"text":"Wi-Fi","action":"tap"}}'
# → {"success":true,"result":{"found":true,"ref":"e0"}}

# Check if an element exists
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"find","params":{"text":"Bluetooth","action":"exists"}}'
# → {"success":true,"result":{"found":true,"ref":"e1"}}

# Find and fill a text field
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"find","params":{"text":"First name","action":"fill","value":"John"}}'

# Get element info without interacting
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"find","params":{"text":"Cancel","action":"get"}}'
# → {"success":true,"result":{"found":true,"ref":"e5","element":{...}}}

Parameters:

  • text (string) — Text to search for (matches label, value, or identifier)
  • action (string) — One of: tap, fill, exists, get
  • value (string, optional) — Text to fill (required for action:"fill")
  • by (string, optional) — Narrow search: label, value, identifier, type, or any (default)

🔄 Native App Workflow Example

Here's a complete flow: open Settings, navigate to Wi-Fi, take a screenshot:

bash
# 1. Open Settings app
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"openApp","params":{"bundleId":"com.apple.Preferences"}}'

# 2. Wait for app to launch
sleep 1

# 3. Take snapshot to see available elements
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"snapshot","params":{"interactiveOnly":true}}' | jq '.result.elements[:5]'

# 4. Find and tap Wi-Fi
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"find","params":{"text":"Wi-Fi","action":"tap"}}'

# 5. Wait for navigation
sleep 0.5

# 6. Take screenshot of Wi-Fi settings
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"screenshot"}' | jq -r '.result.data' | base64 -d > /tmp/wifi-settings.png

# 7. Navigate back (swipe right from left edge)
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"swipe","params":{"direction":"right"}}'

# 8. Close the app
curl -s -X POST http://localhost:9223/command \
  -d '{"method":"closeApp"}'
Helper Script Version
bash
source ~/.openclaw/skills/atl-browser/scripts/atl-helper.sh

atl_openapp "com.apple.Preferences"
sleep 1
atl_find "Wi-Fi" tap
sleep 0.5
atl_screenshot /tmp/wifi-settings.png
atl_swipe right
atl_closeapp

💡 Core Insight: Vision-Free Automation

ATL's killer feature is spatial understanding without vision models:

┌─────────────────────────────────────────────────────────────┐
│  markElements + captureForVision = COMPLETE PAGE KNOWLEDGE  │
└─────────────────────────────────────────────────────────────┘

1. markElements  → Numbers every interactive element [1] [2] [3]
2. captureForVision → PDF with text layer + element coordinates
3. tap x=234 y=567 → Pixel-perfect touch at exact position

Why this matters:

  • No vision API calls — zero token cost for "seeing" the page
  • Faster — no round-trip to GPT-4V/Claude Vision
  • Deterministic — same page = same coordinates, every time
  • Reliable — pixel-perfect coordinates vs. vision interpretation
The Vision-Free Workflow
bash
# 1. Mark elements (adds numbered labels + stores coordinates)
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"1","method":"markElements","params":{}}'

# 2. Capture PDF with text layer (machine-readable, has coordinates)
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"2","method":"captureForVision","params":{"savePath":"/tmp","name":"page"}}' \
  | jq -r '.result.path'
# → /tmp/page.pdf (text-selectable, contains element positions)

# 3. Get specific element's position by mark label
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"3","method":"getMarkInfo","params":{"label":5}}' | jq '.result'
# → {"label":5, "tag":"button", "text":"Add to Cart", "x":187, "y":432, "width":120, "height":44}

# 4. Tap at exact coordinates
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"4","method":"tap","params":{"x":187,"y":432}}'

The marks tell you WHERE everything is. The PDF tells you WHAT everything says. Together = full page understanding.

🎯 The Escalation Ladder

When automation gets stuck, escalate through these levels:

┌─────────────────────────────────────────────────────────────┐
│  Level 1: COORDINATES (fast, cheap, no API calls)          │
│  markElements → getMarkInfo → tap x,y                      │
│                                                             │
│  ↓ If stuck after 2-3 tries...                             │
│                                                             │
│  Level 2: VISION FALLBACK (screenshot to understand state) │
│  screenshot → analyze UI → identify blockers (modals, etc) │
│                                                             │
│  ↓ If still stuck...                                       │
│                                                             │
│  Level 3: JS INJECTION (direct DOM manipulation)           │
│  evaluate → dispatchEvent → force interactions             │
└─────────────────────────────────────────────────────────────┘
When to Escalate
SymptomLikely CauseAction
Tap succeeds but nothing changesModal/overlay openedScreenshot → find new button
Cart count doesn't updateSite needs login or has bot detectionTry JS click with events
Element not found after scrollMarks are page-relative, not viewportUse getBoundingClientRect via evaluate
Same error 3+ timesUI state changed unexpectedlyScreenshot to see actual state
Real-World Pattern: E-commerce Checkout
bash
# 1. Search and find product
atl_goto "https://store.com/search?q=headphones"
atl_mark

# 2. First, dismiss any modals/banners (ALWAYS DO THIS)
# Look for: close, dismiss, continue, accept, no thanks, got it
CLOSE=$(atl_find "close")
[ -n "$CLOSE" ] && atl_click $CLOSE

# 3. Find and click Add to Cart
ATC=$(atl_find "Add to cart")
atl_click $ATC

# 4. Wait, then CHECK if it worked
sleep 2
atl_screenshot /tmp/after-click.png

# 5. If cart didn't update, LOOK at the screenshot
# Maybe a "Choose options" modal opened - find the NEW Add to Cart button
# This is the vision fallback - you need to SEE what happened
Key Insight: Modals Change Everything

When you click "Add to cart" on sites like Target, Amazon, etc., they often:

  1. Open a "Choose options" modal (size, color, quantity)
  2. Show an upsell (protection plans, accessories)
  3. Display a confirmation with "View cart" or "Continue shopping"

Your original tap WORKED — you just can't see the result without a screenshot.

🚀 Quick Start (30 seconds)

bash
# 1. Setup (boots sim, installs ATL)
~/.openclaw/skills/atl-browser/scripts/setup.sh

# 2. Navigate somewhere
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"1","method":"goto","params":{"url":"https://example.com"}}'

# 3. Mark elements (shows [1], [2], [3] labels)
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"2","method":"markElements","params":{}}'

# 4. Take screenshot
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"3","method":"screenshot","params":{}}' | jq -r '.result.data' | base64 -d > /tmp/page.png

# 5. Click element [1]
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"4","method":"clickMark","params":{"label":1}}'

Or use the helper functions:

bash
source ~/.openclaw/skills/atl-browser/scripts/atl-helper.sh
atl_goto "https://example.com"
atl_mark
atl_screenshot /tmp/page.png
atl_click 1

Quick Reference

Base URL: http://localhost:9222

Common Commands
bash
# Check if ATL is running
curl -s http://localhost:9222/ping

# Navigate to URL
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"1","method":"goto","params":{"url":"https://example.com"}}'

# Wait for page ready
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"2","method":"waitForReady","params":{"timeout":10}}'

# Take screenshot (returns base64 PNG)
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"3","method":"screenshot","params":{}}' | jq -r '.result.data' | base64 -d > screenshot.png

# Mark interactive elements (shows numbered labels)
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"4","method":"markElements","params":{}}'

# Click by mark label
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"5","method":"clickMark","params":{"label":3}}'

# Scroll page
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"6","method":"evaluate","params":{"script":"window.scrollBy(0, 500)"}}'

# Type text
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"7","method":"type","params":{"text":"Hello world"}}'

# Click by CSS selector
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"8","method":"click","params":{"selector":"button.submit"}}'

Setup (First Time)

1. Start Simulator
bash
# Boot iPhone 17 simulator (or another device)
xcrun simctl boot "iPhone 17"

# Open Simulator app
open -a Simulator
2. Build & Install AtlBrowser
bash
cd ~/Atl/core/AtlBrowser

# Build for simulator (RECOMMENDED: target by UDID)
# Why: name-based destinations can cause Xcode to pick an older iOS runtime (15/16)
# and fail if AtlBrowser has an iOS 17+ deployment target.
#
# 1) Find a suitable simulator UDID (iOS 17+):
#   xcrun simctl list devices available
#
# 2) Build targeting that UDID:
xcodebuild -workspace AtlBrowser.xcworkspace \
  -scheme AtlBrowser \
  -destination 'id=<SIM_UDID>' \
  -derivedDataPath /tmp/atl-dd \
  build

# Install to a specific simulator (preferred)
xcrun simctl install <SIM_UDID> \
  /tmp/atl-dd/Build/Products/Debug-iphonesimulator/AtlBrowser.app

# Launch the app
xcrun simctl launch <SIM_UDID> com.atl.browser
3. Verify Server
bash
curl -s http://localhost:9222/ping
# Should return: {"status":"ok"}

All Available Methods

App Control (Native Mode)
MethodParamsModeDescription
openApp{bundleId}Any→NativeOpen app, switch to native mode
closeApp-NativeClose current app, return to browser mode
appState-AnyGet current mode and bundleId
openBrowser-Native→BrowserSwitch back to browser mode
Native Accessibility
MethodParamsModeDescription
snapshot{interactiveOnly?, maxDepth?}NativeGet accessibility tree
tapRef{ref}NativeTap element by ref (e.g., "e0")
find{text, action, value?, by?}NativeFind element and interact
fillRef{ref, text}NativeTap element and type text
focusRef{ref}NativeFocus element without typing
Navigation (Browser)
MethodParamsModeDescription
goto{url}BrowserNavigate to URL
reload-BrowserReload page
goBack-BrowserGo back
goForward-BrowserGo forward
getURL-BrowserGet current URL
getTitle-BrowserGet page title
Show full SKILL.md (438 more words)Show less
Interactions (Browser)
MethodParamsModeDescription
click{selector}BrowserClick element
doubleClick{selector}BrowserDouble-click
type{text}BothType text
fill{selector, value}BrowserFill input field
press{key}BothPress key
hover{selector}BrowserHover over element
scrollIntoView{selector}BrowserScroll to element
Mark System (Browser)
MethodParamsModeDescription
markElements-BrowserMark visible interactive elements
markAll-BrowserMark ALL interactive elements
unmarkElements-BrowserRemove marks
clickMark{label}BrowserClick by label number
getMarkInfo{label}BrowserGet element info by label
Screenshots & Capture
MethodParamsModeDescription
screenshot{fullPage?, selector?}BothTake screenshot
captureForVision{savePath?, name?}BrowserFull page PDF
captureJPEG{quality?, fullPage?}BothJPEG capture
captureLight-BrowserText + interactives only
Waiting (Browser)
MethodParamsModeDescription
waitForSelector{selector, timeout?}BrowserWait for element
waitForNavigation-BrowserWait for navigation
waitForReady{timeout?, stabilityMs?}BrowserWait for page ready
waitForAny{selectors, timeout?}BrowserWait for any selector
JavaScript (Browser)
MethodParamsModeDescription
evaluate{script}BrowserRun JavaScript
querySelector{selector}BrowserFind element
querySelectorAll{selector}BrowserFind all elements
getDOMSnapshot-BrowserGet page HTML
Cookies (Browser)
MethodParamsModeDescription
getCookies-BrowserGet all cookies
setCookies{cookies}BrowserSet cookies
deleteCookies-BrowserDelete all cookies
Touch Gestures (Both Modes)
MethodParamsModeDescription
tap{x, y}BothTap at coordinates
longPress{x, y, duration?}BothLong press (default 0.5s)
swipe{direction}BothSwipe up/down/left/right
swipe{fromX, fromY, toX, toY}BothSwipe between points
pinch{scale, duration?}BothPinch zoom (scale > 1 = zoom in)
Swipe Examples
bash
# Swipe up (scroll down)
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"1","method":"swipe","params":{"direction":"up"}}'

# Swipe left (next page in carousel)
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"2","method":"swipe","params":{"direction":"left","distance":400}}'

# Custom swipe path
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"3","method":"swipe","params":{"fromX":200,"fromY":600,"toX":200,"toY":200}}'

# Long press for context menu
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"4","method":"longPress","params":{"x":150,"y":300,"duration":1.0}}'

# Pinch to zoom in
curl -s -X POST http://localhost:9222/command \
  -d '{"id":"5","method":"pinch","params":{"scale":2.0}}'

Typical Workflow

bash
# 1. Navigate to site
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"1","method":"goto","params":{"url":"https://www.apple.com/shop"}}'

# 2. Wait for page to load
sleep 2
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"2","method":"waitForReady","params":{"timeout":10}}'

# 3. Mark elements to see what's clickable
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"3","method":"markElements","params":{}}'

# 4. Take screenshot to see the marks
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"4","method":"screenshot","params":{}}' | jq -r '.result.data' | base64 -d > /tmp/page.png

# 5. Click a marked element (e.g., label 14)
curl -s -X POST http://localhost:9222/command \
  -H "Content-Type: application/json" \
  -d '{"id":"5","method":"clickMark","params":{"label":14}}'

# 6. Repeat as needed

Troubleshooting

Navigation not working (goto returns success but page doesn't change)

Known issue: goto command may return success without navigating. Use JS workaround:

bash
# Instead of goto, use evaluate to navigate
curl -s -X POST http://localhost:9222/command -H "Content-Type: application/json" \
  -d '{"id":"1","method":"evaluate","params":{"script":"location.href = \"https://example.com\"; true"}}'

# Wait for page load
sleep 3
curl -s -X POST http://localhost:9222/command -H "Content-Type: application/json" \
  -d '{"id":"2","method":"waitForReady","params":{"timeout":10}}'
Server not responding
bash
# Check if app is running
xcrun simctl listapps booted | grep atl

# Restart the app
xcrun simctl terminate booted com.atl.browser
xcrun simctl launch booted com.atl.browser

# Check logs
xcrun simctl spawn booted log show --predicate 'process == "AtlBrowser"' --last 1m
Need to rebuild (iOS version changes)
bash
cd ~/Atl/core/AtlBrowser
xcodebuild -workspace AtlBrowser.xcworkspace -scheme AtlBrowser -sdk iphonesimulator build
xcrun simctl install booted ~/Library/Developer/Xcode/DerivedData/AtlBrowser-*/Build/Products/Debug-iphonesimulator/AtlBrowser.app
xcrun simctl launch booted com.atl.browser
Port 9222 in use

The ATL server runs inside the simulator app. If port 9222 is blocked, check for other processes:

bash
lsof -i :9222

Best Practices

1. Clean UI Before Acting

Real users dismiss popups. You should too.

bash
# Before any workflow, check for and dismiss:
# - Cookie consent banners
# - Newsletter popups  
# - Health/privacy consent modals
# - "Download our app" prompts
atl_mark
for KEYWORD in "close" "dismiss" "no thanks" "accept" "got it" "continue"; do
  LABEL=$(atl_find "$KEYWORD")
  [ -n "$LABEL" ] && atl_click $LABEL && sleep 1
done
2. Verify State After Actions

Don't assume — confirm.

bash
atl_click $ADD_TO_CART
sleep 2
# Check if cart updated
CART=$(atl_find "cart [1-9]")
if [ -z "$CART" ]; then
  # Didn't work - take screenshot to see why
  atl_screenshot /tmp/debug.png
  echo "Action may have opened a modal - check screenshot"
fi
3. Use Viewport Coordinates for Taps

Marks give page-relative coordinates. For tap to work, the element must be visible.

bash
# Option A: Scroll element into view first
curl -s -X POST http://localhost:9222/command -H "Content-Type: application/json" \
  -d '{"id":"1","method":"evaluate","params":{"script":"document.querySelector(\"#my-button\").scrollIntoView()"}}'

# Option B: Get viewport-relative coords via JS
curl -s -X POST http://localhost:9222/command -H "Content-Type: application/json" \
  -d '{"id":"2","method":"evaluate","params":{"script":"var r = document.querySelector(\"#my-button\").getBoundingClientRect(); JSON.stringify({x: r.x + r.width/2, y: r.y + r.height/2})"}}'
4. Screenshot is Your Debugging Superpower

When in doubt, look.

bash
atl_screenshot /tmp/current-state.png
# Then analyze with vision or just open the file

Notes

  • ATL runs inside the iOS Simulator, sharing the host's network
  • Port 9222 is the default (matches Chrome DevTools Protocol convention)
  • The mark system shows red numbered labels on interactive elements
  • Screenshots are PNG base64-encoded; use base64 -d to decode
  • iOS 26+ compatible (fixed NWListener binding issue)

Requirements

  • macOS with Xcode installed
  • iOS Simulator (comes with Xcode)
  • That's it!

Examples

See examples/ folder:

  • test-browse.sh - Quick bash test workflow

API Reference

For machine-readable API spec, see openapi.yaml — includes all commands, parameters, and response schemas.

Source

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/atl-mobile of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json
  • scripts/setup.sh

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Atl Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Atl Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Atl Browser this skillLeoYeAI/openclaw-master-skills2.2k—~6.5kAutomated safety check: PassMIT
Connect Apps with ComposioComposioHQ/awesome-claude-skills77k3 repos~557Automated safety check: PassNone
Interceptor iOSHacker-Valley-Media/Interceptor522—~1.9kAutomated safety check: PassCustom licence
Add Connection Typebagofwords1/bagofwords459—~2.3kAutomated safety check: PassCustom licence
Composio Cloud Toolsquarqlabs/argus279—~543Automated safety check: PassApache-2.0
Phone And Simulator GUI ControlCore-Mate/OpenGUI1.8k—~7.9kAutomated safety check: WarnCustom licence

Similar skills

  • Connect Apps with Composio

    ComposioHQ/awesome-claude-skills

    Connects an agent to 1000+ external apps through the Composio Tool Router plugin, so it can actually send emails, create issues and post messages instead of only drafting them.

    77k GitHub starsUsed in 3 repos~557 tokens
    Productivity & AutomationAuto-check passed
  • Interceptor iOS

    Hacker-Valley-Media/Interceptor

    Drive any installed app on an owned, unlocked, Developer-Mode iPhone via interceptor ios : ref-tagged element trees, deterministic coordinate taps (click), reliable text entry (type/keys), scroll…

    522 GitHub stars~1.9k tokensUpdated 7 days ago
    Productivity & AutomationAuto-check passed
  • Add Connection Type

    bagofwords1/bagofwords

    Add a new data source / connection type (e.g. An agent skill from bagofwords1/bagofwords.

    459 GitHub stars~2.3k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Composio Cloud Tools

    quarqlabs/argus

    Routes requests to external SaaS apps such as GitHub, Gmail, Google Calendar, Slack, Notion and Linear through cloud tools, with safeguards on irreversible actions.

    279 GitHub stars~543 tokensUpdated 4 mo ago
    Productivity & AutomationAuto-check passed
  • Completes an already-authorized task on a real Android phone or a macOS iOS simulator one screenshot-driven action at a time.

    1.8k GitHub stars~7.9k tokensUpdated yesterday
    Productivity & AutomationAuto-check: warnings
  • Cloudflare API Key Automation

    ComposioHQ/awesome-claude-skills

    Automate Cloudflare API tasks via Rube MCP (Composio). An agent skill from ComposioHQ/awesome-claude-skills.

    77k GitHub starsUsed in 3 repos~764 tokens
    Productivity & AutomationAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Atl Browser

What does Atl Browser do?

Mobile browser and native app automation via ATL (iOS Simulator). Atl Browser is an agent skill from LeoYeAI/openclaw-master-skills. Mobile browser and native app automation via ATL (iOS Simulator).

When should I use Atl Browser?

Atl Browser fits situations like: tasks that involve App automation through connectors.

How do I install Atl Browser in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill atl-browser -a claude-code`. Or copy the skill folder (skills/atl-mobile in LeoYeAI/openclaw-master-skills) into .claude/skills/atl-browser in your project. Claude Code loads it when a task matches its description.

How do I install Atl Browser in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill atl-browser -a codex`. Or copy the skill folder (skills/atl-mobile in LeoYeAI/openclaw-master-skills) into .agents/skills/atl-browser in your project. Codex loads it when a task matches its description.

Can I use Atl Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill atl-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/atl-browser, .gemini/skills/atl-browser, .github/skills/atl-browser and .opencode/skills/atl-browser in your project.

What does Atl Browser need to run?

Going by SKILL.md and its folder, Atl Browser needs a shell for the scripts in its folder and the command-line tools its instructions call (curl, xcrun, jq and xcodebuild). Our summary lists: A Bash shell.

Does Atl Browser access the network?

SKILL.md names 3 domains. In commands or code: store.com and apple.com; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Atl Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Atl Browser use?

Atl Browser is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Atl Browser use?

About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Atl Browser?

Skills that share tags, products or a category with Atl Browser: Connect Apps with Composio (ComposioHQ/awesome-claude-skills, 77k stars), Interceptor iOS (Hacker-Valley-Media/Interceptor, 522 stars), Add Connection Type (bagofwords1/bagofwords, 459 stars) and Composio Cloud Tools (quarqlabs/argus, 279 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Atl Browser?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.