Automates browser and Electron app interactions for user-flow validation.

Apache-2.0Auto-check passedProductivity & Automation

Install Agent Browser

skills CLI
$ npx skills add Intelligent-Internet/zenith --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Intelligent-Internet/zenith agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
334
Token cost
~6.9k tokens
SKILL.md length
1,210 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Automates browser and Electron app interactions for user-flow validation.

  • Works in 4 steps: Navigate: agent-browser open → Snapshot: agent-browser snapshot -i… → Interact using refs from the snapshot → …
  • Tasks that involve Browser automation
  • SKILL.md covers One-time setup (fresh machines), Quick start, Core workflow and Command chaining, plus 5 more sections
  • Calls openssl, npm and xcrun; reaches proxy.com; needs AGENT_BROWSER_ENCRYPTION_KEY

What it does

Agent Browser is an agent skill from Intelligent-Internet/zenith. Automates browser and Electron app interactions for user-flow validation.

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation and UX design. It works with Electron and Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Browser automation
  • Tasks that involve UX design

Example prompts

  • “Use the agent-browser skill to automate browser and Electron app interactions for user-flow validation”
  • “/agent-browser”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Navigate: agent-browser open
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

What it can do on your machine

Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • openssl
    • npm
    • xcrun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • proxy.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AGENT_BROWSER_ENCRYPTION_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser loads about 6.9k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 1,210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 1,210 words, ~6,925 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Automates browser and Electron app interactions for user-flow validation.

Browser Automation with agent-browser

One-time setup (fresh machines)

If you see an error like "Executable doesn't exist ... chrome-headless-shell", install the bundled Chromium:

bash
agent-browser install

Quick start

bash
agent-browser open <url>        # Navigate to page
agent-browser snapshot -i       # Get interactive elements with refs
agent-browser click @e1         # Click element by ref
agent-browser fill @e2 "text"   # Fill input by ref
agent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Command chaining

Commands can be chained with && in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls.

bash
# Chain open + wait + snapshot in one call
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser snapshot -i

# Chain multiple interactions
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "password123" && agent-browser click @e3

# Navigate and capture
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png

When to chain: Use && when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs).

Commands

Navigation
bash
agent-browser open <url>      # Navigate to URL (aliases: goto, navigate)
                              # Supports: https://, http://, file://, about:, data://
                              # Auto-prepends https:// if no protocol given
agent-browser back            # Go back
agent-browser forward         # Go forward
agent-browser reload          # Reload page
agent-browser close           # Close browser (aliases: quit, exit)
agent-browser connect 9222    # Connect to browser via CDP port
Snapshot (page analysis)
bash
agent-browser snapshot            # Full accessibility tree
agent-browser snapshot -i         # Interactive elements only (recommended)
agent-browser snapshot -i -C      # Include cursor-interactive elements (divs with onclick, cursor:pointer)
agent-browser snapshot -c         # Compact output
agent-browser snapshot -d 3       # Limit depth to 3
agent-browser snapshot -s "#main" # Scope to CSS selector
Interactions (use @refs from snapshot)
bash
agent-browser click @e1           # Click
agent-browser click @e1 --new-tab # Click and open in new tab
agent-browser dblclick @e1        # Double-click
agent-browser focus @e1           # Focus element
agent-browser fill @e2 "text"     # Clear and type
agent-browser type @e2 "text"     # Type without clearing
agent-browser press Enter         # Press key (alias: key)
agent-browser press Control+a     # Key combination
agent-browser keydown Shift       # Hold key down
agent-browser keyup Shift         # Release key
agent-browser hover @e1           # Hover
agent-browser check @e1           # Check checkbox
agent-browser uncheck @e1         # Uncheck checkbox
agent-browser select @e1 "value"  # Select dropdown option
agent-browser select @e1 "a" "b"  # Select multiple options
agent-browser scroll down 500     # Scroll page (default: down 300px)
agent-browser scrollintoview @e1  # Scroll element into view (alias: scrollinto)
agent-browser drag @e1 @e2        # Drag and drop
agent-browser upload @e1 file.pdf # Upload files
Get information
bash
agent-browser get text @e1        # Get element text
agent-browser get text body > page.txt  # Get all page text to file
agent-browser get html @e1        # Get innerHTML
agent-browser get value @e1       # Get input value
agent-browser get attr @e1 href   # Get attribute
agent-browser get title           # Get page title
agent-browser get url             # Get current URL
agent-browser get count ".item"   # Count matching elements
agent-browser get box @e1         # Get bounding box
agent-browser get styles @e1      # Get computed styles (font, color, bg, etc.)
Check state
bash
agent-browser is visible @e1      # Check if visible
agent-browser is enabled @e1      # Check if enabled
agent-browser is checked @e1      # Check if checked
Screenshots & PDF
bash
agent-browser screenshot              # Save to a temporary directory
agent-browser screenshot path.png     # Save to a specific path
agent-browser screenshot --full       # Full page screenshot
agent-browser screenshot --annotate   # Annotated screenshot with numbered element labels
agent-browser pdf output.pdf          # Save as PDF
Video recording
bash
agent-browser record start ./demo.webm    # Start recording (uses current URL + state)
agent-browser click @e1                   # Perform actions
agent-browser record stop                 # Stop and save video
agent-browser record restart ./take2.webm # Stop current + start new recording

Recording creates a fresh context but preserves cookies/storage from your session. If no URL is provided, it automatically returns to your current page. For smooth demos, explore first, then start recording.

Diff (compare page states)
bash
agent-browser diff snapshot                          # Compare current vs last snapshot
agent-browser diff snapshot --baseline before.txt    # Compare current vs saved file
agent-browser diff screenshot --baseline before.png  # Visual pixel diff
agent-browser diff url <url1> <url2>                 # Compare two pages
agent-browser diff url <url1> <url2> --wait-until networkidle  # Custom wait strategy
agent-browser diff url <url1> <url2> --selector "#main"  # Scope to element

Use diff snapshot after performing an action to verify it had the intended effect. This compares the current accessibility tree against the last snapshot taken in the session.

bash
# Typical workflow: snapshot -> action -> diff
agent-browser snapshot -i          # Take baseline snapshot
agent-browser click @e2            # Perform action
agent-browser diff snapshot        # See what changed (auto-compares to last snapshot)

diff snapshot output uses + for additions and - for removals, similar to git diff. diff screenshot produces a diff image with changed pixels highlighted in red, plus a mismatch percentage.

Wait
bash
agent-browser wait @e1                     # Wait for element
agent-browser wait 2000                    # Wait milliseconds
agent-browser wait --text "Success"        # Wait for text (or -t)
agent-browser wait --url "**/dashboard"    # Wait for URL pattern (or -u)
agent-browser wait --load networkidle      # Wait for network idle (or -l)
agent-browser wait --fn "window.ready"     # Wait for JS condition (or -f)
Mouse control
bash
agent-browser mouse move 100 200      # Move mouse
agent-browser mouse down left         # Press button
agent-browser mouse up left           # Release button
agent-browser mouse wheel 100         # Scroll wheel
Semantic locators (alternative to refs)

When refs are unavailable or unreliable, use semantic locators:

bash
agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find text "Sign In" click --exact      # Exact match only
agent-browser find label "Email" fill "user@test.com"
agent-browser find placeholder "Search" type "query"
agent-browser find alt "Logo" click
agent-browser find title "Close" click
agent-browser find testid "submit-btn" click
agent-browser find first ".item" click
agent-browser find last ".item" click
agent-browser find nth 2 "a" hover
Browser settings
bash
agent-browser set viewport 1920 1080          # Set viewport size
agent-browser set device "iPhone 14"          # Emulate device
agent-browser set geo 37.7749 -122.4194       # Set geolocation (alias: geolocation)
agent-browser set offline on                  # Toggle offline mode
agent-browser set headers '{"X-Key":"v"}'     # Extra HTTP headers
agent-browser set credentials user pass       # HTTP basic auth (alias: auth)
agent-browser set media dark                  # Emulate color scheme
agent-browser set media light reduced-motion  # Light mode + reduced motion
Cookies & Storage
bash
agent-browser cookies                     # Get all cookies
agent-browser cookies set name value      # Set cookie
agent-browser cookies clear               # Clear cookies
agent-browser storage local               # Get all localStorage
agent-browser storage local key           # Get specific key
agent-browser storage local set k v       # Set value
agent-browser storage local clear         # Clear all
Network
bash
agent-browser network route <url>              # Intercept requests
agent-browser network route <url> --abort      # Block requests
agent-browser network route <url> --body '{}'  # Mock response
agent-browser network unroute [url]            # Remove routes
agent-browser network requests                 # View tracked requests
agent-browser network requests --filter api    # Filter requests
Tabs & Windows
bash
agent-browser tab                 # List tabs
agent-browser tab new [url]       # New tab
agent-browser tab 2               # Switch to tab by index
agent-browser tab close           # Close current tab
agent-browser tab close 2         # Close tab by index
agent-browser window new          # New window
Frames
bash
agent-browser frame "#iframe"     # Switch to iframe
agent-browser frame main          # Back to main frame
Dialogs
bash
agent-browser dialog accept [text]  # Accept dialog
agent-browser dialog dismiss        # Dismiss dialog
JavaScript (eval)

Use eval to run JavaScript in the browser context. Shell quoting can corrupt complex expressions -- use --stdin or -b to avoid issues.

bash
# Simple expressions work with regular quoting
agent-browser eval 'document.title'
agent-browser eval 'document.querySelectorAll("img").length'

# Complex JS: use --stdin with heredoc (RECOMMENDED)
agent-browser eval --stdin <<'EVALEOF'
JSON.stringify(
  Array.from(document.querySelectorAll("img"))
    .filter(i => !i.alt)
    .map(i => ({ src: i.src.split("/").pop(), width: i.width }))
)
EVALEOF

# Alternative: base64 encoding (avoids all shell escaping issues)
agent-browser eval -b "$(echo -n 'Array.from(document.querySelectorAll("a")).map(a => a.href)' | base64)"

Rules of thumb:

  • Single-line, no nested quotes -> regular eval 'expression' with single quotes is fine
  • Nested quotes, arrow functions, template literals, or multiline -> use eval --stdin <<'EVALEOF'
  • Programmatic/generated scripts -> use eval -b with base64
Profiling
bash
agent-browser profiler start         # Start Chrome DevTools profiling
agent-browser profiler stop trace.json # Stop and save profile (path optional)

Global options

bash
agent-browser --session <name> ...    # Isolated browser session
agent-browser --session-name <name> ...  # Named session with auto-save/restore
agent-browser --json ...              # JSON output for parsing
agent-browser --headed ...            # Show browser window (not headless)
agent-browser --full ...              # Full page screenshot (-f)
agent-browser --cdp <port> ...        # Connect via Chrome DevTools Protocol
agent-browser --auto-connect ...      # Auto-discover running Chrome with remote debugging
agent-browser -p <provider> ...       # Cloud browser provider (--provider)
agent-browser --proxy <url> ...       # Use proxy server
agent-browser --headers <json> ...    # HTTP headers scoped to URL's origin
agent-browser --executable-path <p>   # Custom browser executable
agent-browser --extension <path> ...  # Load browser extension (repeatable)
agent-browser --allow-file-access ... # Allow file:// URL access for local files
agent-browser --ignore-https-errors ...  # Ignore HTTPS certificate errors
agent-browser --config <path> ...     # Custom config file path
agent-browser --help                  # Show help (-h)
agent-browser --version               # Show version (-V)
agent-browser <command> --help        # Show detailed help for a command
Proxy support
bash
agent-browser --proxy http://proxy.com:8080 open example.com
agent-browser --proxy http://user:pass@proxy.com:8080 open example.com
agent-browser --proxy socks5://proxy.com:1080 open example.com

Environment variables

bash
AGENT_BROWSER_SESSION="mysession"            # Default session name
AGENT_BROWSER_EXECUTABLE_PATH="/path/chrome" # Custom browser path
AGENT_BROWSER_EXTENSIONS="/ext1,/ext2"       # Comma-separated extension paths
AGENT_BROWSER_PROVIDER="your-cloud-browser-provider"  # Cloud browser provider (select browseruse or browserbase)
AGENT_BROWSER_STREAM_PORT="9223"             # WebSocket streaming port
AGENT_BROWSER_HOME="/path/to/agent-browser"  # Custom install location (for daemon.js)
AGENT_BROWSER_ENCRYPTION_KEY="<hex>"         # Encrypt session state at rest
AGENT_BROWSER_CONFIG="/path/to/config.json"  # Custom config file path

Configuration file

Create agent-browser.json in the project root for persistent settings:

json
{
  "headed": true,
  "proxy": "http://localhost:8080",
  "profile": "./browser-data"
}

Priority (lowest to highest): ~/.agent-browser/config.json < ./agent-browser.json < env vars < CLI flags. Use --config <path> or AGENT_BROWSER_CONFIG env var for a custom config file. All CLI options map to camelCase keys (e.g., --executable-path -> "executablePath"). Boolean flags accept true/false values. Extensions from user and project configs are merged, not replaced.

Ref lifecycle (important)

Refs (@e1, @e2, etc.) are invalidated when the page changes. Always re-snapshot after:

  • Clicking links or buttons that navigate
  • Form submissions
  • Dynamic content loading (dropdowns, modals)
bash
agent-browser click @e5              # Navigates to new page
agent-browser snapshot -i            # MUST re-snapshot
agent-browser click @e1              # Use new refs

Annotated screenshots (vision mode)

Use --annotate to take a screenshot with numbered labels overlaid on interactive elements. Each label [N] maps to ref @eN. This also caches refs, so you can interact with elements immediately without a separate snapshot.

bash
agent-browser screenshot --annotate
# Output includes the image path and a legend:
#   [1] @e1 button "Submit"
#   [2] @e2 link "Home"
#   [3] @e3 textbox "Email"
agent-browser click @e2              # Click using ref from annotated screenshot

Use annotated screenshots when:

  • The page has unlabeled icon buttons or visual-only elements
  • You need to verify visual layout or styling
  • Canvas or chart elements are present (invisible to text snapshots)
  • You need spatial reasoning about element positions

Timeouts and slow pages

The default Playwright timeout is 60 seconds for local browsers. For slow websites or large pages, use explicit waits instead of relying on the default timeout:

bash
agent-browser wait --load networkidle      # Wait for network activity to settle (best for slow pages)
agent-browser wait "#content"              # Wait for a specific element to appear
agent-browser wait @e1                     # Wait for element by ref
agent-browser wait --url "**/dashboard"    # Wait for URL pattern (useful after redirects)
agent-browser wait --fn "document.readyState === 'complete'"  # Wait for JS condition
agent-browser wait 5000                    # Wait fixed duration (milliseconds) as last resort

When dealing with consistently slow websites, use wait --load networkidle after open to ensure the page is fully loaded before taking a snapshot.

Example: Form submission

bash
agent-browser open https://example.com/form
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result

Example: Authentication with saved state

bash
# Login once
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "username"
agent-browser fill @e2 "password"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Later sessions: load saved state
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard

Session persistence

bash
# Auto-save/restore cookies and localStorage across browser restarts
agent-browser --session-name myapp open https://app.example.com/login
# ... login flow ...
agent-browser close  # State auto-saved to ~/.agent-browser/sessions/

# Next time, state is auto-loaded
agent-browser --session-name myapp open https://app.example.com/dashboard

# Encrypt state at rest
export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32)
agent-browser --session-name secure open https://app.example.com

# Manage saved states
agent-browser state list
agent-browser state show myapp-default.json
agent-browser state clear myapp
agent-browser state clean --older-than 7

Sessions (parallel browsers)

bash
agent-browser --session test1 open site-a.com
agent-browser --session test2 open site-b.com
agent-browser session list

Always close your browser session when done to avoid leaked processes:

bash
agent-browser close                    # Close default session
agent-browser --session test1 close    # Close specific session

If a previous session was not closed properly, the daemon may still be running. Use agent-browser close to clean it up before starting new work.

Data extraction

bash
agent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e5           # Get specific element text
agent-browser get text body > page.txt  # Get all page text

# JSON output for parsing
agent-browser snapshot -i --json
agent-browser get text @e1 --json

Connect to existing Chrome

bash
# Auto-discover running Chrome with remote debugging enabled
agent-browser --auto-connect open https://example.com
agent-browser --auto-connect snapshot

# Or with explicit CDP port
agent-browser --cdp 9222 snapshot

Local files (PDFs, HTML)

bash
# Open local files with file:// URLs
agent-browser --allow-file-access open file:///path/to/document.pdf
agent-browser --allow-file-access open file:///path/to/page.html
agent-browser screenshot output.png

iOS Simulator (Mobile Safari)

bash
# List available iOS simulators
agent-browser device list

# Launch Safari on a specific device
agent-browser -p ios --device "iPhone 16 Pro" open https://example.com

# Same workflow as desktop - snapshot, interact, re-snapshot
agent-browser -p ios snapshot -i
agent-browser -p ios tap @e1          # Tap (alias for click)
agent-browser -p ios fill @e2 "text"
agent-browser -p ios swipe up         # Mobile-specific gesture

# Take screenshot
agent-browser -p ios screenshot mobile.png

# Close session (shuts down simulator)
agent-browser -p ios close

Requirements: macOS with Xcode, Appium (npm install -g appium && appium driver install xcuitest)

Real devices: Works with physical iOS devices if pre-configured. Use --device "<UDID>" where UDID is from xcrun xctrace list devices.

Debugging

bash
agent-browser --headed open example.com   # Show browser window
agent-browser --cdp 9222 snapshot         # Connect via CDP port
agent-browser connect 9222                # Alternative: connect command
agent-browser console                     # View console messages
agent-browser console --clear             # Clear console
agent-browser errors                      # View page errors
agent-browser errors --clear              # Clear errors
agent-browser highlight @e1               # Highlight element
agent-browser trace start                 # Start recording trace
agent-browser trace stop trace.zip        # Stop and save trace
agent-browser record start ./debug.webm   # Record video from current page
agent-browser record stop                 # Save recording
agent-browser profiler start              # Start Chrome DevTools profiling
agent-browser profiler stop trace.json    # Stop and save profile

HTTPS Certificate Errors

For sites with self-signed or invalid certificates:

bash
agent-browser open https://localhost:8443 --ignore-https-errors

Electron App Automation

Automate any Electron desktop app using agent-browser. Electron apps are built on Chromium and expose a Chrome DevTools Protocol (CDP) port that agent-browser can connect to, enabling the same snapshot-interact workflow used for web pages.

Core Workflow
  1. Launch the Electron app with remote debugging enabled
  2. Connect agent-browser to the CDP port
  3. Snapshot to discover interactive elements
  4. Interact using element refs
  5. Re-snapshot after navigation or state changes
bash
# Launch an Electron app with remote debugging
open -a "Slack" --args --remote-debugging-port=9222

# Connect agent-browser to the app
agent-browser connect 9222

# Standard workflow from here
agent-browser snapshot -i
agent-browser click @e5
agent-browser screenshot slack-desktop.png
Show full SKILL.md (479 more words)Show less
Launching Electron Apps with CDP

Every Electron app supports the --remote-debugging-port flag since it's built into Chromium.

macOS
bash
# Slack
open -a "Slack" --args --remote-debugging-port=9222

# VS Code
open -a "Visual Studio Code" --args --remote-debugging-port=9223

# Discord
open -a "Discord" --args --remote-debugging-port=9224

# Figma
open -a "Figma" --args --remote-debugging-port=9225

# Notion
open -a "Notion" --args --remote-debugging-port=9226

# Spotify
open -a "Spotify" --args --remote-debugging-port=9227
Linux
bash
slack --remote-debugging-port=9222
code --remote-debugging-port=9223
discord --remote-debugging-port=9224
Windows
bash
"C:\Users\%USERNAME%\AppData\Local\slack\slack.exe" --remote-debugging-port=9222
"C:\Users\%USERNAME%\AppData\Local\Programs\Microsoft VS Code\Code.exe" --remote-debugging-port=9223

Important: If the app is already running, quit it first, then relaunch with the flag. The --remote-debugging-port flag must be present at launch time.

Connecting to Electron Apps
bash
# Connect to a specific port
agent-browser connect 9222

# Or use --cdp on each command
agent-browser --cdp 9222 snapshot -i

# Auto-discover a running Chromium-based app
agent-browser --auto-connect snapshot -i

After connect, all subsequent commands target the connected app without needing --cdp.

Tab Management

Electron apps often have multiple windows or webviews. Use tab commands to list and switch between them:

bash
# List all available targets (windows, webviews, etc.)
agent-browser tab

# Switch to a specific tab by index
agent-browser tab 2

# Switch by URL pattern
agent-browser tab --url "*settings*"
Common Electron Patterns
Inspect and Navigate an App
bash
open -a "Slack" --args --remote-debugging-port=9222
sleep 3  # Wait for app to start
agent-browser connect 9222
agent-browser snapshot -i
# Read the snapshot output to identify UI elements
agent-browser click @e10  # Navigate to a section
agent-browser snapshot -i  # Re-snapshot after navigation
Take Screenshots of Desktop Apps
bash
agent-browser connect 9222
agent-browser screenshot app-state.png
agent-browser screenshot --full full-app.png
agent-browser screenshot --annotate annotated-app.png
Extract Data from a Desktop App
bash
agent-browser connect 9222
agent-browser snapshot -i
agent-browser get text @e5
agent-browser snapshot --json > app-state.json
Fill Forms in Desktop Apps
bash
agent-browser connect 9222
agent-browser snapshot -i
agent-browser fill @e3 "search query"
agent-browser press Enter
agent-browser wait 1000
agent-browser snapshot -i
Run Multiple Apps Simultaneously

Use named sessions only when you need to control multiple Electron apps at the same time. Each session spawns a separate Chromium process (~300MB RAM), so only create multiple sessions when explicitly required and close each one when done.

bash
# Connect to Slack
agent-browser --session slack connect 9222

# Connect to VS Code
agent-browser --session vscode connect 9223

# Interact with each independently
agent-browser --session slack snapshot -i
agent-browser --session vscode snapshot -i

# Close sessions when done
agent-browser --session slack close
agent-browser --session vscode close
Color Scheme

Playwright overrides the color scheme to light by default when connecting via CDP. To preserve dark mode:

bash
agent-browser connect 9222
agent-browser --color-scheme dark snapshot -i

Or set it globally:

bash
AGENT_BROWSER_COLOR_SCHEME=dark agent-browser connect 9222
Electron Troubleshooting
"Connection refused" or "Cannot connect"
  • Make sure the app was launched with --remote-debugging-port=NNNN
  • If the app was already running, quit and relaunch with the flag
  • Check that the port isn't in use by another process: lsof -i :9222
App launches but connect fails
  • Wait a few seconds after launch before connecting (sleep 3)
  • Some apps take time to initialize their webview
Elements not appearing in snapshot
  • The app may use multiple webviews. Use agent-browser tab to list targets and switch to the right one
  • Use agent-browser snapshot -i -C to include cursor-interactive elements (divs with onclick handlers)
Cannot type in input fields
  • Try agent-browser keyboard type "text" to type at the current focus without a selector
  • Some Electron apps use custom input components; use agent-browser keyboard inserttext "text" to bypass key events

CRITICAL:

  1. Screenshots are required for UI testing. Text snapshots do not catch layout issues, styling problems, alignment, z-index issues, etc. Always use agent-browser screenshot --annotate to get numbered element labels overlaid on the screenshot, which enables both visual verification and immediate interaction via refs.
  2. Use a single browser session. Do not create multiple --session names — each spawns a separate Chromium process (~300MB RAM). Use one session and call agent-browser open <new-url> to navigate between pages within it. If you must create a new session, close the previous one first with agent-browser --session <name> close. Exception: When controlling multiple Electron apps simultaneously on different CDP ports, multiple named sessions are acceptable — but only when the user explicitly needs it, and each session must be closed when no longer needed.
  3. When done, close your browser with agent-browser close (or agent-browser --session <name> close) to free resources.
  4. Electron apps: quit first. If the Electron app is already running, you must quit it completely before relaunching with --remote-debugging-port. The flag only takes effect at launch time.

© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in zenith/src/zenith_harness/bundled/skills/agent-browser of Intelligent-Internet/zenith.

Open the folder on GitHubat commit a8d9b57

Compare with similar skills

Agent Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser this skillIntelligent-Internet/zenith334—~6.9kAutomated safety check: PassApache-2.0
Electron Devtools Testingankitvgupta/exo496—~2.4kAutomated safety check: PassCustom licence
Electron App Automationvercel-labs/agent-browser44k5 repos~1.7kAutomated safety check: PassApache-2.0
Next Dev Loopsanity-io/ui1768 repos~2kAutomated safety check: PassMIT
Skyvern Browser AutomationSkyvern-AI/skyvern23k—~1.9kAutomated safety check: PassAGPL-3.0
Zerotoken OpenclawAMOS144/ZeroToken4551 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Test the Electron app interactively using Chrome DevTools Protocol.

    496 GitHub stars~2.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Electron App Automation

    vercel-labs/agent-browser

    Official

    Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port.

    44k GitHub starsUsed in 5 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Next Dev Loop

    sanity-io/ui

    Official

    Verify Next.js runtime behavior after editing app code. An agent skill from sanity-io/ui.

    176 GitHub starsUsed in 8 repos~2k tokens
    Productivity & AutomationAuto-check passed
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Zerotoken Openclaw

    AMOS144/ZeroToken

    A skill your agent uses when using ZeroToken MCP via OpenClaw for browser automation, trajectory recording and low-token replay, especially for recurring or scheduled browser tasks.

    455 GitHub starsUsed in 1 repo~2.9k tokens
    Productivity & AutomationAuto-check passed
  • AIPex Browser Control

    AIPexStudio/AIPex

    Lets an agent drive Chrome through the AIPex extension and its MCP bridge: navigation, clicking, form filling, screenshots, tab management and downloads.

    1.3k GitHub stars~1.8k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed

More from Intelligent-Internet/zenith

  • Engineering Mission Playbook

    Intelligent-Internet/zenith

    A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…

    334 GitHub stars~8.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Benchmark Validator

    Intelligent-Internet/zenith

    Benchmark validation procedure for one assigned benchmark-related target.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Scrutiny Validator

    Intelligent-Internet/zenith

    Adversarial scrutiny procedure for engineering validation assignments.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • User Testing Validator

    Intelligent-Internet/zenith

    Real-surface validation coordinator for engineering validation assignments.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Optimization Mission Playbook

    Intelligent-Internet/zenith

    Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…

    334 GitHub stars~11k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Agent Browser

What does Agent Browser do?

Automates browser and Electron app interactions for user-flow validation. Agent Browser is an agent skill from Intelligent-Internet/zenith. Automates browser and Electron app interactions for user-flow validation.

When should I use Agent Browser?

Agent Browser fits situations like: tasks that involve Browser automation; tasks that involve UX design.

How do I install Agent Browser in Claude Code?

Run `npx skills add Intelligent-Internet/zenith --skill agent-browser -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/agent-browser in Intelligent-Internet/zenith) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser in Codex?

Run `npx skills add Intelligent-Internet/zenith --skill agent-browser -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/agent-browser in Intelligent-Internet/zenith) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser need to run?

Going by SKILL.md and its folder, Agent Browser needs the command-line tools its instructions call (openssl, npm and xcrun) and credentials named AGENT_BROWSER_ENCRYPTION_KEY.

Does Agent Browser access the network?

SKILL.md names 1 domain. In commands or code: proxy.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Agent Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser use?

Agent Browser is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser?

Skills that share tags, products or a category with Agent Browser: Electron Devtools Testing (ankitvgupta/exo, 496 stars), Electron App Automation (vercel-labs/agent-browser, 44k stars), Next Dev Loop (sanity-io/ui, 176 stars) and Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser?

Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 334 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.

Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.