Agent skill

Agent Browser CLI

by slopus in slopus/happy

Drives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows.

MITAuto-check passedProductivity & Automation

Install Agent Browser CLI

skills CLI
$ npx skills add slopus/happy --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install slopus/happy agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
24k
Token cost
~1.7k tokens
SKILL.md length
191 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Drives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows.

  • Works in 4 steps: Navigate: agent-browser open → Snapshot: agent-browser snapshot -i (get… → Interact: Use refs to click, fill, select → …
  • Testing a web flow in a real browser
  • SKILL.md covers Core Workflow, Command Chaining, Essential Commands and Common Patterns, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill teaches the agent to drive a real browser with the `agent-browser` CLI. Every task follows the same loop: open a URL, take a snapshot with `snapshot -i` that lists interactive elements with short refs such as `@e1`, interact with those refs by clicking, filling or selecting, and snapshot again whenever the page changes, because refs are invalidated by navigation, form submissions and dynamically loaded content.

A background daemon keeps the browser alive between commands, so several can be chained with `&&` in one call, for example open, wait for network idle and snapshot, but they should run separately when the output of one has to be read first. Common patterns shown are form submission, logging in once and saving state for reuse, and extracting text from elements. An essential-commands list covers navigation and more, and the excerpt ends at the annotated-screenshot section for vision mode.

When your agent uses it

  • Testing a web flow in a real browser
  • Filling and submitting a form through the CLI
  • Logging in once and reusing the saved session state
  • Pulling text out of page elements

Example prompts

  • “Open the signup page on localhost, fill in the form with test details and tell me what the confirmation says.”
  • “Log in to the staging app, save the session state and check that the dashboard loads.”
  • “Get the product names and prices from https://example.com/products.”

Requirements

  • The agent-browser CLI, run directly or through npx
  • Pre-approved tools (allowed-tools): Bash(npx agent-browser:*), Bash(agent-browser:*)

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Navigate: agent-browser open
  2. Snapshot: agent-browser snapshot -i (get element refs like @e1, @e2)
  3. Interact: Use refs to click, fill, select
  4. Re-snapshot: After navigation or DOM changes, get fresh refs

What it can do on your machine

Read from SKILL.md and the folder at commit 7ea7017. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(npx agent-browser:*)
    • Bash(agent-browser:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser CLI loads about 1.7k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 191 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from slopus/happy at commit 7ea7017, republished under its MIT licence (© slopus). 191 words, ~1,748 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Browser automation CLI for AI agents. Use this when asked to test something in a real browser.
allowed-tools
Bash(npx agent-browser:*), Bash(agent-browser:*)

Browser Automation with agent-browser

Core Workflow

Every browser automation follows this pattern:

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (get element refs like @e1, @e2)
  3. Interact: Use refs to click, fill, select
  4. Re-snapshot: After navigation or DOM changes, get fresh refs
bash
agent-browser open https://example.com/form
agent-browser snapshot -i
# Output: @e1 [input type="email"], @e2 [input type="password"], @e3 [button] "Submit"

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result

Command Chaining

Commands can be chained with && in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls.

bash
# Chain open + wait + snapshot in one call
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser snapshot -i

# Chain multiple interactions
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "password123" && agent-browser click @e3

# Navigate and capture
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png

When to chain: Use && when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs).

Essential Commands

bash
# Navigation
agent-browser open <url>              # Navigate (aliases: goto, navigate)
agent-browser close                   # Close browser

# Snapshot
agent-browser snapshot -i             # Interactive elements with refs (recommended)
agent-browser snapshot -i -C          # Include cursor-interactive elements (divs with onclick, cursor:pointer)
agent-browser snapshot -s "#selector" # Scope to CSS selector

# Interaction (use @refs from snapshot)
agent-browser click @e1               # Click element
agent-browser click @e1 --new-tab     # Click and open in new tab
agent-browser fill @e2 "text"         # Clear and type text
agent-browser type @e2 "text"         # Type without clearing
agent-browser select @e1 "option"     # Select dropdown option
agent-browser check @e1               # Check checkbox
agent-browser press Enter             # Press key
agent-browser keyboard type "text"    # Type at current focus (no selector)
agent-browser keyboard inserttext "text"  # Insert without key events
agent-browser scroll down 500         # Scroll page
agent-browser scroll down 500 --selector "div.content"  # Scroll within a specific container

# Get information
agent-browser get text @e1            # Get element text
agent-browser get url                 # Get current URL
agent-browser get title               # Get page title

# Wait
agent-browser wait @e1                # Wait for element
agent-browser wait --load networkidle # Wait for network idle
agent-browser wait --url "**/page"    # Wait for URL pattern
agent-browser wait 2000               # Wait milliseconds

# Downloads
agent-browser download @e1 ./file.pdf          # Click element to trigger download
agent-browser wait --download ./output.zip     # Wait for any download to complete
agent-browser --download-path ./downloads open <url>  # Set default download directory

# Capture
agent-browser screenshot              # Screenshot to temp dir
agent-browser screenshot --full       # Full page screenshot
agent-browser screenshot --annotate   # Annotated screenshot with numbered element labels
agent-browser pdf output.pdf          # Save as PDF

# Diff (compare page states)
agent-browser diff snapshot                          # Compare current vs last snapshot
agent-browser diff snapshot --baseline before.txt    # Compare current vs saved file
agent-browser diff screenshot --baseline before.png  # Visual pixel diff
agent-browser diff url <url1> <url2>                 # Compare two pages
agent-browser diff url <url1> <url2> --wait-until networkidle  # Custom wait strategy
agent-browser diff url <url1> <url2> --selector "#main"  # Scope to element

Common Patterns

Form Submission
bash
agent-browser open https://example.com/signup
agent-browser snapshot -i
agent-browser fill @e1 "Jane Doe"
agent-browser fill @e2 "jane@example.com"
agent-browser select @e3 "California"
agent-browser check @e4
agent-browser click @e5
agent-browser wait --load networkidle
Authentication with State Persistence
bash
# Login once and save state
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "$USERNAME"
agent-browser fill @e2 "$PASSWORD"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Reuse in future sessions
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard
Data Extraction
bash
agent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e5           # Get specific element text
agent-browser get text body > page.txt  # Get all page text

# JSON output for parsing
agent-browser snapshot -i --json
agent-browser get text @e1 --json

Ref Lifecycle (Important)

Refs (@e1, @e2, etc.) are invalidated when the page changes. Always re-snapshot after:

  • Clicking links or buttons that navigate
  • Form submissions
  • Dynamic content loading (dropdowns, modals)
bash
agent-browser click @e5              # Navigates to new page
agent-browser snapshot -i            # MUST re-snapshot
agent-browser click @e1              # Use new refs

Annotated Screenshots (Vision Mode)

Use --annotate to take a screenshot with numbered labels overlaid on interactive elements.

bash
agent-browser screenshot --annotate
# Output includes the image path and a legend:
#   [1] @e1 button "Submit"
#   [2] @e2 link "Home"
#   [3] @e3 textbox "Email"
agent-browser click @e2              # Click using ref from annotated screenshot

JavaScript Evaluation

bash
# Simple expressions
agent-browser eval 'document.title'

# Complex JS: use --stdin with heredoc (RECOMMENDED)
agent-browser eval --stdin <<'EVALEOF'
JSON.stringify(
  Array.from(document.querySelectorAll("img"))
    .filter(i => !i.alt)
    .map(i => ({ src: i.src.split("/").pop(), width: i.width }))
)
EVALEOF

Session Management and Cleanup

Always close your browser session when done:

bash
agent-browser close                    # Close default session
agent-browser --session agent1 close   # Close specific session

© slopus, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/agent-browser of slopus/happy.

Open the folder on GitHubat commit 7ea7017

Compare with similar skills

Agent Browser CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser CLI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser CLI this skillslopus/happy24k—~1.7kAutomated safety check: PassMIT
Dev Browser AutomationMemTensor/MemOS12k3 repos~1.7kAutomated safety check: PassApache-2.0
Core Guide for agent-browservercel-labs/agent-browser44k4 repos~9.5kAutomated safety check: PassApache-2.0
Browser Control with Omowrightcode-yeongyu/oh-my-openagent70k—~2.2kAutomated safety check: PassCustom licence
Agent Browsernanocoai/nanoclaw31k3 repos~1.6kAutomated safety check: PassMIT
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone

Similar skills

  • Dev Browser Automation

    MemTensor/MemOS

    Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.

    12k GitHub starsUsed in 3 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Core Guide for agent-browser

    vercel-labs/agent-browser

    Official

    Core usage guide for the agent-browser CLI: the snapshot-and-ref workflow for navigating, clicking, filling forms, extracting data and running parallel sessions.

    44k GitHub starsUsed in 4 repos~9.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser Control with Omowright

    code-yeongyu/oh-my-openagent

    Drives a real browser through the omowright library, either the user's own signed-in browser or a separate browser the code launches, for forms, QA, screenshots and scraping.

    70k GitHub stars~2.2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Agent Browser

    nanocoai/nanoclaw

    Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots.

    31k GitHub starsUsed in 3 repos~1.6k tokens
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Browser Automation

    ynulihao/AgentSkillOS

    Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation.

    617 GitHub stars~2.2k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed

More from slopus/happy

All 8 skills in this repo
  • Searches past Claude Code, Codex and Cursor sessions and summarizes what was worked on, tried or decided, using extraction scripts instead of reading raw logs.

    24k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Traces how an action moves through your code and draws it as a compact ASCII tree: functions called, payload types, state changes and components that re-render.

    24k GitHub stars~654 tokensUpdated today
    Auto-check passed
  • Local development guide for the Happy pnpm monorepo: install, build, test and run the CLI, server, Expo app and Tauri desktop packages.

    24k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Tests interactive CLI and TUI programs with Microsoft's tui-test, driving prompts, arrow keys and screen output in a real pseudo-terminal.

    24k GitHub stars~603 tokensUpdated today
    Auto-check passed
  • Walks you through releasing a component of the Happy monorepo (CLI, mobile, web or server) from a clean local main that matches origin/main.

    24k GitHub stars~6.6k tokensUpdated today
    Auto-check passed

Questions about Agent Browser CLI

What does Agent Browser CLI do?

Drives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows. The skill teaches the agent to drive a real browser with the `agent-browser` CLI. Every task follows the same loop: open a URL, take a snapshot with `snapshot -i` that lists interactive elements with short refs such as `@e1`, interact with those refs by clicking, filling or selecting, and snapshot again whenever the page changes, because refs are invalidated by navigation, form submissions and dynamically loaded content.

When should I use Agent Browser CLI?

Agent Browser CLI fits situations like: testing a web flow in a real browser; filling and submitting a form through the CLI; logging in once and reusing the saved session state; pulling text out of page elements.

How do I install Agent Browser CLI in Claude Code?

Run `npx skills add slopus/happy --skill agent-browser -a claude-code`. Or copy the skill folder (.agents/skills/agent-browser in slopus/happy) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser CLI in Codex?

Run `npx skills add slopus/happy --skill agent-browser -a codex`. Or copy the skill folder (.agents/skills/agent-browser in slopus/happy) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser CLI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slopus/happy --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser CLI need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Browser CLI is instructions for the agent only. Our summary lists: The agent-browser CLI, run directly or through npx. Its frontmatter pre-approves these tools: Bash(npx agent-browser:*), Bash(agent-browser:*).

Does Agent Browser CLI access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Browser CLI safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser CLI use?

Agent Browser CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser CLI use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser CLI?

Skills that share tags, products or a category with Agent Browser CLI: Dev Browser Automation (MemTensor/MemOS, 12k stars), Core Guide for agent-browser (vercel-labs/agent-browser, 44k stars), Browser Control with Omowright (code-yeongyu/oh-my-openagent, 70k stars) and Agent Browser (nanocoai/nanoclaw, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser CLI?

slopus (a GitHub organization) maintains it in slopus/happy, which has 24,039 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: slopus/happy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.