Agent skill

Agent Browser

by davekilleen in davekilleen/Dex

Browser automation using Vercel's agent-browser CLI. An agent skill from davekilleen/Dex.

MITAuto-check passedProductivity & Automation

Install Agent Browser

skills CLI
$ npx skills add davekilleen/Dex --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davekilleen/Dex agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davekilleen/Dex.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/plugins/compound-engineering/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
493
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
167 words
Files
1
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Browser automation using Vercel's agent-browser CLI. An agent skill from davekilleen/Dex.

  • Works in 4 steps: Navigate to URL → Snapshot to get interactive elements… → Interact using refs (@e1, @e2, etc.) → …
  • You need to interact with web pages
  • SKILL.md covers Setup Check, Core Workflow, Key Commands and Semantic Locators (Alternative…, plus 4 more sections
  • Calls npm; reaches site1.com and site2.com

What it does

Agent Browser is an agent skill from davekilleen/Dex. Browser automation using Vercel's agent-browser CLI. Use when you need to interact with web pages, fill forms, take screenshots, or scrape data. Alternative to Playwright MCP - uses Bash commands with ref-based element selection. Triggers on "browse website", "fill form", "click button", "take screenshot", "scrape page", "web automation".

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation and Forms and invoices. It works with Model Context Protocol, Playwright, Bash and Vercel. The repository describes itself as: Your AI Chief of Staff — a personal operating system starter kit that adapts to your role. No coding required. The licence is MIT.

When your agent uses it

  • You need to interact with web pages
  • Take screenshots
  • Take screenshot

Example prompts

  • “browse website”
  • “fill form”
  • “click button”
  • “/agent-browser”

Requirements

  • Node.js

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Navigate to URL
  2. Snapshot to get interactive elements with refs
  3. Interact using refs (@e1, @e2, etc.)
  4. Re-snapshot after navigation or DOM changes

What it can do on your machine

Read from SKILL.md and the folder at commit 227f78e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • site1.com
    • site2.com
    • news.ycombinator.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser loads about 1.6k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 167 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davekilleen/Dex at commit 227f78e, republished under its MIT licence (© davekilleen). 167 words, ~1,592 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Browser automation using Vercel's agent-browser CLI. Use when you need to interact with web pages, fill forms, take screenshots, or scrape data. Alternative to Playwright MCP - uses Bash commands with ref-based element selection. Triggers on "browse website", "fill form", "click button", "take screenshot", "scrape page", "web automation".

agent-browser: CLI Browser Automation

Vercel's headless browser automation CLI designed for AI agents. Uses ref-based selection (@e1, @e2) from accessibility snapshots.

Setup Check

bash
# Check installation
command -v agent-browser >/dev/null 2>&1 && echo "Installed" || echo "NOT INSTALLED - run: npm install -g agent-browser && agent-browser install"
Install if needed
bash
npm install -g agent-browser
agent-browser install  # Downloads Chromium

Core Workflow

The snapshot + ref pattern is optimal for LLMs:

  1. Navigate to URL
  2. Snapshot to get interactive elements with refs
  3. Interact using refs (@e1, @e2, etc.)
  4. Re-snapshot after navigation or DOM changes
bash
# Step 1: Open URL
agent-browser open https://example.com

# Step 2: Get interactive elements with refs
agent-browser snapshot -i --json

# Step 3: Interact using refs
agent-browser click @e1
agent-browser fill @e2 "search query"

# Step 4: Re-snapshot after changes
agent-browser snapshot -i

Key Commands

Navigation
bash
agent-browser open <url>       # Navigate to URL
agent-browser back             # Go back
agent-browser forward          # Go forward
agent-browser reload           # Reload page
agent-browser close            # Close browser
Snapshots (Essential for AI)
bash
agent-browser snapshot              # Full accessibility tree
agent-browser snapshot -i           # Interactive elements only (recommended)
agent-browser snapshot -i --json    # JSON output for parsing
agent-browser snapshot -c           # Compact (remove empty elements)
agent-browser snapshot -d 3         # Limit depth
Interactions
bash
agent-browser click @e1                    # Click element
agent-browser dblclick @e1                 # Double-click
agent-browser fill @e1 "text"              # Clear and fill input
agent-browser type @e1 "text"              # Type without clearing
agent-browser press Enter                  # Press key
agent-browser hover @e1                    # Hover element
agent-browser check @e1                    # Check checkbox
agent-browser uncheck @e1                  # Uncheck checkbox
agent-browser select @e1 "option"          # Select dropdown option
agent-browser scroll down 500              # Scroll (up/down/left/right)
agent-browser scrollintoview @e1           # Scroll element into view
Get Information
bash
agent-browser get text @e1          # Get element text
agent-browser get html @e1          # Get element HTML
agent-browser get value @e1         # Get input value
agent-browser get attr href @e1     # Get attribute
agent-browser get title             # Get page title
agent-browser get url               # Get current URL
agent-browser get count "button"    # Count matching elements
Screenshots & PDFs
bash
agent-browser screenshot                      # Viewport screenshot
agent-browser screenshot --full               # Full page
agent-browser screenshot output.png           # Save to file
agent-browser screenshot --full output.png    # Full page to file
agent-browser pdf output.pdf                  # Save as PDF
Wait
bash
agent-browser wait @e1              # Wait for element
agent-browser wait 2000             # Wait milliseconds
agent-browser wait "text"           # Wait for text to appear

Semantic Locators (Alternative to Refs)

bash
agent-browser find role button click --name "Submit"
agent-browser find text "Sign up" click
agent-browser find label "Email" fill "user@example.com"
agent-browser find placeholder "Search..." fill "query"

Sessions (Parallel Browsers)

bash
# Run multiple independent browser sessions
agent-browser --session browser1 open https://site1.com
agent-browser --session browser2 open https://site2.com

# List active sessions
agent-browser session list

Examples

Login Flow
bash
agent-browser open https://app.example.com/login
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Sign in" [ref=e3]
agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait 2000
agent-browser snapshot -i  # Verify logged in
Search and Extract
bash
agent-browser open https://news.ycombinator.com
agent-browser snapshot -i --json
# Parse JSON to find story links
agent-browser get text @e12  # Get headline text
agent-browser click @e12     # Click to open story
Form Filling
bash
agent-browser open https://forms.example.com
agent-browser snapshot -i
agent-browser fill @e1 "John Doe"
agent-browser fill @e2 "john@example.com"
agent-browser select @e3 "United States"
agent-browser check @e4  # Agree to terms
agent-browser click @e5  # Submit button
agent-browser screenshot confirmation.png
Debug Mode
bash
# Run with visible browser window
agent-browser --headed open https://example.com
agent-browser --headed snapshot -i
agent-browser --headed click @e1

JSON Output

Add --json for structured output:

bash
agent-browser snapshot -i --json

Returns:

json
{
  "success": true,
  "data": {
    "refs": {
      "e1": {"name": "Submit", "role": "button"},
      "e2": {"name": "Email", "role": "textbox"}
    },
    "snapshot": "- button \"Submit\" [ref=e1]\n- textbox \"Email\" [ref=e2]"
  }
}

vs Playwright MCP

Featureagent-browser (CLI)Playwright MCP
InterfaceBash commandsMCP tools
SelectionRefs (@e1)Refs (e1)
OutputText/JSONTool responses
ParallelSessionsTabs
Best forQuick automationTool integration

Use agent-browser when:

  • You prefer Bash-based workflows
  • You want simpler CLI commands
  • You need quick one-off automation

Use Playwright MCP when:

  • You need deep MCP tool integration
  • You want tool-based responses
  • You're building complex automation

© davekilleen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/plugins/compound-engineering/skills/agent-browser of davekilleen/Dex.

Open the folder on GitHubat commit 227f78e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in davekilleen/Dex, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Agent Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser this skilldavekilleen/Dex4931 repos~1.6kAutomated safety check: PassMIT
Browser NavigationFactory-AI/factory-plugins111—~2.7kAutomated safety check: PassNone
Browser Useaiskillstore/marketplace430—~1.1kAutomated safety check: PassNone
Playwright MCPCraftOS-dev/CraftBot3921 repos~1kAutomated safety check: PassMIT
Browsing With Playwrightaiskillstore/marketplace430—~1.2kAutomated safety check: PassNone
Autofillinsundial-org/awesome-openclaw-skills663—~1.8kAutomated safety check: PassNone

Similar skills

  • Browser Navigation

    Factory-AI/factory-plugins

    Automate browser interactions for web testing, form filling, screenshots, and data extraction.

    111 GitHub stars~2.7k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Browser Use

    aiskillstore/marketplace

    Browser automation using Playwright MCP. An agent skill from aiskillstore/marketplace.

    430 GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Playwright MCP

    CraftOS-dev/CraftBot

    Browser automation via Playwright MCP server. An agent skill from CraftOS-dev/CraftBot.

    392 GitHub starsUsed in 1 repo~1k tokens
    Testing & QAAuto-check passed
  • Browsing With Playwright

    aiskillstore/marketplace

    Browser automation using Playwright MCP. An agent skill from aiskillstore/marketplace.

    430 GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Autofillin

    sundial-org/awesome-openclaw-skills

    Automated web form filling and file uploading skill with Playwright browser automation.

    663 GitHub stars~1.8k tokensUpdated 7 mo ago
    Documents & OfficeAuto-check passed
  • Sap Sac Test Automation

    secondsky/sap-skills

    SAP Analytics Cloud (SAC) automated testing skill for designing capability-gated browser discovery and deterministic Playwright test suites for SAC stories, dashboards, reports, planning workflows…

    462 GitHub stars~4.2k tokensUpdated 4 days ago
    Testing & QAAuto-check passed

More from davekilleen/Dex

All 61 skills in this repo
  • Dspy Ruby

    davekilleen/Dex

    This skill should be used when working with DSPy.rb, a Ruby framework for building type-safe, composable LLM applications.

    493 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Diff Adopt Profile

    davekilleen/Dex

    Adopt a full published Heydex profile by handle ('set me up like @davekilleen').

    493 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Diff Generate

    davekilleen/Dex

    Package one workflow — how you use Dex for a specific job — into a shareable DexDiff methodology doc.

    493 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Creating Agent Skills

    davekilleen/Dex

    Expert guidance for creating, writing, and refining Claude Code Skills.

    493 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Feedback

    davekilleen/Dex

    Report a Dex bug to the Dex team with zero homework — Dex investigates locally, builds a privacy-safe report, shows it to you (or auto-sends if you've chosen that), and tracks the ticket until it's…

    493 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Dhh Rails Style

    davekilleen/Dex

    This skill should be used when writing Ruby and Rails code in DHH's distinctive 37signals style.

    493 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Questions about Agent Browser

What does Agent Browser do?

Browser automation using Vercel's agent-browser CLI. An agent skill from davekilleen/Dex. Agent Browser is an agent skill from davekilleen/Dex. Browser automation using Vercel's agent-browser CLI.

When should I use Agent Browser?

Agent Browser fits situations like: you need to interact with web pages; take screenshots; take screenshot.

How do I install Agent Browser in Claude Code?

Run `npx skills add davekilleen/Dex --skill agent-browser -a claude-code`. Or copy the skill folder (.claude/plugins/compound-engineering/skills/agent-browser in davekilleen/Dex) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser in Codex?

Run `npx skills add davekilleen/Dex --skill agent-browser -a codex`. Or copy the skill folder (.claude/plugins/compound-engineering/skills/agent-browser in davekilleen/Dex) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davekilleen/Dex --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser need to run?

Going by SKILL.md and its folder, Agent Browser needs the command-line tools its instructions call (npm). Our summary lists: Node.js.

Does Agent Browser access the network?

SKILL.md names 3 domains. In commands or code: site1.com, site2.com and news.ycombinator.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Agent Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser use?

Agent Browser is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser?

Skills that share tags, products or a category with Agent Browser: Browser Navigation (Factory-AI/factory-plugins, 111 stars), Browser Use (aiskillstore/marketplace, 430 stars), Playwright MCP (CraftOS-dev/CraftBot, 392 stars) and Browsing With Playwright (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser?

davekilleen (a GitHub user) maintains it in davekilleen/Dex, which has 493 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 8, 2026.

Source: davekilleen/Dex on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.