Agent skill

Agent Browser Automation

by withkynam in withkynam/vibecode-pro-max-kit

Drives a browser through the agent-browser CLI, using compact snapshots with element refs to keep context small in long sessions, plus video recording and cloud browsers.

Apache-2.0Auto-check passedProductivity & Automation

Install Agent Browser Automation

skills CLI
$ npx skills add withkynam/vibecode-pro-max-kit --skill vc-agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install withkynam/vibecode-pro-max-kit vc-agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/withkynam/vibecode-pro-max-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/vc-agent-browser .claude/skills/vc-agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vc-agent-browser
GitHub stars
1.1k
Token cost
~2.6k tokens
SKILL.md length
329 words
Files
93 (incl. scripts, references)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Drives a browser through the agent-browser CLI, using compact snapshots with element refs to keep context small in long sessions, plus video recording and cloud browsers.

  • Running a long autonomous browser session without flooding the context
  • SKILL.md covers Quick Start, Core Workflow, Project-Specific Setup and When to Use (vs chrome-devtools), plus 7 more sections
  • Calls npm; needs BROWSERBASE_API_KEY
  • Self-verifying a build by clicking through the running app

What it does

agent-browser is a command-line browser driver built for AI agents around a snapshot-and-refs loop: open a page, take a snapshot of the accessibility tree (interactive-only and compact modes exist), then act on elements by their @e refs with commands such as click, fill, get text and is visible. The skill claims its snapshots are far smaller than those of Playwright MCP, which keeps long sessions within context limits.

A comparison table says when to pick it over chrome-devtools: long autonomous runs, context-constrained workflows, video recording, Browserbase cloud browsers, multi-tab work and self-verifying build loops. Quick screenshots, custom Puppeteer scripts and frame-level debugging fit chrome-devtools better. The skill stays a generic tool reference; project-specific connection patterns and testing policy are left to the consuming repo, and reference files cover Browserbase setup, CDP domains, performance and Puppeteer.

When your agent uses it

  • Running a long autonomous browser session without flooding the context
  • Self-verifying a build by clicking through the running app
  • Recording a video of a browser session for debugging
  • Testing on a cloud browser through Browserbase
  • Working across several tabs in one task

Example prompts

  • “Open https://example.com with agent-browser, take an interactive snapshot and click the sign-in link.”
  • “Fill in the signup form on the staging site and confirm the success message is visible.”
  • “Record a video of the checkout flow so I can see where it fails.”

Requirements

  • Node.js and npm to install agent-browser globally
  • A Browserbase account for cloud browsers, if you use them

What it can do on your machine

Read from SKILL.md and the folder at commit 3bcb2f9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • docs.browserbase.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BROWSERBASE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser Automation loads about 2.6k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 329 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from withkynam/vibecode-pro-max-kit at commit 3bcb2f9, republished under its Apache-2.0 licence (© withkynam). 329 words, ~2,575 tokens.

Download SKILL.mdSave it as .claude/skills/vc-agent-browser/SKILL.md (or your agent's skills folder). This skill also uses 92 other files; get the full folder from GitHub.
name
vc-agent-browser
description
AI-optimized browser automation CLI with context-efficient snapshots. Use for long autonomous sessions, self-verifying workflows, video recording, and cloud browser testing (Browserbase).
license
Apache-2.0
argument-hint
[url or task]
trigger_keywords
browser, screenshot, scrape, automation, web automation, agent-browser, browserbase, cloud browser, headless, playwright, snapshot
layer
helper
metadata.author
vibecode-pro-max-kit
metadata.version
1.0.0

agent-browser Skill

Browser automation CLI designed for AI agents. Uses "snapshot + refs" paradigm for 93% less context than Playwright MCP.

Quick Start

bash
# Install globally
npm install -g agent-browser

# Download Chromium (one-time)
agent-browser install

# Linux: include system deps
agent-browser install --with-deps

# Verify
agent-browser --version

Core Workflow

The 4-step pattern for all browser automation:

bash
# 1. Navigate
agent-browser open https://example.com

# 2. Snapshot (get interactive elements with refs)
agent-browser snapshot -i
# Output: button "Sign In" @e1, textbox "Email" @e2, ...

# 3. Interact using refs
agent-browser fill @e2 "user@example.com"
agent-browser click @e1

# 4. Re-snapshot after page changes
agent-browser snapshot -i

Project-Specific Setup

For project-specific connection patterns, logged-in session reuse through chrome-debug, and when to use agent-browser vs chrome-devtools vs direct non-browser verification, see the project's browser-automation testing notes in the consuming repo, if present. This skill file stays a generic tool reference only and should not redefine the project's broader testing policy.


When to Use (vs chrome-devtools)

Use agent-browserUse chrome-devtools
Long autonomous AI sessionsQuick one-off screenshots
Context-constrained workflowsCustom Puppeteer scripts needed
Video recording for debuggingWebSocket full frame debugging
Cloud browsers (Browserbase)Existing workflow integration
Multi-tab handlingNeed Sharp auto-compression
Self-verifying build loopsSession with auth injection

Token efficiency: ~280 chars/snapshot vs 8K+ for Playwright MCP.

Command Reference

Navigation
bash
agent-browser open <url>       # Navigate to URL
agent-browser back             # Go back
agent-browser forward          # Go forward
agent-browser reload           # Reload page
agent-browser close            # Close browser
Analysis (Snapshot)
bash
agent-browser snapshot         # Full accessibility tree
agent-browser snapshot -i      # Interactive elements only (recommended)
agent-browser snapshot -c      # Compact output
agent-browser snapshot -d 3    # Limit depth
agent-browser snapshot -s "nav" # Scope to CSS selector
Interactions (use @refs from snapshot)
bash
agent-browser click @e1        # Click element
agent-browser dblclick @e1     # Double-click
agent-browser fill @e2 "text"  # Clear and fill input
agent-browser type @e2 "text"  # Type without clearing
agent-browser press Enter      # Press key
agent-browser hover @e1        # Hover over element
agent-browser check @e3        # Check checkbox
agent-browser uncheck @e3      # Uncheck checkbox
agent-browser select @e4 "opt" # Select dropdown option
agent-browser scroll @e1       # Scroll element into view
agent-browser scroll down 500  # Scroll page by pixels
agent-browser drag @e1 @e2     # Drag from e1 to e2
agent-browser upload @e5 file.pdf  # Upload file
Information Retrieval
bash
agent-browser get text @e1     # Get text content
agent-browser get html @e1     # Get HTML
agent-browser get value @e2    # Get input value
agent-browser get attr @e1 href  # Get attribute
agent-browser get title        # Page title
agent-browser get url          # Current URL
agent-browser get count "li"   # Count elements
agent-browser get box @e1      # Bounding box
State Checks
bash
agent-browser is visible @e1   # Check visibility
agent-browser is enabled @e1   # Check if enabled
agent-browser is checked @e3   # Check if checked
Media
bash
agent-browser screenshot           # Capture viewport
agent-browser screenshot --full    # Full page
agent-browser screenshot -o ss.png # Save to file
agent-browser pdf -o page.pdf      # Export PDF
agent-browser record start         # Start video recording
agent-browser record stop          # Stop and save video
agent-browser record restart       # Restart recording
Wait Conditions
bash
agent-browser wait @e1                    # Wait for element
agent-browser wait --text "Success"       # Wait for text to appear
agent-browser wait --url "/dashboard"     # Wait for URL pattern
agent-browser wait --load                 # Wait for page load
agent-browser wait --idle                 # Wait for network idle
agent-browser wait --fn "() => window.ready"  # Wait for JS condition
Browser Configuration
bash
agent-browser viewport 1920 1080   # Set viewport size
agent-browser device "iPhone 14"   # Emulate device
agent-browser geolocation 40.7 -74.0  # Set geolocation
agent-browser offline true         # Enable offline mode
agent-browser headers '{"X-Custom":"val"}'  # Set headers
agent-browser credentials user pass  # HTTP auth
agent-browser color-scheme dark    # Set color scheme
Storage Management
bash
agent-browser cookies              # List cookies
agent-browser cookies set name=val # Set cookie
agent-browser cookies clear        # Clear cookies
agent-browser storage local        # Get localStorage
agent-browser storage session      # Get sessionStorage
agent-browser state save auth.json # Save browser state
agent-browser state load auth.json # Load browser state
Network Control
bash
agent-browser network route "**/*.jpg" --abort    # Block requests
agent-browser network route "**/api/*" --body '{"data":[]}'  # Mock response
agent-browser network unroute "**/*.jpg"          # Remove specific route
agent-browser network requests                    # List intercepted requests
Semantic Finding
bash
agent-browser find role button           # Find by ARIA role
agent-browser find text "Submit"         # Find by text content
agent-browser find label "Email"         # Find by label
agent-browser find placeholder "Search"  # Find by placeholder
agent-browser find testid "login-btn"    # Find by data-testid
agent-browser find first "button"        # First matching element
agent-browser find last "li"             # Last matching element
agent-browser find nth 2 "li"            # Nth element (0-indexed)
Advanced
bash
agent-browser tabs                 # List tabs
agent-browser tab new              # New tab
agent-browser tab 2                # Switch to tab
agent-browser tab close            # Close current tab
agent-browser frame 0              # Switch to frame
agent-browser dialog accept        # Accept dialog
agent-browser dialog dismiss       # Dismiss dialog
agent-browser eval "document.title"  # Execute JS
agent-browser highlight @e1        # Highlight element visually
agent-browser mouse move 100 200   # Move mouse to coordinates
agent-browser mouse down           # Mouse button down
agent-browser mouse up             # Mouse button up

Global Options

OptionDescription
--session <name>Named session for parallel testing
--jsonJSON output for parsing
--headedShow browser window
--cdp <port>Connect via Chrome DevTools Protocol
-p <provider>Cloud browser provider
--proxy <url>Proxy server
--headers <json>Custom HTTP headers
--executable-pathCustom browser binary
--extension <path>Load browser extension

Environment Variables

VariableDescription
AGENT_BROWSER_SESSIONDefault session name
AGENT_BROWSER_PROVIDERCloud provider (e.g., browserbase)
AGENT_BROWSER_EXECUTABLE_PATHBrowser binary location
AGENT_BROWSER_EXTENSIONSComma-separated extension paths
AGENT_BROWSER_STREAM_PORTWebSocket streaming port
AGENT_BROWSER_HOMECustom installation directory
AGENT_BROWSER_PROFILEBrowser profile directory
BROWSERBASE_API_KEYBrowserbase API key
BROWSERBASE_PROJECT_IDBrowserbase project ID

Common Patterns

Form Submission
bash
agent-browser open https://example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3  # Submit button
agent-browser wait url "/dashboard"
State Persistence (Auth)
bash
# Save authenticated state
agent-browser open https://example.com/login
# ... login steps ...
agent-browser state save auth.json

# Reuse in future sessions
agent-browser state load auth.json
agent-browser open https://example.com/dashboard
Video Recording (Debugging)
bash
agent-browser open https://example.com
agent-browser record start
# ... perform actions ...
agent-browser record stop  # Saves to recording.webm
Parallel Sessions
bash
# Terminal 1
agent-browser --session test1 open https://example.com

# Terminal 2
agent-browser --session test2 open https://example.com

Cloud Browsers (Browserbase)

For CI/CD or environments without local browser:

bash
# Set credentials
export BROWSERBASE_API_KEY="your-api-key"
export BROWSERBASE_PROJECT_ID="your-project-id"

# Use cloud browser
agent-browser -p browserbase open https://example.com

See references/browserbase-cloud-setup.md for detailed setup.

Troubleshooting

IssueSolution
Command not foundRun npm install -g agent-browser
Chromium missingRun agent-browser install
Linux deps missingRun agent-browser install --with-deps
Session staleClose browser: agent-browser close
Element not foundRe-run snapshot -i after page changes

Resources

© withkynam, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 92 other files (scripts, references) in .claude/skills/vc-agent-browser of withkynam/vibecode-pro-max-kit.

  • SKILL.md
  • references/.gitkeep
  • references/agent-browser-vs-chrome-devtools.md
  • references/browserbase-cloud-setup.md
  • references/cdp-domains.md
  • references/performance-guide.md
  • references/puppeteer-reference.md
  • screenshots/chat-mock-v1.png
  • screenshots/chat-mock-v2.png
  • screenshots/chat-mock-v3-wide.png
  • screenshots/chat-mock-v4.png
  • screenshots/chat-mock-workspace.png
  • screenshots/ui-audit-20260408-v2/01-dashboard-instances-list-single-running.png
  • screenshots/ui-audit-20260408-v2/02-instance-overview-details-connections-actions.png
  • screenshots/ui-audit-20260408-v2/03-instance-settings-top-details-power-secrets-form.png
  • screenshots/ui-audit-20260408-v2/04-instance-settings-scrolled-secrets-dangerzone.png
  • screenshots/ui-audit-20260408-v2/05-agent-chat-full-layout-sidebar-expanded-3panel.png
  • screenshots/ui-audit-20260408-v2/06-agent-chat-sidebar-collapsed-with-e2e-job.png
  • … and 75 more

Open the folder on GitHubat commit 3bcb2f9

Compare with similar skills

Agent Browser Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser Automation this skillwithkynam/vibecode-pro-max-kit1.1k—~2.6kAutomated safety check: PassApache-2.0
Chrome Devtoolseinverne/dotfiles1211 repos~1.6kAutomated safety check: NotesApache-2.0
Electron Devtools Testingankitvgupta/exo496—~2.4kAutomated safety check: PassCustom licence
Agent Browseroxylabs/agent-skills875—~3kAutomated safety check: PassMIT
AI Search Hubminsight-ai-info/AI-Search-Hub1.3k—~1.3kAutomated safety check: PassNone
Superset Browser Controlsuperset-sh/superset15k—~2.9kAutomated safety check: PassCustom licence

Similar skills

  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check: notes
  • Test the Electron app interactively using Chrome DevTools Protocol.

    496 GitHub stars~2.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Agent Browser

    oxylabs/agent-skills

    Connects to Oxylabs remote agent browsers over the Chrome DevTools Protocol (CDP) with Playwright or Puppeteer.

    875 GitHub stars~3k tokensUpdated 7 days ago
    Productivity & AutomationAuto-check passed
  • AI Search Hub

    minsight-ai-info/AI-Search-Hub

    Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.

    1.3k GitHub stars~1.3k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Superset Browser Control

    superset-sh/superset

    Opens pages, takes screenshots, reads the console, clicks and types in the browser panes of a Superset workspace, with Browser Use as a fallback for other browsers.

    15k GitHub stars~2.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Browser Tools

    981377660LMT/algorithm-study

    Interactive browser automation via Chrome DevTools Protocol.

    278 GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed

More from withkynam/vibecode-pro-max-kit

All 32 skills in this repo
  • Library Documentation Seeker

    withkynam/vibecode-pro-max-kit

    Looks up library and framework documentation through Context7 first, with bundled Node scripts as a fallback that fetch and analyze llms.txt files.

    1.1k GitHub starsUsed in 2 repos~1k tokens
    Auto-check: notes
  • Vc Sequential Thinking

    withkynam/vibecode-pro-max-kit

    Apply step-by-step analysis for complex problems with revision capability.

    1.1k GitHub starsUsed in 2 repos~854 tokens
    Auto-check passed
  • Context Routing Audit

    withkynam/vibecode-pro-max-kit

    Audits a project's context routing, skill discoverability and skill wiring by running a chain of validator scripts and fixing whatever they report.

    1.1k GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Active Plan Audit

    withkynam/vibecode-pro-max-kit

    Reviews a codebase's active plan files for staleness and completion, then archives only the ones confirmed done or obsolete against the real code.

    1.1k GitHub stars~757 tokensUpdated 3 mo ago
    Auto-check passed
  • Systematic Debugging and Investigation

    withkynam/vibecode-pro-max-kit

    Forces root-cause investigation before any fix, combining a four-phase debugging method with log, CI and performance investigation techniques and a rule against unverified completion claims.

    1.1k GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Repository Context Generator

    withkynam/vibecode-pro-max-kit

    Generates or refreshes a project's shared repository context file so both Codex and Claude work from the same up-to-date knowledge layer.

    1.1k GitHub stars~817 tokensUpdated 3 mo ago
    Auto-check passed

Questions about Agent Browser Automation

What does Agent Browser Automation do?

Drives a browser through the agent-browser CLI, using compact snapshots with element refs to keep context small in long sessions, plus video recording and cloud browsers. agent-browser is a command-line browser driver built for AI agents around a snapshot-and-refs loop: open a page, take a snapshot of the accessibility tree (interactive-only and compact modes exist), then act on elements by their @e refs with commands such as click, fill, get text and is visible. The skill claims its snapshots are far smaller than those of Playwright MCP, which keeps long sessions within context limits.

When should I use Agent Browser Automation?

Agent Browser Automation fits situations like: running a long autonomous browser session without flooding the context; self-verifying a build by clicking through the running app; recording a video of a browser session for debugging; testing on a cloud browser through Browserbase.

How do I install Agent Browser Automation in Claude Code?

Run `npx skills add withkynam/vibecode-pro-max-kit --skill vc-agent-browser -a claude-code`. Or copy the skill folder (.claude/skills/vc-agent-browser in withkynam/vibecode-pro-max-kit) into .claude/skills/vc-agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser Automation in Codex?

Run `npx skills add withkynam/vibecode-pro-max-kit --skill vc-agent-browser -a codex`. Or copy the skill folder (.claude/skills/vc-agent-browser in withkynam/vibecode-pro-max-kit) into .agents/skills/vc-agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add withkynam/vibecode-pro-max-kit --skill vc-agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vc-agent-browser, .gemini/skills/vc-agent-browser, .github/skills/vc-agent-browser and .opencode/skills/vc-agent-browser in your project.

What does Agent Browser Automation need to run?

Going by SKILL.md and its folder, Agent Browser Automation needs the command-line tools its instructions call (npm) and credentials named BROWSERBASE_API_KEY. Our summary lists: Node.js and npm to install agent-browser globally; A Browserbase account for cloud browsers, if you use them.

Does Agent Browser Automation access the network?

SKILL.md names 2 domains. As links in the text: github.com and docs.browserbase.com. This is read from the text; nothing was executed.

Is Agent Browser Automation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Browser Automation use?

Agent Browser Automation is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser Automation use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.

What are the alternatives to Agent Browser Automation?

Skills that share tags, products or a category with Agent Browser Automation: Chrome Devtools (einverne/dotfiles, 121 stars), Electron Devtools Testing (ankitvgupta/exo, 496 stars), Agent Browser (oxylabs/agent-skills, 875 stars) and AI Search Hub (minsight-ai-info/AI-Search-Hub, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser Automation?

withkynam (a GitHub user) maintains it in withkynam/vibecode-pro-max-kit, which has 1,144 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on June 21, 2026.

Source: withkynam/vibecode-pro-max-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.