Agent skill

Agent Browser CLI

by sipeed in sipeed/picoclaw

Automates a Chrome or Chromium browser through the agent-browser CLI: navigate, fill forms, click, screenshot and extract data using element refs.

MITAuto-check passedProductivity & Automation

Install Agent Browser CLI

skills CLI
$ npx skills add sipeed/picoclaw --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sipeed/picoclaw agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sipeed/picoclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workspace/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
30k
Token cost
~1.1k tokens
SKILL.md length
138 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Automates a Chrome or Chromium browser through the agent-browser CLI: navigate, fill forms, click, screenshot and extract data using element refs.

  • Works in 4 steps: agent-browser open — navigate → agent-browser snapshot -i — get… → Interact using refs — click @e1, fill… → …
  • Navigating a website and filling in and submitting a form
  • SKILL.md covers Core Workflow, Commands, Authentication and Iframes, plus 3 more sections
  • Calls npm; reaches site-a.com and site-b.com

What it does

The skill drives a browser over CDP with the agent-browser command, installed with npm i -g agent-browser and then agent-browser install. It first checks that the tool exists with which. If it is missing, the agent tells you that the CLI and Chromium are only available in the heavy container image and does not try to install anything at runtime.

The core loop is to open a URL, take an interactive snapshot that lists elements with refs such as @e1, interact by ref with click and fill, and take a new snapshot after any navigation or DOM change because refs are invalidated. Commands can be chained with &&, iframe content is inlined in snapshots so no frame switch is needed, parallel sessions are named with a session flag, and eval runs JavaScript with a stdin option for complex code. Sessions should be closed when the work is done.

For sites that need a login, auth state can be saved from your running Chrome using an auto-connect option and loaded again in later sessions with a state flag.

When your agent uses it

  • Navigating a website and filling in and submitting a form
  • Taking screenshots or extracting data from a web page
  • Testing a web app by clicking through it
  • Reusing a logged-in Chrome session for automation

Example prompts

  • “Open https://example.com/form, fill in the email and password fields and tell me what the page says after submit.”
  • “Take a screenshot of the pricing page and extract the plan names.”
  • “Run two sessions in parallel to compare site-a.com and site-b.com.”

Requirements

  • The agent-browser CLI and Chromium, installed with npm i -g agent-browser and agent-browser install

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. agent-browser open — navigate
  2. agent-browser snapshot -i — get interactive elements with refs (@e1, @e2, ...)
  3. Interact using refs — click @e1, fill @e2 "text"
  4. Re-snapshot after any navigation or DOM change — refs are invalidated

What it can do on your machine

Read from SKILL.md and the folder at commit bbf6893. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • site-a.com
    • site-b.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser CLI loads about 1.1k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 138 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sipeed/picoclaw at commit bbf6893, republished under its MIT licence (© sipeed). 138 words, ~1,059 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Browser automation via agent-browser CLI. Use when the user needs to navigate websites, fill forms, click buttons, take screenshots, extract data, or test web apps.

Agent Browser

CLI browser automation via Chrome/Chromium CDP. Install: npm i -g agent-browser && agent-browser install.

Before using this skill, verify the tool is available by running which agent-browser. If the command is not found, tell the user that browser automation requires the agent-browser CLI and Chromium, which are only available in the heavy container image. Do not attempt to install it at runtime.

Core Workflow

  1. agent-browser open <url> — navigate
  2. agent-browser snapshot -i — get interactive elements with refs (@e1, @e2, ...)
  3. Interact using refs — click @e1, fill @e2 "text"
  4. Re-snapshot after any navigation or DOM change — refs are invalidated
bash
agent-browser open https://example.com/form
agent-browser snapshot -i
# @e1 [input] "Email", @e2 [input] "Password", @e3 [button] "Submit"
agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "secret"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i

Chain commands with && when you don't need intermediate output:

bash
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser snapshot -i

Commands

bash
# Navigation
agent-browser open <url>
agent-browser close

# Snapshot
agent-browser snapshot -i                # Interactive elements with refs
agent-browser snapshot -s "#selector"    # Scope to CSS selector

# Interaction (use @refs from snapshot)
agent-browser click @e1
agent-browser fill @e2 "text"            # Clear + type
agent-browser type @e2 "text"            # Type without clearing
agent-browser select @e1 "option"
agent-browser check @e1
agent-browser press Enter
agent-browser scroll down 500

# Get info
agent-browser get text @e1
agent-browser get url
agent-browser get title

# Wait
agent-browser wait @e1                   # Wait for element
agent-browser wait --load networkidle    # Wait for network idle
agent-browser wait --url "**/dashboard"  # Wait for URL pattern
agent-browser wait --text "Welcome"      # Wait for text
agent-browser wait 2000                  # Wait ms

# Capture
agent-browser screenshot                 # Screenshot to temp dir
agent-browser screenshot --full          # Full page
agent-browser screenshot --annotate      # With numbered element labels ([N] -> @eN)
agent-browser pdf output.pdf

# Semantic locators (when refs unavailable)
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"
agent-browser find role button click --name "Submit"

Authentication

bash
# Option 1: Import from user's running Chrome
agent-browser --auto-connect state save ./auth.json
agent-browser --state ./auth.json open https://app.example.com

# Option 2: Persistent profile
agent-browser --profile ~/.myapp open https://app.example.com/login
# ... login once, all future runs are authenticated

# Option 3: Session name (auto-save/restore)
agent-browser --session-name myapp open https://app.example.com/login
# ... login, close, next run state is restored

# Option 4: State file
agent-browser state save auth.json
agent-browser state load auth.json

Iframes

Iframe content is inlined in snapshots. Interact with iframe refs directly — no frame switch needed.

Parallel Sessions

bash
agent-browser --session s1 open https://site-a.com
agent-browser --session s2 open https://site-b.com
agent-browser session list

JavaScript Eval

bash
agent-browser eval 'document.title'

# Complex JS — use --stdin to avoid shell quoting issues
agent-browser eval --stdin <<'EVALEOF'
JSON.stringify(Array.from(document.querySelectorAll("a")).map(a => a.href))
EVALEOF

Cleanup

Always close sessions when done:

bash
agent-browser close
agent-browser --session s1 close

© sipeed, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in workspace/skills/agent-browser of sipeed/picoclaw.

Open the folder on GitHubat commit bbf6893

Compare with similar skills

Agent Browser CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser CLI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser CLI this skillsipeed/picoclaw30k—~1.1kAutomated safety check: PassMIT
Agent Browser Automationwithkynam/vibecode-pro-max-kit1.1k—~2.6kAutomated safety check: PassApache-2.0
Electron Devtools Testingankitvgupta/exo496—~2.4kAutomated safety check: PassCustom licence
Chrome Devtoolseinverne/dotfiles1212 repos~1.6kAutomated safety check: NotesApache-2.0
Anti Detect Browserantibrow/anti-detect-browser-skills932—~9.8kAutomated safety check: WarnMIT
Reverse Browser Automationsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Agent Browser Automation

    withkynam/vibecode-pro-max-kit

    Drives a browser through the agent-browser CLI, using compact snapshots with element refs to keep context small in long sessions, plus video recording and cloud browsers.

    1.1k GitHub stars~2.6k tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed
  • Test the Electron app interactively using Chrome DevTools Protocol.

    496 GitHub stars~2.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 2 repos~1.6k tokens
    Data & AnalyticsAuto-check: notes
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    932 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings
  • Reverse Browser Automation

    sickn33/agentic-awesome-skills

    Automate browsers (Playwright) and Windows desktop applications (UI automation) for reverse-engineering evidence collection, UI-driven workflows, and network observation during analysis.

    47k GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed
  • UI Preview

    tingly-dev/tingly-box

    Capture headless-Chrome screenshots of the tingly-box frontend (running locally in mock mode) so frontend changes can be visually verified in environments without a real browser.

    351 GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check: notes

More from sipeed/picoclaw

  • Picoclaw Skill Creator

    sipeed/picoclaw

    Guidance for creating, updating and reviewing Picoclaw skills, from the SKILL.md structure to organizing bundled scripts, references and assets.

    30k GitHub stars~4.4k tokensUpdated 14 days ago
    Auto-check passed
  • Reads and controls I2C and SPI peripherals on Sipeed boards such as LicheeRV Nano, MaixCAM and NanoKVM through the i2c and spi tools.

    30k GitHub stars~578 tokensUpdated 14 days ago
    Auto-check passed
  • PicoClaw Agent

    sipeed/picoclaw

    Answers questions about running and changing PicoClaw, from onboarding and model selection to MCP server setup, skill loading and scheduled jobs.

    30k GitHub stars~7.2k tokensUpdated 14 days ago
    Auto-check: notes
  • Weather Lookup

    sipeed/picoclaw

    Looks up current weather and forecasts with curl and verified location matching, using wttr.in for city names and Open-Meteo for structured data.

    30k GitHub stars~657 tokensUpdated 14 days ago
    Auto-check passed

Works with

Questions about Agent Browser CLI

What does Agent Browser CLI do?

Automates a Chrome or Chromium browser through the agent-browser CLI: navigate, fill forms, click, screenshot and extract data using element refs. The skill drives a browser over CDP with the agent-browser command, installed with npm i -g agent-browser and then agent-browser install. It first checks that the tool exists with which.

When should I use Agent Browser CLI?

Agent Browser CLI fits situations like: navigating a website and filling in and submitting a form; taking screenshots or extracting data from a web page; testing a web app by clicking through it; reusing a logged-in Chrome session for automation.

How do I install Agent Browser CLI in Claude Code?

Run `npx skills add sipeed/picoclaw --skill agent-browser -a claude-code`. Or copy the skill folder (workspace/skills/agent-browser in sipeed/picoclaw) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser CLI in Codex?

Run `npx skills add sipeed/picoclaw --skill agent-browser -a codex`. Or copy the skill folder (workspace/skills/agent-browser in sipeed/picoclaw) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser CLI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sipeed/picoclaw --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser CLI need to run?

Going by SKILL.md and its folder, Agent Browser CLI needs the command-line tools its instructions call (npm). Our summary lists: The agent-browser CLI and Chromium, installed with npm i -g agent-browser and agent-browser install.

Does Agent Browser CLI access the network?

SKILL.md names 2 domains. In commands or code: site-a.com and site-b.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Agent Browser CLI safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser CLI use?

Agent Browser CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser CLI use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser CLI?

Skills that share tags, products or a category with Agent Browser CLI: Agent Browser Automation (withkynam/vibecode-pro-max-kit, 1.1k stars), Electron Devtools Testing (ankitvgupta/exo, 496 stars), Chrome Devtools (einverne/dotfiles, 121 stars) and Anti Detect Browser (antibrow/anti-detect-browser-skills, 932 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser CLI?

sipeed (a GitHub organization) maintains it in sipeed/picoclaw, which has 30,010 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 24, 2026.

Source: sipeed/picoclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.