Agent skill

Agent Browser

by nanocoai in nanocoai/nanoclaw

Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots.

MITAuto-check passedProductivity & Automation

Install Agent Browser

skills CLI
$ npx skills add nanocoai/nanoclaw --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nanocoai/nanoclaw agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nanocoai/nanoclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/container/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
31k
Used in
3 other repos
Token cost
~1.6k tokens
SKILL.md length
253 words
Files
1
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots.

  • Works in 4 steps: Navigate: agent-browser open → Snapshot: agent-browser snapshot -i… → Interact using refs from the snapshot → …
  • Researching a topic or reading articles that need a live browser
  • SKILL.md covers Quick start, Core workflow, Commands and Example: Form submission, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The agent works from a short loop: open a URL, take a snapshot that lists interactive elements with references such as @e1, act on those references, and snapshot again after navigation or large page changes. Commands cover navigation, clicking, filling fields, reading text, HTML or input values, waiting, and saving screenshots or PDFs.

Snapshots can be limited to interactive elements or a compact tree. The skill also warns against unbounded polling when waiting on a custom page condition, because a loop that never ends can freeze the whole agent turn. The built-in wait commands are preferred, and any fallback loop needs a limit. The allowed tool is Bash restricted to the agent-browser command.

When your agent uses it

  • Researching a topic or reading articles that need a live browser
  • Filling in a web form or clicking through a web app
  • Capturing a full-page screenshot of a site
  • Extracting data from a page that has no API
  • Checking that a web page behaves as expected

Example prompts

  • “Open the pricing page of example.com and save a full-page screenshot to ./pricing.png.”
  • “Fill the contact form on the demo site with a test message and tell me what the confirmation says.”
  • “Read the headline and first paragraph of this article and summarize them.”

Requirements

  • The agent-browser CLI installed
  • Permission to run Bash commands for agent-browser
  • Pre-approved tools (allowed-tools): Bash(agent-browser:*)

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Navigate: agent-browser open
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

What it can do on your machine

Read from SKILL.md and the folder at commit 66f0823. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(agent-browser:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser loads about 1.6k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 253 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nanocoai/nanoclaw at commit 66f0823, republished under its MIT licence (© nanocoai). 253 words, ~1,584 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use whenever a browser would be useful, not just when the user explicitly asks.
allowed-tools
Bash(agent-browser:*)

Browser Automation with agent-browser

Quick start

bash
agent-browser open <url>        # Navigate to page
agent-browser snapshot -i       # Get interactive elements with refs
agent-browser click @e1         # Click element by ref
agent-browser fill @e2 "text"   # Fill input by ref
agent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Commands

Navigation
bash
agent-browser open <url>      # Navigate to URL
agent-browser back            # Go back
agent-browser forward         # Go forward
agent-browser reload          # Reload page
agent-browser close           # Close browser
Snapshot (page analysis)
bash
agent-browser snapshot            # Full accessibility tree
agent-browser snapshot -i         # Interactive elements only (recommended)
agent-browser snapshot -c         # Compact output
agent-browser snapshot -d 3       # Limit depth to 3
agent-browser snapshot -s "#main" # Scope to CSS selector
Interactions (use @refs from snapshot)
bash
agent-browser click @e1           # Click
agent-browser dblclick @e1        # Double-click
agent-browser fill @e2 "text"     # Clear and type
agent-browser type @e2 "text"     # Type without clearing
agent-browser press Enter         # Press key
agent-browser hover @e1           # Hover
agent-browser check @e1           # Check checkbox
agent-browser uncheck @e1         # Uncheck checkbox
agent-browser select @e1 "value"  # Select dropdown option
agent-browser scroll down 500     # Scroll page
agent-browser upload @e1 file.pdf # Upload files
Get information
bash
agent-browser get text @e1        # Get element text
agent-browser get html @e1        # Get innerHTML
agent-browser get value @e1       # Get input value
agent-browser get attr @e1 href   # Get attribute
agent-browser get title           # Get page title
agent-browser get url             # Get current URL
agent-browser get count ".item"   # Count matching elements
Screenshots & PDF
bash
agent-browser screenshot          # Save to temp directory
agent-browser screenshot path.png # Save to specific path
agent-browser screenshot --full   # Full page
agent-browser pdf output.pdf      # Save as PDF
Wait
bash
agent-browser wait @e1                     # Wait for element
agent-browser wait 2000                    # Wait milliseconds
agent-browser wait --text "Success"        # Wait for text
agent-browser wait --url "**/dashboard"    # Wait for URL pattern
agent-browser wait --load networkidle      # Wait for network idle
Waiting for a custom condition — ALWAYS bound it

Prefer the built-in wait subcommands above. Only fall back to eval-polling when you must wait on a custom JS condition (e.g. a spinner disappearing or a "Send" button re-enabling in a chat UI).

Never write an unbounded wait loop. A bare until … do sleep; done that polls a page condition will loop forever if the condition never becomes true (page failed to load, selector changed, network stalled). That does not just fail the command — it wedges the entire agent turn: the runner keeps the model stream open, later messages get silently swallowed, and the container can hang for hours without the host's stuck-detection firing.

Always cap the wait with BOTH a wall-clock timeout and a max-attempts counter, and always exit the loop (never leave a sleep loop as the last thing running):

bash
# Bounded wait: succeeds when the condition is met, gives up after ~90s.
timeout 90 bash -c '
  for i in $(seq 1 30); do
    if agent-browser eval "document.querySelector(\".loading\") === null" 2>/dev/null | grep -q true; then
      echo READY; exit 0
    fi
    sleep 3
  done
  echo TIMEOUT; exit 1
'
# Check the exit status / output: on TIMEOUT, snapshot the page and decide —
# do NOT re-enter another unbounded wait.

If the wait times out, treat it as a real failure: take a snapshot -i or screenshot to see the actual page state, report what you found, and move on. Retrying the same unbounded wait is what causes the hang.

Semantic locators (alternative to refs)
bash
agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"
agent-browser find placeholder "Search" type "query"
Authentication with saved state
bash
# Login once
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "username"
agent-browser fill @e2 "password"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Later: load saved state
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard
Cookies & Storage
bash
agent-browser cookies                     # Get all cookies
agent-browser cookies set name value      # Set cookie
agent-browser cookies clear               # Clear cookies
agent-browser storage local               # Get localStorage
agent-browser storage local set k v       # Set value
JavaScript
bash
agent-browser eval "document.title"   # Run JavaScript

Example: Form submission

bash
agent-browser open https://example.com/form
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result

Example: Data extraction

bash
agent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e1  # Get product title
agent-browser get attr @e2 href  # Get link URL
agent-browser screenshot products.png

© nanocoai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in container/skills/agent-browser of nanocoai/nanoclaw.

Open the folder on GitHubat commit 66f0823

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in nanocoai/nanoclaw, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser this skillnanocoai/nanoclaw31k3 repos~1.6kAutomated safety check: PassMIT
Dev Browser AutomationMemTensor/MemOS12k3 repos~1.7kAutomated safety check: PassApache-2.0
Core Guide for agent-browservercel-labs/agent-browser44k4 repos~9.5kAutomated safety check: PassApache-2.0
Browser Control with Omowrightcode-yeongyu/oh-my-openagent70k—~2.2kAutomated safety check: PassCustom licence
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone
Browser Automationynulihao/AgentSkillOS617—~2.2kAutomated safety check: PassNone

Similar skills

  • Dev Browser Automation

    MemTensor/MemOS

    Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.

    12k GitHub starsUsed in 3 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Core Guide for agent-browser

    vercel-labs/agent-browser

    Official

    Core usage guide for the agent-browser CLI: the snapshot-and-ref workflow for navigating, clicking, filling forms, extracting data and running parallel sessions.

    44k GitHub starsUsed in 4 repos~9.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser Control with Omowright

    code-yeongyu/oh-my-openagent

    Drives a real browser through the omowright library, either the user's own signed-in browser or a separate browser the code launches, for forms, QA, screenshots and scraping.

    70k GitHub stars~2.2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Browser Automation

    ynulihao/AgentSkillOS

    Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation.

    617 GitHub stars~2.2k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Use

    aiskillstore/marketplace

    Browser automation using Playwright MCP. An agent skill from aiskillstore/marketplace.

    430 GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from nanocoai/nanoclaw

All 59 skills in this repo
  • Installs or refreshes Iron Proxy and its Iron Control web console for NanoClaw, with a local Docker setup, database, credentials and a human approval bridge.

    31k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check: notes
  • Installs or refreshes OneCLI as the gateway provider for NanoClaw, copying the adapter files, registering the provider and running the setup script.

    31k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Guides a conversational migration from an OpenClaw install to NanoClaw v2, carrying over identity, channel credentials, scheduled tasks and workspace files.

    31k GitHub stars~6k tokensUpdated yesterday
    Auto-check: notes
  • Wires up an additional phone number onto an already-installed Dial channel, so one NanoClaw install answers SMS and AI voice calls on more than one line.

    31k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • NanoClaw LLM Wiki Setup

    nanocoai/nanoclaw

    Adds a persistent wiki knowledge base to a NanoClaw group following Karpathy's LLM Wiki pattern, with folders, a tailored container skill and a CLAUDE.md section.

    31k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • NanoClaw Ollama Provider

    nanocoai/nanoclaw

    Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts.

    31k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed

Questions about Agent Browser

What does Agent Browser do?

Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots. The agent works from a short loop: open a URL, take a snapshot that lists interactive elements with references such as @e1, act on those references, and snapshot again after navigation or large page changes. Commands cover navigation, clicking, filling fields, reading text, HTML or input values, waiting, and saving screenshots or PDFs.

When should I use Agent Browser?

Agent Browser fits situations like: researching a topic or reading articles that need a live browser; filling in a web form or clicking through a web app; capturing a full-page screenshot of a site; extracting data from a page that has no API.

How do I install Agent Browser in Claude Code?

Run `npx skills add nanocoai/nanoclaw --skill agent-browser -a claude-code`. Or copy the skill folder (container/skills/agent-browser in nanocoai/nanoclaw) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser in Codex?

Run `npx skills add nanocoai/nanoclaw --skill agent-browser -a codex`. Or copy the skill folder (container/skills/agent-browser in nanocoai/nanoclaw) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nanocoai/nanoclaw --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Browser is instructions for the agent only. Our summary lists: The agent-browser CLI installed; Permission to run Bash commands for agent-browser. Its frontmatter pre-approves these tools: Bash(agent-browser:*).

Does Agent Browser access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser use?

Agent Browser is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser?

Skills that share tags, products or a category with Agent Browser: Dev Browser Automation (MemTensor/MemOS, 12k stars), Core Guide for agent-browser (vercel-labs/agent-browser, 44k stars), Browser Control with Omowright (code-yeongyu/oh-my-openagent, 70k stars) and Playwright Bowser (disler/bowser, 265 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser?

nanocoai (a GitHub organization) maintains it in nanocoai/nanoclaw, which has 30,883 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on October 6, 2026.

Source: nanocoai/nanoclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.