Agent skill

Browser Automation

by openclaw in openclaw/openclaw

A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

MITAuto-check passedProductivity & Automation

Install Browser Automation

skills CLI
$ npx skills add openclaw/openclaw --skill browser-automation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openclaw/openclaw browser-automation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openclaw/openclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/extensions/browser/skills/browser-automation .claude/skills/browser-automation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-automation
GitHub stars
392k
Token cost
~2.9k tokens
SKILL.md length
1,460 words
Files
1
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

  • Works in 5 steps: Check browser state before acting → Prefer stable tab handles → Read before you click → …
  • Controlling web pages with the OpenClaw browser tool
  • SKILL.md covers Operating Loop, Browser batch CLI, Code Mode Loop and Tab Hygiene, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Browser Automation is an agent skill from openclaw/openclaw. Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation. The repository describes itself as: The AI that really does things. Any OS. Any Platform. The lobster way. 🦞. The licence is MIT.

When your agent uses it

  • Controlling web pages with the OpenClaw browser tool
  • Especially multi-step flows
  • Recovery from stale refs/timeouts

Example prompts

  • “/browser-automation”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Check browser state before acting
  2. Prefer stable tab handles
  3. Read before you click
  4. Act narrowly
  5. Report real blockers

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb5970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Automation loads about 2.9k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 1,460 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openclaw/openclaw at commit 1eb5970, republished under its MIT licence (© openclaw). 1,460 words, ~2,903 tokens.

Download SKILL.mdSave it as .claude/skills/browser-automation/SKILL.md (or your agent's skills folder).
name
browser-automation
description
Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.
user-invocable
false

Browser Automation

Use this skill when you need the browser tool for anything beyond a single page check.

Operating Loop

  1. Check browser state before acting:
    • openclaw browser doctor or action="status" when the browser/plugin setup itself may be broken.
    • action="status" for availability.
    • action="profiles" if login state or profile choice matters.
    • action="tabs" before opening a new tab if retries/timeouts may have left windows behind.
  2. Prefer stable tab handles:
    • Open important tabs with label, for example label="meet".
    • After action="tabs" or action="open", store suggestedTargetId and pass it as targetId in later calls.
    • suggestedTargetId is the label when one exists, otherwise the stable tabId handle like t1.
    • Avoid relying on raw DevTools targetId except for immediate diagnostics; it can change under Chromium target replacement.
  3. Read before you click:
    • For “read the page and answer X,” use action="text" with optional selector and maxChars for bounded visible prose (first selector match, otherwise article/main/body). On existing-session profiles, use snapshot instead. Efficient snapshots omit most prose.
    • For virtualized lists, scroll through each segment, capture only the relevant rows, then merge the results.
    • Use action="snapshot" on the intended targetId.
    • Add snapshot query to find lines containing all query tokens, ignoring case; matching lines keep their refs.
    • Use the same targetId for follow-up actions so refs stay on the same tab.
    • For durable Playwright refs, request refs="aria" when supported. If you receive axN refs from snapshotFormat="aria", use them only after that same snapshot call; stale or unbound axN refs fail fast and need a fresh snapshot.
    • Use urls=true when link text is ambiguous or a direct navigation target would avoid brittle clicks.
    • Use labels=true on snapshot or screenshot when visual position matters. On Playwright-backed profiles, the response includes an annotations array ({ref, number, role, name?, box}) with each ref's bounding box in the captured image's coordinate space, so you can reason about position without re-snapshotting; screenshot labels can also combine with fullPage=true (CLI: --full-page) to label the whole document, or ref / element to clip to one element. profile="user" and other existing-session (chrome-mcp) profiles render an overlay into page screenshots but do not attach annotations or use the Playwright full-page/ref/element projection helper, so read positions from the labeled image itself on those profiles. The raw-CDP fallback (no Playwright) does not support labeled screenshots at all and returns a 501, so only request labels when Playwright is available.
  4. Act narrowly:
    • Prefer action="act" with a ref from the latest snapshot.
    • navigate returns the loaded page's compact snapshot inline, and batch act results that report a cross-document navigation include fresh page state; use those refs directly instead of a follow-up snapshot call.
    • After a single act that triggers navigation, and after modal changes or form submissions, snapshot again before the next action.
    • Avoid blind waits. Wait for visible UI state when possible.
    • Use action="emulate" with device, colorScheme, timezoneId, or locale when testing those settings; snapshot again afterward. Existing-session profiles do not support emulation.
  5. Report real blockers:
    • Debug network failures with action="requests", optional URL/type filter, and limit (default 50 recent entries). clear=true clears the collected log after reading. Use a managed profile; existing-session profiles do not support this log.
    • Debug page errors with action="errors" and limit (default 50 recent entries). clear=true clears the collected log after reading. Existing-session profiles do not support this log.
    • If the page needs login, permission, captcha, 2FA, camera/microphone approval, or another manual step, stop and tell the user exactly what is needed.
    • Do not claim the browser is not logged in just because the current page shows a permission or onboarding dialog. Inspect the visible UI first.

Browser batch CLI

openclaw browser batch runs an array of nested /act actions in one /act call (the same kind="batch" runtime reached through the agent tool), so CLI users and scripts can combine actions like wait, click, type, and evaluate into a single replayable plan without per-action round trips. Each entry in actions[] is a BrowserActRequest — the closed union the /act route accepts — not arbitrary openclaw browser subcommands. batch is not supported on profile="user" and other existing-session (chrome-mcp) profiles; send actions individually there.

  • CLI: openclaw browser batch --actions '<json>', --actions-file plan.json, or --actions-file - for stdin. --actions-file and stdin input are capped at 1,000,000 bytes; split larger plans into multiple batch commands. --continue sets stopOnError=false; default stops on first error.
  • Ref lifecycle: refs come from a snapshot run before the batch (snapshot is not a nested action). A nested action that changes page state — such as a click that triggers navigation, or an evaluate that mutates the DOM — can invalidate earlier refs for the rest of the batch; put state-changing actions first, or split into a follow-up batch after re-snapshotting. Navigation and re-snapshotting happen outside the batch, since open, navigate, and snapshot are not /act kinds.
  • Target id: nested actions share the request's tab; an explicit nested targetId that resolves to a different tab is rejected with ACT_TARGET_ID_MISMATCH.
  • Response: { "results": [{ "ok": true } | { "ok": false, "error": "..." }, ...] } in order; with default stopOnError the array ends at the first failure. Any failed entry exits nonzero; use --json to preserve the full response in scripts.
Show full SKILL.md (619 more words)Show less

Code Mode Loop

When tools.codeMode is enabled, the Browser tool has no normal turn — it is cataloged behind exec/wait. Call it from exec cells as an async global, using the callable name the exec quick index advertises for the Browser tool (normally browser; colliding names get suffixed, and a client tool can win an identical name). An exact catalog.search("browser") returns a handle already bound to the effective callable name, so resolve the handle in each cell and call it instead of hard-coding the literal global; an empty result means the Browser tool is not cataloged in this run.

Keep the same labeled tab through the loop, and alternate reads with actions. Each exec cell starts a fresh VM — bindings from a completed cell are gone in the next, and only runs left waiting keep their state until wait resumes them — so carry comparison state across cells by returning it and re-embedding the returned values in the next cell:

javascript
// previous = the url/newElements returned by the last completed cell (a fresh
// VM runs this cell, so prior bindings do not exist here).
const previous = { url: "https://example.com/inbox", newElements: 0 };
const [browser] = await catalog.search("browser", { limit: 1 });
const details = await browser({
  action: "snapshot",
  snapshotFormat: "ai",
  targetId: "task",
  refs: "aria",
  interactive: true,
});
const changed =
  details?.url !== previous.url ||
  (details?.newElements ?? 0) > 0 ||
  details?.blockedByDialog === true;
return {
  targetId: details?.targetId,
  url: details?.url,
  newElements: details?.newElements,
  stats: details?.stats,
  changed,
};
  • Code-mode calls return the tool's structured details directly (targetId, url, newElements, stats, blockedByDialog); rendered page text is not returned to code cells.
  • To read text inside code mode, run a targeted act evaluate (requires the evaluate capability; browser.evaluateEnabled can disable it) and keep the returned value bounded, because page-script output is untrusted:
javascript
const [browser] = await catalog.search("browser", { limit: 1 });
const read = await browser({
  action: "act",
  kind: "evaluate",
  fn: "() => document.body.innerText.slice(0, 2000)",
  targetId: "task",
});
return { url: read?.url, text: read?.result };

When evaluate is unavailable, keep the loop on structured state only.

  • Return only the fields the next step needs; never return the whole details object.
  • Completed cells share no state: re-embed the previous cell's returned url/newElements in the next cell, or keep the comparison inside one cell. Only waiting runs persist, resumed by wait.
  • Interleave each act with a URL or tabs check before the next dependent act.
  • If a batch returns aborted, take a fresh snapshot before continuing.
  • If newElements is positive, inspect those elements first, then update the re-embedded state.
  • Use separate act calls when navigation is expected between steps.

Tab Hygiene

Before creating a tab for a named task, list tabs and reuse an existing matching label or URL when it is still usable.

Example:

json
{ "action": "tabs" }

If no suitable tab exists:

json
{ "action": "open", "url": "https://example.com", "label": "task" }

Then target it by label:

json
{ "action": "snapshot", "targetId": "task", "refs": "aria" }

If a retry creates duplicates, close the extras by tabId:

json
{ "action": "close", "targetId": "t3" }

Do not pass bare numbers like "2" as targetId. Numeric tab positions are only for the CLI openclaw browser tab select 2 helper; browser tool calls need a suggestedTargetId, label, tabId, or raw target id.

Stale Ref Recovery

If an action fails with a missing or stale ref:

  1. Snapshot the same targetId again.
  2. Find the current visible control.
  3. Retry once with the new ref.
  4. If the UI moved to a blocker state, report the blocker instead of looping.

Existing User Browser

Use profile="user" only when existing cookies/login matter. This attaches to the user's running Chromium-based browser.

On macOS, action="importprofile" is the alternative when the agent should use an isolated managed browser with cookies copied from a real Chrome-family profile. First use action="profiles" and inspect systemProfiles, then import into a fresh managed profile name. Import asks for one Keychain/Touch ID consent prompt. It copies cookies, not local storage or IndexedDB; device-bound session credentials (DBSC) mean some Google sessions may still require re-authentication.

For profile="user" and other existing-session profiles, omit timeoutMs on act:type, hover, scrollIntoView, drag, select, and fill; that driver rejects per-call timeout overrides for those actions. act:evaluate accepts timeoutMs.

Google Meet Notes

When creating or joining a Meet:

  • Treat camera/microphone permission screens as progress, not login failure.
  • If asked whether people can hear you, click the microphone option when voice is required.
  • If Google asks for sign-in, 2FA, account chooser confirmation, or permission that needs user approval, report the exact manual action.
  • Use one labeled tab per meeting flow, for example label="meet", and reuse it during retries.

© openclaw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in extensions/browser/skills/browser-automation of openclaw/openclaw.

Open the folder on GitHubat commit 1eb5970

Compare with similar skills

Browser Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Automation this skillopenclaw/openclaw392k—~2.9kAutomated safety check: PassMIT
Agent Browserquran/quran.com-frontend-next1.9k42 repos~3.3kAutomated safety check: PassNone
Dev-Browser CLI AutomationSawyerHood/dev-browser6.7k1 repos~455Automated safety check: PassMIT
Agent Browsersuperagent-ai/grok-cli3.5k1 repos~633Automated safety check: PassMIT
Camoufox CLIBin-Huang/camoufox-cli3501 repos~4.5kAutomated safety check: PassMIT
BrowserVibiumDev/vibium2.9k—~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 42 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Dev-Browser CLI Automation

    SawyerHood/dev-browser

    Browser automation with persistent named pages via the dev-browser CLI. Use when users ask to navigate websites, fill forms, take screenshots, extract web…

    6.7k GitHub starsUsed in 1 repo~455 tokens
    Productivity & AutomationAuto-check passed
  • Agent Browser

    superagent-ai/grok-cli

    Use the host-side agent-browser CLI for local browser smoke tests, screenshots, snapshots, and simple UI validation against forwarded localhost URLs.

    3.5k GitHub starsUsed in 1 repo~633 tokens
    Productivity & AutomationAuto-check passed
  • Camoufox CLI

    Bin-Huang/camoufox-cli

    Anti-detect browser automation CLI & Skills for AI agents. An agent skill from Bin-Huang/camoufox-cli.

    350 GitHub starsUsed in 1 repo~4.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser

    VibiumDev/vibium

    Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.

    2.9k GitHub stars~4.8k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts.

    50k GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from openclaw/openclaw

All 93 skills in this repo
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Tmux

    openclaw/openclaw

    Control tmux sessions/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    392k GitHub starsUsed in 2 repos~640 tokens
    Auto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Auto-check passed
  • Openclaw PR Maintainer

    openclaw/openclaw

    Review, triage, repair, or land OpenClaw issues and pull requests with current-source evidence and the native maintainer workflow.

    392k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Clawsweeper

    openclaw/openclaw

    A skill your agent uses for all ClawSweeper work: OpenClaw issue/PR sweep reports, repair jobs, cloud fix PRs, @clawsweeper maintainer mention commands, trusted ClawSweeper-reviewed…

    392k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Control UI E2E

    openclaw/openclaw

    A skill your agent uses when designing, testing, fixing, or extending the OpenClaw Control UI GUI, including UI stress-test galleries with feedback inputs, Vitest + Playwright end-to-end checks…

    392k GitHub stars~2.9k tokensUpdated today
    Auto-check passed

Questions about Browser Automation

What does Browser Automation do?

A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts. Browser Automation is an agent skill from openclaw/openclaw. Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

When should I use Browser Automation?

Browser Automation fits situations like: controlling web pages with the OpenClaw browser tool; especially multi-step flows; recovery from stale refs/timeouts.

How do I install Browser Automation in Claude Code?

Run `npx skills add openclaw/openclaw --skill browser-automation -a claude-code`. Or copy the skill folder (extensions/browser/skills/browser-automation in openclaw/openclaw) into .claude/skills/browser-automation in your project. Claude Code loads it when a task matches its description.

How do I install Browser Automation in Codex?

Run `npx skills add openclaw/openclaw --skill browser-automation -a codex`. Or copy the skill folder (extensions/browser/skills/browser-automation in openclaw/openclaw) into .agents/skills/browser-automation in your project. Codex loads it when a task matches its description.

Can I use Browser Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openclaw/openclaw --skill browser-automation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-automation, .gemini/skills/browser-automation, .github/skills/browser-automation and .opencode/skills/browser-automation in your project.

What does Browser Automation need to run?

SKILL.md names no scripts, command-line tools or credentials: Browser Automation is instructions for the agent only.

Does Browser Automation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Browser Automation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browser Automation use?

Browser Automation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Automation use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Automation?

Skills that share tags, products or a category with Browser Automation: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Dev-Browser CLI Automation (SawyerHood/dev-browser, 6.7k stars), Agent Browser (superagent-ai/grok-cli, 3.5k stars) and Camoufox CLI (Bin-Huang/camoufox-cli, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Automation?

openclaw (a GitHub organization) maintains it in openclaw/openclaw, which has 391,610 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 8, 2026.

Source: openclaw/openclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.