Agent skill

Browser Use

by xuzhougeng in xuzhougeng/wisp-science

A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…

AGPL-3.0Auto-check passedProductivity & Automation

Install Browser Use

skills CLI
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xuzhougeng/wisp-science browser-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser-use .claude/skills/browser-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-use
GitHub stars
1k
Token cost
~2.7k tokens
SKILL.md length
1,308 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
AGPL-3.0

At a glance

A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…

  • Works in 3 steps: web_open_tab {url} — open the page… → web_scan — read the page (after waiting… → web_execute_js — act, then re-scan to…
  • Drive Wisp Browser Runtime sessions (shared daily Chrome
  • SKILL.md covers Before anything: confirm the…, The loop, Recipes (web_execute_js script) and In-browser chat —…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Browser Use is an agent skill from xuzhougeng/wisp-science. Use this skill to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape content that needs the user's existing cookies and login state. Triggers when the user asks to do something in their browser, log into a site and act inside it, fill out a web form, click through a flow, or extract data from a page that requires being signed in. Tools: browsersetup (check/connect the extension), webopentab (open a…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation and Web scraping. The repository describes itself as: Open-source, local-first desktop AI research workbench for scientific computing with Python/R, MCP bioinformatics tools, SSH/WSL/GPU runtimes, and OpenAI/Anthropic models. The licence is AGPL-3.0.

When your agent uses it

  • Drive Wisp Browser Runtime sessions (shared daily Chrome
  • Workspace Chrome) — open pages
  • Fill and submit forms
  • Scrape content that needs the users existing cookies and login state

Example prompts

  • “/browser-use”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. web_open_tab {url} — open the page (works even with no tab
  2. web_scan — read the page (after waiting for document complete).
  3. web_execute_js — act, then re-scan to confirm the effect. The

What it can do on your machine

Read from SKILL.md and the folder at commit 5eb95c9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Use loads about 2.7k tokens when it runs. Until then it costs about 214 tokens; SKILL.md has 1,308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~214
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xuzhougeng/wisp-science at commit 5eb95c9, republished under its AGPL-3.0 licence (© xuzhougeng). 1,308 words, ~2,676 tokens.

Download SKILL.mdSave it as .claude/skills/browser-use/SKILL.md (or your agent's skills folder).
name
browser-use
description
Use this skill to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape content that needs the user's existing cookies and login state. Triggers when the user asks to do something in their browser, log into a site and act inside it, fill out a web form, click through a flow, or extract data from a page that requires being signed in. Tools: browser_setup (check/connect the extension), web_open_tab (open a URL), web_scan (read visible content + actionable elements with ready-made selectors), web_execute_js (click/type/navigate, or a JSON command for tabs/CDP), web_screenshot (see what the tab is showing — layout, charts, canvas, QR codes). Not for the built-in read-only web fetch — this is for interacting with a live browser.
fold_cue
instead_of=guessing-selectors use=web_scan first — it returns a unique CSS selector and rect for every actionable element; never invent selectors

Browser Use — act inside the user's real Chrome

Wisp talks to the Browser Runtime. Shared mode uses the user's daily Chrome via the unpacked extension — every action runs in their real profile: existing cookies, logins, extensions, and normal fingerprint all apply. Workspace mode can launch a separate Chrome profile. Omitting session always uses shared, even when workspace is connected. Pass session: "workspace" only when the user explicitly requests isolation. If Settings → Browser has Open browser automatically enabled (the default) and the shared extension is disconnected, Wisp may start the installed Chrome/Chromium/Edge so the extension can reconnect. That is still the user's profile, not Playwright or Selenium.

Shared is the default; workspace is not a fallback. Google Chrome 137 and later ignore --load-extension, so on a machine with only branded Chrome the workspace window cannot load the Wisp extension at all. browser_setup {"action":"start_workspace"} therefore returns only once the workspace extension has connected, and otherwise closes the window and fails with WORKSPACE_EXTENSION_BLOCKED. On that error: relay the message, do not retry start_workspace, do not claim any workspace page was opened or read, and get the shared session working instead.

Start with browser_setup without an action (an empty action is also a status check). If the task supplies a target URL, pass it as url so a disconnected daily browser opens that page directly; otherwise startup uses its new-tab page. A connected shared browser is reused without opening another window. Startup does not close existing tabs, including pre-existing blank or new-tab pages.

For figures/code extraction use web_scan with mode: "article" then web_save_assets. For an already-logged-in in-browser chat (ChatGPT, Gemini, or Google AI Mode at google.com/search?udm=50) use web_agent_send, web_agent_wait, web_agent_read on that tab.

Every web_scan and web_execute_js call needs the user's approval by design. Do not treat that as a bug to route around.

Before anything: confirm the bridge is live

Call browser_setup. If status is not connected (or live_retrieval is false), relay its steps (load the unpacked extension from extension_path, verbatim) and stop. Do not answer live, latest, current, or URL-specific questions from prior knowledge. Tell the user this turn contains no live web retrieval and wait until the popup shows Connected to Wisp. Only continue from memory if they explicitly ask for a knowledge-only answer. Never invent the path.

Two fields say why a browser that looks connected is not usable — never report a bare "not connected" when either is set:

  • refused_connection — something reached the bridge port and Wisp refused it (usually a different extension id, or another loopback bridge holding the port). Its popup can still read Connected to Wisp. Relay refused_connection.explanation.
  • update_required / reload_required — a connected extension is older than this build. Call browser_setup {"action":"update_extension"} first. If it returns updated, call browser_setup again and continue only when update_required=false. If it returns manual_reload_required, relay the current and bundled versions plus extension_path verbatim, ask the user to Reload Wisp Real Browser Bridge on chrome://extensions, and wait. Older unpacked extensions cannot accept Wisp's automatic service-worker reload.

One exception: the user says the extension is already installed. Chrome suspends its service worker when idle and reconnects on a one-minute alarm, so disconnected can just be a sleeping worker. Try web_open_tab or web_scan once — a successful call proves the bridge is live — and relay the install steps only if that call fails too.

The loop

  1. web_open_tab {url} — open the page (works even with no tab open yet). Waits until the document is complete, then returns the new tab id plus ready. If ready is false, the load timed out — call web_scan before acting.
  2. web_scan — read the page (after waiting for document complete). Returns page.text, page.title, page.ready_state, and page.elements[], where each element carries a unique selector, its visible text/aria_label, and a rect [x,y,w,h]. Use these selectors directly — do not guess. If ready is false, scan again; do not click a partial page. Use tabs_only:true first when you are unsure which tab to target; pass switch_tab_id:<id> to pin one.
  3. web_execute_js — act, then re-scan to confirm the effect. The extension waits for complete before running the script, and again if the script navigates.

Recipes (web_execute_js script)

Goalscript
Clickdocument.querySelector('<selector>').click()
Type into a fieldconst e=document.querySelector('<sel>'); e.value='text'; e.dispatchEvent(new Event('input',{bubbles:true})); e.dispatchEvent(new Event('change',{bubbles:true}))
Submit a formclick the submit control by its selector, then re-scan
Navigate current tablocation.href='https://example.com'
Read a valuedocument.querySelector('<sel>').textContent

script may instead be a JSON command:

GoalJSON command
Switch to & focus a tab (so the user sees it){"cmd":"tabs","method":"switch","tabId":<id>}
List tabs{"cmd":"tabs"} (or just web_scan tabs_only)
Close tabs you opened{"cmd":"tabs","method":"close","tabIds":[<id>,...]} — returns closed + remaining
Trusted click when .click() is ignored{"cmd":"cdp","method":"Input.dispatchMouseEvent","params":{"type":"mousePressed","x":<x>,"y":<y>,"button":"left","clickCount":1}} then the same with "type":"mouseReleased" — use the element's rect centre from web_scan

Prefer plain JS. Reach for cmd:cdp only when a page blocks synthetic events or you truly need trusted input.

Show full SKILL.md (519 more words)Show less

In-browser chat — web_agent_send / web_agent_wait / web_agent_read

Use these on an already signed-in tab. They are not a new Wisp agent; they drive the chat composer in the user's Chrome.

Supported tabs (HTTPS, exact host, no lookalikes):

  • ChatGPT: chatgpt.com / chat.openai.com
  • Gemini: gemini.google.com
  • Google AI Mode: google.com/search?udm=50 (plain Google Search without udm=50 is refused)

Flow: web_agent_send {prompt} → web_agent_wait → web_agent_read. The read result is {answer_text, citations, status, site}. If the page is login or CAPTCHA, stop and let the user finish it in that tab.

Seeing the page — web_screenshot

web_scan gives text and elements; web_screenshot gives sight. Use it when structure isn't enough: rendered layout, a chart or diagram, a canvas/WebGL page, a QR code, a PDF or image viewer, or a page that looks broken. It captures the visible viewport of the tab — to see below the fold, scroll first (web_execute_js scrollTo(0, 1200)) and capture again. Pass question to say what to read out of it, e.g. {"question":"is the login QR code visible and not expired?"}.

It goes through the configured vision model, so web_scan stays the cheaper default — screenshot when you need eyes, not for every step.

Tab hygiene — the app tracks what you open

Browsing tasks (searching papers, opening a dozen results) used to leave the user with a pile of tabs. Do not ask in chat whether to close them. The desktop records every tab web_open_tab (and tab-create commands) opened in this turn, including after URL changes, and never includes tabs that were already open.

  • If Settings → Browser has Automatically close browser tabs on (browser_setup.auto_close_tabs=true), the app closes this turn's tabs when the turn ends. Do not also close them yourself unless the user asks mid-task.
  • If that setting is off, the app shows a confirmation after the turn (default: close all this-turn tabs; the user can uncheck pages to keep).
  • You may still close a tab mid-task with {"cmd":"tabs","method":"close","tabIds":[...]} if a later step does not need it, or if the user explicitly asks now.

Close only ids you opened yourself. Tabs the user had open, or ones they opened during the task, are theirs.

Stop conditions (do not automate through these)

  • Human verification / CAPTCHA: if web_scan returns human_intervention.required=true, a Wisp prompt has already been shown. Stop browser automation. Do not open another in-app question about the challenge. Do not click, solve, or bypass it. End your turn. The user confirms in the app after completing it in the visible tab; the next message is their confirmation.
  • Credentials: never type passwords, card numbers, or one-time codes yourself. If a step needs a password, have the user sign in directly in the browser and continue once they confirm.
  • Irreversible / outward actions (send, pay, post, delete): confirm with the user before clicking the control.
  • Downloads: for multiple-file downloads, first surface the browser settings from browser_setup (download_automation) and wait for the user to confirm; until then trigger at most one download.
  • Blocked sites: if web_open_tab or a navigational web_execute_js fails with blocked by user URL filter, do not retry that site. Read browser_setup.url_filters.block for the current list. Prefer entries in url_filters.prefer for literature search and similar retrieval; other sites are still allowed.

© xuzhougeng, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/browser-use of xuzhougeng/wisp-science.

Open the folder on GitHubat commit 5eb95c9

Compare with similar skills

Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Use this skillxuzhougeng/wisp-science1k—~2.7kAutomated safety check: PassAGPL-3.0
Browser Harnessdavidondrej/skills4.1k2 repos~3kAutomated safety check: PassMIT
Neo4ier/neo756—~1.8kAutomated safety check: PassNone
Camofox Browserredf0x1/camofox-browser412—~4.6kAutomated safety check: PassMIT
Actionbookactionbook/actionbook1.6k—~1.5kAutomated safety check: PassApache-2.0
Browser Automationalirezarezvani/claude-skills28k—~3.4kAutomated safety check: NotesMIT

Similar skills

  • Browser Harness

    davidondrej/skills

    Direct browser control via CDP. An agent skill from davidondrej/skills.

    4.1k GitHub starsUsed in 2 repos~3k tokens
    Productivity & AutomationAuto-check passed
  • Neo

    4ier/neo

    Browse websites, read web pages, interact with web apps, call website APIs, and automate web tasks.

    756 GitHub stars~1.8k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Camofox Browser

    redf0x1/camofox-browser

    Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.

    412 GitHub stars~4.6k tokensUpdated 17 days ago
    Productivity & AutomationAuto-check passed
  • Actionbook

    actionbook/actionbook

    Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.

    1.6k GitHub stars~1.5k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Automation

    alirezarezvani/claude-skills

    A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check: notes
  • Go Rod Master

    aiskillstore/marketplace

    Comprehensive guide for browser automation and web scraping with go-rod (Chrome DevTools Protocol) including stealth anti-bot-detection patterns.

    430 GitHub starsUsed in 4 repos~4.5k tokens
    Productivity & AutomationAuto-check passed

More from xuzhougeng/wisp-science

All 25 skills in this repo
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Distill Concept Books

    xuzhougeng/wisp-science

    将概念、理论或分析方法类图书蒸馏为证据可追溯、经人工门禁审核且不暴露书名、作者、出版社等来源身份的任务型 Skill 候选。用于新建或恢复图书蒸馏、以本地 Tesseract 扫描 DOCX 全部内嵌图像或 Poppler 渲染的扫描 PDF 全页、建立 source map 与 evidence/claim/relation/capability…

    1k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Skill Creator

    xuzhougeng/wisp-science

    Create, update, validate, and evaluate Wisp skills. An agent skill from xuzhougeng/wisp-science.

    1k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Word Zotero Citations

    xuzhougeng/wisp-science

    Build, audit, authorize, recover, or finalize dynamic Zotero citations and bibliographies in Microsoft Word DOCX files with a protected-source, digest-bound workflow.

    1k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Compute Env Setup

    xuzhougeng/wisp-science

    Set up and validate a reproducible Python or R environment on a Wisp execution context.

    1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Browser Use

What does Browser Use do?

A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…. Browser Use is an agent skill from xuzhougeng/wisp-science. Use this skill to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape content that needs the user's existing cookies and login state.

When should I use Browser Use?

Browser Use fits situations like: drive Wisp Browser Runtime sessions (shared daily Chrome; workspace Chrome) — open pages; fill and submit forms; scrape content that needs the users existing cookies and login state.

How do I install Browser Use in Claude Code?

Run `npx skills add xuzhougeng/wisp-science --skill browser-use -a claude-code`. Or copy the skill folder (skills/browser-use in xuzhougeng/wisp-science) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.

How do I install Browser Use in Codex?

Run `npx skills add xuzhougeng/wisp-science --skill browser-use -a codex`. Or copy the skill folder (skills/browser-use in xuzhougeng/wisp-science) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.

Can I use Browser Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xuzhougeng/wisp-science --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.

What does Browser Use need to run?

SKILL.md names no scripts, command-line tools or credentials: Browser Use is instructions for the agent only.

Does Browser Use access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Browser Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browser Use use?

Browser Use is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Use use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Use?

Skills that share tags, products or a category with Browser Use: Browser Harness (davidondrej/skills, 4.1k stars), Neo (4ier/neo, 756 stars), Camofox Browser (redf0x1/camofox-browser, 412 stars) and Actionbook (actionbook/actionbook, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Use?

xuzhougeng (a GitHub user) maintains it in xuzhougeng/wisp-science, which has 1,022 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.

Source: xuzhougeng/wisp-science on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.