Browser Harness
davidondrej/skills
Direct browser control via CDP. An agent skill from davidondrej/skills.
A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install xuzhougeng/wisp-science browser-use --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser-use .claude/skills/browser-use && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .claude/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-useType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install xuzhougeng/wisp-science browser-use --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/browser-use .agents/skills/browser-use && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .agents/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install xuzhougeng/wisp-science browser-use --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/browser-use .cursor/skills/browser-use && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .cursor/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/xuzhougeng/wisp-science.git --path skills/browser-use--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install xuzhougeng/wisp-science browser-use --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/browser-use .gemini/skills/browser-use && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .gemini/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install xuzhougeng/wisp-science browser-useInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/browser-use .github/skills/browser-use && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .github/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xuzhougeng/wisp-science --skill browser-use -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install xuzhougeng/wisp-science browser-use --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/browser-use .opencode/skills/browser-use && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/xuzhougeng/wisp-science/tree/main/skills/browser-use into .opencode/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser-useA skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…
Browser Use is an agent skill from xuzhougeng/wisp-science. Use this skill to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape content that needs the user's existing cookies and login state. Triggers when the user asks to do something in their browser, log into a site and act inside it, fill out a web form, click through a flow, or extract data from a page that requires being signed in. Tools: browsersetup (check/connect the extension), webopentab (open a…
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Productivity & Automation, covering Browser automation and Web scraping. The repository describes itself as: Open-source, local-first desktop AI research workbench for scientific computing with Python/R, MCP bioinformatics tools, SSH/WSL/GPU runtimes, and OpenAI/Anthropic models. The licence is AGPL-3.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5eb95c9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser Use loads about 2.7k tokens when it runs. Until then it costs about 214 tokens; SKILL.md has 1,308 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from xuzhougeng/wisp-science at commit 5eb95c9, republished under its AGPL-3.0 licence (© xuzhougeng). 1,308 words, ~2,676 tokens.
.claude/skills/browser-use/SKILL.md (or your agent's skills folder).Wisp talks to the Browser Runtime. Shared mode uses the user's daily
Chrome via the unpacked extension — every action runs in their real
profile: existing cookies, logins, extensions, and normal fingerprint all
apply. Workspace mode can launch a separate Chrome profile. Omitting
session always uses shared, even when workspace is connected. Pass
session: "workspace" only when the user explicitly requests isolation. If
Settings → Browser has Open browser automatically enabled (the
default) and the shared extension is disconnected, Wisp may start the installed
Chrome/Chromium/Edge so the extension can reconnect. That is still the
user's profile, not Playwright or Selenium.
Shared is the default; workspace is not a fallback. Google Chrome 137
and later ignore --load-extension, so on a machine with only branded
Chrome the workspace window cannot load the Wisp extension at all.
browser_setup {"action":"start_workspace"} therefore returns only once
the workspace extension has connected, and otherwise closes the window
and fails with WORKSPACE_EXTENSION_BLOCKED. On that error: relay the
message, do not retry start_workspace, do not claim any workspace page
was opened or read, and get the shared session working instead.
Start with browser_setup without an action (an empty action is also a status
check). If the task supplies a target URL, pass it as url so a disconnected
daily browser opens that page directly; otherwise startup uses its new-tab page.
A connected shared browser is reused without opening another window. Startup
does not close existing tabs, including pre-existing blank or new-tab pages.
For figures/code extraction use web_scan with mode: "article" then
web_save_assets. For an already-logged-in in-browser chat (ChatGPT,
Gemini, or Google AI Mode at google.com/search?udm=50) use
web_agent_send, web_agent_wait, web_agent_read on that tab.
Every web_scan and web_execute_js call needs the user's approval by
design. Do not treat that as a bug to route around.
Call browser_setup. If status is not connected (or live_retrieval
is false), relay its steps (load the unpacked extension from
extension_path, verbatim) and stop. Do not answer live, latest,
current, or URL-specific questions from prior knowledge. Tell the user
this turn contains no live web retrieval and wait until the popup shows
Connected to Wisp. Only continue from memory if they explicitly ask
for a knowledge-only answer. Never invent the path.
Two fields say why a browser that looks connected is not usable — never report a bare "not connected" when either is set:
refused_connection — something reached the bridge port and Wisp
refused it (usually a different extension id, or another loopback
bridge holding the port). Its popup can still read Connected to Wisp.
Relay refused_connection.explanation.update_required / reload_required — a connected extension is older than
this build. Call browser_setup {"action":"update_extension"} first. If it
returns updated, call browser_setup again and continue only when
update_required=false. If it returns manual_reload_required, relay the
current and bundled versions plus extension_path verbatim, ask the user to
Reload Wisp Real Browser Bridge on chrome://extensions, and wait. Older
unpacked extensions cannot accept Wisp's automatic service-worker reload.One exception: the user says the extension is already installed. Chrome
suspends its service worker when idle and reconnects on a one-minute
alarm, so disconnected can just be a sleeping worker. Try web_open_tab
or web_scan once — a successful call proves the bridge is live — and
relay the install steps only if that call fails too.
web_open_tab {url} — open the page (works even with no tab
open yet). Waits until the document is complete, then returns the new
tab id plus ready. If ready is false, the load timed out — call
web_scan before acting.web_scan — read the page (after waiting for document complete).
Returns page.text, page.title, page.ready_state, and
page.elements[], where each element carries a unique selector,
its visible text/aria_label, and a rect [x,y,w,h]. Use these
selectors directly — do not guess. If ready is false, scan again;
do not click a partial page. Use tabs_only:true first when you are
unsure which tab to target; pass switch_tab_id:<id> to pin one.web_execute_js — act, then re-scan to confirm the effect. The
extension waits for complete before running the script, and again if
the script navigates.web_execute_js script)| Goal | script |
|---|---|
| Click | document.querySelector('<selector>').click() |
| Type into a field | const e=document.querySelector('<sel>'); e.value='text'; e.dispatchEvent(new Event('input',{bubbles:true})); e.dispatchEvent(new Event('change',{bubbles:true})) |
| Submit a form | click the submit control by its selector, then re-scan |
| Navigate current tab | location.href='https://example.com' |
| Read a value | document.querySelector('<sel>').textContent |
script may instead be a JSON command:
| Goal | JSON command |
|---|---|
| Switch to & focus a tab (so the user sees it) | {"cmd":"tabs","method":"switch","tabId":<id>} |
| List tabs | {"cmd":"tabs"} (or just web_scan tabs_only) |
| Close tabs you opened | {"cmd":"tabs","method":"close","tabIds":[<id>,...]} — returns closed + remaining |
Trusted click when .click() is ignored | {"cmd":"cdp","method":"Input.dispatchMouseEvent","params":{"type":"mousePressed","x":<x>,"y":<y>,"button":"left","clickCount":1}} then the same with "type":"mouseReleased" — use the element's rect centre from web_scan |
Prefer plain JS. Reach for cmd:cdp only when a page blocks synthetic
events or you truly need trusted input.
web_agent_send / web_agent_wait / web_agent_readUse these on an already signed-in tab. They are not a new Wisp agent; they drive the chat composer in the user's Chrome.
Supported tabs (HTTPS, exact host, no lookalikes):
chatgpt.com / chat.openai.comgemini.google.comgoogle.com/search?udm=50 (plain Google Search without
udm=50 is refused)Flow: web_agent_send {prompt} → web_agent_wait → web_agent_read. The
read result is {answer_text, citations, status, site}. If the page is
login or CAPTCHA, stop and let the user finish it in that tab.
web_screenshotweb_scan gives text and elements; web_screenshot gives sight. Use it
when structure isn't enough: rendered layout, a chart or diagram, a
canvas/WebGL page, a QR code, a PDF or image viewer, or a page that looks
broken. It captures the visible viewport of the tab — to see below the
fold, scroll first (web_execute_js scrollTo(0, 1200)) and capture again.
Pass question to say what to read out of it, e.g.
{"question":"is the login QR code visible and not expired?"}.
It goes through the configured vision model, so web_scan stays the cheaper
default — screenshot when you need eyes, not for every step.
Browsing tasks (searching papers, opening a dozen results) used to leave the
user with a pile of tabs. Do not ask in chat whether to close them. The
desktop records every tab web_open_tab (and tab-create commands) opened in
this turn, including after URL changes, and never includes tabs that were
already open.
browser_setup.auto_close_tabs=true), the app closes this turn's tabs
when the turn ends. Do not also close them yourself unless the user asks
mid-task.{"cmd":"tabs","method":"close","tabIds":[...]} if a later step does not
need it, or if the user explicitly asks now.Close only ids you opened yourself. Tabs the user had open, or ones they opened during the task, are theirs.
web_scan returns
human_intervention.required=true, a Wisp prompt has already been shown.
Stop browser automation. Do not open another in-app question about the
challenge. Do not click, solve, or bypass it. End your turn. The user
confirms in the app after completing it in the visible tab; the next
message is their confirmation.browser_setup (download_automation) and wait for the
user to confirm; until then trigger at most one download.web_open_tab or a navigational web_execute_js
fails with blocked by user URL filter, do not retry that site. Read
browser_setup.url_filters.block for the current list. Prefer entries in
url_filters.prefer for literature search and similar retrieval; other
sites are still allowed.© xuzhougeng, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/browser-use of xuzhougeng/wisp-science.
Open the folder on GitHubat commit 5eb95c9
Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser Use this skillxuzhougeng/wisp-science | 1k | — | ~2.7k | Automated safety check: Pass | AGPL-3.0 | |
| Browser Harnessdavidondrej/skills | 4.1k | 2 repos | ~3k | Automated safety check: Pass | MIT | |
| Neo4ier/neo | 756 | — | ~1.8k | Automated safety check: Pass | None | |
| Camofox Browserredf0x1/camofox-browser | 412 | — | ~4.6k | Automated safety check: Pass | MIT | |
| Actionbookactionbook/actionbook | 1.6k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Browser Automationalirezarezvani/claude-skills | 28k | — | ~3.4k | Automated safety check: Notes | MIT |
davidondrej/skills
Direct browser control via CDP. An agent skill from davidondrej/skills.
4ier/neo
Browse websites, read web pages, interact with web apps, call website APIs, and automate web tasks.
redf0x1/camofox-browser
Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.
actionbook/actionbook
Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.
alirezarezvani/claude-skills
A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.
aiskillstore/marketplace
Comprehensive guide for browser automation and web scraping with go-rod (Chrome DevTools Protocol) including stealth anti-bot-detection patterns.
xuzhougeng/wisp-science
A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.
xuzhougeng/wisp-science
学术审查 / research-integrity screening of a manuscript's figures and reported numbers.
xuzhougeng/wisp-science
将概念、理论或分析方法类图书蒸馏为证据可追溯、经人工门禁审核且不暴露书名、作者、出版社等来源身份的任务型 Skill 候选。用于新建或恢复图书蒸馏、以本地 Tesseract 扫描 DOCX 全部内嵌图像或 Poppler 渲染的扫描 PDF 全页、建立 source map 与 evidence/claim/relation/capability…
xuzhougeng/wisp-science
Create, update, validate, and evaluate Wisp skills. An agent skill from xuzhougeng/wisp-science.
xuzhougeng/wisp-science
Build, audit, authorize, recover, or finalize dynamic Zotero citations and bibliographies in Microsoft Word DOCX files with a protected-source, digest-bound workflow.
xuzhougeng/wisp-science
Set up and validate a reproducible Python or R environment on a Wisp execution context.
A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…. Browser Use is an agent skill from xuzhougeng/wisp-science. Use this skill to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape content that needs the user's existing cookies and login state.
Browser Use fits situations like: drive Wisp Browser Runtime sessions (shared daily Chrome; workspace Chrome) — open pages; fill and submit forms; scrape content that needs the users existing cookies and login state.
Run `npx skills add xuzhougeng/wisp-science --skill browser-use -a claude-code`. Or copy the skill folder (skills/browser-use in xuzhougeng/wisp-science) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.
Run `npx skills add xuzhougeng/wisp-science --skill browser-use -a codex`. Or copy the skill folder (skills/browser-use in xuzhougeng/wisp-science) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xuzhougeng/wisp-science --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.
SKILL.md names no scripts, command-line tools or credentials: Browser Use is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Browser Use is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browser Use: Browser Harness (davidondrej/skills, 4.1k stars), Neo (4ier/neo, 756 stars), Camofox Browser (redf0x1/camofox-browser, 412 stars) and Actionbook (actionbook/actionbook, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
xuzhougeng (a GitHub user) maintains it in xuzhougeng/wisp-science, which has 1,022 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.
Source: xuzhougeng/wisp-science on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.