Agent skill

Browser Harness

by davidondrej in davidondrej/skills

Direct browser control via CDP. An agent skill from davidondrej/skills.

MITAuto-check passedProductivity & Automation

Install Browser Harness

skills CLI
$ npx skills add davidondrej/skills --skill browser-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davidondrej/skills browser-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davidondrej/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-and-web/browser-harness .claude/skills/browser-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-harness
GitHub stars
4.1k
Used in
2 other repos
Token cost
~3k tokens
SKILL.md length
1,278 words
Files
2 (incl. references)
Skills in repo
51
Repo updated
First seen
Licence
MIT

At a glance

Direct browser control via CDP. An agent skill from davidondrej/skills.

  • The user wants to automate
  • SKILL.md covers Usage, Tool call shape, Interaction skills and What actually works, plus 7 more sections
  • Calls uv; reaches x.com and docs.browser-use.com; needs BROWSER_USE_API_KEY
  • Interact with web pages

What it does

Browser Harness is an agent skill from davidondrej/skills. Direct browser control via CDP. Use when the user wants to automate, scrape, test, or interact with web pages. Connects to the user's already-running Chrome.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/install.md`).

It sits in Productivity & Automation, covering Browser automation and Web scraping. The repository describes itself as: access to david ondrej's personal agent skills. The licence is MIT.

When your agent uses it

  • The user wants to automate
  • Interact with web pages

Example prompts

  • “/browser-harness”

Requirements

  • A credential in BROWSER_USE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 7874889. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • x.com
    • docs.browser-use.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BROWSER_USE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Harness loads about 3k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 43 tokens; SKILL.md has 1,278 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davidondrej/skills at commit 7874889, republished under its MIT licence (© davidondrej). 1,278 words, ~2,958 tokens.

Download SKILL.mdSave it as .claude/skills/browser-harness/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
browser-harness
description
Direct browser control via CDP. Use when the user wants to automate, scrape, test, or interact with web pages. Connects to the user's already-running Chrome.

browser-harness

Direct browser control via CDP. For task-specific edits, use agent-workspace/agent_helpers.py. For setup, install, or connection problems, read install.md.

Routing check first: if the task needs no interaction (no clicks, logins, or forms) and you just want page content, use DeepAPI POST /v1/scrape/website instead of driving a browser — see the deepapi skill. Use browser-harness when the task needs a real browser: interaction, JS-heavy flows, logged-in sessions, or visual verification.

Domain skills (community-contributed per-site playbooks under agent-workspace/domain-skills/) are off by default. Set BH_DOMAIN_SKILLS=1 to enable them; see the bottom section.

If BH_DOMAIN_SKILLS=1 and the task is site-specific, read every file in the matching agent-workspace/domain-skills/<site>/ directory before inventing an approach.

Usage

bash
browser-harness -c '
new_tab("https://docs.browser-use.com")
wait_for_load()
print(page_info())
'
  • Invoke as browser-harness — it's on $PATH. No cd, no uv run.
  • First navigation is new_tab(url), not goto_url(url) — goto runs in the user's active tab and clobbers their work.

Tool call shape

bash
browser-harness -c '
# any python. helpers pre-imported. daemon auto-starts.
'

run.py calls ensure_daemon() before exec — you never start/stop manually unless you want to.

Remote browsers

Use remote for parallel sub-agents (each gets its own isolated browser via a distinct BU_NAME) or on a headless server. BROWSER_USE_API_KEY must be set. start_remote_daemon, list_cloud_profiles, list_local_profiles, sync_local_profile are pre-imported.

When supervising those sub-agents, after each check send the user one very short status line: what they are doing and whether they are on track.

Claude Code cmux note: after Claude finishes, it may prefill a predicted next user message; that draft is Claude, not the user speaking.

bash
browser-harness -c '
start_remote_daemon("work")                               # default — clean browser, no profile
# start_remote_daemon("work", profileName="my-work")      # reuse a cloud profile (already logged in)
# start_remote_daemon("work", profileId="<uuid>")         # same, but by UUID
# start_remote_daemon("work", proxyCountryCode="de", timeout=120)   # DE proxy, 2-hour timeout
# start_remote_daemon("work", proxyCountryCode=None)      # disable the Browser Use proxy
'

BU_NAME=work browser-harness -c '
new_tab("https://example.com")
print(page_info())
'

start_remote_daemon prints liveUrl and auto-opens it in the local browser (if a GUI is detected) so the user can watch along. Headless servers print only — share the URL with the user. The daemon PATCHes the cloud browser to stop on shutdown, which persists profile state. Running remote daemons bill until timeout.

Profiles (cookies-only login state) live in interaction-skills/profile-sync.md — covers list_cloud_profiles(), the chat-driven "which profile?" pattern, and sync_local_profile() for uploading a local Chrome profile.

Interaction skills

If you start struggling with a specific mechanic while navigating, look in interaction-skills/ for helpers. They cover reusable UI mechanics like dialogs, tabs, dropdowns, iframes, and uploads. The available interaction skills are:

  • connection.md
  • cookies.md
  • cross-origin-iframes.md
  • dialogs.md
  • downloads.md
  • drag-and-drop.md
  • dropdowns.md
  • iframes.md
  • network-requests.md
  • print-as-pdf.md
  • profile-sync.md
  • screenshots.md
  • scrolling.md
  • shadow-dom.md
  • tabs.md
  • uploads.md
  • viewport.md

What actually works

  • Screenshots first: use capture_screenshot() to understand the current page quickly, find visible targets, and decide whether you need a click, a selector, or more navigation.
  • Clicking: capture_screenshot() → read the pixel off the image → click_at_xy(x, y) → capture_screenshot() to verify. Suppress the Playwright-habit reflex of "locate first, then click" — no getBoundingClientRect, no selector hunt. Drop to DOM only when the target has no visible geometry (hidden input, 0×0 node). Hit-testing happens in Chrome's browser process, so clicks go through iframes / shadow DOM / cross-origin without extra work.
  • Bulk HTTP: http_get(url) + ThreadPoolExecutor. No browser for static pages (249 Netflix pages in 2.8s).
  • After goto: wait_for_load().
  • Wrong/stale tab: ensure_real_tab(). Use it when the current tab is stale or internal; the daemon also auto-recovers from stale sessions on the next call.
  • Verification: print(page_info()) is the simplest "is this alive?" check, but screenshots are the default way to verify whether a visible action actually worked.
  • DOM reads: use js(...) for inspection and extraction when the screenshot shows that coordinates are the wrong tool.
  • Iframe sites (Azure blades, Salesforce): click_at_xy(x, y) passes through; only drop to iframe DOM work when coordinate clicks are the wrong tool.
  • Auth wall: redirected to login → stop and ask the user. Don't type credentials from screenshots.
  • Raw CDP for anything helpers don't cover: cdp("Domain.method", params).

Design constraints

  • Coordinate clicks default. Input.dispatchMouseEvent goes through iframes/shadow/cross-origin at the compositor level.
  • Connect to the user's running Chrome. Don't launch your own browser.
  • cdp-use is only for CDPClient.send_raw. Prefer raw CDP strings over typed wrappers.
  • run.py stays tiny. No argparse, subcommands, or extra control layer.
  • Core helpers stay short. Put task-specific helper additions in agent-workspace/agent_helpers.py; daemon/bootstrap and remote session admin live in the core package.
  • Don't add a manager layer. No retries framework, session manager, daemon supervisor, config system, or logging framework.

Hermes Agent integration

Installed at ~/Developer/browser-harness as editable uv tool install -e .. Binary at ~/.local/bin/browser-harness. Skill at ~/.hermes/skills/browser-harness/.

Frontmatter pitfall: The upstream SKILL.md ships with name: browser in frontmatter, which collides with Hermes's built-in browser toolset. When copying into ~/.hermes/skills/, rename to name: browser-harness in the frontmatter or Hermes will shadow/conflict with its own browser tools.

Brave Browser: Works identically to Chrome. Enable remote debugging at brave://inspect/#remote-debugging (same checkbox). The harness auto-discovers Brave's profile directory.

Show full SKILL.md (548 more words)Show less

Authenticated content extraction (proven pattern)

browser-harness connects to the user's real browser with their active sessions — ideal for extracting content from login-walled sites where web_extract or Hermes's built-in browser_navigate fail (e.g. X/Twitter articles, LinkedIn, paywalled sites).

Pattern:

bash
browser-harness -c '
new_tab("https://x.com/user/status/123456")
wait_for_load()
import time
time.sleep(5)  # let JS-heavy pages render
text = js("""
    const article = document.querySelector("article");
    if (article) return article.innerText;
    return document.body.innerText;
""")
with open("/tmp/extracted.txt", "w") as f:
    f.write(text)
print("Written", len(text), "chars")
'
  • Write to a temp file to avoid shell escaping issues with large text
  • Use time.sleep() generously for JS-heavy SPAs (X, LinkedIn need 3-5s)
  • X/Twitter articles render inline — just scroll/extract via DOM, no extra click needed
  • For very long pages, js(...) with innerText grabs everything including below-fold content

Hermes Agent integration

Installed at ~/Developer/browser-harness as editable uv tool install -e .. Binary at ~/.local/bin/browser-harness. Skill at ~/.hermes/skills/browser-harness/.

Frontmatter pitfall: The upstream SKILL.md ships with name: browser in frontmatter, which collides with Hermes's built-in browser toolset. When copying into ~/.hermes/skills/, rename to name: browser-harness in the frontmatter.

Brave Browser: Works identically to Chrome. Enable remote debugging at brave://inspect/#remote-debugging (same checkbox). The harness auto-discovers Brave's profile directory.

Authenticated content extraction (proven pattern)

browser-harness connects to the user's real browser with active sessions — ideal for login-walled sites where web_extract or Hermes's built-in browser_navigate fail (X/Twitter articles, LinkedIn, paywalled sites).

bash
browser-harness -c '
new_tab("https://x.com/user/status/123456")
wait_for_load()
import time
time.sleep(5)  # let JS-heavy pages render
text = js("""
    const article = document.querySelector("article");
    if (article) return article.innerText;
    return document.body.innerText;
""")
with open("/tmp/extracted.txt", "w") as f:
    f.write(text)
print("Written", len(text), "chars")
'
  • Write to a temp file to avoid shell escaping issues with large text
  • Use time.sleep() generously for JS-heavy SPAs (X, LinkedIn need 3-5s)
  • X/Twitter articles render inline — just scroll/extract via DOM, no extra click needed
  • js(...) with innerText grabs everything including below-fold content

Gotchas (field-tested)

  • Brave Browser uses brave://inspect/#remote-debugging instead of chrome://inspect/.... The harness auto-discovers Brave's data dir.
  • Login-walled content extraction (e.g. X/Twitter articles): navigate with new_tab(url), wait_for_load(), then extract via js("document.querySelector('article').innerText"). Write to a temp file to avoid shell escaping: with open('/tmp/out.txt', 'w') as f: f.write(text). The user's existing browser session handles auth automatically.
  • Omnibox popups are fake page targets. Filter chrome://omnibox-popup... and other internals when you need a real tab.
  • CDP target order != Chrome's visible tab-strip order. Use UI automation when the user means "the first/second tab I can see"; Target.activateTarget only shows a known target.
  • Default daemon sessions can go stale. ensure_real_tab() re-attaches to a real page.
  • Browser Use API is camelCase on the wire. cdpUrl, proxyCountryCode, etc.
  • Remote cdpUrl is HTTPS, not ws. Resolve the websocket URL via /json/version.
  • Stop cloud browsers with PATCH /browsers/{id} + {"action":"stop"}.
  • After every meaningful action, re-screenshot before assuming it worked. Use the image to verify changed state, open menus, navigation, visible errors, and whether the page is in the state you expected.
  • Use screenshots to drive exploration. They are often the fastest way to find the next click target, notice hidden blockers, and decide if a selector is even worth writing.
  • Prefer compositor-level actions over framework hacks. Try screenshots, coordinate clicks, and raw key input before adding DOM-specific workarounds.
  • If you need framework-specific DOM tricks, check interaction-skills/ first. That is where dropdown, dialog, iframe, shadow DOM, and form-specific guidance belongs.

Domain skills (opt-in)

Only applies when BH_DOMAIN_SKILLS=1. Otherwise ignore — agent-workspace/domain-skills/ is dormant and goto_url won't surface skill files.

When enabled, search agent-workspace/domain-skills/<host>/ before inventing an approach. goto_url returns up to 10 skill filenames for the navigated host.

If you learn anything non-obvious — a private API, stable selector, framework quirk, URL pattern, hidden wait, or site-specific trap — open a PR to agent-workspace/domain-skills/<site>/. Capture the durable shape of the site (the map, not the diary). Don't write pixel coordinates (break on layout), task narration, or secrets — the directory is public.

© davidondrej, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/research-and-web/browser-harness of davidondrej/skills.

  • SKILL.md
  • references/install.md

Open the folder on GitHubat commit 7874889

Used in 2 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in davidondrej/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Browser Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Harness this skilldavidondrej/skills4.1k2 repos~3kAutomated safety check: PassMIT
Neo4ier/neo756—~1.8kAutomated safety check: PassNone
Camofox Browserredf0x1/camofox-browser412—~4.6kAutomated safety check: PassMIT
Actionbookactionbook/actionbook1.6k—~1.5kAutomated safety check: PassApache-2.0
Browser Automationalirezarezvani/claude-skills28k—~3.4kAutomated safety check: NotesMIT
Browser Usexuzhougeng/wisp-science1k—~2.7kAutomated safety check: PassAGPL-3.0

Similar skills

  • Neo

    4ier/neo

    Browse websites, read web pages, interact with web apps, call website APIs, and automate web tasks.

    756 GitHub stars~1.8k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Camofox Browser

    redf0x1/camofox-browser

    Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.

    412 GitHub stars~4.6k tokensUpdated 17 days ago
    Productivity & AutomationAuto-check passed
  • Actionbook

    actionbook/actionbook

    Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.

    1.6k GitHub stars~1.5k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Automation

    alirezarezvani/claude-skills

    A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check: notes
  • Browser Use

    xuzhougeng/wisp-science

    A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape…

    1k GitHub stars~2.7k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Go Rod Master

    aiskillstore/marketplace

    Comprehensive guide for browser automation and web scraping with go-rod (Chrome DevTools Protocol) including stealth anti-bot-detection patterns.

    430 GitHub starsUsed in 4 repos~4.5k tokens
    Productivity & AutomationAuto-check passed

More from davidondrej/skills

All 51 skills in this repo
  • Nagent

    davidondrej/skills

    Launch a new bb worker thread with the right project, model, worktree, and task brief.

    4.1k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check: notes
  • Persistent Localhost

    davidondrej/skills

    Manage persistent dev servers, APIs, and other local processes on a port using macOS LaunchAgents.

    4.1k GitHub stars~618 tokensUpdated yesterday
    Auto-check passed
  • Reset Cursor Acp

    davidondrej/skills

    Reset a stuck Cursor ACP thread in <chat-system and reload its configuration.

    4.1k GitHub stars~728 tokensUpdated yesterday
    Auto-check passed
  • Anti Sleep

    davidondrej/skills

    Keep a Mac awake for a set duration or while a process runs.

    4.1k GitHub stars~640 tokensUpdated yesterday
    Auto-check: warnings
  • Bb CLI

    davidondrej/skills

    Use this when controlling bb. An agent skill from davidondrej/skills.

    4.1k GitHub stars~823 tokensUpdated yesterday
    Auto-check passed
  • Boat

    davidondrej/skills

    Manage Boat (boat.dev, formerly Box by Ascii) cloud sandboxes via API, CLI, and SSH.

    4.1k GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes

Questions about Browser Harness

What does Browser Harness do?

Direct browser control via CDP. An agent skill from davidondrej/skills. Browser Harness is an agent skill from davidondrej/skills. Direct browser control via CDP.

When should I use Browser Harness?

Browser Harness fits situations like: the user wants to automate; interact with web pages.

How do I install Browser Harness in Claude Code?

Run `npx skills add davidondrej/skills --skill browser-harness -a claude-code`. Or copy the skill folder (skills/research-and-web/browser-harness in davidondrej/skills) into .claude/skills/browser-harness in your project. Claude Code loads it when a task matches its description.

How do I install Browser Harness in Codex?

Run `npx skills add davidondrej/skills --skill browser-harness -a codex`. Or copy the skill folder (skills/research-and-web/browser-harness in davidondrej/skills) into .agents/skills/browser-harness in your project. Codex loads it when a task matches its description.

Can I use Browser Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davidondrej/skills --skill browser-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-harness, .gemini/skills/browser-harness, .github/skills/browser-harness and .opencode/skills/browser-harness in your project.

What does Browser Harness need to run?

Going by SKILL.md and its folder, Browser Harness needs the command-line tools its instructions call (uv) and credentials named BROWSER_USE_API_KEY. Our summary lists: A credential in BROWSER_USE_API_KEY.

Does Browser Harness access the network?

SKILL.md names 2 domains. In commands or code: x.com and docs.browser-use.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Browser Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browser Harness use?

Browser Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Harness use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Browser Harness?

Skills that share tags, products or a category with Browser Harness: Neo (4ier/neo, 756 stars), Camofox Browser (redf0x1/camofox-browser, 412 stars), Actionbook (actionbook/actionbook, 1.6k stars) and Browser Automation (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Harness?

davidondrej (a GitHub user) maintains it in davidondrej/skills, which has 4,112 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on October 8, 2026.

Source: davidondrej/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.