Agent skill

Open Browser

by JasonHonKL in JasonHonKL/Openbrowser

A skill your agent uses whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows.

MITAuto-check passedTesting & QA

Install Open Browser

skills CLI
$ npx skills add JasonHonKL/Openbrowser --skill open-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JasonHonKL/Openbrowser open-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JasonHonKL/Openbrowser.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/open-browser-only .claude/skills/open-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
open-browser
GitHub stars
114
Token cost
~1.6k tokens
SKILL.md length
419 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows.

  • Works in 4 steps: Use persistent sessions for multi-step… → Semantic-first. Plan from the semantic… → Always check [action: ...] tags before… → …
  • The task involves browsing web pages
  • SKILL.md covers Overview, Hard Rules, Canonical Stateful Workflow and JavaScript Mode (--js), plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Open Browser is an agent skill from JasonHonKL/Openbrowser. Use this skill whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows. Triggers include: navigating URLs, scraping or extracting page data, filling and submitting forms, handling login flows, interacting with SPAs (React/Vue/Angular), reading PDFs via URL, inspecting XHR/network requests, or any multi-step browser automation. Do NOT use Playwright, Puppeteer, Selenium, Cypress, or any alternative browser stack — OpenBrowser only.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Browser testing, Browser automation and End-to-end testing. It works with Cypress, Playwright, Puppeteer and Selenium. The repository describes itself as: A browser designed for agent. The licence is MIT.

When your agent uses it

  • The task involves browsing web pages
  • Extracting page content
  • Completing web workflows
  • Include: navigating URLs

Example prompts

  • “/open-browser”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Use persistent sessions for multi-step tasks. Each open-browser interact CLI call wipes all cookies, auth tokens, form state, and…
  2. Semantic-first. Plan from the semantic tree and element IDs — never pixel coordinates.
  3. Always check [action: ...] tags before interacting. Valid actions: click, fill, toggle, select. Do not guess.
  4. IDs are ephemeral. Re-read state after every navigation or DOM mutation. Never cache IDs across page loads.

What it can do on your machine

Read from SKILL.md and the folder at commit 35ae8c2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Open Browser loads about 1.6k tokens when it runs. Until then it costs about 127 tokens; SKILL.md has 419 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~127
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JasonHonKL/Openbrowser at commit 35ae8c2, republished under its MIT licence (© JasonHonKL). 419 words, ~1,595 tokens.

Download SKILL.mdSave it as .claude/skills/open-browser/SKILL.md (or your agent's skills folder).
name
open-browser
description
Use this skill whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows. Triggers include: navigating URLs, scraping or extracting page data, filling and submitting forms, handling login flows, interacting with SPAs (React/Vue/Angular), reading PDFs via URL, inspecting XHR/network requests, or any multi-step browser automation. Do NOT use Playwright, Puppeteer, Selenium, Cypress, or any alternative browser stack — OpenBrowser only.

OpenBrowser Automation

Overview

OpenBrowser is the only permitted browser engine. Never use Playwright, Puppeteer, Selenium, Cypress, or raw Chromium scripts.

ScenarioUse
Multi-step automationbrowser_* tools from ai-agent/open-browser (preferred)
Persistent CLI sessionopen-browser repl
CDP client connectionopen-browser serve
One-off page readopen-browser navigate <url>
Site structure discoveryopen-browser map <url>

Hard Rules

  1. Use persistent sessions for multi-step tasks. Each open-browser interact <url> CLI call wipes all cookies, auth tokens, form state, and localStorage. Multi-step flows (e.g. login → form submit) will fail with repeated one-shot calls.
    • Prefer: browser_new → browser_navigate → … → browser_close
    • Or: open-browser repl for persistent CLI sessions
  2. Semantic-first. Plan from the semantic tree and element IDs — never pixel coordinates.
  3. Always check [action: ...] tags before interacting. Valid actions: click, fill, toggle, select. Do not guess.
  4. IDs are ephemeral. Re-read state after every navigation or DOM mutation. Never cache IDs across page loads.

Canonical Stateful Workflow

1. Open/create session
2. Navigate to target URL
3. Read semantic state (browser_get_state or page output)
4. Identify target by [#ID] + [action: navigate/click/fill/toggle/select]
5. Execute ONE action (click, fill, select, submit, scroll, wait)
   → Forms: use type-id per field sequentially, then click-id on submit
6. Re-read state after every mutation or navigation
7. Repeat until success criteria met
8. Close session

JavaScript Mode (--js)

Enable when:

  • Semantic tree has very few/no interactive elements on a page that should have many
  • A wait selector never resolves without JS
  • Site is a known SPA (React, Vue, Angular)
Site typeWait setting
Default--wait-ms 2000
Slow / heavy SPA--wait-ms 5000

Only inline <script> tags execute. External scripts are not fetched. setTimeout/setInterval are no-ops.


Output Format Selection

GoalFlag
Structured data extraction--format json
Navigation graph--format json --with-nav
Tree structure inspection--format tree
General reading (default)(omit — markdown is default)

--format llm does not exist. Always include source URL and relevant section in extracted answers.


Show full SKILL.md (170 more words)Show less

Advanced Strategies

Site Exploration — use before interacting with unknown or complex sites:

bash
open-browser map https://example.com --depth 2 --output kg.json
# Read kg.json → states (url, title, semantic_tree) + transitions (verified edges)

PDF Handling — navigate directly, no external parser needed:

bash
open-browser navigate https://example.com/report.pdf
# Auto-detects application/pdf, returns parsed semantic tree

XHR / Dynamic Data — inspect raw API responses:

bash
open-browser navigate https://example.com --network-log --format json

若需要 auth,使用 browser_* tool API 的 headers 參數(CLI 目前不支援)

New Tab Handling — after a flow opens a new tab:

bash
open-browser tab list          # find the new tab ID
open-browser tab switch <id>   # switch context explicitly
# In REPL: tab list → tab switch <id>

Quick CLI Reference

bash
# Navigate and read semantic tree (HTML or PDF)
open-browser navigate "https://example.com"
open-browser navigate "https://example.com/report.pdf"

# Output formats
open-browser navigate "https://example.com" --format json
open-browser navigate "https://example.com" --format json --with-nav
open-browser navigate "https://example.com" --format tree

# Interactive elements only
open-browser navigate "https://example.com" --interactive-only

# JavaScript mode
open-browser navigate "https://example.com" --js --wait-ms 2000

# Network debug / XHR extraction
open-browser navigate "https://example.com" --network-log --format json

# Auth bypass
open-browser navigate "https://api.example.com/data" --header "Authorization: Bearer <token>"

# Form filling by element ID
open-browser interact "https://example.com/login" type-id 1 "my_username"
open-browser interact "https://example.com/login" type-id 2 "my_password"
open-browser interact "https://example.com/login" click-id 3

# Map site structure
open-browser map "https://example.com" --depth 2 --output kg.json

# Persistent REPL session
open-browser repl
open> visit https://example.com
open [https://example.com]> click #3
open [https://example.com]> type #5 "search query"
open [https://example.com]> tab open https://example.com/page2
open [https://example.com/page2]> back
open [https://example.com]> exit

# CDP server
open-browser serve --host 0.0.0.0 --port 9222

Error Recovery

ErrorFix
Element ID stale after navigationRe-read state before retrying — IDs reset on every page load
Semantic tree nearly emptyAdd --js --wait-ms 3000 and re-read
method not found from CDPFall back to open-browser CLI or REPL
External scripts not executingBy design — use --wait-ms to let inline-script async content settle
Tab state lost between commandsTab state doesn't persist across CLI calls — use open-browser repl

Safety & Limitations

  • Private/loopback/link-local/metadata endpoints may be blocked — report the constraint, do not bypass with another browser stack.
  • Page.printToPDF is unsupported in semantic-only mode.
  • Screenshot support requires optional build feature — assume unavailable unless confirmed.
  • Prefer core primitives: semanticTree, interact, wait. If a higher-level tool returns method not found, fall back to these.

© JasonHonKL, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/open-browser-only of JasonHonKL/Openbrowser.

Open the folder on GitHubat commit 35ae8c2

Compare with similar skills

Open Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Open Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Open Browser this skillJasonHonKL/Openbrowser114—~1.6kAutomated safety check: PassMIT
Test Framework Migration Skillsickn33/agentic-awesome-skills47k1 repos~2.2kAutomated safety check: PassMIT
Browser Automationynulihao/AgentSkillOS617—~2.2kAutomated safety check: PassNone
Playwright Coretestdino-hq/playwright-skill3861 repos~1.4kAutomated safety check: PassMIT
Brightdata Proxybrightdata/skills264—~5.1kAutomated safety check: PassMIT
Agenttinyfish-io/tinyfish-cookbook2.2k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Test Framework Migration Skill

    sickn33/agentic-awesome-skills

    Migrates and converts test automation scripts between Selenium, Playwright, Puppeteer, and Cypress.

    47k GitHub starsUsed in 1 repo~2.2k tokens
    Testing & QAAuto-check passed
  • Browser Automation

    ynulihao/AgentSkillOS

    Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation.

    617 GitHub stars~2.2k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Playwright Core

    testdino-hq/playwright-skill

    Battle-tested Playwright patterns for writing and debugging reliable E2E, API, component, visual, accessibility, and security tests.

    386 GitHub starsUsed in 1 repo~1.4k tokens
    Testing & QAAuto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Agent

    tinyfish-io/tinyfish-cookbook

    Default browser automation agent — click, fill forms, navigate, log in, and extract structured data from any website using a natural-language goal, or run the same task across multiple sites in…

    2.2k GitHub stars~1.1k tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • React Testing

    affaan-m/ECC

    React component testing with React Testing Library, Vitest/Jest, MSW for network mocking, accessibility assertions with axe, and the decision boundary between component tests and Playwright/Cypress…

    275k GitHub starsUsed in 1 repo~3.3k tokens
    Testing & QAAuto-check passed

Categories

Questions about Open Browser

What does Open Browser do?

A skill your agent uses whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows. Open Browser is an agent skill from JasonHonKL/Openbrowser. Use this skill whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows.

When should I use Open Browser?

Open Browser fits situations like: the task involves browsing web pages; extracting page content; completing web workflows; include: navigating URLs.

How do I install Open Browser in Claude Code?

Run `npx skills add JasonHonKL/Openbrowser --skill open-browser -a claude-code`. Or copy the skill folder (.github/skills/open-browser-only in JasonHonKL/Openbrowser) into .claude/skills/open-browser in your project. Claude Code loads it when a task matches its description.

How do I install Open Browser in Codex?

Run `npx skills add JasonHonKL/Openbrowser --skill open-browser -a codex`. Or copy the skill folder (.github/skills/open-browser-only in JasonHonKL/Openbrowser) into .agents/skills/open-browser in your project. Codex loads it when a task matches its description.

Can I use Open Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JasonHonKL/Openbrowser --skill open-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-browser, .gemini/skills/open-browser, .github/skills/open-browser and .opencode/skills/open-browser in your project.

What does Open Browser need to run?

SKILL.md names no scripts, command-line tools or credentials: Open Browser is instructions for the agent only.

Does Open Browser access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Open Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Open Browser use?

Open Browser is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Open Browser use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Open Browser?

Skills that share tags, products or a category with Open Browser: Test Framework Migration Skill (sickn33/agentic-awesome-skills, 47k stars), Browser Automation (ynulihao/AgentSkillOS, 617 stars), Playwright Core (testdino-hq/playwright-skill, 386 stars) and Brightdata Proxy (brightdata/skills, 264 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Open Browser?

JasonHonKL (a GitHub user) maintains it in JasonHonKL/Openbrowser, which has 114 GitHub stars. The repository was last updated on April 26, 2026.

Source: JasonHonKL/Openbrowser on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.