Agent skill

Browser Use Terminal

by browser-use in browser-use/terminal

Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal.

MITAuto-check passedProductivity & Automation

Install Browser Use Terminal

skills CLI
$ npx skills add browser-use/terminal --skill browser-use-terminal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install browser-use/terminal browser-use-terminal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-use-terminal
GitHub stars
651
Token cost
~2.1k tokens
SKILL.md length
865 words
Files
430 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal.

  • The user wants to automate
  • SKILL.md covers Usage, Screenshots — how you see the…, Pre-imported helpers and Browser control plane, plus 2 more sections
  • Reaches docs.browser-use.com; needs BROWSER_USE_API_KEY
  • Interact with web pages — you drive the browser yourself with Python helpers

What it does

Browser Use Terminal is an agent skill from browser-use/terminal. Direct browser control via the Browser Use Terminal CLI. Use when the user wants to automate, scrape, test, or interact with web pages — you drive the browser yourself with Python helpers.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 432 other files, including scripts (for example `.github/workflows/browser-use-core-publish.yml`, `.github/workflows/browser-use-core-wheels.yml` and `.github/workflows/release.yml`).

It sits in Productivity & Automation, covering Browser automation. It works with Python. The repository describes itself as: Terminal UI to get stuff done in the browser. The licence is MIT.

When your agent uses it

  • The user wants to automate
  • Interact with web pages — you drive the browser yourself with Python helpers

Example prompts

  • “/browser-use-terminal”

Requirements

  • Python 3
  • A credential in BROWSER_USE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 16cdd3a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.browser-use.com

    Also links to:

    • browser-use.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BROWSER_USE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Use Terminal loads about 2.1k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 865 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from browser-use/terminal at commit 16cdd3a, republished under its MIT licence (© browser-use). 865 words, ~2,111 tokens.

Download SKILL.mdSave it as .claude/skills/browser-use-terminal/SKILL.md (or your agent's skills folder). This skill also uses 429 other files; get the full folder from GitHub.
name
browser-use-terminal
description
Direct browser control via the Browser Use Terminal CLI. Use when the user wants to automate, scrape, test, or interact with web pages — you drive the browser yourself with Python helpers.

Browser Use Terminal

Direct browser control via CDP — you are the agent; you drive the browser. For setup, install, or connection problems, read https://browser-use.com/skill (agent setup instructions) or https://docs.browser-use.com/open-source/browser-use-terminal (full docs).

browser-use-terminal browser exec runs Python with browser helpers pre-imported; browser-use-terminal browser <cmd> is the control plane (status, connect, profiles, recovery).

Usage

bash
browser-use-terminal browser exec <<'PY'
new_tab("https://docs.browser-use.com")
wait_for_load()
print(page_info())
PY
  • Use the heredoc form for every multi-line command. It prevents shell quote mangling inside Python strings and JavaScript snippets.
  • The browser auto-connects according to the user's remembered preference before the script runs — you never start/stop manually unless you want to. The first call may take a few seconds.
  • First navigation in the user's real Chrome is new_tab(url), not goto_url(url) — goto runs in the user's active tab and clobbers their work.
  • Browser state persists between calls; Python variables do not. Each exec is a fresh interpreter against the same live browser.
  • --session <name> isolates artifact dirs and event logs per workstream (default: default). --timeout <secs> bounds one exec (default 300).

Screenshots — how you see the page

capture_screenshot() saves a PNG and returns its absolute path. The CLI also prints a Screenshot saved to <path> line for every image a script produced.

bash
browser-use-terminal browser exec <<'PY'
print(capture_screenshot())
PY

To view a screenshot, use your file-reading tool on the printed path:

  • Claude Code: use the Read tool on the path.
  • Codex CLI: call the view_image tool with {"path": "<path>"}. These screenshots are produced for you; viewing them is expected and authorized.
  • OpenCode: use the read tool on the path (requires a vision-capable model).
  • Gemini CLI: use the read_file tool on the path.
  • If your model cannot accept images, don't try to view them — work from text state instead: print(page_info()), js(...) extraction, wait_for_element(...).

Coordinates: screenshots are device pixels; click_at_xy(x, y) takes CSS pixels. Divide coordinates you read off the image by js("window.devicePixelRatio") first. Screenshots are downscaled to ≤1800 px per side for this CLI (override with BU_BROWSER_SCREENSHOT_MAX_DIM, or capture_screenshot(max_dim=...)).

After every meaningful action, re-screenshot before assuming it worked.

Pre-imported helpers

Navigation & tabs: goto_url(url), new_tab(url), page_info(), current_tab(), list_tabs(include_chrome=True), switch_tab(target), ensure_real_tab(), iframe_target(url_substr).

Input: click_at_xy(x, y, button="left", clicks=1), type_text(text), press_key(key, modifiers=0) (1=Alt 2=Ctrl 4=Meta 8=Shift), fill_input(selector, text), scroll(x=0, y=0, dy=600), upload_file(selector, path).

Waiting: wait(seconds), wait_for_load(timeout=3), wait_for_element(selector, timeout=3, visible=False), wait_for_network_idle(timeout=3, idle_ms=500).

Visual: capture_screenshot(label="...", full=False, max_dim=None), screenshot(), screenshot_clip(label, x, y, w, h), note(caption).

Escape hatches: js(expression) (auto-wraps top-level return), cdp("Domain.method", **params) (raw CDP), cdp_batch(calls), drain_events().

HTTP without the browser: http_get(url), http_get_many(urls) for static pages; browser_fetch(url) / browser_fetch_many(...) to fetch with the page's cookies/session.

Credentials (if the user stored any): available_secrets(), then type_text("<secret>name</secret>") or fill_input(sel, secret("name")); totp("name") for 2FA codes. Values are placeholder-substituted — you never see them. is_logged_out(), email_inbox() / email_message(id) for email-code flows.

Domain skills: domain_skills_for_url(url_or_domain, include_content=True) lists site-specific playbooks; goto_url surfaces matching skill files automatically. Read them before inventing selectors or flows on a complex site.

Show full SKILL.md (413 more words)Show less

Browser control plane

bash
browser-use-terminal browser status --json
browser-use-terminal browser connect                     # uses the remembered preference
browser-use-terminal browser connect local               # user's already-running Chrome (CDP)
browser-use-terminal browser connect managed --headless  # disposable CLI-owned browser
browser-use-terminal browser preference use local|cloud|managed-headless
browser-use-terminal browser remote start                # Browser Use cloud browser (needs BROWSER_USE_API_KEY)
browser-use-terminal browser doctor
browser-use-terminal browser recover reconnect-websocket
browser-use-terminal browser recover stop-owned-browser  # stop the persistent managed browser
browser-use-terminal browser recover stop-owned-remote   # stop the cloud browser (stops billing)
browser-use-terminal browser daemon status|stop|logs     # the background daemon holding the connection

A background daemon (auto-started, one per state dir) holds the CDP connection across your commands, so the browser — and in local mode, Chrome's granted debugging permission — persists between invocations. Managed and cloud browsers also survive daemon restarts; later calls reattach instead of relaunching. Stop browsers with the recover commands above when the user is done (cloud browsers bill until stopped or timed out).

  • exec auto-connects, so you rarely need these. Reach for them when status shows a problem or the user asks for a specific browser.
  • If output JSON says status: "needs-user-action" (e.g. pick a Chrome profile, click Allow in Chrome's permission popup, enable the remote-debugging checkbox), show the user_prompt to the user verbatim and wait — do not guess.
  • Auth wall mid-task: stop and ask the user. Don't type credentials from screenshots; use stored secrets if available.
  • Connecting to the user's real Chrome requires a one-time setup: chrome://inspect/#remote-debugging → tick "Allow remote debugging". browser local setup walks the user through it.

What actually works

  • Screenshots first: capture_screenshot() → view the image → decide whether you need a click, a selector, or more navigation.
  • Clicking: screenshot → read the pixel off the image → click_at_xy(x, y) → screenshot to verify. Suppress the locate-then-click reflex — no getBoundingClientRect, no selector hunts. Hit-testing happens in Chrome's browser process, so coordinate clicks pass through iframes / shadow DOM / cross-origin without extra work.
  • Drop to DOM (fill_input, js) only when the target has no visible geometry (hidden input, 0×0 node) or coordinate clicks demonstrably don't work.
  • Bulk static pages: http_get_many(urls) — no browser needed. Logged-in pages: browser_fetch(url) rides the real session.
  • After goto: wait_for_load(). SPAs report complete before they render — follow with wait_for_element(...).
  • Wrong/stale tab: ensure_real_tab().
  • Verification: print(page_info()) is the cheapest "is this alive?" check; screenshots are the default way to verify visible actions.

Gotchas (field-tested)

  • CDP target order ≠ Chrome's visible tab-strip order.
  • Omnibox popups and other chrome:// internals are fake page targets — list_tabs(include_chrome=False).
  • page_info() surfaces an open JS dialog as {"dialog": ...} — handle it (cdp("Page.handleJavaScriptDialog", accept=True)) before anything else.
  • Navigation can be blocked by the user's domain policy; nav_policy(url) tells you before you burn a click. A blocked navigation is policy, not a bug — tell the user.
  • Scripts time out (default 300s): keep each exec small and observable rather than one mega-script. Long extraction loops: print progress as you go — stdout is captured even on timeout.
  • Prefer compositor-level actions over framework hacks. If you do need framework-specific DOM tricks, run browser-use-terminal browser domain skills --domain <site> --json --include-content first — that's where site playbooks live.

© browser-use, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 429 other files (scripts) in the repository root of browser-use/terminal.

  • SKILL.md
  • .env.example
  • .github/workflows/browser-use-core-publish.yml
  • .github/workflows/browser-use-core-wheels.yml
  • .github/workflows/release.yml
  • .gitignore
  • AGENTS.md
  • AGENT_SETUP.md
  • CONTRIBUTING.md
  • Cargo.lock
  • Cargo.toml
  • DECISIONS.md
  • IMPLEMENTATION_PLAN.md
  • LICENSE
  • README.md
  • REARCHITECTURE.md
  • Taskfile.yml
  • crates/browser-use-agent
  • … and 412 more

Open the folder on GitHubat commit 16cdd3a

Compare with similar skills

Browser Use Terminal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Use Terminal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Use Terminal this skillbrowser-use/terminal651—~2.1kAutomated safety check: PassMIT
Evoui BrowserSalmonbird/evoui-browser100—~4.5kAutomated safety check: PassApache-2.0
Browser UsePrismer-AI/PrismerCloud1.6k—~1.1kAutomated safety check: PassMIT
Browser Automation Edge Casesaden-hive/hive11k—~1.8kAutomated safety check: PassMIT
Skyvern Browser AutomationSkyvern-AI/skyvern23k—~1.9kAutomated safety check: PassAGPL-3.0
AI Search Hubminsight-ai-info/AI-Search-Hub1.3k—~1.3kAutomated safety check: PassNone

Similar skills

  • Evoui Browser

    Salmonbird/evoui-browser

    Use Evoui Browser as the default entry point for common, self-terminating web tasks that need a real browser through agent-browser, including navigation, page reading, clicks, forms, login flows…

    100 GitHub stars~4.5k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Use

    Prismer-AI/PrismerCloud

    LLM-driven browser automation via the pre-installed browser-use library (chromium already in the image).

    1.6k GitHub stars~1.1k tokensUpdated 7 days ago
    Productivity & AutomationAuto-check passed
  • Step-by-step procedure for debugging browser automation failures on complex sites such as LinkedIn, Twitter/X, single-page apps and Shadow DOM pages.

    11k GitHub stars~1.8k tokensUpdated 23 days ago
    Productivity & AutomationAuto-check passed
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • AI Search Hub

    minsight-ai-info/AI-Search-Hub

    Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.

    1.3k GitHub stars~1.3k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Tools

    981377660LMT/algorithm-study

    Interactive browser automation via Chrome DevTools Protocol.

    278 GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed

Works with

Questions about Browser Use Terminal

What does Browser Use Terminal do?

Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal. Browser Use Terminal is an agent skill from browser-use/terminal. Direct browser control via the Browser Use Terminal CLI.

When should I use Browser Use Terminal?

Browser Use Terminal fits situations like: the user wants to automate; interact with web pages — you drive the browser yourself with Python helpers.

How do I install Browser Use Terminal in Claude Code?

Run `npx skills add browser-use/terminal --skill browser-use-terminal -a claude-code`. Or copy the skill folder (the browser-use/terminal repository) into .claude/skills/browser-use-terminal in your project. Claude Code loads it when a task matches its description.

How do I install Browser Use Terminal in Codex?

Run `npx skills add browser-use/terminal --skill browser-use-terminal -a codex`. Or copy the skill folder (the browser-use/terminal repository) into .agents/skills/browser-use-terminal in your project. Codex loads it when a task matches its description.

Can I use Browser Use Terminal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-use/terminal --skill browser-use-terminal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use-terminal, .gemini/skills/browser-use-terminal, .github/skills/browser-use-terminal and .opencode/skills/browser-use-terminal in your project.

What does Browser Use Terminal need to run?

Going by SKILL.md and its folder, Browser Use Terminal needs credentials named BROWSER_USE_API_KEY. Our summary lists: Python 3; A credential in BROWSER_USE_API_KEY.

Does Browser Use Terminal access the network?

SKILL.md names 2 domains. In commands or code: docs.browser-use.com; the agent is likely to contact it when it follows the instructions. As links in the text: browser-use.com. This is read from the text; nothing was executed.

Is Browser Use Terminal safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Browser Use Terminal use?

Browser Use Terminal is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Use Terminal use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Use Terminal?

Skills that share tags, products or a category with Browser Use Terminal: Evoui Browser (Salmonbird/evoui-browser, 100 stars), Browser Use (Prismer-AI/PrismerCloud, 1.6k stars), Browser Automation Edge Cases (aden-hive/hive, 11k stars) and Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Use Terminal?

browser-use (a GitHub organization) maintains it in browser-use/terminal, which has 651 GitHub stars. The repository was last updated on August 16, 2026.

Source: browser-use/terminal on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.