Evoui Browser
Salmonbird/evoui-browser
Use Evoui Browser as the default entry point for common, self-terminating web tasks that need a real browser through agent-browser, including navigation, page reading, clicks, forms, login flows…
Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal.
$ npx skills add browser-use/terminal --skill browser-use-terminal -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install browser-use/terminal browser-use-terminal --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .claude/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/terminal --skill browser-use-terminal -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install browser-use/terminal browser-use-terminal --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .agents/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/terminal --skill browser-use-terminal -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install browser-use/terminal browser-use-terminal --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .cursor/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/terminal --skill browser-use-terminal -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install browser-use/terminal browser-use-terminal --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .gemini/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install browser-use/terminal browser-use-terminalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add browser-use/terminal --skill browser-use-terminal -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .github/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/terminal --skill browser-use-terminal -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install browser-use/terminal browser-use-terminal --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser-use-terminal" agent skill from https://github.com/browser-use/terminal/tree/main into .opencode/skills/browser-use-terminal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use-terminal", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser-use-terminalDirect browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal.
Browser Use Terminal is an agent skill from browser-use/terminal. Direct browser control via the Browser Use Terminal CLI. Use when the user wants to automate, scrape, test, or interact with web pages — you drive the browser yourself with Python helpers.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 432 other files, including scripts (for example `.github/workflows/browser-use-core-publish.yml`, `.github/workflows/browser-use-core-wheels.yml` and `.github/workflows/release.yml`).
It sits in Productivity & Automation, covering Browser automation. It works with Python. The repository describes itself as: Terminal UI to get stuff done in the browser. The licence is MIT.
Read from SKILL.md and the folder at commit 16cdd3a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
docs.browser-use.comAlso links to:
browser-use.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BROWSER_USE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser Use Terminal loads about 2.1k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 865 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from browser-use/terminal at commit 16cdd3a, republished under its MIT licence (© browser-use). 865 words, ~2,111 tokens.
.claude/skills/browser-use-terminal/SKILL.md (or your agent's skills folder). This skill also uses 429 other files; get the full folder from GitHub.Direct browser control via CDP — you are the agent; you drive the browser. For setup, install, or connection problems, read https://browser-use.com/skill (agent setup instructions) or https://docs.browser-use.com/open-source/browser-use-terminal (full docs).
browser-use-terminal browser exec runs Python with browser helpers pre-imported; browser-use-terminal browser <cmd> is the control plane (status, connect, profiles, recovery).
browser-use-terminal browser exec <<'PY'
new_tab("https://docs.browser-use.com")
wait_for_load()
print(page_info())
PYnew_tab(url), not goto_url(url) — goto runs in the user's active tab and clobbers their work.exec is a fresh interpreter against the same live browser.--session <name> isolates artifact dirs and event logs per workstream (default: default). --timeout <secs> bounds one exec (default 300).capture_screenshot() saves a PNG and returns its absolute path. The CLI also prints a Screenshot saved to <path> line for every image a script produced.
browser-use-terminal browser exec <<'PY'
print(capture_screenshot())
PYTo view a screenshot, use your file-reading tool on the printed path:
Read tool on the path.view_image tool with {"path": "<path>"}. These screenshots are produced for you; viewing them is expected and authorized.read tool on the path (requires a vision-capable model).read_file tool on the path.print(page_info()), js(...) extraction, wait_for_element(...).Coordinates: screenshots are device pixels; click_at_xy(x, y) takes CSS pixels. Divide coordinates you read off the image by js("window.devicePixelRatio") first. Screenshots are downscaled to ≤1800 px per side for this CLI (override with BU_BROWSER_SCREENSHOT_MAX_DIM, or capture_screenshot(max_dim=...)).
After every meaningful action, re-screenshot before assuming it worked.
Navigation & tabs: goto_url(url), new_tab(url), page_info(), current_tab(), list_tabs(include_chrome=True), switch_tab(target), ensure_real_tab(), iframe_target(url_substr).
Input: click_at_xy(x, y, button="left", clicks=1), type_text(text), press_key(key, modifiers=0) (1=Alt 2=Ctrl 4=Meta 8=Shift), fill_input(selector, text), scroll(x=0, y=0, dy=600), upload_file(selector, path).
Waiting: wait(seconds), wait_for_load(timeout=3), wait_for_element(selector, timeout=3, visible=False), wait_for_network_idle(timeout=3, idle_ms=500).
Visual: capture_screenshot(label="...", full=False, max_dim=None), screenshot(), screenshot_clip(label, x, y, w, h), note(caption).
Escape hatches: js(expression) (auto-wraps top-level return), cdp("Domain.method", **params) (raw CDP), cdp_batch(calls), drain_events().
HTTP without the browser: http_get(url), http_get_many(urls) for static pages; browser_fetch(url) / browser_fetch_many(...) to fetch with the page's cookies/session.
Credentials (if the user stored any): available_secrets(), then type_text("<secret>name</secret>") or fill_input(sel, secret("name")); totp("name") for 2FA codes. Values are placeholder-substituted — you never see them. is_logged_out(), email_inbox() / email_message(id) for email-code flows.
Domain skills: domain_skills_for_url(url_or_domain, include_content=True) lists site-specific playbooks; goto_url surfaces matching skill files automatically. Read them before inventing selectors or flows on a complex site.
browser-use-terminal browser status --json
browser-use-terminal browser connect # uses the remembered preference
browser-use-terminal browser connect local # user's already-running Chrome (CDP)
browser-use-terminal browser connect managed --headless # disposable CLI-owned browser
browser-use-terminal browser preference use local|cloud|managed-headless
browser-use-terminal browser remote start # Browser Use cloud browser (needs BROWSER_USE_API_KEY)
browser-use-terminal browser doctor
browser-use-terminal browser recover reconnect-websocket
browser-use-terminal browser recover stop-owned-browser # stop the persistent managed browser
browser-use-terminal browser recover stop-owned-remote # stop the cloud browser (stops billing)
browser-use-terminal browser daemon status|stop|logs # the background daemon holding the connectionA background daemon (auto-started, one per state dir) holds the CDP connection across your commands, so the browser — and in local mode, Chrome's granted debugging permission — persists between invocations. Managed and cloud browsers also survive daemon restarts; later calls reattach instead of relaunching. Stop browsers with the recover commands above when the user is done (cloud browsers bill until stopped or timed out).
exec auto-connects, so you rarely need these. Reach for them when status shows a problem or the user asks for a specific browser.status: "needs-user-action" (e.g. pick a Chrome profile, click Allow in Chrome's permission popup, enable the remote-debugging checkbox), show the user_prompt to the user verbatim and wait — do not guess.chrome://inspect/#remote-debugging → tick "Allow remote debugging". browser local setup walks the user through it.capture_screenshot() → view the image → decide whether you need a click, a selector, or more navigation.click_at_xy(x, y) → screenshot to verify. Suppress the locate-then-click reflex — no getBoundingClientRect, no selector hunts. Hit-testing happens in Chrome's browser process, so coordinate clicks pass through iframes / shadow DOM / cross-origin without extra work.fill_input, js) only when the target has no visible geometry (hidden input, 0×0 node) or coordinate clicks demonstrably don't work.http_get_many(urls) — no browser needed. Logged-in pages: browser_fetch(url) rides the real session.wait_for_load(). SPAs report complete before they render — follow with wait_for_element(...).ensure_real_tab().print(page_info()) is the cheapest "is this alive?" check; screenshots are the default way to verify visible actions.chrome:// internals are fake page targets — list_tabs(include_chrome=False).page_info() surfaces an open JS dialog as {"dialog": ...} — handle it (cdp("Page.handleJavaScriptDialog", accept=True)) before anything else.nav_policy(url) tells you before you burn a click. A blocked navigation is policy, not a bug — tell the user.exec small and observable rather than one mega-script. Long extraction loops: print progress as you go — stdout is captured even on timeout.browser-use-terminal browser domain skills --domain <site> --json --include-content first — that's where site playbooks live.© browser-use, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 429 other files (scripts) in the repository root of browser-use/terminal.
Open the folder on GitHubat commit 16cdd3a
Browser Use Terminal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser Use Terminal this skillbrowser-use/terminal | 651 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Evoui BrowserSalmonbird/evoui-browser | 100 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Browser UsePrismer-AI/PrismerCloud | 1.6k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Browser Automation Edge Casesaden-hive/hive | 11k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Skyvern Browser AutomationSkyvern-AI/skyvern | 23k | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | |
| AI Search Hubminsight-ai-info/AI-Search-Hub | 1.3k | — | ~1.3k | Automated safety check: Pass | None |
Salmonbird/evoui-browser
Use Evoui Browser as the default entry point for common, self-terminating web tasks that need a real browser through agent-browser, including navigation, page reading, clicks, forms, login flows…
Prismer-AI/PrismerCloud
LLM-driven browser automation via the pre-installed browser-use library (chromium already in the image).
aden-hive/hive
Step-by-step procedure for debugging browser automation failures on complex sites such as LinkedIn, Twitter/X, single-page apps and Shadow DOM pages.
Skyvern-AI/skyvern
Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.
minsight-ai-info/AI-Search-Hub
Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.
981377660LMT/algorithm-study
Interactive browser automation via Chrome DevTools Protocol.
Works with
Categories
Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal. Browser Use Terminal is an agent skill from browser-use/terminal. Direct browser control via the Browser Use Terminal CLI.
Browser Use Terminal fits situations like: the user wants to automate; interact with web pages — you drive the browser yourself with Python helpers.
Run `npx skills add browser-use/terminal --skill browser-use-terminal -a claude-code`. Or copy the skill folder (the browser-use/terminal repository) into .claude/skills/browser-use-terminal in your project. Claude Code loads it when a task matches its description.
Run `npx skills add browser-use/terminal --skill browser-use-terminal -a codex`. Or copy the skill folder (the browser-use/terminal repository) into .agents/skills/browser-use-terminal in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-use/terminal --skill browser-use-terminal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use-terminal, .gemini/skills/browser-use-terminal, .github/skills/browser-use-terminal and .opencode/skills/browser-use-terminal in your project.
Going by SKILL.md and its folder, Browser Use Terminal needs credentials named BROWSER_USE_API_KEY. Our summary lists: Python 3; A credential in BROWSER_USE_API_KEY.
SKILL.md names 2 domains. In commands or code: docs.browser-use.com; the agent is likely to contact it when it follows the instructions. As links in the text: browser-use.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Browser Use Terminal is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browser Use Terminal: Evoui Browser (Salmonbird/evoui-browser, 100 stars), Browser Use (Prismer-AI/PrismerCloud, 1.6k stars), Browser Automation Edge Cases (aden-hive/hive, 11k stars) and Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
browser-use (a GitHub organization) maintains it in browser-use/terminal, which has 651 GitHub stars. The repository was last updated on August 16, 2026.
Source: browser-use/terminal on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.