Agent skill

Browser Use

by cacity in cacity/VideoHub

Automates browser interactions for web testing, form filling, screenshots, and data extraction.

MITAuto-check: warningsProductivity & Automation

Install Browser Use

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add cacity/VideoHub --skill browser-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cacity/VideoHub browser-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/browser-use .claude/skills/browser-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-use
GitHub stars
167
Used in
3 other repos
Token cost
~2.2k tokens
SKILL.md length
340 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Automates browser interactions for web testing, form filling, screenshots, and data extraction.

  • Works in 6 steps: Navigate: browser-use open — starts… → Inspect: browser-use state — returns… → Interact: use indices from state… → …
  • The user needs to navigate websites
  • SKILL.md covers Prerequisites, Core Workflow, Browser Modes and Commands, plus 9 more sections
  • Reaches github.com and abc.trycloudflare.com; needs BROWSER_USE_API_KEY

What it does

Browser Use is an agent skill from cacity/VideoHub. Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation and Forms and invoices. The repository describes itself as: VideoHub 是一款本地化多平台视频处理与智能剪辑工具,支持 YouTube、抖音/TikTok、Instagram、Bilibili 和 Twitter/X,提供视频下载、Whisper 转写、字幕翻译与润色、多模型 AI 配音、影视解说、故事剪辑、音乐卡点及剧集批量处理,并可通过 Codex、Claude Code… The licence is MIT.

When your agent uses it

  • The user needs to navigate websites
  • Interact with web pages
  • Take screenshots
  • Extract information from web pages

Example prompts

  • “Use the browser-use skill to automate browser interactions for web testing, form filling, screenshots, and data extraction”
  • “/browser-use”

Requirements

  • Python 3
  • A credential in BROWSER_USE_API_KEY
  • Pre-approved tools (allowed-tools): Bash(browser-use:*)

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Navigate: browser-use open — starts browser if needed
  2. Inspect: browser-use state — returns clickable elements with indices
  3. Interact: use indices from state (browser-use click 5, browser-use input 3 "text")
  4. Verify: browser-use state or browser-use screenshot to confirm
  5. Repeat: browser stays open between commands
  6. Cleanup: browser-use close when done

What it can do on your machine

Read from SKILL.md and the folder at commit d1da59c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(browser-use:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • abc.trycloudflare.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BROWSER_USE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Use loads about 2.2k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 340 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:33
    rofile "Default" open <url>      # Real Chrome with Default profile (existing logins/cookies)

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cacity/VideoHub at commit d1da59c, republished under its MIT licence (© cacity). 340 words, ~2,176 tokens.

Download SKILL.mdSave it as .claude/skills/browser-use/SKILL.md (or your agent's skills folder).
name
browser-use
description
Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.
allowed-tools
Bash(browser-use:*)

Browser Automation with browser-use CLI

The browser-use command provides fast, persistent browser automation. A background daemon keeps the browser open across commands, giving ~50ms latency per call.

Prerequisites

bash
browser-use doctor    # Verify installation

For setup details, see https://github.com/browser-use/browser-use/blob/main/browser_use/skill_cli/README.md

Core Workflow

  1. Navigate: browser-use open <url> — starts browser if needed
  2. Inspect: browser-use state — returns clickable elements with indices
  3. Interact: use indices from state (browser-use click 5, browser-use input 3 "text")
  4. Verify: browser-use state or browser-use screenshot to confirm
  5. Repeat: browser stays open between commands
  6. Cleanup: browser-use close when done

Browser Modes

bash
browser-use open <url>                         # Default: headless Chromium
browser-use --headed open <url>                # Visible window
browser-use --profile "Default" open <url>      # Real Chrome with Default profile (existing logins/cookies)
browser-use --profile "Profile 1" open <url>   # Real Chrome with named profile
browser-use --connect open <url>               # Auto-discover running Chrome via CDP
browser-use --cdp-url ws://localhost:9222/... open <url>  # Connect via CDP URL

--connect, --cdp-url, and --profile are mutually exclusive.

Commands

bash
# Navigation
browser-use open <url>                    # Navigate to URL
browser-use back                          # Go back in history
browser-use scroll down                   # Scroll down (--amount N for pixels)
browser-use scroll up                     # Scroll up
browser-use switch <tab>                  # Switch to tab by index
browser-use close-tab [tab]              # Close tab (current if no index)

# Page State — always run state first to get element indices
browser-use state                         # URL, title, clickable elements with indices
browser-use screenshot [path.png]         # Screenshot (base64 if no path, --full for full page)

# Interactions — use indices from state
browser-use click <index>                 # Click element by index
browser-use click <x> <y>                 # Click at pixel coordinates
browser-use type "text"                   # Type into focused element
browser-use input <index> "text"          # Click element, then type
browser-use keys "Enter"                  # Send keyboard keys (also "Control+a", etc.)
browser-use select <index> "option"       # Select dropdown option
browser-use upload <index> <path>         # Upload file to file input
browser-use hover <index>                 # Hover over element
browser-use dblclick <index>              # Double-click element
browser-use rightclick <index>            # Right-click element

# Data Extraction
browser-use eval "js code"                # Execute JavaScript, return result
browser-use get title                     # Page title
browser-use get html [--selector "h1"]    # Page HTML (or scoped to selector)
browser-use get text <index>              # Element text content
browser-use get value <index>             # Input/textarea value
browser-use get attributes <index>        # Element attributes
browser-use get bbox <index>              # Bounding box (x, y, width, height)

# Wait
browser-use wait selector "css"           # Wait for element (--state visible|hidden|attached|detached, --timeout ms)
browser-use wait text "text"              # Wait for text to appear

# Cookies
browser-use cookies get [--url <url>]     # Get cookies (optionally filtered)
browser-use cookies set <name> <value>    # Set cookie (--domain, --secure, --http-only, --same-site, --expires)
browser-use cookies clear [--url <url>]   # Clear cookies
browser-use cookies export <file>         # Export to JSON
browser-use cookies import <file>         # Import from JSON

# Python — persistent session with browser access
browser-use python "code"                 # Execute Python (variables persist across calls)
browser-use python --file script.py       # Run file
browser-use python --vars                 # Show defined variables
browser-use python --reset                # Clear namespace

# Session
browser-use close                         # Close browser and stop daemon
browser-use sessions                      # List active sessions
browser-use close --all                   # Close all sessions

The Python browser object provides: browser.url, browser.title, browser.html, browser.goto(url), browser.back(), browser.click(index), browser.type(text), browser.input(index, text), browser.keys(keys), browser.upload(index, path), browser.screenshot(path), browser.scroll(direction, amount), browser.wait(seconds).

Cloud API

bash
browser-use cloud connect                 # Provision cloud browser and connect
browser-use cloud connect --timeout 120 --proxy-country US  # With options
browser-use cloud login <api-key>         # Save API key (or set BROWSER_USE_API_KEY)
browser-use cloud logout                  # Remove API key
browser-use cloud v2 GET /browsers        # REST passthrough (v2 or v3)
browser-use cloud v2 POST /tasks '{"task":"...","url":"..."}'
browser-use cloud v2 poll <task-id>       # Poll task until done
browser-use cloud v2 --help               # Show API endpoints

cloud connect provisions a cloud browser, connects via CDP, and prints a live URL. browser-use close disconnects AND stops the cloud browser.

Tunnels

bash
browser-use tunnel <port>                 # Start Cloudflare tunnel (idempotent)
browser-use tunnel list                   # Show active tunnels
browser-use tunnel stop <port>            # Stop tunnel
browser-use tunnel stop --all             # Stop all tunnels

Profile Management

bash
browser-use profile list                  # List detected browsers and profiles
browser-use profile sync --all            # Sync profiles to cloud
browser-use profile update                # Download/update profile-use binary

Command Chaining

Commands can be chained with &&. The browser persists via the daemon, so chaining is safe and efficient.

bash
browser-use open https://example.com && browser-use state
browser-use input 5 "user@example.com" && browser-use input 6 "password" && browser-use click 7

Chain when you don't need intermediate output. Run separately when you need to parse state to discover indices first.

Common Workflows

Authenticated Browsing

When a task requires an authenticated site (Gmail, GitHub, internal tools), use Chrome profiles:

bash
browser-use profile list                           # Check available profiles
# Ask the user which profile to use, then:
browser-use --profile "Default" open https://github.com  # Already logged in
Connecting to Existing Chrome
bash
browser-use --connect open https://example.com     # Auto-discovers Chrome's CDP endpoint

Requires Chrome with remote debugging enabled. Falls back to probing ports 9222/9229.

Exposing Local Dev Servers
bash
browser-use tunnel 3000                            # → https://abc.trycloudflare.com
browser-use open https://abc.trycloudflare.com     # Browse the tunnel

Global Options

OptionDescription
--headedShow browser window
--profile [NAME]Use real Chrome (bare --profile uses "Default")
--connectAuto-discover running Chrome via CDP
--cdp-url <url>Connect via CDP URL (http:// or ws://)
--session NAMETarget a named session (default: "default")
--jsonOutput as JSON
--mcpRun as MCP server via stdin/stdout

Tips

  1. Always run state first to see available elements and their indices
  2. Use --headed for debugging to see what the browser is doing
  3. Sessions persist — browser stays open between commands
  4. CLI aliases: bu, browser, and browseruse all work

Troubleshooting

  • Browser won't start? browser-use close then browser-use --headed open <url>
  • Element not found? browser-use scroll down then browser-use state
  • Run diagnostics: browser-use doctor

Cleanup

bash
browser-use close                         # Close browser session
browser-use tunnel stop --all             # Stop tunnels (if any)

© cacity, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/browser-use of cacity/VideoHub.

Open the folder on GitHubat commit d1da59c

Used in 3 other repositories

We found 13 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in cacity/VideoHub, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Use this skillcacity/VideoHub1673 repos~2.2kAutomated safety check: WarnMIT
Agent Browserquran/quran.com-frontend-next1.9k41 repos~3.3kAutomated safety check: PassNone
Browse Nownowledge-co/community185—~619Automated safety check: PassNone
Sortedglebis/claude-skills388—~2kAutomated safety check: PassMIT
BrowserVibiumDev/vibium2.9k—~4.8kAutomated safety check: PassApache-2.0
Actionbookactionbook/actionbook1.6k—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 41 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Browse Now

    nowledge-co/community

    Control the user's actual browser through the browse-now CLI when a task needs authenticated pages, dynamic interaction, form filling, screenshots, or other browser automation that web search cannot…

    185 GitHub stars~619 tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Sorted

    glebis/claude-skills

    Automate getSorted.de (Sorted) for freelancer invoicing, expense tracking, and German tax submissions (VAT, ZM, annual returns).

    388 GitHub stars~2k tokensUpdated 11 days ago
    Productivity & AutomationAuto-check passed
  • Browser

    VibiumDev/vibium

    Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.

    2.9k GitHub stars~4.8k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Actionbook

    actionbook/actionbook

    Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.

    1.6k GitHub stars~1.5k tokensUpdated 29 days ago
    Productivity & AutomationAuto-check passed
  • Browser Use

    letta-ai/letta-code

    Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.

    3.5k GitHub stars~3.3k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from cacity/VideoHub

All 12 skills in this repo
  • Videohub Beat Editor

    cacity/VideoHub

    根据一段音频、歌曲或参考视频的节拍,为一个长视频、多个视频或素材目录自动建立镜头候选库,生成卡点剪辑计划,并批量渲染 16:9、3:4、4:3、9:16 等多画幅成片。支持固定镜头数量、强拍切换、歌词字幕、镜头替换、封面、标题、caption、hashtags 和完整 QA。用于“按音乐卡点剪视频”“给音频和素材批量做卡点视频”“检测强拍并自动选镜头”“同一计划输出多个画幅”等任务。

    167 GitHub stars~626 tokensUpdated 5 days ago
    Auto-check passed
  • 为 VideoHub 的影视解说、连续剧、电影、卡点视频和短视频制作可在个人主页小缩略图中辨认的封面。输入剧照、视频帧或已有底图,突出剧名、集数和简短看点,统一生成 9:16、3:4、4:3、16:9 封面与缩略图预览。用于“做封面”“修改封面”“加大集数”“生成横版和竖版封面”“沿用上一集封面模板”“制作抖音或视频号缩略图”等任务。

    167 GitHub stars~453 tokensUpdated 5 days ago
    Auto-check passed
  • 把电影、电视剧或短剧素材制作成第三者旁白主导、关键影视原声点睛的中文解说视频,并生成抖音竖版封面、标题候选、50-100 字文案、话题和完整发布包。复用 videohub-story-editor 的证据提取、剧情理解、剪辑、后置翻译、TTS…

    167 GitHub stars~1.6k tokensUpdated 5 days ago
    Auto-check passed
  • Videohub Story Editor

    cacity/VideoHub

    把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…

    167 GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check: notes
  • Videohub Youtube

    cacity/VideoHub

    处理 YouTube、Twitter(X)、Bilibili 和本地音视频/文本的转写、字幕、翻译与总结。优先复用 src/youtubetranscriber.py 现有 CLI。

    167 GitHub stars~550 tokensUpdated 5 days ago
    Auto-check passed
  • Videohub Douyin

    cacity/VideoHub

    下载抖音单视频或用户主页作品,复用 src/douyincli.py。适合处理抖音分享链接、短链接、标准视频链接和用户主页链接。

    167 GitHub stars~213 tokensUpdated 5 days ago
    Auto-check passed

Questions about Browser Use

What does Browser Use do?

Automates browser interactions for web testing, form filling, screenshots, and data extraction. Browser Use is an agent skill from cacity/VideoHub. Automates browser interactions for web testing, form filling, screenshots, and data extraction.

When should I use Browser Use?

Browser Use fits situations like: the user needs to navigate websites; interact with web pages; take screenshots; extract information from web pages.

How do I install Browser Use in Claude Code?

Run `npx skills add cacity/VideoHub --skill browser-use -a claude-code`. Or copy the skill folder (.agents/skills/browser-use in cacity/VideoHub) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.

How do I install Browser Use in Codex?

Run `npx skills add cacity/VideoHub --skill browser-use -a codex`. Or copy the skill folder (.agents/skills/browser-use in cacity/VideoHub) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.

Can I use Browser Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cacity/VideoHub --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.

What does Browser Use need to run?

Going by SKILL.md and its folder, Browser Use needs credentials named BROWSER_USE_API_KEY. Our summary lists: Python 3; A credential in BROWSER_USE_API_KEY. Its frontmatter pre-approves these tools: Bash(browser-use:*).

Does Browser Use access the network?

SKILL.md names 2 domains. In commands or code: github.com and abc.trycloudflare.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Browser Use safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Browser Use use?

Browser Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Use use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Use?

Skills that share tags, products or a category with Browser Use: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Browse Now (nowledge-co/community, 185 stars), Sorted (glebis/claude-skills, 388 stars) and Browser (VibiumDev/vibium, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Use?

cacity (a GitHub user) maintains it in cacity/VideoHub, which has 167 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 2, 2026.

Source: cacity/VideoHub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.