Agent Browser
quran/quran.com-frontend-next
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add cacity/VideoHub --skill browser-use -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cacity/VideoHub browser-use --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/browser-use .claude/skills/browser-use && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .claude/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-useType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cacity/VideoHub --skill browser-use -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cacity/VideoHub browser-use --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/browser-use .agents/skills/browser-use && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .agents/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cacity/VideoHub --skill browser-use -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cacity/VideoHub browser-use --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/browser-use .cursor/skills/browser-use && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .cursor/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cacity/VideoHub.git --path .agents/skills/browser-use--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cacity/VideoHub --skill browser-use -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cacity/VideoHub browser-use --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/browser-use .gemini/skills/browser-use && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .gemini/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cacity/VideoHub browser-useInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cacity/VideoHub --skill browser-use -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/browser-use .github/skills/browser-use && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .github/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cacity/VideoHub --skill browser-use -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cacity/VideoHub browser-use --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cacity/VideoHub.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/browser-use .opencode/skills/browser-use && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/cacity/VideoHub/tree/main/.agents/skills/browser-use into .opencode/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser-useAutomates browser interactions for web testing, form filling, screenshots, and data extraction.
Browser Use is an agent skill from cacity/VideoHub. Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Productivity & Automation, covering Browser automation and Forms and invoices. The repository describes itself as: VideoHub 是一款本地化多平台视频处理与智能剪辑工具,支持 YouTube、抖音/TikTok、Instagram、Bilibili 和 Twitter/X,提供视频下载、Whisper 转写、字幕翻译与润色、多模型 AI 配音、影视解说、故事剪辑、音乐卡点及剧集批量处理,并可通过 Codex、Claude Code… The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d1da59c. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(browser-use:*)From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comabc.trycloudflare.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BROWSER_USE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser Use loads about 2.2k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 340 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
rofile "Default" open <url> # Real Chrome with Default profile (existing logins/cookies)Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cacity/VideoHub at commit d1da59c, republished under its MIT licence (© cacity). 340 words, ~2,176 tokens.
.claude/skills/browser-use/SKILL.md (or your agent's skills folder).The browser-use command provides fast, persistent browser automation. A background daemon keeps the browser open across commands, giving ~50ms latency per call.
browser-use doctor # Verify installationFor setup details, see https://github.com/browser-use/browser-use/blob/main/browser_use/skill_cli/README.md
browser-use open <url> — starts browser if neededbrowser-use state — returns clickable elements with indicesbrowser-use click 5, browser-use input 3 "text")browser-use state or browser-use screenshot to confirmbrowser-use close when donebrowser-use open <url> # Default: headless Chromium
browser-use --headed open <url> # Visible window
browser-use --profile "Default" open <url> # Real Chrome with Default profile (existing logins/cookies)
browser-use --profile "Profile 1" open <url> # Real Chrome with named profile
browser-use --connect open <url> # Auto-discover running Chrome via CDP
browser-use --cdp-url ws://localhost:9222/... open <url> # Connect via CDP URL--connect, --cdp-url, and --profile are mutually exclusive.
# Navigation
browser-use open <url> # Navigate to URL
browser-use back # Go back in history
browser-use scroll down # Scroll down (--amount N for pixels)
browser-use scroll up # Scroll up
browser-use switch <tab> # Switch to tab by index
browser-use close-tab [tab] # Close tab (current if no index)
# Page State — always run state first to get element indices
browser-use state # URL, title, clickable elements with indices
browser-use screenshot [path.png] # Screenshot (base64 if no path, --full for full page)
# Interactions — use indices from state
browser-use click <index> # Click element by index
browser-use click <x> <y> # Click at pixel coordinates
browser-use type "text" # Type into focused element
browser-use input <index> "text" # Click element, then type
browser-use keys "Enter" # Send keyboard keys (also "Control+a", etc.)
browser-use select <index> "option" # Select dropdown option
browser-use upload <index> <path> # Upload file to file input
browser-use hover <index> # Hover over element
browser-use dblclick <index> # Double-click element
browser-use rightclick <index> # Right-click element
# Data Extraction
browser-use eval "js code" # Execute JavaScript, return result
browser-use get title # Page title
browser-use get html [--selector "h1"] # Page HTML (or scoped to selector)
browser-use get text <index> # Element text content
browser-use get value <index> # Input/textarea value
browser-use get attributes <index> # Element attributes
browser-use get bbox <index> # Bounding box (x, y, width, height)
# Wait
browser-use wait selector "css" # Wait for element (--state visible|hidden|attached|detached, --timeout ms)
browser-use wait text "text" # Wait for text to appear
# Cookies
browser-use cookies get [--url <url>] # Get cookies (optionally filtered)
browser-use cookies set <name> <value> # Set cookie (--domain, --secure, --http-only, --same-site, --expires)
browser-use cookies clear [--url <url>] # Clear cookies
browser-use cookies export <file> # Export to JSON
browser-use cookies import <file> # Import from JSON
# Python — persistent session with browser access
browser-use python "code" # Execute Python (variables persist across calls)
browser-use python --file script.py # Run file
browser-use python --vars # Show defined variables
browser-use python --reset # Clear namespace
# Session
browser-use close # Close browser and stop daemon
browser-use sessions # List active sessions
browser-use close --all # Close all sessionsThe Python browser object provides: browser.url, browser.title, browser.html, browser.goto(url), browser.back(), browser.click(index), browser.type(text), browser.input(index, text), browser.keys(keys), browser.upload(index, path), browser.screenshot(path), browser.scroll(direction, amount), browser.wait(seconds).
browser-use cloud connect # Provision cloud browser and connect
browser-use cloud connect --timeout 120 --proxy-country US # With options
browser-use cloud login <api-key> # Save API key (or set BROWSER_USE_API_KEY)
browser-use cloud logout # Remove API key
browser-use cloud v2 GET /browsers # REST passthrough (v2 or v3)
browser-use cloud v2 POST /tasks '{"task":"...","url":"..."}'
browser-use cloud v2 poll <task-id> # Poll task until done
browser-use cloud v2 --help # Show API endpointscloud connect provisions a cloud browser, connects via CDP, and prints a live URL. browser-use close disconnects AND stops the cloud browser.
browser-use tunnel <port> # Start Cloudflare tunnel (idempotent)
browser-use tunnel list # Show active tunnels
browser-use tunnel stop <port> # Stop tunnel
browser-use tunnel stop --all # Stop all tunnelsbrowser-use profile list # List detected browsers and profiles
browser-use profile sync --all # Sync profiles to cloud
browser-use profile update # Download/update profile-use binaryCommands can be chained with &&. The browser persists via the daemon, so chaining is safe and efficient.
browser-use open https://example.com && browser-use state
browser-use input 5 "user@example.com" && browser-use input 6 "password" && browser-use click 7Chain when you don't need intermediate output. Run separately when you need to parse state to discover indices first.
When a task requires an authenticated site (Gmail, GitHub, internal tools), use Chrome profiles:
browser-use profile list # Check available profiles
# Ask the user which profile to use, then:
browser-use --profile "Default" open https://github.com # Already logged inbrowser-use --connect open https://example.com # Auto-discovers Chrome's CDP endpointRequires Chrome with remote debugging enabled. Falls back to probing ports 9222/9229.
browser-use tunnel 3000 # → https://abc.trycloudflare.com
browser-use open https://abc.trycloudflare.com # Browse the tunnel| Option | Description |
|---|---|
--headed | Show browser window |
--profile [NAME] | Use real Chrome (bare --profile uses "Default") |
--connect | Auto-discover running Chrome via CDP |
--cdp-url <url> | Connect via CDP URL (http:// or ws://) |
--session NAME | Target a named session (default: "default") |
--json | Output as JSON |
--mcp | Run as MCP server via stdin/stdout |
state first to see available elements and their indices--headed for debugging to see what the browser is doingbu, browser, and browseruse all workbrowser-use close then browser-use --headed open <url>browser-use scroll down then browser-use statebrowser-use doctorbrowser-use close # Close browser session
browser-use tunnel stop --all # Stop tunnels (if any)© cacity, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/browser-use of cacity/VideoHub.
Open the folder on GitHubat commit d1da59c
We found 13 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in cacity/VideoHub, which our catalogue first saw on October 7, 2026.
Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser Use this skillcacity/VideoHub | 167 | 3 repos | ~2.2k | Automated safety check: Warn | MIT | |
| Agent Browserquran/quran.com-frontend-next | 1.9k | 41 repos | ~3.3k | Automated safety check: Pass | None | |
| Browse Nownowledge-co/community | 185 | — | ~619 | Automated safety check: Pass | None | |
| Sortedglebis/claude-skills | 388 | — | ~2k | Automated safety check: Pass | MIT | |
| BrowserVibiumDev/vibium | 2.9k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| Actionbookactionbook/actionbook | 1.6k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 |
quran/quran.com-frontend-next
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
nowledge-co/community
Control the user's actual browser through the browse-now CLI when a task needs authenticated pages, dynamic interaction, form filling, screenshots, or other browser automation that web search cannot…
glebis/claude-skills
Automate getSorted.de (Sorted) for freelancer invoicing, expense tracking, and German tax submissions (VAT, ZM, annual returns).
VibiumDev/vibium
Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.
actionbook/actionbook
Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.
letta-ai/letta-code
Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.
cacity/VideoHub
根据一段音频、歌曲或参考视频的节拍,为一个长视频、多个视频或素材目录自动建立镜头候选库,生成卡点剪辑计划,并批量渲染 16:9、3:4、4:3、9:16 等多画幅成片。支持固定镜头数量、强拍切换、歌词字幕、镜头替换、封面、标题、caption、hashtags 和完整 QA。用于“按音乐卡点剪视频”“给音频和素材批量做卡点视频”“检测强拍并自动选镜头”“同一计划输出多个画幅”等任务。
cacity/VideoHub
为 VideoHub 的影视解说、连续剧、电影、卡点视频和短视频制作可在个人主页小缩略图中辨认的封面。输入剧照、视频帧或已有底图,突出剧名、集数和简短看点,统一生成 9:16、3:4、4:3、16:9 封面与缩略图预览。用于“做封面”“修改封面”“加大集数”“生成横版和竖版封面”“沿用上一集封面模板”“制作抖音或视频号缩略图”等任务。
cacity/VideoHub
把电影、电视剧或短剧素材制作成第三者旁白主导、关键影视原声点睛的中文解说视频,并生成抖音竖版封面、标题候选、50-100 字文案、话题和完整发布包。复用 videohub-story-editor 的证据提取、剧情理解、剪辑、后置翻译、TTS…
cacity/VideoHub
把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…
cacity/VideoHub
处理 YouTube、Twitter(X)、Bilibili 和本地音视频/文本的转写、字幕、翻译与总结。优先复用 src/youtubetranscriber.py 现有 CLI。
cacity/VideoHub
下载抖音单视频或用户主页作品,复用 src/douyincli.py。适合处理抖音分享链接、短链接、标准视频链接和用户主页链接。
Automates browser interactions for web testing, form filling, screenshots, and data extraction. Browser Use is an agent skill from cacity/VideoHub. Automates browser interactions for web testing, form filling, screenshots, and data extraction.
Browser Use fits situations like: the user needs to navigate websites; interact with web pages; take screenshots; extract information from web pages.
Run `npx skills add cacity/VideoHub --skill browser-use -a claude-code`. Or copy the skill folder (.agents/skills/browser-use in cacity/VideoHub) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cacity/VideoHub --skill browser-use -a codex`. Or copy the skill folder (.agents/skills/browser-use in cacity/VideoHub) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cacity/VideoHub --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.
Going by SKILL.md and its folder, Browser Use needs credentials named BROWSER_USE_API_KEY. Our summary lists: Python 3; A credential in BROWSER_USE_API_KEY. Its frontmatter pre-approves these tools: Bash(browser-use:*).
SKILL.md names 2 domains. In commands or code: github.com and abc.trycloudflare.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.
Browser Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browser Use: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Browse Now (nowledge-co/community, 185 stars), Sorted (glebis/claude-skills, 388 stars) and Browser (VibiumDev/vibium, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cacity (a GitHub user) maintains it in cacity/VideoHub, which has 167 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 2, 2026.
Source: cacity/VideoHub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.