Agent skill

Browser Use

by aAAaqwq in aAAaqwq/AGI-Super-Team

AI驱动的智能浏览器自动化工具。使用LLM理解页面并自动执行任务,比传统Playwright更智能、更省token。适用于复杂交互、动态页面、需要智能决策的浏览器操作。Chrome浏览器优先。

MITAuto-check: notesProductivity & Automation

Install Browser Use

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill browser-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team browser-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser-use .claude/skills/browser-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-use
GitHub stars
105
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
128 words
Files
1
Skills in repo
161
Repo updated
First seen
Licence
MIT

At a glance

AI驱动的智能浏览器自动化工具。使用LLM理解页面并自动执行任务,比传统Playwright更智能、更省token。适用于复杂交互、动态页面、需要智能决策的浏览器操作。Chrome浏览器优先。

  • Works in 4 steps: 任务描述要清晰 → 使用结构化输出 → 错误处理 → …
  • Tasks that involve Browser automation
  • SKILL.md covers 概述, ⚠️ 资源清理原则(强制), 安装状态 and 快速开始, plus 7 more sections
  • Calls python3 and pip; reaches polymarket.com and ai.9w7.cn; needs XSC_API_KEY

What it does

Browser Use is an agent skill from aAAaqwq/AGI-Super-Team. AI驱动的智能浏览器自动化工具。使用LLM理解页面并自动执行任务,比传统Playwright更智能、更省token。适用于复杂交互、动态页面、需要智能决策的浏览器操作。Chrome浏览器优先。

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation and Browser testing. It works with Playwright. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • Tasks that involve Browser automation
  • Tasks that involve Browser testing

Example prompts

  • “/browser-use”

Requirements

  • Python 3
  • A credential in XSC_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Exec, Read, Write

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 任务描述要清晰
  2. 使用结构化输出
  3. 错误处理
  4. 复用登录态

What it can do on your machine

Read from SKILL.md and the folder at commit 331ecd3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Exec
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • polymarket.com
    • ai.9w7.cn

    Also links to:

    • docs.browser-use.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • XSC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Use loads about 1.6k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 128 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Exec, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 331ecd3, republished under its MIT licence (© aAAaqwq). 128 words, ~1,593 tokens.

Download SKILL.mdSave it as .claude/skills/browser-use/SKILL.md (or your agent's skills folder).
name
browser-use
description
AI驱动的智能浏览器自动化工具。使用LLM理解页面并自动执行任务,比传统Playwright更智能、更省token。适用于复杂交互、动态页面、需要智能决策的浏览器操作。Chrome浏览器优先。
allowed-tools
Bash, Exec, Read, Write
author
Daniel Li

browser-use 智能浏览器自动化

概述

browser-use 是一个 AI 驱动的浏览器自动化工具,它使用 LLM 来:

  • 理解网页内容
  • 智能决策下一步操作
  • 自动完成任务

与 Playwright 的区别:

特性Playwrightbrowser-use
控制方式预编程脚本AI智能决策
适应性页面变化需重写自动适应
Token消耗较低但需调试智能精简
复杂交互需精确选择器自然语言描述
维护成本高低

⚠️ 资源清理原则(强制)

所有涉及浏览器的 cron 任务完成后,必须自动关闭 Chrome 进程!

python
import asyncio
from browser_use import Agent

async def main():
    agent = Agent(task="...", llm=llm)
    result = await agent.run()

    # ⚠️ 任务结束后必须显式关闭浏览器
    if hasattr(agent, 'browser') and agent.browser:
        await agent.browser.close()

    # ⚠️ 推荐在脚本结束时强制清理残留进程
    import subprocess
    subprocess.run(['pkill', '-f', 'chrome'], capture_output=True)

    return result

原因: 避免内存泄漏和资源占用,防止 Gateway CPU 100% 过载

安装状态

✅ 已安装:

  • browser-use 0.11.11
  • browser-use-sdk 2.0.15

快速开始

基本用法
python
import asyncio
from browser_use import Agent
from langchain_openai import ChatOpenAI

async def main():
    agent = Agent(
        task="打开 polymarket.com,查看 Fed 利率市场",
        llm=ChatOpenAI(model="gpt-4o"),
    )
    result = await agent.run()
    print(result)

asyncio.run(main())
使用自定义 LLM(推荐配置)

⚠️ 重要:your-provider API 是 Anthropic 格式,不是 OpenAI 格式!必须使用 ChatAnthropic。

python
from browser_use.llm.anthropic.chat import ChatAnthropic

# your-provider API(Anthropic 兼容)✅ 推荐
llm = ChatAnthropic(
    model="claude-sonnet-4-6",
    base_url="https://your-anthropic-proxy.example.com",  # 注意:不加 /v1
    api_key="your-api-key",  # 或从 pass show api/your-provider 获取
)

agent = Agent(
    task="你的任务",
    llm=llm,
)
python
# ❌ 错误用法:不要用 ChatOpenAI + your-provider
# from browser_use.llm.openai.chat import ChatOpenAI  # 这个不行!your-provider 不支持 OpenAI 格式
python
# 如果使用 OpenAI 兼容 API(如 Provider-B),用 ChatOpenAI:
from browser_use.llm.openai.chat import ChatOpenAI
llm = ChatOpenAI(
    model="gpt-4o",
    base_url="https://ai.9w7.cn/v1",
    api_key="your-api-key",
)
使用 Chrome 浏览器
python
from browser_use import Agent, BrowserProfile, BrowserSession

# 配置使用 Chrome
profile = BrowserProfile(
    executable_path="/usr/bin/google-chrome-stable",
    headless=True,
    disable_security=True,
)
session = BrowserSession(browser_profile=profile)

agent = Agent(
    task="任务描述",
    llm=llm,
    browser_session=session,
)
保存/加载登录态
python
# 保存登录态
await browser_context.save_storage_state(path="polymarket_auth.json")

# 加载登录态
context = BrowserContextConfig(
    storage_state="polymarket_auth.json"
)

常用配置

最小化 Token 消耗
python
agent = Agent(
    task="任务",
    llm=llm,
    use_vision=False,  # 禁用视觉,减少token
    max_actions_per_step=3,  # 限制每步操作数
    message_compaction=True,  # 消息压缩
)
调试模式
python
agent = Agent(
    task="任务",
    llm=llm,
    headless=False,  # 显示浏览器
    slow_mo=1000,  # 慢放,每步延迟1秒
    save_conversation_path="debug_log/",  # 保存日志
)
提取结构化数据
python
from pydantic import BaseModel

class MarketData(BaseModel):
    question: str
    yes_price: float
    no_price: float

agent = Agent(
    task="获取 Polymarket Fed 利率市场数据",
    llm=llm,
    output_model_schema=MarketData,
)
result = await agent.run()
# result 将是 MarketData 类型

Polymarket 集成

查看市场
python
agent = Agent(
    task="""
    1. 打开 https://polymarket.com/event/fed-decision-in-march-885
    2. 提取以下信息:
       - 市场问题
       - Yes 价格
       - No 价格
       - 交易量
    3. 返回 JSON 格式数据
    """,
    llm=llm,
)
执行交易
python
agent = Agent(
    task="""
    1. 打开 Polymarket
    2. 连接钱包(如果需要)
    3. 导航到 Fed 利率市场
    4. 买入 $0.60 的 No(价格 ≥ 0.85)
    5. 确认交易
    """,
    llm=llm,
    sensitive_data={
        "wallet_address": "0x...",
    }
)

最佳实践

1. 任务描述要清晰
python
# 好 ✅
task="打开 polymarket.com,找到 Fed 利率市场,提取 Yes/No 价格"

# 差 ❌
task="帮我看看那个市场"
2. 使用结构化输出
python
from pydantic import BaseModel

class TradingResult(BaseModel):
    success: bool
    market: str
    action: str  # "buy_yes", "buy_no"
    amount: float
    price: float
    tx_hash: str | None

agent = Agent(
    task="执行交易...",
    output_model_schema=TradingResult,
)
3. 错误处理
python
try:
    result = await agent.run()
except Exception as e:
    print(f"任务失败: {e}")
    # 可以使用 browser tool 作为后备
4. 复用登录态
python
# 第一次登录后保存
# 后续直接加载,避免重复登录

browser_context = BrowserContextConfig(
    storage_state="~/.playwright-data/polymarket/auth.json"
)

Token 消耗优化

对比(相同任务)
工具平均 Token原因
browser tool~5000-10000每次快照全页面
Playwright~1000-2000需多次调试
browser-use~2000-4000AI精简决策
优化技巧
  1. 禁用视觉:use_vision=False
  2. 限制历史:max_history_items=10
  3. 压缩消息:message_compaction=True
  4. 减少步骤:max_actions_per_step=3
  5. 使用 Flash 模式:flash_mode=True(快速模式)

完整示例

python
import asyncio
from browser_use import Agent
from langchain_openai import ChatOpenAI
from pydantic import BaseModel
import os

class MarketInfo(BaseModel):
    question: str
    yes_price: float
    no_price: float
    volume: str

async def check_polymarket():
    # 使用本地 LLM API
    llm = ChatOpenAI(
        model="claude-3-5-sonnet-20241022",
        base_url="https://your-anthropic-proxy.example.com",
        api_key=os.environ.get("XSC_API_KEY"),
    )
    
    agent = Agent(
        task="""
        访问 Polymarket Fed 利率市场:
        https://polymarket.com/event/fed-decision-in-march-885
        
        提取并返回:
        - 市场问题
        - Yes 价格(0-1)
        - No 价格(0-1)
        - 24h 交易量
        """,
        llm=llm,
        output_model_schema=MarketInfo,
        use_vision=False,
        max_actions_per_step=5,
    )
    
    result = await agent.run()
    return result

if __name__ == "__main__":
    asyncio.run(check_polymarket())

资源

快速命令

bash
# 安装(已完成)
pip install browser-use

# 运行脚本
python3 script.py

# 测试
python3 -c "from browser_use import Agent; print('✅ OK')"

记住: browser-use 让浏览器操作更智能,省去调试选择器的痛苦!🚀

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/browser-use of aAAaqwq/AGI-Super-Team.

Open the folder on GitHubat commit 331ecd3

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Use this skillaAAaqwq/AGI-Super-Team1051 repos~1.6kAutomated safety check: NotesMIT
AI Search Hubminsight-ai-info/AI-Search-Hub1.3k—~1.3kAutomated safety check: PassNone
Ego Browserkwakseongjae/oh-my-design5311 repos~4.9kAutomated safety check: PassMIT
Browser Controlanomalyco/browser-control438—~8.2kAutomated safety check: PassMIT
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone
Vrboborski/travel-hacking-toolkit689—~1.7kAutomated safety check: PassMIT

Similar skills

  • AI Search Hub

    minsight-ai-info/AI-Search-Hub

    Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.

    1.3k GitHub stars~1.3k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Ego Browser

    kwakseongjae/oh-my-design

    When you need a browser, read this Skill by default. An agent skill from kwakseongjae/oh-my-design.

    531 GitHub starsUsed in 1 repo~4.9k tokens
    Productivity & AutomationAuto-check passed
  • Browser Control

    anomalyco/browser-control

    Drive the user's existing Chromium-family browser with deterministic Playwright.

    438 GitHub stars~8.2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Vrbo

    borski/travel-hacking-toolkit

    Search VRBO (Vrbo / Expedia Group) vacation rentals including entire homes, condos, and cabins via Patchright browser automation.

    689 GitHub stars~1.7k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Sap Browser Automation

    secondsky/sap-skills

    A skill your agent uses when an agent must inspect or operate an authenticated SAP web UI through an in-app Browser, Microsoft Edge CDP, or an existing Playwright client, especially when SAP SSO…

    460 GitHub stars~3.3k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 161 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Performing Security Code Review

    aAAaqwq/AGI-Super-Team

    This skill provides automated assistance for security agent tasks Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    105 GitHub starsUsed in 1 repo~920 tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Sysadmin Toolbox

    aAAaqwq/AGI-Super-Team

    Tool discovery and shell one-liner reference for sysadmin, DevOps, and security tasks.

    105 GitHub starsUsed in 2 repos~775 tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Browser Use

What does Browser Use do?

AI驱动的智能浏览器自动化工具。使用LLM理解页面并自动执行任务,比传统Playwright更智能、更省token。适用于复杂交互、动态页面、需要智能决策的浏览器操作。Chrome浏览器优先。. Browser Use is an agent skill from aAAaqwq/AGI-Super-Team.

When should I use Browser Use?

Browser Use fits situations like: tasks that involve Browser automation; tasks that involve Browser testing.

How do I install Browser Use in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill browser-use -a claude-code`. Or copy the skill folder (skills/browser-use in aAAaqwq/AGI-Super-Team) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.

How do I install Browser Use in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill browser-use -a codex`. Or copy the skill folder (skills/browser-use in aAAaqwq/AGI-Super-Team) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.

Can I use Browser Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.

What does Browser Use need to run?

Going by SKILL.md and its folder, Browser Use needs the command-line tools its instructions call (python3 and pip) and credentials named XSC_API_KEY. Our summary lists: Python 3; A credential in XSC_API_KEY. Its frontmatter pre-approves these tools: Bash, Exec, Read, Write.

Does Browser Use access the network?

SKILL.md names 4 domains. In commands or code: polymarket.com and ai.9w7.cn; the agent is likely to contact these when it follows the instructions. As links in the text: docs.browser-use.com and github.com. This is read from the text; nothing was executed.

Is Browser Use safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Browser Use use?

Browser Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Use use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Use?

Skills that share tags, products or a category with Browser Use: AI Search Hub (minsight-ai-info/AI-Search-Hub, 1.3k stars), Ego Browser (kwakseongjae/oh-my-design, 531 stars), Browser Control (anomalyco/browser-control, 438 stars) and Playwright Bowser (disler/bowser, 265 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Use?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 161 skills in this directory. The repository was last updated on September 27, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.