Agent skill

QwenPaw Browser SDK

by agentscope-ai in agentscope-ai/QwenPaw

Drives a live browser with async Python through QwenPaw's built-in Browser SDK, in a perceive, act, verify loop with handoff for logins and captchas.

Apache-2.0Auto-check passedProductivity & Automation

Install QwenPaw Browser SDK

skills CLI
$ npx skills add agentscope-ai/QwenPaw --skill browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/QwenPaw browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/QwenPaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/qwenpaw/agents/skills/browser-en .claude/skills/browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser
GitHub stars
35k
Token cost
~2.6k tokens
SKILL.md length
1,260 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
Apache-2.0

At a glance

Drives a live browser with async Python through QwenPaw's built-in Browser SDK, in a perceive, act, verify loop with handoff for logins and captchas.

  • Works in 3 steps: semantic page.get_by_role/label/text… → css page.locator(css) role… → coordinates use locator.bounding_box()…
  • Automating a task inside QwenPaw with a live browser page
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Reading and verifying page content step by step before reporting results

What it does

The skill is the full reference for QwenPaw's own Browser SDK, which is not Playwright and has a closed API surface: anything not listed does not exist. The agent writes async Python, connects once with Browser.connect, opens a page, takes a snapshot to read the page text, acts through role-based locators such as a textbox or button, and snapshots again to verify. Large pages are read selectively rather than dumped, and a snapshot can take a query for a focused count.

The discipline rules are strict: state only facts observed this turn, report a stuck step plus verified partial results instead of filling gaps, and never present data from another channel such as web search as a browser result. Logins, captchas, two-factor prompts and other human-only steps call browser.handoff and stop. Sessions are stateful, so variables persist, and with the Chrome backend a session shares cookies and logins with the user's real browser profile, so identity isolation cannot be assumed.

When your agent uses it

  • Automating a task inside QwenPaw with a live browser page
  • Reading and verifying page content step by step before reporting results
  • Handing a login or captcha back to the user mid-task

Example prompts

  • “Open the product site, search for laptop and tell me what the results page shows.”
  • “Check the order status page and report only what you can see on screen.”
  • “That site needs a login, so pause and hand control back to me.”

Requirements

  • QwenPaw with its built-in Browser SDK
  • Async Python

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. semantic page.get_by_role/label/text first choice
  2. css page.locator(css) role missing/unstable
  3. coordinates use locator.bounding_box() first for an exact, low-cost viewport rectangle; use a screenshot to explore only when the element…

What it can do on your machine

Read from SKILL.md and the folder at commit 80e412d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QwenPaw Browser SDK loads about 2.6k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,260 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/QwenPaw at commit 80e412d, republished under its Apache-2.0 licence (© agentscope-ai). 1,260 words, ~2,612 tokens.

Download SKILL.mdSave it as .claude/skills/browser/SKILL.md (or your agent's skills folder).
name
browser
description
Drive a live browser with async Python against QwenPaw's builtin Browser SDK. The full reference is below; re-load this browser skill after context compaction.
metadata.builtin_skill_version
0.3

Browser

Work with discipline: perceive the current page, act through the documented surface, then re-perceive before claiming success. State only facts observed in this turn. When stuck, a complete delivery is the step you are stuck on plus the partial results you have verified — never fill in content you did not observe just to produce a full answer.

Respect the human boundary: for login, captcha, 2FA, or any human-only step, call await browser.handoff(...) and stop. Never automate those flows. When the browser cannot finish, do not substitute data from another channel (such as web_search) and still call it a browser result — state each fact's real source.

This is QwenPaw's builtin Browser SDK, not Playwright. Its surface is closed: anything not listed does not exist. The complete reference is below; re-load this browser skill with the Skill tool if it is no longer in context.

<!-- BEGIN GENERATED: browser-manual -->

QwenPaw Browser SDK — complete reference. This is QwenPaw's OWN internal SDK and this is the ENTIRE API; these are all the entrypoints. The SDK is already in scope as Browser — call the methods below directly. Write async Python. Work in a loop: perceive → act → verify.

Copy this shape:

browser = await Browser.connect() # connect once; reused all session page = await browser.open("https://example.com") # open a page obs = await page.snapshot() # PERCEIVE — page text is obs.text if len(obs.text) < 6000: print(obs.text) else: # Large page: read selectively instead of dumping everything. lines = [line for line in obs.text.splitlines() if "keyword" in line] print(f"{len(obs.text)} chars total; {len(lines)} matching lines:") print("\n".join(lines[:80]))

For a focused count, use: await page.snapshot(query="keyword")

await page.get_by_role("textbox", name="Search").fill("laptop") # ACT await page.get_by_role("button", name="Search").click() # ACT obs = await page.snapshot() # VERIFY — re-perceive to confirm print("Verified; inspect obs.text with the selective pattern above.")

Session state: this is a stateful session — variables you assign (browser, page) persist across calls, so connect once and reuse them. If a call reports the session was reset, re-run await Browser.connect().

Chrome backend caveat: with backend=chrome you operate inside the user's real browser. A session is a tab-ownership group — tabs are isolated per session, but identity (cookies, logins, storage) is shared with the user's profile and with every other session. Do not rely on session-level identity isolation on this backend.

browser (orchestration): await Browser.connect(*, identity: "auto"|"user"|"avatar"|"guest" = "auto") -> browser Connect as an identity: user, avatar, guest, or auto.

``auto`` picks ``user`` when Chrome is connected, otherwise ``guest``.
An unavailable explicit identity raises instead of substituting.

await browser.open(url: str | None = None) -> page Open a page at url and return it.

Reuses this session's active page when one exists; otherwise a
new page is created. Pages are released when the response cycle ends;
start each cycle by calling ``open(url)`` again.

await browser.pages() -> list of page ref (.id, .url, .title, .active) List open pages with URL, title, and active-state details. await browser.switch_page(page: page ref (.id, .url, .title, .active)) -> none Make the given page ref active for later operations. await browser.close_page(page: page ref (.id, .url, .title, .active)) -> none Close the given page ref in this session. await browser.session_status() -> session status (.owner, .variant, .context, .connected) Report the owner, variant, context, and connected state. await browser.handoff(reason: str, instructions: str = "") -> a result dict Hand a step back to a human (captcha, login, 2FA).

Pass a short reason and instructions; the run stops on this signal —
never automate these flows. The active cycle-scoped page is retained
for one extra response cycle after the handoff.

await browser.present(url: str | None = None) -> page Open a page retained for the chat lifetime. await browser.close() -> none Close this session's browser and release its context.

page (operation): await page.goto(url: str) -> a result dict Navigate this page to url and return raw navigation facts. await page.go_back() -> a result dict Navigate back to the previous page in history. await page.go_forward() -> a result dict Navigate forward to the next page in history. await page.reload() -> a result dict Reload the current page. await page.keep() -> none Retain this page across response cycles for the current chat. await page.wait_for_load_state(state: str = "load", *, timeout: float | None = None) -> none Wait until the page reaches the requested load state.

``networkidle`` semantics depend on the backend: the Playwright
backend waits for true network quiescence, while CDP-based
backends (cdp, chrome) degrade to ``document.readyState ==
"complete"`` plus a fixed 500 ms quiet delay and do NOT track
in-flight requests — content loaded by late XHR may still be
missing when this returns.

await page.wait_for_timeout(timeout: float) -> none Sleep unconditionally for timeout milliseconds (capped at 30 000).

Prefer :py:meth:`locator.wait_for(state, timeout)
<LocatorView.wait_for>` when waiting for a specific DOM condition
— it returns as soon as the condition is met and is both faster
and more reliable than an unconditional sleep.

await page.screenshot() -> a result dict Capture this page to a PNG file in the active workspace. page.get_by_role(role: str, *, name: str | None = None) -> locator Locate elements by accessible role and optional name. page.get_by_text(text: str) -> locator Locate elements by their visible text. page.get_by_label(text: str) -> locator Locate a form control by its associated label text. page.get_by_placeholder(text: str) -> locator Locate an input by its placeholder text. page.locator(selector: str) -> locator Locate elements by a CSS selector when no semantic locator fits. page.frame_locator(selector: str) -> locator Scope subsequent locators to the iframe matching selector. await page.snapshot(query: str | None = None) -> observation (read .text; .match_count when you pass a query) Perceive the page and return readable content in .text.

Show full SKILL.md (377 more words)Show less
Pass ``query`` to also report ``.match_count``.

await page.current_surface() -> surface facts (.url, .title, .load_state) Return this page's current URL, title, and load facts. page.mouse -> coordinate/keyboard input surface (see methods below) Viewport-coordinate input surface.

click(x, y) -> a result dict; verify the effect with snapshot().

page.keyboard -> coordinate/keyboard input surface (see methods below) Keyboard input surface.

press(key) -> a result dict; verify the effect with snapshot().

page.get_by_* / page.locator(...) return a locator that mirrors a SUBSET of Playwright's Python locator API — the Playwright-shaped part of this SDK: compose/scope (chainable): get_by_role/get_by_text/get_by_label/ get_by_placeholder, locator(sel), filter(...), nth(i), first, last (properties) iframe scope: page.frame_locator(sel).locator(...) (one frame; no nested frames) read (await): count()->int, inner_text()->str, text_content()->str|None, all_text_contents()->list, get_attribute(name)->str|None, input_value()->str, is_visible()->bool, is_enabled()->bool act (await; returns a short evidence line — read .evidence): click(), fill(v), type(t), press(key), check(), uncheck(), set_checked(b), select_option(*v), hover(), dblclick(), scroll(), focus(), blur(), clear(), wait_for(state), screenshot(), bounding_box()->dict|None (viewport ["x"] ["y"] ["width"] ["height"]; None when the element is not visible) strict-mode uniqueness is enforced — act only when the locator resolves to exactly one element (use count() to check). element_handle / raw CDP unavailable.

Backend differences (chrome/cdp vs playwright): on chrome/cdp the accessible name is a heuristic (aria-labelledby > aria-label > alt > title > text content) - container elements may match get_by_role(name=) more broadly than under playwright, so strict-mode errors are more likely there; narrow with filter(has_text=) or a more specific role. is_enabled() reflects only the disabled property, not aria-disabled. press() supports a fixed key set: printable characters, Enter, Tab, Escape, Backspace, Delete, Arrow keys, Home/End/PageUp/PageDown, and Control/Shift/Alt/Meta combos - anything else fails with guidance. type() sets the value directly and fires an input event; editors that need real per-key events may not react - prefer fill() where possible.

Reading results (read these fields; the type names don't matter): snapshot() -> .text (page text), .match_count (when you pass query) current_surface() -> .url, .title, .load_state page refs -> .id, .url, .title, .active screenshot() -> result dict; read ["path"] bounding_box() -> viewport ["x"] ["y"] ["width"] ["height"]; None when the element is not visible mouse.click()/keyboard.press() -> result dict; fields depend on the backend, so verify with snapshot() actions -> .evidence (a short line saying what happened) locator reads return plain str/int/bool/list directly.

If a locator fails, step DOWN one rung (don't jump):

  1. semantic page.get_by_role/label/text first choice
  2. css page.locator(css) role missing/unstable
  3. coordinates use locator.bounding_box() first for an exact, low-cost viewport rectangle; use a screenshot to explore only when the element is absent from snapshot() For captcha/login/2FA or any human-only step: await browser.handoff(reason, instructions) and stop — never automate them.
<!-- END GENERATED: browser-manual -->

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/qwenpaw/agents/skills/browser-en of agentscope-ai/QwenPaw.

Open the folder on GitHubat commit 80e412d

Compare with similar skills

QwenPaw Browser SDK next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QwenPaw Browser SDK compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QwenPaw Browser SDK this skillagentscope-ai/QwenPaw35k—~2.6kAutomated safety check: PassApache-2.0
Browser Automation Edge Casesaden-hive/hive11k—~1.8kAutomated safety check: PassMIT
Skyvern Browser AutomationSkyvern-AI/skyvern23k—~1.9kAutomated safety check: PassAGPL-3.0
AI Search Hubminsight-ai-info/AI-Search-Hub1.3k—~1.3kAutomated safety check: PassNone
Browser Use Terminalbrowser-use/terminal651—~2.1kAutomated safety check: PassMIT
Browser Tools981377660LMT/algorithm-study2781 repos~1.3kAutomated safety check: PassNone

Similar skills

  • Step-by-step procedure for debugging browser automation failures on complex sites such as LinkedIn, Twitter/X, single-page apps and Shadow DOM pages.

    11k GitHub stars~1.8k tokensUpdated 23 days ago
    Productivity & AutomationAuto-check passed
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • AI Search Hub

    minsight-ai-info/AI-Search-Hub

    Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.

    1.3k GitHub stars~1.3k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Use Terminal

    browser-use/terminal

    Direct browser control via the Browser Use Terminal CLI. An agent skill from browser-use/terminal.

    651 GitHub stars~2.1k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Tools

    981377660LMT/algorithm-study

    Interactive browser automation via Chrome DevTools Protocol.

    278 GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed
  • Evoui Browser

    Salmonbird/evoui-browser

    Use Evoui Browser as the default entry point for common, self-terminating web tasks that need a real browser through agent-browser, including navigation, page reading, clicks, forms, login flows…

    100 GitHub stars~4.5k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed

More from agentscope-ai/QwenPaw

All 20 skills in this repo
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    35k GitHub starsUsed in 6 repos~4.2k tokens
    Auto-check passed
  • Make Skill

    agentscope-ai/QwenPaw

    Turns reusable decisions, templates or workflows from the current conversation into a new workspace skill through plan, approval, draft, validation and publication.

    35k GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • QwenPaw Make Skill

    agentscope-ai/QwenPaw

    Creates a focused workspace skill from the current conversation in QwenPaw, moving through a planning, approval, drafting, validation and publishing script pipeline.

    35k GitHub stars~1.2k tokensUpdated 7 days ago
    Auto-check passed
  • QwenPaw Scheduled Tasks

    agentscope-ai/QwenPaw

    Creates and manages scheduled or recurring jobs with the qwenpaw cron commands, always tied to an explicit agent ID and a confirmed target channel.

    35k GitHub stars~2.3k tokensUpdated 7 days ago
    Auto-check passed
  • DOCX Creation and Editing

    agentscope-ai/QwenPaw

    Creates, reads and edits Word .docx files, including tracked changes and comments, using docx-js for new files and XML editing for existing ones.

    35k GitHub stars~3.3k tokensUpdated 7 days ago
    Auto-check passed
  • Mailbox Operations Hub

    agentscope-ai/QwenPaw

    Connects, registers, and operates a personal mailbox, reading, searching, sending, and organizing, through a managed mail server for nine domains.

    35k GitHub stars~2.6k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about QwenPaw Browser SDK

What does QwenPaw Browser SDK do?

Drives a live browser with async Python through QwenPaw's built-in Browser SDK, in a perceive, act, verify loop with handoff for logins and captchas. The skill is the full reference for QwenPaw's own Browser SDK, which is not Playwright and has a closed API surface: anything not listed does not exist.connect, opens a page, takes a snapshot to read the page text, acts through role-based locators such as a textbox or button, and snapshots again to verify.

When should I use QwenPaw Browser SDK?

QwenPaw Browser SDK fits situations like: automating a task inside QwenPaw with a live browser page; reading and verifying page content step by step before reporting results; handing a login or captcha back to the user mid-task.

How do I install QwenPaw Browser SDK in Claude Code?

Run `npx skills add agentscope-ai/QwenPaw --skill browser -a claude-code`. Or copy the skill folder (src/qwenpaw/agents/skills/browser-en in agentscope-ai/QwenPaw) into .claude/skills/browser in your project. Claude Code loads it when a task matches its description.

How do I install QwenPaw Browser SDK in Codex?

Run `npx skills add agentscope-ai/QwenPaw --skill browser -a codex`. Or copy the skill folder (src/qwenpaw/agents/skills/browser-en in agentscope-ai/QwenPaw) into .agents/skills/browser in your project. Codex loads it when a task matches its description.

Can I use QwenPaw Browser SDK in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/QwenPaw --skill browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser, .gemini/skills/browser, .github/skills/browser and .opencode/skills/browser in your project.

What does QwenPaw Browser SDK need to run?

SKILL.md names no scripts, command-line tools or credentials: QwenPaw Browser SDK is instructions for the agent only. Our summary lists: QwenPaw with its built-in Browser SDK; Async Python.

Does QwenPaw Browser SDK access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is QwenPaw Browser SDK safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QwenPaw Browser SDK use?

QwenPaw Browser SDK is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QwenPaw Browser SDK use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QwenPaw Browser SDK?

Skills that share tags, products or a category with QwenPaw Browser SDK: Browser Automation Edge Cases (aden-hive/hive, 11k stars), Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars), AI Search Hub (minsight-ai-info/AI-Search-Hub, 1.3k stars) and Browser Use Terminal (browser-use/terminal, 651 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QwenPaw Browser SDK?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/QwenPaw, which has 35,464 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on September 30, 2026.

Source: agentscope-ai/QwenPaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.