Agent skill

Agent Browser

by paperclipai in paperclipai/paperclip

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

MITAuto-check passedProductivity & Automation

Install Agent Browser

skills CLI
$ npx skills add paperclipai/paperclip --skill agent-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install paperclipai/paperclip agent-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/paperclipai/paperclip.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills-catalog/catalog/optional/browser/agent-browser .claude/skills/agent-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-browser
GitHub stars
99k
Token cost
~1.3k tokens
SKILL.md length
709 words
Files
1
Skills in repo
60
Repo updated
First seen
Licence
MIT

At a glance

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

  • Works in 7 steps: Launch with a real-looking user agent… → Set a sane viewport (e.g., 1366×768… → Navigate and wait for the right signal.… → …
  • Tasks that involve Browser automation
  • SKILL.md covers When to use, When not to use, Before launching the browser and Driving the browser, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Browser is an agent skill from paperclipai/paperclip. Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation. The repository describes itself as: The open-source app everyone uses to manage agents at work. The licence is MIT.

When your agent uses it

  • Tasks that involve Browser automation

Example prompts

  • “/agent-browser”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Launch with a real-looking user agent when the target is the public internet; an unrealistic UA flags automation traffic.
  2. Set a sane viewport (e.g., 1366×768 desktop, 390×844 iPhone-ish).
  3. Navigate and wait for the right signal. Prefer waiting for a specific selector or network-idle over arbitrary sleeps.
  4. Capture evidence immediately after the wait condition succeeds, before any interaction perturbs the state.
  5. Interact deliberately. One click at a time, with a wait between actions; re-screenshot after each meaningful state change.
  6. Read the console and network panels for unexpected errors, 4xx/5xx responses, or slow requests.
  7. Close the browser cleanly when done. Long-running browser sessions leak memory and hold ports.

What it can do on your machine

Read from SKILL.md and the folder at commit b9750b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Browser loads about 1.3k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 709 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from paperclipai/paperclip at commit b9750b1, republished under its MIT licence (© paperclipai). 709 words, ~1,281 tokens.

Download SKILL.mdSave it as .claude/skills/agent-browser/SKILL.md (or your agent's skills folder).
name
agent-browser
description
Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.
key
paperclipai/optional/browser/agent-browser
recommendedForRoles
qa, engineer, researcher
tags
browser, puppeteer, playwright, verification

Agent Browser

Use a controlled browser to verify behavior, capture evidence, or extract information from web pages that a static fetch cannot reach (SPAs, login-gated pages, dynamic content). This skill is about supervised verification, not unattended scraping.

When to use

  • You need a screenshot of a deployed page or a local dev server to confirm a UI change.
  • You need to read JavaScript-rendered content that curl/wget will not see.
  • A user reports a UI bug and you need to reproduce it interactively to capture console errors, network requests, or layout state.
  • You need to walk through a short flow (load page, click, observe) to verify acceptance criteria.

When not to use

  • The page is reachable as static HTML. Use curl/HTTP fetch — it is cheaper, faster, and more reliable.
  • The task is unattended large-scale scraping. That belongs to a dedicated scraper with rate limits, robots.txt handling, and a real user agent policy — not this skill.
  • The site is behind authentication you do not own credentials for, or whose terms of service prohibit automation.
  • The site involves sensitive accounts (banking, healthcare, government) where automation risks lockout or compliance issues.

Before launching the browser

  • Confirm the URL and what state should be true after navigation.
  • Decide what evidence is needed: full-page screenshot, viewport screenshot, console log, network trace, HTML snapshot, extracted text.
  • Decide the viewport size that matters for the task (mobile vs desktop). Default to a desktop size unless the task is mobile-specific.
  • For local dev servers, confirm the server is running and the port is what you expect.

Driving the browser

A typical verification session:

  1. Launch with a real-looking user agent when the target is the public internet; an unrealistic UA flags automation traffic.
  2. Set a sane viewport (e.g., 1366×768 desktop, 390×844 iPhone-ish).
  3. Navigate and wait for the right signal. Prefer waiting for a specific selector or network-idle over arbitrary sleeps.
  4. Capture evidence immediately after the wait condition succeeds, before any interaction perturbs the state.
  5. Interact deliberately. One click at a time, with a wait between actions; re-screenshot after each meaningful state change.
  6. Read the console and network panels for unexpected errors, 4xx/5xx responses, or slow requests.
  7. Close the browser cleanly when done. Long-running browser sessions leak memory and hold ports.

What evidence to record

For a verification task, deliver:

  • A full-page or viewport screenshot of each meaningful state.
  • The console log, filtered to warnings/errors.
  • Any non-2xx network response with the URL, status, and a short response body excerpt.
  • A short narration: "Navigated to X, observed Y, clicked Z, observed W."

For a UI bug repro, also record:

  • The exact reproduction steps the user can follow.
  • Viewport size and (where relevant) device pixel ratio.
  • Whether the bug reproduces on first load vs after interaction.
Show full SKILL.md (249 more words)Show less

Login-gated pages

  • Prefer programmatic auth (API token, magic link) over UI login.
  • If UI login is the only path, the user must provide credentials explicitly for this run. Never reuse credentials outside the session.
  • Do not store credentials in the session log, screenshot, or returned output.

Performance and politeness

  • Throttle to one navigation per few seconds when touching shared infra.
  • Respect robots.txt for public sites you are inspecting at any volume.
  • Cancel navigations if a page exceeds a reasonable timeout (e.g., 30s); the page is broken or rate-limiting you.
  • Do not retry forever on failure. Retry once with a longer timeout, then escalate.

Common failure modes

  • Selector not found. Page changed, or you are waiting before render. Take a screenshot to see actual state; adjust the selector.
  • Click does nothing. The element is offscreen, covered by a modal, or in a shadow DOM. Scroll into view or pierce the shadow root.
  • Headless detection. Some sites detect headless Chrome and serve a different page. Use a non-headless mode or a fingerprint-realistic configuration only when authorized.
  • Cross-origin iframe blocking. Iframes you do not own cannot be inspected; the page must offer the data outside the iframe or the task is infeasible.

Anti-patterns

  • Long unsupervised browser sessions that drift from the original task.
  • Scraping behind authentication you do not own.
  • Captioning a screenshot with "looks good" without saying what state was loaded and what selectors confirmed it.
  • Treating a passing screenshot as proof of correctness across viewports you did not actually test.

© paperclipai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills-catalog/catalog/optional/browser/agent-browser of paperclipai/paperclip.

Open the folder on GitHubat commit b9750b1

Compare with similar skills

Agent Browser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Browser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Browser this skillpaperclipai/paperclip99k—~1.3kAutomated safety check: PassMIT
Agent Browserquran/quran.com-frontend-next1.9k40 repos~3.3kAutomated safety check: PassNone
Dev-Browser CLI AutomationSawyerHood/dev-browser6.7k1 repos~455Automated safety check: PassMIT
Browser Automationopenclaw/openclaw392k—~2.9kAutomated safety check: PassMIT
Camoufox CLIBin-Huang/camoufox-cli3501 repos~4.5kAutomated safety check: PassMIT
BrowserVibiumDev/vibium2.9k—~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 40 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Dev-Browser CLI Automation

    SawyerHood/dev-browser

    Browser automation with persistent named pages via the dev-browser CLI. Use when users ask to navigate websites, fill forms, take screenshots, extract web…

    6.7k GitHub starsUsed in 1 repo~455 tokens
    Productivity & AutomationAuto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Camoufox CLI

    Bin-Huang/camoufox-cli

    Anti-detect browser automation CLI & Skills for AI agents. An agent skill from Bin-Huang/camoufox-cli.

    350 GitHub starsUsed in 1 repo~4.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser

    VibiumDev/vibium

    Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.

    2.9k GitHub stars~4.8k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts.

    50k GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from paperclipai/paperclip

All 60 skills in this repo
  • Garden Inbox

    paperclipai/paperclip

    Scan a Paperclip user's Mine inbox, classify reversible archive candidates, request checkbox confirmation, and archive only accepted selections.

    99k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    Interact with the Paperclip control plane API for task coordination and governance.

    99k GitHub stars~9.6k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    A skill your agent uses for Paperclip-managed tasks and heartbeats: reading task context, delivering task documents or files, updating completion or blockers, coordinating or delegating work, and…

    99k GitHub stars~17k tokensUpdated today
    Auto-check passed
  • Design Guide

    paperclipai/paperclip

    Paperclip UI design system guide for building consistent, reusable frontend components.

    99k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Paperclip Page

    paperclipai/paperclip

    Publish static HTML pages and asset folders to the Paperclip S3/CloudFront page host.

    99k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paperclip Create Agent

    paperclipai/paperclip

    Create new agents in Paperclip with governance-aware hiring.

    99k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Questions about Agent Browser

What does Agent Browser do?

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation. Agent Browser is an agent skill from paperclipai/paperclip. Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

When should I use Agent Browser?

Agent Browser fits situations like: tasks that involve Browser automation.

How do I install Agent Browser in Claude Code?

Run `npx skills add paperclipai/paperclip --skill agent-browser -a claude-code`. Or copy the skill folder (packages/skills-catalog/catalog/optional/browser/agent-browser in paperclipai/paperclip) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.

How do I install Agent Browser in Codex?

Run `npx skills add paperclipai/paperclip --skill agent-browser -a codex`. Or copy the skill folder (packages/skills-catalog/catalog/optional/browser/agent-browser in paperclipai/paperclip) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.

Can I use Agent Browser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add paperclipai/paperclip --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.

What does Agent Browser need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Browser is instructions for the agent only.

Does Agent Browser access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Browser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Browser use?

Agent Browser is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Browser use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Browser?

Skills that share tags, products or a category with Agent Browser: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Dev-Browser CLI Automation (SawyerHood/dev-browser, 6.7k stars), Browser Automation (openclaw/openclaw, 392k stars) and Camoufox CLI (Bin-Huang/camoufox-cli, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Browser?

paperclipai (a GitHub organization) maintains it in paperclipai/paperclip, which has 98,967 GitHub stars. The repository holds 60 skills in this directory. The repository was last updated on October 9, 2026.

Source: paperclipai/paperclip on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.