Agent skill

Vellum Browser Use

by vellum-ai in vellum-ai/vellum-assistant

Browse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant.

MITAuto-check passedProductivity & Automation

Install Vellum Browser Use

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vellum-browser-use .claude/skills/vellum-browser-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vellum-browser-use
GitHub stars
1.4k
Token cost
~3k tokens
SKILL.md length
1,325 words
Files
1
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Browse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant.

  • Works in 2 steps: Install the Vellum Assistant Chrome… → Open the extension in Chrome and pair it…
  • Tasks that involve Browser automation
  • SKILL.md covers Getting Started — Check…, Browser Modes, When a Page Cannot Be Reached and Targeting a Specific Client, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vellum Browser Use is an agent skill from vellum-ai/vellum-assistant. Browse the web using assistant browser CLI commands

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Vellum personal assistants

It sits in Productivity & Automation, covering Browser automation. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

When your agent uses it

  • Tasks that involve Browser automation

Example prompts

  • “/vellum-browser-use”

Requirements

  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Install the Vellum Assistant Chrome Extension from the Chrome Web Store: https://chromewebstore.google.com/detail/vellum-assistant-browser/…
  2. Open the extension in Chrome and pair it with the assistant.

What it can do on your machine

Read from SKILL.md and the folder at commit 33cc983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • chromewebstore.google.com
    • vellum.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Vellum Browser Use loads about 3k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 1,325 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit 33cc983, republished under its MIT licence (© vellum-ai). 1,325 words, ~2,993 tokens.

Download SKILL.mdSave it as .claude/skills/vellum-browser-use/SKILL.md (or your agent's skills folder).
name
vellum-browser-use
description
Browse the web using `assistant browser` CLI commands
compatibility
Designed for Vellum personal assistants
metadata.emoji
🌐

Use this skill to browse the web. All browser operations are executed through the assistant browser CLI, invoked via bash or host_bash. Each operation is a subcommand:

CommandDescription
assistant browser navigateNavigate to a URL
assistant browser snapshotList interactive elements on the current page
assistant browser screenshotTake a visual screenshot
assistant browser clickClick an element
assistant browser typeType text into an input
assistant browser press-keyPress a keyboard key
assistant browser scrollScroll the page or a specific element
assistant browser select-optionSelect an option from a native <select> element
assistant browser hoverHover over an element to reveal menus/tooltips
assistant browser wait-forWait for a condition
assistant browser extractExtract page text content
assistant browser wait-for-downloadWait for a file download to complete
assistant browser fill-credentialFill a stored credential into a form field
assistant browser attachAttach the Chrome debugger to the active tab
assistant browser detachDetach the Chrome debugger from the active tab
assistant browser closeClose the browser page
assistant browser statusDiagnose browser backend readiness and setup steps

Getting Started — Check Browser Readiness

Before using any browser commands, run assistant browser --json status first to check which browser backends are available. The status command returns JSON with readiness information for each backend mode:

bash
assistant browser --json status

The response includes:

  • recommendedMode — the best available backend (use this)
  • modes[] — per-mode status with available, summary, and userActions (remediation steps)

For --virtual-desktop, Chrome and desktop components are included in the assistant image. Browser and computer use start the existing desktop without installing dependencies. If components are missing, report the image problem; do not install packages or launch a separate Chrome process.

Browser Modes

Use --browser-mode <mode> on the assistant browser parent command to pin the browser backend:

ValueBackendDescription
autoAutomaticDefault. Picks the best available backend based on context.
extensionChrome extensionRoutes through the user's Chrome browser via the extension debugger.
cdp-inspectCDP inspectConnects to an already-running Chrome instance via DevTools Protocol.
localPlaywrightDrives a dedicated Playwright-managed Chromium instance.
bash
assistant browser --browser-mode extension navigate --url http://www.example.com
Prefer the Chrome Extension

The Chrome extension (extension mode) is the preferred browser backend. It is:

  • More secure than Chrome's native remote debugging
  • Uses the user's real browser profile (cookies, sessions, saved logins)
  • Best experience for interacting with the user's actual browsing context

If the status check shows the extension is not available, encourage the user to install and pair it:

  1. Install the Vellum Assistant Chrome Extension from the Chrome Web Store: https://chromewebstore.google.com/detail/vellum-assistant-browser/hphbdmpffeigpcdjkckleobjmhhokpne
  2. Open the extension in Chrome and pair it with the assistant.

The status response's userActions array for the extension mode provides these same steps when the extension is not connected.

When a Page Cannot Be Reached

If navigate, curl, or any fetch times out, hits an auth wall, or cannot reach a host (VPN, company login, internal dashboard):

  1. Tell the user a connected desktop app or Chrome extension can open the page in a browser where they are already logged in.
  2. Give the install links:
  3. Offer those first. Only ask for a screenshot or pasted page content if they cannot install either.
  4. On iOS or Android there is no in-app browser and no extension to install on the phone. Offer the desktop app or Chrome extension on a computer. Do not describe a browser panel.
Fallback Modes

If the user declines to install the extension:

  • cdp-inspect — Connects to an already-running Chrome instance via DevTools Protocol (Chrome 146+). Requires enabling remote debugging in Chrome settings.
  • local — Drives a dedicated Playwright-managed Chromium instance. Last resort — does not use the user's browser profile.

Only fall back to these if the user explicitly indicates they do not want to install the extension. Prefer cdp-inspect over local.

Targeting a Specific Client

When multiple clients support host_browser (e.g. two Chrome profiles, a macOS client and a Chrome extension), use --target-client-id <id> on the assistant browser parent command to pin all operations in the invocation to one specific client:

bash
assistant browser --target-client-id <client-id> navigate --url https://example.com

Obtain client IDs from:

bash
assistant clients list --capability host_browser

Omit --target-client-id when only one client is connected — the default interface-preference order (chrome-extension first, then macos) picks the best available client automatically.

Tab Handling (Chrome extension)

On the Chrome extension backend, navigate opens a dedicated tab the first time it runs in a conversation and pins subsequent operations to it, so browsing never disturbs the tab the user is on (often the tab they're chatting with the assistant from). Later navigates reuse that pinned tab.

  • --new-tab — force a brand-new tab even when one is already pinned.
  • --use-active-tab — navigate the user's currently-active tab instead of a dedicated one.

Both flags are ignored on the local and cdp-inspect backends, which manage their own browser context.

Show full SKILL.md (543 more words)Show less

Session Management

Use --session <id> on the assistant browser parent command to group sequential operations so they share browser state (same page, cookies, etc.). Different session IDs create independent browser contexts.

bash
assistant browser --session myflow navigate --url https://example.com
assistant browser --session myflow snapshot
assistant browser --session myflow click --element-id e3

Omitting --session uses the default session.

Machine-Readable Output

Use --json on the assistant browser parent command to get structured JSON output suitable for parsing in scripts:

bash
assistant browser --json navigate --url https://example.com
# {"ok":true,"content":"Page title: Example Domain"}

assistant browser --json snapshot
# {"ok":true,"content":"...element list..."}

assistant browser --json screenshot
# {"ok":true,"content":"...","screenshots":[{"mediaType":"image/jpeg","data":"<base64>"}]}

Error responses use {"ok":false,"error":"..."}.

Screenshots

To save a screenshot to disk, use --output <path>:

bash
assistant browser screenshot --output page.jpg
assistant browser screenshot --full-page --output full.jpg

To receive base64 screenshot data in JSON output:

bash
assistant browser --json screenshot

The response includes a screenshots array with mediaType and data (base64) fields.

Typical Workflow

  1. assistant browser --json status to check backend readiness — if the extension is not available, help the user install it
  2. (Optional) assistant browser attach to establish the session
  3. assistant browser navigate --url <url> to load a page
  4. assistant browser snapshot to discover interactive elements
  5. Use click, type, press-key, scroll, select-option, or hover to interact
  6. assistant browser extract or assistant browser screenshot --output <path> to capture results
  7. Always assistant browser detach when you are done — this releases the debugger so the user can browse freely

Human verification

Treat every CAPTCHA and bot-detection challenge as a request for human help, including drag-to-verify sliders, press-and-hold checks, verification checkboxes and image puzzles. Stop before interacting with the challenge, even if its controls look easy to automate. Do not try it yourself, retry it, or script a solution.

In the virtual desktop, request the desktop-help card immediately, using one short sentence for the needed action. Wait for Done or Skip. After Done, take a fresh snapshot; if verification remains, ask for help again. On other browser backends, ask the user to complete verification in their browser and wait for confirmation.

For ordinary logins, use saved credentials or securely prompt for missing credentials, then fill the form yourself. A CAPTCHA on a login page still requires human help.

Interaction Strategies

Date pickers / calendars: Click the date input to open the picker, re-snapshot to see calendar controls, click month navigation arrows to reach the target month, then click the target date. For <input type="date">, use type with YYYY-MM-DD format.

Native <select> elements: Use select-option with --value, --label, or --index. Do not try to click individual <option> elements.

ARIA / custom dropdowns: Click to open, take a new snapshot, then click the desired option by --element-id.

Autocomplete inputs: Type the search text, wait 500-1000ms (wait-for --duration), re-snapshot for suggestions, then click the suggestion or use press-key --key ArrowDown + press-key --key Enter.

Multi-step forms: Complete each step, wait for the next section to load, re-snapshot to discover new elements, then proceed.

Dynamic content: After interactions that change the page, use wait-for (with --selector or --text) or re-snapshot to see updated elements before continuing.

Scrolling: Use scroll --direction down to reveal below-the-fold content before snapshotting. Long pages may require multiple scrolls.

Hover menus / tooltips: Use hover to reveal hidden menus or tooltips, then re-snapshot to see newly revealed elements.

Verification

After critical actions (form submission, booking confirmation, checkout), take a screenshot and then read the saved image to visually verify results before reporting success to the user:

bash
assistant browser screenshot --output /tmp/verify.jpg

Then read the saved image to inspect it before reporting success. Use file_read if the screenshot was taken via bash, or host_file_read if it was taken via host_bash.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/vellum-browser-use of vellum-ai/vellum-assistant.

Open the folder on GitHubat commit 33cc983

Compare with similar skills

Vellum Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vellum Browser Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vellum Browser Use this skillvellum-ai/vellum-assistant1.4k—~3kAutomated safety check: PassMIT
Agent Browserquran/quran.com-frontend-next1.9k40 repos~3.3kAutomated safety check: PassNone
Dev-Browser CLI AutomationSawyerHood/dev-browser6.7k1 repos~455Automated safety check: PassMIT
Browser Automationopenclaw/openclaw392k—~2.9kAutomated safety check: PassMIT
Camoufox CLIBin-Huang/camoufox-cli3501 repos~4.5kAutomated safety check: PassMIT
BrowserVibiumDev/vibium2.9k—~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 40 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Dev-Browser CLI Automation

    SawyerHood/dev-browser

    Browser automation with persistent named pages via the dev-browser CLI. Use when users ask to navigate websites, fill forms, take screenshots, extract web…

    6.7k GitHub starsUsed in 1 repo~455 tokens
    Productivity & AutomationAuto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Camoufox CLI

    Bin-Huang/camoufox-cli

    Anti-detect browser automation CLI & Skills for AI agents. An agent skill from Bin-Huang/camoufox-cli.

    350 GitHub starsUsed in 1 repo~4.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser

    VibiumDev/vibium

    Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.

    2.9k GitHub stars~4.8k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts.

    50k GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated 2 days ago
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated 2 days ago
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check: warnings

Questions about Vellum Browser Use

What does Vellum Browser Use do?

Browse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant. Vellum Browser Use is an agent skill from vellum-ai/vellum-assistant.

When should I use Vellum Browser Use?

Vellum Browser Use fits situations like: tasks that involve Browser automation.

How do I install Vellum Browser Use in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a claude-code`. Or copy the skill folder (skills/vellum-browser-use in vellum-ai/vellum-assistant) into .claude/skills/vellum-browser-use in your project. Claude Code loads it when a task matches its description.

How do I install Vellum Browser Use in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a codex`. Or copy the skill folder (skills/vellum-browser-use in vellum-ai/vellum-assistant) into .agents/skills/vellum-browser-use in your project. Codex loads it when a task matches its description.

Can I use Vellum Browser Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vellum-browser-use, .gemini/skills/vellum-browser-use, .github/skills/vellum-browser-use and .opencode/skills/vellum-browser-use in your project.

What does Vellum Browser Use need to run?

SKILL.md names no scripts, command-line tools or credentials: Vellum Browser Use is instructions for the agent only. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Vellum Browser Use access the network?

SKILL.md names 2 domains. As links in the text: chromewebstore.google.com and vellum.ai. This is read from the text; nothing was executed.

Is Vellum Browser Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vellum Browser Use use?

Vellum Browser Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vellum Browser Use use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vellum Browser Use?

Skills that share tags, products or a category with Vellum Browser Use: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Dev-Browser CLI Automation (SawyerHood/dev-browser, 6.7k stars), Browser Automation (openclaw/openclaw, 392k stars) and Camoufox CLI (Bin-Huang/camoufox-cli, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vellum Browser Use?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,408 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.