Agent Browser
quran/quran.com-frontend-next
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
Browse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant.
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vellum-browser-use .claude/skills/vellum-browser-use && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .claude/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-useType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/vellum-browser-use .agents/skills/vellum-browser-use && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .agents/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/vellum-browser-use .cursor/skills/vellum-browser-use && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .cursor/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vellum-ai/vellum-assistant.git --path skills/vellum-browser-use--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/vellum-browser-use .gemini/skills/vellum-browser-use && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .gemini/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-useInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/vellum-browser-use .github/skills/vellum-browser-use && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .github/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vellum-ai/vellum-assistant vellum-browser-use --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/vellum-browser-use .opencode/skills/vellum-browser-use && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vellum-browser-use" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/vellum-browser-use into .opencode/skills/vellum-browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vellum-browser-use", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vellum-browser-useBrowse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant.
Vellum Browser Use is an agent skill from vellum-ai/vellum-assistant. Browse the web using assistant browser CLI commands
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Vellum personal assistants
It sits in Productivity & Automation, covering Browser automation. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 33cc983. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
chromewebstore.google.comvellum.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Vellum personal assistants
From compatibility in the SKILL.md frontmatter.
Vellum Browser Use loads about 3k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 1,325 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vellum-ai/vellum-assistant at commit 33cc983, republished under its MIT licence (© vellum-ai). 1,325 words, ~2,993 tokens.
.claude/skills/vellum-browser-use/SKILL.md (or your agent's skills folder).Use this skill to browse the web. All browser operations are executed through the assistant browser CLI, invoked via bash or host_bash. Each operation is a subcommand:
| Command | Description |
|---|---|
assistant browser navigate | Navigate to a URL |
assistant browser snapshot | List interactive elements on the current page |
assistant browser screenshot | Take a visual screenshot |
assistant browser click | Click an element |
assistant browser type | Type text into an input |
assistant browser press-key | Press a keyboard key |
assistant browser scroll | Scroll the page or a specific element |
assistant browser select-option | Select an option from a native <select> element |
assistant browser hover | Hover over an element to reveal menus/tooltips |
assistant browser wait-for | Wait for a condition |
assistant browser extract | Extract page text content |
assistant browser wait-for-download | Wait for a file download to complete |
assistant browser fill-credential | Fill a stored credential into a form field |
assistant browser attach | Attach the Chrome debugger to the active tab |
assistant browser detach | Detach the Chrome debugger from the active tab |
assistant browser close | Close the browser page |
assistant browser status | Diagnose browser backend readiness and setup steps |
Before using any browser commands, run assistant browser --json status first to check which browser backends are available. The status command returns JSON with readiness information for each backend mode:
assistant browser --json statusThe response includes:
recommendedMode — the best available backend (use this)modes[] — per-mode status with available, summary, and userActions (remediation steps)For --virtual-desktop, Chrome and desktop components are included in the assistant image. Browser and computer use start the existing desktop without installing dependencies. If components are missing, report the image problem; do not install packages or launch a separate Chrome process.
Use --browser-mode <mode> on the assistant browser parent command to pin the browser backend:
| Value | Backend | Description |
|---|---|---|
auto | Automatic | Default. Picks the best available backend based on context. |
extension | Chrome extension | Routes through the user's Chrome browser via the extension debugger. |
cdp-inspect | CDP inspect | Connects to an already-running Chrome instance via DevTools Protocol. |
local | Playwright | Drives a dedicated Playwright-managed Chromium instance. |
assistant browser --browser-mode extension navigate --url http://www.example.comThe Chrome extension (extension mode) is the preferred browser backend. It is:
If the status check shows the extension is not available, encourage the user to install and pair it:
The status response's userActions array for the extension mode provides these same steps when the extension is not connected.
If navigate, curl, or any fetch times out, hits an auth wall, or cannot reach a host (VPN, company login, internal dashboard):
If the user declines to install the extension:
cdp-inspect — Connects to an already-running Chrome instance via DevTools Protocol (Chrome 146+). Requires enabling remote debugging in Chrome settings.local — Drives a dedicated Playwright-managed Chromium instance. Last resort — does not use the user's browser profile.Only fall back to these if the user explicitly indicates they do not want to install the extension. Prefer cdp-inspect over local.
When multiple clients support host_browser (e.g. two Chrome profiles, a macOS client and a Chrome extension), use --target-client-id <id> on the assistant browser parent command to pin all operations in the invocation to one specific client:
assistant browser --target-client-id <client-id> navigate --url https://example.comObtain client IDs from:
assistant clients list --capability host_browserOmit --target-client-id when only one client is connected — the default interface-preference order (chrome-extension first, then macos) picks the best available client automatically.
On the Chrome extension backend, navigate opens a dedicated tab the first time it runs in a conversation and pins subsequent operations to it, so browsing never disturbs the tab the user is on (often the tab they're chatting with the assistant from). Later navigates reuse that pinned tab.
--new-tab — force a brand-new tab even when one is already pinned.--use-active-tab — navigate the user's currently-active tab instead of a dedicated one.Both flags are ignored on the local and cdp-inspect backends, which manage their own browser context.
Use --session <id> on the assistant browser parent command to group sequential operations so they share browser state (same page, cookies, etc.). Different session IDs create independent browser contexts.
assistant browser --session myflow navigate --url https://example.com
assistant browser --session myflow snapshot
assistant browser --session myflow click --element-id e3Omitting --session uses the default session.
Use --json on the assistant browser parent command to get structured JSON output suitable for parsing in scripts:
assistant browser --json navigate --url https://example.com
# {"ok":true,"content":"Page title: Example Domain"}
assistant browser --json snapshot
# {"ok":true,"content":"...element list..."}
assistant browser --json screenshot
# {"ok":true,"content":"...","screenshots":[{"mediaType":"image/jpeg","data":"<base64>"}]}Error responses use {"ok":false,"error":"..."}.
To save a screenshot to disk, use --output <path>:
assistant browser screenshot --output page.jpg
assistant browser screenshot --full-page --output full.jpgTo receive base64 screenshot data in JSON output:
assistant browser --json screenshotThe response includes a screenshots array with mediaType and data (base64) fields.
assistant browser --json status to check backend readiness — if the extension is not available, help the user install itassistant browser attach to establish the sessionassistant browser navigate --url <url> to load a pageassistant browser snapshot to discover interactive elementsclick, type, press-key, scroll, select-option, or hover to interactassistant browser extract or assistant browser screenshot --output <path> to capture resultsassistant browser detach when you are done — this releases the debugger so the user can browse freelyTreat every CAPTCHA and bot-detection challenge as a request for human help, including drag-to-verify sliders, press-and-hold checks, verification checkboxes and image puzzles. Stop before interacting with the challenge, even if its controls look easy to automate. Do not try it yourself, retry it, or script a solution.
In the virtual desktop, request the desktop-help card immediately, using one short sentence for the needed action. Wait for Done or Skip. After Done, take a fresh snapshot; if verification remains, ask for help again. On other browser backends, ask the user to complete verification in their browser and wait for confirmation.
For ordinary logins, use saved credentials or securely prompt for missing credentials, then fill the form yourself. A CAPTCHA on a login page still requires human help.
Date pickers / calendars: Click the date input to open the picker, re-snapshot to see calendar controls, click month navigation arrows to reach the target month, then click the target date. For <input type="date">, use type with YYYY-MM-DD format.
Native <select> elements: Use select-option with --value, --label, or --index. Do not try to click individual <option> elements.
ARIA / custom dropdowns: Click to open, take a new snapshot, then click the desired option by --element-id.
Autocomplete inputs: Type the search text, wait 500-1000ms (wait-for --duration), re-snapshot for suggestions, then click the suggestion or use press-key --key ArrowDown + press-key --key Enter.
Multi-step forms: Complete each step, wait for the next section to load, re-snapshot to discover new elements, then proceed.
Dynamic content: After interactions that change the page, use wait-for (with --selector or --text) or re-snapshot to see updated elements before continuing.
Scrolling: Use scroll --direction down to reveal below-the-fold content before snapshotting. Long pages may require multiple scrolls.
Hover menus / tooltips: Use hover to reveal hidden menus or tooltips, then re-snapshot to see newly revealed elements.
After critical actions (form submission, booking confirmation, checkout), take a screenshot and then read the saved image to visually verify results before reporting success to the user:
assistant browser screenshot --output /tmp/verify.jpgThen read the saved image to inspect it before reporting success. Use file_read if the screenshot was taken via bash, or host_file_read if it was taken via host_bash.
© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/vellum-browser-use of vellum-ai/vellum-assistant.
Open the folder on GitHubat commit 33cc983
Vellum Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vellum Browser Use this skillvellum-ai/vellum-assistant | 1.4k | — | ~3k | Automated safety check: Pass | MIT | |
| Agent Browserquran/quran.com-frontend-next | 1.9k | 40 repos | ~3.3k | Automated safety check: Pass | None | |
| Dev-Browser CLI AutomationSawyerHood/dev-browser | 6.7k | 1 repos | ~455 | Automated safety check: Pass | MIT | |
| Browser Automationopenclaw/openclaw | 392k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Camoufox CLIBin-Huang/camoufox-cli | 350 | 1 repos | ~4.5k | Automated safety check: Pass | MIT | |
| BrowserVibiumDev/vibium | 2.9k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 |
quran/quran.com-frontend-next
Automates browser interactions for web testing, form filling, screenshots, and data extraction.
SawyerHood/dev-browser
Browser automation with persistent named pages via the dev-browser CLI. Use when users ask to navigate websites, fill forms, take screenshots, extract web…
openclaw/openclaw
A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.
Bin-Huang/camoufox-cli
Anti-detect browser automation CLI & Skills for AI agents. An agent skill from Bin-Huang/camoufox-cli.
VibiumDev/vibium
Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.
moeru-ai/airi
Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts.
vellum-ai/vellum-assistant
Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.
vellum-ai/vellum-assistant
Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration
vellum-ai/vellum-assistant
Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity
vellum-ai/vellum-assistant
Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.
vellum-ai/vellum-assistant
A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.
vellum-ai/vellum-assistant
Connect a Slack app to the Vellum Assistant via Socket Mode.
Categories
Browse the web using assistant browser CLI commands. An agent skill from vellum-ai/vellum-assistant. Vellum Browser Use is an agent skill from vellum-ai/vellum-assistant.
Vellum Browser Use fits situations like: tasks that involve Browser automation.
Run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a claude-code`. Or copy the skill folder (skills/vellum-browser-use in vellum-ai/vellum-assistant) into .claude/skills/vellum-browser-use in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a codex`. Or copy the skill folder (skills/vellum-browser-use in vellum-ai/vellum-assistant) into .agents/skills/vellum-browser-use in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill vellum-browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vellum-browser-use, .gemini/skills/vellum-browser-use, .github/skills/vellum-browser-use and .opencode/skills/vellum-browser-use in your project.
SKILL.md names no scripts, command-line tools or credentials: Vellum Browser Use is instructions for the agent only. Compatibility (from SKILL.md): Designed for Vellum personal assistants.
SKILL.md names 2 domains. As links in the text: chromewebstore.google.com and vellum.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vellum Browser Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vellum Browser Use: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Dev-Browser CLI Automation (SawyerHood/dev-browser, 6.7k stars), Browser Automation (openclaw/openclaw, 392k stars) and Camoufox CLI (Bin-Huang/camoufox-cli, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,408 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.
Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.