Dev Browser Automation
MemTensor/MemOS
Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.
Drives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows.
$ npx skills add slopus/happy --skill agent-browser -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install slopus/happy agent-browser --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agent-browser .claude/skills/agent-browser && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .claude/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/slopus/happy/tree/main/.agents/skills/agent-browserType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add slopus/happy --skill agent-browser -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install slopus/happy agent-browser --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/agent-browser .agents/skills/agent-browser && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .agents/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slopus/happy --skill agent-browser -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install slopus/happy agent-browser --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/agent-browser .cursor/skills/agent-browser && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .cursor/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/slopus/happy.git --path .agents/skills/agent-browser--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add slopus/happy --skill agent-browser -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install slopus/happy agent-browser --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/agent-browser .gemini/skills/agent-browser && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .gemini/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install slopus/happy agent-browserInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add slopus/happy --skill agent-browser -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/agent-browser .github/skills/agent-browser && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .github/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slopus/happy --skill agent-browser -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install slopus/happy agent-browser --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slopus/happy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/agent-browser .opencode/skills/agent-browser && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-browser" agent skill from https://github.com/slopus/happy/tree/main/.agents/skills/agent-browser into .opencode/skills/agent-browser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-browser", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-browserDrives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows.
The skill teaches the agent to drive a real browser with the `agent-browser` CLI. Every task follows the same loop: open a URL, take a snapshot with `snapshot -i` that lists interactive elements with short refs such as `@e1`, interact with those refs by clicking, filling or selecting, and snapshot again whenever the page changes, because refs are invalidated by navigation, form submissions and dynamically loaded content.
A background daemon keeps the browser alive between commands, so several can be chained with `&&` in one call, for example open, wait for network idle and snapshot, but they should run separately when the output of one has to be read first. Common patterns shown are form submission, logging in once and saving state for reuse, and extracting text from elements. An essential-commands list covers navigation and more, and the excerpt ends at the annotated-screenshot section for vision mode.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 7ea7017. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(npx agent-browser:*)Bash(agent-browser:*)From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Browser CLI loads about 1.7k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 191 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from slopus/happy at commit 7ea7017, republished under its MIT licence (© slopus). 191 words, ~1,748 tokens.
.claude/skills/agent-browser/SKILL.md (or your agent's skills folder).Every browser automation follows this pattern:
agent-browser open <url>agent-browser snapshot -i (get element refs like @e1, @e2)agent-browser open https://example.com/form
agent-browser snapshot -i
# Output: @e1 [input type="email"], @e2 [input type="password"], @e3 [button] "Submit"
agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i # Check resultCommands can be chained with && in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls.
# Chain open + wait + snapshot in one call
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser snapshot -i
# Chain multiple interactions
agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "password123" && agent-browser click @e3
# Navigate and capture
agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser screenshot page.pngWhen to chain: Use && when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs).
# Navigation
agent-browser open <url> # Navigate (aliases: goto, navigate)
agent-browser close # Close browser
# Snapshot
agent-browser snapshot -i # Interactive elements with refs (recommended)
agent-browser snapshot -i -C # Include cursor-interactive elements (divs with onclick, cursor:pointer)
agent-browser snapshot -s "#selector" # Scope to CSS selector
# Interaction (use @refs from snapshot)
agent-browser click @e1 # Click element
agent-browser click @e1 --new-tab # Click and open in new tab
agent-browser fill @e2 "text" # Clear and type text
agent-browser type @e2 "text" # Type without clearing
agent-browser select @e1 "option" # Select dropdown option
agent-browser check @e1 # Check checkbox
agent-browser press Enter # Press key
agent-browser keyboard type "text" # Type at current focus (no selector)
agent-browser keyboard inserttext "text" # Insert without key events
agent-browser scroll down 500 # Scroll page
agent-browser scroll down 500 --selector "div.content" # Scroll within a specific container
# Get information
agent-browser get text @e1 # Get element text
agent-browser get url # Get current URL
agent-browser get title # Get page title
# Wait
agent-browser wait @e1 # Wait for element
agent-browser wait --load networkidle # Wait for network idle
agent-browser wait --url "**/page" # Wait for URL pattern
agent-browser wait 2000 # Wait milliseconds
# Downloads
agent-browser download @e1 ./file.pdf # Click element to trigger download
agent-browser wait --download ./output.zip # Wait for any download to complete
agent-browser --download-path ./downloads open <url> # Set default download directory
# Capture
agent-browser screenshot # Screenshot to temp dir
agent-browser screenshot --full # Full page screenshot
agent-browser screenshot --annotate # Annotated screenshot with numbered element labels
agent-browser pdf output.pdf # Save as PDF
# Diff (compare page states)
agent-browser diff snapshot # Compare current vs last snapshot
agent-browser diff snapshot --baseline before.txt # Compare current vs saved file
agent-browser diff screenshot --baseline before.png # Visual pixel diff
agent-browser diff url <url1> <url2> # Compare two pages
agent-browser diff url <url1> <url2> --wait-until networkidle # Custom wait strategy
agent-browser diff url <url1> <url2> --selector "#main" # Scope to elementagent-browser open https://example.com/signup
agent-browser snapshot -i
agent-browser fill @e1 "Jane Doe"
agent-browser fill @e2 "jane@example.com"
agent-browser select @e3 "California"
agent-browser check @e4
agent-browser click @e5
agent-browser wait --load networkidle# Login once and save state
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "$USERNAME"
agent-browser fill @e2 "$PASSWORD"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json
# Reuse in future sessions
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboardagent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e5 # Get specific element text
agent-browser get text body > page.txt # Get all page text
# JSON output for parsing
agent-browser snapshot -i --json
agent-browser get text @e1 --jsonRefs (@e1, @e2, etc.) are invalidated when the page changes. Always re-snapshot after:
agent-browser click @e5 # Navigates to new page
agent-browser snapshot -i # MUST re-snapshot
agent-browser click @e1 # Use new refsUse --annotate to take a screenshot with numbered labels overlaid on interactive elements.
agent-browser screenshot --annotate
# Output includes the image path and a legend:
# [1] @e1 button "Submit"
# [2] @e2 link "Home"
# [3] @e3 textbox "Email"
agent-browser click @e2 # Click using ref from annotated screenshot# Simple expressions
agent-browser eval 'document.title'
# Complex JS: use --stdin with heredoc (RECOMMENDED)
agent-browser eval --stdin <<'EVALEOF'
JSON.stringify(
Array.from(document.querySelectorAll("img"))
.filter(i => !i.alt)
.map(i => ({ src: i.src.split("/").pop(), width: i.width }))
)
EVALEOFAlways close your browser session when done:
agent-browser close # Close default session
agent-browser --session agent1 close # Close specific session© slopus, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/agent-browser of slopus/happy.
Open the folder on GitHubat commit 7ea7017
Agent Browser CLI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Browser CLI this skillslopus/happy | 24k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Dev Browser AutomationMemTensor/MemOS | 12k | 3 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Core Guide for agent-browservercel-labs/agent-browser | 44k | 4 repos | ~9.5k | Automated safety check: Pass | Apache-2.0 | |
| Browser Control with Omowrightcode-yeongyu/oh-my-openagent | 70k | — | ~2.2k | Automated safety check: Pass | Custom licence | |
| Agent Browsernanocoai/nanoclaw | 31k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Playwright Bowserdisler/bowser | 265 | — | ~1.1k | Automated safety check: Notes | None |
MemTensor/MemOS
Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.
vercel-labs/agent-browser
Core usage guide for the agent-browser CLI: the snapshot-and-ref workflow for navigating, clicking, filling forms, extracting data and running parallel sessions.
code-yeongyu/oh-my-openagent
Drives a real browser through the omowright library, either the user's own signed-in browser or a separate browser the code launches, for forms, QA, screenshots and scraping.
nanocoai/nanoclaw
Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots.
disler/bowser
Headless browser automation using Playwright CLI. An agent skill from disler/bowser.
ynulihao/AgentSkillOS
Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation.
slopus/happy
Searches past Claude Code, Codex and Cursor sessions and summarizes what was worked on, tried or decided, using extraction scripts instead of reading raw logs.
slopus/happy
Traces how an action moves through your code and draws it as a compact ASCII tree: functions called, payload types, state changes and components that re-render.
slopus/happy
Local development guide for the Happy pnpm monorepo: install, build, test and run the CLI, server, Expo app and Tauri desktop packages.
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
slopus/happy
Tests interactive CLI and TUI programs with Microsoft's tui-test, driving prompts, arrow keys and screen output in a real pseudo-terminal.
slopus/happy
Walks you through releasing a component of the Happy monorepo (CLI, mobile, web or server) from a clean local main that matches origin/main.
Categories
Drives a real browser from the shell with the agent-browser CLI: open pages, snapshot elements by ref, click, fill and extract, for testing web flows. The skill teaches the agent to drive a real browser with the `agent-browser` CLI. Every task follows the same loop: open a URL, take a snapshot with `snapshot -i` that lists interactive elements with short refs such as `@e1`, interact with those refs by clicking, filling or selecting, and snapshot again whenever the page changes, because refs are invalidated by navigation, form submissions and dynamically loaded content.
Agent Browser CLI fits situations like: testing a web flow in a real browser; filling and submitting a form through the CLI; logging in once and reusing the saved session state; pulling text out of page elements.
Run `npx skills add slopus/happy --skill agent-browser -a claude-code`. Or copy the skill folder (.agents/skills/agent-browser in slopus/happy) into .claude/skills/agent-browser in your project. Claude Code loads it when a task matches its description.
Run `npx skills add slopus/happy --skill agent-browser -a codex`. Or copy the skill folder (.agents/skills/agent-browser in slopus/happy) into .agents/skills/agent-browser in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slopus/happy --skill agent-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-browser, .gemini/skills/agent-browser, .github/skills/agent-browser and .opencode/skills/agent-browser in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Browser CLI is instructions for the agent only. Our summary lists: The agent-browser CLI, run directly or through npx. Its frontmatter pre-approves these tools: Bash(npx agent-browser:*), Bash(agent-browser:*).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Browser CLI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Browser CLI: Dev Browser Automation (MemTensor/MemOS, 12k stars), Core Guide for agent-browser (vercel-labs/agent-browser, 44k stars), Browser Control with Omowright (code-yeongyu/oh-my-openagent, 70k stars) and Agent Browser (nanocoai/nanoclaw, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
slopus (a GitHub organization) maintains it in slopus/happy, which has 24,039 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.
Source: slopus/happy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.