Agent skill

Dev Browser Automation

by MemTensor in MemTensor/MemOS

Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.

Apache-2.0Auto-check passedProductivity & Automation

Install Dev Browser Automation

skills CLI
$ npx skills add MemTensor/MemOS --skill dev-browser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MemTensor/MemOS dev-browser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MemTensor/MemOS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/apps/openwork-memos-integration/apps/desktop/skills/dev-browser .claude/skills/dev-browser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dev-browser
GitHub stars
12k
Used in
3 other repos
Token cost
~1.7k tokens
SKILL.md length
516 words
Files
17 (incl. scripts, references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data.

  • Works in 2 steps: Scripts call client.page("name") just… → Automation runs on the user's actual…
  • Navigating a website, clicking through it and filling out forms
  • SKILL.md covers Choosing Your Approach, Setup, Writing Scripts and Workflow Loop, plus 5 more sections
  • Runs TypeScript and Shell scripts from its folder; calls npm and npx

What it does

The skill keeps named pages alive across script executions, so the agent writes small scripts that each do one thing, such as navigate, click, fill or check, and logs state at the end to decide the next step. It picks an approach by situation: read the source of local sites to write selectors directly, use getAISnapshot() and selectSnapshotRef() to find and operate elements on unfamiliar pages, and take screenshots for visual feedback.

There are two modes. Standalone mode, the default, launches a new Chromium browser through server.sh and can run headless, and the agent waits for a Ready message before running scripts. Extension mode connects to your existing Chrome through a relay server, so automation runs in a session where you are already logged in, provided the extension is installed and active. Scripts run with npx tsx from the skill folder, and a reference note covers scraping.

When your agent uses it

  • Navigating a website, clicking through it and filling out forms
  • Taking screenshots of a page to check how it looks
  • Extracting data from a site that needs several steps to reach
  • Testing a local web app by driving it in a browser
  • Automating something behind a login using your own Chrome session

Example prompts

  • “Open my staging site, fill in the signup form with a test user and screenshot the confirmation page.”
  • “Go to the pricing page and pull the plan names and prices into a table.”
  • “Using my logged-in Chrome, open the analytics dashboard and screenshot the weekly chart.”

Requirements

  • Node.js with npx and tsx
  • Chromium, or Chrome with the dev-browser extension for extension mode

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Scripts call client.page("name") just like the normal mode to create new pages / connect to existing ones.
  2. Automation runs on the user's actual browser session

What it can do on your machine

Read from SKILL.md and the folder at commit a7367d0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (TypeScript and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dev Browser Automation loads about 1.7k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 516 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from MemTensor/MemOS at commit a7367d0, republished under its Apache-2.0 licence (© MemTensor). 516 words, ~1,722 tokens.

Download SKILL.mdSave it as .claude/skills/dev-browser/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
dev-browser
description
Browser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.

Dev Browser Skill

Browser automation that maintains page state across script executions. Write small, focused scripts to accomplish tasks incrementally. Once you've proven out part of a workflow and there is repeated work to be done, you can write a script to do the repeated work in a single execution.

Choosing Your Approach

  • Local/source-available sites: Read the source code first to write selectors directly
  • Unknown page layouts: Use getAISnapshot() to discover elements and selectSnapshotRef() to interact with them
  • Visual feedback: Take screenshots to see what the user sees

Setup

Two modes available. Ask the user if unclear which to use.

Standalone Mode (Default)

Launches a new Chromium browser for fresh automation sessions.

bash
./skills/dev-browser/server.sh &

Add --headless flag if user requests it. Wait for the Ready message before running scripts.

Extension Mode

Connects to user's existing Chrome browser. Use this when:

  • The user is already logged into sites and wants you to do things behind an authed experience that isn't local dev.
  • The user asks you to use the extension

Important: The core flow is still the same. You create named pages inside of their browser.

Start the relay server:

bash
cd skills/dev-browser && npm i && npm run start-extension &

Wait for Waiting for extension to connect... followed by Extension connected in the console. To know that a client has connected and the browser is ready to be controlled. Workflow:

  1. Scripts call client.page("name") just like the normal mode to create new pages / connect to existing ones.
  2. Automation runs on the user's actual browser session

If the extension hasn't connected yet, tell the user to launch and activate it. Download link: https://github.com/SawyerHood/dev-browser/releases

Writing Scripts

Run all scripts from skills/dev-browser/ directory. The @/ import alias requires this directory's config.

Execute scripts inline using heredocs:

bash
cd skills/dev-browser && npx tsx <<'EOF'
import { connect, waitForPageLoad } from "@/client.js";

const client = await connect();
// Create page with custom viewport size (optional)
const page = await client.page("example", { viewport: { width: 1920, height: 1080 } });

await page.goto("https://example.com");
await waitForPageLoad(page);

console.log({ title: await page.title(), url: page.url() });
await client.disconnect();
EOF

Write to tmp/ files only when the script needs reuse, is complex, or user explicitly requests it.

Show full SKILL.md (219 more words)Show less
Key Principles
  1. Small scripts: Each script does ONE thing (navigate, click, fill, check)
  2. Evaluate state: Log/return state at the end to decide next steps
  3. Descriptive page names: Use "checkout", "login", not "main"
  4. Disconnect to exit: await client.disconnect() - pages persist on server
  5. Plain JS in evaluate: page.evaluate() runs in browser - no TypeScript syntax

Workflow Loop

Follow this pattern for complex tasks:

  1. Write a script to perform one action
  2. Run it and observe the output
  3. Evaluate - did it work? What's the current state?
  4. Decide - is the task complete or do we need another script?
  5. Repeat until task is done
No TypeScript in Browser Context

Code passed to page.evaluate() runs in the browser, which doesn't understand TypeScript:

typescript
// ✅ Correct: plain JavaScript
const text = await page.evaluate(() => {
  return document.body.innerText;
});

// ❌ Wrong: TypeScript syntax will fail at runtime
const text = await page.evaluate(() => {
  const el: HTMLElement = document.body; // Type annotation breaks in browser!
  return el.innerText;
});

Scraping Data

For scraping large datasets, intercept and replay network requests rather than scrolling the DOM. See references/scraping.md for the complete guide covering request capture, schema discovery, and paginated API replay.

Client API

typescript
const client = await connect();

// Get or create named page (viewport only applies to new pages)
const page = await client.page("name");
const pageWithSize = await client.page("name", { viewport: { width: 1920, height: 1080 } });

const pages = await client.list(); // List all page names
await client.close("name"); // Close a page
await client.disconnect(); // Disconnect (pages persist)

// ARIA Snapshot methods
const snapshot = await client.getAISnapshot("name"); // Get accessibility tree
const element = await client.selectSnapshotRef("name", "e5"); // Get element by ref

The page object is a standard Playwright Page.

Waiting

typescript
import { waitForPageLoad } from "@/client.js";

await waitForPageLoad(page); // After navigation
await page.waitForSelector(".results"); // For specific elements
await page.waitForURL("**/success"); // For specific URL

Inspecting Page State

Screenshots
typescript
await page.screenshot({ path: "tmp/screenshot.png" });
await page.screenshot({ path: "tmp/full.png", fullPage: true });
ARIA Snapshot (Element Discovery)

Use getAISnapshot() to discover page elements. Returns YAML-formatted accessibility tree:

yaml
- banner:
  - link "Hacker News" [ref=e1]
  - navigation:
    - link "new" [ref=e2]
- main:
  - list:
    - listitem:
      - link "Article Title" [ref=e8]
      - link "328 comments" [ref=e9]
- contentinfo:
  - textbox [ref=e10]
    - /placeholder: "Search"

Interpreting refs:

  • [ref=eN] - Element reference for interaction (visible, clickable elements only)
  • [checked], [disabled], [expanded] - Element states
  • [level=N] - Heading level
  • /url:, /placeholder: - Element properties

Interacting with refs:

typescript
const snapshot = await client.getAISnapshot("hackernews");
console.log(snapshot); // Find the ref you need

const element = await client.selectSnapshotRef("hackernews", "e2");
await element.click();

Error Recovery

Page state persists after failures. Debug with:

bash
cd skills/dev-browser && npx tsx <<'EOF'
import { connect } from "@/client.js";

const client = await connect();
const page = await client.page("hackernews");

await page.screenshot({ path: "tmp/debug.png" });
console.log({
  url: page.url(),
  title: await page.title(),
  bodyText: await page.textContent("body").then((t) => t?.slice(0, 200)),
});

await client.disconnect();
EOF

© MemTensor, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files (scripts, references) in apps/openwork-memos-integration/apps/desktop/skills/dev-browser of MemTensor/MemOS.

  • SKILL.md
  • .gitignore
  • package.json
  • references/scraping.md
  • scripts/start-relay.ts
  • scripts/start-server.ts
  • server.sh
  • src/client.ts
  • src/index.ts
  • src/relay.ts
  • src/snapshot/__tests__/snapshot.test.ts
  • src/snapshot/browser-script.ts
  • src/snapshot/index.ts
  • src/snapshot/inject.ts
  • src/types.ts
  • tsconfig.json
  • … and 1 more

Open the folder on GitHubat commit a7367d0

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in MemTensor/MemOS, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Dev Browser Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dev Browser Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dev Browser Automation this skillMemTensor/MemOS12k3 repos~1.7kAutomated safety check: PassApache-2.0
Core Guide for agent-browservercel-labs/agent-browser44k4 repos~9.5kAutomated safety check: PassApache-2.0
Browser Control with Omowrightcode-yeongyu/oh-my-openagent70k—~2.2kAutomated safety check: PassCustom licence
Agent Browsernanocoai/nanoclaw31k3 repos~1.6kAutomated safety check: PassMIT
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone
Browser Automationynulihao/AgentSkillOS617—~2.2kAutomated safety check: PassNone

Similar skills

  • Core Guide for agent-browser

    vercel-labs/agent-browser

    Official

    Core usage guide for the agent-browser CLI: the snapshot-and-ref workflow for navigating, clicking, filling forms, extracting data and running parallel sessions.

    44k GitHub starsUsed in 4 repos~9.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser Control with Omowright

    code-yeongyu/oh-my-openagent

    Drives a real browser through the omowright library, either the user's own signed-in browser or a separate browser the code launches, for forms, QA, screenshots and scraping.

    70k GitHub stars~2.2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Agent Browser

    nanocoai/nanoclaw

    Drives a web browser from the shell with the agent-browser CLI: open pages, read an element snapshot, click and fill by reference, grab text and screenshots.

    31k GitHub starsUsed in 3 repos~1.6k tokens
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Browser Automation

    ynulihao/AgentSkillOS

    Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation.

    617 GitHub stars~2.2k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Use

    aiskillstore/marketplace

    Browser automation using Playwright MCP. An agent skill from aiskillstore/marketplace.

    430 GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from MemTensor/MemOS

  • Controls a browser through BrowserWing's local HTTP API: navigation, clicks, typing, data extraction, accessibility snapshots, screenshots and batch operations.

    12k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check passed
  • Ask User Question

    MemTensor/MemOS

    Shows a question as a modal in the interface to clarify a task, collect a preference or get approval, since the user cannot see terminal output.

    12k GitHub stars~1k tokensUpdated 8 days ago
    Auto-check passed
  • BrowserWing Admin

    MemTensor/MemOS

    Installs, configures and operates BrowserWing, a browser automation platform, covering Chrome setup, LLM provider configuration, and creating or running automation scripts.

    12k GitHub stars~4.1k tokensUpdated 8 days ago
    Auto-check: notes
  • Safe File Deletion

    MemTensor/MemOS

    Enforces explicit user permission before any file deletion. Activates when you're about to use rm, unlink, fs.rm, or any operation that removes files from…

    12k GitHub stars~291 tokensUpdated 8 days ago
    Auto-check passed
  • Explains when to call MemOS's own memory tools to search past conversations, after the automatic per-turn recall hook comes up empty.

    12k GitHub stars~3.5k tokensUpdated 8 days ago
    Auto-check: warnings

Questions about Dev Browser Automation

What does Dev Browser Automation do?

Automates a real browser through short TypeScript scripts that keep page state between runs, for navigating, filling forms, taking screenshots and extracting data. The skill keeps named pages alive across script executions, so the agent writes small scripts that each do one thing, such as navigate, click, fill or check, and logs state at the end to decide the next step. It picks an approach by situation: read the source of local sites to write selectors directly, use getAISnapshot() and selectSnapshotRef() to find and operate elements on unfamiliar pages, and take screenshots for visual feedback.

When should I use Dev Browser Automation?

Dev Browser Automation fits situations like: navigating a website, clicking through it and filling out forms; taking screenshots of a page to check how it looks; extracting data from a site that needs several steps to reach; testing a local web app by driving it in a browser.

How do I install Dev Browser Automation in Claude Code?

Run `npx skills add MemTensor/MemOS --skill dev-browser -a claude-code`. Or copy the skill folder (apps/openwork-memos-integration/apps/desktop/skills/dev-browser in MemTensor/MemOS) into .claude/skills/dev-browser in your project. Claude Code loads it when a task matches its description.

How do I install Dev Browser Automation in Codex?

Run `npx skills add MemTensor/MemOS --skill dev-browser -a codex`. Or copy the skill folder (apps/openwork-memos-integration/apps/desktop/skills/dev-browser in MemTensor/MemOS) into .agents/skills/dev-browser in your project. Codex loads it when a task matches its description.

Can I use Dev Browser Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MemTensor/MemOS --skill dev-browser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dev-browser, .gemini/skills/dev-browser, .github/skills/dev-browser and .opencode/skills/dev-browser in your project.

What does Dev Browser Automation need to run?

Going by SKILL.md and its folder, Dev Browser Automation needs TypeScript and a shell for the scripts in its folder and the command-line tools its instructions call (npm and npx). Our summary lists: Node.js with npx and tsx; Chromium, or Chrome with the dev-browser extension for extension mode.

Does Dev Browser Automation access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Dev Browser Automation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dev Browser Automation use?

Dev Browser Automation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dev Browser Automation use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Dev Browser Automation?

Skills that share tags, products or a category with Dev Browser Automation: Core Guide for agent-browser (vercel-labs/agent-browser, 44k stars), Browser Control with Omowright (code-yeongyu/oh-my-openagent, 70k stars), Agent Browser (nanocoai/nanoclaw, 31k stars) and Playwright Bowser (disler/bowser, 265 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dev Browser Automation?

MemTensor (a GitHub organization) maintains it in MemTensor/MemOS, which has 11,738 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 29, 2026.

Source: MemTensor/MemOS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.