C Browser
daxaur/openpaw
Headless browser automation — navigate pages, click elements, fill forms, take screenshots, scrape content using agent-browser or playwright-cli.
Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.
$ npx skills add letta-ai/letta-code --skill browser-use -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install letta-ai/letta-code browser-use --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/builtin/browser-use .claude/skills/browser-use && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .claude/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-useType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add letta-ai/letta-code --skill browser-use -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install letta-ai/letta-code browser-use --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/builtin/browser-use .agents/skills/browser-use && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .agents/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add letta-ai/letta-code --skill browser-use -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install letta-ai/letta-code browser-use --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/builtin/browser-use .cursor/skills/browser-use && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .cursor/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/letta-ai/letta-code.git --path src/skills/builtin/browser-use--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add letta-ai/letta-code --skill browser-use -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install letta-ai/letta-code browser-use --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/builtin/browser-use .gemini/skills/browser-use && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .gemini/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install letta-ai/letta-code browser-useInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add letta-ai/letta-code --skill browser-use -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/builtin/browser-use .github/skills/browser-use && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .github/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add letta-ai/letta-code --skill browser-use -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install letta-ai/letta-code browser-use --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/builtin/browser-use .opencode/skills/browser-use && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser-use" agent skill from https://github.com/letta-ai/letta-code/tree/main/src/skills/builtin/browser-use into .opencode/skills/browser-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-use", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser-useControl a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.
Browser Use is an agent skill from letta-ai/letta-code. Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video. Load only when the user asks to open or automate a browser, interact with or test rendered page UI, scrape a site that needs browser execution, or capture a browser screenshot or video. Do not load for backend logs, traces, API or stream events, source-code inspection, or plain HTTP or web research that does not require a browser.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/recording.md`).
It sits in Productivity & Automation, covering Browser automation, Web scraping and Forms and invoices. The repository describes itself as: Stateful agents that are like people, with memory, identity, and the ability to learn and adapt. The licence is Apache-2.0.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 253a3bc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and bash).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
chromedevtools.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser Use loads about 3.3k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 1,112 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from letta-ai/letta-code at commit 253a3bc, republished under its Apache-2.0 licence (© letta-ai). 1,112 words, ~3,295 tokens.
.claude/skills/browser-use/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Drive the browser through its native Chrome DevTools Protocol over the
remote-debugging WebSocket. This works with zero dependencies: launch the
browser with --remote-debugging-port, then talk JSON over fetch and the
built-in WebSocket global (available in Bun and Node ≥ 22 — no ws package).
If the project already has Playwright or Puppeteer installed, using it is usually simpler — reach for raw CDP when no automation library is available, when protocol-level control is needed, or when recording a deterministic visual demo.
Protocol reference: https://chromedevtools.github.io/devtools-protocol/.
The running browser's exact schema is at http://127.0.0.1:<port>/json/protocol;
tip-of-tree docs can differ from the installed version.
When the computer has a display, prefer a visible (headful) browser for any task the user might watch or take over: clicking or typing, forms, sign-in, checkout/payment, CAPTCHAs or bot protection, and user handoff. Most browser tasks exist because plain HTTP is not enough; a headless browser is more likely to trigger bot protection and gives the user no way to observe or step in. Visible does not mean pixel-driven: keep operating the page over CDP, and the user sees every action in the window.
Use headless mode only for work the user explicitly wants in the background and that cannot require interaction or handoff, such as read-only scraping, CI, or screenshot/PDF generation, or when no display exists. A headless page does not satisfy a request to open or reopen a site in a browser the user can see.
When the user asks to review, watch, or take over, leave that browser window open after the task. Do not kill or close it before replying.
/json/list; pick the "page" target by URL or title.webSocketDebuggerUrl and enable only the domains you need
(usually Page, Runtime, DOM, Input; add Network, Log when debugging).Input.* for user-like interactions; use Runtime.evaluate for
inspection, coordinate math, and setup with no meaningful user interaction.Any Chromium-based browser supports CDP (Chrome, Chromium, Edge, Brave). Probe in order:
# macOS
for c in "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
"/Applications/Chromium.app/Contents/MacOS/Chromium" \
"/Applications/Microsoft Edge.app/Contents/MacOS/Microsoft Edge" \
"/Applications/Brave Browser.app/Contents/MacOS/Brave Browser"; do
[ -x "$c" ] && { echo "$c"; break; }
done
# Linux
for c in google-chrome google-chrome-stable chromium chromium-browser microsoft-edge brave-browser; do
command -v "$c" && break
doneOn Windows, check %ProgramFiles%\Google\Chrome\Application\chrome.exe,
%ProgramFiles(x86)%\..., %LocalAppData%\Google\Chrome\Application\chrome.exe,
and the same patterns for Microsoft\Edge.
Any browser found by the probe above works identically — use it. If truly no Chromium-based browser exists, do not install or download one automatically. Tell the user that browser use requires Chrome or another Chromium-based browser and recommend either:
Wait for the user to choose. Do not silently replace the browser task with plain HTTP or claim browser automation succeeded.
Use a disposable profile and a fixed port. Chrome refuses to run as root
without --no-sandbox, so add that flag when id -u is 0:
chrome_args=( \
--remote-debugging-port=9222 \
--user-data-dir=/tmp/cdp-profile \
--window-size=1440,900 \
--force-device-scale-factor=1 \
--no-first-run \
--no-default-browser-check \
)
[ "$(id -u)" -eq 0 ] && chrome_args+=(--no-sandbox)
"$CHROME" "${chrome_args[@]}" https://example.comAdd --headless=new only for explicitly invisible work or when no display
exists (see "Visible by default" above).
With --remote-debugging-port=0, read the chosen port from
<user-data-dir>/DevToolsActivePort. Launch in the background and poll
http://127.0.0.1:9222/json/version until it responds.
HTTP endpoints: /json/version (browser metadata + browser-level WebSocket URL),
/json/list (targets), PUT /json/new?<url> (open tab),
/json/activate/<id>, /json/close/<id>, /json/protocol (schema).
Attach to the page target for Page/DOM/Runtime/Input work; use the
browser target only for browser-wide commands (target control, downloads,
browser contexts).
CDP messages are JSON with monotonically increasing request ids. Run with bun:
const targets = await (await fetch("http://127.0.0.1:9222/json/list")).json();
const target = targets.find((t: any) => t.type === "page");
if (!target) throw new Error("No page target");
const ws = new WebSocket(target.webSocketDebuggerUrl);
await new Promise((resolve, reject) => {
ws.onopen = resolve;
ws.onerror = reject;
});
let nextId = 0;
const pending = new Map();
ws.onmessage = (event) => {
const msg = JSON.parse(String(event.data));
if (!msg.id) return handleEvent(msg); // Page.loadEventFired, Log.entryAdded, ...
const p = pending.get(msg.id);
if (!p) return;
pending.delete(msg.id);
msg.error ? p.reject(new Error(JSON.stringify(msg.error))) : p.resolve(msg.result);
};
ws.onclose = () => {
for (const p of pending.values()) p.reject(new Error("socket closed"));
pending.clear();
};
function send(method: string, params: object = {}): Promise<any> {
return new Promise((resolve, reject) => {
const id = ++nextId;
pending.set(id, { resolve, reject });
ws.send(JSON.stringify({ id, method, params }));
});
}
await send("Page.enable");
await send("Runtime.enable");Navigations destroy execution contexts and can invalidate in-flight
Runtime.evaluate calls; retry after observing the new document.
Start with a concise UI inventory:
const result = await send("Runtime.evaluate", {
expression: `JSON.stringify({
buttons: [...document.querySelectorAll('button')].map((el) => ({
text: el.innerText.trim(), aria: el.getAttribute('aria-label'), title: el.title
})).filter((x) => x.text || x.aria || x.title),
inputs: [...document.querySelectorAll('input, textarea, [contenteditable=true]')].map((el) => ({
tag: el.tagName, type: el.type, placeholder: el.placeholder,
aria: el.getAttribute('aria-label'), value: el.value
}))
})`,
returnByValue: true,
});Use awaitPromise: true for async expressions and userGesture: true when the
page requires user activation. Treat exceptionDetails in the result as an
error even though the CDP command itself succeeded.
document.querySelector does not cross shadow boundaries — traverse open
shadow roots explicitly; for closed shadow roots or remote-object work use the
DOM domain (DOM.getDocument, DOM.querySelector, DOM.getBoxModel).
Compute coordinates in CSS pixels immediately before acting, then send native input events:
async function point(expr: string) {
const r = await send("Runtime.evaluate", {
expression: `(() => {
const el = ${expr};
if (!el) return null;
el.scrollIntoView({ block: 'center', inline: 'center' });
const b = el.getBoundingClientRect();
return { x: b.left + b.width / 2, y: b.top + b.height / 2 };
})()`,
returnByValue: true,
});
if (!r.result.value) throw new Error(`Element not found: ${expr}`);
return r.result.value;
}
async function click(expr: string) {
const { x, y } = await point(expr);
await send("Input.dispatchMouseEvent", { type: "mouseMoved", x, y });
await send("Input.dispatchMouseEvent", { type: "mousePressed", x, y, button: "left", buttons: 1, clickCount: 1 });
await send("Input.dispatchMouseEvent", { type: "mouseReleased", x, y, button: "left", buttons: 0, clickCount: 1 });
}Focus an editable element (click it), then insert text:
await click(`document.querySelector('input[aria-label="Search"]')`);
await send("Input.insertText", { text: "search terms" });Use Input.insertText for text and Unicode; use paired Input.dispatchKeyEvent
(keyDown + keyUp with key, code, windowsVirtualKeyCode) for Enter,
Escape, arrows, Tab, and shortcuts. Modifier bits: Alt=1, Ctrl=2, Meta=4, Shift=8.
Native <select> and framework-controlled inputs may need the prototype setter
plus bubbling events:
const setter = Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value").set;
setter.call(input, "new value");
input.dispatchEvent(new Event("input", { bubbles: true }));
input.dispatchEvent(new Event("change", { bubbles: true }));Prefer real Input.* events for the behavior being demonstrated or tested;
direct DOM mutation is fine for deterministic setup and inspection.
A returned command does not mean the action completed. Poll the state that proves completion:
async function waitFor(expr: string, timeoutMs = 30_000) {
const start = Date.now();
while (Date.now() - start < timeoutMs) {
const r = await send("Runtime.evaluate", { expression: `Boolean(${expr})`, returnByValue: true });
if (r.result.value) return;
await new Promise((res) => setTimeout(res, 250));
}
throw new Error(`Timed out waiting for ${expr}`);
}
await waitFor(`document.body.innerText.includes('Saved')`);For navigation, wait on Page.loadEventFired or a lifecycle networkIdle
event — but SPA route changes may emit neither, so prefer the UI condition
that actually matters.
const shot = await send("Page.captureScreenshot", { format: "png" });
await Bun.write("screenshot.png", Buffer.from(shot.data, "base64"));captureBeyondViewport: true for full-page; Page.getLayoutMetrics + clip
for exact regions; Page.printToPDF for PDFs. For video recording with
Page.startScreencast and demo-polish tips, read
references/recording.md.
Enable Network and Log, then watch Network.requestWillBeSent,
Network.responseReceived, Network.loadingFailed, Runtime.consoleAPICalled,
Runtime.exceptionThrown, and Log.entryAdded. Fetch bodies with
Network.getResponseBody. Never log authorization headers, cookies, API keys,
passwords, or response bodies containing secrets — scrub before returning
output to context.
Common failure modes:
/json/version first./json/list; match by URL/title, don't take
the first page blindly.Input.insertText,
or the prototype-setter pattern above.Browser automation acts with the user's browser authority.
127.0.0.1) unless the user
explicitly needs remote access and has authentication in place.© letta-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in src/skills/builtin/browser-use of letta-ai/letta-code.
Open the folder on GitHubat commit 253a3bc
Browser Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser Use this skillletta-ai/letta-code | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| C Browserdaxaur/openpaw | 174 | — | ~383 | Automated safety check: Pass | MIT | |
| Browser Automationaiskillstore/marketplace | 430 | — | ~1.5k | Automated safety check: Notes | None | |
| Browserwingbrowserwing/browserwing | 1.4k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Actionbookactionbook/actionbook | 1.6k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Gsd Browsergsd-build/gsd-browser | 268 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
daxaur/openpaw
Headless browser automation — navigate pages, click elements, fill forms, take screenshots, scrape content using agent-browser or playwright-cli.
aiskillstore/marketplace
Enterprise-grade browser automation using WebDriver protocol.
browserwing/browserwing
Browser automation platform with 78 built-in scripts and full CLI.
actionbook/actionbook
Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.
gsd-build/gsd-browser
Native Rust browser automation CLI for AI agents. An agent skill from gsd-build/gsd-browser.
justrach/kuri
Use kuri-server to automate Chrome via HTTP API — navigate pages, get a11y snapshots, interact with elements, capture network traffic (HAR), extract cookies, and bypass bot protection.
letta-ai/letta-code
Guide for creating effective skills. An agent skill from letta-ai/letta-code.
letta-ai/letta-code
Generates and reviews mod learning env JSON files for Letta Code local mods.
letta-ai/letta-code
Comprehensive guide for initializing or reorganizing agent memory.
letta-ai/letta-code
Inspect or modify Letta Code's own memory, model, context window, system prompt, compaction, permissions, toolsets, mods, skills, channels, schedules, agent secrets, and local runtime settings.
letta-ai/letta-code
Creates and edits trusted local Letta Code mods, including tools, slash commands, local-only model providers, lifecycle/turn events, scoped conversation helpers, panels, and capability-gated behavior.
letta-ai/letta-code
Creates, edits, and migrates Letta Code statusline mods. An agent skill from letta-ai/letta-code.
Categories
Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video. Browser Use is an agent skill from letta-ai/letta-code. Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.
Browser Use fits situations like: automate a browser; test rendered page UI; scrape a site that needs browser execution; capture a browser screenshot.
Run `npx skills add letta-ai/letta-code --skill browser-use -a claude-code`. Or copy the skill folder (src/skills/builtin/browser-use in letta-ai/letta-code) into .claude/skills/browser-use in your project. Claude Code loads it when a task matches its description.
Run `npx skills add letta-ai/letta-code --skill browser-use -a codex`. Or copy the skill folder (src/skills/builtin/browser-use in letta-ai/letta-code) into .agents/skills/browser-use in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add letta-ai/letta-code --skill browser-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-use, .gemini/skills/browser-use, .github/skills/browser-use and .opencode/skills/browser-use in your project.
SKILL.md names no scripts, command-line tools or credentials: Browser Use is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: chromedevtools.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Browser Use is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 494 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Browser Use: C Browser (daxaur/openpaw, 174 stars), Browser Automation (aiskillstore/marketplace, 430 stars), Browserwing (browserwing/browserwing, 1.4k stars) and Actionbook (actionbook/actionbook, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
letta-ai (a GitHub organization) maintains it in letta-ai/letta-code, which has 3,552 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.
Source: letta-ai/letta-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.