AI Search Hub
minsight-ai-info/AI-Search-Hub
Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.
Drive the user's existing Chromium-family browser with deterministic Playwright.
$ npx skills add anomalyco/browser-control --skill browser-control -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install anomalyco/browser-control browser-control --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser-control .claude/skills/browser-control && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .claude/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/anomalyco/browser-control/tree/main/skills/browser-controlType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add anomalyco/browser-control --skill browser-control -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install anomalyco/browser-control browser-control --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/browser-control .agents/skills/browser-control && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .agents/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anomalyco/browser-control --skill browser-control -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install anomalyco/browser-control browser-control --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/browser-control .cursor/skills/browser-control && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .cursor/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/anomalyco/browser-control.git --path skills/browser-control--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add anomalyco/browser-control --skill browser-control -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install anomalyco/browser-control browser-control --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/browser-control .gemini/skills/browser-control && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .gemini/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install anomalyco/browser-control browser-controlInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add anomalyco/browser-control --skill browser-control -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/browser-control .github/skills/browser-control && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .github/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anomalyco/browser-control --skill browser-control -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install anomalyco/browser-control browser-control --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anomalyco/browser-control.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/browser-control .opencode/skills/browser-control && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser-control" agent skill from https://github.com/anomalyco/browser-control/tree/main/skills/browser-control into .opencode/skills/browser-control/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser-control", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser-controlDrive the user's existing Chromium-family browser with deterministic Playwright.
Browser Control is an agent skill from anomalyco/browser-control. Drive the user's existing Chromium-family browser with deterministic Playwright. Use when asked to inspect, automate, test, or interact with a visible browser tab; continue an authenticated browser workflow; handle 2FA, passkeys, CAPTCHAs, or payment confirmation; record browser behavior; or capture an authenticated network flow.
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Productivity & Automation, covering Browser automation and Browser testing. It works with Playwright. The repository describes itself as: Local browser driver for trusted agents: control your existing Chromium browser through a small extension and local relay. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 22c0f77. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ubereats.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser Control loads about 8.2k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 3,830 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from anomalyco/browser-control at commit 22c0f77, republished under its MIT licence (© anomalyco). 3,830 words, ~8,211 tokens.
.claude/skills/browser-control/SKILL.md (or your agent's skills folder).Browser Control is a driver, not an agent. The calling agent decides what to do; Browser Control runs deterministic Playwright code in the user's visible browser.
Use one loop throughout: inspect, act, verify. Inspect the real page before choosing locators, act through the narrowest stable control, then verify the result through a URL or fresh page read. Never treat a successful click or human acknowledgment as proof that the task succeeded.
Start with the requested browser work. Relay-backed commands start the detached
relay and wait for the extension; do not start browser-control serve first.
browser-control execute 'return { url: page.url(), title: await page.title() }'Use browser-control doctor only when setup or runtime behavior is unclear.
status and doctor are observational and never start the relay.
MCP startup, tool discovery, skill, and session_current do not contact the
relay. The first operational tool call starts it if needed; relay-backed
observational tools report unavailability instead of starting it.
For an externally supervised relay, set BROWSER_CONTROL_AUTOSTART=false to
make ordinary calls fail when that relay is absent instead of launching another
process. Existing relay connections and explicit relay restart still work.
Ordinary CLI/MCP/SDK calls never replace a running relay. On a build mismatch,
coordinate with other agents before running browser-control relay restart.
It preserves browser tabs and durable sessions but resets JavaScript state and
snapshot refs. A busy or timed-out drain leaves the old relay running; finish
recordings/captures and disconnect raw CDP clients rather than forcing a stop.
browser-control doctor
browser-control status --jsonCompletion: one execute returns a page result and a readable session id, or
doctor identifies the concrete setup failure.
The bundled shim 0.0.25 verifies debugger ownership during reconnection. Reload
the unpacked extension after installing this shim update. DevTools-attached tabs
are excluded from Browser Control's inventory. page.title() reads time out
after five seconds if the page execution context remains unavailable; the read
timeout does not close or replace the tab.
A bare CLI execute creates a fresh session-owned page and prints the exact
--session <id> continuation command, or you can pass a descriptive session
name (including an optional emoji, which appears directly on the browser tab
group and in-page status pill). Every later CLI call must pass that id or set
BROWSER_CONTROL_SESSION; bare execute never guesses from human-shell current
state.
browser-control session new "🎙️ elevenlabs"
browser-control execute --session "🎙️ elevenlabs" 'return page.url()'
browser-control execute 'return page.url()'MCP keeps one implicit process session. Omit session for that normal path, or
call session_new and pass an explicit id when one MCP process needs multiple
sessions.
To control a tab already open in the user's browser, ask the user to click the
Browser Control toolbar button on that tab. Select it for one execute or adopt
it for sticky reuse (omit --target-url / targetUrl when only one user tab is
attached):
browser-control session adopt --session github
browser-control session adopt --target-url github.com --session github
browser-control execute --target-url github.com 'return page.url()'execute --target-url selects a page for that call only. Continuing with just
--session uses the session's default page, which may still be about:blank.
For a multi-step task in an existing user tab, adopt it first. Always include
page.url() when diagnosing an empty snapshot. Snapshot labels are compact
descriptions; use ref() for actions rather than assuming their text is an
exact Playwright accessible name.
targetUrl and targetIndex select existing attached pages; they never
navigate. A URL selector must match exactly one page, and URL and index selectors
cannot be combined. Adoption makes that tab the session default, closes the
session's previous relay-created page, and is exclusive to one Browser Control
session. Reset or delete releases an adopted user tab without closing it.
Adoption binds ownership and exact target identity before initializing automation.
Until the first execute resolves the page, session status reports connected: false and pageUrl: null; adoptedUrl is the registry-selected URL, not a fresh
page read. A busy page can produce session-page/adopted-initialization-timeout
on execute: user code did not run, and the exact adopted target is retained.
Retry after the page settles; do not reset or adopt a different tab to recover.
Prefer adoption for authenticated browser state rather than reproducing login in a fresh page.
Each relay controls one browser/profile at a time. A second extension connection
cannot replace a healthy active connection. If status or doctor reports
rejected competing connections, keep the extension enabled only in the intended
browser/profile. To switch browsers, disconnect the incumbent extension first;
creating a new execute session does not switch browsers.
Completion: the selected page URL is the intended page, and later work either retains the returned session id or intentionally uses the MCP process session.
Inspect before guessing roles or selectors:
return await snapshot()Then act from the returned structure and verify the destination:
await ref("e12").click()
await page.getByRole("heading", { name: "Settings" }).waitFor()
if (!page.url().includes("/settings")) {
throw new Error(`Unexpected destination: ${page.url()}`)
}
return { url: page.url(), heading: await page.getByRole("heading").first().innerText() }Use normal Playwright first. Keep dependent interactions in one execute when they rely on transient UI such as an open menu, selected rows, hover state, or an in-progress form.
If native locator.fill() hangs, inspect the failure before using the explicit
input, textarea, or contenteditable fallback. A target/cross-extension-page
diagnostic requires human dismissal or completion first; fillInput is not a
way around protected extension UI:
await fillInput(page.getByPlaceholder("Username"), "standard_user")Completion: the final return value contains evidence of the requested outcome, not merely evidence that an action was attempted.
A click returning, HTTP 200, a redirect, or cleared fields do not prove delivery. Register any expected navigation before clicking once, then read the destination and assert the workflow's specific receipt or confirmation. Coordinate rejection and confirmation checks with the actual page you inspected; the driver cannot infer business success from generic page text.
await Promise.all([
page.waitForURL(expectedDestination, { waitUntil: "domcontentloaded", timeout: 15_000 }),
ref(sendRef).click({ timeout: 15_000 }),
])
const observed = await snapshot()
await expectedReceipt.waitFor({ state: "visible", timeout: 5_000 })
return { url: page.url(), observed, receipt: await expectedReceipt.innerText() }expectedDestination, sendRef, and expectedReceipt come from the inspected
workflow, not guessed selectors. For same-URL navigation, register a main-frame
navigation or response wait instead. A mouse coordinate click need not wait for
navigation; an immediate snapshot can still be the old document.
If submission timed out or the destination cannot be read, report unverified and inspect once without clicking again. Re-read only after the exact tab/context settles. Never replay an uncertain send, booking, purchase, or payment. An explicit site rejection means rejected, not delivered; protection requiring human interaction means hand off the ordinary page to the user. Do not modify protection tokens, spoof human signals, or dispatch DOM clicks to evade a security boundary.
Named sessions preserve their default page across short-lived CLI and MCP
processes. They also survive relay restarts: Browser Control restores the id,
read-only mode, and exact default target. JavaScript state and snapshot refs
are process-local and reset after a relay restart with an explicit warning.
browser-control session list
browser-control session reset github
browser-control session delete githubDeletion is idempotent for an explicit session id, so cleanup can be safely retried when that session is already absent.
Every execute is journaled under
~/.browser-control/sessions/<id>/journal.jsonl. The journal records code,
status, duration, URL movement, warnings, handoffs, and bounded diagnostics.
Never place credentials directly in execute source.
browser-control journal --session github --limit 50Completion: retain the session only when follow-up work is expected; otherwise reset or delete session-owned pages and report any warnings that affect later work.
The distinguishing Browser Control workflow is an authenticated tab plus a human-only prompt:
Attach and adopt the existing tab, inspect its real UI, fill ordinary fields,
then register handoff before triggering WebAuthn, 2FA, CAPTCHA, or payment UI.
After the user completes it, verify the authenticated destination. The same
session can continue after an MCP process or relay restart.
When the prompt-triggering action may itself block, put only that action in
start. Browser Control presents and acknowledges WAIT before invoking it:
await handoff("Complete the security-key prompt, then continue", {
timeoutMs: 600_000,
start: () => page
.getByRole("button", { name: /passkey|security key|sign in/i })
.click({ timeout: 600_000 }),
})
await page.waitForURL((url) => !url.pathname.startsWith("/login"))
await page.getByRole("heading", { name: /account|dashboard/i }).waitFor()
return { authenticatedUrl: page.url(), title: await page.title() }After a resolved handoff, Browser Control waits through transient destination context replacement before returning, so this verification can remain in the same execute.
With start, the handoff deadline and target cancellation remain active until
the action settles, even if the user has already pressed Continue. An early
acknowledgment does not authorize an indefinitely pending action.
For a handoff on another page, pass { page: otherPage }. Readiness checks that
page, not the session default. If a non-default page was replaced or closed,
inspect the remaining pages rather than assuming an old Playwright reference
now identifies its replacement.
Tell the user what action is waiting. Human acknowledgment is not verification:
always assert the expected URL or stable element after handoff. If the action
was already completed and only the human step remains, call handoff(message)
without start. The default timeout is ten minutes.
Completion: the prompt was presented only after WAIT was registered, the action settled, and the authenticated result was independently verified.
To turn a user-demonstrated flow into reusable Playwright, use demonstrate().
It uses the same exact-tab handoff, records clicks, edits, checkbox/select
changes, and same-tab navigations, then returns editable code:
return await demonstrate("Perform the workflow once, then continue")Password fields become explicit secret-source comments rather than copied values. Review selectors and add outcome assertions before reusing generated code; a demonstration records actions, not proof that the workflow succeeded.
Ordinary webpage fields and accessible open-shadow-root controls remain usable. 1Password's inline menus are extension-owned iframes, not ordinary webpage DOM. Chromium blocks one extension from debugging another extension's pages; toolbar popups and native unlock, Touch ID, and Windows Hello prompts are not supported Playwright control surfaces.
Focusing or filling a card-number or credential field can open the inline menu
by itself, even inside a third-party payment iframe. While it is open, Chrome
rejects every automation command for that tab; Playwright shows this as
"Execution context was destroyed" or a locator timeout, so the execute result
carries the target/cross-extension-page diagnostic and warning, and
browser-control status marks the tab protected-ui=true.
target/cross-extension-page means a permission boundary. Ask the user to finish
or dismiss the prompt and retry; do not reset the page, read vault contents, or
weaken browser security to get around it. The page itself is healthy: do not
treat the failure as an unresponsive tab or create a new page. Register handoff on the originating
webpage before triggering a human-only prompt when possible. If the prompt
already prevents attachment, give the user the required action directly rather
than assuming an in-page handoff can be displayed. Verify the intended webpage
state after the prompt is completed.
A page's leftover password-manager status text is not evidence that its menu is still open. Use current relay diagnostics and a fresh ordinary-page read after human dismissal, without inspecting the extension's private frame contents.
Use the least expensive view that answers the question:
snapshot() is the compact read-before-act default. It prioritizes semantic
groups, alerts, lists, tables, headings, links, and controls. Text input and
textarea values are omitted.main remain discoverable. Use an explicit within scope
when you intentionally need background content. Repeated list wrappers have
a bounded reservation so product links can still fit in a dense page.spinbutton/searchbox. Native disclosure
controls are labeled summary, which is an element kind rather than an ARIA
role; use the returned ref() to operate them.ref("e12") resolves a control from the latest snapshot. Refs fail closed
after navigation or incompatible DOM drift. Compatible refs keep the same id
across repeated same-document captures.snapshot({ diff: true }) reports semantic changes from the compatible prior
baseline. snapshot({ delta: true }) returns a full first baseline and compact
deltas afterward. Existing compatible refs remain usable.snapshot({ find: "checkout", context: 2 }) searches the bounded semantic
snapshot and returns matching lines with nearby context and actionable refs.ariaSnapshot(target?, { timeout }) returns Playwright's detailed YAML aria
tree when the compact snapshot omits needed structure. Native text-control
values, custom ARIA range values, and editable content are omitted so they do
not enter tool output. Await it separately; do not run other operations on
the same page concurrently.screenshotWithLabels({ page?, path? }) adds visual labels and metadata when
layout matters, and registers its e1..eN labels for ref().screenshotDiff({ baseline, path?, threshold?, fullPage? }) compares a saved
PNG (absolute path or Buffer) with the current session page at CSS-pixel scale.
It returns matches, changedPixels, changedRatio (0..1), dimensions, and a
red-highlighted PNG. Omit path to return the image as execute media; otherwise
supply a fresh absolute .png path. Existing output files are never overwritten.return await snapshot({ within: "main", maxItems: 200 })
return await snapshot({ find: /checkout|payment/i, context: 2 })
return await snapshot({ delta: true })
// When layout matters, return the image through MCP so it can be inspected.
return await screenshotWithLabels({ page })Saving an image and returning only "ok" proves file creation, not visual
correctness. Return screenshot buffers through MCP when visual evidence matters.
For visual regression checks, save a baseline before changing the UI:
await page.screenshot({ path: "/absolute/before.png", scale: "css" })
// After the intended UI change, in the same viewport:
return await screenshotDiff({ baseline: "/absolute/before.png" })The threshold defaults to 0.1 and controls per-pixel color tolerance, not the
allowed changed area. Set it to 0 for exact pixels. Antialiasing changes count.
Settle animations yourself and use the same viewport and fullPage setting for
both captures. Dimension mismatches fail explicitly; images are never resized.
Comparisons are limited to PNGs of 32 MiB / 16 megapixels each. Screenshots and
diffs include visible page content: inspect for private information before sharing.
Execute code can use page, context, browser, persistent state, selected
Node modules through modules and aliases such as fs and path, plus the
Browser Control helpers documented here. Execute code runs in Node. Use
page.evaluate for window, document, storage, and same-origin fetch
with page cookies. Single expressions auto-return;
multi-statement scripts need return. Use --file for longer scripts:
browser-control execute --session github --file ./perform-flow.jsHuman CLI output includes logs, warnings, and a concise aftermath. Use --json
when another command needs to branch on ok, value, error, warnings, or
aftermath:
browser-control execute --json --session github '({ url: page.url() })' | jq .value.urlPlaywright downloads are unavailable through extension-backed tabs because
Chromium blocks download artifact control through chrome.debugger. If the
page exposes the payload through fetch or an API response, read the bytes in the
page and write them with fs. Do not retry page.waitForEvent("download").
Pages with WebMCP enabled can expose structured page tools. Discover and call them through the execute helper; names, descriptions, schemas, and results come from the page:
const tools = await webmcp.list()
const result = await webmcp.call("search_catalog", { query: "adapter" })
return { tools, result }Discovery covers all frames. If the same name appears in multiple frames, pass
the exact reported frame label as { frame }. Browser Control re-discovers the
tool immediately before invoking it, so stale registrations fail directly.
Browser Control blocks CDP commands that would destroy shared browser state, including browser close and cookie/cache clearing. Never work around those guardrails.
For inspect-only work, use a read-only session:
browser-control session new inspect-prod --read-only
browser-control execute --session inspect-prod 'await page.goto("https://example.com"); return page.title()'Read-only sessions reject Input.*, so they cannot click or type through
Playwright. page.evaluate can still mutate the DOM; read-only prevents trusted
mistakes, not malicious code.
For destructive UI work, use a two-phase read, confirm, verify flow:
Do not discover and confirm destructive candidates in one script unless the user already approved exact stable identifiers. Never globally auto-accept native dialogs; wait for the expected dialog and assert its type and message before accepting it.
Applications and CLI tools can use BrowserControlClient.origin() for direct
same-origin JSON requests authenticated by a session page (with optional default
headers and automatic in-page handoff recovery when a session expires or
redirects to login):
import { BrowserControlClient } from "@opencode-ai/browser-control"
const ubereats = BrowserControlClient.origin("https://www.ubereats.com", {
session: "🥤 karma-cafe",
startUrl: "/feed",
headers: { "x-csrf-token": "x" },
handoffOnAuthFailure: true,
})
const store = await ubereats.post("/_p/api/getStoreV1", {
storeUuid: "78cb1602-9f58-57bb-ad08-f6b8f80bb788",
diningMode: "DELIVERY",
})For Effect-native applications with schema decoding or sensitive: true
Redacted responses, use BrowserControlClient.Service:
import { BrowserControlClient } from "@opencode-ai/browser-control"
const sensitive = yield* origin.json({
path: "/api/session",
method: "POST",
body: {},
response: SessionResponse,
sensitive: true,
})
const session = BrowserControlClient.reveal(sensitive)Use network capture when the browser is needed to authenticate or discover a workflow, but repeated direct HTTP calls would be faster and more reliable. Capture each flow at least twice with different inputs so constants and parameters can be distinguished.
browser-control network start --session github --url /api/ \
--resource-type fetch --resource-type xhr
browser-control execute --session github --file ./perform-flow.js
browser-control network status --session github
browser-control network stop --session github \
--output ./github.har --secrets githubWritten artifacts replace credential-bearing headers, cookies, query fields,
OAuth fragment fields, and structured body fields with stable references such as ${BC_SECRET_1}.
--secrets github stores lossless values separately in a mode-0600 Secret
Profile. Never copy profile values into source, output, diagnostics, or journals,
and never deliberately return or log credentials.
Inspect the redacted artifact offline, generate one typed function per observed flow, then verify each function with a harmless live request. Run generated clients without exposing values:
browser-control secrets status github
browser-control secrets run github -- ./github-cli repositoriesGenerated TypeScript applications can own that wrapper internally through the public SDK:
import { SecretProfile } from "@opencode-ai/browser-control"
import { Effect } from "effect"
const result = await Effect.runPromise(SecretProfile.run({
name: "github",
command: process.execPath,
args: ["./github-cli.js", "repositories"],
}))The trusted worker receives BC_SECRET_N variables and its bounded output is
redacted. The public SDK exposes profile metadata but never raw profile values.
Refresh credentials normally renewed by a page reload with:
browser-control secrets refresh github --session github --url /api/If refresh requires login or a human prompt, reauthenticate in the adopted tab
and repeat capture with the same profile. MCP exposes equivalent network_*
and secrets_* tools.
Completion: the artifact contains references rather than credential values, the generated operation passes a harmless live check, and no secret value appears in source or output.
Record an attached or session-owned tab with:
browser-control recording start ./tmp/demo.mp4 --session github --mode cdp
browser-control recording status --session github
browser-control recording stop --session githubStart, stop, and status accept --json. CDP stop/status results include a
quality receipt: output dimensions/rate, source and retained-image counts/rates,
coalesced/dropped frames, and screenshotFallback. The fallback means no
compositor frames arrived and the video holds one stop-time screenshot; do not
present that as recorded motion. Source counters count compositor events, not
visually distinct frames. Low rates can be normal on a static page. Tab capture
and older relays omit quality telemetry rather than inventing measurements.
Explicit frame rates must be integers from 1 through 60 in either mode; invalid values fail instead of silently clamping. The start result reports the chosen rate.
--mode auto uses tab capture for user-owned tabs and CDP for relay-owned tabs.
Tab capture can include audio; CDP requires ffmpeg and has no audio. Use the
command's --help for format and cursor options.
Playwright mouse actions automatically reveal the on-page Ghost Cursor
(distance-glide motion + tactile-bloom click shockwave). When recording a
user-facing proof or PR walkthrough video, opt into the ghostCursor helpers
inside execute to focus attention on key steps and verified postconditions:
await showGhostCursor()
await ghostCursor.caption("Verify cluster & promote release", { step: "01", tone: "neutral" })
await ghostCursor.zoom("#release-card", { scale: 1.45 })
await page.locator("#promote-btn").click()
await ghostCursor.keys("⌘+⇧+P", "Promote Release")
await ghostCursor.resetZoom()
await ghostCursor.spotlight("#status-badge", { label: "Verified", detail: "200 OK", tone: "success" })CDP recordings preserve the starting CSS viewport (not a fixed 720p canvas),
use high-quality source frames, and default to 60 fps. Use --frame-rate 30
for smaller files. Actual motion still depends on Chrome delivering new frames;
60 fps output does not guarantee 60 distinct frames. Larger viewports cost more
CPU, transport bandwidth, and storage. Set the viewport before recording and do
not change viewport/emulation mid-recording. Odd dimensions round down to even.
For failures that are hard to reproduce, keep a rolling CDP frame buffer and save recent history after the problem occurs:
browser-control flight-recorder start --session github --retention-ms 60000
browser-control flight-recorder status --session github
browser-control flight-recorder save-last ./tmp/failure.mp4 --session github --duration-ms 30000
browser-control flight-recorder cancel --session githubSaving does not stop buffering. The recorder is memory-bounded, reports retained duration/frame/byte/drop counters, writes a JSON receipt beside each clip, and is mutually exclusive with ordinary recording on the same tab. CLI recording and flight-recorder lifecycle operations are also available as MCP tools.
Inspect an encoded frame at native size before sharing: the whole viewport must fill the frame, small text must be readable, and motion must not be a repeated still image. Do not crop and upscale a low-resolution capture to call it HD. On an older installed relay that shrinks the page into a padded corner, record the defect and coordinate a recorder update; changing the file's resolution is not a repair.
Completion: stop the recorder, inspect the resulting media rather than only its existence, and report the viewport, state, and interaction path actually tested.
browser-control doctor; it checks package metadata, CLI/relay build
identity, extension protocol compatibility, sessions, targets, and artifacts.status --json to inspect exact sessions and target ownership.Common diagnoses:
session-page/context-read-timeout; operation=page.title; timeoutMs=5000:
the title read did not finish within its budget. The execution context may be
unavailable or busy; this does not prove a frozen renderer. The watchdog does
not cancel the underlying read or trigger page replacement. A cached
page.url() read can still work; retry the page read after the page settles.
Ordinary missing-locator timeouts do not receive this diagnostic.connected:false: run a relay-backed command and allow the extension startup
or alarm wake-up to reconnect. A sleeping extension wakes on a 30-second
alarm, so the command waits up to 35 seconds. Reload the unpacked extension
only if that loop does not recover.doctor, then coordinate an explicit
browser-control relay restart. It requires an exact managed instance and
safe shutdown protocol 2. Legacy relays need a one-time coordinated manual
stop; foreground/source or newer relays are never force-killed or downgraded.
MCP observational tools remain available on a mismatch.~/.browser-control/relays/<port>/lifecycle.jsonl for requester/build/instance
metadata. Preparing or selecting a development candidate must not restart the
daemon. Use isolated runtime:prepare / runtime:select, not a live checkout
link, when developing Browser Control itself.Target not found: attach the intended tab, then select or adopt it using a
unique URL substring or explicit index.state
and snapshot refs reset. Continue after the warning.session-page/*-unresponsive
diagnosis if the page still does not answer. Blank or unknown URLs are preserved:
they can contain unsaved content. Only a crashed or
chrome-error:// relay-owned page is closed and recreated. It never replaces
an adopted user tab. Main-frame navigation clears an earlier crash diagnosis;
child-frame navigation does not. When a page stays unresponsive (bot-protected sites can
stall the main world for automation while rendering normally for the human),
open a fresh tab with context.newPage() or hand the tab to the user.fillInput after
confirming the selector or locator resolves. String selectors search open
shadow roots recursively; closed shadow roots remain unavailable.await locator.evaluate((element) => element.click()),
then verify the result. Do not make dispatched clicks the default.isChecked(). Do not force-click an invisible input or infer an
arbitrary ancestor.fs; extension-backed Playwright cannot
retain a native download artifact.getBoundingClientRect() and use
page.mouse.move() when coordinate input is appropriate.chrome.debugger. Do not retry or bypass that boundary; open the page
for manual operation.For deeper relay diagnosis, restart with BROWSER_CONTROL_DEBUG=1. Debug traces
must never include expressions, arguments, results, headers, cookies, or form
values.
Whenever Browser Control fails, wedges, replaces a page/session, or behaves
unexpectedly, create or update a browser-control project todo with the Browser
Control version, safe session/page context, exact error, deterministic
reproduction, expected versus actual behavior, and recovery attempted. Never
include credentials, form values, or private account data.
© anomalyco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/browser-control of anomalyco/browser-control.
Open the folder on GitHubat commit 22c0f77
Browser Control next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser Control this skillanomalyco/browser-control | 438 | — | ~8.2k | Automated safety check: Pass | MIT | |
| AI Search Hubminsight-ai-info/AI-Search-Hub | 1.3k | — | ~1.3k | Automated safety check: Pass | None | |
| Ego Browserkwakseongjae/oh-my-design | 531 | 1 repos | ~4.9k | Automated safety check: Pass | MIT | |
| Playwright Bowserdisler/bowser | 265 | — | ~1.1k | Automated safety check: Notes | None | |
| Vrboborski/travel-hacking-toolkit | 689 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Sap Browser Automationsecondsky/sap-skills | 460 | — | ~3.3k | Automated safety check: Pass | GPL-3.0 |
minsight-ai-info/AI-Search-Hub
Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.
kwakseongjae/oh-my-design
When you need a browser, read this Skill by default. An agent skill from kwakseongjae/oh-my-design.
disler/bowser
Headless browser automation using Playwright CLI. An agent skill from disler/bowser.
borski/travel-hacking-toolkit
Search VRBO (Vrbo / Expedia Group) vacation rentals including entire homes, condos, and cabins via Patchright browser automation.
secondsky/sap-skills
A skill your agent uses when an agent must inspect or operate an authenticated SAP web UI through an in-app Browser, Microsoft Edge CDP, or an existing Playwright client, especially when SAP SSO…
ZhixiangLuo/10xProductivity
Walk through a UI flow once manually and capture a durable interaction map — which DOM elements to click, which network requests they trigger, and what field shapes they expose.
anomalyco/browser-control
Run a real-world navigation and capability gauntlet on Browser Control across live websites, fix general DOM/ARIA/CDP root causes, benchmark speed and token efficiency, and finish with a mandatory…
anomalyco/browser-control
Dogfood a project-owned tool in a realistic user flow and evaluate the agent experience.
Works with
Categories
Drive the user's existing Chromium-family browser with deterministic Playwright. Browser Control is an agent skill from anomalyco/browser-control. Drive the user's existing Chromium-family browser with deterministic Playwright.
Browser Control fits situations like: asked to inspect; interact with a visible browser tab; continue an authenticated browser workflow; payment confirmation.
Run `npx skills add anomalyco/browser-control --skill browser-control -a claude-code`. Or copy the skill folder (skills/browser-control in anomalyco/browser-control) into .claude/skills/browser-control in your project. Claude Code loads it when a task matches its description.
Run `npx skills add anomalyco/browser-control --skill browser-control -a codex`. Or copy the skill folder (skills/browser-control in anomalyco/browser-control) into .agents/skills/browser-control in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anomalyco/browser-control --skill browser-control -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-control, .gemini/skills/browser-control, .github/skills/browser-control and .opencode/skills/browser-control in your project.
Going by SKILL.md and its folder, Browser Control needs the command-line tools its instructions call (jq).
SKILL.md names 1 domain. In commands or code: ubereats.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Browser Control is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browser Control: AI Search Hub (minsight-ai-info/AI-Search-Hub, 1.3k stars), Ego Browser (kwakseongjae/oh-my-design, 531 stars), Playwright Bowser (disler/bowser, 265 stars) and Vrbo (borski/travel-hacking-toolkit, 689 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
anomalyco (a GitHub organization) maintains it in anomalyco/browser-control, which has 438 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: anomalyco/browser-control on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.