Agent skill

Browser Ops Skill

by OpenLoaf in OpenLoaf/OpenLoaf

Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or…

AGPL-3.0Auto-check passedDocuments & Office

Install Browser Ops Skill

skills CLI
$ npx skills add OpenLoaf/OpenLoaf --skill browser-ops-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenLoaf/OpenLoaf browser-ops-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenLoaf/OpenLoaf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/apps/server/src/ai/builtin-skills/browser-ops/en .claude/skills/browser-ops-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-ops-skill
GitHub stars
107
Token cost
~1.5k tokens
SKILL.md length
664 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or…

  • Works in 5 steps: OpenUrl → BrowserSnapshot to inspect… → For each field: BrowserAct { action:… → Submit: BrowserAct { action:… → …
  • Asks for page-level interaction with a specific webpage: login
  • SKILL.md covers Tool Inventory, Core Mental Model, OpenUrl Opening Modes and BrowserSnapshot Details, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Browser Ops Skill is an agent skill from OpenLoaf/OpenLoaf. Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or anti-bot measures, accessing SPA dynamic content; also used as a fallback when WebFetch returns an empty shell or gets blocked. Typical phrasing: "log me into X then scrape Y", "take a screenshot". Not for: factual questions or "what is xx" (→ WebSearch), or simply reading static webpage text (try WebFetch first).

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Web scraping and Forms and invoices. The repository describes itself as: 🍞Open-source, local-first AI workspace with Agents, multi-model chat (GPT/Claude/Gemini/DeepSeek), Notion-like docs, AI image & video generation, email, calendar & terminal…. The licence is AGPL-3.0.

When your agent uses it

  • Asks for page-level interaction with a specific webpage: login
  • Pagination scraping
  • Downloading page images
  • Handling CAPTCHAs

Example prompts

  • “log me into X then scrape Y”
  • “take a screenshot”
  • “what is xx”
  • “/browser-ops-skill”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. OpenUrl → BrowserSnapshot to inspect form structure and selectors
  2. For each field: BrowserAct { action: "fill", selector: "...", text: "..." }
  3. Submit: BrowserAct { action: "click-css", selector: "button[type=submit]" }
  4. BrowserWait { type: "urlIncludes", url: "/success" } to confirm success
  5. BrowserSnapshot for final confirmation

What it can do on your machine

Read from SKILL.md and the folder at commit f7eccf6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Ops Skill loads about 1.5k tokens when it runs. Until then it costs about 135 tokens; SKILL.md has 664 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenLoaf/OpenLoaf at commit f7eccf6, republished under its AGPL-3.0 licence (© OpenLoaf). 664 words, ~1,487 tokens.

Download SKILL.mdSave it as .claude/skills/browser-ops-skill/SKILL.md (or your agent's skills folder).
name
browser-ops-skill
description
Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or anti-bot measures, accessing SPA dynamic content; also used as a fallback when WebFetch returns an empty shell or gets blocked. Typical phrasing: "log me into X then scrape Y", "take a screenshot". **Not for**: factual questions or "what is xx" (→ `WebSearch`), or simply reading static webpage text (try `WebFetch` first).

Browser Operations Guide

Tool Inventory

ToolPurposeRead-only
OpenUrlOpen a webpage (headless / tab / window — three modes)Yes
BrowserSnapshotGet page text + interactive element list + screenshot + rawHtml in one callYes
BrowserActPage interaction (click / input / scroll / keypress)No
BrowserWaitWait for page state (load / networkidle / url / text / timeout)Yes
BrowserDownloadImageDownload an image from the page to local diskNo

Loading: All are deferred tools — call ToolSearch(names: "OpenUrl,BrowserSnapshot,BrowserAct,BrowserWait,BrowserDownloadImage") to activate their schemas before invoking.

Core Mental Model

Browser operations = observe-act-verify loop. After every action you must snapshot to confirm state, because webpages are stateful — a click may trigger navigation, a popup, or an AJAX load, and you cannot predict the outcome. Blindly chaining actions is the single most common failure mode.

OpenUrl → BrowserSnapshot → analyze → BrowserAct → BrowserWait → BrowserSnapshot → ...

OpenUrl Opening Modes

  • headless — First choice for pure automation. No UI; suited for scraping, background agents, and batch operations.
  • tab (default) — Use when the user needs to see/operate the page. Embedded panel.
  • window — Use when the user needs a standalone window for deep interaction.

Rule of thumb: user doesn't need to see it → headless; user needs to see it → tab; user wants a standalone window → window.

BrowserSnapshot Details

A single BrowserSnapshot call returns:

  • Page info: URL, title, readyState
  • Full text: body.innerText (truncated at 32KB)
  • Interactive element list: up to 120 clickable/inputtable elements with their selectors
  • iframe contents: text and elements from same-origin iframes
  • Screenshot: captures the full page by default (fullPage), saved to the session asset directory
  • rawHtmlPath: disk path to the complete outerHTML

What if text exceeds 32KB? Use Read/Grep on rawHtmlPath to fetch the full DOM — don't re-snapshot repeatedly.

Screenshot control: fullPage: false captures only the current viewport.

Three Core Workflows

Workflow 1: Information Extraction

OpenUrl → BrowserWait { type: "load" } → BrowserSnapshot

Extract the required information directly from the returned text.

Pagination scraping: BrowserSnapshot → BrowserAct { action: "click-text", text: "Next" } → BrowserWait { type: "networkidle" } → loop. Snapshot on every page.

Workflow 2: Form Filling and Login
  1. OpenUrl → BrowserSnapshot to inspect form structure and selectors
  2. For each field: BrowserAct { action: "fill", selector: "...", text: "..." }
  3. Submit: BrowserAct { action: "click-css", selector: "button[type=submit]" }
  4. BrowserWait { type: "urlIncludes", url: "/success" } to confirm success
  5. BrowserSnapshot for final confirmation

fill vs type: fill atomically clears and inputs — best for forms; type appends characters at the current focus — best for search boxes.

Login caveats: Cookies persist for the session, so after a single login subsequent requests carry auth automatically. Do not snapshot password fields.

Show full SKILL.md (266 more words)Show less
Workflow 3: Screenshots and Image Downloads
  • BrowserSnapshot — snapshot + full-page screenshot (fullPage by default)
  • BrowserSnapshot { fullPage: false } — viewport only
  • BrowserDownloadImage { selector: ".product-image" } — download an image from the page

Screenshots are a debugging superpower: when text doesn't reveal the issue, the screenshot often does at a glance.

Selector Selection Strategy

Pick selectors from the element list returned by BrowserSnapshot — don't guess on your own.

Priority: #id > [data-testid] > input[name] > .class > click-text

Wait Strategies

  • load: Use after traditional page navigation.
  • networkidle: First choice for SPA/AJAX pages — safest but slowest.
  • urlIncludes: Wait for redirect after form submission.
  • textIncludes: Wait for asynchronously loaded content.
  • timeout: Last resort.

Error Diagnosis

Element not found? → BrowserSnapshot to confirm page state → may have navigated elsewhere → may be below the viewport (scroll and retry) → may be inside an iframe (cross-iframe not supported)

Page load timeout? → BrowserSnapshot to see how far loading got; the content may already be sufficient.

Action has no effect? → Page not fully loaded (add BrowserWait) → selector matched the wrong element (check via snapshot) → a popup is blocking (close it first)

SPA content empty? → Snapshot again after BrowserWait { type: "networkidle" } → if still empty, wait on specific content with textIncludes

CAPTCHAs and Anti-Bot

When you encounter a CAPTCHA, 403/429, or an anti-bot page, stop immediately and notify the user — do not retry blindly.

Iron Rules

  1. BrowserSnapshot after every action to verify state
  2. Snapshot before acting — don't guess selectors
  3. Wait for the page to be ready before acting
  4. When text isn't enough, look at the screenshot
  5. Don't expose sensitive info (don't snapshot after filling a password)
  6. Stop immediately on anti-bot encounters

© OpenLoaf, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in apps/server/src/ai/builtin-skills/browser-ops/en of OpenLoaf/OpenLoaf.

Open the folder on GitHubat commit f7eccf6

Compare with similar skills

Browser Ops Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Ops Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Ops Skill this skillOpenLoaf/OpenLoaf107—~1.5kAutomated safety check: PassAGPL-3.0
Paginated Reportdata-goblin/power-bi-agentic-development1k—~3.5kAutomated safety check: PassGPL-3.0
Apify Marketing Agency Databaseapify/awesome-skills262—~2.5kAutomated safety check: PassMIT
Extract Document Datasickn33/agentic-awesome-skills47k1 repos~1.1kAutomated safety check: PassApache-2.0
Champion Trackergooseworks-ai/goose-skills1.2k1 repos~1.1kAutomated safety check: NotesMIT
Power DesignItsssssJack/power-design713—~2.8kAutomated safety check: PassMIT

Similar skills

  • Paginated Report

    data-goblin/power-bi-agentic-development

    Author, validate, publish, and test Power BI paginated reports in the RDL format.

    1k GitHub stars~3.5k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Official

    Build a marketing agency database from Clutch.co with the Clutch.co Agency API Actor (johnvc/clutch-agency-api).

    262 GitHub stars~2.5k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed
  • Extract Document Data

    sickn33/agentic-awesome-skills

    Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating.

    47k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check passed
  • Champion Tracker

    gooseworks-ai/goose-skills

    Track product champions for job changes and qualify their new companies against ICP.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check: notes
  • Power Design

    ItsssssJack/power-design

    Generate beautiful, on-brand HTML — presentation decks or full responsive websites — in any brand's design language, combining brand DNA extracted via Firecrawl with codified, research-backed design…

    713 GitHub stars~2.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Actionbook

    actionbook/actionbook

    Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.

    1.6k GitHub stars~1.5k tokensUpdated 29 days ago
    Productivity & AutomationAuto-check passed

More from OpenLoaf/OpenLoaf

All 34 skills in this repo
  • Agent Orchestration Skill

    OpenLoaf/OpenLoaf

    Triggers when the master Agent faces a multi-step complex task and is deciding whether / how to outsource sub-tasks to built-in subagents (browser / doc-editor / data-analyst / extractor /…

    107 GitHub stars~1.9k tokensUpdated 4 mo ago
    Auto-check passed
  • Canvas Ops Skill

    OpenLoaf/OpenLoaf

    Triggered when the user wants lifecycle management of OpenLoaf canvases / whiteboards: create, open, filter, duplicate, delete, rename, or change ownership.

    107 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Reads, edits, converts and reviews Word documents through three dedicated tools, covering tracked changes, comments, tables, images and format conversion.

    107 GitHub stars~1.9k tokensUpdated 4 mo ago
    Auto-check passed
  • Email Operations

    OpenLoaf/OpenLoaf

    Handles a real email account through query and mutate tools: check the inbox, read, search, reply, forward, compose and organize, with sending always confirmed first.

    107 GitHub stars~1.8k tokensUpdated 4 mo ago
    Auto-check passed
  • macOS Desktop Control

    OpenLoaf/OpenLoaf

    Guides an agent to operate native macOS apps by surveying an app first, acting through intents, menus or keystrokes, and verifying each step, in OpenLoaf Desktop only.

    107 GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed
  • PDF Skill

    OpenLoaf/OpenLoaf

    All-in-one PDF read / write / convert / OCR. An agent skill from OpenLoaf/OpenLoaf.

    107 GitHub stars~2k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Browser Ops Skill

What does Browser Ops Skill do?

Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or…. Browser Ops Skill is an agent skill from OpenLoaf/OpenLoaf. Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or anti-bot measures, accessing SPA dynamic content; also used as a fallback when WebFetch returns an empty shell or gets blocked.

When should I use Browser Ops Skill?

Browser Ops Skill fits situations like: asks for page-level interaction with a specific webpage: login; pagination scraping; downloading page images; handling CAPTCHAs.

How do I install Browser Ops Skill in Claude Code?

Run `npx skills add OpenLoaf/OpenLoaf --skill browser-ops-skill -a claude-code`. Or copy the skill folder (apps/server/src/ai/builtin-skills/browser-ops/en in OpenLoaf/OpenLoaf) into .claude/skills/browser-ops-skill in your project. Claude Code loads it when a task matches its description.

How do I install Browser Ops Skill in Codex?

Run `npx skills add OpenLoaf/OpenLoaf --skill browser-ops-skill -a codex`. Or copy the skill folder (apps/server/src/ai/builtin-skills/browser-ops/en in OpenLoaf/OpenLoaf) into .agents/skills/browser-ops-skill in your project. Codex loads it when a task matches its description.

Can I use Browser Ops Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenLoaf/OpenLoaf --skill browser-ops-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-ops-skill, .gemini/skills/browser-ops-skill, .github/skills/browser-ops-skill and .opencode/skills/browser-ops-skill in your project.

What does Browser Ops Skill need to run?

SKILL.md names no scripts, command-line tools or credentials: Browser Ops Skill is instructions for the agent only.

Does Browser Ops Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Browser Ops Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browser Ops Skill use?

Browser Ops Skill is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Ops Skill use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Ops Skill?

Skills that share tags, products or a category with Browser Ops Skill: Paginated Report (data-goblin/power-bi-agentic-development, 1k stars), Apify Marketing Agency Database (apify/awesome-skills, 262 stars), Extract Document Data (sickn33/agentic-awesome-skills, 47k stars) and Champion Tracker (gooseworks-ai/goose-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Ops Skill?

OpenLoaf (a GitHub organization) maintains it in OpenLoaf/OpenLoaf, which has 107 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on May 14, 2026.

Source: OpenLoaf/OpenLoaf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.