Agent skill

Browse

by SeifBenayed in SeifBenayed/cloclo

Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives.

MITAuto-check passedProductivity & Automation

Install Browse

skills CLI
$ npx skills add SeifBenayed/cloclo --skill browse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SeifBenayed/cloclo browse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SeifBenayed/cloclo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/browse .claude/skills/browse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browse
GitHub stars
114
Token cost
~2.7k tokens
SKILL.md length
822 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives.

  • Works in 3 steps: Navigate to the target URL → Observe the page with get_state — this… → Screenshot to see the visual layout
  • Asked to open a site
  • SKILL.md covers Core Principle, Starting a Session, Interacting with Elements and Keyboard Input (send_keys), plus 17 more sections
  • Reaches site-a.com and site-b.com

What it does

Browse is an agent skill from SeifBenayed/cloclo. Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives. Navigate sites, interact with elements, manage tabs/sessions, handle file uploads, dropdowns, iframes, and verify results using the screenshot-analyze-act-verify loop. Use when asked to "open a site", "test a page", "fill a form", "check a deployment", "browse to", "click on", "verify the UI", "compare pages side by side", or any task involving web interaction. Also use proactively when a task would…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering File uploads and storage and Browser automation. The repository describes itself as: Open-source Claude Code SDK — single-file CLIs in Node.js, Python, Go, Rust that use directly. The licence is MIT.

When your agent uses it

  • Asked to open a site
  • Check a deployment
  • Compare pages side by side
  • Any task involving web interaction

Example prompts

  • “open a site”
  • “test a page”
  • “fill a form”
  • “/browse”

Requirements

  • Pre-approved tools (allowed-tools): Browser, Read, Grep, Glob

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Navigate to the target URL
  2. Observe the page with get_state — this returns the DOM with indexed interactive elements AND the page text
  3. Screenshot to see the visual layout

What it can do on your machine

Read from SKILL.md and the folder at commit 00f5195. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Browser
    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • site-a.com
    • site-b.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browse loads about 2.7k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SeifBenayed/cloclo at commit 00f5195, republished under its MIT licence (© SeifBenayed). 822 words, ~2,734 tokens.

Download SKILL.mdSave it as .claude/skills/browse/SKILL.md (or your agent's skills folder).
name
browse
description
Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives. Navigate sites, interact with elements, manage tabs/sessions, handle file uploads, dropdowns, iframes, and verify results using the screenshot-analyze-act-verify loop. Use when asked to "open a site", "test a page", "fill a form", "check a deployment", "browse to", "click on", "verify the UI", "compare pages side by side", or any task involving web interaction. Also use proactively when a task would benefit from checking a live URL.
allowed-tools
Browser, Read, Grep, Glob

Browse — Visual Feedback Browser Automation

You have a Browser tool with an action-based interface. This skill teaches you how to use it effectively with a screenshot → analyze → act → verify loop that prevents blind clicking.

Core Principle

Never act without seeing. Never assume an action worked. Always verify.

OBSERVE  →  DECIDE  →  ACT  →  VERIFY
(get_state   (analyze    (click/    (get_state
+ screenshot  elements)   type)     + screenshot
                                    + compare)

Starting a Session

Every browser session starts the same way:

  1. Navigate to the target URL
  2. Observe the page with get_state — this returns the DOM with indexed interactive elements AND the page text
  3. Screenshot to see the visual layout
Browser { action: "navigate", url: "https://example.com" }
Browser { action: "get_state" }
Browser { action: "screenshot" }

Read the screenshot to understand the visual layout. The get_state output gives you:

  • URL and title
  • Scroll position (how far down the page you are)
  • Interactive element count (links, inputs, buttons)
  • Indexed elements: [0] <button> "Submit", [1] <input:text> name="email"
  • Page text (first 3000 chars)
JSON Format

For structured processing, use format: "json":

Browser { action: "get_state", format: "json" }

Returns: { url, title, scroll, stats, elements: [{index, tag, text, ...}], text, session_id, active_tab_id }

Interacting with Elements

Always prefer element indices over CSS selectors. The indices come from get_state and are reliable. CSS selectors can break if the page structure changes.

Browser { action: "click_element", index: 5 }
Browser { action: "type_element", index: 3, value: "hello@example.com" }

Fall back to CSS selectors only when:

  • The element doesn't appear in the indexed list
  • You need to target something very specific (like input[name="csrf_token"])
Browser { action: "click", selector: "#submit-btn" }
Browser { action: "fill", selector: "input[name='email']", value: "test@example.com" }

Keyboard Input (send_keys)

For named keys and key combinations:

Browser { action: "send_keys", keys: "Enter" }
Browser { action: "send_keys", keys: "Tab Tab Enter" }
Browser { action: "send_keys", keys: "Ctrl+a" }
Browser { action: "send_keys", keys: "Shift+Tab" }

Supported named keys: Enter, Tab, Escape, Backspace, Delete, Space, ArrowUp/Down/Left/Right, Home, End, PageUp, PageDown, F1-F12. Modifiers: Ctrl, Alt, Shift, Meta/Cmd. Sequences: space-separated.

File Upload

Browser { action: "upload_file", selector: "input[type='file']", file_path: "/path/to/file.pdf" }

Dropdowns

Browser { action: "dropdown_options", selector: "#my-select" }
Browser { action: "select_dropdown", selector: "#my-select", value: "option_value" }

Structured Data Extraction

Extract structured data from the page using CSS selectors:

Browser { action: "extract", schema: {"title": "h1", "price": ".price", "description": ".desc"} }

Returns JSON: {"title": "Product Name", "price": "$29.99", "description": "..."}

Multi-Tab Workflows

Open, switch between, and close tabs for side-by-side comparison:

Browser { action: "navigate", url: "https://site-a.com" }
Browser { action: "new_tab", url: "https://site-b.com" }
Browser { action: "list_tabs" }
Browser { action: "switch_tab", tab_id: "PREVIOUS_TAB_ID" }
Browser { action: "close_tab", tab_id: "TAB_TO_CLOSE" }

Tab IDs are strings (CDP target IDs). Get them from list_tabs.

Multi-Session Workflows

Use separate sessions for isolated browser instances (different profiles, auth states):

Browser { action: "new_session", session_id: "admin", profile_name: "admin-profile" }
Browser { action: "navigate", url: "https://app.com/admin", session_id: "admin" }
Browser { action: "new_session", session_id: "user", profile_name: "user-profile" }
Browser { action: "navigate", url: "https://app.com/dashboard", session_id: "user" }
Browser { action: "list_sessions" }
Browser { action: "close_session", session_id: "admin" }
Named Profiles

Profiles persist cookies and state across sessions:

  • profile_name: "my-profile" → stored in ~/.claude/browser-profiles/my-profile/
  • user_data_dir: "/custom/path" → raw Chrome user data directory
  • profile_dir: "Profile 1" → Chrome's --profile-directory flag

Attach Mode (Remote Chrome)

Connect to an already-running Chrome instance:

Browser { action: "new_session", session_id: "remote", cdp_url: "http://localhost:9222" }

Or set BROWSER_CDP_URL=http://localhost:9222 env var for the default session.

In attach mode, close disconnects without killing Chrome.

Iframe Support

Browser { action: "list_frames" }
Browser { action: "click_element", index: 0, frame_id: "FRAME_ID" }
Browser { action: "fill", selector: "input", value: "text", frame_id: "FRAME_ID" }

Frame IDs come from list_frames. Use frame_id on any DOM action to target elements inside iframes.

Events & Dialogs

The browser captures events (dialogs, navigations, crashes, downloads):

Browser { action: "get_events" }
Browser { action: "set_dialog_auto_dismiss", enabled: true }

Dialogs (alert/confirm/prompt) are auto-dismissed by default. Disable with enabled: false to handle manually.

The Verify Loop

After every action that changes the page, you must verify:

1. Take action (click, type, navigate, submit)
2. Wait briefly if needed:  Browser { action: "wait_for", selector: ".result", timeout: 3000 }
3. Observe again:           Browser { action: "get_state" }
4. Screenshot again:        Browser { action: "screenshot" }
5. Compare: did the page change as expected?
   - YES → continue to next step
   - NO  → try a different approach (different element, different selector, scroll first)

Iteration Budget

You have a maximum of 10 iterations (observe-act-verify cycles) per task. This prevents infinite loops. If you haven't achieved the goal in 10 iterations:

  1. Stop
  2. Report what you accomplished and what failed
  3. Include the last screenshot as evidence

Count your iterations. Mention the count when reporting.

Show full SKILL.md (371 more words)Show less

Error Recovery

When something doesn't work:

  1. Element not found — the page may have changed. Run get_state again to refresh the element indices.
  2. Click didn't work — the element might be obscured. Try scroll_to first, then click again.
  3. Page didn't load — try wait_for with a key selector, or reload.
  4. Form submission failed — screenshot to see error messages, read them, adjust input.
  5. Same action 3 times — the Browser tool has loop detection. If you get a loop warning, you MUST try a completely different approach.

Fallback chain for clicking:

click_element by index  →  click by CSS selector  →  evaluate with JS click  →  scroll + retry

Scrolling

Pages are often longer than the viewport. The get_state output shows scroll position. If you need elements below the fold:

Browser { action: "scroll_to", value: "500" }     // scroll down 500px
Browser { action: "scroll_to", selector: "#footer" }  // scroll to element
Browser { action: "get_state" }                     // refresh elements after scroll

Always get_state after scrolling — the element indices change.

Evidence and Reporting

When completing a task, provide evidence:

  1. Before state — screenshot of the page before your actions
  2. Actions taken — list of what you did (clicked X, typed Y, navigated to Z)
  3. After state — screenshot showing the result
  4. Verification — what you checked to confirm the task is done

Format your report:

## Browser Task: [description]

**URL:** https://example.com/page
**Session:** default
**Iterations:** 4/10

### Actions
1. Navigated to https://example.com
2. Clicked [3] <button> "Login"
3. Typed email into [5] <input:email>
4. Typed password into [7] <input:password>
5. Sent keys: Enter

### Verification
- Page title changed to "Dashboard"
- User avatar visible in header
- No error messages present

### Screenshots
- Before: [screenshot 1]
- After: [screenshot 2]

JavaScript Evaluation

For complex checks or actions not covered by built-in actions:

Browser { action: "evaluate", value: "document.querySelectorAll('.error').length" }
Browser { action: "evaluate", value: "window.localStorage.getItem('token')" }

This is powerful but use it sparingly — prefer the built-in actions.

Cookies

For authenticated sessions:

Browser { action: "cookies_get" }
Browser { action: "cookies_set", cookie: { name: "session", value: "abc123", domain: ".example.com" } }
Browser { action: "cookies_clear" }

Closing

Always close the browser when done:

Browser { action: "close" }

This frees system resources. The browser will auto-launch again if needed.

Quick Reference

GoalActionNotes
See the pageget_state + screenshotAlways first
See page (structured)get_state with format: "json"For parsing
Click a buttonclick_element with indexFrom get_state
Type texttype_element with index and value
Press keyssend_keys with keys"Enter", "Tab", "Ctrl+a"
Navigatenavigate with url
Upload a fileupload_file with selector, file_path
Read dropdowndropdown_options with selectorReturns [{value,text,selected}]
Select dropdownselect_dropdown with selector, value
Extract dataextract with schemaReturns structured JSON
Wait for contentwait_for with selector
Scroll downscroll_to with value or selector
Open new tabnew_tab with optional url
Switch tabswitch_tab with tab_id
List tabslist_tabs
Close tabclose_tab with tab_id
New sessionnew_session with session_idOptional: profile_name, cdp_url
List sessionslist_sessions
List frameslist_framesFor iframe inspection
Get eventsget_eventsDialog, crash, download, navigation
Run JSevaluate with value
Go backback
Save page as PDFpdf
Check cookiescookies_get
Doneclose

© SeifBenayed, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/browse of SeifBenayed/cloclo.

Open the folder on GitHubat commit 00f5195

Compare with similar skills

Browse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browse this skillSeifBenayed/cloclo114—~2.7kAutomated safety check: PassMIT
Javascript SDKaiskillstore/marketplace4301 repos~3.3kAutomated safety check: PassNone
Autofillinsundial-org/awesome-openclaw-skills663—~1.8kAutomated safety check: PassNone
Agent Browserquran/quran.com-frontend-next1.9k41 repos~3.3kAutomated safety check: PassNone
Dev-Browser CLI AutomationSawyerHood/dev-browser6.7k1 repos~455Automated safety check: PassMIT
Camoufox CLIBin-Huang/camoufox-cli3501 repos~4.5kAutomated safety check: PassMIT

Similar skills

  • Javascript SDK

    aiskillstore/marketplace

    JavaScript/TypeScript SDK for inference.sh - run AI apps, build agents, integrate with all models.

    430 GitHub starsUsed in 1 repo~3.3k tokens
    Backend & APIsAuto-check passed
  • Autofillin

    sundial-org/awesome-openclaw-skills

    Automated web form filling and file uploading skill with Playwright browser automation.

    663 GitHub stars~1.8k tokensUpdated 7 mo ago
    Backend & APIsAuto-check passed
  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 41 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Dev-Browser CLI Automation

    SawyerHood/dev-browser

    Browser automation with persistent named pages via the dev-browser CLI. Use when users ask to navigate websites, fill forms, take screenshots, extract web…

    6.7k GitHub starsUsed in 1 repo~455 tokens
    Productivity & AutomationAuto-check passed
  • Camoufox CLI

    Bin-Huang/camoufox-cli

    Anti-detect browser automation CLI & Skills for AI agents. An agent skill from Bin-Huang/camoufox-cli.

    350 GitHub starsUsed in 1 repo~4.5k tokens
    Productivity & AutomationAuto-check passed
  • Browser

    VibiumDev/vibium

    Automate browsers with the Vibium CLI. An agent skill from VibiumDev/vibium.

    2.9k GitHub stars~4.8k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from SeifBenayed/cloclo

  • Chatgpt Search

    SeifBenayed/cloclo

    Search ChatGPT and extract the full response + hydration JSON that powers the UI.

    114 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check: notes
  • Simplify

    SeifBenayed/cloclo

    Review code for unnecessary complexity and simplify it. An agent skill from SeifBenayed/cloclo.

    114 GitHub stars~430 tokensUpdated 6 mo ago
    Auto-check: notes
  • Commit

    SeifBenayed/cloclo

    Create a git commit with a good message. An agent skill from SeifBenayed/cloclo.

    114 GitHub stars~299 tokensUpdated 6 mo ago
    Auto-check: notes
  • Debug

    SeifBenayed/cloclo

    Troubleshoot and debug issues. An agent skill from SeifBenayed/cloclo.

    114 GitHub stars~419 tokensUpdated 6 mo ago
    Auto-check: notes
  • Review PR

    SeifBenayed/cloclo

    Review a pull request for bugs, security issues, and improvements.

    114 GitHub stars~376 tokensUpdated 6 mo ago
    Auto-check: notes

Questions about Browse

What does Browse do?

Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives. Browse is an agent skill from SeifBenayed/cloclo. Visual feedback browser automation with multi-session, multi-tab, iframe support, and enterprise primitives.

When should I use Browse?

Browse fits situations like: asked to open a site; check a deployment; compare pages side by side; any task involving web interaction.

How do I install Browse in Claude Code?

Run `npx skills add SeifBenayed/cloclo --skill browse -a claude-code`. Or copy the skill folder (.claude/skills/browse in SeifBenayed/cloclo) into .claude/skills/browse in your project. Claude Code loads it when a task matches its description.

How do I install Browse in Codex?

Run `npx skills add SeifBenayed/cloclo --skill browse -a codex`. Or copy the skill folder (.claude/skills/browse in SeifBenayed/cloclo) into .agents/skills/browse in your project. Codex loads it when a task matches its description.

Can I use Browse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SeifBenayed/cloclo --skill browse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browse, .gemini/skills/browse, .github/skills/browse and .opencode/skills/browse in your project.

What does Browse need to run?

SKILL.md names no scripts, command-line tools or credentials: Browse is instructions for the agent only. Its frontmatter pre-approves these tools: Browser, Read, Grep, Glob.

Does Browse access the network?

SKILL.md names 2 domains. In commands or code: site-a.com and site-b.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Browse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browse use?

Browse is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browse use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browse?

Skills that share tags, products or a category with Browse: Javascript SDK (aiskillstore/marketplace, 430 stars), Autofillin (sundial-org/awesome-openclaw-skills, 663 stars), Agent Browser (quran/quran.com-frontend-next, 1.9k stars) and Dev-Browser CLI Automation (SawyerHood/dev-browser, 6.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browse?

SeifBenayed (a GitHub user) maintains it in SeifBenayed/cloclo, which has 114 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 5, 2026.

Source: SeifBenayed/cloclo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.