Agent skill

Browser Intent

by ruvnet in ruvnet/ruflo

Execute a natural-language browser intent via page-agent (browseract) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider…

MITAuto-check: notesProductivity & Automation

Install Browser Intent

skills CLI
$ npx skills add ruvnet/ruflo --skill browser-intent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo browser-intent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-browser/skills/browser-intent .claude/skills/browser-intent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser-intent
GitHub stars
74k
Token cost
~1.2k tokens
SKILL.md length
517 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Execute a natural-language browser intent via page-agent (browseract) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider…

  • Works in 4 steps: Call browser_act with a task string, and… → Read the response contract → On contentFlagged: true, the returned… → …
  • Tasks that involve Browser automation
  • SKILL.md covers When to use, Steps, Provider requirements (why… and Caveats
  • Calls npm; needs ANTHROPIC_API_KEY and OPENROUTER_API_KEY

What it does

Browser Intent is an agent skill from ruvnet/ruflo. Execute a natural-language browser intent via page-agent (browseract) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider isn't configured

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation. It works with OpenAI. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Browser automation

Example prompts

  • “/browser-intent”

Requirements

  • Node.js
  • A credential in ANTHROPIC_API_KEY
  • A credential in OPENROUTER_API_KEY
  • Pre-approved tools (allowed-tools): mcp__plugin_ruflo-core_ruflo__browser_act, mcp__plugin_ruflo-core_ruflo__browser_open, mcp__plugin_ruflo-core_ruflo__browser_snapshot, mcp__plugin_ruflo-core_ruflo__browser_close, mcp__plugin_ruflo-core_ruflo__aidefence_has_pii, mcp__plugin_ruflo-core_ruflo__aidefence_is_safe, Bash, Read

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Call browser_act with a task string, and optionally url (navigates first) and session (default "default")
  2. Read the response contract
  3. On contentFlagged: true, the returned result has already been redacted by AIDefence (PII or a prompt-injection/threat pattern was detected…
  4. Prefer a recorded session (browser-record) when the interaction matters enough to replay later; browser_act itself does not open an RVF…

What it can do on your machine

Read from SKILL.md and the folder at commit de590e1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • mcp__plugin_ruflo-core_ruflo__browser_act
    • mcp__plugin_ruflo-core_ruflo__browser_open
    • mcp__plugin_ruflo-core_ruflo__browser_snapshot
    • mcp__plugin_ruflo-core_ruflo__browser_close
    • mcp__plugin_ruflo-core_ruflo__aidefence_has_pii
    • mcp__plugin_ruflo-core_ruflo__aidefence_is_safe
    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • OPENROUTER_API_KEY
    • OLLAMA_API_KEY
    • CLAUDE_FLOW_PAGE_AGENT_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser Intent loads about 1.2k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 517 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: mcp__plugin_ruflo-core_ruflo__browser_act, mcp__plugin_ruflo-core_ruflo__browser_open, mcp__plugin_r

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit de590e1, republished under its MIT licence (© ruvnet). 517 words, ~1,197 tokens.

Download SKILL.mdSave it as .claude/skills/browser-intent/SKILL.md (or your agent's skills folder).
name
browser-intent
description
Execute a natural-language browser intent via page-agent (browser_act) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider isn't configured
allowed-tools
mcp__plugin_ruflo-core_ruflo__browser_act, mcp__plugin_ruflo-core_ruflo__browser_open, mcp__plugin_ruflo-core_ruflo__browser_snapshot, mcp__plugin_ruflo-core_ruflo__browser_close, mcp__plugin_ruflo-core_ruflo__aidefence_has_pii, mcp__plugin_ruflo-core_ruflo__aidefence_is_safe, Bash, Read
argument-hint
<task-description> [--url <url>] [--session <id>]

Browser Intent

Natural-language layer on top of the low-level browser_* selector tools. Where browser-extract and browser-form-fill compose selector-based primitives (browser_click, browser_fill, browser_snapshot), browser-intent lets the caller say what they want ("Click the login button", "Fill the search box with cats and submit") and delegates execution to page-agent — in-page injected JS that turns the DOM into text and drives an LLM tool-call loop against it.

When to use

  • The target element is easier to describe in words than to select reliably (dynamic class names, ambiguous structure, A/B-tested markup).
  • A one-shot interaction where writing out a selector chain isn't worth it.
  • Prefer browser_click / browser_fill / browser_snapshot directly when you already know the exact selector or ref (@e1) — browser_act adds LLM latency + cost that a direct selector call doesn't.

Steps

  1. Call browser_act with a task string, and optionally url (navigates first) and session (default "default"):
    mcp__plugin_ruflo-core_ruflo__browser_act({
      task: "Click the login button",
      url: "https://example.com/account",
      session: "my-session"
    })
  2. Read the response contract:
    • { success: true, result, steps, history, contentFlagged, llmSource } — the intent executed. result is the AIDefence-gated final text page-agent produced; history is the full step trace (reflection + action + tool result per step); steps is history.length.
    • { success: true, degraded: true, reason, hint } — page-agent isn't installed, or no OpenAI-compatible LLM provider is configured. Never treat degraded: true as an error to retry — surface the hint and fall back to selector-based browser_* tools instead.
    • { success: false, error, ... } — a real failure (browser open failed, injection failed, execution timed out, or page-agent's own execute() reported success:false).
  3. On contentFlagged: true, the returned result has already been redacted by AIDefence (PII or a prompt-injection/threat pattern was detected in the page-agent output) — do not attempt to recover the original text.
  4. Prefer a recorded session (browser-record) when the interaction matters enough to replay later; browser_act itself does not open an RVF container — it operates on whatever session id you pass (or "default").
Show full SKILL.md (217 more words)Show less

Provider requirements (why this degrades so often)

page-agent calls its LLM directly from the browser page context via a plain OpenAI-compatible POST {baseURL}/chat/completions. That means:

  • A bare ANTHROPIC_API_KEY is not sufficient — Anthropic's native API is a different shape (/v1/messages).
  • Configure one of: OPENROUTER_API_KEY (OpenRouter, OpenAI-compatible), OLLAMA_API_KEY (Ollama Cloud, OpenAI-compatible), or CLAUDE_FLOW_PAGE_AGENT_BASE_URL + CLAUDE_FLOW_PAGE_AGENT_API_KEY for a custom OpenAI-compatible endpoint.
  • The real provider key never enters the page: browser_act starts a short-lived loopback HTTP proxy that holds the key server-side and injects the real Authorization header itself. The page only ever sees a 127.0.0.1 URL and a placeholder key string.

Caveats

  • page-agent is an optionalDependencies entry (npm i page-agent if the doctor/degraded hint asks for it) — this plugin stays fully operational without it; you simply lose the natural-language layer and fall back to selector-based tools.
  • The npm bundle's demo auto-init tail (which would otherwise construct a second PageAgent instance against Alibaba's public test endpoint) is stripped before injection — you should never see traffic to a page-ag-testing-* host from this tool.
  • Every successful browser_act call best-effort records the intent + resulting trajectory into the browser memory namespace (ADR-174 distillation loop). This is fire-and-forget — a memory-store failure never fails the tool call.
  • timeoutMs (default 120000) bounds how long browser_act polls for execute() to settle; a slow multi-step intent may need a higher value.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-browser/skills/browser-intent of ruvnet/ruflo.

Open the folder on GitHubat commit de590e1

Compare with similar skills

Browser Intent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser Intent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser Intent this skillruvnet/ruflo74k—~1.2kAutomated safety check: NotesMIT
Codex Bridge Chatgptanightmonarch/codex-bridge-chatgpt145—~1.1kAutomated safety check: PassMIT
Consult Oracletobihagemann/turbo4061 repos~1.1kAutomated safety check: PassMIT
BrowserWing AdminMemTensor/MemOS12k—~4.1kAutomated safety check: NotesApache-2.0
Scrapingbee CLIScrapingBee/scrapingbee-cli108—~3.1kAutomated safety check: PassMIT
Page AgentTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT

Similar skills

  • Codex Bridge Chatgpt

    anightmonarch/codex-bridge-chatgpt

    A skill your agent uses when a complex repository task needs ChatGPT web reasoning while Codex remains responsible for local evidence, edits, and tests.

    145 GitHub stars~1.1k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Consult Oracle

    tobihagemann/turbo

    Consult ChatGPT Pro via ChatGPT browser automation for problems that resist standard approaches.

    406 GitHub starsUsed in 1 repo~1.1k tokens
    Productivity & AutomationAuto-check passed
  • BrowserWing Admin

    MemTensor/MemOS

    Installs, configures and operates BrowserWing, a browser automation platform, covering Chrome setup, LLM provider configuration, and creating or running automation scripts.

    12k GitHub stars~4.1k tokensUpdated 8 days ago
    Productivity & AutomationAuto-check: notes
  • Scrapingbee CLI

    ScrapingBee/scrapingbee-cli

    Fetch and read any web page, search the web, crawl a site, or pull structured data out of pages.

    108 GitHub stars~3.1k tokensUpdated 26 days ago
    Productivity & AutomationAuto-check passed
  • Page Agent

    Tommy-yw/RunbookHermes

    Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

    546 GitHub starsUsed in 3 repos~2.3k tokens
    Productivity & AutomationAuto-check: notes
  • Surf

    w-winter/dot314

    Control Chrome browser via CLI for testing, automation, and debugging.

    139 GitHub stars~7.6k tokensUpdated 4 days ago
    Productivity & AutomationAuto-check: warnings

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 2 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Works with

Questions about Browser Intent

What does Browser Intent do?

Execute a natural-language browser intent via page-agent (browseract) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider…. Browser Intent is an agent skill from ruvnet/ruflo.

When should I use Browser Intent?

Browser Intent fits situations like: tasks that involve Browser automation.

How do I install Browser Intent in Claude Code?

Run `npx skills add ruvnet/ruflo --skill browser-intent -a claude-code`. Or copy the skill folder (plugins/ruflo-browser/skills/browser-intent in ruvnet/ruflo) into .claude/skills/browser-intent in your project. Claude Code loads it when a task matches its description.

How do I install Browser Intent in Codex?

Run `npx skills add ruvnet/ruflo --skill browser-intent -a codex`. Or copy the skill folder (plugins/ruflo-browser/skills/browser-intent in ruvnet/ruflo) into .agents/skills/browser-intent in your project. Codex loads it when a task matches its description.

Can I use Browser Intent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill browser-intent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser-intent, .gemini/skills/browser-intent, .github/skills/browser-intent and .opencode/skills/browser-intent in your project.

What does Browser Intent need to run?

Going by SKILL.md and its folder, Browser Intent needs the command-line tools its instructions call (npm) and credentials named ANTHROPIC_API_KEY, OPENROUTER_API_KEY, OLLAMA_API_KEY and CLAUDE_FLOW_PAGE_AGENT_API_KEY. Our summary lists: Node.js; A credential in ANTHROPIC_API_KEY; A credential in OPENROUTER_API_KEY. Its frontmatter pre-approves these tools: mcp__plugin_ruflo-core_ruflo__browser_act, mcp__plugin_ruflo-core_ruflo__browser_open, mcp__plugin_ruflo-core_ruflo__browser_snapshot, mcp__plugin_ruflo-core_ruflo__browser_close, mcp__plugin_ruflo-core_ruflo__aidefence_has_pii, mcp__plugin_ruflo-core_ruflo__aidefence_is_safe, Bash, Read.

Does Browser Intent access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Browser Intent safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Browser Intent use?

Browser Intent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser Intent use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser Intent?

Skills that share tags, products or a category with Browser Intent: Codex Bridge Chatgpt (anightmonarch/codex-bridge-chatgpt, 145 stars), Consult Oracle (tobihagemann/turbo, 406 stars), BrowserWing Admin (MemTensor/MemOS, 12k stars) and Scrapingbee CLI (ScrapingBee/scrapingbee-cli, 108 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser Intent?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,012 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 7, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.