Agent skill

Benchmark Browsing Skills

by browsing-skills in browsing-skills/browsing-skills

A skill your agent uses when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls…

MITAuto-check passedAI & LLM Engineering

Install Benchmark Browsing Skills

skills CLI
$ npx skills add browsing-skills/browsing-skills --skill benchmark-browsing-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install browsing-skills/browsing-skills benchmark-browsing-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/browsing-skills/browsing-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark-browsing-skills .claude/skills/benchmark-browsing-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-browsing-skills
GitHub stars
117
Token cost
~596 tokens
SKILL.md length
260 words
Files
2 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls…

  • Works in 6 steps: Identify the website, action, target… → Read only the matching site SKILL.md and… → Run the with-skill API pass with the… → …
  • Benchmarking browsing-skills actions
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Comparing a maintained site skill against a no-skill browser agent

What it does

Benchmark Browsing Skills is an agent skill from browsing-skills/browsing-skills. Use when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls, browser/tool calls, and result quality. Trigger whenever the user asks to benchmark a browsing skill, compare with/without skill, measure token or time savings, or update benchmark tables in this repo.

Its SKILL.md is about 600 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api-benchmark-protocol.md`).

It sits in AI & LLM Engineering, covering LLM API integration and Browser automation. It works with OpenAI. The repository describes itself as: Open-source registry of browser-automation skills for AI agents. The licence is MIT.

When your agent uses it

  • Benchmarking browsing-skills actions
  • Comparing a maintained site skill against a no-skill browser agent
  • Measuring OpenAI API token usage
  • Browser/tool calls

Example prompts

  • “/benchmark-browsing-skills”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify the website, action, target URL/query, browser layer, expected output fields, session requirements, and success criteria.
  2. Read only the matching site SKILL.md and action reference from skills//references/.md; exclude benchmark history and other non-action…
  3. Run the with-skill API pass with the skill index/reference available.
  4. Run the without-skill API pass without skill files; provide only browser tools and live page observations.
  5. Capture tokens, API calls, wall time, browser calls, and result quality for both passes.
  6. Record results in

What it can do on your machine

Read from SKILL.md and the folder at commit 07855b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Browsing Skills loads about 596 tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 260 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~596
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from browsing-skills/browsing-skills at commit 07855b5, republished under its MIT licence (© browsing-skills). 260 words, ~596 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-browsing-skills/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
benchmark-browsing-skills
description
Use when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls, browser/tool calls, and result quality. Trigger whenever the user asks to benchmark a browsing skill, compare with/without skill, measure token or time savings, or update benchmark tables in this repo.

Benchmark Browsing Skills

Use this skill to run and document fair benchmarks for actions in this repo.

The benchmark compares two agents on the same live browser task:

  • With skill: loads the site action index plus only the action-facing parts of one action reference, then runs the maintained action code.
  • Without skill: receives browser access but no skill/reference, inspects the live page, derives extraction logic, and iterates until success or timeout.

For official numbers, prefer OpenAI API-separated runs over Codex subagents. Subagents are useful for quick qualitative comparisons, but API runs give explicit token accounting and cleaner cost comparison.

Use the same browser layer and session state for both branches. For actions that depend on login, personalization, region, cart state, or existing cookies, prefer Chrome Bridge /run-action against the user's loaded Chrome session. Use Playwright only for public, unauthenticated actions or when you intentionally create an equivalent browser profile for both branches.

Workflow

  1. Identify the website, action, target URL/query, browser layer, expected output fields, session requirements, and success criteria.
  2. Read only the matching site SKILL.md and action reference from skills/<domain>/references/<action>.md; exclude benchmark history and other non-action prose from the measured with-skill prompt.
  3. Run the with-skill API pass with the skill index/reference available.
  4. Run the without-skill API pass without skill files; provide only browser tools and live page observations.
  5. Capture tokens, API calls, wall time, browser calls, and result quality for both passes.
  6. Record results in:
    • BENCHMARKS.md
    • the site skills/<domain>/SKILL.md benchmark table
    • the action reference ## Benchmark section

Read references/api-benchmark-protocol.md before designing or running a benchmark.

© browsing-skills, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/benchmark-browsing-skills of browsing-skills/browsing-skills.

  • SKILL.md
  • references/api-benchmark-protocol.md

Open the folder on GitHubat commit 07855b5

Compare with similar skills

Benchmark Browsing Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Browsing Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Browsing Skills this skillbrowsing-skills/browsing-skills117—~596Automated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Agnes Free Textkangarooking/agnes-free-model-skills199—~630Automated safety check: PassMIT
Openai Docstheowenyoung/home1151 repos~1.4kAutomated safety check: PassApache-2.0
Instructor Structured LLM OutputsOrchestra-Research/AI-Research-SKILLs13k6 repos~4.2kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k—~2.8kAutomated safety check: PassMIT

Similar skills

  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agnes Free Text

    kangarooking/agnes-free-model-skills

    Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.

    199 GitHub stars~630 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Openai Docs

    theowenyoung/home

    A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

    115 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Instructor Structured LLM Outputs

    Orchestra-Research/AI-Research-SKILLs

    Shows how to pull validated, typed data out of LLM responses with Instructor and Pydantic models, including retries on failure and partial streaming.

    13k GitHub starsUsed in 6 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from browsing-skills/browsing-skills

All 24 skills in this repo
  • Website Browsing Skills Index

    browsing-skills/browsing-skills

    Umbrella skill for a library of website-specific browsing skills. Use when the user's request targets one of these specific websites: <!-- DOMAINS:START…

    117 GitHub stars~1.6k tokensUpdated 4 mo ago
    Auto-check passed
  • Browsing Airbnb

    browsing-skills/browsing-skills

    A skill your agent uses when the user wants to interact with Airbnb — search public stays, extract listing details, extract visible reviews, or read availability and price breakdowns.

    117 GitHub stars~464 tokensUpdated 4 mo ago
    Auto-check passed
  • Browsing Aliexpress

    browsing-skills/browsing-skills

    A skill your agent uses when the user wants to interact with AliExpress — search products, inspect product detail pages, read the shopping cart, or add explicitly requested items to cart.

    117 GitHub stars~583 tokensUpdated 4 mo ago
    Auto-check passed
  • Amazon.com Browsing Actions

    browsing-skills/browsing-skills

    Searches Amazon products and groceries, reads product pages and cart contents, and adds explicitly named items to cart, stopping short of checkout.

    117 GitHub stars~702 tokensUpdated 4 mo ago
    Auto-check passed
  • Capterra Data Extraction

    browsing-skills/browsing-skills

    Extracts software product ratings, pricing and written user reviews from Capterra pages with a real browser, since the site has no public API.

    117 GitHub stars~532 tokensUpdated 4 mo ago
    Auto-check passed
  • Facebook Marketplace Browsing

    browsing-skills/browsing-skills

    Use when the user wants to interact with Facebook Marketplace — search visible listings, extract listing details, and summarize seller/profile information…

    117 GitHub stars~459 tokensUpdated 4 mo ago
    Auto-check passed

Works with

Questions about Benchmark Browsing Skills

What does Benchmark Browsing Skills do?

A skill your agent uses when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls…. Benchmark Browsing Skills is an agent skill from browsing-skills/browsing-skills. Use when benchmarking browsing-skills actions, comparing a maintained site skill against a no-skill browser agent, measuring OpenAI API token usage, wall time, API calls, browser/tool calls, and result quality.

When should I use Benchmark Browsing Skills?

Benchmark Browsing Skills fits situations like: benchmarking browsing-skills actions; comparing a maintained site skill against a no-skill browser agent; measuring OpenAI API token usage; browser/tool calls.

How do I install Benchmark Browsing Skills in Claude Code?

Run `npx skills add browsing-skills/browsing-skills --skill benchmark-browsing-skills -a claude-code`. Or copy the skill folder (.agents/skills/benchmark-browsing-skills in browsing-skills/browsing-skills) into .claude/skills/benchmark-browsing-skills in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Browsing Skills in Codex?

Run `npx skills add browsing-skills/browsing-skills --skill benchmark-browsing-skills -a codex`. Or copy the skill folder (.agents/skills/benchmark-browsing-skills in browsing-skills/browsing-skills) into .agents/skills/benchmark-browsing-skills in your project. Codex loads it when a task matches its description.

Can I use Benchmark Browsing Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browsing-skills/browsing-skills --skill benchmark-browsing-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-browsing-skills, .gemini/skills/benchmark-browsing-skills, .github/skills/benchmark-browsing-skills and .opencode/skills/benchmark-browsing-skills in your project.

What does Benchmark Browsing Skills need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Browsing Skills is instructions for the agent only.

Does Benchmark Browsing Skills access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Browsing Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Browsing Skills use?

Benchmark Browsing Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Browsing Skills use?

About 596 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Browsing Skills?

Skills that share tags, products or a category with Benchmark Browsing Skills: Dingo Verify (MigoXLab/dingo, 757 stars), Agnes Free Text (kangarooking/agnes-free-model-skills, 199 stars), Openai Docs (theowenyoung/home, 115 stars) and Instructor Structured LLM Outputs (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Browsing Skills?

browsing-skills (a GitHub organization) maintains it in browsing-skills/browsing-skills, which has 117 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on June 10, 2026.

Source: browsing-skills/browsing-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.