Best of

Best Web Scraping and Web Search Skills for Claude Code and Codex

Compare web scraping and web search skills for Claude Code and Codex: API keys needed, what each fetches, licence, safety results and robots caveats.

By Updated 7 min read

If you need a web scraping skill for Claude Code or Codex, the choice comes down to how the skill reaches the web. A hosted API such as Firecrawl or Tavily is the simplest route and needs a key. A real-browser skill handles logins and clicks without a key but can act as you. A platform reader pulls from named sites through one command line.

This guide compares the skills in the web scraping topic and the web search topic. Each entry says what the skill fetches, which keys it needs, its licence as the directory records it and the automated check result, plus a trade-off. The details come from each skill's SKILL.md and repository, and the editorial team did not run them against live sites. A plain note on robots rules and platform terms sits in its own section below, and it applies to everything here.

How these skills reach the web

Three patterns cover almost every skill in the two topics.

  • Hosted APIs. The skill tells the agent to call a service that searches, fetches or crawls on its behalf. You pay the service or use its free tier, and the key lives in an environment variable.
  • A real browser. The skill drives Chrome or Chromium, often over the Chrome DevTools Protocol, so JavaScript runs and logged-in sessions work.
  • Routing rules. The skill adds no tool of its own. It tells the agent when to search, when to fetch and how to cite, using whatever search tool you already provide.

Pick by the page, not the brand. Static article pages rarely need a browser, and dashboards behind a login always do.

Best web scraping skills for Claude compared

SkillBest forCredentialsLicenceAutomated check
firecrawl/firecrawl-buildAdding scrape, search and interact to an appFirecrawl API keyISCInfo finding
firecrawl/deep-researchCited multi-source researchSearch and scrape toolsMITPassed
andrewyng/tavily-best-practicesBuilding Tavily search into codeTavily API keyMITPassed
vercel-labs/core (official)agent-browser CLI for real browsingNoneApache-2.0Passed
memtensor/dev-browserScripted browsing with persistent pagesNoneApache-2.0 recordedPassed
eze-is/web-accessTiered search, fetch and browser routingNoneMITPassed
panniantong/agent-reachReading many named platformsOptional loginsMITPassed
yixinforu/web-searchRules for citing search resultsA search toolApache-2.0Passed

"Passed" means the static check reported no findings. An "info finding" is a low-severity note, and the skill's page lists it.

Firecrawl skills

Best for: teams building web data into a product, and agents that need search, scrape and page interaction behind one key.

firecrawl/firecrawl-build is published in Firecrawl's own repository under the ISC licence. It routes between three endpoints: scrape for a known URL, search when a feature starts from a query, and interact for pages that need clicks, forms or pagination. It expects FIRECRAWL_API_KEY in a .env file or the environment, and an optional setting points it at a self-hosted deployment. The directory's check raised an informational finding for the .env reference, which is normal for a skill that handles a key.

The same repository has narrower skills, firecrawl/firecrawl-build-scrape, firecrawl/firecrawl-build-search and firecrawl/firecrawl-build-interact, each loading around 700 to 1,100 tokens. Use one of them when you only need one endpoint.

Trade-off: the skill is about integrating Firecrawl into application code. Its own text sends one-off terminal tasks to the Firecrawl command-line tool instead.

Firecrawl's web agent repository also holds playbooks such as firecrawl/deep-research, firecrawl/structured-extraction and firecrawl/pricing-tracker, all MIT-licensed. They plan searches, extract data into a fixed JSON shape and rate confidence by how many sources agree. They carry no tool of their own and need search and scrape tools from the host. Each also appears in three template folders of the same repository, which are the same files, not independent copies.

One skill, many copies: Firecrawl Scraper

davila7/firecrawl-scraper and sickn33/firecrawl-scraper are the same short skill, at about 250 to 400 tokens each, and both are MIT-licensed. The copy in the davila7 repository names its creator as BenedictKing and gives that author's repository for installation with npx skills add. It needs a Firecrawl API key and Node.js. Since the original repository is not listed in the directory, either use Firecrawl's own skills or install from the author's repository named in the file.

Tavily, Perplexity and other search services

andrewyng/tavily-best-practices is MIT-licensed and lives in Andrew Ng's context-hub repository. It is documentation for developers: how to call Tavily's search, extract, crawl and research methods from Python or JavaScript, choose a search depth and pair them with frameworks such as LangChain. It needs TAVILY_API_KEY and the matching SDK. Use it when you are writing code that searches, not when you want the agent to search for you right now.

For the second case, allenpeng0705/tavily runs searches through a bundled Python script, with domain filters, news mode and optional AI-written answers. The directory records no licence for it, and the check raised an informational finding for the key handling. lichamnesia/tavily-search is an MIT-licensed alternative with a similar job.

davila7/perplexity-search is MIT-licensed and runs Perplexity's Sonar models through OpenRouter, so it needs an OpenRouter key with credit and Python's litellm package. jxxghp/anysearch needs an AnySearch key, ships a bundled command line and is GPL-3.0 licensed, which matters if you plan to redistribute it.

Trade-off for all of these: your queries and the pages you fetch pass through a third-party service. Do not send private text to a search API you have not reviewed.

Real-browser skills

vercel-labs/core (agent-browser)

Best for: driving a browser from the shell with compact page snapshots.

vercel-labs/core is the guide for the agent-browser command line, published by Vercel Labs under Apache-2.0 and marked official in the directory. The workflow takes an accessibility snapshot, assigns short references to elements, then clicks, fills or extracts by reference, which keeps token use low. Sessions persist by name, and credentials go in an auth vault instead of shell history. The skill tells the agent to treat page content as untrusted data.

Requirements: the agent-browser CLI and Chrome or Chromium, which its installer fetches.

Many repositories carry copies of this guide under names such as agent-browser, for example slopus/agent-browser. Prefer the official one.

memtensor/dev-browser and eze-is/web-access

memtensor/dev-browser is a copy, inside the MemOS repository, of the dev-browser skill by SawyerHood, whose file credits that author. Browsing happens through short TypeScript scripts, and named pages keep state between runs, either in a fresh Chromium or an existing Chrome in extension mode. The directory records Apache-2.0 for this copy, while the original repository states MIT, so check the terms that apply to the files you install.

eze-is/web-access is MIT-licensed and tiered: search first, then fetch, then curl, then a real Chrome or Edge session over CDP when a page demands it. It needs Node.js 22 or newer.

For heavier workflows, skyvern-ai/skyvern is AGPL-3.0 and needs a Skyvern API key or a self-hosted instance, and browser-use/browser-harness is MIT-licensed, connects to Chrome over CDP and mentions a Browser Use API key.

Trade-off for browser skills: they can click, type and submit as you. Use a dedicated browser profile, and review what a skill does before pointing it at an account that matters.

Reading named platforms: Agent Reach

panniantong/agent-reach is MIT-licensed. It routes lookups across 16 sites through one command-line tool. Web pages, YouTube subtitles, RSS, public GitHub, Bilibili and V2EX work without a login, while platforms such as Twitter/X, Reddit, Facebook, Instagram and LinkedIn use cookies from your own session. The README warns that cookie logins can be detected and lead to account suspension, and it advises using a dedicated secondary account.

Requirements: the agent-reach tool, plus yt-dlp and other helpers for some sources.

Robots rules, terms and the law in plain words

A skill does not give you the right to collect a site's content. Before you scrape, check three things. First, the site's robots.txt file states which paths automated clients should avoid, and respecting it is the baseline. Second, the terms of service often forbid automated collection or logged-in access, and breaking them can get an account banned or lead to legal claims. Third, copyright and privacy law can apply to what you store, especially personal data.

Practical habits help: use official APIs where they exist, keep request rates low, identify your client, cache instead of refetching, and avoid skills that describe getting past blocks or bot protection. This is general information, not legal advice. Ask a lawyer before building a commercial data product.

How to choose

Start with the question. If the agent must answer from current pages, a search skill with a key is the shortest path. If you are building a product, use the Firecrawl or Tavily skills as integration guides. If the page needs a login or a click, use agent-browser or dev-browser on a throwaway profile. If you only want better search habits, the rules-only skill costs almost nothing.

For install steps, see the pages for Claude Code and Codex. The same sources feed the stock analysis skills comparison, and the top skills list shows what else is popular.

Frequently asked questions

What is the difference between a web scraping skill and a web search skill?

A search skill takes a query and returns a list of results, often with snippets or full text. A scraping skill starts from a known page and extracts its content or structured data, sometimes after clicking through it. Many services in the directory offer both, so the skill you pick usually depends on whether you start from a question or from a URL.

Do I need an API key to scrape websites with an agent?

Not always. Skills that drive a local browser or fetch public pages need no key. Hosted services such as Firecrawl, Tavily and Perplexity do, and the skill reads the key from an environment variable or a .env file. Never paste a key into a prompt or commit it to a repository.

Is web scraping legal?

It depends on the site, the data and where you are. Robots rules, a site's terms of service, copyright and privacy laws can all apply, and logging in with an account usually brings stricter terms. A skill does not grant permission to collect anything. Use official APIs where they exist and ask a lawyer about commercial projects.

Which skill is best for pages that need a login or clicking?

A real-browser skill such as agent-browser, dev-browser or the Firecrawl interact guide fits that case, because it drives a browser session and keeps page state. These skills can act as you on a site, so test on a throwaway account first and keep the agent away from accounts you cannot afford to lose.

Are several of these skills copies of each other?

Yes. Some skills, including the Firecrawl Scraper and the agent-browser guide, appear in many repositories with small edits. This guide points to the original or official source where one is listed, because it is the version most likely to stay current.