Agent skill

Webcrawler Deep Crawl

by browser-act in browser-act/skills

Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links.

MITAuto-check passedDatabases

Install Webcrawler Deep Crawl

skills CLI
$ npx skills add browser-act/skills --skill webcrawler-deep-crawl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install browser-act/skills webcrawler-deep-crawl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/browser-act/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/solutions/search-research/webcrawler-deep-crawl .claude/skills/webcrawler-deep-crawl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
webcrawler-deep-crawl
GitHub stars
6.1k
Token cost
~4.3k tokens
SKILL.md length
1,880 words
Files
5 (incl. scripts)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links.

  • Works in 2 steps: Tool Readiness → Login Verification (when prerequisites…
  • User mentions deep crawl website
  • SKILL.md covers Language, Objective, Prerequisites and Pre-execution Checks, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Webcrawler Deep Crawl is an agent skill from browser-act/skills. Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links. Use when user mentions deep crawl website, recursive crawl, crawl a whole site, scrape entire website, scrape docs site, scrape documentation, scrape knowledge base, scrape blog, build RAG corpus, build vector database from website, knowledge base for chatbot, GPT knowledge files, llms.txt, sitemap crawl, BFS crawl, scrape with…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/discover-links.py`, `scripts/discover-llms-txt.py` and `scripts/discover-sitemap.py`).

It sits in Databases, covering Web scraping, Vector databases and AI search optimization. The repository describes itself as: Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session… The licence is MIT.

When your agent uses it

  • User mentions deep crawl website
  • Recursive crawl
  • Crawl a whole site
  • Scrape entire website

Example prompts

  • “/webcrawler-deep-crawl”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Tool Readiness
  2. Login Verification (when prerequisites include login requirement)

What it can do on your machine

Read from SKILL.md and the folder at commit 11c057b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Webcrawler Deep Crawl loads about 4.3k tokens when it runs. Until then it costs about 257 tokens; SKILL.md has 1,880 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~257
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from browser-act/skills at commit 11c057b, republished under its MIT licence (© browser-act). 1,880 words, ~4,285 tokens.

Download SKILL.mdSave it as .claude/skills/webcrawler-deep-crawl/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
webcrawler-deep-crawl
description
Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links. Use when user mentions deep crawl website, recursive crawl, crawl a whole site, scrape entire website, scrape docs site, scrape documentation, scrape knowledge base, scrape blog, build RAG corpus, build vector database from website, knowledge base for chatbot, GPT knowledge files, llms.txt, sitemap crawl, BFS crawl, scrape with depth or page limit, include exclude URL globs, remove boilerplate, strip navigation header footer, website to markdown, website to text, multi-page extraction, bulk page scraping, clean markdown from URL, docs site to markdown corpus, site to clean corpus. Also applies to building RAG pipelines, indexing a customer site, syncing docs into a vector store, generating training corpora from any docs hub, or expanding a single start URL into a clean corpus of every reachable in-scope page.

Website Deep Crawl

Input: one or more start URLs (+ optional scope, depth, page-count, globs, removal selectors). Output: per-page records {url, crawl, metadata, text, markdown, html, outboundLinks} for every page reached within scope.

Language

All process output to user (progress updates, process notifications) follows the user's language.

Objective

From a small set of start URLs, breadth-first crawl every reachable in-scope page, strip boilerplate (navigation, header, footer, cookie banners, etc.), and emit per-page LLM-ready content (text / markdown / HTML) plus structured metadata — suitable for feeding RAG pipelines, vector databases, or chatbot knowledge bases.

Prerequisites

  • One or more start URLs are provided by the caller.
  • Target pages are publicly reachable, OR the running browser is already logged in for any pages behind authentication.
  • A working directory is available for writing per-page JSON records and the crawl state file.

Pre-execution Checks

1. Tool Readiness

If browser-act has been confirmed available in the current session → skip this step.

Invoke browser-act via Skill tool to load usage. If installation or configuration issues arise, follow its guidance to resolve then retry.

2. Login Verification (when prerequisites include login requirement)

If login status for the target site has been confirmed in the current session → skip this step.

Otherwise: open the target site and observe the page login status:

  • Logout/sign-out entry, user avatar, or username exists → logged in, continue execution
  • Login/register entry exists with no logout entry → not logged in, inform the user that login is needed first, assist the user in completing the login flow

User refuses or cannot log in → terminate execution.

Capability Components

This Skill's operational boundary = what the user can manually do in their browser. It only reads data already displayed to the user on the page, never bypassing authentication or access controls. Its role is equivalent to copy-pasting on the user's behalf — the data is already on screen, automation merely saves time. JS code is encapsulated in Python files under the scripts/ directory, invoked via browser-act --session <name> eval "$(python scripts/xxx.py {params})". $(...) is bash syntax; use the bash tool for execution. The eval token below refers to the browser-act CLI eval subcommand — always include the browser-act --session <name> prefix and the "$(...)" substitution.

Below are all atomic capabilities discovered and verified during the exploration phase, listed by command template with parameters. Simply invoke them as needed — no need to read scripts/*.py source code or re-verify. Only inspect scripts when execution fails for troubleshooting. Combine freely as needed during execution.

API: discover URLs from /llms.txt

Probes {origin}/llms.txt (a convention used by LLM-friendly documentation sites) and returns every URL it lists. Fast path — try this first.

eval "$(python scripts/discover-llms-txt.py 'https://example.com')"

Parameters:

  • positional origin: the site origin (scheme + host), e.g. https://docs.example.com

Output example:

json
{
  "error": false,
  "source": "llms.txt",
  "count": 88,
  "urls": ["https://docs.example.com/intro", "https://docs.example.com/install"]
}

On failure (file missing or HTTP error): {"error": true, "message": "llms.txt not available (HTTP 404)", "urls": []} — move on to sitemap discovery.

API: discover URLs from /sitemap.xml

Probes {origin}/sitemap.xml and {origin}/sitemap_index.xml, follows nested sitemap indexes, and collects every <loc> URL.

eval "$(python scripts/discover-sitemap.py 'https://example.com' --max-urls 5000)"

Parameters:

  • positional origin: site origin
  • --max-urls: hard cap to stop ballooning sitemap indexes, default 5000

Output example:

json
{
  "error": false,
  "source": "sitemap.xml",
  "count": 86,
  "urls": ["https://example.com/page-a", "https://example.com/page-b"]
}

On failure: {"error": true, "message": "No sitemap found at standard paths", "urls": []} — fall back to DOM link discovery.

DOM: discover URLs from the current page

Reads <a href> from the currently loaded page DOM, normalizes them to absolute URLs, drops fragments, asset extensions, and out-of-scope links, and optionally applies include / exclude glob filters. Use when llms.txt and sitemap are both unavailable, or to extend the queue with links discovered while crawling.

eval "$(python scripts/discover-links.py 'https://example.com/docs/' --include-globs '[]' --exclude-globs '["**/changelog/**"]')"

Parameters:

  • positional start_url: scope URL. Only links under its directory (or equal to it) are kept.
  • --include-globs: JSON array of glob patterns; if non-empty, a link must match at least one to be kept. Default [] (no include filter).
  • --exclude-globs: JSON array of glob patterns; matching links are dropped. Default [].

Glob semantics: ** matches any characters (including /), * matches any except /, ? matches one character. Example: https://example.com/{docs,api}/**.

Output example:

json
{
  "error": false,
  "source": "dom",
  "page": "https://example.com/docs/intro",
  "scope_base": "https://example.com/docs/",
  "count": 12,
  "links": ["https://example.com/docs/intro", "https://example.com/docs/install"]
}

The core extractor. Run this on every crawled page after wait stable. Returns the page body in the requested format(s), structured metadata, and the in-scope outbound links found on this page (so callers can extend the BFS queue without a second DOM pass).

eval "$(python scripts/extract-page-content.py 'https://example.com/docs/' --output-format markdown --remove-selectors '.cookie-banner,#chat-widget' --include-globs '[]' --exclude-globs '[]')"

Parameters:

  • positional start_url: scope URL — used to filter the outboundLinks array to in-scope links only.
  • --output-format: one of markdown, text, html, all. Default markdown. all includes every body field.
  • --remove-selectors: comma-separated CSS selectors to delete from the chosen content root before extraction (in addition to the built-in boilerplate list). Use this to strip site-specific chrome (e.g. .cookie-banner, #chat-widget).
  • --keep-selector: a single CSS selector identifying the main content area. If set, only this element's content is extracted (overrides the built-in content-root heuristic). Use this when the site has a known main wrapper, e.g. article.docs-content.
  • --include-globs / --exclude-globs: same semantics as discover-links; applied to the returned outboundLinks array.

Content-root heuristic (used when --keep-selector is not provided), in priority order: <main>, [role="main"], <article>, #content, .content, <body>.

Built-in boilerplate removal (always applied) includes: nav, header, footer, aside, script, style, noscript, iframe, [role="navigation"], [role="banner"], [role="contentinfo"], .cookie*, .advertisement, .modal, .popup, .share, .social, .breadcrumb, .toc, [aria-hidden="true"], etc.

Output example:

json
{
  "error": false,
  "url": "https://example.com/docs/intro",
  "crawl": {
    "loadedUrl": "https://example.com/docs/intro",
    "loadedTime": "2026-06-25T04:37:23.643Z",
    "referrerUrl": null
  },
  "metadata": {
    "canonicalUrl": "https://example.com/docs/intro",
    "title": "Introduction — Example Docs",
    "description": "Get started with Example.",
    "author": null,
    "keywords": [],
    "languageCode": "en",
    "publishedAt": null,
    "modifiedAt": null,
    "ogImage": "https://example.com/og.png",
    "ogType": "website"
  },
  "text": "Introduction\n\nGet started with Example…",
  "markdown": "# Introduction\n\nGet started with Example…",
  "outboundLinks": [
    "https://example.com/docs/install",
    "https://example.com/docs/quick-start"
  ]
}

[AI Intervention] On pages with infinite scroll or lazy-loaded sections, before invoking this script: scroll down repeatedly (until page height stops growing or a max-scroll cap is hit) so the dynamic content is in the DOM. The script reads what is currently rendered — it cannot trigger lazy loading on its own.

Composite: full deep crawl from start URL(s)

End-to-end flow. The Agent orchestrates discovery → BFS queue → per-page extraction → persistence. Records are written one per page so that crashes can resume from where they stopped.

Step 1 — seed the queue:

For each start URL, perform discovery in this priority order and merge results. Stop discovery once the queue has enough URLs to honor max_pages.

a. eval "$(python scripts/discover-llms-txt.py '{origin}')" — instant full list when available. b. eval "$(python scripts/discover-sitemap.py '{origin}' --max-urls {cap})" — broad coverage. c. If both fail or return zero in-scope URLs: navigate to the start URL, wait stable, then eval "$(python scripts/discover-links.py '{start_url}' --include-globs '{globs}' --exclude-globs '{globs}')".

Filter all discovered URLs to scope: every URL must start with the start URL's origin + dirname/, and must satisfy include / exclude globs.

Step 2 — initialize state:

Create the following in the working directory:

  • crawl_state.json — {visited: [], queue: [...seedUrls], output_dir: "...", config: {...}}
  • pages/ directory — one JSON file per successfully crawled page, named by URL hash

Step 3 — BFS loop (one URL at a time, in queue order, until max_pages reached or queue empty):

For each url popped from the queue:

a. Skip if url is in visited or its metadata.canonicalUrl (from a prior page) is already in visited. b. navigate {url} → wait stable (use --timeout 60000 for slow sites). c. (Optional, only when the target has lazy-loaded content) scroll down until height stable or 10 scrolls done. d. eval "$(python scripts/extract-page-content.py '{start_url}' --output-format {format} --remove-selectors '{selectors}' --include-globs '{globs}' --exclude-globs '{globs}')". e. If result error: true → record the failure into crawl_state.json#failed and continue. Do NOT retry blindly. f. Write the JSON record to pages/{hash}.json. g. Append url and metadata.canonicalUrl to visited. h. For each link in outboundLinks: if not in visited and not already in queue, append to queue. Cap queue size at max_pages * 4 to bound memory. i. Persist crawl_state.json after every page (resume on next run if interrupted).

Step 4 — finalize:

Emit a summary {total_pages, success_count, failed_count, duration_seconds, output_dir}. Optionally concatenate all pages/*.json into a single dataset.jsonl for downstream loading.

Configuration parameters (set by the Agent before Step 1 based on user request):

  • start_urls: list of seed URLs (one or more)
  • max_pages: hard cap on pages crawled, default 100
  • max_depth: hard cap on link depth from start URL, default unlimited (-1)
  • include_globs: JSON array, default []
  • exclude_globs: JSON array, default []
  • output_format: markdown | text | html | all, default markdown
  • remove_selectors: comma-separated site-specific selectors to strip, default ""
  • keep_selector: optional content-root selector, default ""
  • output_dir: where to write pages/ and crawl_state.json, default ./output/{site}-crawl/

Output example (per-page record, written to pages/{hash}.json):

json
{
  "url": "https://example.com/docs/intro",
  "crawl": { "loadedUrl": "...", "loadedTime": "...", "referrerUrl": null, "depth": 0 },
  "metadata": { "title": "...", "description": "...", "languageCode": "en", "canonicalUrl": "..." },
  "text": "...",
  "markdown": "# ...",
  "outboundLinks": ["..."]
}

Summary example:

json
{
  "total_pages": 86,
  "success_count": 84,
  "failed_count": 2,
  "duration_seconds": 412,
  "output_dir": "./output/example-crawl/"
}
Show full SKILL.md (545 more words)Show less

Pagination

This is a recursive crawler, not a list with pages. Boundary control is by max_pages (queue length cap) and max_depth (links-away-from-start cap), not by API pagination. Termination: queue empty OR max_pages reached OR no more in-scope outbound links discovered.

Success Criteria

success_count >= 1 AND success_count / total_pages >= 0.8 AND extracted markdown body length per page > 100 chars for at least 80% of pages

Known Limitations

  • Pages behind authentication require the running browser to be logged in beforehand — this Skill does not handle login flows.
  • Pages whose content is rendered after async user interaction beyond simple scroll (e.g. clicking "Load more", expanding accordions to reveal content) need the Agent to add the relevant click before invoking extract-page-content; otherwise the hidden content will be missing.
  • <iframe> content is removed by default (treated as boilerplate). If a page's main content lives inside an iframe, the Agent must first navigate into the iframe URL and crawl it separately.
  • File downloads (PDF, DOCX, XLSX) are not handled. URLs with these extensions are intentionally filtered from the crawl queue.
  • Some sites' boilerplate is structurally indistinguishable from main content (e.g. "Was this page helpful?" footers placed inside <main>). The Agent should pass site-specific patterns via --remove-selectors to strip them.
  • Single-page applications that load content via JS after navigation may need a longer wait stable --timeout. Anti-scraping CAPTCHAs are not bypassed by this Skill; the calling browser must already pass them.

Execution Efficiency

  • Batch orchestration: Write a bash script to loop through the command templates serially within a single session; do not parallelize within one browser (prone to triggering anti-scraping restrictions). Refer to rate information in "Known Limitations" above to add appropriate intervals. To increase throughput, open multiple stealth browser sessions and distribute work across them — each session has an independent fingerprint so rate limits apply per session
  • Test before batch execution: After writing a batch script, you must first test with 1-2 items to verify the script runs correctly; only then run the full batch. Never skip testing and execute in batch directly
  • Reduce redundant pre-operations: When multiple steps depend on the same prerequisite state, complete them in batch under that state to avoid repeatedly establishing the same state
  • Error resumption: Save results item by item during batch processing; on failure, resume from the breakpoint rather than starting over
  • Prefer llms.txt / sitemap.xml over DOM discovery: A single fetch returns the full URL list; DOM discovery requires loading every page first. Always try the two API discovery routes first and only fall back to DOM when both return empty.
  • Polite delay between pages: Default to 500–1500 ms between page navigations to avoid burst patterns. Tighten only on sites you control or own.

Experience Notes

Path: {working-directory}/browser-act-skill-forge-memories/webcrawler-deep-crawl.memory.md (working directory is determined by the Agent running the Skill, typically the project root or current working directory)

Before execution: If the file exists, read it first — it records unexpected situations encountered during past executions (e.g., a strategy has become ineffective); adjust strategy order accordingly.

After execution: If an unexpected situation is encountered (strategy became ineffective, page redesigned, anti-scraping upgraded, better path discovered), append a line: {YYYY-MM-DD}: {what happened} → {conclusion}

Normal execution does not write to the file. Do not record what keywords were used or how many results were returned — those are task outputs, not experience.

© browser-act, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in solutions/search-research/webcrawler-deep-crawl of browser-act/skills.

  • SKILL.md
  • scripts/discover-links.py
  • scripts/discover-llms-txt.py
  • scripts/discover-sitemap.py
  • scripts/extract-page-content.py

Open the folder on GitHubat commit 11c057b

Compare with similar skills

Webcrawler Deep Crawl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Webcrawler Deep Crawl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Webcrawler Deep Crawl this skillbrowser-act/skills6.1k—~4.3kAutomated safety check: PassMIT
Firecrawl Knowledge Ingestfirecrawl/skills115—~565Automated safety check: PassISC
Agent Readiness Auditindranilbanerjee/digital-marketing-pro8541 repos~3.9kAutomated safety check: PassMIT
Setup Workspaceprobabl-ai/skills135—~1.5kAutomated safety check: PassBSD-3-Clause
Name Framework Migration Third Stepopensanctions/opensanctions831—~904Automated safety check: PassMIT
Cf Crawldavila7/claude-code-templates32k—~2.6kAutomated safety check: NotesMIT

Similar skills

  • Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser.

    115 GitHub stars~565 tokensUpdated today
    Sales & SupportAuto-check passed
  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended)…

    854 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    135 GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Name Framework Migration Third Step

    opensanctions/opensanctions

    Complete the name framework migration in a crawler (Step 3) by removing all custom name cleaning/splitting logic and the Step 1 review scaffolding, replacing it with a single…

    831 GitHub stars~904 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Cf Crawl

    davila7/claude-code-templates

    Crawl entire websites using Cloudflare Browser Rendering /crawl API.

    32k GitHub stars~2.6k tokensUpdated today
    Knowledge ManagementAuto-check: notes
  • Rendering Strategies

    kostja94/marketing-skills

    When the user wants to choose or optimize rendering strategy for SEO.

    1k GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from browser-act/skills

All 87 skills in this repo
  • Amazon ASIN Lookup

    browser-act/skills

    Fetches structured Amazon product details such as title, price, ratings and availability for a given ASIN through BrowserAct's lookup API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Best Sellers Finder

    browser-act/skills

    Extracts structured Amazon product data for a keyword and marketplace through the BrowserAct API, including titles, prices, ratings, reviews, sales volume and promotions.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Buy Box Monitor

    browser-act/skills

    Pulls Amazon product details, competing seller prices and seller ratings for a given ASIN through the BrowserAct API, without browser automation.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed
  • Analyzes a competitor's Amazon listing by ASIN with BrowserAct data extraction, then reports what it does well, where the market has gaps and opportunity points for your own listing.

    6.1k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Pulls structured Amazon search results (titles, ASINs, prices, ratings, specifications) for a keyword and brand through BrowserAct's Amazon Product API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Collects structured product data from Amazon search results for a keyword and optional brand, using a BrowserAct script, for market and competitor research.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about Webcrawler Deep Crawl

What does Webcrawler Deep Crawl do?

Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links. Webcrawler Deep Crawl is an agent skill from browser-act/skills. Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links.

When should I use Webcrawler Deep Crawl?

Webcrawler Deep Crawl fits situations like: user mentions deep crawl website; recursive crawl; crawl a whole site; scrape entire website.

How do I install Webcrawler Deep Crawl in Claude Code?

Run `npx skills add browser-act/skills --skill webcrawler-deep-crawl -a claude-code`. Or copy the skill folder (solutions/search-research/webcrawler-deep-crawl in browser-act/skills) into .claude/skills/webcrawler-deep-crawl in your project. Claude Code loads it when a task matches its description.

How do I install Webcrawler Deep Crawl in Codex?

Run `npx skills add browser-act/skills --skill webcrawler-deep-crawl -a codex`. Or copy the skill folder (solutions/search-research/webcrawler-deep-crawl in browser-act/skills) into .agents/skills/webcrawler-deep-crawl in your project. Codex loads it when a task matches its description.

Can I use Webcrawler Deep Crawl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-act/skills --skill webcrawler-deep-crawl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/webcrawler-deep-crawl, .gemini/skills/webcrawler-deep-crawl, .github/skills/webcrawler-deep-crawl and .opencode/skills/webcrawler-deep-crawl in your project.

What does Webcrawler Deep Crawl need to run?

Going by SKILL.md and its folder, Webcrawler Deep Crawl needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Webcrawler Deep Crawl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Webcrawler Deep Crawl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Webcrawler Deep Crawl use?

Webcrawler Deep Crawl is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Webcrawler Deep Crawl use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Webcrawler Deep Crawl?

Skills that share tags, products or a category with Webcrawler Deep Crawl: Firecrawl Knowledge Ingest (firecrawl/skills, 115 stars), Agent Readiness Audit (indranilbanerjee/digital-marketing-pro, 854 stars), Setup Workspace (probabl-ai/skills, 135 stars) and Name Framework Migration Third Step (opensanctions/opensanctions, 831 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Webcrawler Deep Crawl?

browser-act (a GitHub organization) maintains it in browser-act/skills, which has 6,108 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on August 24, 2026.

Source: browser-act/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.