Agent skill

Crawl4AI Web Scraping

by smallnest in smallnest/goclaw

Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

MITAuto-check passedData & Analytics

Install Crawl4AI Web Scraping

skills CLI
$ npx skills add smallnest/goclaw --skill crawl4ai -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install smallnest/goclaw crawl4ai --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .claude/skills/crawl4ai && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawl4ai
GitHub stars
598
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
439 words
Files
16 (incl. scripts, references)
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

  • Works in 2 steps: Schema-Based CSS Extraction (Most… → LLM-Based Extraction
  • Scraping a website into clean markdown
  • SKILL.md covers Overview, Quick Start, Core Concepts and Markdown Generation (Primary…, plus 4 more sections
  • Runs Python scripts from its folder; calls python and pip; reaches shop.com and site1.com

What it does

The skill documents two interfaces to Crawl4AI: the crwl command line, recommended for quick scriptable jobs, and the Python SDK for full programmatic control, with separate CLI and SDK guides. Both share the same configuration layers: browser settings such as headless mode, viewport, user agent and proxy; crawler settings such as page timeout, wait conditions, cache mode, JavaScript to run and CSS selector focus; extraction from a schema; and content filters.

Every crawl returns markdown, raw HTML, discovered links, media and any extracted content. Clean markdown generation is the primary use case, with relevance filters such as BM25. Bundled Python scripts cover a basic crawler, batch crawling, an extraction pipeline and a Google search helper, and a tests folder checks crawling, extraction and markdown generation. Installation is pip install crawl4ai followed by crawl4ai-setup.

When your agent uses it

  • Scraping a website into clean markdown
  • Extracting structured data from JavaScript-heavy pages
  • Crawling many URLs in a batch
  • Building an automated web data pipeline

Example prompts

  • “Crawl our docs site and save each page as markdown.”
  • “Extract product names and prices from this listing page with a schema, without using an LLM.”
  • “Crawl these twenty URLs in a batch and give me the links found on each.”

Requirements

  • Python with the crawl4ai package, installed with pip and crawl4ai-setup

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Schema-Based CSS Extraction (Most Efficient)
  2. LLM-Based Extraction

What it can do on your machine

Read from SKILL.md and the folder at commit e05c79d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • shop.com
    • site1.com
    • site2.com
    • site3.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawl4AI Web Scraping loads about 2.5k tokens when it runs, and up to ~63k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 439 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~63k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from smallnest/goclaw at commit e05c79d, republished under its MIT licence (© smallnest). 439 words, ~2,466 tokens.

Download SKILL.mdSave it as .claude/skills/crawl4ai/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
crawl4ai
description
This skill should be used when users need to scrape websites, extract structured data, handle JavaScript-heavy pages, crawl multiple URLs, or build automated web data pipelines. Includes optimized extraction patterns with schema generation for efficient, LLM-free extraction.
version
1.0.0
author
Claude
always
true

Crawl4AI

Overview

Crawl4AI provides comprehensive web crawling and data extraction capabilities. This skill supports both CLI (recommended for quick tasks) and Python SDK (for programmatic control).

Choose your interface:

  • CLI (crwl) - Quick, scriptable commands: CLI Guide
  • Python SDK - Full programmatic control: SDK Guide

Quick Start

Installation
bash
pip install crawl4ai
crawl4ai-setup

# Verify installation
crawl4ai-doctor
bash
# Basic crawling - returns markdown
crwl https://example.com

# Get markdown output
crwl https://example.com -o markdown

# JSON output with cache bypass
crwl https://example.com -o json -v --bypass-cache

# See more examples
crwl --example
Python SDK
python
import asyncio
from crawl4ai import AsyncWebCrawler

async def main():
    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun("https://example.com")
        print(result.markdown[:500])

asyncio.run(main())

For SDK configuration details: SDK Guide - Configuration (lines 61-150)


Core Concepts

Configuration Layers

Both CLI and SDK use the same underlying configuration:

ConceptCLISDK
Browser settings-B browser.yml or -b "param=value"BrowserConfig(...)
Crawl settings-C crawler.yml or -c "param=value"CrawlerRunConfig(...)
Extraction-e extract.yml -s schema.jsonextraction_strategy=...
Content filter-f filter.ymlmarkdown_generator=...
Key Parameters

Browser Configuration:

  • headless: Run with/without GUI
  • viewport_width/height: Browser dimensions
  • user_agent: Custom user agent
  • proxy_config: Proxy settings

Crawler Configuration:

  • page_timeout: Max page load time (ms)
  • wait_for: CSS selector or JS condition to wait for
  • cache_mode: bypass, enabled, disabled
  • js_code: JavaScript to execute
  • css_selector: Focus on specific element

For complete parameters: CLI Config | SDK Config

Output Content

Every crawl returns:

  • markdown - Clean, formatted markdown
  • html - Raw HTML
  • links - Internal and external links discovered
  • media - Images, videos, audio found
  • extracted_content - Structured data (if extraction configured)

Markdown Generation (Primary Use Case)

Crawl4AI excels at generating clean, well-formatted markdown:

CLI
bash
# Basic markdown
crwl https://docs.example.com -o markdown

# Filtered markdown (removes noise)
crwl https://docs.example.com -o markdown-fit

# With content filter
crwl https://docs.example.com -f filter_bm25.yml -o markdown-fit

Filter configuration:

yaml
# filter_bm25.yml (relevance-based)
type: "bm25"
query: "machine learning tutorials"
threshold: 1.0
Python SDK
python
from crawl4ai.content_filter_strategy import BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator

bm25_filter = BM25ContentFilter(user_query="machine learning", bm25_threshold=1.0)
md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)

config = CrawlerRunConfig(markdown_generator=md_generator)
result = await crawler.arun(url, config=config)

print(result.markdown.fit_markdown)  # Filtered
print(result.markdown.raw_markdown)  # Original

For content filters: Content Processing (lines 2481-3101)


Data Extraction

1. Schema-Based CSS Extraction (Most Efficient)

No LLM required - fast, deterministic, cost-free.

CLI:

bash
# Generate schema once (uses LLM)
python scripts/extraction_pipeline.py --generate-schema https://shop.com "extract products"

# Use schema for extraction (no LLM)
crwl https://shop.com -e extract_css.yml -s product_schema.json -o json

Schema format:

json
{
  "name": "products",
  "baseSelector": ".product-card",
  "fields": [
    {"name": "title", "selector": "h2", "type": "text"},
    {"name": "price", "selector": ".price", "type": "text"},
    {"name": "link", "selector": "a", "type": "attribute", "attribute": "href"}
  ]
}
2. LLM-Based Extraction

For complex or irregular content:

CLI:

yaml
# extract_llm.yml
type: "llm"
provider: "openai/gpt-4o-mini"
instruction: "Extract product names and prices"
api_token: "your-token"
bash
crwl https://shop.com -e extract_llm.yml -o json

For extraction details: Extraction Strategies (lines 4522-5429)


Advanced Patterns

Dynamic Content (JavaScript-Heavy Sites)

CLI:

bash
crwl https://example.com -c "wait_for=css:.ajax-content,scan_full_page=true,page_timeout=60000"

Crawler config:

yaml
# crawler.yml
wait_for: "css:.ajax-content"
scan_full_page: true
page_timeout: 60000
delay_before_return_html: 2.0
Multi-URL Processing

CLI (sequential):

bash
for url in url1 url2 url3; do crwl "$url" -o markdown; done

Python SDK (concurrent):

python
urls = ["https://site1.com", "https://site2.com", "https://site3.com"]
results = await crawler.arun_many(urls, config=config)

For batch processing: arun_many() Reference (lines 1057-1224)

Show full SKILL.md (173 more words)Show less
Session & Authentication

CLI:

yaml
# login_crawler.yml
session_id: "user_session"
js_code: |
  document.querySelector('#username').value = 'user';
  document.querySelector('#password').value = 'pass';
  document.querySelector('#submit').click();
wait_for: "css:.dashboard"
bash
# Login
crwl https://site.com/login -C login_crawler.yml

# Access protected content (session reused)
crwl https://site.com/protected -c "session_id=user_session"

For session management: Advanced Features (lines 5429-5940)

Anti-Detection & Proxies

CLI:

yaml
# browser.yml
headless: true
proxy_config:
  server: "http://proxy:8080"
  username: "user"
  password: "pass"
user_agent_mode: "random"
bash
crwl https://example.com -B browser.yml

Common Use Cases

Google Search Scraping
bash
# Search Google and get results as JSON
python scripts/google_search.py "your search query" 20

# Example
python scripts/google_search.py "2026年Go语言展望" 20

The script extracts:

  • Search result titles
  • URLs (cleaned, removes Google redirects)
  • Descriptions/snippets
  • Site names

Output is saved to google_search_results.json and printed to stdout.

Documentation to Markdown
bash
crwl https://docs.example.com -o markdown > docs.md
E-commerce Product Monitoring
bash
# Generate schema once
python scripts/extraction_pipeline.py --generate-schema https://shop.com "extract products"

# Monitor (no LLM costs)
crwl https://shop.com -e extract_css.yml -s schema.json -o json
News Aggregation
bash
# Multiple sources with filtering
for url in news1.com news2.com news3.com; do
  crwl "https://$url" -f filter_bm25.yml -o markdown-fit
done
Interactive Q&A
bash
# First view content
crwl https://example.com -o markdown

# Then ask questions
crwl https://example.com -q "What are the main conclusions?"
crwl https://example.com -q "Summarize the key points"

Resources

Provided Scripts
  • scripts/google_search.py - Google search scraper with JSON output
  • scripts/extraction_pipeline.py - Schema generation and extraction
  • scripts/basic_crawler.py - Simple markdown extraction
  • scripts/batch_crawler.py - Multi-URL processing
Reference Documentation
DocumentPurpose
CLI GuideCommand-line interface reference
SDK GuidePython SDK quick reference
Complete SDK ReferenceFull API documentation (5900+ lines)

Best Practices

  1. Start with CLI for quick tasks, SDK for automation
  2. Use schema-based extraction - 10-100x more efficient than LLM
  3. Enable caching during development - --bypass-cache only when needed
  4. Set appropriate timeouts - 30s normal, 60s+ for JS-heavy sites
  5. Use content filters for cleaner, focused markdown
  6. Respect rate limits - Add delays between requests

Troubleshooting

JavaScript Not Loading
bash
crwl https://example.com -c "wait_for=css:.dynamic-content,page_timeout=60000"
Bot Detection Issues
bash
crwl https://example.com -B browser.yml
yaml
# browser.yml
headless: false
viewport_width: 1920
viewport_height: 1080
user_agent: "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
Content Not Extracted
bash
# Debug: see full output
crwl https://example.com -o all -v

# Try different wait strategy
crwl https://example.com -c "wait_for=js:document.querySelector('.content')!==null"
Session Issues
bash
# Verify session
crwl https://site.com -c "session_id=test" -o all | grep -i session

For comprehensive API documentation, see Complete SDK Reference.

© smallnest, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts, references) in internal/builtin_skills/crawl4ai-skill of smallnest/goclaw.

  • SKILL.md
  • .claude/settings.local.json
  • README.md
  • references/cli-guide.md
  • references/complete-sdk-reference.md
  • references/sdk-guide.md
  • scripts/basic_crawler.py
  • scripts/batch_crawler.py
  • scripts/extraction_pipeline.py
  • scripts/google_search.py
  • tests/README.md
  • tests/run_all_tests.py
  • tests/test_advanced_patterns.py
  • tests/test_basic_crawling.py
  • tests/test_data_extraction.py
  • tests/test_markdown_generation.py

Open the folder on GitHubat commit e05c79d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in smallnest/goclaw, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Crawl4AI Web Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawl4AI Web Scraping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawl4AI Web Scraping this skillsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
Authoritative Data Harvesteryushui2022/MathModel-Skill4521 repos~1.1kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT
Google Maps ScraperMahanaicoach/google-maps-scraper-kit1.3k—~2.8kAutomated safety check: PassMIT

Similar skills

  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    452 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Google Maps Scraper

    Mahanaicoach/google-maps-scraper-kit

    Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.

    1.3k GitHub stars~2.8k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed

More from smallnest/goclaw

  • Reference for OpenClaw's discord tool: sending messages, reactions, stickers, emoji uploads, polls, threads and moderation in Discord channels and DMs.

    598 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed
  • Looks up action manuals with verified selectors for multi-step website tasks, then drives the browser with the actionbook CLI in a dedicated or an existing Chrome session.

    598 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check: warnings
  • Controls an Android device over ADB in a loop of screenshot, vision analysis, tap and verification, so the agent acts on real on-screen coordinates.

    598 GitHub stars~436 tokensUpdated 6 mo ago
    Auto-check passed
  • Feishu Upload Image

    smallnest/goclaw

    Upload images to Feishu/Lark. An agent skill from smallnest/goclaw.

    598 GitHub stars~556 tokensUpdated 6 mo ago
    Auto-check passed
  • goclaw Skill Finder

    smallnest/goclaw

    Helps you find, install, update and remove skills in the goclaw agent framework, using the goclaw skills commands and outside skill sources.

    598 GitHub stars~1.2k tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Crawl4AI Web Scraping

What does Crawl4AI Web Scraping do?

Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM. The skill documents two interfaces to Crawl4AI: the crwl command line, recommended for quick scriptable jobs, and the Python SDK for full programmatic control, with separate CLI and SDK guides. Both share the same configuration layers: browser settings such as headless mode, viewport, user agent and proxy; crawler settings such as page timeout, wait conditions, cache mode, JavaScript to run and CSS selector focus; extraction from a schema; and content filters.

When should I use Crawl4AI Web Scraping?

Crawl4AI Web Scraping fits situations like: scraping a website into clean markdown; extracting structured data from JavaScript-heavy pages; crawling many URLs in a batch; building an automated web data pipeline.

How do I install Crawl4AI Web Scraping in Claude Code?

Run `npx skills add smallnest/goclaw --skill crawl4ai -a claude-code`. Or copy the skill folder (internal/builtin_skills/crawl4ai-skill in smallnest/goclaw) into .claude/skills/crawl4ai in your project. Claude Code loads it when a task matches its description.

How do I install Crawl4AI Web Scraping in Codex?

Run `npx skills add smallnest/goclaw --skill crawl4ai -a codex`. Or copy the skill folder (internal/builtin_skills/crawl4ai-skill in smallnest/goclaw) into .agents/skills/crawl4ai in your project. Codex loads it when a task matches its description.

Can I use Crawl4AI Web Scraping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add smallnest/goclaw --skill crawl4ai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawl4ai, .gemini/skills/crawl4ai, .github/skills/crawl4ai and .opencode/skills/crawl4ai in your project.

What does Crawl4AI Web Scraping need to run?

Going by SKILL.md and its folder, Crawl4AI Web Scraping needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python with the crawl4ai package, installed with pip and crawl4ai-setup.

Does Crawl4AI Web Scraping access the network?

SKILL.md names 4 domains. In commands or code: shop.com, site1.com, site2.com and site3.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Crawl4AI Web Scraping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Crawl4AI Web Scraping use?

Crawl4AI Web Scraping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawl4AI Web Scraping use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 61k tokens, read only when the agent opens those files.

What are the alternatives to Crawl4AI Web Scraping?

Skills that share tags, products or a category with Crawl4AI Web Scraping: Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars) and Ax (yusukebe/ax, 719 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawl4AI Web Scraping?

smallnest (a GitHub user) maintains it in smallnest/goclaw, which has 598 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on March 27, 2026.

Source: smallnest/goclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.