Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .claude/skills/web-scraping && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .claude/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scrapingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .agents/skills/web-scraping && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .agents/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .cursor/skills/web-scraping && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .cursor/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ginlix-ai/LangAlpha.git --path plugins/alternative_data/skills/web-scraping--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .gemini/skills/web-scraping && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .gemini/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ginlix-ai/LangAlpha web-scrapingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .github/skills/web-scraping && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .github/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .opencode/skills/web-scraping && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "web-scraping" agent skill from https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/alternative_data/skills/web-scraping into .opencode/skills/web-scraping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraping", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
web-scrapingWeb scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…
Web Scraping is an agent skill from ginlix-ai/LangAlpha. Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors, sessions, and spiders
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api-reference.md`).
It sits in Data & Analytics, covering Web scraping. It works with Python. The repository describes itself as: Claude Code for Financial Market. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 2855e43. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
protected-site.comspa-site.comspa-website.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Web Scraping loads about 2.1k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 491 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ginlix-ai/LangAlpha at commit 2855e43, republished under its MIT licence (© ginlix-ai). 491 words, ~2,120 tokens.
.claude/skills/web-scraping/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Two ways to scrape in the sandbox:
scrape_page, scrape_pages) — recommended for straight "give me this page's content". Synchronous, return dicts..css() / .xpath().Quick fetches can run inline via ExecuteCode. For spiders, multi-URL crawls, or anything you'll iterate on, write the scraper to <task_name>/scraper.py and run it via Bash — edit-and-rerun beats resubmitting code.
Import from tools.scrape. Synchronous — no await.
from tools.scrape import scrape_page, scrape_pagesscrape_page(url: str, mode: str = "fast", extraction: str = "markdown",
timeout: float = 30.0, solve_cloudflare: bool = False) -> dict
scrape_pages(urls: list[str], mode: str = "fast", extraction: str = "markdown",
timeout: float = 30.0, solve_cloudflare: bool = False) -> dict| Param | Default | Notes |
|---|---|---|
mode | "fast" | "fast" plain HTTP · "browser" JS rendering · "stealth" bot-protected sites |
extraction | "markdown" | "markdown" (article text, cleaned) · "html" (raw) · "text" (plain) |
timeout | 30.0 | Per-fetch seconds, 1–60 — seconds in every mode, not ms |
solve_cloudflare | False | Only meaningful with mode="stealth" |
urls | — | scrape_pages only; max 10 per call |
Escalate modes only as needed: start fast, go to browser when the page needs JavaScript, stealth when you're getting blocked, and add solve_cloudflare=True only if stealth still returns a challenge page.
scrape_page returns a flat dict:
{
"url": "https://example.com",
"status": 200,
"title": "Example Domain",
"content": "# Example Domain\n\nThis domain is for use in...", # str
"extraction": "markdown",
"mode": "fast",
}content is a plain string, not a list — use it directly, never content[0] (that yields a single character).content is truncated to 400,000 chars..css() / .xpath() / .body / .headers / .cookies — for selectors use the direct Python API below, or parse extraction="html" with BeautifulSoup.scrape_pages wraps them:
{
"results": [ ... ], # one entry per input URL, in input order
"count": 3,
}Errors are returned, never raised. Always check for "error" before reading content.
res = scrape_page(url="https://example.com")
if "error" in res:
print(res["error"], res["detail"])
else:
print(res["content"])Per-URL errors — appear as {"error", "detail", "url"} entries inside scrape_pages["results"], or as the whole return of scrape_page:
| Code | Meaning |
|---|---|
invalid_url | Not an http:// / https:// URL |
fetch_failed | Network, DNS, timeout, or browser failure |
extract_failed | Page fetched but the extractor failed on the markup; the entry still carries status |
scrape_failed | Unexpected internal failure for that one URL |
Whole-call errors — the entire return is {"error", "detail"}, no results:
| Code | Meaning |
|---|---|
invalid_mode / invalid_extraction / invalid_timeout | Bad argument value |
invalid_urls | scrape_pages got an empty list or more than 10 URLs |
One bad URL never sinks a batch. scrape_pages always returns one entry per input URL, in input order — failures come back as error entries alongside the successes.
from tools.scrape import scrape_page, scrape_pages
# Single page → markdown
res = scrape_page(url="https://example.com")
if "error" not in res:
print(res["title"], res["status"], len(res["content"]))
# JS-rendered page
res = scrape_page(url="https://spa-site.com", mode="browser", timeout=60)
# Bot-protected page
res = scrape_page(url="https://protected-site.com", mode="stealth", solve_cloudflare=True)
# Batch — split successes from failures
batch = scrape_pages(urls=[...], mode="fast") # <= 10 URLs
pages = [r for r in batch["results"] if "error" not in r]
failed = [(r["url"], r["error"]) for r in batch["results"] if "error" in r]
# Raw HTML when you need to parse structure yourself
res = scrape_page(url="https://example.com", extraction="html")
from bs4 import BeautifulSoup
soup = BeautifulSoup(res["content"], "html.parser")
titles = [h1.get_text() for h1 in soup.find_all("h1")]Batches run concurrently — 8 at a time in fast mode, 2 at a time in browser / stealth (browser sessions are memory-heavy). More than 10 URLs means more than one call.
For selectors, sessions, spiders, or when you need the full Page object. Requires imports. Async.
from scrapling.fetchers import AsyncFetcher
page = await AsyncFetcher.get("https://example.com", stealthy_headers=True)
print(page.status) # 200
print(page.body) # Raw bytes
print(page.headers) # Response headers
# CSS selectors (Scrapy-style pseudo-elements)
titles = page.css("h1::text").getall()
links = page.css("a::attr(href)").getall()
# XPath
items = page.xpath("//div[@class='item']/text()").getall()
# BeautifulSoup-style
divs = page.find_all("div", class_="content")from scrapling.fetchers import DynamicFetcher
page = await DynamicFetcher.async_fetch(
"https://spa-website.com",
headless=True,
network_idle=True,
disable_resources=True,
timeout=30000, # milliseconds here, unlike the MCP tools
wait_selector=".data-table",
)
rows = page.css("table.data-table tr")
for row in rows:
cells = row.css("td::text").getall()from scrapling.fetchers import StealthyFetcher
page = await StealthyFetcher.async_fetch(
"https://protected-site.com",
headless=True,
solve_cloudflare=True,
network_idle=True,
)from scrapling.fetchers import FetcherSession
with FetcherSession(impersonate="chrome") as session:
login_page = session.post("https://site.com/login", data={...})
dashboard = session.get("https://site.com/dashboard")
data = dashboard.css(".user-data::text").getall()from scrapling.spiders import Spider, Request, Response
class PriceScraper(Spider):
name = "prices"
start_urls = ["https://example.com/products"]
concurrent_requests = 5
async def parse(self, response: Response):
for product in response.css(".product"):
yield {
"name": product.css(".name::text").get(),
"price": product.css(".price::text").get(),
}
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield Request(next_page)
spider = PriceScraper()
result = spider.start()
result.items.to_json("<task_name>/data/prices.json")Only needed when you fetched HTML yourself — scrape_page(extraction="markdown") already does this.
import html_to_markdown
markdown = html_to_markdown.convert(
html_string, html_to_markdown.ConversionOptions(extract_metadata=False)
).content
# Article-only extraction (strips nav/ads/boilerplate)
import trafilatura
article = trafilatura.extract(html_string, output_format="markdown", favor_recall=True)| Need | Use |
|---|---|
| Quick page content as markdown | scrape_page() |
| Several known URLs at once | scrape_pages() (≤10 per call) |
| Extract specific elements (CSS/XPath) | Direct Python API with selectors |
| Login + scrape authenticated pages | Direct Python API with sessions |
| Crawl many pages with pagination | Direct Python API with Spider |
| Bypass Cloudflare | scrape_page(mode="stealth", solve_cloudflare=True) or direct StealthyFetcher |
| Save results to file | Direct Python API (spider .to_json()) |
© ginlix-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/alternative_data/skills/web-scraping of ginlix-ai/LangAlpha.
Open the folder on GitHubat commit 2855e43
Web Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Scraping this skillginlix-ai/LangAlpha | 1.8k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper | 1.5k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Axyusukebe/ax | 719 | 1 repos | ~918 | Automated safety check: Pass | MIT | |
| Google Maps ScraperMahanaicoach/google-maps-scraper-kit | 1.3k | — | ~2.8k | Automated safety check: Pass | MIT | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
eatmoreduck/boss-zhipin-scraper
Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.
yusukebe/ax
Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.
Mahanaicoach/google-maps-scraper-kit
Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
Cedriccmh/claude-code-skill-scrapling
使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…
ginlix-ai/LangAlpha
Quality-checks an investment deck in .pptx form before it goes out: number consistency, chart and narrative alignment, source coverage, language and a circulation verdict.
ginlix-ai/LangAlpha
Produces a first-time equity research initiation report in five tasks: company research, financial model, valuation, charts and a DOCX report.
ginlix-ai/LangAlpha
Builds or repairs an integrated income statement, balance sheet and cash flow model in Excel with live formulas, supporting schedules, scenarios and a Checks sheet.
ginlix-ai/LangAlpha
Audits an existing Excel financial model without editing it, checking structure, formulas, integrity identities and source tie-out, and ends in a prioritized issue log.
ginlix-ai/LangAlpha
Builds a live Excel DCF valuation workbook with free cash flow projections, WACC, terminal value, three scenarios, sensitivity grids and a reverse DCF.
ginlix-ai/LangAlpha
Builds Word files with python-docx, edits existing ones in place with tracked changes and comments, then renders and validates the result.
Works with
Categories
Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…. Web Scraping is an agent skill from ginlix-ai/LangAlpha.
Web Scraping fits situations like: tasks that involve Web scraping.
Run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a claude-code`. Or copy the skill folder (plugins/alternative_data/skills/web-scraping in ginlix-ai/LangAlpha) into .claude/skills/web-scraping in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a codex`. Or copy the skill folder (plugins/alternative_data/skills/web-scraping in ginlix-ai/LangAlpha) into .agents/skills/web-scraping in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraping, .gemini/skills/web-scraping, .github/skills/web-scraping and .opencode/skills/web-scraping in your project.
SKILL.md names no scripts, command-line tools or credentials: Web Scraping is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 3 domains. In commands or code: protected-site.com, spa-site.com and spa-website.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Web Scraping is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Web Scraping: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars), Ax (yusukebe/ax, 719 stars) and Google Maps Scraper (Mahanaicoach/google-maps-scraper-kit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ginlix-ai (a GitHub organization) maintains it in ginlix-ai/LangAlpha, which has 1,811 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 10, 2026.
Source: ginlix-ai/LangAlpha on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.